Hi!
This week, I’m writing the newsletter from Baltimore, where I’m attending the Machine Learning for Healthcare conference at Johns Hopkins.
Today was the first day, and it was great to hear presentations from the AMA, NVIDIA, Anthropic, and others. I’m looking forward to meeting more interesting people over the next few days and, of course, continuing to learn.
One of the researchers from the first paper featured in this week’s AIBytes will also be presenting here, which makes this issue feel especially timely.
Now I need to hit send and get some sleep.
Let’s dive into today’s issue.
🤖AIBytes
Two clinical studies that deserve a closer look.
Can AI Match Doctors in Video Consultations?
In this paper, researchers tested AMIE (Video). It is a Google AI system designed to conduct real-time medical consultations using video, audio, and clinical reasoning.
Methods
Researchers compared AMIE with 10 board-certified primary care physicians across 100 simulated clinical cases.
The study included 15 patient actors. Another 20 physicians independently evaluated the consultations. AMIE was also compared with its text-only version.
Results
AMIE scored 83% overall vs. 68% for physicians on case-specific clinical evaluations.
Its top diagnosis was correct in 91% of cases vs. 77% for physicians.
For video-based perception and examination, AMIE scored 74% vs. 47% for physicians.
Patient actors preferred AMIE for assessing and explaining conditions. Physicians were preferred for rapport and partnership, although these differences were not significant.
Patients also preferred AMIE Video over text for communicating concerns, convenience, and feeling understood.

Key Takeaways
This study moves medical AI beyond text. AMIE could see, listen, reason, and guide parts of a physical examination during a video visit.
Performance was strong across several clinical tasks. But this was an OSCE with simulated patients, not real-world clinical care.
AMIE still struggled with subtle cues, fine visual details, and fast movements such as tremor.
The authors state that the system is not ready for clinical deployment. The paper is also a Google Research/DeepMind preprint and has not yet been peer reviewed.
🔗 Nagda M, Lee J, Thompson M, et al. Towards Expert-level Medical AI for Real-time Video Consultations. arXiv. 2026. https://arxiv.org/abs/2608.09861
When AI Support Reinforces Mental Health Risk
In this paper, researchers developed SIM-VAIL. It tests how AI chatbots respond to users with different mental-health vulnerabilities over several conversation turns.
Methods
Researchers simulated 810 conversations across 30 user profiles and nine AI chatbots.
Profiles included depression, psychosis, obsessive-compulsive disorder, mania, and insecure attachment.
The team focused on 13 mental-health risk measures. A group of 27 clinicians also reviewed a sample of the conversations.
Results
Concerning behavior increased as conversations continued.
Risk increased fastest in simulated users with psychosis and mania.
Risk was higher when users sought emotional dependence, glorification of distress, or help with risky actions.
Changing one concerning response early reduced risk in later turns. The effect remained detectable for five more turns.

Key Takeaways
AI safety should be tested across whole conversations, not single responses.
Responses that seem supportive can still reinforce unhealthy beliefs or behaviors.
Detecting risk early may help prevent harmful conversations from escalating.
The study used simulated users and model APIs, not real patients using consumer chatbot interfaces.
🔗 Weilnhammer V, Hou KYC, Luettgau L, et al. A clinically validated framework for auditing AI chatbot behavior in mental health interactions. Nature Medicine. 2026. https://doi.org/10.1038/s41591-026-04577-2
🧬AIMedily Snaps
Fast updates clinicians should not miss.
Stanford: Researchers used generative AI to design novel bacteriophages that can infect and kill E. coli. (Link).
CHAI: The Coalition for Health AI launched a cybersecurity work group focused on healthcare AI security (Link).
OpenAI: ChatGPT for Academic Researchers will give 100,000 researchers free access to frontier AI models and tools. (Link).
FDA: The FDA will convene an advisory panel to review GRAIL’s AI-powered Galleri multi-cancer blood test (Link).
National Academies: Experts are examining how AI could enable new biological threats—and help accelerate medical defenses. (Link).
OpenEvidence: OpenEvidence launched MedMini and Synapses, new medical knowledge games for clinicians (Link).
🧪Research Signals
New papers worth your time.
Nature Medicine: PRISM2 uses 2.3 million pathology slides and clinical dialogue for cancer diagnosis and prediction. (Paper)
npj: In ICU outcome prediction, specialists often outperformed AI, while clinician with AI performed best. (Paper)
npj: Deep learning uses routine 12-lead ECGs to detect hypertrophic cardiomyopathy. (Paper)
JAMA: Updated guidance for disclosure, transparency, and accountability when using AI in medical publishing. (Paper)
npj: Systematic review of OpenEvidence found generally evidence-supported answers, but limited evaluation data. (Paper)
medRxiv: Framework for choosing fairness metrics in clinical AI based on intended use and clinical impact. (Paper)
🦾TechTools
Useful tools
Uses AI on routine coronary CT angiography to measure coronary inflammation and plaque, helping clinicians identify cardiovascular risk that stenosis alone may not reveal.
Remote patient monitoring platform that combines connected health data with AI-guided care coordination to help clinical teams manage chronic conditions between visits.
📈 Productivity AI tool of the week: Fyxer
AI email assistant that organizes your inbox, drafts replies in your writing style, takes meeting notes, and helps with scheduling directly inside Gmail or Outlook.
That’s it for today.
Thank you for taking the time to read!
You’re already ahead of the curve in medical AI — don’t keep it to yourself. Forward AIMedily to one friend.
Itzel Fer, MD PM&R
Forwarded this email? Subscribe free.
