Welcome!

A few days ago, I presented at a conference about the Open-Source Leg and how important it is to connect engineering and medicine if we want innovation to actually reach patients.

This week, I came across a much bigger example of that same idea.

NIH, BioHub, Google DeepMind, Meta, and others announced a $1.8 billion effort to build biological datasets that could help AI better predict how cells respond to disease and treatments (Link).

The hope is that these models could help researchers test biological questions virtually, identify new treatment possibilities, and potentially move discoveries faster.

Very different area and a much bigger scale, but the same principle: progress is stronger when these worlds work together.

Now it’s time to check this week’s newsletter.

🤖AIBytes

Two clinical studies that deserve a closer look.

Can an Open Medical AI Model Compete With Much Larger Models?

MedGemma is Google’s open medical AI model, built to work with both text and medical images.

This new Nature Medicine study takes a closer look at how well it performs across clinical questions, reasoning, and several types of medical imaging.

🔬 Methods

  • Researchers evaluated three versions of MedGemma: two that work with text and images, and one focused on text.

  • They compared MedGemma with the original Gemma 3 models to see whether medical training improved performance.

  • Some tests used completely new datasets that the models had not seen during training.

  • They also fine-tuned MedGemma for specific tasks to see how well it could adapt, including when only a small amount of training data was available.

📊 Results

  • On new datasets, MedGemma improved performance by 2.6–10% for medical image questions, 15.5–18.1% for chest X-ray classification, and 10.8% on simulated agent tasks compared with the original Gemma models.

  • The 27B text model scored 86.2% on MedQA benchmark, one of the highest results reported for an open model.

  • In physician-reviewed medical responses, inaccurate answers decreased from 18.5% with Gemma 3 to 15.4% with MedGemma 4B.

  • For chest X-rays, 81% of AI reports were judged likely to lead to the same or better clinical decisions.

🔑 Key Takeaways

  • Training the model specifically for medicine helped.

  • The same model could handle very different tasks, from clinical questions to several types of medical imaging.

  • Because it is open, developers can adapt it for specific uses and potentially run it locally, which may matter for privacy, cost, and control over data.

  • The results give a strong foundation for the next step: testing how MedGemma performs in real clinical settings.

🔗 Sellergren A, Kazemzadeh S, Mahvar F, et al. An open vision-language model for diverse medical applications. Nature Medicine. 2026. doi:10.1038/s41591-026-04626-

AI vs Residents for Pre-Visit History Taking

Taking a complete history takes time, especially in busy outpatient clinics.

Researchers tested whether an AI agent could handle the pre-consultation history before patients saw the ophthalmologist.

🔬 Methods

  • 172 patients with non-emergency eye conditions were randomized to either an AI agent or an ophthalmology resident for pre-consultation.

  • The AI asked patients questions using voice or text and automatically created a clinical note.

  • They evaluated history-taking, communication, documentation, patient satisfaction, and test recommendations.

  • A senior ophthalmologist then reviewed the chart, examined the patient, and completed the assessment.

📊 Results

  • The AI scored higher for overall pre-consultation quality, with an adjusted difference of 17.7 points on a 100-point scale.

  • Patients rated the AI higher for patience and empathy, while residents were better at reducing anxiety. Overall satisfaction was similar.

  • The AI took much longer: 11.2 minutes compared with 3.1 minutes for residents.

  • Using history alone, the AI had higher diagnostic accuracy. But once physical examination findings were added the difference became much smaller.

🔑 Key Takeaway

  • AI collected a more complete history, while residents were much faster and more focused.

  • This could make AI useful before the visit, helping gather information so physicians can spend more time on the exam and clinical decisions.

  • But the physical exam still made a difference. Residents gained more diagnostic value from the ocular findings than the AI did.

🔗 Luo M-J, Lai Y, Wang W, et al. Standardized pre-consultation by a large language model agent vs ophthalmology residents: a randomized clinical trial. npj Digital Medicine. 2026. doi:10.1038/s41746-026-03232-x

🧬AIMedily Snaps

Fast updates clinicians should not miss.

  • The U.S. Department of Health and Human Services awarded funding to a project focused on bringing agentic AI into clinical care. (Link)

  • Google introduced new Earth AI work that combines environmental and health data to help public-health teams anticipate outbreaks and identify vulnerable communities. (Link)

  • CHAI released a new playbook to help health systems evaluate AI vendors more consistently before adoption. (Link)

  • Stanford received up to $14.9 million to develop a system that monitors other AI agents being used in cardiovascular care. (Link)

  • Teladoc added AI tools for ambient documentation and contactless vital signs to its hospital platform. (Link)

  • Six major U.S. physician organizations said physician judgment should remain central as AI becomes more involved in patient care. (Link)

🧪Research Signals

New papers worth your time.

  • Nature: A randomized study found that a respiratory AI chatbot helped people answer simulated diagnosis and triage questions more accurately than regular web search. (Paper)

  • npj: An AI system for breast ultrasound improved radiologist accuracy and reduced missed cancers in a multinational study. (Paper)

  • npj: A multicenter randomized trial, AI-assisted telerehabilitation was tested in people with early Parkinson’s disease to see whether remotely delivered, AI-supported rehabilitation could improve care. (Paper)

  • npj: Clinicians preferred AI-generated discharge summaries in longer hospital stays and rated them higher. (Paper)

  • npj: Researchers found publicly available performance evidence for only 39% of authorized AI diagnostic tools in pathology and hematology. (Paper)

  • Stanford: An Evidence API that can search more than 100 million scientific papers and quickly surface evidence supporting or challenging a scientific claim. (Link)

🦾TechTools

AI medical tools

Us2.ai (Link)
An FDA-cleared tool that turns any echo study into a complete, guideline-based report with automated measurements and disease detection in seconds.

Regard (Link)
Reviews the full medical record, surfaces possible diagnoses with the supporting evidence, and can generate a draft clinical note before the encounter.

📈 Productivity AI tool of the week:
Julius AI (Link)
Lets you analyze spreadsheets and other datasets with plain-language questions, create visualizations, and run statistical analyses.

That’s it for today.

Thanks for spending a few minutes with me.

P.S. If you’ve been enjoying AIMedily, I’d love to hear what’s been most useful to you. Your feedback helps me make it better for this community. Share a quick note — it takes less than a minute.

Itzel Fer, MD PM&R

Follow me on LinkedIn | Substack | X | Instagram

Forwarded this email? Subscribe free.