Hello!

This week, a JAMA Perspective asking whether autonomous AI could provide better medical care than physicians using AI created a lot of controversy.

And I understand why.

The question is bigger than whether AI can get the right diagnosis or outperform physicians on a benchmark. It is really about what “better medical care” means.

Medicine is inherently human and messy. We work with uncertainty, incomplete information, patient preferences, competing priorities, and decisions that are not always black and white. (See the paper in 🧪 Research Signals).

Thank you to all our new subscribers this week. I’m so glad you’re here.

Let’s dive into today’s issue.

🤖AIBytes

Two clinical studies that deserve a closer look.

Can AI Check Medical AI for Errors?

In this paper, researchers developed MedVAL. It trains language models to check whether AI-generated medical text is safe and factually consistent.

Methods

Researchers tested MedVAL across 10 language models and six medical tasks.

They created MedVAL-Bench with 840 physician-reviewed examples.

Tasks included medication questions, radiology summaries, hospital-course translation, and doctor-patient dialogue summaries.

Results

  • MedVAL improved average safe/unsafe classification F1 scores from 66.2% to 82.8%.

  • It also improved four-level medical risk classification, from 36.7% to 51.0%.

  • Performance improved on both familiar and previously unseen tasks.

  • On a smaller subset reviewed by multiple physicians, MedVAL performed statistically non-inferior to a single human expert.

Key Takeaways

  • AI may eventually help check other AI-generated medical text before it reaches clinical workflows.

  • MedVAL does not require physician-labeled training data or a reference answer for every case.

  • The approach could help identify outputs that need human review.

  • This is still a proof-of-concept. Real-world testing in clinical workflows is needed.

🔗 Aali A, Bikia V, Varma M, et al. Toward expert-level medical text validation with language models. npj Digital Medicine. 2026. https://doi.org/10.1038/s41746-026-03084-5

AI Helps Detect Subclinical Basal Cell Carcinoma

In this paper, researchers tested whether AI combined with line-field confocal optical coherence tomography (LC-OCT) could detect basal cell carcinomas (BCCs) before they were visible on a routine skin exam.

Methods

The study included 150 patients at increased risk for BCC in Germany. Researchers excluded facial lesions that already looked suspicious for cancer. They then scanned areas of normal-looking facial skin using AI-assisted LC-OCT.

Lesions identified as possible BCCs were confirmed with biopsy when available.

Results

AI-assisted LC-OCT identified 18 lesions as possible BCCs. Of these, 15 were confirmed by biopsy, giving a positive predictive value of 83.3% (95% CI, 58.1%–96.4%).

Overall, researchers found 17 subclinical BCCs in 14 patients, or 9.3% of the group.

Most were superficial BCCs (76.5%), followed by nodular BCCs (17.6%) and infiltrative BCCs (5.9%).

The system also correctly identified the BCC subtype in 13 of 15 biopsy-confirmed cases (86.7%).

Key Takeaways

  • AI-assisted imaging detected BCCs in skin that looked normal on clinical examination.

  • About 1 in 11 high-risk patients had a subclinical BCC detected.

  • The study measured how often positive AI findings were correct, but not how many cancers the system may have missed.

  • Larger studies are needed to determine whether earlier detection improves outcomes and is cost-effective.

🔗 Ronicke M, Dürr L, Höner MW, et al. AI-Assisted Line-Field Confocal Optical Coherence Tomography to Detect Subclinical Basal Cell Carcinoma. JAMA Dermatol. Published online August 19, 2026. doi:10.1001/jamadermatol.2026.2992

🧬AIMedily Snaps

Fast updates clinicians should not miss.

  • New Health Guardian from Google features for Pixel Watch and Fitbit will track trends in insulin resistance, blood pressure, and sleep breathing quality (Link).

  • The AMA and Digital Medicine Society defining the physician’s role in the digital and AI era of medicine (Link).

  • Researchers from Microsoft and Stanford discuss on a Podcast when AI can improve clinician performance and when behaviors can get in the way (Link).

  • The FDA released a paper asking for feedback on how generative AI medical devices should be evaluated before and after they reach the market (Link).

  • Abridge is expanding its AI tools across partner health systems, with features like pre-visit summaries and answers based on the patient’s chart. (Link).

  • OpenAI launched AI for Civil Society and Philanthropy with a $100 million commitment focused first on helping improve health outcomes and delivery (Link).

🧪Research Signals

New papers worth your time.

  • JAMA: Will autonomous AI exceed AI-aided physicians as the best medical care? (Paper).

  • NEJM AI: A classification of safety risks in medical AI (Paper).

  • npj: Diagnostic accuracy of large language models for rare diseases: a systematic review and meta-analysis (Paper).

  • BMC: AI safety evaluation of clinical decision support and frontier language models for Medicaid patient messaging triage (Paper).

  • Health Policy: A systematic review of the economic evidence for AI in healthcare (Paper).

  • NEJM AI: Parental perspectives on using LLM-powered chatbots in rare disease care (Paper).

🦾TechTools

Medical AI tools

DeepHealth - AI-powered breast imaging tools for cancer detection, breast density assessment, risk assessment, and workflow support.

Oracle Health - Healthcare platform that connects the electronic health record, clinical AI, and tools designed to support clinical and administrative workflows.

📈 Productivity AI tool of the week:

Avoma - AI meeting assistant that records, transcribes, summarizes meetings, and helps automate follow-ups.