🤖AIBytes
Two clinical studies that deserve a closer look.
Proactive AI Support for Mental Health
This NEJM AI randomized trial tested whether a generative AI mobile app could support mental well-being before distress becomes more severe.
The app, called Flourish, was designed to give short, personalized well-being support through an AI coach, evidence-based activities, and weekly insights.
Methods
Researchers enrolled 486 undergraduate students from three U.S. institutions.
Students were randomized to:
access to the Flourish app
or care as usual
The study lasted 6 weeks.
Students assigned to Flourish were asked to use the app at least twice per week.
The app included:
emotion check-ins
conversations with an AI well-being coach
brief activities such as journaling, gratitude, cognitive reappraisal, social connection, and mindfulness
daily and weekly mood insights
Results
Students using Flourish reported better outcomes in several well-being measures.
They had greater:
resilience
social well-being
belonging
closeness to their community
They also had lower loneliness and were buffered against declines in:
positive affect
mindfulness
flourishing
But there were no significant differences in depression, anxiety, or stress in this general student population.
The authors note that this was a nonclinical undergraduate sample, and the app was designed as a preventive, strengths-based tool, not as a treatment for mental illness.

Key Takeaways
🔹 A purpose-built generative AI app may help provide low-intensity, preventive mental well-being support.
🔹 The strongest signal was in well-being, resilience, and social connection, not depression or anxiety reduction.
🔹 This should not be framed as AI replacing therapy or clinical care.
🔹 More studies are needed in clinical populations, more diverse groups, and with longer follow-up.
🔗 Cachia JYA, Zhao X, Hunter J, et al. AI for Proactive Mental Health: A Multi-Institutional, Longitudinal Randomized Controlled Trial. NEJM AI. Published July 22, 2026. doi:10.1056/AIoa2501293.
Can LLMs Help Detect Missed Diagnoses in the Emergency Department?
This study tested whether LLMs could help screen emergency department records for missed opportunities for diagnosis.
Methods
Researchers reviewed 288 encounters from 9 emergency departments (ED) in one US health system:
191 patients discharged from the ED who returned and were admitted within 72 hours
97 patients admitted to a hospital floor who required intensive care within 24 hours
Two emergency physicians independently reviewed each case using the Safer Dx framework.
Six models were tested: Claude Sonnet 4, Claude Sonnet 4.6, Claude Opus 4.6, Gemini 3 Pro, GPT-5, and GPT-5 mini.
Results
Physicians identified 39 missed diagnostic opportunities, or 13.5% of cases.
72-hour return cohort
21 of 191 cases had a missed diagnostic opportunity
Highest sensitivity: Claude Sonnet 4, 85.7%
Highest specificity: GPT-5 mini, 82.9%
Floor-to-ICU cohort
18 of 97 cases had a missed diagnostic opportunity
Highest sensitivity: Claude Sonnet 4, 55.6%
Highest specificity: GPT-5 mini, 97.5%
Adjusting each model’s threshold could have reduced review time from 24 hours to 12.2 hours while still finding at least 80% of missed cases.

Key Takeaways
LLMs may help prioritize records for physician review, but they should not independently decide whether a diagnosis was missed.
Models showed different sensitivity and specificity tradeoffs, even when their overall discrimination was similar.
Some models caught more true cases but created more false alarms; others reduced false alarms but missed more cases.
The best model depended on the clinical group and whether the priority was sensitivity or fewer unnecessary reviews.
🔗 Marks CM, Gibney S, Stenson B, et al. Screening for missed opportunities for diagnosis in the ED using eTriggers and large language models. JAMA Netw Open. 2026;9(6):e2620939. June 29, 2026. doi:10.1001/jamanetworkopen.2026.20939.
🧬AIMedily Snaps
Fast updates clinicians should not miss.
Coalition for Health AI (CHAI) Launches PULSE, a National Initiative to Help Public Health Agencies Responsibly Implement AI at Scale. (Link)
Boston Children’s and OpenEvidence will study AI-powered analysis of clinical practice patterns. (Link)
Human Health Services joins the Genesis Mission to use AI for chronic disease research and biomedical discovery. (Link)
Mayo Clinic and Scale report work on AI for chart review, safety detection, and admin burden. (Link)
🧪Research Signals
New papers worth your time.
Nature: Generating guideline-concordant and safe AI recommendations for diabetic kidney disease (Paper).
CLEAR: an auditable foundation model for radiology based in clinical concepts (Paper).
Nature: Why high scores do not mean application readiness for health AI (Paper).
Nature: Multimodal AI for predicting comorbid REM sleep behavior disorder in depresion (Paper).
BMJ: Predicting cardiac index from coronary angiogram videos (Paper).
NEJM AI: Will AI widen global health inequalities? (Paper)
🦾TechTools
Useful tools
Medscape AI
Clinician-facing AI tool that combines Medscape content, peer-reviewed literature, and medical news to support evidence search.
Bunkerhill Health
Helps health systems build and deploy AI agents across clinical, operational, and administrative workflows.
📈Productivity AI tool of the week:
Kimi K3 by Moonshot AI
AI model for writing, coding, research, and long-context work. It has drawn controversy because it comes from China-based Moonshot AI and competes with leading U.S. AI models. Not for patient data.
That’s all for today.
Thank you for taking the time to read.
If AIMedily helped you keep up this week, please forward it to one clinician who may find it useful too.
Itzel Fer, MD PM&R
Forwarded this email? Subscribe free.
