←Back to NewsAI News/BenchmarkspaperBenchmarksDataCMU Builds a Synthetic Hospital Where AI Fails 45% of Medical SummariesA fully synthetic longitudinal EHR benchmark from Carnegie Mellon fools physicians and stumps frontier models on real clinical reasoning tasks.SourceAlphaSignalPublishedSep 24, 2026, 4:00 PMAuthorAlphaSignal NewsroomRead1 min readA fully synthetic longitudinal EHR benchmark from Carnegie Mellon fools physicians and stumps frontier models on real clinical reasoning tasks.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsGoogle DeepMind · newsGoogle DeepMind's AlphaGenome Atlas Maps 9 Billion DNA Mutations Without Any GPUTencent Hy · newsTencent's EvolveScaler Exposes How AI Fails Tracking Changing Facts Over 1,200 EventsMicrosoft Research · newsMicrosoft Research's Skala 1.1 Beats Top Chemistry AI at Half the Cost