←Back to NewsAI News/BenchmarksnewsBenchmarksAgentsStanford's Terminal-Bench-Science Exposes GPT-6 Astra Leading at 63% on Real Science TasksA new 70-task agentic benchmark tests frontier models on real scientific research workflows, with top scores stuck in the low 60s.SourceArtificial AnalysisPublishedSep 24, 2026, 11:30 PMAuthorAlphaSignal NewsroomRead1 min readA new 70-task agentic benchmark tests frontier models on real scientific research workflows, with top scores stuck in the low 60s.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsQwen · newsAlibaba's Qwen Intelligence Splits Phone Automation Across Three Specialized AgentsVals AI · newsVals AI's Terminal-Bench 4.0 Shows Most AI Agents Fail 70% of Expert TasksArtificial Analysis · newsArtificial Analysis Rebuilds AI Leaderboards Around Real Legal and Medical Jobs