←Back to NewsAI News/BenchmarksnewsBenchmarksAgentsVals AI's Terminal-Bench 4.0 Shows Most AI Agents Fail 70% of Expert TasksA fresh 66-task suite pushes agents beyond software into hardware, science, and media, with only three models clearing 30 percent.SourceVals AIPublishedSep 16, 2026, 11:19 PMAuthorAlphaSignal NewsroomRead1 min readA fresh 66-task suite pushes agents beyond software into hardware, science, and media, with only three models clearing 30 percent.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsArtificial Analysis · newsArtificial Analysis Rebuilds AI Leaderboards Around Real Legal and Medical JobsMicrosoft Developer · newsMicrosoft's ThinkingBox Catches AI Agents Lying About Database ChangesQwen · newsQwen's E-Commerce Bench Exposes How Badly AI Agents Fail at Running a Business