Long-Horizon Terminal-Bench: Grok 4.5 Leads 21-Model Ranking but Top Model Solves Only 28% of Tasks
Summary
- • LHTB tests 21 frontier models on 46 long-horizon terminal tasks with hidden verifiers; Grok 4.5 leads with 13/46 tasks solved — the best any model achieves is 28%
- • 29 of 46 tasks unsolved by any model; ~55% of all runs score below 0.25 — agents frequently get stuck, loop, or quit before the 90-minute budget expires
- • Cost does not predict capability: Hy3 (Tencent) at $2.47/task and MiniMax M3 at $6.13/task outperform Claude Fable 5 at $73.11/task on overall score
- • Anthropic models hold ranks 2–4 (Sonnet 5, Opus 4.8, Fable 5) but at 3–7x the per-task cost of leading Grok 4.5
Details
LHTB benchmark design
46 tasks spanning interactive games/puzzles, multimodal analysis, software/reverse engineering, scientific computing, earth/energy systems, security/performance, research reproduction, and professional APEX-style workflows. Companion to Terminal-Bench 2.0. Evaluated with Harbor harness, 90-minute budget per task, 21 frontier models.
Grok 4.5 leads — 13/46 tasks solved
Top leaderboard position: 0.505 mean reward, 13 of 46 tasks solved at strict threshold (R≥0.95), at $11.19/task average cost. No model exceeds a 28% solve rate.
Top 5 leaderboard
1. Grok 4.5 (0.505, xAI, $11.19/task) 2. Claude Sonnet 5 (0.497, Anthropic, $60.37/task) 3. Claude Opus 4.8 (0.492, Anthropic, $39.11/task) 4. Claude Fable 5 (0.487, Anthropic, $73.11/task) 5. GPT-5.6-sol (0.451, OpenAI, $21.14/task)
29/46 tasks unsolved by any model
The median task is not solved by any of the 21 frontier models tested. Only 17 tasks are solved by at least one model. The benchmark is far from saturation.
~55% of all runs score below 0.25
Across all model×task combinations, the majority of agents get stuck, loop, or quit early before the 90-minute budget expires — revealing a fundamental failure mode for long-horizon agentic work.
Cost-capability mismatch across frontier models
Hy3 (Tencent) at $2.47/task and MiniMax M3 at $6.13/task achieve competitive mean rewards with models costing 10–30x more. Claude Fable 5 at $73.11/task ranks 4th; GPT-5.4 at $27.57/task ranks 17th.
Hidden rebuild-from-artifact verifiers
LHTB grades agents by whether they can reconstruct specified artifacts, not by self-reported progress. Prevents score inflation through false claims of completion — a key advance over prior terminal benchmarks.
Stuck/loop/quit failure mode is pervasive
55% of runs fail to reach reward 0.25. Agents across the board struggle to sustain productive work over hundreds of steps — the defining challenge separating current models from reliable long-horizon agentic deployment.
All data from the LHTB GitHub repository (zli12321/LHTB) leaderboard snapshot as reported in TLDR AI. 21 models evaluated through one identical Terminus-2 harness; costs are estimated at list prices per task.
What This Means
LHTB exposes a stark reality about current frontier model agents: even the best — Grok 4.5 — succeeds on barely more than a quarter of long-horizon terminal tasks, and the median task defeats every model tested. The hidden verifier design prevents score inflation, making this one of the most rigorous agentic benchmarks publicly available and lending weight to its sobering results. The cost-efficiency finding is particularly striking: Tencent's Hy3 at $2.47/task and MiniMax M3 at $6.13/task achieve competitive outcomes with models costing 10–30x more per task, suggesting premium pricing does not buy proportionally better long-horizon performance. For practitioners deploying agentic AI on complex multi-step workflows, LHTB underscores that the "stuck, loop, or quit early" failure mode affects the majority of runs across the entire frontier — and that significant capability headroom remains before agents can reliably handle real-world long-horizon tasks.
