RT by @_akhaliq: NEW BENCHMARK ON THE @huggingface HUB Your favorite LLM agent can write a script…
TL;DR - Long-Horizon Terminal-Bench (LHTB), a new benchmark on Hugging Face, evaluates whether LLM agents can reliably complete extended terminal workflows. It targets a key agent capability: maintaining coherence and producing correct system state across more than 300 steps.
- The benchmark contains 46 tasks designed to resist contamination.
- Hidden verifiers assess actual terminal state rather than relying on subjective output evaluation.
- MiniMax M3 leads the reported leaderboard, followed by Kimi K2.7 Code and GLM 5.2.
- The benchmark may also help compare smaller, locally deployable agent models, though no such results are provided.