Perhimpunan Mahasiswa SUTD Indonesia (PADI
TERMINALWORLD: Benchmarking Agents on Terminal Tasks
Pages
18
Time to read
52 mins
Publication
Language
English
Pages
18
Time to read
52 mins
Publication
Language
English
This document is a technical report introducing TERMINALWORLD, a scalable data engine that automates the reverse-engineering of high-fidelity evaluation tasks from real-world terminal recordings. The engine processes a substantial dataset of 80,870 terminal recordings, resulting in a comprehensive benchmark of 1,530 validated tasks across 18 categories. These tasks encompass a range of operations, from simple commands to complex workflows involving over 50 steps, and include 1,280 unique commands. A curated subset of 200 tasks, known as TERMINALWORLD-VERIFIED, has been manually reviewed to serve as a rigorous testbed for evaluating the performance of various terminal agents and models. The report outlines the challenges of evaluating terminal agents using traditional benchmarks and presents the methodology employed by TERMINALWORLD to create authentic and scalable evaluation tasks. It also discusses the limitations of existing benchmarks and highlights the need for a more accurate assessment of agent capabilities in real-world terminal environments.