T1-Bench: Benchmarking multi-scenario agents in large-scale real-world domains

A benchmark for evaluating agentic LLMs in realistic multi-domain environments with interconnected scenarios.


Latest publications