T1-Bench: Benchmarking multi-scenario agents in large-scale real-world domains
A benchmark for evaluating agentic LLMs in realistic multi-domain environments with interconnected scenarios.
Recent advances in large language models (LLMs) have enabled increasingly capable agentic systems, yet existing benchmarks remain limited in scale, realism, domain diversity, and temporal complexity, making it difficult to evaluate long-horizon reasoning and reliability. We introduce T1-BENCH, a high-fidelity benchmark for evaluating agentic LLMs in realistic multi-domain environments with interconnected scenarios that require reasoning over prior customer interactions and historical context. We also present a comprehensive and reproducible evaluation framework combining automatic metrics, fine-grained behavioral analysis, and LLM-as-a-judge assessment. Benchmarking a broad range of proprietary and open-weight models reveals that even state-of-the-art systems struggle with long-horizon reasoning, context retention, and reliable decision-making in interconnected environments. We release both the benchmark and evaluation framework to support future research on scalable and trustworthy agentic AI systems.
Latest publications
RECAP: Regression evaluation for continual adaptation of prompts
A benchmark that measures continual-learning phenomena at the constraint level for prompt-level adaptation methods.
EMNLPGRAID: Synthetic data generation with geometric constraints and multi-agentic reflection for harmful content detection
A novel pipeline that leverages Large Language Models (LLMs) for dataset augmentation.
EMNLPseqBench: A tunable benchmark to quantify sequential reasoning limits of LLMs
A parametrized benchmark for probing sequential reasoning limits in LLMs.
EMNLP