The next step in the τ-bench lineage from Sierra and Princeton's Karthik Narasimhan, whose τ² and τ³ variants feed the Artificial Analysis Intelligence Index: "hyper-tau-bench" makes building the agent the task. A developer agent receives what a real client engagement provides, the records a business keeps, a client who holds the requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models, and must deliver a deployable agent from them. Current systems reach 23.9% against an 82.2% expert baseline.

Paper

evalbenchmarkagentsagentic