ByteDance Seed (with Nanjing University, M-A-P and TokenWave.AI) grounds an agent benchmark in market demand rather than researcher intuition: tasks are derived from AI startup products with demonstrated adoption, reverse-engineering their product workflows into 97 deliverable-oriented end-to-end tasks across six professional domains, each scored against ~25 fine-grained rubrics.

The framing — "does agent progress extend to work users actually pay for?" — and the low headline scores (best models around 30%) position it alongside GDPval-style economically-grounded evals rather than capability suites. Filed as a paper pending broader adoption.

Paper

agentsbenchmarkevaluationresearch

Related