August 28, 2026, (Inside AI) — Baidu has released DuMateBench, a new evaluation leaderboard that measures whether AI agents can finish real-world office tasks and deliver usable outputs, not just generate answers. The benchmark covers more than 200 tasks across six categories and tests agents in complex operating environments.
The launch addresses a persistent gap in agent evaluation. Most benchmarks score models on answer accuracy. DuMateBench scores agents on task understanding, tool use, continuous execution, and final-result delivery. That shift matters because enterprise adoption depends on reliable completion, not conversational fluency.
Why Completion Metrics Lag Behind Model Benchmarks
Baidu said DuMateBench uses a general evaluation framework and open interfaces. Different models and agents can be tested under the same criteria. That design allows direct comparison across proprietary and open-source systems, a rarity in agent evaluation.
The six task categories include scheduling, document processing, data retrieval, and multi-step workflows. Each task requires an agent to interpret instructions, select tools, execute actions, and produce a final artifact. Partial progress earns no credit.
Industry analysts note that agent benchmarks often inflate scores by testing narrow skills. DuMateBench’s focus on deliverable outputs mirrors real business requirements. A travel agent that books a flight but fails to send a confirmation email is not useful.
Agent Evaluation Faces Fragmentation and Gaming
Baidu’s move follows similar efforts from research labs and cloud providers. OpenAI, Google DeepMind, and Microsoft have published internal agent evaluation suites. None has become a universal standard. Fragmentation slows enterprise procurement and makes vendor claims hard to verify.
DuMateBench’s open interfaces could attract third-party submissions. Baidu has not disclosed whether it will publish a public leaderboard with live rankings or restrict access to partners. The company also did not specify how it prevents benchmark contamination or gaming.
Benchmark gaming remains a serious concern. Agents can memorize task templates or exploit interface quirks. Without rigorous anti-contamination protocols, leaderboard scores may not reflect real-world reliability. Baidu has not detailed its mitigation strategy.
The launch signals Baidu’s ambition to shape agent standards beyond China. The company competes with Alibaba and Tencent in domestic enterprise AI. A widely adopted benchmark could give Baidu influence over how global buyers evaluate agent platforms.
DuMateBench’s real test will be adoption. Benchmarks gain credibility only when independent researchers and enterprises submit models. Baidu’s own ERNIE agent will likely score well, but third-party validation is essential for trust.
Baidu has not announced a timeline for public leaderboard access or third-party evaluation results. The company said the benchmark is available to partners and researchers through open interfaces.