Benchmark for document agents

DocOps

A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

Jiazhen Jiang1,2 Boxi Cao1 Lingyong Yan3 Yaojie Lu1 Hongyu Lin1

Shuaiqiang Wang3 Dawei Yin3 Xianpei Han1 Le Sun1

1Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences

2University of Chinese Academy of Sciences   3Baidu Inc.

{jiangjiazhen2025, caoboxi}@iscas.ac.cn

Introduction

Autonomous LLM agents are moving from passive chat interfaces toward active participants in digital workspaces. Yet current document evaluations remain constrained by two dominant paradigms: (1) static document understanding, which treats documents as read-only repositories for extraction or question answering, and (2) workflow-oriented software evaluation, which treats the document as a passive payload moved between applications.

Consequently, it remains an open research question whether current agents can reliably execute end-to-end document-centric tasks while maintaining global document state consistency, and achieving user objectives without introducing destructive modifications.

To this end, we introduce DocOps, a rigorously verifiable evaluation framework designed to systematically assess LLM agents on complex end-to-end document manipulation. A primary contribution of this work is the formulation of a principled taxonomy that deconstructs the operational black box of document manipulation.

As illustrated in the overview figure, our taxonomy maps the operational space across two orthogonal axes: atomic capabilities and workflow depth. First, we isolate document interaction into three interdependent dimensions: content, format, and structure. Second, we establish a four-tiered complexity gradient: L1/L2 isolate single- and multi-dimensional atomic edits, while L3/L4 escalate to long-horizon, cross-document workflows inspired by real-world interaction scenarios.

This hierarchical taxonomy serves not merely as a categorization scheme, but as a fine-grained agentic diagnostic framework, enabling researchers to localize whether failures stem from localized structural disruption or a collapse in long-term state tracking. Each task is paired with a deterministic native-document verifier that checks target postconditions, native validity, and preservation constraints on the final artifact.

Overview framework of DocOps.

Experimental Results

LLM Agent DocTools Terminus-2 Codex w/ Skill Codex w/o Skill Claude Code w/ Skill Claude Code w/o Skill
Closed-source models
GPT-5.50.1380.5240.6710.648----
GPT-5.40.1190.4520.6620.638----
Claude Sonnet 4.60.1190.419----0.5190.552
Open-source models
DeepSeek-V4-Pro0.0570.4670.4240.4000.4000.433
Gemma4-31B0.0330.2810.3380.3240.3000.233
Qwen3.6-27B0.0950.4760.2100.2480.4240.386
Qwen3.6-35B-A3B0.0860.3860.2190.2050.3380.329
Qwen3.5-122B-A10B0.1000.3620.2810.2900.2900.295
Qwen3.5-27B0.1140.3620.2950.2710.3710.300
Qwen3.5-35B-A3B0.1050.3000.1760.1570.2950.262
Qwen3.5-9B0.0810.2330.0860.0620.1760.181
GLM-4.5-Air0.0710.162----0.1670.129