# ExtractBench: Document-Agent Evaluation Guide

> Run ExtractBench to measure document-agent completeness, accuracy, grounding, perception, table structure, and cost before production.

- **Published**: 2026-08-13
- **Category**: AI Infrastructure
- **URL**: https://agentpedia.codes/blog/llamaindex-extractbench-document-agent-evaluation-guide

---

**ExtractBench** is an open benchmark for schema-guided enterprise document extraction. It tests the part of a document agent that is easy to overestimate: returning plausible fields is not enough. A useful system must find long lists, survive scans and handwriting, preserve table structure, attach evidence to values, and do so at a cost the business can afford.

> **Note callout**

**Practical verdict:** use ExtractBench as a starting regression suite and as a design checklist. It is unusually broad, but its published leaderboard is still a source-produced evaluation. Reproduce the harness, pin versions, add your own documents, and separate extraction quality from authorization and workflow safety.

## The practical verdict

LlamaIndex says the benchmark contains **370 enterprise documents and 4,869 pages**, spanning **8 business domains and 67 document types**. Its dataset, evaluation harness, schemas, and methodology are public.

## What ExtractBench measures

ExtractBench is designed around schema-guided extraction: the system receives a document and a user-defined schema rather than only a fixed set of labels. Its challenge axes are independent, so a low result can point to a particular production failure.

| Axis | What it exposes |
| --- | --- |
| Task challenge | Long repeated records, sparse values, dense forms, reconciliation, and nested fields |
| Perception | Born-digital pages, scans, handwriting, rotation, and image-only documents |
| Table structure | Merged headers, cross-page continuation, pivoted tables, huge tables, and tables inside cells |
| Document length | Short, medium, and long records, including documents where truncation is the main failure |
| Business domain | Coverage across filings, procurement, healthcare, valuations, bankruptcy, energy, and other enterprise paperwork |
| Grounding | Whether extracted values can be traced to word- or page-level evidence |
| Cost | Measured cost per page where the system exposes a meaningful price |

The benchmark reports more than one kind of extraction quality. **Value accuracy** asks whether the returned values are right. **Completeness** asks whether the system returned the full record set. **Grounding** asks whether a reader can trace a value back to the source. These are different failure modes: a system can have high precision on the fields it returns while silently omitting most rows in a 100-page schedule.

> Introducing ExtractBench: an open benchmark for schema-guided enterprise document extraction.
>
> -- [@llama_index, August 11, 2026](https://x.com/llama_index/status/2087195936225108171)

## How the ground truth works

LlamaIndex describes three ground-truth pipelines:

1. **Real documents:** multiple systems produce candidate values; agreement becomes a candidate truth, and disagreements receive human review.
2. **Synthetic long lists:** values are generated before rendering the document, so completeness can be scored at large row counts.
3. **Forms:** 169 regulatory and tax forms are reviewed field by field, with bounding boxes placed for most fields.

This is important for fair evaluation. A benchmark answer key should not simply trust one extractor's output, especially when comparing that extractor against competitors. The dataset should be treated as a versioned artifact: record its revision, schema, split, and any local modifications.

The benchmark also separates page perception from schema reasoning. A model that understands a form but reads a handwritten `5` as an `8` has a perception problem; a model that reads the value but attaches it to the wrong row has a structure problem.

## Read the published results

LlamaIndex reports results for 14 systems, including frontier VLMs, self-hosted open-weight models, coding agents running in isolated sandboxes, and specialized extraction APIs. The launch page reports these headline overall value-F1 and cost figures:

| System | Value F1 | Reported cost/page |
| --- | ---: | ---: |
| LlamaExtract Agentic Plus | 95.6% | 8.1c |
| LlamaExtract Agentic | 89.5% | 3.1c |
| LlamaExtract Cost Effective | 86.8% | 1.0c |
| Codex (GPT-5.5) | 93.6% | 27.8c |
| Claude Code (Opus 4.8) | 87.1% | 16.2c |
| Reducto Deep Extract | 90.4% | 34.4c |
| Gemini 3.5 Flash | 79.8% | 1.0c |
| GPT-5.4 Nano | 74.9% | 0.21c |

These are first-party benchmark results from the ExtractBench announcement. They are not an independent audit, and a single mean hides the shape of the workload. LlamaIndex's own analysis says short documents flatten the leaderboard while long documents expose truncation and completeness failures.

ExtractBench's deterministic scoring is worth preserving in a reproduction: repeated records are aligned as an unordered set with globally optimal one-to-one matching; dates are normalized; missing keys score like explicit `null`; failed or missing documents score zero; and the unified value metric does not use an LLM judge. Grounding is evaluated where verified box ground truth exists.

The benchmark also reports that coding agents can perform strongly but at much higher cost, and that grounding remains difficult. A system that returns no evidence receives no grounding credit by design, even if some values happen to be correct.

At enterprise scale, cost differences compound: one cent per page is **$10,000 per million pages**. Compare cost with completeness, grounding, retries, review time, and downstream correction--not with value F1 alone.

## Reproduce the benchmark

The official repository documents a `uv`-based runner. A minimal starting sequence is:

```bash
git clone https://github.com/run-llama/ExtractBench.git
cd ExtractBench
uv sync --extra runners
cp .env.example .env
# Add only the API key for the selected pipeline
uv run extract-bench pipelines
uv run extract-bench download --test
uv run extract-bench run pymupdf_text --test

```

The exact pipeline names, provider requirements, splits, and output locations can change. Pin the repository revision and keep the generated results outside the source checkout when building a repeatable CI job.

Before running a costly suite:

1. list available pipelines;
2. inspect the pipeline's input and output contract;
3. configure credentials through environment variables or a secrets manager;
4. run one small smoke sample;
5. record model/provider versions, schema revision, dataset revision, concurrency, and cost assumptions;
6. run the full split only after the smoke output validates.

The launch leaderboard's price figures use provider prices documented at the benchmark's measurement date, not a permanent price promise. Freeze the provider price sheet and model version when reproducing it; a cost comparison can change without any model-quality change.

The official dataset is available from [Hugging Face](https://huggingface.co/datasets/llamaindex/ExtractBench), and the paper is [arXiv:2607.29677](https://arxiv.org/abs/2607.29677). The repository is Apache 2.0 according to its public metadata; verify the current repository license before redistributing modified evaluation code or data.

## Adapt it to your document agent

ExtractBench is valuable even when you cannot use its documents in your own CI. Map its axes to your private corpus:

| Your test slice | Include examples for |
| --- | --- |
| Long records | 50+ page filings, repeated schedules, continuation tables, and end-of-document values |
| Perception | Scans, rotation, handwriting, low contrast, stamps, and image-only pages |
| Structure | Merged headers, nested tables, row spans, cross-page tables, and repeated groups |
| Grounding | Exact value boxes, page references, source snippets, and an evidence confidence field |
| Schema generalization | New schemas that were not used to tune prompts or parsers |
| Cost | Token, page, OCR, tool, storage, and human-review costs |

Use order-insensitive field scoring when row order is not meaningful, but also measure row completeness and duplicate rate. A document agent that returns the right values in the wrong records can be more dangerous than one that refuses.

A practical output contract should include:

```json
{
  "records": [],
  "evidence": [],
  "warnings": [],
  "pages_seen": 0,
  "complete": false,
  "review_required": true
}
```

Keep `complete` separate from "the parser returned valid JSON." Make the agent state which pages it read, which fields lack evidence, and whether a human must review the result.

## Limitations and fair-use rules

ExtractBench does not prove that an extractor is safe for your organization. It does not replace:

- document access control and tenant isolation;
- retention, deletion, and regional-data policy;
- prompt-injection handling in document text;
- malware and active-content scanning;
- human approval for financial, legal, medical, or account-changing actions;
- monitoring of provider outages and model drift;
- evaluation on your languages, layouts, and private document distribution.

Benchmark documents can also become familiar to models or providers over time. Preserve the benchmark revision, avoid tuning directly on the test set, and maintain a private holdout. If you publish a comparison, identify whether the result is reproduced, vendor-reported, or an internal measurement.

The strongest adoption pattern is layered: use ExtractBench for broad regression coverage, use a private corpus for deployment fitness, and require evidence-backed human review before extracted values trigger consequential actions.

## FAQ

## FAQ

### What is ExtractBench?

ExtractBench is an open benchmark for schema-guided enterprise document extraction. It evaluates value accuracy, completeness on long records, word- and page-level grounding, perception challenges, table structure, and measured cost.

### How large is ExtractBench?

The public benchmark covers 370 enterprise documents and 4,869 pages across 8 business domains and 67 document types, with each type using its own schema.

### Can I run ExtractBench locally?

Yes. The official repository documents a uv-based setup and runner interface. Clone the repository, install its runner extras, list pipelines, and run a named pipeline; provider credentials and compute requirements depend on the pipeline.

### Does a high extraction score prove a production document agent is safe?

No. Benchmark scores do not cover your access controls, retention policy, prompt injection defenses, workflow approvals, or every document distribution. Re-run representative private samples and inspect evidence links before automation.


## Sources and links

- [LlamaIndex: Introducing ExtractBench](https://www.llamaindex.ai/blog/introducing-extractbench)
- [ExtractBench GitHub repository](https://github.com/run-llama/ExtractBench)
- [ExtractBench dataset on Hugging Face](https://huggingface.co/datasets/llamaindex/ExtractBench)
- [ExtractBench paper](https://arxiv.org/abs/2607.29677)
- [ExtractBench pipeline documentation](https://github.com/run-llama/ExtractBench/blob/main/docs/pipelines.md)
- [AgentPedia: LlamaIndex Parse Gateway guide](/blog/llamaindex-parse-gateway-document-routing-guide)
- [AgentPedia: scientific software validation guide](/blog/coding-agents-scientific-software-validation-guide)
- [AgentPedia: Liquid LFM2.5 local-agent guide](/blog/liquid-lfm2-5-2-6b-on-device-agent-guide)


---

- [All articles](https://agentpedia.codes/blog)