We measure whether the answers are actually right.
Most document AI ships on vibes. ChatPDF is graded against a fixed set of 60 questions written from 5 real public documents - 237 pages of federal frameworks, financial reports and policy circulars - with every answer scored for correctness, grounding and citation accuracy.
91.5%
answered end to end
54 of 59 questions cleared every gate
90.9%
citations verified
100 of 110 quotes matched the source text exactly
1.6s
median answer time
194ms to find the evidence
$0.53
to run the suite
2.34M tokens across 353 API calls
How we score
Four independent gates, not one blended number
A single accuracy score hides which part of the system failed. We grade retrieval, content, citations and behaviour separately, so a weakness in one never disguises itself as a weakness in another. End to end means all four passed on the same question.
Pass rate by gate
Hover any bar for the 95% confidence interval.
Retrieval
98.3%
The evidence needed to answer reached the model's context window.
Content
96.6%
The answer is correct, complete, and faithful to the retrieved sources.
Citations
93.2%
Every claim is attributable to a quote that exists in the document.
Behaviour
96.6%
The system abstains when it should and respects page scoping.
Where it falls short
Every miss is classified by root cause
When a question fails, it lands in exactly one bucket - attributed to the earliest stage responsible, so a retrieval problem is never mislabelled as a writing problem. This is the list we work from.
Outcome distribution
60 questions in the set
Passed every gate
Correct, grounded, cited and behaving as intended.
54
System error
The case never produced a judgeable answer.
1
Abstention failure
Answered a question the document cannot support, or refused an answerable one.
2
Citation failure
The answer itself is right; the citations are missing, wrong, or unverifiable.
3
Answer quality
Graded on seven dimensions, not a thumbs up
Two independent judge passes score every answer from 0 to 4 on each dimension. When they disagree, a third adjudication pass settles it. Further from the centre is better.
Quality profile
Mean score across every evaluated question, out of 4.
Citations that hold up
Quotes matched to source text
90.9%
Every citation is checked deterministically against the document text, not taken on the model's word. 100 of 110 quotes were found verbatim in the source.
Context precision
How much of what we retrieve is actually on-topic.
Irrelevant context in the window
2.40 / 4
Scored 0 to 4 where lower is better, so a wider bar means a tighter, less padded context window.
Speed
Median and 95th percentile.
- Find the evidence
- 194ms / 459ms p95
- Write the answer
- 1.6s / 3.3s p95
The test set
Real documents, not toy examples
237 pages across 5 public documents - dense tables, appendices, footnotes and multi-column layouts, the material that actually breaks document AI.
End-to-end pass rate by document
Every document carries the same number of questions.
What's in the corpus
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Published 2023-01-26
48 pp
The NIST Cybersecurity Framework (CSF) 2.0
Published 2024-02-26
32 pp
Zero Trust Maturity Model Version 2.0
Published 2023-04-01
32 pp
Economic Well-Being of U.S. Households in 2024
Published 2025-05-28
90 pp
OMB Circular No. A-123: Management's Responsibility for Internal Control
Published 2026-03-10
35 pp
91%
single turn
34 questions
100%
page scoped
15 questions
80%
unanswerable
10 questions
What it costs
Every token is accounted for
The harness keeps a ledger of each billable API call, split by pipeline stage, so cost per answered question is a measured number rather than an estimate.
Cost by pipeline stage
Grading an answer costs more than producing one, deliberately.
$0.0089
per question
$0.0183 at p95
2.34M
tokens used
353 API calls
899.6k
tokens served from cache
42% of all input
127.7k
reasoning tokens
hidden thinking, billed at output rate
Method
How the grading works
The harness runs against the real production pipeline on real models. Nothing is mocked.
01
Fixed corpus
5 public documents are checksummed and pinned, so a change in results can only come from a change in the system.
02
Authored questions
60 questions with reference answers and exact supporting quotes, each verified against the source text before entering the set.
03
Real pipeline
Each question runs through the same retrieval and generation path the product uses, capturing the full trace.
04
Independent judging
2 independent gradings at high reasoning effort - well above what answering uses - with a third adjudication pass whenever they disagree.
Ask your own documents the hard questions.
Every answer arrives with citations you can click straight through to the page it came from.