Est.

Confidence Scoring to Decide Human-in-the-Loop Routing

Calibrate confidence scores to real accuracy before routing documents to automation or review.

Reporter · · 12 min read
Cover illustration for “Confidence Scoring to Decide Human-in-the-Loop Routing”
Manual Review Elimination · September 29, 2026 · 12 min read · 2,714 words

Confidence scoring only works as a routing mechanism when the numbers behind it are calibrated to real accuracy, and the practical answer is a three-lane system built on per-field thresholds, hard business rule overrides, and feedback loops that keep the whole thing honest as documents change. The instinct most teams start with is simple: if extraction accuracy is high, automation is safe. That instinct fails in a specific and expensive way.

Take a document pipeline running at 97% accuracy, a number most vendors would call excellent dev.to. Processing 500 invoices a month at that rate still produces roughly 15 invoices carrying at least one extraction error dev.to. Fifteen sounds manageable until you notice that errors don't distribute themselves evenly across a document set dev.to. They cluster on the fields that matter most, the invoice total, the payment terms, the diagnosis code, precisely because those are the fields most likely to sit inside a messy scan, a nonstandard layout, or a stamped-over signature block.

This is the proof-of-concept cliff, and it's a well-documented one. Systems tested in a proof-of-concept environment routinely clear 90% recognition and look production-ready velt.dev parseur.com. Then real operational volume arrives: poor scans, physical stamps, vendors who format invoices in ways no one anticipated, and the score drops immediately velt.dev parseur.com. A July 2026 study of regulated financial workflows put a number on this erosion directly arxiv.org. Of 72 configurations that cleared a demonstration bar, only 32 survived a production bar, a 56.1% survival rate, marking a structural gap between lab and field conditions arxiv.org. That gap means neither full automation nor blanket manual review is a defensible strategy at scale. Full automation inherits errors nobody catches; universal review burns reviewer hours on documents that never needed a human eye. What's needed instead is a routing layer, and that layer needs to know, field by field, where the model is actually right and where it only sounds right.

What confidence scores are

Diagram: Lab vs. Production: The 56% Survival Gap. Visualizes: Show the erosion from demonstration performance to production performance documented in a July 2026 study of regulated financial workflows.

A confidence score is a number between 0 and 1.0 attached to each extracted field, reflecting the model's internal read on how correctly it interpreted that piece of content. An invoice total pulled from a clean PDF might be 0.99. The same field pulled from a blurry scan with a smudged decimal point might be 0.71. In principle, this is what a routing system needs: a signal that lets high-scoring fields pass straight through while low-scoring fields get flagged for a human, instead of forcing every document through the same all-or-nothing gate.

There are two broad families of confidence estimation running under the hood of most large language models doing this work. One is verbalized confidence, where the model simply states a number it believes reflects its certainty. The other is token-level log-probability estimation, a more mechanical measure derived from how the model actually generated its output.

Confidence is not correctness. A score tells you how certain the model felt about its own answer, not that the answer matches reality. The evidence for this gap is not theoretical. In financial sentiment HITL data, researchers found no statistically meaningful relationship between confidence score and whether a human reviewer ended up correcting the output dev.to. Extractions scored at 70% confidence got corrected at roughly the same rate as those scored at 60% arxiv.org dev.to Infrrd. Everything that follows in a well-built routing system exists to close that gap.

Calibration closes the gap between a score's claim and its delivery

Calibration is the statistical process that forces a model's stated confidence to line up with its actual accuracy. A properly calibrated 90% score should mean the field is correct 90% of the time across a real document population, not just on the samples the model happened to train against velt.dev parseur.com. This sounds like a modest technical adjustment. It is not, and teams cannot assume a vendor's confidence scores are reliable without testing them directly. A 2026 benchmark of vision-language models on document extraction found calibration quality ranging from near-perfect to severely overconfident, depending on the model. It has to be measured directly, against the specific mix of vendors, formats, and scan quality a given operation actually processes.

Three diagnostic tools do that measuring. Expected Calibration Error, or ECE, buckets predictions by confidence level and measures the gap between stated confidence and observed accuracy in each bucket; a lower ECE means tighter alignment. Precision-recall curves map the trade-off between how many correct positives a threshold captures and how much coverage it sacrifices, which is how a team finds the specific cutoff that balances accuracy against however much reviewer capacity actually exists. Reliability diagrams plot the same relationship visually, predicted confidence against observed accuracy, where perfect calibration traces a diagonal line and any bulge above or below it shows exactly where a model is overconfident or underconfident.

This kind of measurement used to be difficult simply because there was nowhere near enough low-accuracy data to test against. Most benchmark sets skew toward clean, well-scanned documents, leaving the messy end of the distribution, the exact region where calibration failures matter most, thin on evidence. ConfBench, released in August 2026, was built specifically to fix that arxiv.org dev.to. It applies 20 controlled degradation pipelines to a diverse document set, producing 1,346 variants and more than 70,000 entity-level evaluations spanning the full accuracy spectrum, from clean to badly degraded arxiv.org dev.to. A companion effort, VerifyDocBench, pushes the same idea to field-level granularity, assessing calibration, selective risk, and grounding on a per-field basis rather than treating a document as one aggregate score. Calibration isn't a switch a team flips once during setup. As document sources shift, so does the accuracy behind every stated confidence number, which means calibration needs to be re-measured on a schedule, not assumed to hold indefinitely.

The three-lane routing architecture: straight-through, field review, and full exception

Diagram: The Three-Lane Routing Architecture. Visualizes: Illustrate the three-lane document routing system described in the article.

The most common design mistake in HITL routing is treating "uncertain" as a single category dev.to. Every document that fails to clear a threshold gets dumped into one undifferentiated review queue, and that queue grows until it overwhelms whatever reviewer capacity exists, defeating the entire point of triage in the first place. The pile wins.

A better architecture splits routing into three lanes. Lane two is field-level review, where a document is otherwise processable but one or two specific fields fell below threshold; only those fields get surfaced to a reviewer, and the rest of the document keeps moving. Lane three is full exception: documents where confidence collapsed broadly, or where a critical field is missing entirely, get pulled out for full manual handling or escalation to a specialist.

The distinction between lanes two and three isn't cosmetic, it changes what the reviewer's screen actually needs to show. A lane-two reviewer needs a side-by-side view of the original document against the extracted values, with the specific flagged fields highlighted so attention goes exactly where it's needed. A lane-three reviewer is often re-keying an entire record or kicking it up to someone with domain expertise, a fundamentally different task with different tooling requirements. Collapsing both into one queue with one interface means building for neither job well.

There's a useful industry yardstick here. Accounts payable, an industry with a long history of measuring this exact metric, reports an average touchless processing rate of 32.6%, with best-in-class operations reaching 49.2% Ardent Partners. A properly tuned three-lane system, after roughly three months of adjustment, should be aiming for 60 to 75% straight-through Ardent Partners. That gap is the entire argument for building the three-lane structure carefully rather than treating confidence routing as a bolt-on feature. Full human review on its own can reach roughly 99.9% accuracy, against roughly 80% for systems running fully automated with no oversight velt.dev. The three-lane structure is what makes that 99.9% figure reachable without paying the cost of reviewing every single document to get there velt.dev. Lane 1 (straight-through): high-confidence fields are auto-approved and passed downstream without human involvement, typically comprising 70–90% of document volume when thresholds are well-tuned arxiv.org dev.to velt.dev parseur.com.

Per-field threshold configuration

Not every field carries the same risk if it's wrong, so applying one confidence threshold across an entire document is close to guaranteed to get the balance wrong somewhere. Set it too low and financially material fields slip through unchecked; set it too high and low-stakes fields flood the review queue for no good reason.

A workable starting framework looks something like this, built around what actually happens downstream when a field is wrong. Invoice totals and payment amounts are 0.92 or higher, because an error there is directly financial. Invoice numbers and reference numbers are around 0.90, since downstream matching systems depend on exact values, not approximate ones. Date fields, due dates and contract dates especially, are also near 0.90, because a wrong due date creates a payment timing failure. Vendor and party names can run a little lower, around 0.85, since errors there tend to be visually obvious to a reviewer scanning the document anyway. Line item quantities are also around 0.85, given that three-way matching depends on them being right. General description fields can drop to 0.75, since they're lower stakes and can be checked by periodic sampling rather than individual review. Document classification stays high, around 0.90, because misclassifying a document type routes it into the wrong workflow entirely.

None of these numbers are universal law. They're starting points, and the right threshold for any given field depends on the document type, the quality of the input channel, and how much error tolerance the downstream system actually has. The sane way to tune them is to start conservative, meaning higher thresholds and more human review than might feel necessary, then measure straight-through rate and error rate over the first month and loosen thresholds as the model's real performance on that specific document type gets confirmed.

If more than half of all extractions are falling into review, that's not a threshold problem, it's a model quality problem, and no amount of threshold tuning fixes an extractor that's fundamentally struggling with the input parseur.com. HITL exists to handle the genuine edge cases in an otherwise strong system, not to compensate for a weak one. The sweet spot where human review adds the most value sits in a fairly narrow band: 93 to 98% raw accuracy on a given document type. Below that, around 85%, the problem is document quality or model selection, and review can't efficiently solve it. Above roughly 99.5%, review starts adding friction without adding much benefit, since there's almost nothing left to catch.

Deterministic business rule overrides that bypass confidence scoring entirely

Confidence scoring has a hard ceiling: a model can be entirely confident and entirely wrong, because confidence reflects the model's self-assessment, not any external check against ground truth. Certain risks simply cannot be caught by asking the model how sure it feels, and those risks need rules that fire regardless of what the confidence score says.

A duplicate invoice number extracted at 0.99 confidence is still a duplicate invoice, and no amount of extraction certainty changes that fact. A failed three-way match, where quantity, price, and receipt data disagree with each other, is something basic arithmetic catches instantly but that the extraction model has no built-in way to notice on its own. A vendor's bank details changing between invoices is a fraud signal that has nothing to do with how cleanly the new account number was read off the page. Certain fields, particularly in healthcare, finance, and legal contexts, require documented human verification as a matter of policy, independent of how well the AI is performing. And under the EU AI Act's Article 14, effective in 2026, human oversight is now a legal requirement for high-risk AI systems including credit scoring, employment screening, and medical devices. Some of this routing exists because compliance demands it, not because a confidence score dipped below threshold.

These rules sit downstream of both extraction and confidence scoring, evaluated as a final gate before anything moves forward. They're not competing with calibration work, they're covering the exact territory calibration cannot reach. Confidence scoring routes on uncertainty about what the model read. Business rules route on known risk that exists independent of what the model read. A system needs both, and neither one substitutes for the other.

Feedback loops that keep routing decisions accurate as document distributions shift

Thresholds calibrated against last quarter's document mix don't stay accurate forever. New suppliers show up with unfamiliar invoice templates, seasonal filings change the document types running through the queue, scan quality drifts as someone switches office printers.

The fix is treating every human correction as a labeled data point worth capturing, including the field type, the model's original confidence score, what it output, and what the reviewer actually determined was correct. That correction data does three distinct jobs. It drives threshold adjustment, so if 0.88-confidence invoice totals are getting corrected often, the threshold for that specific field needs to move up. It feeds calibration re-measurement, giving ECE and reliability diagram analysis fresh ground truth to work from and surfacing whether the model has quietly become overconfident on some new document subtype. It surfaces recurring failure patterns that should be addressed through retraining or prompt changes, rather than papering over the same mistake indefinitely with review queues.

There's a second, separate loop that matters just as much: auditing a random sample of documents that went straight through, not the ones that got flagged. This is the only mechanism that catches a threshold set too low before a downstream system finds the error on its own, and it's a genuinely different check from reviewing what got flagged, since flagged items already got a human's attention. MIT Sloan research on human-AI collaboration found a team combining both achieved 90% accuracy on a bird classification task, against 81% for humans working alone and 73% for AI working alone, though the broader body of that research also found human-AI combinations don't reliably beat the best of either working solo velt.dev parseur.com. Pairing doesn't automatically help, but a feedback loop that actually teaches the model where its blind spots sit compounds over time in a way a static system never does velt.dev parseur.com. That compounding is what moves a team off the industry average touchless rate of 32.6%, past the 49.2% best-in-class mark, toward the 60 to 75% range a genuinely tuned three-lane system can sustain Ardent Partners.

Benchmarking infrastructure that makes calibration and routing claims verifiable

Before 2026, there was no standardized way to evaluate confidence calibration for key information extraction at all, which meant every vendor's calibration claim was effectively unverifiable against anyone else's. That's changed, and it matters because a routing architecture's trustworthiness depends directly on the calibration claims it is built on.

ConfBench, released in August 2026, fills a specific hole arxiv.org dev.to. Rather than testing against clean documents where nearly everything scores well, it applies 20 controlled degradation pipelines to generate 1,346 document variants and over 70,000 entity-level evaluations, deliberately populating the low-accuracy region of the distribution that older benchmarks left mostly empty arxiv.org dev.to. That's the exact region where a routing threshold either earns its keep or fails.

A second benchmark, ExtractBench, released between July and August 2026, takes a broader enterprise view: 4,869 pages across 370 documents spanning eight business domains and 67 document types, evaluating value accuracy, record completeness, grounding, and cost together rather than in isolation. It closes two gaps that earlier evaluation work left open: no end-to-end benchmark existed for PDF-to-JSON extraction at real enterprise schema breadth, and no benchmark treated nested extraction correctness with any real precision, exact match for identifiers, tolerance bands for quantities, semantic equivalence for names, rather than one blunt accuracy number covering everything. Its ground truth was built from real documents checked through multi-system agreement with adjudication on disagreements, synthetic long-form lists, and 169 regulatory and tax forms reviewed field by field with bounding boxes.

None of this benchmarking work replaces measuring calibration against a team's own documents directly. But it gives the field something it didn't have before: a shared, independently reproducible way to check whether a confidence score, and the routing architecture built on top of it, actually deserves the trust being placed in it. RealDocBench, dated June 2026, serves as benchmarking infrastructure that makes calibration and routing claims verifiable.

Sources

  1. Human-in-the-Loop Document Review: When to Use It and How to Set It Up (2026)
  2. Human-in-the-Loop Workflows for AI (June 2026)
  3. Human in the Loop AI Review Layer (May 2026)
  4. Human-in-the-Loop Workflow | Parseur®
  5. arxiv.org

More in Manual Review Elimination