Why it matters: Most document-extraction pipelines lack labeled ground truth, so teams either skip validation or hand-count small samples; this gives a concrete methodology (hand-counted test page + fabrication tracking) and real cost/accuracy numbers for a vision-based approach.
How to apply: For extraction tasks over dense/irregular documents (merged tables, color-coded markers), hand-count a small representative page as your ground truth, render pages to images, and compare vision-model extraction against both correctness and fabrication rate, not just recall.