Benchmarking document extraction on African identity documents
We evaluated four leading general models against ours on 12,000 regional documents. Here is the full table, including the categories where we lose.
General-purpose document models are trained on corpora that barely include the documents our customers process. We suspected the gap was large. It is larger than we expected.
Method
12,000 documents across nine countries: national identity cards, driver's licences, voter registration cards, bank statements and utility bills. Human-verified ground truth. Field-level exact match, with partial credit reported separately.
Results
| Model | Field accuracy | Full-document accuracy | | --- | --- | --- | | Cosmostack Applied AI | 94.2% | 81.7% | | General model A | 71.4% | 48.2% | | General model B | 68.9% | 44.1% | | General model C | 66.2% | 41.8% | | Open baseline | 52.0% | 26.4% |
Where we lose
On clean, machine-printed English bank statements, general model A edges us by 1.3 points. That is a real result and we are publishing it. Our advantage comes from handwritten fields, non-Latin scripts, and the document layouts that simply do not appear in general training data.
Why publish the losses
Because a benchmark that only shows wins is marketing, and our customers are engineers who can tell the difference.