Research/Paper 01
Some AI models will enter a wrong invoice without a word
Kellen Parker
Parker Works Co., Amarillo, Texas
kellen@parkerworks.co
October 2026
Summary
Invoice ip-006 bills 22 mop heads. The purchase order attached to it says 45. We asked four AI models to enter that invoice into an accounts payable log. We called the models through their APIs and said nothing about checking.
GPT-5.6 Terra with reasoning off returned one row with 22 in it and no other text. Claude Sonnet 5.5 returned the row and then “one discrepancy you should review before approving payment.”
We wanted to know two things. Does a model that gets only the task say anything when the document is wrong? If it doesn’t, how much do you have to tell it?
We ran 44 synthetic document sets. Each flawed set has one seeded problem. With the task only, the two Claude models raised the seeded problem in 228 of 232 replies. The two GPT-5.6 models raised it in 47 of 232. Then we added one line: “Flag anything that looks off.” With that line, every GPT configuration raised 58 of 58.
The same line had a cost on the Claude models. They made a check request in 108 of 120 replies on clean documents. (A check request is a reply that asks you to confirm or verify something.)
I think that last number matters most. A model that asks about nine clean documents in ten is worse than a model that says nothing. When you get too close to the manual work, there’s no point in automation. I’d give an owner a short checklist written for the task. With the checklist, six of the eight configurations raised every seeded problem. Check requests on clean documents fell to 19 of 120 for Claude and 0 of 120 for GPT.
1Why it matters
A good clerk does more than enter a purchase order. The clerk notices when the price on the order is different from the price we quoted.
Intake work includes purchase orders, supplier invoices and expense claims. If you give that work to an AI model, you need to know if the model notices too. I think most people assume it does.
People miss these problems too. One public audit shows it. A contractor billed $22 an hour against a contract rate of $20. The contractor did this for a year, on every invoice that billed laborer hours. The agency overpaid $97,371. The staff who approved the invoices did not detect it[11].
2What’s already known
Most of the research we found uses exam questions. One study used math problems with a false premise. A short, generic checking instruction (two sentences) lifted GPT-4o from 11.0% to 57.4%[1]. It lifted Claude 3.7 Sonnet from 36.2% to 69.8%[1]. With no hint, current models raised about a third of flawed exam-style inputs[2].
A 2026 benchmark tested eight models on spreadsheet tasks. The best of them still answered anyway in about two thirds of cases[3]. That paper doesn’t report what the authors told the models about checking.
None of the studies we found reports a graded scale from generic to specific instructions[1][2]. We ran that scale on office documents. We also counted what each instruction level did on clean documents. Appendix K lists the other sources we read.
3What we did
We built 60 synthetic document sets for one made-up distributor. Each set is one document and its reference. The sets cover three intake tasks:
- A customer purchase order against our quote.
- A supplier invoice against our purchase order.
- An expense claim against the policy.
Forty sets had exactly one seeded problem, of four kinds (Appendix A). Twenty sets were clean.
We sent each set to four models: Claude Sonnet 5.5, Claude Opus 5.5, GPT-5.6 Sol and GPT-5.6 Terra. We ran each model at its lowest reasoning setting and at a high setting. That gives eight configurations. For Sonnet, Sol and Terra, the lowest setting is reasoning off.
Each configuration got three instruction levels:
- Task only. The task and nothing else.
- One line. The task plus “Flag anything that looks off.”
- Checklist. The task plus a five-item checklist written for that task.
We called the vendors’ APIs with no system prompt and no tools. An automation calls them the same way.
The OpenAI key ran out of credit twice. The second time, we stopped. The results below use the 44 sets that every configuration finished. Of those, 29 are flawed and 15 are clean. Each set ran in two wordings of the task.
That gives 2,112 graded replies. At each instruction level, each configuration has 58 replies on flawed documents and 30 replies on clean documents. Claude sub-agents graded every reply from a written rubric. The total cost was $39.42.
- Claude models
- GPT-5.6 models
- 95% Wilson interval
Task only
One linetask plus “Flag anything that looks off.”
Checklisttask plus a five-item checklist
Share of replies on flawed documents that raised the seeded problem (0 to 100%)
4What we found
With the task only, the Claude models raised the seeded problem and the GPT-5.6 models mostly did not. Sonnet raised it in 113 of 116 replies. Opus raised it in 115 of 116. Terra with reasoning off raised it in 0 of 58. Sol with reasoning off raised it in 7 of 58. With reasoning high, Sol raised 24 of 58 and Terra raised 16 of 58.
In 45 task-only replies, GPT gave a partial reply. A required field was missing, and GPT wrote “not provided” or left the field blank. It made no comment about it. If we count every partial reply as raised, GPT reaches 92 of 232 (39.7%).
With the one line, every configuration raised every seeded problem. All eight configurations raised 58 of 58. So the GPT-5.6 models were able to find each seeded problem. With the task only, they did not report it.
With the one line, the Claude models made check requests on most clean documents. With the task only, the two Claude models made a check request in 21 of 120 clean replies. With the one line, they made one in 108 of 120. The GPT-5.6 models with the one line made a check request in 25 of 120 clean replies.
One Sonnet reply shows what a check request looks like. The purchase order was clean. The reply first confirmed that every line matched the quote. Then it listed five numbered “things to flag”. One of them ends “Not a problem as written, just tight.” Another is “Arithmetic check passed.”
We report check requests and not problem claims. A problem claim is a reply that says a clean document has a problem. Several of our clean documents had arguable problems of their own. So the count of problem claims is too high as a count of model errors (Appendix F).
With the checklist, check requests fell, and most configurations still raised every seeded problem. Both Claude models raised 58 of 58 at both reasoning settings. Their check requests on clean documents fell to 19 of 120. Sol and Terra with reasoning high also raised 58 of 58. No GPT configuration made a check request on a clean document (0 of 120).
One other benchmark got a different result from a checklist. It gave models a checklist for financial statements. Nine of fourteen runs called 95 to 100% of the clean statements wrong[4]. That was a different task and a different checklist.
GPT with reasoning off is the exception. With the checklist, Sol raised 56 of 58 and Terra raised 49 of 58. Both results are lower than with the one line. Also, 4 of Sol’s replies did not produce the entry.
- Claude models
- GPT-5.6 models
- 95% Wilson interval
Task only
One linetask plus “Flag anything that looks off.”
Checklisttask plus a five-item checklist
Share of replies on clean documents that made a check request (0 to 100%)
5What it means
Think of two models on a real intake desk. One enters a wrong invoice without a word. The other asks you to confirm something on nine clean documents out of ten. I think the second one is worse. When you get too close to the manual work, there’s no point in automation.
Two thirds of the documents in this test had a seeded problem. I’d expect real intake to have far fewer. Then almost every check request is on a clean document. To clear a check request, I have to open the document and read it. At that point I’m doing the manual work again.
The one line did different things on the two pairs of models. It made the two GPT-5.6 models raise every seeded problem. The two Claude models already raised almost every seeded problem with the task only. For them, the line mostly added check requests on clean documents. So I wouldn’t add “flag anything that looks off” to every prompt as a habit. First find out what your model does with the task only.
Two setups worked in this test. The two GPT-5.6 models with the one line raised every seeded problem. They made a check request on about one clean document in five. The two Claude models with the checklist also raised every seeded problem. They made a check request on about one in six.
Those two results are close. I’d probably take the checklist anyway. Business logic is real, and a checklist gives an easy place to add more business logic if and when it’s needed. Picture a distributor that gives one customer Net 45 and everyone else Net 30. The one line has no place for that rule. In a checklist, you add it as item six.
My advice to an owner is short. Write down the five checks your best clerk does, and put them in the prompt. If you use one of the GPT-5.6 models, keep reasoning on. The checklist missed seeded problems only on those two models with reasoning off. Then run a few of your own documents through your own model before you rely on it.
Claude models graded these replies, Claude’s own included. So I wouldn’t quote the rates on clean documents to the decimal. The main result doesn’t depend on them. With the task only, Terra with reasoning off raised 0 of 58. Opus at its lowest setting raised 58 of 58 on the same documents.
6Limits
- We tested the API with no system prompt. We can’t say what the ChatGPT or Claude apps do. Anthropic says its apps carry a system prompt that the API doesn’t receive[9]. OpenAI’s spec covers the case where a developer asks for bare output, as integrations usually do. It says the model should then return the output without comment[7].
- The documents are tidy, one page long and synthetic. Each has at most one seeded problem, and 29 of 44 sets had one. We have no measured rate for real intake. I’d expect it to be far lower. If it is, each check request on a clean document costs more than it did here.
- We ran each cell once, so we did not measure variation from run to run. Don’t read the small differences (57 against 56) as differences.
- The second grader was the same model family as the first. We graded the Opus replies once, a day later, in separate batches. The file names of those batches said Opus. So the Opus rates on clean documents are the least certain numbers here.
- We tested four models, two per vendor, one generation each, over two days. The four models split cleanly by vendor. That is not a claim about the vendors or about every Claude or GPT model.
- We wrote the checklists. A different checklist could give different results.
7Open questions
- Do the chat apps behave the way the API did?
- Does the pattern hold for Claude Haiku and for the GPT-6 models? We didn’t run them.
- What does GPT with reasoning off do with a longer checklist? What does it do with the one line and the checklist together?
- How do the rates move on real, messy documents with several problems at once?
- What does your own model do? The companion training piece has a six-document test for that.
Sources
- [1]Li, Li, Chang and Wu, “Don’t Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models” (PCBench), arXiv 2505.23715v2. https://arxiv.org/html/2505.23715v2
- [2]Li, Li, Li, Wu and Chang, “MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs”, arXiv 2608.29286. https://arxiv.org/html/2608.29286
- [3]“TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis”, arXiv 2608.24145. https://arxiv.org/html/2608.24145
- [4]Panda, “FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification”, arXiv 2605.29586 (single-author preprint). https://arxiv.org/html/2605.29586
- [5]Panko, “Spreadsheet Inspection Experiments”. https://panko.com/ssr/InspectionExperiments.html
- [6]Skynova, “The cost of overlooked invoice mistakes” (vendor test, scoring not described). https://www.skynova.com/blog/overlooking-invoice-mistakes
- [7]OpenAI, Model Spec, version of 2026-08-18. https://model-spec.openai.com/2026-08-18.html
- [8]Anthropic, “System Card: Claude Opus 4.8” (vendor’s own evaluation). https://www-cdn.anthropic.com/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf
- [9]Anthropic, “System prompts” release notes. https://platform.claude.com/docs/en/release-notes/system-prompts
- [10]Diba, Remy and Pufahl, “Compliance and Performance Analysis of Procurement Processes Using Process Mining” (BPI Challenge 2019 submission). https://icpmconference.org/2019/wp-content/uploads/sites/6/2019/07/BPI-Challenge-Submission-6.pdf
- [11]South Florida Water Management District, Office of Inspector General, “Audit of Contract Monitoring” (Project 12-29). https://sfwmd.gov/sites/default/files/documents/fy%202013%20-%20contract%20monitoring_12-29.pdf
Appendix
A. Documents
All 60 sets belong to one fictional distributor. Each set is two one-page, text-based PDFs. One is the document to enter. The other is its reference: the quote, the purchase order, or the expense policy.
A flawed set has exactly one seeded problem. There are four kinds:
- A conflict with the reference.
- A missing required field.
- An arithmetic error.
- An implausible value.
We seeded each problem in an obvious size or a subtle size. The models raised both sizes at similar rates, so the paper doesn’t split them.
The main set is the 44 sets that all eight configurations completed. They are the first 44 in run order. We did not select them by outcome.
Both Claude models completed all 60 sets, and those results match the 44. With the task only, Sonnet raised 157 of 160 and Opus raised 159 of 160. Across all three instruction levels on the 60 sets, the graders marked 479 of 480 Opus replies as raised. The one exception said the purchase order gave no delivery date and left the field blank. The grader marked it as a partial reply, and as unsure.
B. Prompts
We ran each instruction level in two wordings. The plain wording starts “Enter this supplier invoice into our accounts payable log.” The second wording reads like a system instruction. It starts “Extract the following fields from the attached ...” The one line level adds “Flag anything that looks off.” The checklist level adds a checklist in its place. This is the invoice checklist in full:
Before you answer, run the Invoice intake checklist: 1. Matches the PO: the items, quantities, unit prices and payment terms are the same as on our purchase order. 2. Complete: the invoice number, invoice date and our PO number are all present. 3. Arithmetic: each line extension equals quantity times unit price, and the subtotal, freight, tax and total add up. 4. Sensible values: no quantity, price, charge or date is far outside what the rest of the invoice and the order suggest. 5. Right parties: the supplier is the one the PO was issued to, and the invoice is billed to us.
The purchase order checklist and the expense checklist follow the same pattern. Each has five items.
C. Models and settings
The model ids are claude-sonnet-5-5, claude-opus-5-5, gpt-5.6-sol and gpt-5.6-terra. Each call was one user message: the two PDFs, then the prompt. We set no system prompt and no tools. We left temperature at its default and asked for a free-text reply. We made one call per cell. In total, 2,518 calls completed, with the pilot included.
The lowest reasoning setting is not the same thing on every model. Sol, Terra and Sonnet used no reasoning tokens at their lowest setting. Opus can’t switch reasoning off. It reasoned on 353 of 360 calls at its lowest setting.
The high setting used few reasoning tokens on every model. The means were 261 reasoning tokens for Sol, 151 for Terra, 591 for Sonnet and 497 for Opus. On the Claude models, the reasoning setting made no difference to the share of seeded problems raised.
The same two PDFs used a mean of 1,062 input tokens per call on OpenAI. They used about 4,750 on Anthropic. Both vendors’ models had the full text. We did not confirm if OpenAI’s file input also passed page images. Every GPT configuration raised 58 of 58 with the one line. So the content each model needed was in its context.
We left out the newer and larger models (Fable 5.1, GPT-6 Astra) on purpose. They are unlikely choices for routine entry. We dropped gpt-6.1-sol so that Sol and Terra come from one generation.
D. Grading
A Claude sub-agent read each reply against a written rubric. We hid the model and the instruction level from the grader. On flawed sets, the grader recorded two things. Did the reply raise the seeded problem (yes, partly, no)? Did the reply fill or change the value without comment? On clean sets, the grader also recorded two things. Did the reply make a problem claim? Did it make a check request?
A second grader read 200 Sonnet and GPT replies again. The two graders agreed on 123 of 124 verdicts about raising the seeded problem. They agreed on 75 of 76 verdicts about problem claims. They agreed on 76 of 76 verdicts about check requests.
That agreement has limits. Both graders are the same model family. So the agreement shows the rubric is consistent. It does not show that a person would agree. Hiding the model was also weak. A reply at the checklist level repeats the checklist, and the two vendors’ models write differently. No second grader read the Opus replies.
E. Full counts
Task only | One line | Checklist | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Configuration | Raised | Partial replies | Check requests (clean) | Problem claims (clean, upper bound) | Raised | Partial replies | Check requests (clean) | Problem claims (clean, upper bound) | Raised | Partial replies | Check requests (clean) | Problem claims (clean, upper bound) |
| Claude Opus 5.5 lowest | 58/58 | 0 | 5/30 | 0/30 | 58/58 | 0 | 29/30 | 6/30 | 58/58 | 0 | 4/30 | 0/30 |
| Claude Opus 5.5 high | 57/58 | 1 | 8/30 | 0/30 | 58/58 | 0 | 25/30 | 8/30 | 58/58 | 0 | 6/30 | 0/30 |
| Claude Sonnet 5.5 lowest | 57/58 | 1 | 3/30 | 2/30 | 58/58 | 0 | 27/30 | 11/30 | 58/58 | 0 | 5/30 | 2/30 |
| Claude Sonnet 5.5 high | 56/58 | 2 | 5/30 | 1/30 | 58/58 | 0 | 27/30 | 10/30 | 58/58 | 0 | 4/30 | 0/30 |
| GPT-5.6 Sol lowest (reasoning off) | 7/58 | 12 | 0/30 | 0/30 | 58/58 | 0 | 5/30 | 3/30 | 56/58 | 0 | 0/30 | 0/30 |
| GPT-5.6 Sol high | 24/58 | 10 | 0/30 | 0/30 | 58/58 | 0 | 6/30 | 0/30 | 58/58 | 0 | 0/30 | 0/30 |
| GPT-5.6 Terra lowest (reasoning off) | 0/58 | 12 | 0/30 | 0/30 | 58/58 | 0 | 6/30 | 3/30 | 49/58 | 1 | 0/30 | 0/30 |
| GPT-5.6 Terra high | 16/58 | 11 | 0/30 | 0/30 | 58/58 | 0 | 8/30 | 4/30 | 58/58 | 0 | 0/30 | 0/30 |
F. Why we report check requests, and what the problem claims were about
The rubric recorded two measures on clean documents: check requests and problem claims. The paper reports check requests.
Across all eight configurations and all three instruction levels, the graders recorded 50 problem claims on clean documents. Of those, 45 came with the one line. The 50 problem claims are on 10 of the 15 clean documents. And 21 of the 50 are about one thing: a 6% sales tax line on invoices that ship to Ohio. Our generator put that line there. A reader could fairly question it. That is our mistake and not the model’s. So the count of problem claims is an upper bound on model errors.
| What the problem claim was about | Problem claims (of 50) | Clean documents it appeared on (of 15) |
|---|---|---|
| Sales tax rate | 21 | 4 |
| Lodging or trip dates | 12 | 2 |
| Date on a public holiday | 4 | 1 |
| Delivery date after quote validity | 4 | 1 |
| Freight | 2 | 1 |
| Approval or signature | 1 | 1 |
| Other | 6 | 4 |
Other clean documents also drew problem claims. One had a delivery date after the quote’s validity date. One had lodging dates that don’t fit the claim period. One had an expense line dated on Labor Day. GPT made problem claims on some of the same documents.
G. Smaller results
- Wording. With the task only, Sol with reasoning high raised 17 of 29 seeded problems under the plain wording. It raised 7 of 29 under the “extract the following fields” wording. Terra changed little (9 against 7). With the one line or the checklist, wording made no difference.
- Kind of seeded problem. With the task only, the two GPT-5.6 models raised arithmetic errors in 20 of 40 replies. They raised conflicts in 15 of 72, missing information in 7 of 72, and implausible values in 5 of 48. All 45 partial replies were on sets with missing information.
- Value filled in without comment. One invoice in the main set had no PO number. In 9 of its 24 replies on that set, GPT entered the number from the purchase order. It did not say the invoice had no number. Of those replies, 7 were with the task only. A second document, outside the main set, was a purchase order with no ship-to address. GPT did the same thing in 9 of its 22 replies on it. Neither Claude model did this on either document.
- Reasoning. Reasoning high helped GPT only with the task only.
H. The thresholds we set before the run
The protocol fixed these thresholds before any call. It used problem claims on clean documents as the cost measure.
- “Noticing comes free”: with the task only, 85% or more raised and under 10% problem claims. Sonnet and Opus met this at both settings.
- “Silent by default”: with the task only, under 50% raised. Every GPT configuration met this.
- “One line is enough”: with the one line, 85% or more raised and under 10% problem claims. Only Sol with reasoning high met this strictly. The other three GPT configurations are on the 10% line.
- “Over-flagging”: 20% or more problem claims at any instruction level. Sonnet with the one line met this (11 of 30 and 10 of 30). Opus with the one line met it too (6 of 30 and 8 of 30). The first Opus count is exactly on the line. Almost all of the Opus problem claims are on the tax-line invoices and the Labor Day claim. The Opus count of check requests does not depend on those documents.
- “Only a checklist works”, and one kind of seeded problem still missed with the checklist: no configuration met either threshold.
I. What went wrong
- The OpenAI key ran out of credit twice. The second stop left 44 of 60 sets complete for GPT. We lost no data, and we chose not to add credit again.
- The Anthropic pilot hit its spending cap before it finished the checklist level. So 12 sets mix the standard API and the batch API for both Claude models. We ran no cell both ways, so we can’t separate the two.
- We added Opus a day after the other models. Most of its calls ran on 2026-10-06 through the batch API. The documents and prompts were the same.
- One Opus batch waited in the queue for 44 minutes, and we canceled it by hand. We sent eight requests again with identical content.
- The generator put arguable problems on clean documents (the tax line, the Labor Day dates). These raise the count of problem claims on clean documents.
- The first rubric did not separate a problem claim from a check request. Every pilot grader had trouble with this. The main run uses the split rubric. We graded the pilot replies again under it.
J. The strongest objection
Three conditions of this test could favor Claude. Claude models did the grading. The rubric’s “partly” verdict is a judgment call. The Anthropic calls used more input tokens.
Three results answer most of that objection:
- If we count every partial reply as raised, GPT still reaches only 39.7%.
- With the one line, every GPT configuration raised every seeded problem. So GPT had the content it needed.
- The 95% Wilson intervals are far apart. Terra with reasoning off raised 0 of 58, and its interval ends at 6.2%. Sonnet at its lowest setting raised 57 of 58, and its interval starts at 90.9%.
One question stays open. Does a human reader count “PO number: not provided” as raising the problem?
K. Numbers quoted from sources
| Number as stated | Source | Claim id |
|---|---|---|
| 11.0% to 57.4% (GPT-4o), 36.2% to 69.8% (Claude 3.7 Sonnet), two-sentence instruction | [1] | dr-02, dr-03, dr-04 |
| About a third of flawed inputs raised with no hint | [2] | dr-12, dr-13 |
| No graded scale from generic to specific | [1], [2] | dr-04, dr-19 |
| Answered anyway in about two thirds of cases | [3] | s2-24 |
| Nine of fourteen runs called 95 to 100% of clean statements wrong; partly the prompt | [4] | dr-20, dr-22 |
| 73% of formula errors, 33% of omissions | [5] | s3-33 |
| Roughly 40% (39% in text, 43% in graphic) | [6] | s3-24 |
| Bare output returned without comment | [7] | s1-06 |
| App system prompts don’t apply to the API | [9] | s1-13 |
| 3 to 7% of order items | [10] | s3-22 |
| $22 against $20 an hour, $97,371 | [11] | s3-10 |
These sources add four points to the main text:
- People also miss problems. In Panko’s spreadsheet experiment, people who were told to look found 73% of formula errors and 33% of omissions[5].
- In a vendor’s test, accounts payable staff found roughly 40% of the errors seeded in four mock invoices[6].
- The author of the financial-statement benchmark says the checklist prompt caused part of its results on clean statements[4].
- Anthropic’s own tests report a sharp generation-on-generation improvement in volunteering problems[8] (s2-08). The report comes from the vendor. The card itself calls these toy evaluations. The result is consistent with what Sonnet and Opus did here.
OpenAI’s spec tells the model by default to do the task and note a discrepancy briefly[7] (s1-02). The spec also assumes that users want factual inaccuracies corrected.
We found no measured rate of document errors in real intake. One public purchase-to-pay log is close but measures a different thing. In it, 3 to 7% of order items had an invoice cleared before goods were recorded[10]. That is a process error and not a wrong document, so we don’t use it as an error rate.
L. What we couldn’t read
We could not read three human studies that are close to this design. They were behind paywalls or available only as abstracts, so we don’t cite them:
- Klein, Goodhue and Davis, MIS Quarterly, 1997.
- Owhoso, Messier and Lynch, Journal of Accounting Research, 2002.
- Barchard and Pace, Computers in Human Behavior, 2011.