Evidence and AI
A fast decision model is not decision governance.
A fast decision model can tell you whether words support a claim. It cannot tell you whether the evidence behind those words is real.
Outweigh research · 7 October 2026 · 8 minute read
The naming trap
“Decision” sounds bigger than the job being done.
In these systems, a decision is usually a narrow, typed judgement: is a statement probably true, which option from a fixed list should be selected, or what score should be assigned? That can be fast and useful. It is not the same as examining a company commitment whose evidence may be incomplete, stale or wrong.
The distinction matters because a precise percentage can look like verification. But a model can only judge the material it receives and the tools it is allowed to use. If it cannot open the source, it cannot know whether the receipt is genuine.
The controlled test
Same words. Two source receipts. Only one was real.
We gave TypeSafe Jev and OpenAI Decisions the same customer quote and the same claim. One record pointed to a public page containing the quote. The other pointed to a plausible-looking address that returned a 404. Neither model had a browser or a retrieval tool. The test ran once, without changing the prompt after seeing results.
“The staff are very helpful, polite, and genuinely willing to assist.”
Genuine receipt
exa.ai/library/place/4ddl4h19ssb
- Page response
- HTTP 200
- Exact quote
- Found
Counterfeit receipt
pinergy.ie/customer-information/independent-customer-review-2026/
- Page response
- HTTP 404
- Exact quote
- Not found
The result
Both models understood the meaning. Neither could see the provenance.
Both said the quote strongly supported the claim. Both also correctly declined to call either record verified and recommended treating both as unverified. The crucial point is that their treatment of the genuine and counterfeit records was almost identical. The separate source check—not the model score—established which receipt was real.
- Model
- TypeSafe Jev
- Genuine
- 86% supports claim
- Counterfeit
- 89% supports claim
- Handling
- Use unverified
- Model
- OpenAI Decisions
- Genuine
- 98% supports claim
- Counterfeit
- 100% supports claim
- Handling
- Use unverified
What this demonstrates
A text-only judgement cannot authenticate provenance it was not equipped to inspect. Source verification must happen outside that judgement.
What it does not demonstrate
It does not prove either model is generally inaccurate. Nor do these percentages forecast customer behaviour or business success.
The company risk
A sensible score can still sit on top of unsupported evidence.
Imagine an internal proposal quotes a customer review, a competitor offer or a market statistic. A model may accurately say that the quoted words support the proposal. That is useful—but incomplete. If nobody checks that the words really appeared at the stated source, the workflow can give weak evidence the appearance of having passed a control.
The answer is not to reject fast models. It is to give each part of the process the right job, in the right order.
The governed sequence
The model is a component. The process is the product.
01
Capture
Fetch the actual page and retain when it was captured.
02
Verify
Check that the exact words appear at the stated source.
03
Judge
Assess whether the verified evidence supports the claim.
04
Challenge
Expose objections, assumptions and missing evidence.
05
Approve
Keep the consequential choice with an accountable person.
06
Record
Freeze what was known, assumed and decided before the result.
Jev or OpenAI Decisions may help after verification: classifying evidence, applying a bounded rule or sending a question to the right workflow. They do not replace source capture, independent challenge, human approval or a frozen record that can be revisited when the result is known.
Evidence note
Reproducible, bounded and deliberately narrow.
Test date: 7 October 2026. One blind pass. The quote and claim were byte-for-byte identical in both records; only the source receipt changed. Jev ran as typesafe/jev-1.13-20260917 and OpenAI Decisions as gpt-6-luna. The genuine source returned HTTP 200 and contained the quote verbatim. The counterfeit address returned HTTP 404 and did not contain the quote.
See the distinction in practice