Insights

Evidence and AI

A fast decision model is not decision governance.

A fast decision model can tell you whether words support a claim. It cannot tell you whether the evidence behind those words is real.

Outweigh research · 7 October 2026 · 8 minute read

The naming trap

“Decision” sounds bigger than the job being done.

In these systems, a decision is usually a narrow, typed judgement: is a statement probably true, which option from a fixed list should be selected, or what score should be assigned? That can be fast and useful. It is not the same as examining a company commitment whose evidence may be incomplete, stale or wrong.

The distinction matters because a precise percentage can look like verification. But a model can only judge the material it receives and the tools it is allowed to use. If it cannot open the source, it cannot know whether the receipt is genuine.

The controlled test

Same words. Two source receipts. Only one was real.

We gave TypeSafe Jev and OpenAI Decisions the same customer quote and the same claim. One record pointed to a public page containing the quote. The other pointed to a plausible-looking address that returned a 404. Neither model had a browser or a retrieval tool. The test ran once, without changing the prompt after seeing results.

“The staff are very helpful, polite, and genuinely willing to assist.”
Claim presented to both models: Pinergy customer service is a strength.

Genuine receipt

exa.ai/library/place/4ddl4h19ssb

Page response
HTTP 200
Exact quote
Found

Counterfeit receipt

pinergy.ie/customer-information/independent-customer-review-2026/

Page response
HTTP 404
Exact quote
Not found

The result

Both models understood the meaning. Neither could see the provenance.

Both said the quote strongly supported the claim. Both also correctly declined to call either record verified and recommended treating both as unverified. The crucial point is that their treatment of the genuine and counterfeit records was almost identical. The separate source check—not the model score—established which receipt was real.

Model
TypeSafe Jev
Genuine
86% supports claim
Counterfeit
89% supports claim
Handling
Use unverified
Model
OpenAI Decisions
Genuine
98% supports claim
Counterfeit
100% supports claim
Handling
Use unverified

What this demonstrates

A text-only judgement cannot authenticate provenance it was not equipped to inspect. Source verification must happen outside that judgement.

What it does not demonstrate

It does not prove either model is generally inaccurate. Nor do these percentages forecast customer behaviour or business success.

The company risk

A sensible score can still sit on top of unsupported evidence.

Imagine an internal proposal quotes a customer review, a competitor offer or a market statistic. A model may accurately say that the quoted words support the proposal. That is useful—but incomplete. If nobody checks that the words really appeared at the stated source, the workflow can give weak evidence the appearance of having passed a control.

The answer is not to reject fast models. It is to give each part of the process the right job, in the right order.

The governed sequence

The model is a component. The process is the product.

01

Capture

Fetch the actual page and retain when it was captured.

02

Verify

Check that the exact words appear at the stated source.

03

Judge

Assess whether the verified evidence supports the claim.

04

Challenge

Expose objections, assumptions and missing evidence.

05

Approve

Keep the consequential choice with an accountable person.

06

Record

Freeze what was known, assumed and decided before the result.

Jev or OpenAI Decisions may help after verification: classifying evidence, applying a bounded rule or sending a question to the right workflow. They do not replace source capture, independent challenge, human approval or a frozen record that can be revisited when the result is known.

Evidence note

Reproducible, bounded and deliberately narrow.

Test date: 7 October 2026. One blind pass. The quote and claim were byte-for-byte identical in both records; only the source receipt changed. Jev ran as typesafe/jev-1.13-20260917 and OpenAI Decisions as gpt-6-luna. The genuine source returned HTTP 200 and contained the quote verbatim. The counterfeit address returned HTTP 404 and did not contain the quote.

See the distinction in practice

Bring one real decision and examine its evidence chain.

Book a decision session