Part 2 of 3

AI Data Readiness

What Happened When We Changed the Complete AI Data Configuration?

We gave two versions of the same production system the same difficult exam — same governed source, same questions, judged against the same ground truth. What changed was the complete configuration through which the AI received and worked with the information. The result was not close: 60.0% versus 35.1% fully correct.

James Watt

James Watt

24 July 2026·2 min read
What Happened When We Changed the Complete AI Data Configuration?

We gave two versions of the same production system the same difficult exam.

They used the same governed source, answered the same questions and were judged against the same ground truth.

What changed was the complete configuration through which the AI received and worked with the information.

The result was not close.

ConfigurationFully correct
Revised configuration60.0%
Earlier configuration35.1%
Difference+24.9 percentage points

That means about 25 more correct answers for every 100 questions tested.

The main analysis contained 1,300 matched records. The estimated 95% interval was 21.5 to 28.0 points. Even the cautious end was seven times the study's equivalence margin.

What did the test include?

The evaluation used 1,181 question items over a governed transport dataset containing approximately 380,000 objects.

It included straightforward retrieval and demanding tasks involving multi-step reasoning, complex relationships, aggregation, completeness, ambiguity, missing information and conversation context.

An answer received credit only when fully correct.

The direction held under pressure

On a fixed harder subset with tighter operating constraints, the revised configuration scored 41.8% and the earlier configuration 8.3%: a 33.5-point difference.

With a smaller model, both performed worse, but the revised configuration still led 30.0% to 17.5%.

The advantage was broad:

Battery blockDifference
Complex reasoning under load+26.4 points
Behavioural and ambiguity probes+22.9 points
Tasks designed to favour explicit structure+20.0 points
Single-fact questions+24.5 points

How was the evidence checked?

All 272 low-confidence grades received three independent judgements. Forty-four automated outcomes changed and are reflected in the final figures.

A separate audit re-marked 102 high-confidence decisions. Seven permissive verdicts were balanced across configurations and could not explain the gap.

A 100-question control found no environmental drift large enough to account for the result.

The claim has boundaries

This was a comparison of complete configurations, not one isolated technical component.

It covered one domain, read-only questions and a difficult battery. The 60.0% score is not a forecast of routine accuracy.

The defensible conclusion is narrower: in this evaluated system, the revised complete configuration substantially outperformed the earlier one on a paired stress test.


Full methods, results, limitations and independent research are in our public white paper, AI Data Readiness and Answer Quality.

This is Part 2 of Graphshare's AI Data Readiness series.

James Watt

Written by

James Watt

Graphex Software — the affordance-driven data platform.

Want to learn more?

Get in touch to see how we can help your organisation harness the power of connected data.

Get in Touch

Keep Reading

Related Articles

Five Lessons for Leaders Connecting AI to Business Data

Users don't experience a model, data platform and operating layer as separate technologies. They experience one service and judge whether its answer is useful, complete and trustworthy. Five practical lessons for leaders connecting AI to real business data.

Read more
Why Good Data Is Not Enough for AI

Two executives receive the same facts — one as a dense export, one as a clear briefing. The facts are identical; the ability to use them is not. AI assistants face the same challenge, and the handover between trusted data and the model can quietly decide whether the answer is right.

Read more
Your AI Agents Don't Need the Map. They Need to Know Where They're Standing — Part 3

Part 3 of the series. The quiet inversion that changes everything: an affordance engine doesn't give the agent the full graph — it gives it the current menu. The agent's choice is real. The intelligence is real. And the structure that makes both possible is entirely invisible. This isn't hypothetical. We've been building it.

Read more