We gave two versions of the same production system the same difficult exam.
They used the same governed source, answered the same questions and were judged against the same ground truth.
What changed was the complete configuration through which the AI received and worked with the information.
The result was not close.
| Configuration | Fully correct |
|---|---|
| Revised configuration | 60.0% |
| Earlier configuration | 35.1% |
| Difference | +24.9 percentage points |
That means about 25 more correct answers for every 100 questions tested.
The main analysis contained 1,300 matched records. The estimated 95% interval was 21.5 to 28.0 points. Even the cautious end was seven times the study's equivalence margin.
What did the test include?
The evaluation used 1,181 question items over a governed transport dataset containing approximately 380,000 objects.
It included straightforward retrieval and demanding tasks involving multi-step reasoning, complex relationships, aggregation, completeness, ambiguity, missing information and conversation context.
An answer received credit only when fully correct.
The direction held under pressure
On a fixed harder subset with tighter operating constraints, the revised configuration scored 41.8% and the earlier configuration 8.3%: a 33.5-point difference.
With a smaller model, both performed worse, but the revised configuration still led 30.0% to 17.5%.
The advantage was broad:
| Battery block | Difference |
|---|---|
| Complex reasoning under load | +26.4 points |
| Behavioural and ambiguity probes | +22.9 points |
| Tasks designed to favour explicit structure | +20.0 points |
| Single-fact questions | +24.5 points |
How was the evidence checked?
All 272 low-confidence grades received three independent judgements. Forty-four automated outcomes changed and are reflected in the final figures.
A separate audit re-marked 102 high-confidence decisions. Seven permissive verdicts were balanced across configurations and could not explain the gap.
A 100-question control found no environmental drift large enough to account for the result.
The claim has boundaries
This was a comparison of complete configurations, not one isolated technical component.
It covered one domain, read-only questions and a difficult battery. The 60.0% score is not a forecast of routine accuracy.
The defensible conclusion is narrower: in this evaluated system, the revised complete configuration substantially outperformed the earlier one on a paired stress test.
Full methods, results, limitations and independent research are in our public white paper, AI Data Readiness and Answer Quality.
This is Part 2 of Graphshare's AI Data Readiness series.




