+24.9pp
The difference in fully correct answers between two complete system configurations in the main paired condition. The revised configuration achieved 60.0% strict accuracy, compared with 35.1% for the earlier configuration across 1,300 matched records.
The model is only one part of the answer
Think of a data export and a decision briefing. Both can contain the same underlying facts. Only one may make the important points easy to recognise and use. AI assistants face the same challenge when they work with complex organisational information.
Access to trusted data is necessary. Making that data usable for the task is a separate design and evaluation problem.
Graphshare compared two complete configurations of the same production system. Both worked from the same governed data and answered the same questions. The main evaluation found a 24.9 percentage-point difference in fully correct answers.
01 — Result
Large difference
60.0% strict accuracy for the revised configuration versus 35.1% for the earlier one.
02 — Support
Direction held
The revised configuration also led under tighter constraints and with a smaller model.
03 — Boundary
System-specific
The evidence supports the evaluated complete configuration, not a universal technical law.
Key Takeaways
- Evaluate the complete route from governed information to a useful answer — not the model alone.
- Use representative business questions, not demonstrations selected for success.
- Define correctness before testing, and report adverse findings alongside headline metrics.
- Repeat evaluation when models, data or workflows change.
A paired test of complete production configurations
The study compared the revised and earlier system configurations as they were intended to operate. Their associated instructions and operating choices remained part of each configuration. This makes the result realistic for deployment decisions, while also limiting the claim: the study does not identify the isolated causal contribution of any one technical element.
1,181
generated question items
380,248
objects in the evaluation space
4,301
measured records across all phases
Question battery
The governed public-sector transport dataset included routes, stop points, infrastructure, incidents and other connected operational information. The battery combined:
Complex work
Multi-step reasoning, relationships, fan-out, aggregation, comparison and complete enumeration.
Behavioural probes
Ambiguity, missing information, completeness, direction, absence and conversation context.
Structure-favouring tasks
Questions selected to give the earlier configuration a plausible home-field advantage.
Single-fact baseline
Straightforward questions used to separate basic retrieval from complex reasoning.
Evaluation phases
| Phase | Purpose | Scope | Measured records |
|---|---|---|---|
| Main paired comparison | Primary accuracy result | Full battery, primary model | 2,601 |
| Constraint pressure | Tighter operating allowance | Fixed hard subset, 3 repeats | 1,200 |
| Model pressure | Smaller-model comparison | Same fixed hard subset | 400 |
| Drift control | Check environmental movement | 100-question repeat | 100 |
The difference was large and precisely estimated
Pairs
1,300
Matched records in the main analysis after one documented run error.
95% interval
+21.5 to +28.0
The plausible range for the paired difference, allowing for related questions.
Discordant pairs
399 vs 76
Revised-only correct versus earlier-only correct; McNemar p < 1 × 10⁻¹⁵.
The advantage was broad
| Battery block | Revised | Earlier | Difference |
|---|---|---|---|
| Complex reasoning under load | 60.5% | 34.1% | +26.4pp |
| Behavioural and ambiguity probes | 50.4% | 27.5% | +22.9pp |
| Structure-favouring tasks | 51.7% | 31.7% | +20.0pp |
| Single-fact baseline | 66.3% | 41.8% | +24.5pp |
Percentage points are the direct difference between two percentages. A 24.9-point gap means approximately 25 more fully correct answers for every 100 questions tested.
The direction held, but the interactions challenged expectations
| Condition | Pairs | Revised | Earlier | Difference | 95% interval |
|---|---|---|---|---|---|
| Main model, standard allowance | 1,300 | 60.0% | 35.1% | +24.9pp | +21.5 to +28.0 |
| Main model, constrained hard subset | 600 | 41.8% | 8.3% | +33.5pp | +27.0 to +40.8 |
| Smaller model, hard subset | 200 | 30.0% | 17.5% | +12.5pp | +5.7 to +19.4 |
Constraint pressure
The planned expectation was that tighter operating constraints would widen the difference. On the fixed 200-question hard subset, however, the difference was already 33.5 points under the standard allowance: 43.0% versus 9.5%. Under the tighter allowance it remained 33.5 points: 41.8% versus 8.3%.
Model pressure
The study also expected a smaller model to depend more heavily on the revised configuration. The opposite occurred. Accuracy declined in both configurations and the difference compressed to 12.5 points.
These findings strengthen the system-specific conclusion while rejecting two stronger predictions in the original study design. Reporting both outcomes helps distinguish evidence from post-hoc storytelling.
The headline survived adjudication, audit and drift checks
Blind adjudication
272 low-confidence decisions
Each received three independent judgements: 243 were unanimous, 28 were decided by a 2-of-3 majority, and one three-way split was resolved by the stated rubric. Forty-four automated outcomes changed in the final statistics.
Confident-grade audit
102 decisions re-marked
Agreement was 93.1%. All seven disagreements were overly permissive automated decisions. They were balanced across configurations and far too few to explain the headline difference.
A/B/A drift control
100 questions repeated
Seventy-six verdicts reproduced exactly. The changed verdicts were balanced 13 to 11, and efficiency measures moved only a few percent. No material drift was found at the control's resolution.
Record accounting
4,301 measured records
The full study inventory reconciled across all phases. One documented run error accounts for the single missing record from the 4,302-record plan.
Observed failure patterns
Outcome-level analysis found that the configurations differed both in whether the system completed an answer and in accuracy when an answer was produced. Because the associated instructions changed with the complete configuration, this paper does not attribute that behaviour to one isolated technical cause. No internal decision logic is disclosed.
Where the revised configuration did not win
- The earlier configuration narrowly led on pagination totals: 7 correct versus 6.
- The earlier configuration narrowly led on deliberately ambiguous questions: 28 versus 26.
- The revised configuration produced more output-hygiene flags involving internal references: 184 versus 108.
- Both configurations largely failed absent-entity questions: 1 correct answer out of 40 in each.
Other research supports the broader thesis, not a universal winner
Tool-output processing. Kate and colleagues evaluated 15 models on 1,298 questions derived from complex tool responses. They found that structured outputs remained difficult even for frontier models and that different processing strategies could shift performance substantially.
Graph encoding. Talk Like a Graph reported gains ranging from 4.8% to 61.8% from selecting a suitable graph encoding.
Model and task dependence. KG-LLM-Bench compared five textualisation strategies across seven models, finding spreads of up to 17.5 percentage points and no one best choice for every model and task.
Specialised interfaces. Retrieve-Rewrite-Answer and StructGPT provide further evidence that the interface between structured sources and a language model can affect question-answering performance.
Scope and limitations
- Complete-configuration comparison: connected elements changed together, so the contribution of any one element is not isolated.
- Single domain: one governed transport dataset containing approximately 380,000 objects.
- Read-only scope: the study did not test business actions or writes.
- Limited model coverage: one primary and one smaller model condition.
- Stress battery: absolute accuracy is not a forecast of routine production performance.
- Sequential execution: the drift control bounds, but cannot eliminate, a disturbance that appeared and disappeared between its two measurements.
- Residual grading risk: audit and adjudication reduced error but do not make automated grading perfect.
- Generalisation: replication is required across more organisations, domains, models and workflows.
Evaluate AI in the context that matters
For executive sponsors
- Judge the finished service, not the model alone.
- Use representative, decision-relevant work.
- Require an explicit definition of fully correct.
- Demand limitations and adverse results alongside headline metrics.
For evaluation teams
- Pair questions across candidate configurations.
- Freeze ground truth and scoring before opening outcomes.
- Audit confident grades as well as uncertain ones.
- Check environmental drift when phases run sequentially.
Recommended follow-up studies
- Replicate across additional domains and model families.
- Separate connected configuration changes in controlled ablations.
- Extend from read-only questions to bounded business actions.
- Measure user effort and complete operational cost alongside answer quality.
- Repeat representative tests after material model, data or workflow changes.
References
- Kate, K. et al. (2026). "How Good Are LLMs at Processing Tool Outputs?" EACL 2026. aclanthology.org/2026.eacl-long.134
- Markowitz, E. et al. (2025). "KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs." arxiv.org/abs/2504.07087
- Fatemi, B., Halcrow, J. and Perozzi, B. (2024). "Talk Like a Graph: Encoding Graphs for Large Language Models." ICLR 2024.
- Wu, Y. et al. (2023). "Retrieve-Rewrite-Answer: A KG-to-Text Enhanced LLMs Framework for Knowledge Graph Question Answering." arxiv.org/abs/2309.11206
- Jiang, J. et al. (2023). "StructGPT: A General Framework for Large Language Model to Reason over Structured Data." EMNLP 2023. aclanthology.org/2023.emnlp-main.574