Back to Benchmark Library
BENCHMARK CASE
Carillion 2016 Annual Report
Could a prudent decision maker identify structural concerns requiring escalation?
- Document
Carillion plc Annual Report 2016
- Objective
Could a prudent decision maker identify structural concerns requiring escalation?
- Methodology
Identical prompt, identical document, identical context window across NDOR, ChatGPT Plus, and Claude Pro. Scored on executive decision-relevance: did the output produce a usable basis for committee-level action under the question asked?
FINDINGS PER SYSTEM
What each system surfaced
NDOR
- Margin deterioration across core contracting segments — disclosed but not contextualised in the narrative.
- Weak forward revenue visibility relative to reported order-book size.
- Pension scheme exposure approaching balance-sheet-material proportions.
- Cash-conversion ratio diverging materially from reported operating profit.
- Goodwill and intangible-asset balances concentrated against deteriorating performance — impairment risk underdisclosed.
ChatGPT Plus
- Internal contradictions between narrative and accounts surfaced.
- Escalation concluded to be warranted.
Claude Pro
- Strongest contradiction-testing pass against the narrative.
- Output confidence rating of 38/100 on the going-concern position.
- High structural risk concluded.
VERDICT
NDOR produced the strongest decision-control framework — explicitly naming the metrics, the contradictions, and the escalation thresholds. Each of the three systems read the report; only NDOR built the case a non-executive director could act on.
Run the same kind of analysis on a document of your choice.
NDOR applies the same structured validation workflow to contracts, models, proposals, and reports.