Why a fluent deck is harder to check than a rough one
A practitioner protocol for reviewing an AI-drafted strategy deck before it goes to a board or investment committee: what to check, in what order, the five distortions to look for in every number, and what to do with claims you cannot verify.
A rough draft invites scrutiny. The prose stumbles, the numbers look provisional, and anyone reading it naturally asks whether the analysis holds. A fluent deck does the opposite: it answers questions before they are asked. The sentences are complete, the transitions are smooth, and the citations sit exactly where a careful analyst would put them. That polish is precisely what makes the deck dangerous to review, because the reviewer's instinct, calibrated on human work, is to trust what reads well. Fluency is not evidence of accuracy, but it reliably dampens the scrutiny that accuracy requires.
Confabulation is the named failure mode
The failure mode has a formal name. The US National Institute of Standards and Technology, in its Generative AI Profile (NIST AI 600-1, July 2024), describes confabulation as a phenomenon in which generative systems generate and confidently present erroneous or false content in response to prompts, known colloquially as hallucinations or fabrications.[1] The same document makes two observations that matter directly to anyone reviewing an AI-drafted deck. First, generative outputs may include confabulated logic or citations that purport to justify or explain the system's answer, which may further mislead humans into inappropriately trusting that output. Second, NIST lists automation bias and over-reliance among the risks that arise from how humans and AI systems are configured together: the more reliable the technology appears, the less it is questioned.
Neither observation is an argument against using AI in strategy work. It is an argument for a specific reviewing posture. If the failure mode is confident, well-sourced-looking error, then the reviewer cannot rely on the instinct that works for human drafts, which is to probe where the writing is weak. The probe has to run where the writing is strong.
- Fluency dampens scrutiny: a polished deck suppresses the questions a rough draft provokes.
- Confabulation is the named failure mode: confidently stated but erroneous or false content, including fabricated justifications and citations.
- Impressionistic review is therefore insufficient: verification must be procedural, a fixed sequence of checks run the same way every time, not a judgement based on how the deck reads.
The rest of this article sets out that procedure. It is written as a protocol someone can run today on a deck that already exists: triage the claims, open each cited source, compare the sentence, name the distortion, disposition what cannot be verified, and record who signed off.
A claim that carries a source versus a claim that sounds sourced
The first sorting task is to separate two categories of statement that look identical on a slide. An attributable claim has a named, real, openable document behind it: a specific publication, from a specific organisation, that a reviewer can locate and read. Source-flavored prose merely borrows the register of citation. It name-checks plausible institutions, plausible years and plausible attributions, in the grammar of evidence, without a document that opens at the other end. On the slide, the two are indistinguishable. Both read as researched. Only one survives a click.
| Attributable claim | Source-flavored prose | |
|---|---|---|
| What sits behind it | A named, real, openable document | No document, or one that does not match the claim |
| How it reads on the slide | Researched and specific | Researched and specific, identical register |
| What a skim does with it | Passes it as credible | Passes it as credible |
| The only reliable test | Open the source and compare the sentence | Open the source, discover there is nothing to open |
Source-flavored prose survives a skim because it is engineered to. Generative models are trained on text in which claims are followed by attributions, so their output reproduces the shape of attribution fluently, whether or not a document exists behind it. This is not a hypothetical concern. Nature's news team has reported that hallucinated citations, references generated by AI that do not correspond to any real publication, are appearing in the published scientific literature,[2] and an audit of 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN and PubMed Central found a sharp rise in non-existent references following widespread adoption of large language models.[3] The scientific literature is the canary: the audit's authors note that preprint moderation and journal publication processes capture only a fraction of these errors.
The practical conclusion for a reviewer is blunt. There is no reliable discriminator between the two categories other than opening the source. Reading the sentence again does not help. Asking whether the attribution sounds right does not help, because plausibility is exactly what the prose is optimised for. The test is mechanical: for every load-bearing claim, locate the document, open it, and confirm that it says what the slide says it says. Sections three and four describe that comparison in detail.
The source check: open the document and compare the sentence
The core of the protocol is a side-by-side comparison, performed one claim at a time. For each load-bearing claim in the deck: open the cited source, locate the passage the claim rests on, and read the source's wording against the slide's wording. Not the topic against the topic, the sentence against the sentence. Most distortions in AI-drafted decks live at the sentence level, in the space between what a source states and what a summary of that source implies. A paragraph-level or theme-level comparison will not catch them.
The triage order that keeps the pass finite
A full deck contains more claims than anyone can verify in an afternoon, so the pass needs an order. Work through it in three tiers:
- Numbers first. Every figure that carries the argument: market sizes, capacities, percentages, growth rates, dates attached to quantities. These are the claims a board will repeat, and the claims where a single changed word silently changes the meaning.
- Attributions second. Every claim that credits a named institution, author or study with a finding. Confirm the named party exists, is credited with this finding in the source, and is characterised correctly.
- Narrative last. The causal chains and qualitative judgements: that A led to B, that a market is consolidating, that a capability is a differentiator. These are harder to falsify and softer in consequence, so they come after the claims that can be checked to a letter.
When the source is paywalled, offline, or does not exist
Three things can happen when the reviewer opens a citation. The source is accessible and the comparison can proceed. The source is real but inaccessible: paywalled, offline, or available only in a version the reviewer cannot read in the time available. Or the source does not exist at the stated location, or at all. Treat these differently. An inaccessible source is a scheduling problem: the claim stays in the deck only if someone with access performs the comparison, or the claim is downgraded to an assumption until they have. A nonexistent source is not a scheduling problem. The claim has no support, and it moves to the disposition step described later in this article. What the reviewer must never do is let an unopened citation stand because the attribution sounded credible. That is the exact failure the protocol exists to prevent.
Worked example: one rewrite, two distortions
The comparison step is best learned on a real case, and one published source illustrates it completely. The Federal Reserve Bank of Dallas, writing about the February 2021 Texas winter storm and its effect on the petrochemical industry, states: 'As much as 80 percent of U.S. basic organic chemicals capacity was offline after the storm, and up to 60 percent was still offline in mid-March.' The same sentence attributes those figures to estimates from Wood Mackenzie, an energy industry consultancy, and adds that capacity was largely restored by April.[4] A draft summarising that source rendered the first half as: 'roughly 80 percent of basic organic chemical capacity across the region'.
The draft sentence reads perfectly well. It is grammatical, it is close to the source, and nothing in it looks wrong. It contains two distortions, and a reviewer who has learned to name them will find both in under a minute.
| Distortion | Source says | Draft says | What changed |
|---|---|---|---|
| Scope shift | 'U.S. basic organic chemicals capacity', a national figure | 'basic organic chemical capacity across the region', a regional figure | A United States population became a regional one, changing what the number covers |
| Bound shift | 'As much as 80 percent', an upper bound | 'roughly 80 percent', a central estimate | A ceiling on the figure became a typical value for it |
The scope shift first. The Dallas Fed is describing capacity in the United States, and it can do so credibly because Texas alone accounts for the large majority of it. The draft converts that national figure into a regional one. Depending on which region a reader imagines, the number now describes a different population than the source describes, and the slide's argument may rest on the regional reading. The bound shift is subtler and, in a board context, more consequential. 'As much as 80 percent' is an upper bound: the worst case, the most the capacity offline reached. 'Roughly 80 percent' is a central estimate: the typical value, the number to plan around. Presenting a worst case as a central estimate makes a bad scenario read like the expected one, which is exactly the kind of quiet optimism a board should not absorb without knowing it happened.
Note what the check itself consisted of. No forensic technique, no statistical test. The reviewer opened the cited document, found the sentence, and compared the wording. Every distortion in this example was visible in that comparison, and none was visible in the draft sentence alone.
The five distortions to check on every number
The worked example demonstrated two of five recurring distortions. Together they form the checklist a reviewer runs against every load-bearing figure: scope, bound, date, unit, and attribution. Each is a single dimension on which a summary can drift from its source while the sentence continues to read cleanly.
| Distortion | The question to ask | Why it survives a skim |
|---|---|---|
| Scope | What population does the figure cover, and has that population changed in the summary? | The number is unchanged, so the eye sees consistency where the meaning has moved |
| Bound | Is the figure an upper bound, a lower bound, or a central estimate, and does the summary preserve that? | Boundaries are carried by small phrases such as 'as much as' or 'at least', which summaries drop without breaking grammar |
| Date | What period does the figure describe, and does the summary state the same period? | A figure that was true of one year reads as a general fact once the period is dropped |
| Unit | What is actually measured, and is the summary measuring the same thing? | Adjacent measures such as capacity and output, or value and volume, sound interchangeable though they are not |
| Attribution | Who is credited with the figure, and does the source credit the same party? | A plausible institution name satisfies the reader even when the source credits someone else |
Three of these deserve a word of illustration beyond the table. A date distortion occurs when a figure tied to a specific period, a single quarter or a particular month, is restated without that anchor, so a snapshot becomes a trend in the reader's mind. A unit distortion occurs when the measured quantity changes character in the summary: capacity becomes output, a share of one base becomes a share of another, a stock becomes a flow. An attribution distortion occurs when the summary credits an institution with a figure that the source attributes elsewhere, or characterises the source's role more strongly than the source does, for example turning an industry estimate into an official statistic.
Finally, remember that the distortions travel together. The worked example carried two in a single clause, and nothing prevents a rewrite from carrying three. This is why the check is a sentence-level comparison rather than a per-number glance: one comparison can surface several failures at once, and a per-number glance will catch only the first.
What to do with a claim you cannot verify
Some claims will not survive the pass: the source does not exist, it says something else, or it cannot be opened in time. For each such claim there are exactly three defensible dispositions, and the reviewer should choose among them deliberately rather than by default.
- Remove the claim. If it does not carry the argument on its own, the cleanest disposition is deletion. The deck proceeds on what survived checking.
- Flag it as an assumption. If the argument needs the claim, restate it explicitly as an assumption in the deck, visible to the audience as such, and carry it to the assumptions register described below.
- Re-source it. If a document the reviewer has opened supports the claim, replace the citation and the wording with what that document actually says, and add the new source to the verification record.
Why vagueness is not safety
The disposition to resist is the fourth one that occurs to every busy reviewer: softening. An unverifiable number is rewritten as 'significant', 'substantial' or 'a large share', and the slide now reads as confident prose with nothing to check. This is the softening trap. Vagueness is not a reduced claim, it is an uncheckable one, and it reads in a deck as established knowledge rather than as the uncertainty it is hiding. A board member who probes a vague claim finds nothing to probe, which is the problem. If a figure cannot be verified, the honest presentations are removal or an explicit assumption, not a synonym.
The assumptions register is what makes the second disposition workable. It is a single, maintained list of every claim in the deck that stands as an assumption rather than a verified fact, each with its origin and, where possible, what evidence would confirm it. The register belongs in the appendix or the backup material, and it changes the deck's epistemic structure: its credibility now rests visibly on the claims that survived checking, with the remainder labelled. A presenter who can distinguish, on demand, between what was verified and what was assumed is in a materially stronger position in the room than one presenting a deck where every sentence sounds equally certain.
Who signs off, and on what
Verification ends with a named person. The presenter owns the verification record: the list of load-bearing claims that were checked, the source opened for each, and the comparison performed, including the distortions found and the dispositions applied. This is not bureaucracy. It is what converts the protocol from a private habit into a management act, and it is what allows the deck to be defended when a board member asks, weeks later, where a figure came from. The answer is a document, not a memory.
- Each load-bearing claim, in the triage order used: numbers, then attributions, then narrative.
- The source opened for each claim, and the sentence compared.
- Distortions found, named by type, and how each was resolved: corrected, removed, or flagged as an assumption.
- The assumptions register, listing every claim the deck carries as an assumption.
- The name of the person who performed the comparison and the date it was completed.
Sign-off is a management act, not a formality. The World Economic Forum's Empowering AI Leadership oversight toolkit for boards of directors is built on that division of labour: management creates and executes strategy in a world shaped by AI, while the board's role is oversight of management's decisions and their results.[5] A deck that reaches a board or an investment committee is exactly such a decision input, and the person presenting it is the one accountable for what in it is verified and what is assumed. Signing the verification record is how that accountability is made explicit rather than implied.
The protocol described here is deliberately tool-agnostic: it can be run on any deck, from any drafting process. But the effort it requires depends heavily on how the deck was built. Decisity is an AI-native strategy platform whose output is built to be checked: every claim and number in a Decisity deck clicks through to its source, so the comparison step of this protocol runs as a fast pass rather than a forensic exercise, and the verification record assembles itself from the traceable links rather than being reconstructed after the fact. For teams presenting to boards and investment committees, that difference is the difference between verification being a project and verification being part of the workflow. The underlying standard is the same either way, and it is the one this article has applied throughout: a deck is ready when every load-bearing claim in it has been opened, compared, and signed.
Sources
- NIST AI 600-1: Generative Artificial Intelligence Profile
- Nature: Hallucinated citations are polluting the scientific literature. What can be done?
- LLM hallucinations in the wild: Large-scale evidence from non-existent citations
- Dallas Fed: Texas winter deep freeze broke refining, petrochemical supply chains
- World Economic Forum: Empowering AI Leadership, Oversight Toolkit for Boards of Directors



