Can We Reproduce an AI Strategy Analysis Six Months Later?

An AI-generated strategy analysis can be reproduced six months later only if the record was kept on the day: the exact deck version presented, the frozen set of documents and sources it drew on, the brief that was asked, a claim-to-source map, and a named sign-off record. A later re-run gives a new answer, not the one the board saw.

Can We Reproduce an AI Strategy Analysis Six Months Later?

Image: Decisity

Key Takeaways

  • Reproducibility belongs to the record kept around the tool, not the platform: a re-run on today's documents gives today's answer.
  • A reproducibility pack has seven parts: frozen deck, frozen source set, brief, claim-to-source map, sign-off record, model note, exclusions.
  • UK companies must keep directors' meeting minutes for at least ten years under section 248 of the Companies Act 2006.
  • EU AI Act Article 12 requires event logging only for high-risk AI systems; it does not cover internal strategy decks.
  • In Nature's 2016 survey of researchers, more than 70% had tried and failed to reproduce another scientist's experiment.

Why reproducibility is a property of the record, not the tool

Whether an AI-assisted analysis can be reconstructed six months later depends on what was retained on the day of sign-off. This guide sets out the seven artefacts that make up a reproducibility pack, a fifteen-minute cold reconstruction test, and the vendor questions to ask before you buy.

Yes: an AI-assisted strategy analysis can be reproduced six months later, but only if the record was retained on the day the decision was made. Reproducibility does not sit inside the platform. It sits in the surrounding record: the documents the platform saw, the brief it was given, the version of the analysis that was presented, the sources behind each figure, and the names of the people who checked and approved it. If that record was not kept, no amount of platform capability will reconstruct it later. A board decision is only as defensible as the record retained on the day.

The question usually arrives in a specific form. A board member, an audit committee, an investment committee or a limited partner asks: what exactly did the platform see, what did it produce, and who signed it off? Each part of that question maps to a distinct artefact. What the platform saw is the source set. What it produced is a specific, versioned deck. Who signed it off is a named review record. An organisation that cannot answer all three from retained documents is relying on memory, staff turnover and a vendor's login state, which is not a defensible position.

Reproducibility failure is not a niche risk. In science, where reproduction is a founding norm, Nature's 2016 survey of 1,576 researchers reported that more than 70% had tried and failed to reproduce another scientist's experiments, and more than half had failed to reproduce their own.[1] If trained researchers working with methods sections and peer review struggle to reproduce work months later, a strategy team should not assume its own record will survive unaided. The governance equivalent is a written policy that states which document is the official record of a decision, so that the question does not have to be settled retrospectively.

What follows is the practical form of that discipline: why re-running a platform is not reproduction, the seven artefacts to retain at sign-off, a fifteen-minute test of whether the record actually works, the questions to put to a vendor, and the retention and failure modes to avoid.

Reproduction is not re-running: capture inputs, not just outputs

A common and costly assumption is that reproduction is guaranteed by the platform: if the analysis is ever questioned, the team can simply log in and run it again. That assumption fails for a structural reason. A re-run works on today's documents, and today is not the day the board decided. Web sources change or disappear. Internal documents are superseded by later versions, reorganised into new folder structures or deleted. The underlying models and prompts behind the platform's behaviour change as vendors ship updates. A re-run therefore produces today's answer to today's inputs, not the answer the board saw. It is a new analysis, not a reproduction of the old one.

This is why the record must freeze the inputs, not just the final deck. The deck PDF tells a reviewer what the conclusion was. It does not tell them which version of a market report the revenue figure came from, which draft of the internal cost model was ingested, or what the platform was asked to do. Reproduction requires the ability to walk from the recommendation back to the exact sources as they existed on the day, and that walk is only possible if the inputs were captured at the time.

It is worth being precise about what regulation does and does not require here, because the EU AI Act is often invoked in these conversations. Article 12 of the regulation requires high-risk AI systems to technically allow for the automatic recording of events (logs) over the lifetime of the system, and that provision comes into force on 2 December 2027 for high-risk systems under Annex III.[3][2] An internal strategy deck produced with an AI platform does not fall within that scope. The logging duties bind providers and deployers of systems classified as high-risk under the Act's annexes; they are not a compliance regime for board analytics. The practical implication is plain: no regulator keeps this record for you. The organisation must keep its own.

  • Capture the documents as ingested, with dates, because a later retrieval cannot prove what the earlier one contained.
  • Capture the brief, because the analysis can only be judged against the question it was given.
  • Capture the presented version of the deck, because the output alone cannot support a walk-back from a figure to its source.

The reproducibility pack: seven artefacts to retain at sign-off

The reproducibility pack is the concrete answer to the buyer question: a defined set of artefacts, assembled and frozen at the moment of sign-off, that allows someone outside the engagement to reconstruct what the platform saw, what it produced and who approved it. Each artefact below earns its place by what fails later without it. The quality bar for the pack as a whole comes from the records management standard ISO 15489, under which a record remains authoritative through four characteristics: authenticity, reliability, integrity and useability.[5][4]

The exact deck version presented

Retain the precise version of the deck that was put in front of the board or committee, frozen so it cannot drift. Version-locking by hash or an equivalent immutable identifier is preferable to a filename, because filenames are edited and copies circulate. Without this, a later challenge begins with an unresolvable argument about which numbers were actually on the table.

The frozen source set

Retain every document the platform ingested and every external source cited, each with a retrieval date. Without the date, a citation degrades into a dead end precisely when it is needed, because the version consulted on the day can no longer be distinguished from whatever sits at that location later. Without the frozen source set, the analysis becomes unverifiable in principle, not merely inconvenient to check.

The scope and brief record

Record the question actually asked of the platform and the constraints given: the market, the time horizon, the assumptions to hold fixed, the options in and out of scope. The same documents answer different questions differently, so an analysis cannot be judged without knowing what it was asked to do. Without the brief, a reviewer cannot distinguish an analytical error from a differently framed question.

The claim-to-source map

For each load-bearing figure, record the source and the page, table or cell it came from. This is the artefact that turns a citation into evidence: it lets a reviewer go from the number on the slide to the exact location in the source as retrieved. Without it, verifying a single figure means redoing the analysis, which is the one thing a challenged decision cannot wait for.

The review and sign-off record

Name who verified which claims, how they verified them, and who approved the recommendation. Review that cannot be attributed to a person is indistinguishable from no review. Without this record, the organisation cannot show that any human judgement stood between the platform's output and the board's decision.

The platform and model version note

Record which platform and model versions produced the analysis, and the limitations known at the time. The practice is the same one that documentation standards for released machine learning models already follow: state the intended use, the evaluation and the known limits. Without a version note, a later reviewer cannot tell whether the analysis predated a known weakness or a material change in the tooling.

What was excluded or could not be verified

Record the known gaps: sources that could not be confirmed, data that was unavailable, lines of analysis that were considered and dropped. This is the artefact people most often omit, because it feels like admitting weakness. In fact it is the opposite: a decision record that acknowledges its limits is more credible under challenge than one that claims completeness it cannot demonstrate.

ArtefactWhat it capturesWhat fails later without it
Deck versionThe exact version presented, frozen or hashedDisputes over which numbers were on the table
Frozen source setEvery ingested and cited document, with retrieval datesCitations become unverifiable dead ends
Scope and briefThe question asked and the constraints givenNo way to judge the analysis against its question
Claim-to-source mapSource, page or cell behind each load-bearing figureVerifying one figure means redoing the analysis
Review and sign-offWho verified what, how, and who approvedNo evidence of human judgement in the decision
Version notePlatform and model versions, known limitationsCannot tell if the analysis predated a known weakness
ExclusionsWhat was excluded or could not be verifiedThe record overclaims completeness it cannot show

The cold reconstruction test: a fifteen-minute protocol

A pack that has never been tested is a hypothesis, not a control. The cold reconstruction test is a short, repeatable protocol that an organisation can run before buying a platform, using a pilot engagement's material, and again after any significant meeting. Its purpose is to find out whether the record works for someone with no context, which is exactly the position of a reviewer, auditor or incoming director six months later.

  1. Hand the reproducibility pack to a colleague who was not involved in the engagement.
  2. Give them fifteen minutes and no access to the platform. The pack must stand alone.
  3. Ask them three questions, drawn from the pack alone: what was the recommendation and its three load-bearing figures; where did each figure come from and what date does that source carry; who verified the figures and who signed off the recommendation.
  4. Record what they could and could not answer, and how long each answer took to find.

A pass is unambiguous: all three questions answered from the pack alone, within the fifteen minutes, with no recourse to the people who ran the engagement or to the platform itself. That standard is not arbitrary. It is the useability characteristic of ISO 15489 applied in practice, since the standard's guidance is that records be managed so that they remain authentic, reliable, complete, unaltered and usable: a record that cannot be located, retrieved and interpreted by someone outside the engagement is not serving as a record, however carefully it was stored.[6]

The failures are as diagnostic as the pass, because each one maps to a missing artefact. Dead source links point to a source set that was never frozen with retrieval dates. An unclear deck version points to a missing version lock. A reviewer who cannot be named points to a missing sign-off record. An analysis that cannot be judged against its question points to a brief that was never written down. Running the test before purchase tells you whether a vendor's exports can support a pack at all; running it after the meeting tells you whether your own discipline held.

Test outcomeUnderlying gapMissing artefact
Source links deadSources not frozen with retrieval datesFrozen source set
Deck version unclearNo version lock on the presented deckFrozen deck version
Reviewer unnamedNo attribution of verificationReview and sign-off record
Analysis cannot be judged against its questionBrief never recordedScope and brief record

What to ask a vendor before you buy

Vendor diligence on reproducibility comes down to export capability, and it is worth asking the questions in a form that produces checkable answers rather than assurances. The core questions are whether the platform exports the source set, the version history and the citation map; in what format those exports come; and whether the exports survive termination of the contract. The last question is the one most often left unasked and most expensive to have skipped: an archive that dies with the login is not an archive, because the moment the subscription lapses or the vendor relationship ends, the record the organisation is relying on may cease to exist in any retrievable form.

  • Can the platform export the complete source set, including documents ingested and external sources cited, with retrieval dates?
  • Can it export the version history of a deck, so that the version presented can be distinguished from later edits?
  • Can it export the citation or claim-to-source map, linking each figure to its source and location?
  • In what format are these exports delivered, and can they be read without the vendor's software?
  • Do the exports remain accessible after termination of the contract, and what does the agreement say happens to the data?

The documentation precedent here is one the software industry already knows. Model cards, proposed by Mitchell and colleagues, are short documents that accompany trained models with disclosure of intended use, evaluation conditions and limitations.[7] A strategy platform that exports its sources, versions and citations is offering the same kind of transparency one level up, at the level of the analysis rather than the model.

One caution keeps this section honest: export capability is a commercial question to ask, not a compliance guarantee to assume. No regulatory requirement obliges a vendor to preserve your record, so the contractual and practical questions above are the only protection.

The practical test ties this section back to the protocol above: run a pilot engagement, request the exports as part of the pilot, and put the exported material through the cold reconstruction test before signing. A vendor demonstration shows what the platform can produce; the test shows whether what it produces can be reconstructed by a stranger with fifteen minutes and no login.

Retention and the three failures to avoid

Keep the pack for as long as the decision can be revisited, and align that period with the organisation's existing document retention policy for board papers rather than inventing a new schedule for AI-assisted work. The reference point most organisations already have is statutory: under section 248 of the Companies Act 2006, every UK company must cause minutes of all proceedings at meetings of its directors to be recorded, and those records must be kept for at least ten years from the date of the meeting.[8] A strategy analysis that fed a board decision has a natural home alongside those papers, under the same retention discipline. The related governance task is to state in a written policy which document counts as the official record of the decision. No specific retention period is proposed here beyond that alignment, because the right period depends on the organisation's own policy and the jurisdictions it answers to.

Three failures account for most of the reproducibility gaps this article is meant to prevent, and each is avoidable at negligible cost if caught early.

  • Retaining only the PDF of the deck. The PDF captures the output and nothing else: no source set, no brief, no claim-to-source map, no sign-off record. It is the artefact most likely to exist and the least able to answer the buyer question on its own.
  • Relying on the platform's login as the archive. A login is access to a live system, not a record of a past one. Sources change, versions are superseded, contracts end, and when they do, the login answers none of the questions it was assumed to answer.
  • Reconstructing the record after a challenge has been raised. A record assembled under pressure reads as after-the-fact justification rather than evidence, and no one can verify that the reconstructed source set matches what the platform actually saw. The record has value only if it was kept before it was needed.

Where Decisity fits

Decisity is an AI-native strategy consulting platform. Its decks carry every claim, number and recommendation clickable to its source, and its source set and version history can be exported. In the terms of this article, the platform generates much of the reproducibility pack as a by-product of the work: the claim-to-source map exists because every claim is linked to its origin, and the frozen source set and version history exist because they can be exported at sign-off. The pack is therefore a short assembly task at sign-off rather than an excavation six months later.

The platform does not remove the organisation's duty to keep the record. The brief still has to be written down, the reviewers still have to be named, the exclusions still have to be stated, and the pack still has to be retained under the organisation's own policy. What the platform removes is the manual labour of assembling the parts it touches, and the cold reconstruction test still applies to the result. A strategy analysis deserves the same standard of self-description that documented machine learning models already carry, and the organisation that demands it, from any vendor, is the one that can answer the board's question six months later.

Related reading:

Sources

Frequently Asked Questions

DECISITY

AI Summary

Ask an AI assistant to summarise Decisity.