How to Pilot an AI Strategy Platform: The Back-Test

To evaluate an AI strategy platform before committing, run a back-test: give the platform the original brief and document set from an engagement the team has already completed, then score its output against the deck the board actually saw on six fixed dimensions: problem framing, structure, source verifiability, so-what, omissions and errors.

How to Pilot an AI Strategy Platform: The Back-Test

Image: Decisity

Key Takeaways

  • A back-test reruns a completed engagement with the original brief and documents, then scores the platform output against the deck the board or investment committee actually saw.
  • Score on six fixed dimensions: problem framing, structure, source verifiability, so-what, omissions and errors.
  • ICO guidance under Article 28(3) sets eight minimum contract terms for a data processing agreement, including named sub-processor authorisation, end-of-contract deletion and audit rights.
  • The FTC's Operation AI Comply announced five enforcement actions on 25 September 2024 against unsupported AI claims, so treat vendor performance claims as unproven until back-tested.
  • Fluent, well-structured output whose sources do not hold is a fail, not a near-pass: verifiability outranks polish.

Why a demo cannot answer the procurement question

Before committing to an AI strategy platform, back-test it: rerun an engagement your team has already completed, using the original brief and document set, then score the output against the deck the board actually saw. This guide sets out the protocol, a six-dimension rubric, the vendor artifacts to request, and what a pass and a fail look like.

How do you evaluate an AI strategy platform before you commit, and what should a pilot actually test? The direct answer is this: test the platform on work your team has already completed, using the same brief and the same documents it had at the time, and score the output against the deck your board or investment committee actually saw. A pilot built on that principle tells you what no demonstration can, because you already know the right answer and where the difficulties were. Everything else in this article is a protocol for doing exactly that.

The reason the burden falls on the buyer is that a vendor demonstration is constructed to prove fluency, not judgement. The vendor chooses the sample data, frames the question and controls the narrative around the output. On curated material, a language model will nearly always look impressive, because the test conditions have removed the two things that actually break strategy work: messy, incomplete documents and a question that is harder than it first appears. Fluency under controlled conditions is a property of the technology; correctness on your problem is a property of the specific system applied to your documents, and only you can measure the second.

Regulators and standards bodies point in the same direction. In September 2024 the United States Federal Trade Commission announced Operation AI Comply, a co-ordinated sweep of five enforcement actions against operations that relied on AI hype or made performance claims without evidence, including a company that had not tested whether its AI output matched the expertise of the professionals it claimed to replace[1]. The commission's position was that there is no AI exemption from existing law on deceptive claims[1]. A buyer who accepts a vendor's performance narrative without evidence is exposed to exactly the gap the FTC went after: a claim nobody tested.

The measurement framing also has an established home in standards work. The United States National Institute of Standards and Technology released its Artificial Intelligence Risk Management Framework in January 2023 for voluntary use by organisations developing, deploying or using AI systems, and built it around four functions: govern, map, measure and manage[2]. The framework describes a flexible, structured and measurable process for addressing AI risk, and treats measurement and monitoring as preconditions for trust rather than optional extras[2]. A back-test is that measurement function applied to one concrete procurement decision.

The sections that follow give you the protocol in full: how to select and run the back-test engagement, the six-dimension rubric for scoring the output, the artifacts to demand from the vendor, who should run and sign off the pilot, and the two pilot designs that reliably produce a wrong answer. The aim throughout is that the decision to adopt rests on evidence you generated, not on a narrative the vendor supplied.

The back-test: replay a completed engagement

A back-test reruns strategy work your team has already finished and grades the platform against a known result. It borrows its logic from quantitative finance, where a strategy is tested against historical data before capital is committed. The protocol has four steps.

  1. Select one completed engagement and freeze the inputs. Retrieve the original brief exactly as it was given, and the document set the team actually had at the time. Nothing later: no final data, no board feedback, no documents produced after the recommendation was made. The test is only valid if the platform sees precisely what the team saw.
  2. Run the platform to a finished deck and record the run conditions. Note which model configuration was used, how long the run took, what prompts or scoping inputs were required, and where humans intervened. These conditions are part of the evidence: a result that only holds after heavy manual correction is a different product from one that holds on a clean run.
  3. Score the output against the final human deck, using the fixed six-dimension rubric set out in the next section. Score every dimension, including the ones the platform does well, so the profile is comparable across vendors and across time.
  4. Compare and decide. Where the platform output differs from the human deck, ask which version the board would have preferred, and where the platform found something the team missed. Those differences, not the overall polish, are the finding.

The value of the protocol lies in the ground truth it gives the evaluator: you already know which data points were traps and which slides the board contested, so every claim can be checked against a result the vendor does not control. That is what makes a fluent but hollow output impossible to pass off as analysis.

One caution on scope. A single back-test is one data point on one problem class. If the pilot budget allows, run two engagements of different types, for example a market entry screen and a portfolio review, so you can see whether performance holds across the kind of work you actually commission. The protocol is cheap enough to repeat, because the inputs already exist.

The six-dimension scoring rubric

Score the back-test output on six fixed dimensions. Fixing them in advance matters: it stops the evaluation drifting towards whatever the platform happens to do well, and it makes results comparable if you test more than one vendor. The rubric also has a regulatory echo. The EU AI Act requires that the instructions for high-risk AI systems disclose the level of accuracy, including its metrics, against which the system has been tested and validated, so that deployers can interpret output and use it appropriately[3]. A buyer applying this rubric is doing, in miniature, what that provision expects of the deployer: testing against stated metrics before use.

DimensionWhat you are testingWhat a good result looks like
1. Problem framingDid the platform define the same question the team did, or a better one?The core question is stated sharply, with scope and decision owner clear; any reframing is an improvement the team would accept, not a dodge
2. StructureIs the storyline MECE and conclusion-first, with one message per slide?The issue tree does not double-count or leave gaps; each slide title states the finding; the storyline reads as an argument from title page to recommendation
3. Source verifiabilityDo the cited sources actually state what the slides say?Open a sample of at least ten cited claims; every one traces to a source that supports the claim as written
4. So-whatDoes every exhibit carry an implication for the decision?No chart or table is presented without a stated consequence; implications match the evidence, not generic commentary
5. OmissionsWhat did the original deck contain that the output lacks?Gaps are limited to what the frozen document set could not support; nothing the board relied on is missing without a traceable reason
6. ErrorsIs anything stated that the source does not support?No unsourced number, no overreach beyond the evidence, no confident assertion where the documents are silent

Two dimensions deserve emphasis because they are where fluent systems fail. Source verifiability is the dimension that separates an analysis tool from a well-written text generator. Omissions is the dimension buyers most often forget to score: a platform can produce a correct, well-sourced deck that misses the point the board actually cared about, and only a comparison against the real deck reveals that. Score all six every time, and record the evidence for each score, because the scoring sheet is itself the artifact that sign-off rests on.

Choosing the right engagement to back-test

The engagement you choose determines how much the back-test can tell you. Three criteria govern the selection, and they pull against each other, so expect a trade-off rather than a perfect case.

  • Recent enough to reconstruct. The original brief, the document set as it stood at the time, and the final deck must all still exist and be retrievable. This mirrors the storage limitation principle in UK GDPR, Article 5(1)(e), under which personal data must not be kept in identifiable form for longer than necessary for the purposes of processing[4]. If your retention policy has already done its work, the materials are gone and the engagement cannot be back-tested, however good a candidate it otherwise is.
  • Contested enough to be informative. Choose an engagement where the board or investment committee pushed back: where a key assumption was challenged, an exhibit was queried, or the recommendation was reworked. A smooth engagement tests nothing, because there were no traps. A contested one shows how the platform handles ambiguity, weak evidence and a question that changed shape mid-stream, which is precisely where AI-generated strategy work tends to fail.
  • Not confidential beyond the pilot's data terms. The document set is uploaded to the vendor under the pilot data processing agreement, so the engagement must be one you are permitted to process on external infrastructure. Anything touching personal data, live transactions or restricted deal material is out of scope for a first pilot unless the agreement explicitly covers it.

A practical shortlist looks like this: a market or competitor assessment from the last one to two years, a make-or-buy or portfolio decision that went through at least two board iterations, or a due diligence screen where the team later learned what the answer should have been. Avoid the most recent engagement if its lessons are still politically live, and avoid the easiest one, because a back-test that asks nothing teaches nothing.

The artifact checklist to request from the vendor

Before any of your documents are shared, the vendor should be able to produce a set of artifacts that let you inspect how the platform behaves and how your data would be handled. Request them in writing at the start of the pilot, and treat refusal or delay on any of them as a finding in itself.

  • A sample deck with clickable sources. This is the single most revealing artifact, and it requires no data from you. Open the sources on a sample of claims and check whether each source states what the slide says. If the vendor's own sample deck does not survive that inspection, the pilot's verifiability dimension is already answered.
  • The data processing agreement. Under UK GDPR Article 28(3), a controller-processor contract must contain eight minimum terms: processing only on the controller's documented instructions, a duty of confidence, appropriate security measures, terms governing sub-processors, assistance with data subjects' rights, assistance with the controller's security and breach obligations, end-of-contract deletion or return of data, and audits and inspections[5]. Check the vendor's draft against that list rather than assuming it is complete.
  • The named sub-processor list. A processor may not engage a sub-processor without the controller's prior specific or general written authorisation, and remains liable for the sub-processor's compliance[5]. You need to know which model providers, hosting providers and other third parties would touch your documents, by name, before the pilot starts.
  • The data residency statement. Where your documents are stored and processed, in which jurisdictions, and under what transfer mechanisms. For a European corporate this is frequently a gating item, and it should be stated precisely enough to verify rather than as a marketing phrase.
  • The retention and deletion policy. What is kept, for how long, and what happens at the end of the contract. Article 28(3) requires that, at the controller's choice, the processor deletes or returns all personal data at the end of the contract and deletes existing copies unless storage is legally required[5]. Ask the same question for your confidential business documents, not only for personal data.
  • An export of the audit trail. The record of what the platform did: which documents were read, what was produced, what was changed and by whom. Request a sample export during the pilot and confirm it is complete enough that a reviewer could reconstruct how a conclusion was reached.

Two of these artifacts do double duty. The sample deck with clickable sources previews the platform's verifiability before you have shared anything, and the audit trail export previews whether the platform's working can be reviewed at all. A vendor who produces both readily is telling you something about how the platform was built; a vendor who cannot has answered your procurement question in the negative.

Who runs the pilot, who signs off, and what pass and fail look like

A pilot without governance produces an opinion, not a decision. Three roles need to be fixed before the back-test runs.

  • The evaluator should be someone who worked on the original engagement, because only they can supply the ground truth the scoring depends on: which assumption the board challenged, and which slide took the team three rewrites. Without that first-hand knowledge the scoring of omissions and errors collapses into guesswork, which is exactly the failure mode a back-test exists to prevent.
  • A second reviewer should score the output independently before scores are compared. This does not need to be elaborate: two evaluators, the same rubric, and a short reconciliation session where divergent scores are argued out against the evidence. It keeps the result from resting on one person's recollection.
  • Sign-off sits with the budget holder or, in an advisory firm, the managing partner. The person who owns the purchase decision should see the scoring sheet, the run conditions and the sampled source checks, and should be able to explain the result to a board or investment committee without the evaluator in the room.

Define pass and fail before the run, in writing, so the criteria cannot be renegotiated after the result is known. A pass means the output scores acceptably across all six rubric dimensions, with every sampled source verified: the framing is sound or better, the structure would survive a board read-through, the so-what is present on the exhibits that matter, and the omissions and errors are within the tolerance you set in advance. Partial results should be recorded as such, because a platform that passes five dimensions and fails one is a different procurement conversation from one that passes all six.

The most important fail case to name explicitly is the fluent failure: the output reads well, the storyline is clean, and the sources do not hold. When you open the cited documents and find that they do not state what the slides claim, that is a fail on verifiability and errors regardless of how polished the deck is, and it is disqualifying rather than a deduction. Other fail cases include an output that required substantial undocumented human correction to reach its final state, and an audit trail too thin to reconstruct how conclusions were reached.

Two pilot designs to avoid, and what a back-testable platform looks like

Most failed pilots fail at the design stage, before the platform has done anything. Two designs account for most of the waste.

  • Judging on a vendor demo. A demonstration is useful for one thing only: deciding whether the platform is worth a back-test.
  • Judging on speed alone. Time-to-deck is the easiest metric to measure and the least informative, because speed measures throughput, not whether the analysis is right. A platform that produces a fluent deck in an afternoon has solved the easy half of the problem; the half that matters is whether every claim survives contact with its source. Measure speed only after verifiability has passed, and never instead of it.

What the pilot should leave you with is a platform that can be back-tested at all. That means a system that accepts your brief and your documents, produces a board-ready deck, and exposes its evidence: every claim traceable to the source it rests on, so that the six-dimension rubric can actually be applied.

Decisity is an AI-native strategy consulting platform that turns a brief and documents into board-ready decks with every claim traceable to its source. It is designed to be evaluated with exactly this protocol: run it on a completed engagement, score the output against the deck the board saw, and inspect the sources behind every claim. If you want to see that workflow end to end, the how it works page describes the three-step process from document ingestion to a finished deck, and you can request a demo as the first step towards a back-test on your own material.

Related reading:

Sources

Frequently Asked Questions

DECISITY

AI Summary

Ask an AI assistant to summarise Decisity.