Confidential Documents and AI Strategy Tools: What to Verify

Before uploading confidential documents to an AI strategy tool, ask for documents rather than assurances. Request the data processing agreement, the named sub-processor list, the stated data residency region, the written retention and deletion policy, and the security measures document. A vendor who can answer only in conversation has already told you something.

Confidential Documents and AI Strategy Tools: What to Verify

Image: Decisity

Key Takeaways

  • Replace 'is it secure?' with document requests: each concern maps to an artifact a serious vendor should already publish.
  • GDPR Article 28(3) requires a written processing contract covering instructions, confidentiality, security, deletion and audit rights.
  • A named, published sub-processor list beats a verbal description: GDPR Article 28(2) requires the controller's written authorisation.
  • Under GDPR Article 28(3)(g) the processor must delete or return your data, at your choice, when the services end.
  • European buyers should verify a stated data residency region; transfers outside the EEA need a Chapter V safeguard such as the 2021 SCCs.

Why 'is it secure?' invites the wrong answer

Before uploading a board pack or diligence material to an AI strategy platform, replace the question 'is it secure?' with a document-request checklist: a data processing agreement, a named sub-processor list, a stated data residency region, a written retention and deletion policy, and a security measures document.

A board pack is about to leave the building. So is a management presentation, a set of internal financials, or a diligence file that only four people were meant to see. The question that naturally precedes the upload is some version of: is it secure? The problem with that question is that it is built to be answered badly. It asks the vendor for a feeling, and vendors are practised at giving feelings. Reassuring is not the same as verified, and a confident conversation about encryption tells you nothing about what happens to your documents after they arrive.

The UK Information Commissioner's Office makes the underlying point in the contracts and third parties section of its data protection audit framework for artificial intelligence: without appropriate contracts in place, breaches of controller and processor requirements cannot be assessed or attributed, and with only verbal agreements there is a lack of recourse if there is a breach of UK GDPR requirements.[1] A verbal assurance has no failure mode you can point to. A document does.

So this article takes a different approach. Instead of asking whether a vendor is secure, it converts each buyer concern into a specific artifact: a document that either exists, in a form you can read, or does not. For every concern that matters when confidential material is at stake, there is a named artifact a serious vendor should already publish. The sections that follow cover five of them:

  • The data processing agreement and the contract's data-use clauses, which answer the question of whether the vendor merely processes your documents or trains on them.
  • A stated data residency region, which answers where the data physically sits and what happens if it moves.
  • A named, published sub-processor list, which answers who else touches the material.
  • A written retention and deletion policy, which answers what happens to the data during the engagement and after it ends.
  • A security measures document, which answers who inside the vendor can see your material and how that access is controlled.

Two caveats before the checklist. First, nothing here is legal advice: the regulatory points below are tied to named, in-force instruments, but counsel should review the actual contract terms before confidential material is uploaded. Second, the checklist is deliberately vendor-neutral. It works for any AI strategy platform you are evaluating, including the one this article is published by.

Training on your data versus processing it

The single most consequential distinction in an AI vendor contract is between processing your documents to deliver your analysis and retaining rights to train or fine-tune models on them. Processing means the vendor reads your board pack, cross-references it, produces the deliverable you asked for, and uses the material for nothing else. Training means the vendor may incorporate what it learned from your documents into models that outlive the engagement and serve other customers. These are different relationships with your confidential information, and contracts do not always make the difference obvious.

Morgan Lewis, writing on vendors increasingly seeking express rights to use customer data to train their AI models, notes that in practice the critical issue is often not who owns the underlying data but what rights the vendor receives to use, retain and derive value from that data over time, and that even if a company retains ownership of its raw data, broad training rights may permit a vendor to create models that incorporate learnings derived from that data indefinitely.[2] Ownership of raw data does not equal control of what a vendor derives from it. The same analysis describes the potential exposure of confidential information and trade secrets as one of the most significant concerns, and lists the questions to consider: what categories of data will be used for training, whether confidential or trade secret protections are maintained, and whether training can be restricted to specific datasets.[2]

The artifact to request here is the contract itself, read in two places. First, the data-use and training clauses in the main agreement: do they grant the vendor rights to improve services, develop products or create aggregated datasets, and can training be excluded or limited to an opt-in? Second, the data processing agreement, which governs how personal data within your documents is handled and should state that the processor acts only on your documented instructions. Morgan Lewis also notes that the legal terms are only part of the analysis and that technical safeguards are equally important, listing what organisations should understand: whether data is segregated between customers, whether training occurs on shared or dedicated models, whether the provider uses third-party foundational models, and whether customer data can be excluded from future training cycles.[2] A vendor that states plainly, in writing, that it never uses client data to train AI models has answered the question. A vendor that answers with a paragraph about anonymisation and aggregation has told you to keep reading.

Where the data physically sits, and why it matters in Europe

Data residency is a fact, not a phrase. 'Bank-grade security' and 'enterprise-ready' say nothing about which country's servers hold your documents, which jurisdictions can reach them, or what happens when a support engineer in another region opens a session. For a European buyer, the location question has a specific legal edge, because the GDPR, Regulation (EU) 2016/679, restricts transfers of personal data to third countries unless the safeguards of its Chapter V are met. If your documents contain personal data, and a board pack or an internal financial file almost always does, where the vendor hosts and processes it determines whether a transfer restriction applies at all.

The artifacts to request are correspondingly concrete. First, a stated data residency region in the contract or the vendor's documentation: a named region, in writing, not a marketing claim. Second, if any processing or access outside the EEA occurs, the transfer mechanism relied on. The standard instrument is Commission Implementing Decision (EU) 2021/914 of 4 June 2021 on standard contractual clauses for the transfer of personal data to third countries pursuant to Regulation (EU) 2016/679, which remains in force.[3] If the vendor's answer to 'where does the data sit?' is a conversation about global infrastructure rather than a named region in a document, treat that as an answer too.

One nuance worth knowing before you read the vendor's documentation. The European Data Protection Board's Recommendations 01/2020, on measures that supplement transfer tools, adopted in final form on 18 June 2021, set out a step-by-step roadmap for data exporters that begins with knowing your transfers, and state that remote access from a third country, for example in support situations, and storage in a cloud situated outside the EEA are also considered to be a transfer.[4] A vendor can host in the EU and still create a transfer if staff or sub-processors outside the EEA can reach the data. Ask about access, not only about storage.

Sub-processors: ask for a named, published list

Behind every AI strategy platform sits a supply chain: hosting, model providers, support tooling, analytics services. Each of these is another party that may touch your documents, and each is a point where your confidentiality depends on someone you did not contract with. The question 'who else can see our data?' is therefore not paranoid; it is the supply-chain question, and it has a specific legal shape under the GDPR.

Article 28(2) of the GDPR provides that the processor shall not engage another processor without prior specific or general written authorisation of the controller, and that in the case of general written authorisation the processor shall inform the controller of any intended changes concerning the addition or replacement of other processors, thereby giving the controller the opportunity to object to such changes.[5] Article 28(4) requires the same data protection obligations to be imposed on any other processor engaged for carrying out specific processing activities, and provides that where that other processor fails to fulfil its data protection obligations, the initial processor shall remain fully liable to the controller for the performance of that other processor's obligations.[5] In plain terms: the vendor cannot outsource your documents without your authorisation, must tell you when the chain changes, and remains on the hook for what its sub-processors do.

The artifact to request is a named sub-processor list that is published and dated, together with the notice-and-objection mechanism that governs changes to it. 'Published and dated' matters more than it sounds. A list described in conversation cannot be checked later, cannot be compared against what the vendor said last quarter, and gives you nothing to object to. A published list with a change-notice period lets you see exactly who is in the chain today and gives you a defined window to object when it changes. A checkable answer looks like this: the data processing agreement commits the vendor to notify subscribers of new sub-processors by updating the published list, sets a stated objection window in writing, and names the region in which processing may occur.

Retention and deletion, including at the end of the contract

Most buyers think about what happens to their documents when they are uploaded. Fewer think about what happens when the engagement ends, and that second moment is where confidentiality is quietly won or lost. A document that sits in a vendor's backups three years after the project closed is still a document that has left your building. The evaluation question is not only 'is the data safe now?' but 'who decides when it stops existing, and can I verify that it did?'

The artifact to request is a written retention and deletion policy stating retention periods, how backups are handled, and deletion timelines. Under the GDPR, this is not an optional courtesy: Article 28(3)(g) requires that the processor, at the choice of the controller, deletes or returns all the personal data to the controller after the end of the provision of services relating to processing, and deletes existing copies unless Union or Member State law requires storage of the personal data.[5] The end-of-life terms in the contract should align with the published policy: if the contract says deletion on termination and the policy says archival for an unspecified period, you have found a gap worth resolving before signature, not after.

Two details deserve specific attention when you read the policy. First, derived artefacts. Your raw documents are not the only thing a platform creates: embeddings, analysis outputs, logs and exports all derive from your material, and a deletion policy that covers only source files leaves the derivatives behind. Ask what the policy says about each. Second, the mechanics of the end of the contract. A well-drafted data processing agreement provides that, before the agreement expires, the processor shall at the choice and instruction of the customer securely delete or return all personal data, and that the customer may request written notice of the measures taken on completion of processing. A written confirmation of deletion is the difference between believing the data is gone and knowing someone committed in writing to tell you so.

Access control inside the vendor

Encryption protects your documents from outsiders. Access control protects them from insiders, and for confidential material the insider question is often the sharper one: which employees of the vendor can open a board pack, under what conditions, and with what record of having done so? A platform can be well defended at the perimeter and still allow broad internal access by default. The artifact that answers this is a security measures document describing role-based access, least-privilege permissions and audit logging in operational detail.

The GDPR gives this document a legal backbone. Article 28(3)(b) requires that the processing contract stipulate that the processor ensures that persons authorised to process the personal data have committed themselves to confidentiality or are under an appropriate statutory obligation of confidentiality.[5] Article 28(3)(h) requires that the processor makes available to the controller all information necessary to demonstrate compliance with the obligations laid down in that Article and allows for and contributes to audits, including inspections, conducted by the controller or another auditor mandated by the controller.[5] Read together, these provisions mean you are entitled to see how the vendor restricts and records internal access, not merely to hear that access is restricted.

The practical advice is to read the security measures document alongside the audit-rights clause in the contract, rather than accepting a summary of certifications. A certification claim is a claim about a process that happened at a point in time; the security document describes the controls that operate on your data every day. A well-prepared vendor publishes this material, for example on an integrations and security page that describes role-based access control with least-privilege permissions, exportable audit logs, and configurable retention policies per engagement. And where a vendor does claim a certification, verify the claim against the vendor's own published documentation rather than assuming it: ask what the certification covers, which entity holds it, and whether the scope includes the platform you are actually being sold. A serious vendor will have already published the answers.

When the answer is verbal, and what a prepared vendor looks like

Sooner or later in an evaluation, a vendor will answer one of these questions in conversation only. The response is not to argue; it is to request the answer in writing before signature. The ICO's audit framework for AI systems makes the reason explicit: with only verbal agreements there is a lack of recourse if there is a breach of UK GDPR requirements, and without appropriate contracts in place breaches of controller and processor requirements cannot be assessed or attributed.[1] A vendor who cannot produce the document has answered the question anyway. The answer is that the artifact does not exist, or exists in a form the vendor prefers you not read. Either way, you have learned something usable before the upload, not after.

The full checklist, in the order a serious evaluation should run:

  1. Data processing agreement and data-use clauses: does the vendor train on your data, and can training be excluded?
  2. Stated data residency region: a named region in the contract or documentation, plus the transfer mechanism if any processing or access occurs outside the EEA.
  3. Named, published, dated sub-processor list: who is in the chain, and what notice-and-objection mechanism governs changes.
  4. Written retention and deletion policy: retention periods, backup handling, deletion timelines, treatment of derived artefacts, and end-of-contract return or deletion with written confirmation.
  5. Security measures document: role-based access, least-privilege permissions and audit logging, read alongside the contract's audit-rights clause.

This is the standard Decisity was built against. As an AI-native strategy platform, not a consultancy and not a security vendor, it expects to be asked these questions and answers them with published documents rather than reassurance: a published data processing agreement with EU/EEA processing commitments, a sub-processor list with a defined objection window, and a security page that states its access, logging and retention controls in operational detail. For European buyers handling confidential strategy material, that posture extends to European data residency as a stated design principle rather than a sales talking point, including for Mittelstand and mid-sized companies whose board documents demand it. The same standard applies to every vendor you evaluate, including this one: read the documents, and where a claim matters, verify it against what the vendor has actually published.

One closing point belongs to your counsel, not to this article: have them review the contract terms, including the data processing agreement, the audit-rights clause and the liability provisions, before any confidential material is uploaded. The checklist above tells you which documents to put in front of them. Their signature tells you the terms hold.

Sources

Frequently Asked Questions

DECISITY

AI Summary

Ask an AI assistant to summarise Decisity.