Catch Advisors
AI Strategy

AI Vendor Data Retention: 12 Questions to Ask Before Employees Upload Sensitive Data

Your employees are going to paste sensitive data into AI tools.

Some already have. Contract language, customer notes, source code, meeting transcripts, HR documents, financial forecasts. Telling people to “use good judgment” is not a control. It is a hope dressed up as policy.

Decide which data each AI service may receive, learn what happens after upload, and put the vendor’s promises into enforceable terms. Do that before rollout, not after legal discovers that a chat was retained somewhere nobody expected.

One warning up front: “We do not train on your data” is not the same as “we do not retain your data.” A vendor may keep prompts and responses for chat history, security review, support, legal holds, product feedback, or abuse detection without using them to train a foundation model. Those details affect privacy, discovery, breach exposure, and your ability to honor a deletion request.

Map the data before comparing vendors

Do not start with a vendor questionnaire. Start with the work.

Pick the actual use cases employees want. For each one, list what enters the tool, what the tool can retrieve, what it creates, and where the output goes. A meeting assistant might receive audio, names, calendar details, a transcript, and follow-up tasks. A contract assistant could receive confidential terms, personal information, and attorney comments. A coding tool may see proprietary source code, secrets accidentally left in a file, and internal system names.

Classify that data using your existing policy. If your company already labels information as public, internal, confidential, and restricted, use those labels. Do not invent a separate AI classification scheme that employees have to learn.

Then set an allowed-data rule for each use case:

  • Public data may be allowed in approved business tools.
  • Internal data may require an enterprise account and company login.
  • Confidential data may require specific contract terms, retention settings, and logging.
  • Restricted data may stay prohibited until legal, security, and the data owner approve a narrow exception.

A general AI license does not mean every employee can upload every kind of company data.

Ask these 12 questions and require proof

A useful vendor review gets past the privacy-page headline. Ask for written answers, contract references, and a live admin demonstration where a setting matters.

1. What data do you collect?

Ask about prompts, uploaded files, retrieved records, generated responses, chat titles, feedback, usage metadata, IP addresses, device information, connector data, and support tickets.

“Customer content” is often a defined term in the contract. Read the definition. If transcripts, system logs, or feedback sit outside it, they may follow different rules.

2. Which data is stored, and for how long?

Get a retention period for each data type. “For as long as necessary” does not help you configure policy or answer an auditor.

Google’s current Workspace Gemini Privacy Hub shows why product-level detail matters. It lists different retention behavior for Gemini in Workspace prompts and responses, the Gemini app, and Gemini Notebook. Administrators can set some periods, while other data follows Workspace deletion terms.

A vendor may offer several AI experiences under one brand. Do not assume they share one retention schedule.

3. Can administrators shorten retention?

Ask whether your team can set a company-wide period, whether users can override it, and which license tier includes the control. Have the vendor show the setting.

Also ask what “delete” means. Does deletion remove the item from the user interface, active systems, backups, safety systems, and subprocessors? How long does each step take?

4. What happens when a user deletes a chat?

A missing chat window does not prove the underlying record is gone. Microsoft explains that Copilot messages can be stored in hidden Exchange mailbox folders for compliance. Its Purview retention documentation also explains how retention policies, litigation holds, and eDiscovery holds can delay permanent deletion.

That may be exactly what a regulated company needs. It can also surprise a company that expected a user deletion to remove the record. Your legal and records teams should choose the outcome.

5. Is our data used to train or improve models?

Split this into several questions:

  • Is customer content used to train a shared foundation model?
  • Is it used to tune a model dedicated to our company?
  • Is feedback handled differently from normal prompts?
  • Can a user opt in without an administrator?
  • Does the answer change by product, plan, region, or model provider?

Microsoft states that prompts, responses, and Microsoft Graph data used by Microsoft 365 Copilot are not used to train its foundation models. Google states that Workspace Gemini prompts, Workspace content, webpage context, and responses are not used to train generative AI models outside the customer’s domain without permission.

Good. Still verify that your contract covers the exact service you are buying. A consumer account, developer API, embedded feature, and enterprise workspace can have different terms.

6. Who can review our content?

Ask whether vendor employees or contractors can view content for support, safety, abuse investigation, quality review, or feedback analysis. Require the vendor to explain approval, access logging, location, and confidentiality controls.

If human review cannot be disabled, decide whether the proposed data belongs in that service at all.

7. Which subprocessors and model providers receive it?

An AI application may sit on top of another model, cloud platform, vector database, logging service, or support system. Ask for the subprocessor list and change-notice process.

You also need to know whether the underlying model provider receives prompts and outputs. AWS states in its Amazon Bedrock data protection guidance that model providers do not have access to Bedrock logs, customer prompts, or completions in its deployment accounts. AWS also documents retention modes separately, including an option that can permit provider data sharing for models that require it. Configuration and model choice matter.

8. Where is the data processed and stored?

Get the regions for primary storage, backups, support access, and subprocessors. Ask whether prompts can leave your selected region when the service calls a model or web search.

“Hosted in the United States” is too broad for companies with contractual residency requirements. Ask for the architecture and the commitment that appears in your agreement.

9. How do connectors change the answer?

Connecting an AI tool to Microsoft 365, Google Workspace, Salesforce, ServiceNow, a file share, or an internal database changes the risk. The tool is no longer seeing only what a user pastes into a box.

Ask whether retrieved data is copied into chat history, cached, indexed, logged, or sent to another provider. Confirm that the service honors source permissions. Then test what happens when access is removed at the source.

If your company is evaluating tools inside an existing ecosystem, Catch partner pages for Microsoft, Google Cloud, and AWS can help frame the platform options. Existing infrastructure may simplify identity and governance. It does not remove the need to inspect the exact AI service and configuration.

10. Can we export records before leaving?

Ask what you can export, in which format, and at what cost. Include chats, files, system instructions, audit logs, evaluations, connector settings, and retention configurations.

Then ask what the vendor deletes after termination, what it must retain, and when it will certify deletion. A ninety-day offboarding promise is not useful if your renewal decision happens ten days before the contract ends.

The contract should state when the vendor notifies you of an incident, what evidence it provides, and how it supports investigation. Ask how legal holds, government requests, policy violations, and abuse reviews change normal retention.

Do not accept a generic incident-response statement if the AI service has separate infrastructure or providers. Follow the data path.

12. Which promises are contractual?

A help-center article can change. Your protection comes from the agreement, data processing addendum, service-specific terms, and documented configuration.

List the promises that matter to your use case: no shared-model training, named retention period, deletion timing, approved regions, subprocessor notice, breach notice, export rights, and support for audit evidence. Put them in the contract or attach the controlling service terms by version and date.

Build a retention scorecard that exposes tradeoffs

Use one page per service. Record the use case, allowed data classes, storage period, admin controls, training terms, human review, subprocessors, residency, deletion process, legal-hold behavior, export path, and contract location.

Score each item as verified, unclear, or unacceptable. Do not average away a deal breaker. If a tool fails a restricted-data requirement, a beautiful interface and a discount should not rescue it.

Run one test with realistic but synthetic data. Upload a file, start a chat, delete the conversation, export the available records, and inspect the admin and audit views. Ask the vendor to explain anything you cannot see.

The result may be a tiered rollout. One tool can handle general productivity work. A more controlled platform can handle approved confidential workflows. Some restricted data may remain off limits. That is a reasonable architecture. Forcing every use case into one AI product usually creates either too much risk or too many limits.

Make the policy easy enough to follow

Employees need a short answer at the moment of use. Tell them which tools are approved, which account to use, what data is prohibited, and where to request an exception. Put the rule in onboarding, the AI tool catalog, and the browser or access workflow employees already use.

Back it with technical controls where possible. Disable unapproved applications, require company identity, control connectors, apply data loss prevention rules, and review logs. Policy matters, but policy without enforcement eventually becomes a suggestion.

AI data retention is a buying decision before it is a compliance problem. Get specific about the data, verify the settings, and put the important promises in writing.

Catch Advisors helps IT leaders compare enterprise AI options without defaulting to the loudest vendor. If your team needs to sort out data handling, platform fit, contract terms, and rollout controls, schedule a vendor-neutral assessment before sensitive data starts moving.

Sources