Guide
Will an AI Tool Train on Your Business Data? What Small Businesses Should Check Before Signing Up
Understand AI training, retention, human review, integrations, and vendor data-use terms before sharing business or customer information.
An AI vendor may tell you, “We don’t train on your data.”
That is useful information. But it does not answer every question a small business should ask before uploading customer records, internal documents, emails, contracts, CRM data, or other confidential information.
A tool could exclude your content from model training while still retaining it for a period of time. Certain information might be accessible for support, safety, abuse prevention, or other defined purposes. An integration might gain access to your email, drive, CRM, or other systems. A vendor may also use subprocessors or separate consumer and business accounts under different terms.
So the practical question is broader:
What is this exact product, on this exact plan, under these exact settings and terms, permitted to do with our data?
The FTC has specifically warned AI providers that they must honor commitments about customer data, including promises related to using data to train or update models. It has also said that important omissions about how customer data is collected or used can matter.
For a small business, that means a marketing sentence should be the beginning of the review—not the end.
1. The Question Is Bigger Than Training
“Will you train on our data?” is an important question, but several different data practices can sit behind an AI service.
Model training generally concerns whether submitted or derived information may be used to develop, fine-tune, evaluate, or update models. A “no training” commitment does not by itself tell you how long data is kept or who might access it.
Product or service improvement is broader. Depending on the vendor, it could refer to analytics, debugging, abuse prevention, quality evaluation, feature development, safety work, or model improvement. If a policy simply says information may be used to “improve services,” find out what that means for your product and plan.
Data retention concerns how long prompts, uploaded files, outputs, logs, metadata, backups, and related information remain in vendor systems. Retention and training are separate questions.
Human review concerns whether authorized personnel or reviewers may access customer content, and under what circumstances. Support, security investigations, feedback, abuse prevention, and incident response can have different rules.
Subprocessors and service providers are third parties that process information for the vendor. They can include hosting, support, analytics, identity, infrastructure, or model providers.
Connected-app data is information an AI tool gains access to after you connect systems such as email, cloud storage, calendars, databases, or a CRM. NIST’s Generative AI Profile specifically notes that third-party generative-AI integrations can increase privacy and information-security risk, making integrations worth evaluating as distinct data-access decisions.
The key point:
“We do not train on your data” does not automatically mean “we do not retain it,” “no human can access it,” or “connected systems expose no additional data.”
2. Nine Things to Check Before You Share Business Data
1. Model training
Ask whether inputs, outputs, uploaded files, retrieved connected-app data, and feedback may be used to train, fine-tune, evaluate, or update models.
Also determine whether the rule is:
- disabled by default,
- enabled by default with an opt-out,
- available only through opt-in,
- or dependent on your plan.
2. Product or service improvement
Look for phrases such as:
- improve our services
- improve our products
- quality improvement
- service development
- research
- evaluation
Do not assume these terms include model training. But do not assume they exclude it either.
If the wording is broad, ask for clarification.
3. Retention
Find out what the provider retains and for how long.
Check separately for:
- prompts
- files
- outputs
- metadata
- logs
- backups
- support records
- abuse-prevention records
A deletion statement may not apply identically to every category.
4. Human review
Ask whether employees, contractors, support staff, safety teams, or other authorized people can review customer content.
The useful questions are:
Who can review it? Why? Under what controls? For how long?
5. Subprocessors
Look for an official subprocessor or service-provider list.
If another AI company, hosting provider, analytics provider, or other processor receives customer content, understand its role and the restrictions that apply.
6. Connected-app access
Treat every integration as a separate permission decision.
For example, connecting an AI assistant to a CRM could expose substantially more information than manually pasting one customer record into a prompt.
Check:
- what it can read
- what it can write
- which folders or objects it can access
- whether access continues after the immediate task
- whether retrieved information is retained
- how access can be revoked
FTC small-business guidance recommends limiting access to sensitive assets to people and parties that need that access, and evaluating third-party risks before formal relationships are established.
7. Consumer versus business plan
Do not assume one vendor has one universal data policy.
Consumer, free, business, enterprise, API, and other offerings may operate under different terms or controls.
The plan distinction is not theoretical: current official documentation from OpenAI, Google, and Anthropic shows materially different handling depending on product and account context.
8. Default versus opt-in or opt-out
A policy stating that you can opt out is different from a policy stating that data sharing is off by default.
Check the actual account settings rather than relying only on a policy summary.
9. Contract terms versus marketing claims
When an important promise affects your decision, identify where it exists in writing.
Depending on the vendor, that may include:
- product terms
- business terms
- privacy documentation
- data-processing terms
- security addenda
- an order form
- incorporated policies
FTC guidance emphasizes evaluating suppliers and third parties, and its AI guidance notes that commitments can arise through product materials and terms—not merely through one type of document.
3. A 10-Step Small-Business Verification Process
Step 1: Identify the exact product and plan
Write down what your business will actually use:
- product name
- plan
- add-ons
- API
- browser extension
- mobile app
- integrations
Avoid evaluating “Vendor X” as one thing when multiple products may operate differently.
Step 2: Classify the data you plan to share
Before reviewing policy language, identify what would actually enter the system.
Examples include:
- internal documents
- customer contact information
- CRM records
- invoices
- HR information
- source code
- contracts
- emails
- support tickets
- trade secrets
- health or financial information
The NIST Generative AI Profile treats data privacy as a distinct AI risk involving issues such as unauthorized use, disclosure, and leakage of sensitive or personally identifiable information.
Step 3: Check official documentation
Start with primary vendor sources rather than reviews or marketing summaries.
Search official documents for terms such as:
train improve content inputs outputs prompts retain delete review human subprocessor API business enterprise feedback
Do not stop at a page that simply describes a product as “secure,” “private,” or “enterprise-grade.”
Step 4: Separate the six core questions
Require distinct answers to:
- Is our content used for model training?
- Can it be used for product improvement, evaluation, diagnostics, analytics, safety, or related purposes?
- How long are content, logs, metadata, and backups retained?
- Can humans access customer content, and under what circumstances?
- Which subprocessors or model providers receive the information?
- What does each connected integration retrieve, write, store, or retain?
Step 5: Inspect the actual account settings
Log into the account you intend to use and check:
- model-improvement controls
- training opt-ins or opt-outs
- feedback controls
- retention options
- workspace policies
- admin settings
- connector permissions
Documentation describes possible behavior. Account settings tell you what is currently configured.
Step 6: Review integrations separately
A permission request to access Gmail, Google Drive, Microsoft 365, Salesforce, a CRM, or another system deserves its own review.
Ask whether the requested permissions match the business purpose.
NIST warns that third-party generative-AI integrations can introduce additional data privacy and information-security considerations.
Step 7: Review the terms governing your account
Depending on the vendor, review relevant:
- service terms
- business terms
- privacy documentation
- data-processing addendum
- security terms
- subprocessor list
- order form
Not every small business can negotiate vendor contracts, but you can still identify what written terms govern your account.
Step 8: Resolve contradictions in writing
Suppose a product page says:
“We don't train on your business data.”
But another document says customer information can be used to “improve our services.”
That is not automatically a contradiction.
Ask whether “improvement” includes:
- model training
- model evaluation
- human review
- feedback datasets
- product analytics
- third-party AI providers
FTC guidance makes clear that material omissions about data use can matter alongside explicit representations.
Step 9: Document what you found
Save:
- plan name
- policy URLs
- screenshots
- settings
- integration scopes
- vendor replies
- contractual documents
- date reviewed
Policies and product settings can change.
Step 10: Begin with lower-risk data
If the documentation and settings look acceptable, consider beginning with public, synthetic, or otherwise lower-risk data.
Avoid making your first experiment a full CRM, mailbox, or confidential document repository.
This does not replace vendor due diligence. It limits unnecessary exposure while you validate the workflow.
4. Questions to Send an AI Vendor
You can adapt this short message:
Subject: Data-use questions for [product and plan]
We are considering [product and plan] for [business use case].
Before connecting or uploading business information, please confirm:
- Are our inputs, outputs, uploaded files, or connected-app data used to train, fine-tune, or update any model?
- Can that information be used for product improvement, evaluation, analytics, safety, or debugging?
- What are your standard retention periods, and what deletion, backup, legal, or abuse-prevention exceptions apply?
- Under what circumstances may employees, contractors, or other reviewers access customer content?
- Which subprocessors or third-party model providers may process our content?
- What information can the proposed integration read, write, retrieve, or retain?
- Which terms or documentation govern these commitments for our specific plan?
Please provide links to the applicable documentation where possible.
A useful answer does not have to say “no” to every question.
It has to be specific enough for you to understand what will happen.
5. How to Read the Vendor’s Answer
| Topic | Clear answer | Needs follow-up | Red flag |
|---|---|---|---|
| Model training | Identifies the plan and clearly explains training rules, defaults, and settings | “We respect your privacy” without answering training | Broad reuse rights while refusing to clarify whether training is included |
| Product improvement | Defines what “improvement” includes | Uses the term without explaining it | Very broad rights to reuse business content for unspecified purposes |
| Retention | Gives retention periods, deletion process, exceptions, and available controls | “As long as necessary” without useful detail | Cannot meaningfully explain retention or deletion |
| Human review | States who can review content, for what reasons, and under what controls | “Authorized staff may access data” | Broad discretionary review with little explanation |
| Subprocessors | Provides a current list or meaningful description of processing roles | Names providers but not their role | Cannot explain meaningful third-party processing |
| Connected apps | Explains permissions, scope, revocation, and retention | Says only that the connection is “secure” | Requests broad access unrelated to the proposed task |
| Plan context | Names the exact plan or account type covered | Uses general terms such as “enterprise-grade” | Important assurance cannot be tied to the actual plan |
| Contractual commitment | Identifies governing written terms | Refers vaguely to “our policy” | Critical commitments exist only in informal claims |
An unclear answer is not proof that a vendor is unsafe.
It means you do not yet have enough information to rely on the claim.
6. Consumer and Business Plans Can Differ
These examples illustrate why plan context matters. They are based on current official documentation and are not Vettlume product tests, endorsements, or recommendations.
OpenAI
OpenAI currently offers business products including ChatGPT Business, ChatGPT Enterprise, and the API Platform.
Its current official documentation states that, by default, inputs and outputs from products for business users are not used to train its models. API organizations may explicitly opt in to share certain data for model improvement or training. Individual services such as consumer ChatGPT operate under separate controls and may use content for training depending on user settings.
That statement does not mean every OpenAI service has identical retention, access, or feature rules.
The correct question is still: Which product and terms govern your account?
Google Workspace with Gemini
Google’s current Workspace documentation says qualifying commercial Workspace customers receive business-oriented data protections.
For covered Workspace with Gemini use, Google states that customer content is not human reviewed or used for generative-AI model training outside the customer’s domain without permission. Users without qualifying Workspace editions may be subject to different terms.
Do not generalize this Workspace rule to consumer Gemini, personal Google accounts, or every Google AI service.
Anthropic / Claude
Anthropic’s current commercial-data documentation states that inputs and outputs from commercial products such as Claude for Work and the Anthropic API are not used to train its models by default.
Explicit feedback or an explicit opt-in may create exceptions. Consumer products have separate terms and settings.
Canva
Canva’s current Trust Center documentation states that content from its Business, Teams, Enterprise, and Education plans is not used to improve AI-powered features under the controls described there. Its current pricing page separately notes that Canva Teams is a legacy plan that is no longer available for new sign-ups or upgrades, and that Business is the team plan for new customers.
So on Canva, data-training protection is tied to the plan tier, and the plan available to new team sign-ups changed during 2026.
For a researched-only look at Canva’s AI features, plan availability, and limits, see Vettlume’s Canva AI review.
Again, the relevant lesson is not that one vendor is better than another.
It is that product, plan, settings, and governing terms matter.
7. When to Pause Before Connecting Data
Consider pausing before broad deployment when the tool will handle:
- customer personal information
- financial information
- health information
- sensitive employee records
- source code
- trade secrets
- confidential contracts
- large CRM databases
- full mailboxes
- shared-drive repositories
- information restricted by a customer agreement
At that point, the decision may involve more than ordinary product evaluation.
Your business may have privacy, contractual, security, or sector-specific obligations that this guide does not determine.
FTC guidance encourages businesses to assess third-party cybersecurity risks and understand contractual and regulatory requirements.
For regulated, contractual, or unusually sensitive uses, qualified legal, privacy, security, or compliance advice may be appropriate.
This guide is not legal advice.
8. What to Document
A simple internal record can make future reviews much easier.
| Field | Record |
|---|---|
| Product | Exact product name |
| Plan | Consumer, business, enterprise, API, etc. |
| Intended use | What the business wants the tool to do |
| Planned data | What information the tool will receive |
| Settings | Training, feedback, retention, admin, and data controls |
| Permissions | Connector scopes and read/write access |
| Privacy documentation | Official URL |
| Governing terms | Applicable agreement or terms |
| Vendor reply | Written clarification received |
| Retention | Periods and known exceptions |
| Human review | Conditions under which access may occur |
| Subprocessors | Relevant third parties |
| Reviewer | Internal person who performed the review |
| Decision date | Date approved, rejected, or paused |
| Next review | When documentation should be checked again |
A dated record is useful because vendor policies, settings, features, and product names can change.
9. What This Review Does Not Prove
Completing this process does not certify that:
- a vendor is secure
- a product is legally compliant
- a model is safe
- your business meets every privacy obligation
- a DPA eliminates all risk
- a security certification proves suitability
- the product is appropriate for regulated information
NIST describes its AI Risk Management Framework as voluntary guidance designed to help organizations incorporate trustworthiness considerations into AI design, development, use, and evaluation. It is not a government certification for an AI vendor or a mandatory Vettlume scoring system.
The purpose of this review is narrower:
understand what the vendor's documentation, account settings, permissions, and governing terms say about your data before you expose more of it.
10. What to Do Next
Data-use review is only one part of buying an AI tool.
Once you understand the vendor's training, retention, human-review, integration, and contractual terms, use Vettlume's full AI-tool evaluation checklist to examine the broader business decision, including fit, cost, implementation, output quality, and operational risk.
After that, test the tool on a limited, lower-risk workflow before broader rollout using Vettlume's AI Tool Trial Scorecard.
A practical sequence is:
Verify data use → evaluate business fit → run a limited trial → expand only when the evidence supports it.
Evidence and sources
- Artificial Intelligence Risk Management Framework (AI RMF 1.0)National Institute of Standards and Technology / official-documentation / accessed 2026-08-06
AI risk management approach organized around governing, mapping, measuring, and managing.
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNational Institute of Standards and Technology / official-documentation / accessed 2026-08-06
Generative-AI reliability, evaluation, monitoring, documentation, and appropriate human oversight.
- AI Companies: Uphold Your Privacy and Confidentiality CommitmentsFederal Trade Commission / regulatory-source / accessed 2026-08-06
Privacy and confidentiality commitments and potential liability when data practices conflict with commitments made to users and customers.
- Cybersecurity for Small BusinessFederal Trade Commission / regulatory-source / accessed 2026-08-14
Vendor-security guidance on assessing third-party risks, limiting access, documenting data-handling expectations, and verifying vendor practices.
- Enterprise privacy at OpenAIOpenAI / official-documentation / accessed 2026-08-14
Business-data defaults for ChatGPT Business, ChatGPT Enterprise, and the API Platform, including explicit opt-in exceptions.
- Gemini for Google Workspace FAQ - Business / EnterpriseGoogle Workspace Admin Help / official-documentation / accessed 2026-08-14
Qualifying Google Workspace plan protections stating that covered content is not human reviewed or used for generative-AI model training outside the customer's domain without permission.
- Is my data used for model training?Anthropic Privacy Center / official-documentation / accessed 2026-08-14
Commercial-product training defaults for Claude for Work and the Anthropic API, with explicit feedback or opt-in exceptions.
- Canva Trust CenterCanva Pty Ltd / official-documentation / accessed 2026-08-17
Canva states that Business, Teams, Enterprise, and Education content is not used to improve AI-powered features under the current controls described in its Trust Center.
- Canva PricingCanva Pty Ltd / official-pricing-page / accessed 2026-08-31
Canva's current US pricing page lists Free, Pro, Business, and Enterprise. It shows Canva Pro at US$144 per year for one person and Canva Business at US$250 per year per person, with higher AI allowances than Free. The page states that Canva Teams is unavailable for new sign-ups or upgrades and that existing subscribers remain on that legacy plan unless they switch or cancel.
- Meet Canva Business: A powerful new plan for small businesses with big ambitionsCanva Pty Ltd / official-documentation / accessed 2026-08-31
Canva Business is the current plan for new team sign-ups and upgrades, including teams of one, with no seat minimum. Canva states a US$20 per-person monthly price in its announcement. Existing Canva Teams subscribers keep their current plan and pricing, but Teams is not available for new sign-ups or upgrades.
Related content
article
How to Evaluate an AI Tool Before Paying: A Practical Small-Business Checklist
Evaluate an AI tool’s usefulness, total cost, accuracy, privacy, integrations, security, and cancellation terms before paying for a business subscription.
article
AI Tool Trial Scorecard: A 9-Point Evaluation Worksheet for Small Businesses
Use this 9-point AI tool trial scorecard to evaluate business fit, value, output quality, privacy, cost, integrations, and vendor reliability before buying.