Do not decide whether to upload a document to AI based on the brand name alone. First establish what the information contains, whether you are authorized to process it this way and which terms apply to the actual product. For a small business or nonprofit, public or fictional information is often a more manageable starting point than client files.

A free account, business subscription and enterprise contract may have different conditions. A commitment not to use data for model training does not mean that data is never retained, accessible to administrators or sent to a connected service.

Separate the types of information

These categories can overlap. A client file may contain personal information, confidential business information and a trade secret at the same time.

  • Public: published material, still subject to applicable rights. Public availability does not remove every restriction on use.
  • Internal: an unpublished procedure or memo. External processing still needs authorization.
  • Personal: information that identifies a person, either alone or alongside other details. An address, role and context may be enough.
  • Confidential: information protected by a contract, professional obligation or internal rule, such as an unannounced proposal.
  • Trade secret: valuable business information whose value depends in part on remaining secret, such as a method or formula.
  • Client information: material entrusted to you by a client. Access does not automatically give you permission to disclose it to an AI provider.

When summarizing a file, ask whether the entire document is necessary. An approved extract or fictional example may answer the question without exposing the real details.

Check the product and the data journey

For ChatGPT, Microsoft Copilot, Gemini or Claude, identify the exact version, account type and governing agreement. “Copilot” can describe different experiences; seeing the logo in a browser does not confirm the protections of a managed Microsoft 365 environment.

Review how inputs and outputs are used, retention, deletion, access, processing locations and the contractual commitments you require. Model training, conversation history and security logs are different purposes and should be assessed separately.

Start with the documentation for your offering: OpenAI business privacy, Microsoft Copilot data and privacy, Google Workspace AI privacy and Anthropic commercial terms. These do not necessarily describe every consumer account. Record when and what you checked rather than assuming the terms will remain unchanged.

Review permissions and connected services

An assistant that searches company files can make information with overly broad permissions easier to discover. Review access before enabling a connection. An assistant used by leadership does not necessarily need access to human resources files.

An extension, action or connector may send information to a different provider. Assess that recipient and its scope separately. Also review shared conversation links and who may see the output.

The ChatGPT, Copilot, Gemini and Claude comparison helps match tools to work. The data decision must then address the specific environment you intend to use.

Make the decision before uploading

For each use case, keep short answers to these questions:

  1. Is the information necessary for the task?
  2. Does the organization have the right to disclose it for this purpose?
  3. Are the tool, account and settings approved for this information category?
  4. Are the people receiving the answer allowed to see the source material?
  5. Who will review the result and handle a possible error?

If an answer is missing, pause the real-data upload and use a fictional example. Put these decisions in an internal AI use policy and practise them during team training.

Is removing names enough?

Not always. Dates, amounts, job titles and events can identify someone when combined. Pseudonymized information, which can be linked back to a person using additional information, is different from genuinely anonymized information. For a trial, inventing a realistic case is often safer than partially obscuring a real file.

Practical AI integration includes these boundaries: a usable workflow needs to protect information as well as produce a useful answer.