Trust & data
Where does your business knowledge actually go?
A straight, sourced answer to every question we get asked about GitHub, Claude, ChatGPT and Gemini: what's stored, what gets used to train a model, and who owns it. Every claim below links to the provider's own policy, not a summary of one.
AI providers change these policies more often than most software terms. GitHub changed its own mid-way through this page being written. Treat this as a current snapshot, not a permanent guarantee, and use the source links throughout if you need the live version for a specific decision.
The short version
Two separate questions, not one.
Most of the worry we hear treats “is my knowledge safe” as a single question. It's actually two, and they have different answers: where is it stored, and who's allowed to process it when an AI tool reads it.
Your repository sits in GitHub.
Your AI Brain is built and stored as a private GitHub repository. Nothing is sent to any AI company automatically just because it exists there. It sits still until something reads it.
An AI tool reads it, on request.
The moment Claude, ChatGPT or Gemini is asked to use it (to answer a question, draft a letter, whatever), that specific tool's own policy takes over. GitHub is no longer involved.
Which policy applies at step 2 depends far more on which plan is doing the reading than on which company built it. That distinction runs through almost every answer below.
Storage
GitHub: where the repository lives.
Does GitHub use our knowledge repository to train its own AI?
No. GitHub's own position, restated in their March 2026 privacy update, is direct: “We do not use private repository content at rest to train AI models.” That's true whether the repository sits on GitHub Free, Team or Enterprise: the plan doesn't change this particular commitment.
Source: GitHub Changelog, 25 March 2026
Who can actually see the repository?
Whoever you grant access to, through GitHub's ordinary private repository permissions, the same access controls used by every private codebase on GitHub. It isn't publicly discoverable, isn't indexed, and isn't shared with GitHub's own products beyond what's described above.
Processing
The AI tool you connect to it.
Once content leaves GitHub and enters a conversation with Claude, ChatGPT or Gemini, treat it exactly as if you'd typed or pasted it in directly. The fact it came from a repository doesn't change which policy applies. For example, Claude's own GitHub connection pulls in “files (names and contents) in a repo on a specific branch”, not commit history or anything else, and from that point it's simply an input to Claude, governed by whichever Claude plan is asking.
Does Anthropic (Claude) train on our content?
On Claude for Work, Team, Enterprise or the API (commercial products): no, not by default. Anthropic's own wording: “By default, we will not use your inputs or outputs from our commercial products to train our models.” Inputs and outputs are automatically deleted from Anthropic's backend within 30 days, and a Zero Data Retention arrangement is available for qualifying organisations, removing storage entirely.
On a personal Claude Free, Pro or Max account: a different, consumer agreement applies. Training is opt-in: decline it and you get 30-day retention, accept it and retention extends to five years. Conversations flagged by Anthropic's automated safety systems can be retained for up to two years and used for model improvement regardless of that choice.
Sources: Anthropic Privacy Center, commercial training policy, Anthropic, consumer terms update
What does “30 days of retention” actually mean?
It means a copy of what was sent and what came back sits on the provider's own servers for up to 30 days, for ordinary operational reasons (debugging, abuse monitoring, service reliability), the same kind of retention almost any cloud software keeps. It is not reviewed by humans as a matter of course, isn't shared outside the company, and, on a business-tier account, isn't used for training. After 30 days, it's automatically deleted. The exceptions worth knowing: content flagged for a genuine policy violation can be kept longer (up to two years, at Anthropic), and a Zero Data Retention arrangement can remove even the 30-day window for organisations that want it.
Does OpenAI (ChatGPT) train on our content?
On ChatGPT Business, Enterprise, Edu or the API: no, not by default. OpenAI's own wording: “By default, we do not use data from ChatGPT Enterprise, ChatGPT Business… or our API platform, including inputs or outputs, for training or improving our models.” Retention is configurable, including a zero data retention option for qualifying organisations on the API.
On a personal or free ChatGPT account, the default settings can include model training use, with an explicit setting to turn it off. Worth checking directly if a free or personal Plus account is ever the one connected to business content.
Source: OpenAI, Business data privacy, security and compliance
Does Google (Gemini) train on our content?
On a licensed Google Workspace with Gemini account: no. Google's own wording: “Submissions aren't used to train models and are never reviewed by humans,” and content stays inside your organisation's domain rather than being shared outward.
On a personal Gemini account attached to an ordinary Gmail login, rather than a licensed Workspace seat, weaker consumer protections apply instead. That's the same pattern as the other three providers.
The part that actually matters
Business-tier accounts vs. personal ones.
Every answer above follows the same shape: strong, default, contractual protection on the business-tier product; weaker, conditional protection on the free or personal version of the same tool. GitHub, Claude, ChatGPT and Gemini all agree with each other on this point. The provider matters less than the plan.
We're a small team: could we just buy everyone a personal Claude Pro account, instead of Team, to save money?
Nothing technically stops it. Claude Pro includes Cowork, and day-to-day it can do the same work. But it's a real downside, not just a smaller privacy setting:
- Training and retention default to the weaker consumer terms above, not the commercial ones.
- The contract with Anthropic belongs to each individual, not the practice. There's no business-to-business agreement, and no Data Processing Agreement to point to if one's ever needed.
- Consumer accounts can't be shared or centrally managed. Anthropic's own terms prohibit sharing account logins, so “a bunch of Pro accounts” really does mean one per person, each fully owned by that person, not the business.
- When someone leaves, their account, and everything they ever discussed in it, leaves with them. No admin console, no export, no way to revoke access or confirm what was in there. That's the exact risk you're trying to avoid, just relocated to the day someone resigns.
It's cheaper on the sticker price, but it re-creates the problem it's meant to solve: business knowledge sitting in accounts the business doesn't own. If cost is the real constraint, Claude Team's per-seat price isn't far above Pro's. Almost all of the gap comes from its five-seat minimum. For a one or two-person admin team, the Claude or OpenAI API (pay-as-you-go, same commercial protections, no seat minimum) is usually the cheaper honest route to the same guarantee.
Sources: Anthropic Consumer Terms of Service, Anthropic, who owns and manages team data
Ownership
Who owns your knowledge, and what AI produces from it.
Our SOPs, templates and business knowledge: who owns that once it's in the repository?
You do. Storing something in a private repository, or having an AI tool read it to answer a question, doesn't transfer ownership to GitHub, Anthropic, OpenAI or Google. Standard business-tier terms across all four grant them only a limited licence to operate the service (process the prompt, generate the response), not ownership of what you put in.
What about the reports, letters and templates the AI helps produce?
Generally that's yours too, under each provider's standard terms. The genuine caveat: AI-generated output doesn't always attract copyright protection the same way human-authored work does, and that's an unsettled area of law rather than something any provider, or we, can promise outright. If a specific piece of output needs watertight ownership, that's worth a conversation with a lawyer, the same caveat we'd put on our own terms of use.
Alternatives
Do we have to accept this, or is there another way?
For a small practice's business and admin knowledge (not clinical records, which sit outside the AI Brain entirely), avoiding cloud AI altogether isn't a proportionate response to what the research above actually shows. On a correctly-configured business-tier account, the protection isn't a hope, it's a contract. That said, if a specific piece of knowledge genuinely warrants more caution, here's the real spectrum of options, from the one we'd recommend by default through to a genuine last resort.
Business-tier accounts
Our default recommendation for every AI Brain: GitHub on any plan, plus the client's AI tool on its Team, Enterprise or API tier. Covers everything above, for a modest per-seat premium over the personal tier.
Zero Data Retention
Available from Anthropic and OpenAI for qualifying business accounts. Removes even the 30-day storage window: nothing is kept after the response is returned.
Keep the most sensitive items out entirely
Anything genuinely irreplaceable or competitively sensitive can simply stay out of the repository, or be redacted, as a human-only reference, the same judgement call any business makes about what goes in a shared drive.
A signed Data Processing Agreement
Available on the business-tier plans above for organisations with formal compliance obligations: a direct, signed commitment rather than relying on the default published terms.
Locally-hosted models: a genuine last resort
Technically possible for an organisation that decides public cloud AI isn't appropriate at all: open-weight models run on infrastructure the business owns. A materially bigger technical and cost undertaking, and not something we currently build or recommend for a small practice's day-to-day AI Brain use. It trades away the quality of the frontier tools for a governance guarantee most businesses can already get contractually, for a fraction of the cost.
Is it “inevitable”? In the sense that using any modern software runs on trusting a contract, not a promise: yes, the same way your practice management software, accounting platform and email already do. In the sense that you're stuck with weak protection and no recourse: no. The actual choice is which contract applies, and which tier you're on decides that. That's within your control, and it's a decision made once, correctly, rather than something to keep trusting blindly.
A note on this page
This page reflects our understanding of each provider's published policies as of August 2026, drawn directly from their own documentation, linked throughout rather than summarised second-hand. It isn't legal advice. Providers change these terms periodically, so for anything that needs a binding guarantee, check the current policy at the source links above, or talk to a lawyer, the same way we'd say it about our own terms of use.
Still have questions?
Ask us directly.
Happy to walk through exactly how this applies to your business, before you commit to anything.