Quick answer. "Nothing leaves your servers" is true in exactly one architecture: the model, the search index and the connector all run on hardware the firm controls. With any hosted model the matter content is sent to the provider to be read, whatever the connector does locally, and the claims that survive are training off and zero retention under the firm's own written agreement.
The test in one sentence: ask where the model runs; if the answer is anywhere outside your walls, the slogan is wrong and the real questions are where the content goes, who can reach it there, and what's kept afterwards.
What we wrote, and where it was wrong
The page is /sr/pravni-softver/, our Serbian hub for law-firm software, written as a landing page for lawyers in Belgrade. The post it links to, on which AI tools a Serbian firm can accept without breaking professional secrecy, was written against vendor documents and closes with this (our translation): "A vendor who tells you nothing ever left the office, while its model runs over the internet, either doesn't understand its own product or is counting on you not asking." The hub was making that exact claim, in front of the one audience professionally trained to notice. The fixing commit changed ten lines, five out and five in.
| Where on the page | What it said | What it says now |
|---|---|---|
| FAQ, "what can the AI do" | The difference from ChatGPT in a browser is that "podaci ne izlaze iz kancelarije" (the data doesn't leave the office) | The AI is connected to your matters and works under the terms the firm contracts for |
| FAQ, confidentiality | "Podaci ostaju kod vas" (your data stays with you), then role-based access and audit logging | The firm sets the processing terms in writing: training on your data switched off, and no storage of inputs, contracted |
| Hero paragraph | "Podaci ostaju kod vas, advokat zadržava kontrolu i potpis" (your data stays with you, the lawyer keeps control and the signature) | Processing terms are agreed up front, the lawyer keeps control and the signature |
| Section heading | "Podaci ostaju kod vas" | "Uslovi obrade su vaša odluka" (the processing terms are your decision) |
| Card under that heading | Audit logging, plus the option to run fully locally on the firm's own equipment | Adds the training-off and no-storage line, "in writing, not as a promise", and links to the post that shows how to verify it. The local-model sentence stayed, because on that path the claim is true |
The FAQPage schema was regenerated against the visible copy at the same time, eleven questions, so the JSON copy that search engines read matched the visible one. How it happened isn't mysterious: the hub was written to sell and the post was written to hold up, both by us, and nobody grepped one against the other before shipping. Publishing a rigorous post next to a loose page exposes the page, so any content that sets a standard now gets the pages linking to it searched for claims that fail that standard before it goes live.
We ran that search across the English site the same day, and again on 27 August. Twenty-one English files carry a sentence of the "never leaves your" family, and every one sits in a category where it's true: credentials that stay on the user's machine because the connector runs locally, self-hosted search where the next sentence says the AI step sends passages to the provider, data migration inside the firm's own cloud tenant, database-level access scoping, on-prem model deployments, and posts quoting the claim in order to reject it. No English page makes it for a hosted model. The Serbian page was written faster and sold harder, and that's the whole explanation.
When "nothing leaves" is actually true
The claim has a home, and we've built in it. LegalSearch is a document search platform we built for a 17-person law firm with fifteen years of files, more than a million documents. The firm's requirement ruled out public-cloud AI, so the index, the search layer and the embedding model all run on its own servers, access is scoped so an associate only sees the matters they're assigned to, and every query lands in an audit log the firm owns. That case study says "documents never leave the firm's servers", and we stand behind it, because there's no model provider in the request path to send anything to. The same page lists a hosted-embedding option for firms without that requirement, and on that option the honest sentence narrows to "never leave your servers for retrieval", which is how our self-hosted search pages already phrase it.
The other home is a local model: an open-weight model on a Mac Studio or a Linux box in the office, with the connector on the same machine, the setup in the on-prem privilege stack. It comes with a check a firm can run itself: capture the network traffic during a session and confirm nothing goes out to any AI provider's domain. If you can pull the network cable and the assistant still answers questions that don't need fresh practice-management data, the claim is true. What you're trusting here is your own IT, meaning patching, backups and who can walk into the server room, and you're accepting that local models sit a step behind hosted ones and cost more per drafting task. It fits a small number of firms with a specific reason, and it's the only category that also takes a foreign cloud provider's legal exposure out of the picture, which the Australian data region post goes into.
Everywhere else, the model reads the content somewhere that isn't your office, and the slogan has to be replaced with something that can be checked.
The three questions that replace the slogan
Where does the content go? It sorts every vendor into a category, and each category asks you to trust something different. Each is a legitimate place to be; describing one row as if it were another is the problem.
| Category | Looks like | Where the content goes | What you're trusting |
|---|---|---|---|
| Your own hardware | Open-weight model in the office, self-hosted search like LegalSearch | Nowhere. It's read on your machine | Your own IT |
| Cloud, in your own account | Claude on Amazon Bedrock, the equivalents on other clouds | A machine in a region you pick, inside your account boundary | The cloud vendor's documentation and contract, plus whoever configured the account |
| Direct API, your own agreement | The model provider's API under the firm's commercial terms, with a zero retention agreement where offered | The provider's infrastructure, for the length of the request | The provider's commercial terms and the retention addendum you signed |
| Vendor-hosted product | Harvey, CoCounsel, Clio Duo, MyCase IQ, an embedded feature, a chat subscription used in the browser | The vendor's infrastructure and its sub-processors, including whichever model provider it calls | The vendor's contract, its sub-processor list, and which model version it calls today |
| Consumer plan | A personal chat subscription | The provider, under consumer terms | Terms written for individuals. After United States v. Heppner, not a place for privileged material |
A local connector doesn't move a vendor up this table. Our own connector pages say your credentials never leave your machine, which is true: the connector reads the practice-management API from your laptop with no relay server. The matter it fetches then goes to whichever model the firm chose, and if that model is hosted, the privacy claim belongs to the model's row.
Who can reach it there? On your own hardware, your staff. In a cloud account, the cloud operator and the model provider, and the useful vendors put both in writing: Amazon's Bedrock documentation states the service "uses a zero operator access (ZOA) data security model", meaning "no operators of the service can access model input or output", and says separately that model providers have no access to customer prompts and completions. On a direct API, the provider's staff under whatever abuse-detection conditions the terms allow. In a vendor-hosted product, the vendor's staff, its sub-processors and the model provider behind them, three organisations to ask instead of one. On a consumer plan, potentially the provider's training pipeline.
What's kept afterwards, and under whose terms? Two separate promises live here and they get conflated constantly. Training is the first. Anthropic's Commercial Terms, Section B, state that "Anthropic may not train models on Customer Content from Services". Retention is the second: Anthropic's retention page says "we automatically delete inputs and outputs on our backend within 30 days of receipt or generation", unless "you and we have agreed otherwise (e.g. zero data retention agreement)". So the API default is 30 days, and zero is a signed agreement granted at the organization level, which the Team vs API post shows a small firm can obtain. On Bedrock the default is stronger, per the same AWS abuse detection page: "Amazon Bedrock uses a zero data retention (ZDR) data security model. This means that by default, Amazon Bedrock does not store model inputs or outputs." The page then carries an exception for one current model, retained for up to 30 days with an opt-in to sharing that traffic with the model provider for abuse detection and potential human review, and notes that eligible customers may request full zero retention through their account team. Model names go stale in a compliance record, so the zero data retention ceiling post keeps the current list with a verification date.
For a vendor-hosted product, all of those answers sit in the vendor's contract and depend on which model version the product calls, which can change without notice. A consumer plan sits under consumer terms, which are different terms; read them rather than assume the enterprise ones carry over.
Access vs retention: the distinction most evaluations miss
Retention is what happens to the content after the request. Access is whether the model read it at all. Firms tend to ask "is it stored?", get a good answer, and stop, and that's the gap. A zero retention agreement constrains what happens later; it doesn't undo the reading. An audit log records that access happened, which is useful, and says nothing about whether it was permissible. For privileged material, being read by a third party is the exposure, and every hosted category on the table above involves it.
That's why even the corrected copy on our Serbian hub doesn't say the content stays home. Read closely, it says: the content left, it was read, nothing was learned from it, nothing was kept, and all of that is in writing. For most matters at most firms that's enough, and it's a stronger sentence to hand a client than a slogan, because every clause points at a document. For some matters it isn't enough, and those justify the first row of the table. The court in United States v. Heppner (S.D.N.Y., February 2026) left privilege inside "a closed, enterprise-grade AI system" as an open question, which is another way of saying a firm on a hosted model is making a reasoned bet and a firm on its own hardware isn't.
The practical follow-on is to narrow what gets read. A well-built integration sends the clause, the field or the single document a task needs. Sending the whole matter file is usually just the easier thing to code.
A vendor test you can run on anyone, including us
Seven questions. A vendor in any of the five categories can answer all of them, and the answers will differ by category without any of them being wrong. Score the answers; the category is context.
- Where does the model run? The answer should place the vendor on one row of the table. "It's secure" is not a row.
- What exactly does the model see per task? Fields, documents, or the whole matter. Then ask whether it could be narrower.
- Which model version, by name, and what does that version do with inputs and outputs? A vendor that can't name it hasn't answered the retention question yet, whatever the brochure says.
- Is training on our content switched off, and where is that written? A contract clause or the provider's commercial terms, quoted. Not a FAQ.
- What's retained, for how long, and under whose agreement? The default window, and whether zero retention is in place and who signed it. A vendor-hosted product has to answer for its model provider too.
- Can a person at the vendor or at the model provider read our content, and under what conditions? Abuse review, support access, cross-region routing. The honest answer for most hosted setups is "under these conditions", written down.
- What changes when you upgrade the model, and will we be told first? An upgrade can move a firm from zero retention into a retention-and-review window without anyone deciding to.
Reserve the red flag for hedging. A vendor on any of the five rows can produce seven clean answers. What fails the test is a slogan in place of an answer, and specifically a page that says "nothing leaves" from any row but the first. When you find one, ask which row they're on and whether they'll correct the page. We were on the second and third rows, the page said row one, and we corrected it in five places. We'd rather you check than take our word for that.
Want the seven questions run against your own setup?
We build the open-source Clio connector, the self-hosted search platform in the case study above, and the hosted-model integrations that sit on the second and third rows of the table. Thirty minutes is enough to tell you which row you're on, which sentence you can put in writing for a client, and which sentence you can't.
Book a 30-minute architecture call →Vendor lines verified 27 August 2026 against Anthropic's Commercial Terms of Service (effective 17 June 2025), Anthropic's organization data retention page (dated 1 July 2026) and the Amazon Bedrock abuse detection page. Terms change. Check the vendor's current page before relying on any of it, including this one.
Frequently asked questions
Is 'nothing leaves your servers' ever a true claim for legal AI?
Yes, in one architecture: the model, the search index and the connector all run on hardware the firm controls. An open-weight model on a server in the office, or a self-hosted document search platform, can honestly say the content never left. With any hosted model, the matter content is sent to the model provider to be read, whatever the connector does locally, so the claim is wrong for cloud-in-your-account setups, direct API setups and vendor-hosted products alike. For those, the defensible claims are that training on your content is switched off and that nothing is stored after the response, both under the firm's own written agreement.
Does an AI vendor see my client data if it says it does not train on it?
With a hosted model, yes: the model has to read the content to answer, and a no-training promise only governs what happens afterwards. Anthropic's Commercial Terms, for example, state that Anthropic may not train models on Customer Content from Services, and by default inputs and outputs are deleted within 30 days unless a zero data retention agreement is in place. Neither line says the content stayed in your office. They say it left, was read, was not learned from, and was not kept. For privileged material, being read is the exposure, so ask what exactly the model sees per task and whether it could be narrower.
What is the difference between zero data retention and no training?
No training means the provider does not use your inputs and outputs to improve its models. On the Anthropic API that is the contractual default. Zero data retention means your prompts and responses are not stored after the request finishes, and it is a separate arrangement: the API default is deletion within 30 days, and zero retention is granted by agreement at the organization level. Amazon Bedrock states that by default it does not store model inputs or outputs and that no operator of the service can access them, with a documented exception for one current model that requires 30-day retention and opt-in to provider review. For privileged content you want both controls, in writing, and you still have not made the content stay home.
If the connector runs locally, does the data stay local?
Only the part the connector handles. A local connector keeps your practice-management credentials on your machine and reads the matter from the practice-management API without a relay server. That is a true and useful claim, and it is what our own connector pages say. The matter content the connector fetches is then sent to whichever model the firm has chosen, and if that model is hosted, the content leaves. A local connector plus a hosted model is a cloud setup with a local front end, and the privacy claim has to be made about the model, not the connector.
What should a law firm ask an AI vendor instead of 'does the data leave'?
Three questions. Where does the content go, meaning where does the model run: your hardware, a cloud account you own, the provider's API under your own agreement, or the vendor's infrastructure. Who can reach it there: your staff, the cloud operator, the model provider's staff under abuse-detection conditions, the vendor's sub-processors. And what is kept afterwards, for how long, under whose signed terms. Every category of vendor can answer all three. The red flag is a vendor that answers with 'it is secure' or with a slogan, and a vendor whose page says 'nothing leaves' while running a hosted model should be asked whether they will correct the page.