You signed an NDA to get the contract. Your data is covered by HIPAA, or a defense clause, or a client agreement that says the records never touch a third party. Then someone on your team asks whether AI could speed up the tedious part — reading intake forms, flagging the risky files, clearing a backlog nobody has hours for. The reflexive answer is no. AI means the cloud, the cloud means your data lands on somebody else's server, and that is the one thing your contracts forbid.
That reflex is understandable, and most of the time it is wrong. We are LegacyForge AI, based in Sioux Falls, and we build automation for businesses that cannot let their data leave the building. The approach has a name — captive — and it is worth understanding before you rule out the whole category.
Where the assumption comes from
Nearly every AI product a normal person has touched runs the same way. You type into a box, your text travels to a vendor's servers, a model there produces an answer, and it travels back. The famous chat tools work like this. So do most of the AI features bolted onto software you already pay for. Your input becomes the vendor's input. For a lot of businesses that is completely fine. For yours, it may be the exact thing your contracts prohibit.
But that arrangement is a deployment choice, not a property of AI itself. Where the computation happens, who controls the hardware, and whether your raw data is ever touched are all knobs. Captive work turns those knobs the other direction.
The captive pattern, plainly
Captive means the intelligence comes to your data instead of your data going to the intelligence. There is a ladder of control, and how far up you climb depends on how strict your rules are:
- De-identified exports. The sensitive fields — names, account numbers, anything that identifies a person — get stripped or masked before anything is processed. The system works on the shape of the data, not the identities.
- Processing on hardware you control. The computation runs on a machine you designate, on your premises or in an environment you own, not on a rented cloud account someone else spins up.
- On-premise deployment. The whole system lives inside your building. Nothing it does requires a connection to the outside world.
- No third-party AI cloud in the pilot path. When we test whether the thing even works, the trial does not quietly route your records through an outside model to find out.
- No training on your data. Whatever we build does not feed your records back into a model that other people's projects then benefit from. Your data trains nothing.
You do not need every rung. A marketing agency de-identifying a lead list needs the first one. A clinic or a defense subcontractor might need all five. The point is that "use AI" and "send our data to a vendor" are two separate decisions. You can say yes to the first while saying a hard no to the second.
The strongest version: nothing generative runs in the building at all
There is a version stricter than on-premise, and it is often the right one. Instead of installing a model in your building, we study the decision you are trying to speed up, work out the actual rules behind it with your people, and encode those rules in ordinary deterministic code. The intelligence goes in at build time. What ships to you is a program that does the same thing every run. No model sits in the loop. Nothing is generated at runtime.
A concrete case. You have a technician who spent fifteen years learning which incoming jobs are quick wins and which are money pits. Most of that judgment is a set of patterns — this combination of factors tends toward this outcome. We sit with them, pull the pattern out, and write it into code that scores a new job in a fraction of a second. The AI did the hard part during the build, distilling the judgment. The thing running in your building afterward is just software following rules you approved. It cannot leak your data to a model, because there is no model there to leak to.
That deterministic core matters for another reason: a human stays in the loop. These systems produce recommendations — this file needs a second look, this job scores low, this record is an outlier. A person on your side approves or overrides. Nothing acts on its own, and every recommendation traces back to a rule you can read, not a black box that shrugs.
Cloud-with-a-contract is legitimate — just not for everyone
Honesty requires the other side of this. Plenty of serious AI vendors will sign a Business Associate Agreement or an equivalent data contract, keep your records out of training, and hand you a defensible paper trail. For a great many regulated businesses that is genuinely enough, and it is cheaper and faster than building something captive. We will say so when it is the smarter buy. We would rather point you at a signed BAA than sell you an on-premise project you did not need.
It stops being enough when the honest answer to "can this data ever sit on infrastructure we do not control" is a flat no. Some NDAs and government clauses are written exactly that way. No contract with an outside vendor satisfies a rule that says the records never leave the building. That is where captive earns its keep, and where a BAA, however well drafted, does not clear the bar.
What it costs, and how we keep you from guessing
Automation work does not get a price list, because pinning a number on a project we have not scoped would be a guess dressed up as a quote. It starts with a free discovery conversation — what the task is, what the data is, which rungs of that ladder your rules actually require. Then, before anyone signs up for a big build, we prove it with a small pilot measured in dollars: one workflow, real numbers, hours or errors saved that you can see on paper. If the pilot does not pay, you have lost nothing and learned where your real bottleneck is. If it does, we scope the rest from evidence instead of a pitch.
We are based in Sioux Falls and we serve businesses anywhere — most captive work can be built and delivered remotely. When a job genuinely needs hands on your hardware inside a secured room, and the strictest air-gapped setups do, we travel for the right price. The constraint on reach is your rules, not our zip code.
So when someone asks whether they can use AI without their data leaving their building, the answer is yes — on hardware you control, or with nothing generative running inside your walls at all. That old instinct to keep everything in-house was protecting you, and it does not have to close the door on the work. Tell us the task and the constraint. We will tell you whether captive is the right shape, whether a signed data contract would serve you better, or whether the job is not worth automating yet. Start with the discovery conversation. It costs nothing, and you walk away knowing more about your own operation either way.
