What if a hospital could give you everything its data knows about a disease — without ever giving you a single patient’s data?
That sounds like a riddle. It’s actually an architecture question, and I think about it more than almost anything else, because the answer decides whose diseases get studied.
The file
Here is the idea in its plainest form. A biobank — a hospital’s collection of biological data, gathered over years from real people — runs a training process on its own machines, inside its own walls. What comes out is not a copy of the data. It is a small file, on the order of fifty to a hundred megabytes, that captures the statistical shape of the dataset: the patterns, the structure, the regularities that make the data scientifically valuable in the first place.
That file contains none of the original records. Not encrypted, not anonymized — absent. But load it into a model somewhere else, and the model can generate synthetic data with the same statistical properties as the originals. The knowledge travels. The people it came from stay home. And synthetic data is already a working tool in this space: the Rare AI Archive publishes a synthetic-patients dataset built to support training, precisely so that real patient records don’t have to travel.
I want to be precise here, because there is a tempting wrong version of this idea. This is not a zip file. Compression keeps the data and makes it smaller; this keeps the structure and discards the data. And that distinction is exactly why it can protect anyone: what leaves the building is the shape of a population, not the story of a person.
Why it works — and when it wouldn’t
Training a model is, at heart, an act of distillation. The more real pattern a dataset has — the more its contents rhyme with each other — the more a model can learn from it. Train on truly random noise and there is nothing to learn; the method would give you nothing back. So this isn’t magic, and it has an honest boundary: it works because biology has structure. Vast libraries of protein data can distill toward a portable file precisely because proteins are not random.
That boundary is worth respecting out loud. If someone tells you a technique like this works on everything, they are selling something. It works where structure exists, and medicine is lucky: structure is what bodies are made of.
Turn the pipeline around
The deeper move hiding in that file is a reversal. The way we usually imagine medical AI, the data travels: your records go to the big computer somewhere, and you hope the people who run it are careful. The file lets the pipeline run the other way. The model — or the distilled essence a model needs — is what travels. The sensitive thing stays where it was collected, under the governance of the people who collected it and the people it describes.
Honesty requires one footnote to that picture: models don’t “travel” like a visiting specialist does — they are copied, and anything a site sends back still needs careful de-identification before it moves. The reversal changes where the risk lives; it does not abolish it. Getting that residual part right is real work, and it is nobody’s afterthought.
But think about what the reversal makes possible. Three cancer institutes, each holding data it cannot ethically or legally hand over, could collaborate on a treatment question — each contributing a distilled file, none emailing a single patient record. Institutions that today cannot work together because sharing is impossible could work together because sharing is no longer required.
The patient in the sentence
I keep saying “a patient’s data,” singular, and that is deliberate. It is easy to talk about privacy at the scale of databases, and at that scale it becomes an engineering abstraction. But data is trust. Every row in a biobank is a person who let medicine look at the most private thing they have — their own biology — usually in the hope that it would help someone like them.
A rare disease patient understands this trade in their bones, because we live on both sides of it. Our data is so scarce and so identifying that sharing it is genuinely risky — and so scarce and so valuable that not sharing it can mean no one ever studies our disease at all. An architecture where the knowledge can travel while the person stays protected isn’t a compliance feature to us. It is the difference between being studied with dignity and not being studied.
That is the standard I want to hold this work to: the patient is a collaborator, not a raw material. The file is one way to build that respect directly into the plumbing.
Where this stands
None of this describes a product you can download today. It describes the architecture I am building toward — the same direction as the verification work I wrote about in the paper behind Lattice, and the same reason: multi-institution research, especially in rare disease, needs infrastructure where trust is engineered rather than assumed. Distillation for privacy, verification for integrity, federation for scale. Each is a hard problem; none is finished; I’d rather show you the blueprints honestly than sell you a building that isn’t there.
But the question at the top has stopped being a riddle for me. Can a hospital share what its data knows without sharing its data? Yes — in principle cleanly, in practice carefully. A fifty-megabyte yes.