In January 2025, Eric Boniardi, Alison Haire and I put out a paper — arXiv 2501.05374. It’s a preprint, not yet through peer review, and the title is deliberately dry: “Validation of GPU Computation in Decentralized, Trustless Networks.” What it actually is: the starting point for why I believe Lattice Protocol can exist at all.
Let me start with the problem, because the problem is more interesting than the solution.
The awkward question
Say you rent time on a GPU somewhere you have never been, owned by someone you will never meet, to run a model that matters to you. The machine sends back a result. Look closer at what you are actually holding: an answer, and nothing whatsoever about how it came to exist. How do you know it actually ran your computation — and not a cheaper approximation, or nothing at all?
This is not paranoia. It is the live question under every GPU marketplace and every federated-compute network. If you cannot answer it, “decentralized AI compute” is just trust with extra steps.
The obvious answer is to run the job again yourself and compare, bit for bit. In an open network, that fails fast. You can force a GPU into deterministic modes — the tooling exists — but a network of strangers runs on whatever hardware and settings it has, and across different machines, floating-point operations reorder and honest results drift apart in the low decimal places. The drift is small enough to miss — you only notice it several decimal places in — and that is exactly what makes it treacherous: exact matching would flag honest strangers as cheaters. It is the wrong mental model for an open network.
What doesn’t work for us, first
I want to be honest about the paths we didn’t take, because each one is a real trade, not a strawman.
You could demand special hardware — trusted execution environments that attest to what they ran. That works where you can require the right silicon and firmware, and it narrows your network to the machines that have them. You could push toward fully homomorphic encryption — but FHE’s job is keeping the data private, which is a different problem from proving the computation was done, and its overhead on model inference is still a serious practical constraint, though where that line sits keeps moving. Each approach buys its guarantee by giving up some of the open network.
We wanted to find out how much guarantee you can get without giving that up.
What the paper actually does
Here is the part I most want to get right, because it is easy to oversell.
We examined three candidate signals, each one a layer deeper into the work than the last — from the shape of the output, to what the output means, to the physical trace of the computation that produced it. Model fingerprinting — does the output carry the statistical signature of the right model? It turned out to be the least reliable of the three on its own; its limits are part of what the paper reports. Semantic similarity — does the result mean what it should, even when the exact bits differ? This became the core of the verification framework we propose. GPU profiling — the physical fingerprint of the work itself, timing and behavior. We describe it and defer it: it is future work, not a component of the tested system.
The framework we actually tested is humbler than a trust-free free-for-all: responses are compared against a trusted reference node — a machine you do trust, checking machines you don’t — first in a binary setup, then in a small network of three response nodes and two verifiers. Evaluated on a thousand-question test set against trusted reference responses, the paper reports 76.5% verification accuracy. Friendly nodes, a handful of machines, a research demo.
I write that number with some pride, because 76.5% on an honest testbed is what a real starting point looks like — and because telling you it “guarantees integrity” would be exactly the kind of claim this paper exists to replace with evidence. It doesn’t yet handle a determined adversary, or colluding nodes, and it still leans on that trusted reference. Those are the open problems, named in the paper, that the next work has to close.
Theory before product
I am an engineer by temperament; I like building the thing. But you cannot build a trustless compute network on the hope that nobody cheats. The argument had to come first — that meaningful verification on ordinary hardware is possible and measurable — before the architecture made any sense.
So the order matters: theory, then architecture, then product. The paper is not marketing for Lattice; it is the first load-bearing brick. And I should be precise about the connection people ask me about most. The rare-disease dream — research that spans many hospitals’ compute without centralizing anyone’s patient data — needs verification like this to be solved, and it needs privacy machinery, governance, and clinical validation that are entirely separate problems. A paper about checking GPUs is not a patient-data system. It is one prerequisite among several, and I’d rather tell you which brick this is than pretend it’s the building.
The title is dry on purpose. The question under it is not: can you check the work of a computer you will never see? What we wrote down is the beginning of a yes — how far it holds today, measured; where it stops, named. Almost everything I am building now is an attempt to extend it.