Obit runs inference across machines we do not own. Those machines can read the pieces of the model they are running, so a worker who keeps a copy has a copy of the model. The defences that genuinely prevent this either cost minutes per token or need hardware most GPUs do not have. RBWO is our attempt at the middle ground.
Serving a large model across many machines is how you use hardware you could not otherwise afford, including idle GPUs on marketplaces. The catch is that each machine holds a piece of the model in a form it can read and run. Anyone with access to that machine can copy the piece, and enough pieces add up to a working model.
The standard answer is confidential computing, where the hardware itself keeps the operator out. That works, and it is the right answer when the hardware supports it. On commodity and marketplace GPUs it usually does not, and that is exactly where cheap capacity lives.
RBWO takes a different angle. Instead of hiding the weights from the machine, we make the weights useless on their own. Each piece is stored in a transformed form that only computes correctly when combined with a secret held somewhere else. A worker who copies a piece has numbers that do not run as a model.
This is not private inference and it is not encryption. It raises the cost of stealing a usable model rather than making theft impossible, and its value depends entirely on how well that holds up against an attacker who is trying.
Cryptographic private inference gives real guarantees about inputs, and sometimes about weights. It is also far too slow for serving. Published systems report token generation for a 7B model in the range of minutes per token, which rules it out for an API.
Confidential GPUs are the strongest direct alternative, and where they work the overhead is now small, in some measurements under a tenth of throughput for large models. They require specific hardware, attestation, and a machine you can trust to report honestly about itself. A permissionless marketplace GPU offers none of those.
Watermarking and fingerprinting help you prove a model was stolen after the fact. They do not stop the copy from being made or used.
Transforming weights so that untrusted hardware cannot use them is not new either. Systems have done it for single layers and for parts of a model, mostly to protect user inputs in a trusted-accelerator setting. What we have not seen is that idea applied to a pipeline spread across independent operators who may collude, which is the situation a distributed inference network is actually in.
The measurement that matters is not accuracy or speed. It is whether an attacker who has the shards can end up with a model worth using.
The attack we expect to decide it is the cheapest one. An attacker holding transformed middle layers but not the embedding or the head can train the missing pieces themselves, using the traffic they serve and any public text. If that works within a realistic budget, the scheme does not do its job, and we would rather find that out early than ship around it.
A second risk is that the protected model is a fine-tune of a public one. In that case the attacker already has a strong template and may be able to line the transformed weights back up against it.
So the success test has to be stated before the experiment, in terms of how good the best recovered model is compared to the original, with a threshold for what counts as commercially useful. A recovered model that scores close to the original is a failure even if the exact original weights were never recovered.
It is not private inference. Workers that see internal states may still be able to infer something about the text being processed, and we are not claiming otherwise.
It is not a cryptographic guarantee. We are not claiming better security than encrypted computation or confidential hardware, and a compromised trusted service defeats the whole scheme.
It does not protect a model whose full transformed pipeline has leaked, because a complete transformed graph runs perfectly well without the original weights. Keeping at least one piece out of reach is the entire mechanism.
What it might be is a way to use hardware that confidential computing cannot reach, at a cost that makes it worth deploying, against an attacker who is not willing to spend much. Whether that is true is an empirical question we have not answered.