// Halo

Run Llama without hosting it — pay per prompt on Halo

Meta's open-weight Llama models, served by independent operators. Open weights without the GPU bill — pay per prompt, no account.

Llama is Meta’s open-weight family, and open weights are exactly why it’s interesting: nobody can take the model away from you. The catch is that “open” doesn’t mean “free to run” — someone still has to hold a GPU.

Halo is where that someone is a market rather than a vendor. Independent operators serve Llama alongside 140+ other models, and you pay per prompt instead of per hour of hardware.

Open weights without the GPU bill

Self-hosting means renting or owning a GPU, keeping a serving stack alive, and paying for it whether you prompt once a day or a thousand times. On Halo you pay per prompt in USDC on Base, for the tokens you actually use. Idle costs nothing.

No account, no API key

Sign in with a social account and prompt. There’s no provider signup, no API key to rotate, and no card on file.

Proof, not trust

Serving an open model honestly and serving a cheaper substitute look identical from the outside — unless you check. Halo checks: every result carries a statistical proof of execution, and the network rejects results that fail it.

Serve it yourself, too

Llama runs on hardware people already own. If you have a machine sitting idle, you can be on the other side of this market — see what to serve and operator pricing and earnings.

Also worth comparing: Mistral, DeepSeek and Qwen.

Frequently asked questions

Why not just self-host Llama?

You can — the weights are open. But self-hosting means a GPU you rent or own, a serving stack to maintain, and a bill that runs whether or not you're prompting. Halo gives you the same open model on demand, priced per prompt, with none of the ops.

Which Llama models are served?

Operators serve models from the Llama line, and availability shifts as they come online. The app shows what's actually being served at any moment, so you're never picking from a stale list.

How do I know an operator ran the real Llama and not a smaller model?

Every result carries a statistical proof of execution. The genuine model produces roughly 90%+ token-distribution overlap with the verifier; a substitute or fabricated output looks like noise and is rejected by the network.

Is it cheaper than a hosted API?

Often, because independent operators compete for each request rather than a single vendor setting one price. You also pay strictly per prompt, with no minimum and no idle GPU time.