solanalm On-Chain AI Federated Learning Solana GPU Networks

Decentralized LLM Inference with Solana Settlement

How SolanaLM runs an OpenAI-compatible inference network on independent GPUs: gateway and registry design, onion-routed privacy, and per-request settlement.

DS
Dipankar Sarkar
• • 18 min read • 3,596 words

A decentralized LLM inference network lets anyone with a GPU serve model requests and get paid per request, without a platform sitting in the middle taking custody of either the money or the prompt. Cryptuon’s SolanaLM is a working instance of that idea: a FastAPI gateway routes requests to independently operated inference nodes, the client SDK speaks the OpenAI schema so existing applications switch with a base-URL change, settlement and node incentives run on Solana, and an optional onion-routed circuit means the operator serving your prompt cannot tell who sent it. The same runtime also coordinates federated learning rounds — FedAvg, FedProx, FedAdam, and SCAFFOLD — so a node earns from both serving inference and contributing training cycles.


TL;DR

  • Three problems have to be solved together: discovery (which node can serve this model?), settlement (how does it get paid per request?), and privacy (who sees the prompt?). Solving only one produces a demo, not a network.
  • OpenAI compatibility is a distribution decision, not a convenience. Speaking an existing schema means an application adopts the network by changing base_url, with no rewrite and no new SDK to learn.
  • Settlement is the part that needs a chain. Per-request payments at fractions of a cent are not viable on rails with per-transaction overhead measured in cents and finality measured in minutes.
  • Privacy is a circuit, not a promise. A three-hop onion route separates the identity of the requester from the content of the request. It defeats a curious operator; it does not defeat a global passive adversary.
  • Federated learning gives idle GPUs a second income. FedAvg, FedProx, FedAdam, and SCAFFOLD each fail differently in production, and picking the wrong one wastes rounds rather than producing a bad model.
  • Where this loses: a hosted provider will beat it on p99 latency and support, and a local runner will beat it on cost for a single machine. The network wins on censorship resistance, monetisation, and privacy topology.
  • See how we compare inference providers on your workload, or describe your project if you are routing production traffic.

The three problems, and why they are one problem

Every attempt at decentralised inference runs into the same three questions, and the reason most attempts stall is that the questions are coupled.

Discovery. A client asks for llama-3.1-70b-instruct at a given context length. Somewhere in the network are nodes with the weights resident, enough VRAM free, and an acceptable queue depth. Finding them is a routing problem — but the routing table has to be trustworthy, because a node that lies about its capacity to win traffic degrades the whole network.

Settlement. Once a node serves a request, it needs paying, and the payment has to be proportional to the work — tokens generated, context length, model size. At the price points inference actually trades at, a single request is worth a fraction of a cent. Any settlement rail whose per-transaction cost approaches the value of the transaction is unusable.

Privacy. The prompt is the product. Enterprise traffic is full of customer records, source code, and internal strategy. A network where the node operator can read every prompt, and correlate prompts to a paying identity, is not a viable destination for that traffic regardless of how cheap it is.

The coupling is what makes this hard. Discovery wants to know who the client is, so it can prioritise and rate-limit. Settlement wants to know who the client is, so it can charge them. Privacy wants nobody to know who the client is. SolanaLM resolves this by separating the payment identity from the request path: the chain knows an account paid, the exit node knows a prompt arrived, and no single component holds both facts.

OpenAI compatibility as a distribution strategy

The most consequential design decision in SolanaLM is also the least technically exciting: the gateway speaks the OpenAI HTTP schema.

The reasoning is about adoption cost, not elegance. Every LLM application already has an OpenAI-shaped client somewhere — the official SDK, LangChain, LlamaIndex, a hand-rolled requests wrapper. Introducing a new schema means every one of those integrations needs rewriting, and rewriting an integration to try an unproven network is a cost nobody pays for a maybe.

Speaking the existing schema reduces adoption to one line:

from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.solanalm.example/v1",
    api_key=SOLANALM_NODE_TOKEN,      # scoped to a funded Solana account
)

resp = client.chat.completions.create(
    model="llama-3.1-70b-instruct",
    messages=[
        {"role": "system", "content": "You extract structured fields from invoices."},
        {"role": "user", "content": invoice_text},
    ],
    temperature=0.0,
    max_tokens=512,
)

print(resp.choices[0].message.content)
print(resp.usage.prompt_tokens, resp.usage.completion_tokens)

That is a production application pointed at a decentralised network with a two-line diff. The tradeoff is real and worth naming: adopting someone else’s schema means inheriting their design decisions, including the parts that fit awkwardly. Node selection, privacy mode, and settlement preferences have no home in the OpenAI request body, so they live in headers and in per-key configuration rather than in the request itself. That is a slightly uglier API than a purpose-built one would be, and it is the correct trade, because an uglier API that existing code can call beats a beautiful one that requires a rewrite.

Architecture: gateway, registry, nodes, privacy layer, settlement

The system decomposes into five components with deliberately narrow responsibilities.

The gateway is a FastAPI service that terminates the OpenAI-compatible API. It authenticates the caller, resolves the requested model to a candidate set of nodes, forwards the request, streams the response back, and emits a usage record. Multiple gateways can run; the gateway is not the network.

The registry is the discovery layer. Nodes announce the models they serve, their hardware class, their context limits, and their pricing. Registry entries are backed by the operator’s Solana account, which is what makes a false advertisement expensive rather than free.

Inference nodes run the models. A node can host PyTorch and Transformers directly, run llama.cpp for quantised local models, or proxy to an upstream provider — OpenAI, Anthropic, Cohere, or a local Ollama instance. That last capability is more important than it looks: it means a node can join the network without owning a GPU at all, which bootstraps liquidity for models the network cannot yet serve natively.

The privacy layer implements onion-routed inference and differential-privacy aggregation for federated rounds. It is optional per request, because the latency cost is real and not every workload needs it.

Settlement is the Solana leg: escrowed balances, per-request accounting, and operator payouts.

Why Solana for the settlement leg

The requirement is unusual. Inference settlement needs high transaction volume, tiny per-transaction value, and fast enough finality that a node is not extending unbounded credit to strangers.

That combination rules out most rails immediately. Card networks have a per-transaction floor far above the value of a request. Ethereum L1 finality is minutes, and base-layer fees regularly exceed the entire value of a completion. Invoicing monthly works, but it reintroduces exactly the trusted intermediary the network exists to remove, and it means an operator carries credit risk against anonymous counterparties.

Solana’s combination of sub-second confirmation and fees measured in small fractions of a cent is what makes per-request settlement arithmetic work at all. In practice the gateway batches usage into periodic settlement rather than writing one transaction per completion — the chain is the clearing layer, not the metering layer — but the option of settling at request granularity is what keeps the credit window short enough that nobody needs to trust anybody for long.

Cryptuon has written about the same latency-versus-finality question from the payments side in How AI Agents Pay: x402 and On-Chain Settlement, where the constraint is identical and the conclusion is the same: finality, not throughput, is the number that decides whether machine-scale payments are feasible.

SolanaLM resources: solanalm.cryptuon.com · Documentation · Source on GitHub

Running a node

A node is a Python service with a model backend and a Solana keypair. The configuration is mostly about declaring honestly what the hardware can do.

# node.yaml
node:
  name: "gpu-edinburgh-01"
  keypair_path: "~/.config/solana/node.json"
  public_endpoint: "https://node01.example:8443"

backends:
  - kind: transformers
    model: "meta-llama/Llama-3.1-70B-Instruct"
    dtype: bfloat16
    max_context: 32768
    max_concurrent: 4

  - kind: llama_cpp
    model_path: "/models/mistral-7b-instruct-q5_k_m.gguf"
    max_context: 8192
    max_concurrent: 12

  - kind: proxy            # no local GPU needed for this one
    upstream: "ollama"
    base_url: "http://127.0.0.1:11434"
    models: ["qwen2.5-coder:14b"]

pricing:
  unit: "per_1k_tokens"
  prompt_lamports: 900
  completion_lamports: 2400

privacy:
  accept_onion_routed: true
  log_prompts: false       # enforced by the daemon; advertised in the registry

federation:
  participate: true
  algorithms: ["fedavg", "fedprox", "scaffold"]
  max_round_minutes: 25

max_concurrent is the field operators most often get wrong. Advertising more concurrency than the GPU can sustain wins traffic in the short term and destroys the node’s reputation when queue depth blows out p99 latency. The registry records advertised capacity; the network measures delivered capacity; the gap is what reputation is built from.

Onion-routed inference: what the circuit actually protects

The privacy design borrows directly from onion routing. A request destined for a privacy-mode completion is wrapped in nested encryption and passed through three hops. The entry hop knows the client’s network address but sees only ciphertext. The middle hop knows neither endpoint. The exit hop — the node that actually runs the model — sees the plaintext prompt but has no idea who sent it.

from solanalm import PrivateClient

client = PrivateClient(
    keypair_path="~/.config/solana/id.json",
    circuit_hops=3,
    exit_policy={"require_no_log": True, "min_reputation": 0.85},
)

resp = client.completions(
    model="llama-3.1-70b-instruct",
    prompt=confidential_contract_text,
    max_tokens=1024,
)

What this stops is concrete and limited, and it is worth being precise about both halves.

It defeats a curious operator. The node serving your prompt cannot link it to your identity, your other prompts, or your payment account. Prompt content and requester identity are held by different parties, and neither can reconstruct the pair alone.

It defeats network-level observation at a single point. An observer watching the entry hop sees encrypted traffic to a relay; an observer watching the exit sees a prompt with no attributable origin.

It does not defeat a global passive adversary who can observe both ends of the circuit and correlate timing and volume. It does not stop an exit node from logging the prompt — the require_no_log policy is an advertised commitment backed by reputation and stake, not a cryptographic guarantee. And it does not make the content private from the model: the exit node runs the inference, so it necessarily sees the plaintext. If the prompt must never be visible to any third party in cleartext, onion routing is the wrong tool and the answer is a local model on hardware you control.

Stating that plainly is the point. A privacy layer that oversells itself is worse than none, because it changes what users are willing to send.

Federated learning: the second income stream

An inference node is idle whenever nobody is asking it questions, and a GPU that is idle is a GPU losing money. SolanaLM’s federation runtime uses that capacity for training rounds where the data never leaves the node that owns it — only model updates travel.

Four algorithms are implemented because they fail differently, and the failure mode is the whole basis for choosing.

FedAvg is the baseline: each participant trains locally, the coordinator averages the weights. It works well when participants’ data is roughly identically distributed. When it is not — and in a permissionless network it never is — local models drift toward their own data and averaging them produces something worse than any individual contributor.

FedProx adds a proximal term that penalises local weights for drifting too far from the global model. It costs convergence speed and buys stability under heterogeneous data. This is usually the right first thing to try when FedAvg is oscillating.

FedAdam moves adaptive optimisation to the server side, treating aggregated client updates as gradients. It helps when clients contribute very unequal amounts of data, because the server-side optimiser can smooth the resulting noise.

SCAFFOLD attacks client drift directly with control variates — correction terms carried per client that cancel the drift rather than merely penalising it. It converges in fewer rounds on badly non-IID data and costs roughly double the communication per round, because the control variates travel alongside the updates.

from solanalm.federation import Round, Aggregator, DifferentialPrivacy

round_cfg = Round(
    base_model="mistral-7b-instruct",
    algorithm="scaffold",          # non-IID node data; drift correction earns its cost
    min_participants=12,
    local_epochs=2,
    client_lr=2e-5,
    dp=DifferentialPrivacy(
        clip_norm=1.0,             # bound each client's contribution first
        noise_multiplier=1.1,      # then add calibrated Gaussian noise
    ),
)

agg = Aggregator(round_cfg)
result = agg.run()                 # blocks until quorum reports or the round times out

print(result.participants, result.dropped)
print(result.epsilon_spent)        # cumulative privacy budget across rounds

The differential-privacy step is not decoration. Model updates leak information about training data, and in a network of mutually untrusted participants that leak is an attack surface. Clipping bounds any single client’s influence — which also limits poisoning — and calibrated noise bounds what can be inferred from the aggregate. The cost is accuracy, and epsilon_spent is the number that says how much privacy budget the project has already consumed. Treating it as a running total rather than a per-round setting is the difference between a privacy claim and a privacy accounting.

How the options compare

OptionModel accessSettlementPrivacy topologyWhere it genuinely winsMaturity
SolanaLMOpenAI-compatible; any model a node hosts, plus proxy backendsPer-request accounting settled on SolanaOptional 3-hop onion circuit; DP aggregation for trainingOpen participation, prompt/identity separation, federated training on the same fleetActive development
BittensorSubnet-defined interfaces, varies by subnetNative token emissions on its own chainSubnet-specific; not a first-class circuitDeep incentive research, large established validator ecosystem, strong communityMature network, established economics
Akash NetworkWhatever you deploy; generic compute leasingNative token, lease-based rather than per-requestWhatever your deployed stack providesGenuinely general-purpose compute, mature marketplace, you control the whole stackMature
Together AICurated, well-optimised hosted modelsTraditional invoicingProvider’s policy; standard cloud trust modelBest p99 latency and throughput, real support, no ops burden at allProduction, commercial
OllamaLocal models on your own machineNone neededPerfect — nothing leaves the hostZero marginal cost, zero latency to the model, complete controlMature, single-host

The honest reading: if you want the lowest latency and somebody to call when it breaks, use a hosted provider. If you want absolute privacy for one workload on one machine, run Ollama locally. If you want general compute rather than an inference API specifically, Akash is purpose-built for that and SolanaLM is not. Bittensor has been iterating on decentralised-AI incentive design longer and has a larger operator base.

SolanaLM’s narrower claim is the combination: an OpenAI-shaped API that existing code can call, settlement fast and cheap enough for per-request economics, a privacy circuit that separates prompt from payer, and a federation runtime that monetises the same hardware between requests. Each of those exists somewhere else. Having all four behind one endpoint is the argument.

Business outcome: where the money actually moves

Two sides of the market, two different calculations.

For an operator, a GPU has a fixed cost whether it is busy or not. Utilisation is the entire economics of the machine. Running an inference node adds a revenue stream that fills gaps between whatever the hardware was bought for, and federated rounds fill the gaps between the gaps. The relevant metric is not headline rate per token but the change in achieved utilisation, because that is where the margin lives.

For a consumer, the calculation is different and less about unit price than teams expect. The interesting saving is avoided commitment: no minimum spend, no annual contract, no renegotiation when volume moves, and no single provider whose terms change unilaterally. For workloads that are spiky, experimental, or politically sensitive about where data lands, optionality is worth more than the last fraction of a cent per thousand tokens.

There is also a category of work that simply cannot use a hosted API — regulated data that may not leave a jurisdiction, or a model trained on data that participants will not pool. Federated rounds address that directly: the data never moves, only updates do, and differential privacy bounds what those updates reveal. That is not a cheaper version of an existing option; it is the only option for that workload.

The verification angle — proving a node actually ran the model it claimed — is a separate problem covered in Verifiable On-Chain AI Without a Trusted Oracle, and the same replicated-compute techniques appear in Cryptuon’s research on decentralised detection networks.

Limitations

Current limitations

  • Tail latency is worse than a hosted provider’s, structurally. Gateway routing, variable node load, and cold model loads all add variance. A workload with a hard p99 requirement should keep a hosted fallback.
  • Onion routing costs latency and does not hide the prompt from the exit node. Three hops add round trips, and the node running inference necessarily sees plaintext. It separates identity from content; it does not encrypt content end to end.
  • Model availability depends on operators. If nobody has loaded the weights you want, the network cannot serve it. Proxy backends paper over this at the cost of reintroducing an upstream provider.
  • Federated rounds are fragile to churn. A participant dropping mid-round wastes the round for everyone. Quorum settings and timeouts are a tuning problem with no universally right answer.
  • Differential privacy costs accuracy, and the budget is finite. Every round spends epsilon. A long-running federation eventually has to choose between stopping and weakening its own guarantee.

Roadmap items

  • Richer routing that accounts for measured rather than advertised capacity, including per-model latency history per node.
  • Stronger guarantees for the no-log exit policy, moving it from an advertised commitment toward something attestable.
  • Better handling of partial federated rounds, so a dropped participant degrades a round rather than voiding it.
  • Settlement ergonomics for teams: sub-accounts, spend caps, and per-project accounting on top of a single funded balance.

Frequently Asked Questions

What is decentralized LLM inference?

It is serving language model requests across independently operated machines rather than a single provider’s datacentre, with a protocol handling discovery, routing, and payment. SolanaLM implements this with a FastAPI gateway, a node registry backed by Solana accounts, operator-run inference nodes, and on-chain settlement, while exposing an OpenAI-compatible API so existing applications need no rewrite.

How do I switch an existing app to SolanaLM?

Point your OpenAI client at the gateway’s base URL and supply a key scoped to a funded Solana account. Because the gateway speaks the OpenAI HTTP schema, the official SDKs, LangChain, LlamaIndex, and hand-rolled clients all work unchanged. Features with no place in the OpenAI request body — privacy mode, node selection policy — are configured through headers and per-key settings.

Can the node operator read my prompt?

In standard mode, yes: the node runs the model, so it sees the plaintext. In privacy mode a three-hop onion circuit separates identity from content — the exit node sees the prompt but not who sent it, and the entry node sees who you are but only ciphertext. If a prompt must never be visible to any third party at all, run the model on hardware you control.

What hardware do I need to run a node?

Anything from a single consumer GPU running quantised models through llama.cpp to a multi-GPU machine hosting a 70B model in bfloat16. A node can also run in proxy mode with no GPU, forwarding to a local Ollama instance or an upstream provider. Declare capacity honestly in node.yaml: over-advertising wins traffic briefly and costs reputation permanently.

Which federated learning algorithm should I use?

Start with FedAvg if participant data is roughly identically distributed. Move to FedProx when local models drift and averaging degrades the global model. Use FedAdam when clients contribute very unequal data volumes. Choose SCAFFOLD for badly non-IID data where drift correction is worth roughly double the per-round communication.

How is this different from Bittensor or Akash?

Bittensor is a broader incentive network organised into subnets with its own chain and emission economics, and it has a larger established operator base. Akash leases generic compute rather than serving an inference API. SolanaLM is narrower than both: an OpenAI-compatible inference and federated-learning network settling on Solana, with onion-routed privacy as a first-class request mode.

The bottom line

Decentralised inference has never been blocked by a lack of GPUs. It has been blocked by three unglamorous problems: finding a node that can serve your model, paying it small amounts quickly enough that nobody carries credit risk, and keeping the prompt away from the party that knows who you are.

SolanaLM’s answer is to solve all three behind an interface that existing applications already speak. The OpenAI schema removes adoption cost. Solana settlement makes per-request economics arithmetically possible. The onion circuit splits payer from prompt. And the federation runtime turns the idle half of a GPU’s day into a second reason to keep it in the network.

None of that beats a hosted provider on latency, and it is not supposed to. It is for workloads where open participation, privacy topology, or the ability to train on data that cannot be pooled matter more than the last few milliseconds. If that describes yours, see how we compare inference providers on your workload or describe your project.

DS

Dipankar Sarkar

Founder, Cryptuon

Blockchain researcher and systems engineer. Author of 5 published papers on cross-chain composability, MEV mitigation, and DePIN protocols. Building production blockchain infrastructure in Rust and Zig.

Have a requirement like this?

Tell us the outcome you need, what is blocking it, and when it must work. We reply with what a scoped assessment would cover and cost.