The short answer
Comparing inference providers properly means replaying your traffic against each candidate and measuring five things: cost per 1M tokens at your real input/output ratio, time-to-first-token and throughput percentiles, error rate, and quality on your own eval set. Cryptuon is provider-neutral. We benchmark managed APIs, open-model hosts, self-hosted serving, and decentralised networks under identical conditions, then implement the routing and reporting that keep the result true as prices and models change.
Most buyers end up on managed providers, self-hosted serving, or a mix. The benchmark decides which.
Decision criteria
What does your workload look like in production?
Cost and latency depend on the shape of the traffic. A retrieval-heavy assistant with 6,000-token prompts and 300-token answers behaves nothing like an agent loop with short prompts, many tool calls, and long outputs. Input and output tokens are usually priced differently, and long prompts dominate time to first token. We profile your traces before choosing candidates: length distributions, concurrency peaks, streaming share, and structured-output and tool use.
What quality bar must a cheaper option clear?
A model that is cheaper per token but fails 8% more of your tasks costs more per resolved task once you count retries and escalations. Quality is scored on an eval set built from your domain, with pass criteria your team agrees. If no such set exists, building a first version is part of the assessment.
Which latency number does your user feel?
Chat users feel time to first token, readers feel output tokens per second, and agent pipelines feel end-to-end task time. Each is measured as p50/p95/p99 at stated concurrency levels, because providers that look identical at one request at a time often diverge under load.
What is your unit of cost?
For APIs, the unit is blended cost per 1M tokens at your measured ratio, adjusted for prompt-caching and batch discounts where your traffic qualifies. For self-hosting, it is GPU cost per hour ÷ (sustained tokens per second × 3,600) × 1,000,000, at the utilisation you will actually achieve. Both are then converted to cost per successful task.
What can your data-handling terms accept?
Retention, training-use, and residency terms differ between providers and between plan tiers. Some workloads rule out any third-party processing, which narrows the field to self-hosting before any benchmark runs. We record these constraints, but the legal reading of provider terms is for your counsel.
Implementation options
| Approach | How it works | Strength | Tradeoff | Maturity |
|---|---|---|---|---|
| SolanaLM (Cryptuon) | OpenAI-compatible gateway routes to registered nodes running Transformers or llama.cpp, or proxying to existing providers; settlement is per request in SOL | Self-hostable, MIT-licensed gateway; a wallet can act as the API key, so agents can pay per call | No production network; compute is not cryptographically verified; you operate the nodes or depend on operators | Declared; source on GitHub |
| Managed frontier model APIs | Provider-hosted proprietary models, priced per input and output token | Strongest models, no infrastructure, mature tooling | Highest unit prices; rate limits; you depend on the provider’s capacity and terms | Production |
| Open-model inference hosts | Hosted open-weights models behind per-token or dedicated-capacity pricing | Low unit cost; choice of models; easy to switch hosts | Quality varies by model and host; quantisation and serving settings vary, so benchmark each host | Production |
| Self-hosted vLLM or SGLang | Your GPUs serve open-weights models with continuous batching behind an OpenAI-compatible API | Data stays with you; lowest unit cost at high utilisation | You own capacity planning, upgrades, and on-call; unit cost rises at low utilisation | Production-grade open source |
| Router or gateway layer | One API in front of several providers, with fallbacks, budgets, and logging | Removes single-provider dependency; per-feature cost attribution | An extra hop and component to run; does not by itself choose the right models | Production |
Other decentralised networks, such as Bittensor subnets or general GPU marketplaces like Akash, can be added as candidates when a buyer requires them. The same harness applies.
The delivery sequence
- Profile. Extract a privacy-safe replay set from your traces and agree the concurrency levels.
- Shortlist. Choose three to five candidates that meet your data-handling constraints.
- Benchmark. Run every candidate in the same window with the same harness, recording raw per-request results.
- Score quality. Run your eval set on every candidate and agree pass rates with your domain owners.
- Recommend. Rank candidates by cost per successful task within the latency SLO, with a primary and a fallback.
- Implement and operate. Deploy routing, failover tests, and telemetry; report monthly and re-benchmark when models are released.
The core of the harness works against any OpenAI-compatible endpoint: a managed API, an open-model host, a self-hosted vLLM server, or a SolanaLM gateway. The version shown here runs sequentially; the delivered harness drives the agreed concurrency levels.
import os
import statistics
import time
from openai import APIError, OpenAI
def run_one(client, model, messages, max_tokens):
start = time.perf_counter()
ttft, usage = None, None
stream = client.chat.completions.create(
model=model,
messages=messages,
max_tokens=max_tokens,
stream=True,
stream_options={"include_usage": True}, # check support per provider
)
for chunk in stream:
if ttft is None and chunk.choices and chunk.choices[0].delta.content:
ttft = time.perf_counter() - start
if chunk.usage is not None:
usage = chunk.usage
return ttft, time.perf_counter() - start, usage
def benchmark(base_url, key_env, model, replay, usd_per_1m_in, usd_per_1m_out):
client = OpenAI(base_url=base_url, api_key=os.environ[key_env], timeout=60)
ttfts, decode_tps, errors, cost, tokens = [], [], 0, 0.0, 0
for req in replay:
try:
ttft, total, usage = run_one(client, model, req["messages"], req["max_tokens"])
except APIError:
errors += 1
continue
if ttft is None or usage is None:
errors += 1
continue
ttfts.append(ttft)
if total > ttft:
decode_tps.append(usage.completion_tokens / (total - ttft))
cost += (usage.prompt_tokens * usd_per_1m_in
+ usage.completion_tokens * usd_per_1m_out) / 1_000_000
tokens += usage.prompt_tokens + usage.completion_tokens
pct = statistics.quantiles(ttfts, n=100)
return {
"ttft_p50_s": pct[49], "ttft_p95_s": pct[94], "ttft_p99_s": pct[98],
"decode_tps_p50": statistics.median(decode_tps),
"error_rate": errors / len(replay),
"blended_usd_per_1m_tokens": cost / tokens * 1_000_000,
}
The self-hosted candidate is served with the same API, so it goes through the same harness unchanged:
# Self-hosted candidate on rented GPUs; the harness points at http://<host>:8000/v1
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--port 8000
Prices are taken from each provider’s published price list on the benchmark date and recorded in the report, so the comparison can be recomputed when they change.
Evidence
- SolanaLM’s gateway, node registry (health checks and circuit breakers), Prometheus metrics, and request lifecycle are documented on the SolanaLM architecture page and in its documentation.
- SolanaLM’s own FAQ states that cryptographic verification of inference is on the roadmap and not shipped. Its comparisons with Together AI and Bittensor set out where hosted and decentralised options differ.
- What replicated inference does and does not prove is covered in verifiable on-chain AI inference and resolution. Agent-paid inference fits the wider agent stack described on agents and in agentic payments and x402 settlement.
- SolanaLM’s evidence level is in the maturity register, and the engagement terms are on services.
- Not yet evidenced: SolanaLM has no production inference network, no published latency or throughput benchmarks, and no third-party audit. Per-request SOL settlement has not been exercised at production volume. Cryptuon does not publish a standing cross-provider benchmark; each engagement measures candidates afresh on the buyer’s workload.
What drives the cost
- Number of candidates and models. Each adds integration work, quality scoring, and pass-through inference spend.
- Eval-set maturity. An existing, trusted eval set shortens the assessment; building one with your domain experts lengthens it.
- Concurrency and volume tested. Load tests at production concurrency cost more in tokens and GPU hours than functional runs.
- Self-hosted candidates. Provisioning GPUs and tuning serving settings for a fair comparison adds days and GPU spend.
- Routing scope. A single fallback pair is simple. Per-feature routing with budgets, caching, and attribution is a larger sprint.
- Operations envelope. How often we re-benchmark, the number of providers managed, and the reporting cadence set the monthly fee.
Limitations and what we won’t do
- Benchmarks are snapshots. Providers change models, prices, and capacity without notice, which is why the harness is handed over and re-run, not quoted for ever.
- We do not recommend SolanaLM, or any decentralised network, for production traffic it has not passed under your criteria. Today that usually means it is a pilot or agent-payment option, not the primary provider.
- We do not certify that a provider meets your data-protection obligations. We record their stated terms and flag where legal review is needed.
- Unit-cost comparisons exclude your own engineering time for self-hosting unless you ask us to model it, and we state that in the report.
- Our assessments are engineering reviews, not independent audits.
Next step
Send us a description of the workload (or a sample of anonymised traces), the providers you use today, your latency target, and any data-handling constraints. The assessment returns a workload profile, a shortlist, measured results per candidate, and a costed recommendation for routing and operations.