Controlled pilot

Inference procurement and benchmarking on your own workload

“How do we compare inference providers on our actual workload, for cost, latency, and quality, instead of on their pricing pages?”

Cryptuon benchmarks managed, self-hosted, and decentralised inference on your real workload, then runs routing and reports cost per 1M tokens and latency each month.

When teams bring us this

  • Your inference bill grew faster than usage, and nobody can say which feature or model drives it
  • You are choosing between a frontier API, an open-model host, and self-hosting, and so far the comparison is list prices
  • Users complain about slow first responses, but your dashboards show only averages
  • A single provider's outage or rate limit took your product down
  • Data-handling terms such as retention and residency now rule out some providers

Usually owned by

  • Head of AI or ML platform
  • CTO
  • Head of infrastructure
  • AI product lead
Scope

What you receive, and how “done” is defined

Deliverables

  • Workload profile from your traces: prompt and completion length distributions, concurrency, streaming share, and tool and structured-output use
  • Benchmark of three to five provider and model combinations on that workload: blended cost per 1M tokens, TTFT and throughput percentiles, error rates, and eval-set quality
  • A reproducible harness, configs, and raw results that your team can re-run when prices or models change
  • Routing and fallback design; in a sprint, the gateway implemented with cost and latency telemetry
  • Optional managed operations: provider management, monthly cost and SLO reporting, and re-benchmarking when models are released

Example acceptance criteria

  • Every candidate runs the same agreed replay set at the same concurrency levels in the same time window, and raw per-request results are delivered
  • The report gives, per candidate: blended cost per 1M tokens at your measured input/output ratio, p50/p95/p99 TTFT, output tokens per second, error and timeout rate, and eval-set pass rate
  • Your team can re-run the handed-over harness and reproduce the headline metrics within an agreed tolerance
  • Where routing is implemented, an injected provider failure triggers fallback within the agreed time, verified in a failover test

Not included unless scoped

  • Inference and GPU spend, which is passed through at cost or billed to your own provider accounts, never hidden in fees
  • Training or fine-tuning models (we evaluate fine-tuned candidates you supply)
  • Uptime beyond what the underlying providers' SLAs cover
  • Legal review of provider data-processing and retention terms
Engagement

How this is bought

Start with a fixed-fee assessment. Every later stage is optional and scoped in writing before it starts.

  1. Step 1 · Assess

    Feasibility or bottleneck assessment

    from $1,500 · 1–2 weeks

    • Current-state analysis against your real workload or codebase
    • Options compared, including ones that do not use Cryptuon technology
    • Risk register and costed, scoped recommendation
  2. Step 2 · Implement

    Implementation sprint

    from $12,000 · 3–8 weeks

    • The agreed migration or integration, delivered against written acceptance criteria
    • Tests, runbooks, and documentation
    • Handover, or transition into managed operations
  3. Step 3 · Operate

    Managed operations and maintenance

    from $1,500/month · Monthly, 3-month minimum

    • Defined monitoring, maintenance, and incident responsibilities
    • Supported changes within an agreed envelope
    • Provider management and monthly reporting

Prices exclude independent audits, substantial infrastructure consumption, and legal advice unless written into the scope. All engagement types →

The short answer

Comparing inference providers properly means replaying your traffic against each candidate and measuring five things: cost per 1M tokens at your real input/output ratio, time-to-first-token and throughput percentiles, error rate, and quality on your own eval set. Cryptuon is provider-neutral. We benchmark managed APIs, open-model hosts, self-hosted serving, and decentralised networks under identical conditions, then implement the routing and reporting that keep the result true as prices and models change.

Most buyers end up on managed providers, self-hosted serving, or a mix. The benchmark decides which.

Decision criteria

What does your workload look like in production?

Cost and latency depend on the shape of the traffic. A retrieval-heavy assistant with 6,000-token prompts and 300-token answers behaves nothing like an agent loop with short prompts, many tool calls, and long outputs. Input and output tokens are usually priced differently, and long prompts dominate time to first token. We profile your traces before choosing candidates: length distributions, concurrency peaks, streaming share, and structured-output and tool use.

What quality bar must a cheaper option clear?

A model that is cheaper per token but fails 8% more of your tasks costs more per resolved task once you count retries and escalations. Quality is scored on an eval set built from your domain, with pass criteria your team agrees. If no such set exists, building a first version is part of the assessment.

Which latency number does your user feel?

Chat users feel time to first token, readers feel output tokens per second, and agent pipelines feel end-to-end task time. Each is measured as p50/p95/p99 at stated concurrency levels, because providers that look identical at one request at a time often diverge under load.

What is your unit of cost?

For APIs, the unit is blended cost per 1M tokens at your measured ratio, adjusted for prompt-caching and batch discounts where your traffic qualifies. For self-hosting, it is GPU cost per hour ÷ (sustained tokens per second × 3,600) × 1,000,000, at the utilisation you will actually achieve. Both are then converted to cost per successful task.

What can your data-handling terms accept?

Retention, training-use, and residency terms differ between providers and between plan tiers. Some workloads rule out any third-party processing, which narrows the field to self-hosting before any benchmark runs. We record these constraints, but the legal reading of provider terms is for your counsel.

Implementation options

ApproachHow it worksStrengthTradeoffMaturity
SolanaLM (Cryptuon)OpenAI-compatible gateway routes to registered nodes running Transformers or llama.cpp, or proxying to existing providers; settlement is per request in SOLSelf-hostable, MIT-licensed gateway; a wallet can act as the API key, so agents can pay per callNo production network; compute is not cryptographically verified; you operate the nodes or depend on operatorsDeclared; source on GitHub
Managed frontier model APIsProvider-hosted proprietary models, priced per input and output tokenStrongest models, no infrastructure, mature toolingHighest unit prices; rate limits; you depend on the provider’s capacity and termsProduction
Open-model inference hostsHosted open-weights models behind per-token or dedicated-capacity pricingLow unit cost; choice of models; easy to switch hostsQuality varies by model and host; quantisation and serving settings vary, so benchmark each hostProduction
Self-hosted vLLM or SGLangYour GPUs serve open-weights models with continuous batching behind an OpenAI-compatible APIData stays with you; lowest unit cost at high utilisationYou own capacity planning, upgrades, and on-call; unit cost rises at low utilisationProduction-grade open source
Router or gateway layerOne API in front of several providers, with fallbacks, budgets, and loggingRemoves single-provider dependency; per-feature cost attributionAn extra hop and component to run; does not by itself choose the right modelsProduction

Other decentralised networks, such as Bittensor subnets or general GPU marketplaces like Akash, can be added as candidates when a buyer requires them. The same harness applies.

The delivery sequence

  1. Profile. Extract a privacy-safe replay set from your traces and agree the concurrency levels.
  2. Shortlist. Choose three to five candidates that meet your data-handling constraints.
  3. Benchmark. Run every candidate in the same window with the same harness, recording raw per-request results.
  4. Score quality. Run your eval set on every candidate and agree pass rates with your domain owners.
  5. Recommend. Rank candidates by cost per successful task within the latency SLO, with a primary and a fallback.
  6. Implement and operate. Deploy routing, failover tests, and telemetry; report monthly and re-benchmark when models are released.

The core of the harness works against any OpenAI-compatible endpoint: a managed API, an open-model host, a self-hosted vLLM server, or a SolanaLM gateway. The version shown here runs sequentially; the delivered harness drives the agreed concurrency levels.

import os
import statistics
import time

from openai import APIError, OpenAI


def run_one(client, model, messages, max_tokens):
    start = time.perf_counter()
    ttft, usage = None, None
    stream = client.chat.completions.create(
        model=model,
        messages=messages,
        max_tokens=max_tokens,
        stream=True,
        stream_options={"include_usage": True},  # check support per provider
    )
    for chunk in stream:
        if ttft is None and chunk.choices and chunk.choices[0].delta.content:
            ttft = time.perf_counter() - start
        if chunk.usage is not None:
            usage = chunk.usage
    return ttft, time.perf_counter() - start, usage


def benchmark(base_url, key_env, model, replay, usd_per_1m_in, usd_per_1m_out):
    client = OpenAI(base_url=base_url, api_key=os.environ[key_env], timeout=60)
    ttfts, decode_tps, errors, cost, tokens = [], [], 0, 0.0, 0
    for req in replay:
        try:
            ttft, total, usage = run_one(client, model, req["messages"], req["max_tokens"])
        except APIError:
            errors += 1
            continue
        if ttft is None or usage is None:
            errors += 1
            continue
        ttfts.append(ttft)
        if total > ttft:
            decode_tps.append(usage.completion_tokens / (total - ttft))
        cost += (usage.prompt_tokens * usd_per_1m_in
                 + usage.completion_tokens * usd_per_1m_out) / 1_000_000
        tokens += usage.prompt_tokens + usage.completion_tokens
    pct = statistics.quantiles(ttfts, n=100)
    return {
        "ttft_p50_s": pct[49], "ttft_p95_s": pct[94], "ttft_p99_s": pct[98],
        "decode_tps_p50": statistics.median(decode_tps),
        "error_rate": errors / len(replay),
        "blended_usd_per_1m_tokens": cost / tokens * 1_000_000,
    }

The self-hosted candidate is served with the same API, so it goes through the same harness unchanged:

# Self-hosted candidate on rented GPUs; the harness points at http://<host>:8000/v1
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.90 \
  --port 8000

Prices are taken from each provider’s published price list on the benchmark date and recorded in the report, so the comparison can be recomputed when they change.

Evidence

  • SolanaLM’s gateway, node registry (health checks and circuit breakers), Prometheus metrics, and request lifecycle are documented on the SolanaLM architecture page and in its documentation.
  • SolanaLM’s own FAQ states that cryptographic verification of inference is on the roadmap and not shipped. Its comparisons with Together AI and Bittensor set out where hosted and decentralised options differ.
  • What replicated inference does and does not prove is covered in verifiable on-chain AI inference and resolution. Agent-paid inference fits the wider agent stack described on agents and in agentic payments and x402 settlement.
  • SolanaLM’s evidence level is in the maturity register, and the engagement terms are on services.
  • Not yet evidenced: SolanaLM has no production inference network, no published latency or throughput benchmarks, and no third-party audit. Per-request SOL settlement has not been exercised at production volume. Cryptuon does not publish a standing cross-provider benchmark; each engagement measures candidates afresh on the buyer’s workload.

What drives the cost

  • Number of candidates and models. Each adds integration work, quality scoring, and pass-through inference spend.
  • Eval-set maturity. An existing, trusted eval set shortens the assessment; building one with your domain experts lengthens it.
  • Concurrency and volume tested. Load tests at production concurrency cost more in tokens and GPU hours than functional runs.
  • Self-hosted candidates. Provisioning GPUs and tuning serving settings for a fair comparison adds days and GPU spend.
  • Routing scope. A single fallback pair is simple. Per-feature routing with budgets, caching, and attribution is a larger sprint.
  • Operations envelope. How often we re-benchmark, the number of providers managed, and the reporting cadence set the monthly fee.

Limitations and what we won’t do

  • Benchmarks are snapshots. Providers change models, prices, and capacity without notice, which is why the harness is handed over and re-run, not quoted for ever.
  • We do not recommend SolanaLM, or any decentralised network, for production traffic it has not passed under your criteria. Today that usually means it is a pilot or agent-payment option, not the primary provider.
  • We do not certify that a provider meets your data-protection obligations. We record their stated terms and flag where legal review is needed.
  • Unit-cost comparisons exclude your own engineering time for self-hosting unless you ask us to model it, and we state that in the report.
  • Our assessments are engineering reviews, not independent audits.

Next step

Send us a description of the workload (or a sample of anonymised traces), the providers you use today, your latency target, and any data-handling constraints. The assessment returns a workload profile, a shortlist, measured results per candidate, and a costed recommendation for routing and operations.

Technology

Cryptuon technology we may use

Open-source components we can bring to this work. They are used only where testing on your workload supports them, and evidence levels come from the public maturity register.

Alternatives

Options we’d recommend when they fit better

The assessment compares these on your actual workload. If one of them wins, the recommendation says so.

Managed frontier model APIs

Only the strongest proprietary models pass your eval set, and per-token pricing fits your volume

Open-model inference hosts

An open-weights model passes your eval set and you want per-token pricing without running GPUs

Self-hosted serving with vLLM or SGLang

Volume is high and steady enough to keep GPUs busy, or data must stay inside your environment

Router or gateway layer (e.g. LiteLLM)

You use several providers and need one API, fallbacks, and cost attribution per team or feature

FAQ

Questions buyers ask

How do we compare LLM inference providers on our own workload?

Replay a sample of production-shaped requests (with personal data removed) against each candidate at the concurrency you actually run. Measure time to first token and throughput as percentiles, count errors and timeouts, and score the outputs on your own eval set. Then compute cost at your real input/output token ratio. Public leaderboards and list prices are a starting point, not an answer.

How much does an inference benchmarking or procurement engagement cost?

A benchmarking assessment starts from $1,500. Implementing routing, fallbacks, and telemetry is an implementation sprint, from $12,000. Ongoing provider management and monthly reporting is managed operations, from $1,500/month. Inference and GPU spend during benchmarking is a pass-through cost, listed separately in the scope.

Do you only use SolanaLM?

No. SolanaLM is one candidate, and it has no production inference network today. Most buyers end up on a managed API, an open-model host, self-hosted serving, or a mix behind a router. SolanaLM is worth benchmarking mainly where per-request on-chain settlement or agents paying from their own wallets is a requirement.

When is self-hosting LLM inference cheaper than an API?

When you can keep the GPUs busy. Self-hosted cost per 1M tokens is the GPU hourly cost divided by the tokens per hour you actually sustain, so it rises quickly at low utilisation. Spiky or low-volume traffic usually favours per-token APIs, while steady high volume on an open model that passes your evals can favour self-hosting. The benchmark measures both on your traffic.

Which latency metrics should an inference SLO use?

For chat and streaming interfaces, time to first token at p95 or p99, plus output tokens per second. For batch or agent pipelines, end-to-end latency per task and the timeout rate. Averages hide the tail that users notice, so every SLO we write is expressed as a percentile at a stated concurrency.

Will switching inference providers break our application?

It can, even between OpenAI-compatible APIs. Tool-calling formats, structured-output support, tokenisers, context limits, and default sampling settings differ. That is why the benchmark scores each candidate on your own eval set and checks the features you rely on, not just whether the endpoint responds.

Bring us the requirement

Describe the outcome, what is blocking it, and when it must work. We reply with what a scoped assessment would cover and cost.