Core delivery

Solana and blockchain data cost reduction

“Our historical blockchain data is becoming too expensive to store and query, and we need to decide whether to change providers or change our architecture.”

Cryptuon benchmarks your real Solana data workload, compares provider, hybrid, and self-managed options, then redesigns, migrates, and operates the one that wins.

When teams bring us this

  • Your RPC or data-provider bill is growing faster than your users or revenue
  • Historical queries (backfills, address histories, old blocks) are slow, rate-limited, or time out
  • Your indexer is a pile of polling jobs nobody wants to own, and re-syncs take days
  • You are about to renew a provider contract, or a provider has changed pricing or limits
  • Finance has asked whether self-hosting would be cheaper and nobody has a measured answer

Usually owned by

  • Infrastructure lead
  • Data lead
  • CTO
Scope

What you receive, and how “done” is defined

Deliverables

  • Workload profile: your real query mix, volumes, history depth, and freshness needs, captured from logs or traffic
  • Benchmark of the current setup and two to three candidate architectures against that workload, with cost modelled under stated assumptions
  • Storage and query redesign: schema, partitioning, retention tiers, and compression choices for the recommended option
  • Migration with dual-running, reconciliation reports, and a rollback path
  • Runbooks and dashboards, with optional managed operations afterwards

Example acceptance criteria

  • Workload coverage: every query class in the agreed workload profile is served by the new design, with none silently dropped
  • Query correctness: results match the reference source for a sampled set of addresses, slots, and time ranges, with zero unexplained mismatches
  • Latency: p50, p95, and p99 for each query class at or under the agreed targets, measured at the agreed load
  • Recovery: the store rebuilds or catches up from a defined failure (node loss, stream gap, corrupt partition) within the agreed time, demonstrated in a drill
  • Total monthly cost at or under the agreed figure under the stated volume, retention, and pricing assumptions

Not included unless scoped

  • Provider fees, cloud consumption, and hardware, which are passed through or paid directly by you
  • Building end-user analytics dashboards or BI reporting beyond operational monitoring
  • Chain coverage outside the agreed networks (Solana first; EVM chains scoped separately)
  • Guaranteeing a third-party provider's uptime or future pricing
Engagement

How this is bought

Start with a fixed-fee assessment. Every later stage is optional and scoped in writing before it starts.

  1. Step 1 · Assess

    Feasibility or bottleneck assessment

    from $1,500 · 1–2 weeks

    • Current-state analysis against your real workload or codebase
    • Options compared, including ones that do not use Cryptuon technology
    • Risk register and costed, scoped recommendation
  2. Step 2 · Implement

    Implementation sprint

    from $12,000 · 3–8 weeks

    • The agreed migration or integration, delivered against written acceptance criteria
    • Tests, runbooks, and documentation
    • Handover, or transition into managed operations
  3. Step 3 · Operate

    Managed operations and maintenance

    from $1,500/month · Monthly, 3-month minimum

    • Defined monitoring, maintenance, and incident responsibilities
    • Supported changes within an agreed envelope
    • Provider management and monthly reporting

Prices exclude independent audits, substantial infrastructure consumption, and legal advice unless written into the scope. All engagement types →

The short answer

Historical blockchain data gets expensive when every read, including reads of data that will never change, is paid for per call, or when a self-built indexer was never designed for the query mix it now serves. Cryptuon measures your actual workload, benchmarks the current setup against provider, hybrid, and self-managed alternatives under stated cost assumptions, and then implements and can operate the winner against latency, correctness, and recovery criteria you sign off before the build.

Often the answer is a better use of the provider you already have. The assessment is designed to find that out before anyone migrates anything.

Decision criteria

The questions that decide cost and approach:

What is your query mix, by class?

“RPC costs” is usually three or four different workloads billed together. Typical classes are point lookups of recent state, transaction fetches by signature, address-history pagination, block-range backfills, and analytical aggregations. Each one has a different cheapest home. An address history that pages through getSignaturesForAddress 1,000 signatures at a time and then fetches each transaction is a fan-out pattern; for busy addresses it can dominate both latency and the bill.

How much of what you read is immutable?

Finalized blocks and transactions do not change. If your system re-fetches them from a provider, a content-addressed cache or a local store can remove that traffic entirely. If most of your reads are of live account state, caching helps less and freshness becomes the binding constraint.

How deep and how hot does history need to be?

The last few days, the last year, and genesis-to-date are three different storage problems. Hot history belongs on fast disk or in a warehouse. Rarely read history can sit compressed in object storage, or be read from an archive such as Old Faithful, usually with slower retrieval than hot storage. Paying hot-archive prices for data you read once a quarter is a common source of overspend.

Who will own ingestion and gap recovery?

Streaming with Yellowstone gRPC (Geyser) gives low-latency, filtered feeds, but streams do not backfill themselves. Someone has to detect missed slots, re-fetch them, and reconcile. If nobody on the team will own that, a managed provider’s indexed APIs may be cheaper in total even when the invoice is larger.

What does a wrong answer cost you?

Wallet balances, accounting exports, and compliance reports need correctness tests and reconciliation. A trend dashboard can tolerate small gaps. This decides how much verification the design carries.

Implementation options

ApproachHow it worksStrengthTradeoffMaturity
Better use of the current providerCache immutable reads, cut polling, trim requested fields, move to the right plan or methodLowest risk; no migrationCeiling on savings; the provider still prices history and fan-outEstablished
Another managed provider (Helius, Triton One, QuickNode)Move some or all traffic to a provider whose pricing and APIs fit your mixFast to adopt, with indexed APIs, webhooks, and gRPC streams availableVendor pricing can change; you are still billed per call or per creditEstablished
HybridProvider for live state and streams; your own store for history and analyticsUsually the best cost and latency balance for read-heavy teamsTwo systems to reconcile; needs gap detectionEstablished pattern
Self-managed pipeline (Geyser plus ClickHouse or Parquet/DuckDB)Run or rent a Geyser-enabled node, stream into a columnar store, backfill from RPC or archiveFull control, flat cost at volume, SQL over your dataYou own ingestion, backfills, upgrades, and on-callEstablished components; integration is bespoke
Cryptuon components (blockchain-compression, SolanaVault, StreamSync designs)Chain-tuned compression for archived blobs; a self-hostable compressed block store with Solana JSON-RPC; sharded DuckDB serving patternsCan reduce stored bytes and serve history through an unchanged RPC interfaceRatios are workload-dependent; SolanaVault and StreamSync networks are not liveblockchain-compression: reproduced (crates.io). SolanaVault, StreamSync: declared, no third-party audit published

Most recommendations are a hybrid: keep a managed provider for live reads and streams, and move history and analytics into a store you control. The Cryptuon components enter the design only when the benchmark shows a gain over the off-the-shelf choice.

The delivery sequence

  1. Capture the workload. Pull a representative window of request logs or provider usage exports, then classify each request by method, parameters, history depth, and caller.
  2. Define the benchmark. Fix the query classes, sample sets (addresses, slots, time ranges), load profile, and latency percentiles that will become acceptance criteria.
  3. Measure the baseline. Run the benchmark against your current setup to establish latency, error rate, and cost per query class.
  4. Test the candidates. Run the same benchmark against two to three candidate architectures, including at least one that uses no Cryptuon technology.
  5. Model cost. Project monthly cost per option under explicit assumptions: volume growth, retention, provider pricing on a stated date, cloud and storage rates, and engineering time.
  6. Build and dual-run. Implement the chosen design and run it alongside the current one, with reconciliation reports, until correctness and latency criteria are met.
  7. Cut over and drill. Shift traffic gradually, rehearse recovery from a stream gap and a lost node, and hand over or move into managed operations.

The benchmark harness is the core artefact. A trimmed version of the kind we use is below; it runs the same sampled query class against two endpoints over raw JSON-RPC, records latency percentiles, and flags any disagreement between them:

// Benchmark one query class (block signatures by slot) against two endpoints.
// Records p50/p95/p99 latency and counts result mismatches against the reference.
type Endpoint = { name: string; url: string };

async function rpc(url: string, method: string, params: unknown[]) {
  const res = await fetch(url, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({ jsonrpc: "2.0", id: 1, method, params }),
  });
  const body = await res.json();
  if (body.error) throw new Error(`${method}: ${body.error.message}`);
  return body.result;
}

function percentile(samples: number[], p: number): number {
  const sorted = [...samples].sort((a, b) => a - b);
  return sorted[Math.min(sorted.length - 1, Math.floor((p / 100) * sorted.length))];
}

export async function benchBlocks(reference: Endpoint, candidate: Endpoint, slots: number[]) {
  const latencies: number[] = [];
  let mismatches = 0;
  const params = (slot: number) => [
    slot,
    { encoding: "json", maxSupportedTransactionVersion: 0, transactionDetails: "signatures", rewards: false },
  ];

  for (const slot of slots) {
    const expected = await rpc(reference.url, "getBlock", params(slot));
    const t0 = performance.now();
    const actual = await rpc(candidate.url, "getBlock", params(slot));
    latencies.push(performance.now() - t0);
    if (JSON.stringify(actual?.signatures) !== JSON.stringify(expected?.signatures)) mismatches++;
  }

  return {
    endpoint: candidate.name,
    samples: slots.length,
    p50: percentile(latencies, 50),
    p95: percentile(latencies, 95),
    p99: percentile(latencies, 99),
    mismatches, // acceptance: 0 unexplained
  };
}

Where archived raw blocks or account snapshots are a large share of storage, we measure compression on a uniform sample of your data rather than quoting a headline ratio. This uses the published blockchain-compression API, and checks the lossless round-trip on every sample:

// Measure ratio and decompression time on a uniform sample of your own blobs.
use blockchain_compression::core::traits::CompressionStrategy;
use blockchain_compression::presets::solana::{SolanaCompressor, SolanaPreset};
use std::time::Instant;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut compressor = SolanaCompressor::new(SolanaPreset::MaxCompression);
    let (mut raw, mut packed) = (0usize, 0usize);
    let mut decode_us = Vec::new();

    for entry in std::fs::read_dir("samples")? {
        let data = std::fs::read(entry?.path())?;
        let compressed = compressor.compress(&data)?;
        let t0 = Instant::now();
        let restored = compressor.decompress(&compressed)?;
        decode_us.push(t0.elapsed().as_micros());
        assert_eq!(data, restored, "lossless round-trip failed");
        raw += data.len();
        packed += compressed.len();
    }

    decode_us.sort_unstable();
    let p95 = decode_us[(decode_us.len() * 95 / 100).min(decode_us.len() - 1)];
    println!("ratio {:.1}:1, p95 decode {} us", raw as f64 / packed as f64, p95);
    Ok(())
}

The same sample also goes through the codec you already use (for example ZSTD in ClickHouse or Parquet). The number that matters is the marginal gain over that baseline, not the ratio against uncompressed bytes.

Evidence

  • blockchain-compression publishes its preset ranges and a method for measuring them honestly, including why you should budget against the lower bound: measuring ratios that matter, the presets, and the documentation. It is published on crates.io, so you can run the benchmark above yourself.
  • SolanaVault documents its architecture, its self-hosting path, and its stated limitations in the architecture overview, the FAQ, and the documentation. Its 15–25:1 compression figures are its own measurements on Solana blocks and have not been independently reproduced.
  • StreamSync’s sharded DuckDB design and its comparisons with managed providers are on the StreamSync site and in the documentation. Its sub-10ms query latency is a design target, not a published field benchmark, and its operator network is not live.
  • The broader lessons from building Solana data tooling are in Solana infrastructure lessons from eleven projects.
  • Evidence levels for all three projects are listed in the maturity register. SolanaVault and StreamSync have no published third-party audit (blockchain-compression is a library, where an audit does not apply), and none of the three has yet been accepted in a paid delivery.

What drives the cost

  • Number of query classes and chains. Each class needs its own benchmark, correctness check, and latency target.
  • History depth and backfill. Backfilling years of history through RPC or from archive files often costs more time than building the live pipeline.
  • Correctness requirements. Balance-grade or accounting-grade data needs full reconciliation, while trend analytics can be sampled.
  • Number of consumers to migrate. Every service, bot, and dashboard that calls the old endpoint must be repointed and tested.
  • Operating model. Whether your team runs the result or Cryptuon operates it under managed operations changes the runbook, monitoring, and handover scope.
  • Pass-through infrastructure. Provider plans, Geyser-enabled nodes, warehouse compute, and storage are priced separately and modelled in the assessment.

Limitations and what we won’t do

  • We will not quote a saving before measuring your workload. Provider pricing, query mix, and history depth vary too much for a credible estimate from a landing page.
  • Compression ratios from any of our projects are workload-dependent. We report the ratio measured on your sample, alongside the codec you already run.
  • SolanaVault’s peer-to-peer storage network and StreamSync’s operator network are not live. We do not route your production traffic through them, and we do not recommend their tokens or incentive layers.
  • Self-managed is not automatically cheaper. Once on-call, upgrades, and backfills are counted, a managed provider often wins at moderate volume, and the assessment will say so.
  • We cannot guarantee a third-party provider’s uptime, rate limits, or future pricing. The design records these as assumptions and builds fallbacks where you need them.

For related work, see trading infrastructure if your data feeds execution, and data verification and provenance if you need to prove where data came from. Engagement terms are on services, and our open research is on research.

Next step

Start a brief and send us a week of RPC or provider usage exports (or request logs), your current monthly bill, and the queries that hurt most. The assessment returns a workload profile, a benchmark of your current setup against two to three alternatives, and a costed recommendation with the acceptance criteria for any migration.

Technology

Cryptuon technology we may use

Open-source components we can bring to this work. They are used only where testing on your workload supports them, and evidence levels come from the public maturity register.

Alternatives

Options we’d recommend when they fit better

The assessment compares these on your actual workload. If one of them wins, the recommendation says so.

Managed Solana RPC and data providers (Helius, Triton One, QuickNode)

Your volume is moderate, your team is small, and a better plan, caching, or cheaper method choices close the cost gap

Yellowstone gRPC (Geyser) streaming into your own store

You need low-latency, filtered account and transaction feeds and are prepared to own ingestion, backfill, and gap recovery

Self-hosted warehouse (ClickHouse, or Parquet with DuckDB)

Your heavy queries are analytical scans and aggregations that per-call RPC pricing makes expensive

Old Faithful archive or public warehouse datasets

You need deep history occasionally and can tolerate slower retrieval or dataset lag instead of paying for hot archive access

FAQ

Questions buyers ask

How much does it cost to reduce our Solana data costs with Cryptuon?

Most engagements start with a fixed-fee assessment (from $1,500) that profiles your workload and benchmarks options. Implementation sprints start from $12,000, and managed operations from $1,500 per month. Provider fees, cloud consumption, and hardware are passed through or paid by you directly, and are modelled explicitly in the assessment.

Should we self-host an indexer or use a managed provider?

It depends on the query mix, not on ideology. Point lookups and recent-state reads are usually cheapest on a managed provider with good caching. Large analytical scans, long address histories, and repeated backfills are often cheaper in a store you own, fed by a stream. Most teams land on a hybrid. The assessment measures which classes of your traffic belong where.

How can we reduce Solana RPC costs without changing providers?

Common wins are caching immutable data (finalized blocks and transactions never change), removing polling in favour of subscriptions or streams, requesting only the fields you need, and moving bulk history reads off per-call RPC. We measure these first, because changing nothing structural is the cheapest migration.

Do you only use SolanaVault and StreamSync?

No. StreamSync’s operator network and SolanaVault’s peer-to-peer storage network are not live, so neither is offered as a hosted provider. Their components and the blockchain-compression library are used only where benchmarks on your data justify them. Many recommendations involve only managed providers, Geyser streaming, and an off-the-shelf warehouse.

Why is historical Solana data so expensive to store and query?

Solana produces a large volume of blocks and transactions, and RPC nodes keep only recent ledger locally; deeper history is served from separate long-term storage. Per-call pricing then multiplies across paginated reads such as address histories, which return at most 1,000 signatures per request. Cost is driven by how far back you read, how often, and how much of each record you actually need.

Can you guarantee our query latency after a migration?

We commit to measured acceptance criteria rather than guarantees: agreed p50, p95, and p99 targets per query class, at an agreed load, demonstrated before cutover. If we operate the system afterwards, the same percentiles are monitored and reported monthly. We cannot guarantee the uptime of third-party providers in the design.

Bring us the requirement

Describe the outcome, what is blocking it, and when it must work. We reply with what a scoped assessment would cover and cost.