Specialist engineering

Embedded EVM runtime evaluation, integration, and performance engineering

“We need an EVM execution engine inside our own system. Which one fits our compatibility and performance requirements — and can someone prove it on our workload?”

Cryptuon tests EVM engines (revm, evmone, geth, Zig-EVM) on your own workload, then embeds and tunes the one that meets your fork and throughput needs.

When teams bring us this

  • Execution, not consensus or I/O, is the measured bottleneck in your sequencer, simulator, or indexer
  • You need EVM execution inside a Python, Node.js, C, or Go host and the obvious engine is written in a different language
  • Your rollup or appchain needs custom precompiles, a modified gas schedule, or execution tracing the stock client does not expose
  • A vendor or internal team is quoting parallel-execution speed-ups and nobody has checked them against your transaction mix
  • A target hard fork is coming and you need to know whether your embedded engine supports it

Usually owned by

  • Rollup or L2 engineering lead
  • Developer-platform lead
  • Research engineering lead
Scope

What you receive, and how “done” is defined

Deliverables

  • Requirements matrix: target fork, required opcodes and precompiles, state backend, tracing, host language, and latency vs throughput priority
  • Engine comparison run on your own transaction traces, with a benchmark report a third party can reproduce
  • Differential compatibility suite that compares the chosen engine against a reference engine on state root, gas, logs, and return data
  • The embedding itself (bindings, state-backend adapter, tracing hooks, custom precompiles) or targeted performance changes to an existing integration
  • Benchmark harness in CI with regression gates, plus a runbook for fork upgrades

Example acceptance criteria

  • Zero unexplained divergences from the reference engine across the agreed fixture set: Ethereum execution-spec state tests for the target fork plus your recorded transactions
  • Throughput and p99 latency meet the agreed targets at your measured conflict rate, not only on a zero-conflict batch
  • The benchmark report reproduces from a pinned commit, container image, and documented hardware within an agreed tolerance
  • The embedded build passes an agreed soak test with no memory growth and clean AddressSanitizer runs on the FFI boundary

Not included unless scoped

  • Building a chain: consensus, networking, mempool, or a validator client
  • Independent security audit of the engine or your precompiles (we scope one and bring in an auditor, priced separately)
  • Committing to a speed-up figure before your workload has been measured
Engagement

How this is bought

Start with a fixed-fee assessment. Every later stage is optional and scoped in writing before it starts.

  1. Step 1 · Specialist review

    Specialist technical assessment

    from $3,000 · 1–3 weeks

    • Signing, cross-chain, runtime, or production-readiness analysis with a written scope
    • Failure-mode and rehearsal plan for the change you intend to make
    • Evidence pack you can share with your own reviewers
  2. Step 2 · Implement

    Implementation sprint

    from $12,000 · 3–8 weeks

    • The agreed migration or integration, delivered against written acceptance criteria
    • Tests, runbooks, and documentation
    • Handover, or transition into managed operations

Prices exclude independent audits, substantial infrastructure consumption, and legal advice unless written into the scope. All engagement types →

The short answer

Choosing an embedded EVM is a compatibility decision first and a performance decision second. Cryptuon lists the opcodes, precompiles, fork rules, and host-language constraints your system actually needs. It then runs the candidate engines (revm, evmone, geth’s interpreter, and, where parallel batch execution matters, Zig-EVM) on your own transactions, and integrates the one that passes, against acceptance criteria you sign off first.

Often the honest answer is revm or evmone. We say so when it is.

Decision criteria

The questions that decide cost and approach:

Which fork, opcodes, and precompiles must you support?

An embedded engine is only useful if it runs your bytecode correctly. Contracts compiled with recent Solidity versions can emit opcodes introduced in Shanghai (PUSH0) and Cancun (MCOPY, TLOAD/TSTORE), and many protocols depend on precompiles such as ecrecover, modexp, and the BN254 pairing. We extract the opcode and precompile set from your deployed bytecode and traces and check each candidate engine against it. A gap is not always disqualifying: if the missing opcode never appears in your workload, it can be documented instead of fixed.

Is your bottleneck latency or batch throughput?

These favour different engines. A simulator answering one eth_call at a time needs low per-transaction latency, which is where revm and evmone are strong. A sequencer or backtester that executes thousands of transactions per block can benefit from parallel scheduling, but only if those transactions rarely touch the same state.

What is your conflict rate?

For parallel execution, this one number matters more than anything else. It is the fraction of transactions in a batch that read or write state written by another transaction in the same batch. Near 0%, speed-up approaches the core count minus scheduling overhead. At high contention it approaches 1x, and a parallel scheduler can be marginally slower than sequential execution because it does the same work plus conflict detection. We measure this on your recorded traces before anyone quotes a speed-up.

What language is the host, and who maintains the boundary?

A Rust sequencer embeds revm natively. A Go stack already has geth’s interpreter. A Python research platform or a Node.js service needs an FFI boundary, and someone has to own memory safety, error mapping, and version pinning across it. That ongoing ownership is often the larger cost.

Implementation options

ApproachHow it worksStrengthTradeoffMaturity
Zig-EVM (Cryptuon)Embeddable Zig engine; dependency analysis groups independent transactions into waves run on a work-stealing pool with speculative execution and rollback; C ABI with Python, Rust, JS, and C bindingsBuilt-in parallel batch execution and multi-language embedding from one core96+ opcodes, with partial coverage relative to the latest fork; slower per opcode than revm/evmone in its own comparison; no published auditOpen source (MIT), public test suite; evidence level reproduced
revmRust EVM library with a pluggable database and inspector interfaceCurrent fork coverage, passes the Ethereum state tests, used in production tooling such as Reth and FoundryRust-first API; non-Rust hosts must build and maintain their own FFI; no built-in parallel schedulerEstablished, widely deployed
evmoneC++ EVM implementing the EVMC interfaceLow per-opcode latency; standard C ABI via EVMCYou supply state and host logic through EVMC; single-threadedEstablished, Ethereum Foundation–maintained
go-ethereum core/vmThe reference client’s interpreter, used as a Go libraryBehaviour matches the most widely run client; natural for Go stacks and geth forksTied to geth’s state database and types; LGPL-3.0 library licence; awkward to embed outside GoReference implementation
py-evmPure-Python EVMReadable, easy to instrument for research and testsOrders of magnitude too slow for throughput workEstablished for testing and research

Many integrations end up mixed. revm or geth stays the reference engine for correctness, and a second engine is used only where it measurably helps, checked continuously against the reference.

The delivery sequence

  1. Requirements. Record the target fork, host language, state backend, tracing needs, and whether latency or throughput is the priority.
  2. Capture the workload. Export representative blocks or batches from your system and measure their conflict rate and opcode and precompile mix.
  3. Shortlist and compatibility check. Run the Ethereum execution-spec state tests for your fork plus your recorded transactions through each candidate, and diff the results against a reference engine.
  4. Benchmark. Measure the engines that pass, using the reporting standard below, across a range of conflict rates rather than a single point.
  5. Integrate. Build the bindings, state adapter, tracing hooks, and any custom precompiles. Add differential tests and benchmark regression gates to CI.
  6. Hand over. Deliver the report, the harness, and a fork-upgrade runbook so your team can re-run everything without us.

Differential testing is the acceptance test that matters most. A sketch of the harness (illustrative: each adapter wraps whichever engine is under evaluation, through its own bindings):

# Differential test: every fixture must produce identical observable results
# on the candidate engine and the reference engine. Illustrative harness.
from dataclasses import dataclass
from typing import Protocol

@dataclass(frozen=True)
class Outcome:
    state_root: bytes
    gas_used: int
    logs: tuple
    return_data: bytes
    status: int

class Engine(Protocol):
    name: str
    def run(self, fixture: dict) -> Outcome: ...

def diff(reference: Engine, candidate: Engine, fixtures: list[dict]) -> list[str]:
    failures = []
    for fx in fixtures:
        ref, got = reference.run(fx), candidate.run(fx)
        for field in ("state_root", "gas_used", "logs", "return_data", "status"):
            if getattr(ref, field) != getattr(got, field):
                failures.append(f"{fx['name']}: {field} differs ({candidate.name})")
    return failures

def test_candidate_matches_reference(reference, candidate, fixtures):
    assert diff(reference, candidate, fixtures) == []

Every benchmark we report is described by a pinned configuration like this, so a third party can rerun it:

# Benchmark manifest — one per published result (illustrative values).
[engine]
name    = "zig-evm"
commit  = "<git sha>"
build   = "zig build -Doptimize=ReleaseFast"

[hardware]
cpu       = "<model>"
cores     = 8
memory_gb = 32
host      = "bare metal"   # or the cloud instance type

[workload]
source        = "client-recorded blocks 19,000,000-19,000,500"
batch_size    = 1000
conflict_rate = [0.0, 0.05, 0.10, 0.25, 0.50, 0.85]
threads       = [1, 2, 4, 8]

[report]
runs        = 10
warmup_runs = 2
metrics     = ["throughput_tps", "p50_latency_ms", "p99_latency_ms", "speedup_vs_sequential"]
reproduce   = "make bench MANIFEST=bench/client-mix.toml"

Evidence

  • Zig-EVM’s architecture, execution pipeline, and FFI surface are documented on the project site and in the Zig-EVM documentation. You can step through bytecode in the playground before any engagement.
  • The parallel-execution design and its known constraints (conservative serialisation of transactions with dynamic storage access, overhead on batches under about 100 transactions, Cancun changes still on the roadmap) are written up in how we built a parallel execution engine in Zig. That article reports 1,000 independent transactions executing in 162 ms versus 970 ms sequentially on an 8-core machine.
  • The benchmarking method, including why conflict rate must be disclosed and why ratio speed-ups cannot be compared with end-to-end chain TPS, is set out in parallel EVM benchmarks 2026.
  • Zig-EVM’s own comparison with revm and evmone states that both lead on single-threaded per-opcode performance.
  • Not yet evidenced: the 5-6x figure is Cryptuon’s own measurement on low-conflict batches and has not been independently replicated. There is no published third-party audit, and no published execution-spec conformance run for the latest fork. The current evidence level is in the maturity register.

What drives the cost

  • Number of engines compared. Each needs a harness adapter and a benchmark run.
  • Fork and precompile gaps. Implementing missing opcodes or precompiles in a candidate engine is real engineering, with its own differential tests.
  • State backend integration. Connecting an engine to your database, trie, or snapshot format is usually more work than the execution call itself.
  • Custom precompiles and gas rules. These need specifications, tests, and usually an independent review.
  • Host-language boundary. Python and Node.js embeddings need memory-safety testing, error mapping, and packaging for your deployment targets.

Limitations and what we won’t do

  • We will not quote a speed-up before measuring your conflict rate. Workloads dominated by shared state gain little from any parallel scheduler.
  • Zig-EVM’s opcode and precompile coverage is partial relative to the latest hard fork. If your bytecode needs something it lacks, we either recommend another engine or price the gap explicitly.
  • An engineering review is not an independent audit. Engines that will process value in production, and any custom precompile, should be audited separately.
  • We do not build chains or consensus. If your real bottleneck is state I/O, networking, or data availability, the assessment says so and points to blockchain data costs or cross-chain integration work instead. Open-ended runtime research can be funded through sponsored research engineering. See /services for engagement terms and /research for our published work.

Next step

Send us your host language, target fork, a sample of recorded transactions or blocks, and the throughput or latency you need. The specialist assessment returns a compatibility matrix, a reproducible comparison of the candidate engines on your workload, and a costed integration plan, or you can start here.

Technology

Cryptuon technology we may use

Open-source components we can bring to this work. They are used only where testing on your workload supports them, and evidence levels come from the public maturity register.

Alternatives

Options we’d recommend when they fit better

The assessment compares these on your actual workload. If one of them wins, the recommendation says so.

revm (Rust)

Your host is Rust, or you need broad, current fork coverage and a large production user base (it is the engine behind Reth and Foundry)

evmone (C++)

You want low single-threaded per-opcode latency and a standard EVMC interface from C or C++

go-ethereum core/vm

Your stack is already a Go client or geth fork and you accept geth's LGPL-3.0 library licence and its state model

py-evm

You need a readable Python reference for research or test tooling and throughput does not matter

FAQ

Questions buyers ask

Which EVM implementation should we embed: revm, evmone, geth, or Zig-EVM?

It depends on your host language, target fork, and whether you need latency or batch throughput. revm is the default choice for Rust hosts and current fork coverage; evmone suits C/C++ hosts that want low per-opcode latency; geth's core/vm suits Go stacks; Zig-EVM is worth testing when you need parallel batch execution behind a C ABI. The assessment measures the shortlisted engines on your own transactions before recommending one.

How much does an EVM runtime evaluation cost?

A specialist assessment starts from $3,000 and returns the requirements matrix, a reproducible engine comparison on your workload, and a recommendation. Embedding or performance work runs as an implementation sprint from $12,000. Cost rises with the number of engines compared, custom precompiles, and how much state-backend integration is needed.

Do you only use Zig-EVM?

No. Zig-EVM is one candidate. revm and evmone are faster per opcode on single-threaded work, according to Zig-EVM's own published comparison, and revm has wider fork coverage. If one of them fits better, we recommend it and integrate it.

Is Zig-EVM production-ready?

Not for every workload. It has a public test suite and FFI bindings, but opcode and precompile coverage is partial relative to the latest hard fork, and no third-party audit has been published. Execution-spec conformance, differential fuzzing, and an audit are open roadmap items. We check coverage against your bytecode before proposing it.

Will parallel EVM execution speed up our workload?

Only as far as your transactions are independent. Zig-EVM reports a 5-6x gain over sequential execution on low-conflict batches with up to 8 threads; batches dominated by one hot AMM pool or a single sender’s nonce chain fall towards 1-2x, and small batches can run slower than sequential. We measure your conflict rate first and give you a curve rather than a single number.

What should a credible EVM benchmark report include?

The engine version and commit, the workload and how it was generated, the hardware and core count, the conflict rate, latency percentiles (at least p50 and p99) alongside throughput, the number of runs and their variance, and a script someone else can use to reproduce it. We treat a single TPS figure without these as unverified.

Bring us the requirement

Describe the outcome, what is blocking it, and when it must work. We reply with what a scoped assessment would cover and cost.