Core delivery

Validator signer migration and signing-safety engineering

“We need to change our validator infrastructure without introducing signing risk.”

Cryptuon reviews, rehearses, and executes validator signer changes with failure drills, policy, and monitoring, so migrations do not create conflicting signatures.

When teams bring us this

  • You are moving keys from validator-client keystores to a remote signer, or between remote signers
  • You are changing validator client, hosting provider, region, or HSM and KMS backend
  • You want automatic failover for signing, and someone has asked how you avoid two signers being live at once
  • You are evaluating distributed validator technology (Obol, SSV) or threshold signing
  • An incident, near miss, or customer due-diligence questionnaire has exposed gaps in your signing controls

Usually owned by

  • Validator operations lead
  • Staking infrastructure lead
  • CTO
Scope

What you receive, and how “done” is defined

Deliverables

  • Signing-topology map: every process, host, backup, and standby that could produce a signature for each key
  • Failure-mode review of the planned change, with a risk register and the rehearsal plan that tests each risk
  • Rehearsed migration runbook (stop, export, verify, import, refusal test, start), executed on a test network with your client versions
  • Signing policy and monitoring integration: refusal alerts, signer availability, and protection-database checks wired into your on-call
  • Evidence pack of drill results, logs, and verification output you can share with your own reviewers

Example acceptance criteria

  • Protection-history completeness: every migrated key is present in the target protection database, with bounds no lower than the newest epoch and slot in the source export, verified by script before any start
  • Refusal drills: 100% of scripted double-vote, surround-vote, and double-proposal requests on the test network are refused, with the reason recorded
  • Single-writer invariant: in every failover drill, no two signers with independent protection state are enabled for the same key, as shown by signer logs
  • Cutover cost: missed duties during the production cutover stay within the agreed budget of epochs, measured from beacon data
  • Alerting: refusal, signer-unavailable, and stale-protection alerts fire during the drills and reach the on-call rota

Not included unless scoped

  • Custody of keys or seed phrases; keys stay in your keystores, HSM, or KMS, and we never ask for them
  • Independent security audit of any signer, including nklave (we can scope one with a third-party auditor, priced separately)
  • Guarantees against staking penalties, including inactivity leaks, offline penalties, and slashing from causes outside the reviewed setup
  • Restaking or AVS-specific slashing conditions, unless explicitly scoped
Engagement

How this is bought

Start with a fixed-fee assessment. Every later stage is optional and scoped in writing before it starts.

  1. Step 1 · Specialist review

    Specialist technical assessment

    from $3,000 · 1–3 weeks

    • Signing, cross-chain, runtime, or production-readiness analysis with a written scope
    • Failure-mode and rehearsal plan for the change you intend to make
    • Evidence pack you can share with your own reviewers
  2. Step 2 · Implement

    Implementation sprint

    from $12,000 · 3–8 weeks

    • The agreed migration or integration, delivered against written acceptance criteria
    • Tests, runbooks, and documentation
    • Handover, or transition into managed operations
  3. Step 3 · Operate

    Managed operations and maintenance

    from $1,500/month · Monthly, 3-month minimum

    • Defined monitoring, maintenance, and incident responsibilities
    • Supported changes within an agreed envelope
    • Provider management and monthly reporting

Prices exclude independent audits, substantial infrastructure consumption, and legal advice unless written into the scope. All engagement types →

The short answer

The risk in changing validator infrastructure is rarely the new signer itself. It is the transition: a stale protection export, a standby that comes up while the primary is still live, or a protection database imported into the wrong backend. Cryptuon reviews the planned change, rehearses it on a test network with your client versions, drills the failure modes, and wires refusal and availability monitoring into your on-call before anything touches production keys.

We do not ask you to replace your signer on the strength of a web page. Often the right outcome is your existing signer, with a better runbook.

Decision criteria

The questions that decide cost and approach:

Where can each key sign today?

Every validator client, remote signer, backup host, warm standby, and “temporary” migration box that holds a key or can reach a signer is a place a conflicting signature can come from. The topology map comes first, because most slashing incidents trace back to a second signing path nobody counted.

Is the change single-writer or multi-writer?

Moving from one client to one remote signer is a sequential handover: stop, export, import, start. Adding failover is a different problem. Two signers must either share one authoritative protection store (for example, a shared PostgreSQL database, which both Web3Signer and nklave support) or be fenced so that only one can sign. Threshold designs such as Dirk or distributed validators replace the question with cluster-level coordination, which brings its own failure modes.

What does EIP-3076 cover, and what does it not?

EIP-3076 defines a JSON interchange format for slashing-protection history and the conditions a signer must respect after importing it. It lets you move protection between clients and signers. It does not:

  • stop two independently configured signers from being live for the same key;
  • address inactivity leaks or offline penalties, which come from not signing;
  • protect beyond what the export contains: an export taken while the old client was still signing, or a missing key, leaves a gap.

Referencing EIP-3076 is not proof of protection. The rehearsal is.

What is your tolerance for missed duties during cutover?

A safe cutover deliberately accepts a short gap: a few missed attestations cost far less than one slashable message. We agree that budget in epochs up front, so nobody is tempted to overlap signers to save a few minutes.

Which custody backend and chains are involved?

Local keystores, YubiHSM, cloud HSM, and cloud KMS each change the runbook. Cosmos and CometBFT validators have their own double-sign rules and tombstoning, and need their own drills.

Implementation options

ApproachHow it worksStrengthTradeoffMaturity
nklave (Cryptuon)Remote signer speaking the Web3Signer HTTP protocol; requests pass an ordered policy chain (EIP-3076 rules, fork allowlist, rate limit) before reaching the key; decisions go to a Merkle-checkpointed audit logExplicit allow or refuse per request with reason codes; tamper-evident log; Ethereum and CometBFT in one signerYoung project; fewer key backends than Web3Signer; existing protection history must be migrated and rehearsed per setupReproduced (public source); no third-party audit published
Client built-in protectionEach validator client keeps its own protection database; EIP-3076 export and import for moves; doppelganger detection on startSimplest, with no extra service; supported by all major clientsProtection lives in the signing process; no shared state for failoverEstablished
Consensys Web3SignerReference remote signer; PostgreSQL slashing-protection database; many key-store backendsBroadest client integration and operator familiarityJVM service to run; few configurable signing policies beyond slashing protectionEstablished
Attestant Dirk + VouchDistributed signer with threshold signing across Dirk instances, driven by the Vouch validator client over gRPCIn threshold mode, no single signer instance holds the full signing capabilityAdopts Attestant’s toolchain; not a drop-in for Web3Signer-style clientsEstablished
DVT (Obol, SSV)Validator key split across a cluster of operators that co-sign with threshold BLSTolerates node and operator failure without a single keyCluster coordination, added latency, new operational and trust modelIn production on mainnet

Teams with a working setup often keep their current signer and add what the review shows is missing: a rehearsed runbook, fencing for standbys, and refusal monitoring.

The delivery sequence

  1. Map the topology. List every key, every process that can sign for it, and every protection database involved.
  2. Review failure modes. For the planned change, enumerate what can produce a conflicting signature or a long outage, and how each will be tested.
  3. Build the rehearsal environment. Use a test network with test keys, and the same client, signer, and custody versions as production.
  4. Rehearse the migration. Stop, export, validate, import, run a refusal test, then start. Repeat until the runbook is boring.
  5. Drill failover. Kill the primary, partition the network, restart the standby with stale state, and confirm the single-writer invariant holds.
  6. Integrate policy and monitoring. Alert on refusals, signer unavailability, and protection-store health. With nklave, also verify the audit log (nklave log verify).
  7. Cut over production. Follow the rehearsed runbook within the agreed missed-duty budget, then hand over or move into managed operations.

The pre-import gate is a script, not a checklist item. This one validates an EIP-3076 file against the network you expect, the keys you expect, and the current epoch:

# Gate before importing an EIP-3076 interchange file. Exit non-zero to block the cutover.
import json
import sys
from collections import defaultdict

def check(path: str, expected_gvr: str, expected_pubkeys: set[str], current_epoch: int,
          max_lag_epochs: int = 2) -> list[str]:
    with open(path) as f:
        doc = json.load(f)
    problems = []
    meta = doc["metadata"]
    if meta["interchange_format_version"] != "5":
        problems.append(f"unexpected format version {meta['interchange_format_version']}")
    if meta["genesis_validators_root"].lower() != expected_gvr.lower():
        problems.append("genesis_validators_root mismatch: export is from a different network")

    targets = defaultdict(list)
    for record in doc["data"]:
        key = record["pubkey"].lower()
        targets[key] += [int(a["target_epoch"]) for a in record.get("signed_attestations", [])]

    for key in sorted(expected_pubkeys - targets.keys()):
        problems.append(f"{key}: no protection history in export")
    for key, epochs in targets.items():
        if not epochs:
            problems.append(f"{key}: empty attestation history; confirm the key never signed")
        elif current_epoch - max(epochs) > max_lag_epochs:
            problems.append(f"{key}: newest target epoch {max(epochs)} is stale; was the old signer stopped first?")
    return problems

if __name__ == "__main__":
    path, gvr, keys_file, epoch = sys.argv[1:5]
    expected = {line.strip().lower() for line in open(keys_file) if line.strip()}
    issues = check(path, gvr, expected, int(epoch))
    print("\n".join(issues) or "interchange file passed")
    sys.exit(1 if issues else 0)

The refusal drill runs on a test network against the signer’s Web3Signer-style endpoint. The two request bodies are fixtures for the same validator and target epoch with different block roots, so the second is a double vote:

# Drill (test network only): sign attestation A, then a conflicting attestation B
# for the same target epoch. B must be refused. Web3Signer signals a slashing
# refusal with HTTP 412; other signers may use a different non-200 status.
import json
import requests

SIGNER = "http://signer.test.internal:9000"
PUBKEY = open("fixtures/test_pubkey.txt").read().strip()

def sign(body_file: str) -> requests.Response:
    with open(body_file) as f:
        body = json.load(f)
    return requests.post(f"{SIGNER}/api/v1/eth2/sign/{PUBKEY}", json=body, timeout=5)

first = sign("fixtures/attestation_a.json")
assert first.status_code == 200, f"baseline signing failed: {first.status_code}"

conflict = sign("fixtures/attestation_b_same_target.json")
assert conflict.status_code != 200, "DOUBLE VOTE WAS SIGNED: stop the rollout"
print(f"refused with {conflict.status_code}: {conflict.text[:200]}")

Refusals in production are rare, and each one deserves a human. For nklave, the project documents a nklave_policy_refusals_total metric, which maps to a Prometheus rule like this:

groups:
  - name: signer-safety
    rules:
      - alert: SignerPolicyRefusal
        expr: increase(nklave_policy_refusals_total[10m]) > 0
        labels:
          severity: page
        annotations:
          summary: "Signer refused a signing request"
          description: "Check the audit log for the refusal reason before restarting or failing over any validator component."

Evidence

What drives the cost

  • Number of keys, clients, and sites. Each combination of client and custody backend needs its own rehearsal.
  • Failover design. A sequential migration is far simpler than adding HA, threshold signing, or DVT.
  • Custody backends. HSM and KMS integrations add provisioning, access-control review, and latency testing.
  • Chains. Ethereum and Cosmos or CometBFT validators have different slashing rules and need separate drills.
  • Evidence requirements. If your customers or reviewers need a formal evidence pack, drills are recorded and packaged accordingly.

Limitations and what we won’t do

  • An assessment is an engineering review, not an independent audit. Where you need one, we scope it with a third-party auditor.
  • No signer, including nklave, prevents every staking penalty. Inactivity leaks, offline penalties, and slashing caused by components outside the reviewed setup remain your operational risk.
  • We will not recommend replacing a working signer unless rehearsal shows a concrete improvement, and we will not run production cutovers that skip the rehearsed steps.
  • nklave does not enforce restaking or AVS-specific slashing conditions out of the box.
  • We never take custody of keys or seed phrases.

Related work: cross-chain integration for bridge and relayer signing, and the full engagement terms on services.

Next step

Start a brief and send us a description of your current signing setup (clients, signers, custody, hosts) and the change you intend to make. The specialist assessment returns a topology map, a failure-mode review, and a rehearsal plan with the acceptance criteria for the migration.

Technology

Cryptuon technology we may use

Open-source components we can bring to this work. They are used only where testing on your workload supports them, and evidence levels come from the public maturity register.

Alternatives

Options we’d recommend when they fit better

The assessment compares these on your actual workload. If one of them wins, the recommendation says so.

Validator-client built-in slashing protection

One client per key on a single host is enough, and you mainly need a clean client or host change using EIP-3076 export and import plus doppelganger detection

Consensys Web3Signer

You want the reference remote signer, broad client integration, PostgreSQL-backed slashing protection, and the widest range of key-store backends

Attestant Dirk with Vouch

You want threshold signing across several signer instances and are comfortable adopting Attestant's gRPC-based toolchain

Distributed validator technology (Obol, SSV)

You want to remove the single key and single operator as points of failure, and can accept cluster coordination, added latency, and a new operational model

FAQ

Questions buyers ask

How much does a validator signer migration review cost?

A specialist assessment of the planned change, covering topology, failure modes, and a rehearsal plan, starts from $3,000. A rehearsed migration delivered as an implementation sprint starts from $12,000, and managed monitoring and operations from $1,500 per month. Cost scales with the number of keys, clients, sites, and custody backends.

How do we test failover without risking conflicting signatures?

Rehearse on a test network with the same client and signer versions, using test keys. Script the conflicting requests (double vote, surround vote, double proposal) and confirm each is refused. Then prove that only one signer with authoritative protection state can be live per key at any time. Production failover is enabled only after those drills pass.

Does EIP-3076 protect us from all slashing and penalties?

No. EIP-3076 defines an interchange format for slashing-protection data and the conditions a signer must respect after importing it. It moves protection history between clients; it does not stop two independently configured signers from signing at once, it does not address inactivity leaks or offline penalties, and it only protects as well as the exported history is complete and correctly imported.

How do we migrate a slashing protection database safely?

Stop the old signer first, export after it has stopped, validate the file (network, key coverage, recency), import into the backend you will actually run in production, test that a known-conflicting request is refused, and only then start the validator client. Accept a few missed attestations rather than overlapping signers.

Do you only use nklave?

No. nklave is one option. Client built-in protection, Web3Signer, Dirk with Vouch, and distributed validator technology from Obol or SSV are all compared on your setup, and many reviews recommend keeping your current signer with better runbooks, policy, and monitoring. We do not ask anyone to replace critical signing infrastructure without a rehearsal that supports it.

Will you hold or manage our validator keys?

No. We do not take custody of keys or seed phrases. Keys stay in your keystores, HSM, or KMS, and production steps are executed by your team or under access you control, following the rehearsed runbook.

Bring us the requirement

Describe the outcome, what is blocking it, and when it must work. We reply with what a scoped assessment would cover and cost.