The short answer
The risk in changing validator infrastructure is rarely the new signer itself. It is the transition: a stale protection export, a standby that comes up while the primary is still live, or a protection database imported into the wrong backend. Cryptuon reviews the planned change, rehearses it on a test network with your client versions, drills the failure modes, and wires refusal and availability monitoring into your on-call before anything touches production keys.
We do not ask you to replace your signer on the strength of a web page. Often the right outcome is your existing signer, with a better runbook.
Decision criteria
The questions that decide cost and approach:
Where can each key sign today?
Every validator client, remote signer, backup host, warm standby, and “temporary” migration box that holds a key or can reach a signer is a place a conflicting signature can come from. The topology map comes first, because most slashing incidents trace back to a second signing path nobody counted.
Is the change single-writer or multi-writer?
Moving from one client to one remote signer is a sequential handover: stop, export, import, start. Adding failover is a different problem. Two signers must either share one authoritative protection store (for example, a shared PostgreSQL database, which both Web3Signer and nklave support) or be fenced so that only one can sign. Threshold designs such as Dirk or distributed validators replace the question with cluster-level coordination, which brings its own failure modes.
What does EIP-3076 cover, and what does it not?
EIP-3076 defines a JSON interchange format for slashing-protection history and the conditions a signer must respect after importing it. It lets you move protection between clients and signers. It does not:
- stop two independently configured signers from being live for the same key;
- address inactivity leaks or offline penalties, which come from not signing;
- protect beyond what the export contains: an export taken while the old client was still signing, or a missing key, leaves a gap.
Referencing EIP-3076 is not proof of protection. The rehearsal is.
What is your tolerance for missed duties during cutover?
A safe cutover deliberately accepts a short gap: a few missed attestations cost far less than one slashable message. We agree that budget in epochs up front, so nobody is tempted to overlap signers to save a few minutes.
Which custody backend and chains are involved?
Local keystores, YubiHSM, cloud HSM, and cloud KMS each change the runbook. Cosmos and CometBFT validators have their own double-sign rules and tombstoning, and need their own drills.
Implementation options
| Approach | How it works | Strength | Tradeoff | Maturity |
|---|---|---|---|---|
| nklave (Cryptuon) | Remote signer speaking the Web3Signer HTTP protocol; requests pass an ordered policy chain (EIP-3076 rules, fork allowlist, rate limit) before reaching the key; decisions go to a Merkle-checkpointed audit log | Explicit allow or refuse per request with reason codes; tamper-evident log; Ethereum and CometBFT in one signer | Young project; fewer key backends than Web3Signer; existing protection history must be migrated and rehearsed per setup | Reproduced (public source); no third-party audit published |
| Client built-in protection | Each validator client keeps its own protection database; EIP-3076 export and import for moves; doppelganger detection on start | Simplest, with no extra service; supported by all major clients | Protection lives in the signing process; no shared state for failover | Established |
| Consensys Web3Signer | Reference remote signer; PostgreSQL slashing-protection database; many key-store backends | Broadest client integration and operator familiarity | JVM service to run; few configurable signing policies beyond slashing protection | Established |
| Attestant Dirk + Vouch | Distributed signer with threshold signing across Dirk instances, driven by the Vouch validator client over gRPC | In threshold mode, no single signer instance holds the full signing capability | Adopts Attestant’s toolchain; not a drop-in for Web3Signer-style clients | Established |
| DVT (Obol, SSV) | Validator key split across a cluster of operators that co-sign with threshold BLS | Tolerates node and operator failure without a single key | Cluster coordination, added latency, new operational and trust model | In production on mainnet |
Teams with a working setup often keep their current signer and add what the review shows is missing: a rehearsed runbook, fencing for standbys, and refusal monitoring.
The delivery sequence
- Map the topology. List every key, every process that can sign for it, and every protection database involved.
- Review failure modes. For the planned change, enumerate what can produce a conflicting signature or a long outage, and how each will be tested.
- Build the rehearsal environment. Use a test network with test keys, and the same client, signer, and custody versions as production.
- Rehearse the migration. Stop, export, validate, import, run a refusal test, then start. Repeat until the runbook is boring.
- Drill failover. Kill the primary, partition the network, restart the standby with stale state, and confirm the single-writer invariant holds.
- Integrate policy and monitoring. Alert on refusals, signer unavailability, and protection-store health. With nklave, also verify the audit log (
nklave log verify). - Cut over production. Follow the rehearsed runbook within the agreed missed-duty budget, then hand over or move into managed operations.
The pre-import gate is a script, not a checklist item. This one validates an EIP-3076 file against the network you expect, the keys you expect, and the current epoch:
# Gate before importing an EIP-3076 interchange file. Exit non-zero to block the cutover.
import json
import sys
from collections import defaultdict
def check(path: str, expected_gvr: str, expected_pubkeys: set[str], current_epoch: int,
max_lag_epochs: int = 2) -> list[str]:
with open(path) as f:
doc = json.load(f)
problems = []
meta = doc["metadata"]
if meta["interchange_format_version"] != "5":
problems.append(f"unexpected format version {meta['interchange_format_version']}")
if meta["genesis_validators_root"].lower() != expected_gvr.lower():
problems.append("genesis_validators_root mismatch: export is from a different network")
targets = defaultdict(list)
for record in doc["data"]:
key = record["pubkey"].lower()
targets[key] += [int(a["target_epoch"]) for a in record.get("signed_attestations", [])]
for key in sorted(expected_pubkeys - targets.keys()):
problems.append(f"{key}: no protection history in export")
for key, epochs in targets.items():
if not epochs:
problems.append(f"{key}: empty attestation history; confirm the key never signed")
elif current_epoch - max(epochs) > max_lag_epochs:
problems.append(f"{key}: newest target epoch {max(epochs)} is stale; was the old signer stopped first?")
return problems
if __name__ == "__main__":
path, gvr, keys_file, epoch = sys.argv[1:5]
expected = {line.strip().lower() for line in open(keys_file) if line.strip()}
issues = check(path, gvr, expected, int(epoch))
print("\n".join(issues) or "interchange file passed")
sys.exit(1 if issues else 0)
The refusal drill runs on a test network against the signer’s Web3Signer-style endpoint. The two request bodies are fixtures for the same validator and target epoch with different block roots, so the second is a double vote:
# Drill (test network only): sign attestation A, then a conflicting attestation B
# for the same target epoch. B must be refused. Web3Signer signals a slashing
# refusal with HTTP 412; other signers may use a different non-200 status.
import json
import requests
SIGNER = "http://signer.test.internal:9000"
PUBKEY = open("fixtures/test_pubkey.txt").read().strip()
def sign(body_file: str) -> requests.Response:
with open(body_file) as f:
body = json.load(f)
return requests.post(f"{SIGNER}/api/v1/eth2/sign/{PUBKEY}", json=body, timeout=5)
first = sign("fixtures/attestation_a.json")
assert first.status_code == 200, f"baseline signing failed: {first.status_code}"
conflict = sign("fixtures/attestation_b_same_target.json")
assert conflict.status_code != 200, "DOUBLE VOTE WAS SIGNED: stop the rollout"
print(f"refused with {conflict.status_code}: {conflict.text[:200]}")
Refusals in production are rare, and each one deserves a human. For nklave, the project documents a nklave_policy_refusals_total metric, which maps to a Prometheus rule like this:
groups:
- name: signer-safety
rules:
- alert: SignerPolicyRefusal
expr: increase(nklave_policy_refusals_total[10m]) > 0
labels:
severity: page
annotations:
summary: "Signer refused a signing request"
description: "Check the audit log for the refusal reason before restarting or failing over any validator component."
Evidence
- nklave’s design, policy chain, and audit log are documented on its features and how it works pages and in the documentation, including a published threat model that lists what it does not protect against.
- Head-to-head comparisons that say when the alternative is the better fit: nklave vs Web3Signer and nklave vs Dirk.
- The migration sequence and failure modes for EIP-3076 import are written up in EIP-3076 import and export during validator migration.
- Background on slashable offences and past incidents is in validator slashing prevention.
- nklave is listed at evidence level reproduced in the maturity register. Not yet evidenced: no third-party security audit is published, and nklave has not yet been accepted in a paid production migration.
What drives the cost
- Number of keys, clients, and sites. Each combination of client and custody backend needs its own rehearsal.
- Failover design. A sequential migration is far simpler than adding HA, threshold signing, or DVT.
- Custody backends. HSM and KMS integrations add provisioning, access-control review, and latency testing.
- Chains. Ethereum and Cosmos or CometBFT validators have different slashing rules and need separate drills.
- Evidence requirements. If your customers or reviewers need a formal evidence pack, drills are recorded and packaged accordingly.
Limitations and what we won’t do
- An assessment is an engineering review, not an independent audit. Where you need one, we scope it with a third-party auditor.
- No signer, including nklave, prevents every staking penalty. Inactivity leaks, offline penalties, and slashing caused by components outside the reviewed setup remain your operational risk.
- We will not recommend replacing a working signer unless rehearsal shows a concrete improvement, and we will not run production cutovers that skip the rehearsed steps.
- nklave does not enforce restaking or AVS-specific slashing conditions out of the box.
- We never take custody of keys or seed phrases.
Related work: cross-chain integration for bridge and relayer signing, and the full engagement terms on services.
Next step
Start a brief and send us a description of your current signing setup (clients, signers, custody, hosts) and the change you intend to make. The specialist assessment returns a topology map, a failure-mode review, and a rehearsal plan with the acceptance criteria for the migration.