The short answer
Verifying external information means answering one of three questions with evidence: where did this come from (provenance), does this look manipulated (detection), or did this system really say this (data attestation). Cryptuon designs the evaluation first. We agree a permitted test set and fix the success criteria before testing, then score two to four options side by side, including commercial APIs and open standards. Only then do we build a controlled pilot that writes a reproducible evidence record for every decision.
We do not promise detection accuracy in advance. The pilot measures it against the criteria you sign off.
Decision criteria
Which question are you actually answering?
- Origin: C2PA manifests answer “who captured or edited this”, but only when the producer signed it.
- Manipulation: detectors give a probability, not proof, and they degrade on generators they were not trained on.
- External facts: oracle networks, zkTLS, and multi-reporter commit-reveal each prove something different: node agreement, server provenance, and independence respectively.
Many pipelines need two of the three. Choosing the wrong one is the most expensive mistake in this area.
What does each kind of error cost you?
A false positive on a newsroom photo damages credibility. A false negative on an onboarding video lets fraud through. That asymmetry sets the threshold, and the threshold sets precision and recall. We agree the operating point, and decide what goes to a human review queue, before measuring anything.
Can you lawfully use a test set that looks like production?
Vendor benchmarks rarely match your compression, resolution, languages, or generator mix. The evaluation needs labelled media that you have the rights to process and to share with each candidate vendor. Building or licensing that set is often the critical path.
Who must be able to check the evidence, and must it be public?
If only your auditors need to check it, a signed append-only log may be enough. If counterparties or the public must check it without trusting you, anchor hashes on a public chain. Public chains are immutable, though, and even a hash or a pseudonymous identifier can count as personal data. Erasure obligations under GDPR then need a design that keeps personal data off-chain, and that design needs legal review. We flag where the review is needed; we do not give legal advice.
What are the volume and latency budgets?
Synchronous upload screening needs seconds. Batch review of flagged items can take minutes. Multi-operator consensus and on-chain anchoring add latency and per-item cost, so they suit high-stakes decisions better than every upload.
Implementation options
| Approach | How it works | Strength | Tradeoff | Maturity |
|---|---|---|---|---|
| DFPN (Cryptuon) | Independent operators run their own detectors, commit then reveal results, and a reputation-weighted verdict is anchored on Solana | Auditable verdict from several models; no single vendor controls the score | Accuracy depends on the operators’ models; adds latency and on-chain cost | Declared; no published accuracy benchmark; network not on mainnet |
| DataMgmt Node (Cryptuon) | Encrypts data shared between nodes and records each share as an EVM transaction, checkable via /verify_data | A trail of who shared what, verifiable without access to your systems | Self-hosted P2P infrastructure; on-chain metadata needs privacy review | Declared; not a certified compliance system |
| C2PA / Content Credentials | Signed manifests attached at capture or edit and verified downstream | Open standard with broad industry backing; proves origin, not just likelihood | Covers only cooperating producers; manifests can be stripped | Established standard, adoption still growing |
| Commercial detection APIs | Vendor-hosted proprietary models return a manipulation score | Managed, multimodal, contractual SLA | Opaque models; scores not comparable across vendors; per-item pricing | Production |
| Oracle networks (e.g. Chainlink) | A staked node set fetches, aggregates, and posts external data on-chain | Widely integrated, long production record for price data | You trust the node set and its sources; custom data needs custom jobs | Production |
| zkTLS (e.g. TLSNotary, Reclaim) | Proves that a named HTTPS server returned a given response | Strong provenance for web data, no server cooperation needed | Proves the source said it, not that it is true; tooling is young | Emerging |
The delivery sequence
- Frame the question. Classify each decision as origin, manipulation, or external fact, and record what an error costs.
- Build the permitted test set. Source and label media or records you have the rights to use, sliced by modality, compression, and generator family.
- Pre-register. Fix the threshold and the pass criteria, and commit to them before any candidate runs.
- Evaluate side by side. Run every candidate on the same set, then publish per-item results and metrics per slice.
- Architect. Decide what is checked in what order (typically provenance, then detection, then human review), what is logged, and where the evidence is anchored.
- Pilot. Integrate the chosen option with a bounded slice of real workflow and an evidence record for each decision, then report against the criteria.
The pre-registration step uses commit-reveal so that thresholds and labels are sealed while candidates are being scored:
# Seal the pilot's success criteria and label file before any candidate runs.
import hashlib
import json
from commit_reveal import CommitRevealScheme
with open("labels.csv", "rb") as f:
labels_sha256 = hashlib.sha256(f.read()).hexdigest()
criteria = json.dumps({
"labels_sha256": labels_sha256,
"threshold": 0.80,
"min_precision": 0.90,
"min_recall": 0.75,
"max_p95_latency_s": 4.0,
"max_cost_per_item_usd": 0.05,
}, sort_keys=True)
scheme = CommitRevealScheme() # SHA-256 by default
commitment, salt = scheme.commit(criteria)
# Share `commitment` with all parties (or anchor it on-chain). Keep `salt` sealed.
# After scoring: reveal, and anyone can check that nothing moved.
assert scheme.reveal(criteria, salt, commitment)
The acceptance test then runs unchanged against every candidate. In this sketch, detector and permitted_test_set are pytest fixtures that wrap a vendor API, an open model, or a DFPN request; the numbers are placeholders for the criteria you agree.
import statistics
import time
def score_all(detector, items):
rows = []
for item in items:
start = time.perf_counter()
score, cost_usd = detector.score(item.path)
rows.append((item.manipulated, score, time.perf_counter() - start, cost_usd))
return rows
def test_candidate_meets_preregistered_criteria(detector, permitted_test_set):
rows = score_all(detector, permitted_test_set)
tp = sum(1 for y, s, _, _ in rows if y and s >= 0.80)
fp = sum(1 for y, s, _, _ in rows if not y and s >= 0.80)
fn = sum(1 for y, s, _, _ in rows if y and s < 0.80)
precision = tp / (tp + fp) if tp + fp else 0.0
recall = tp / (tp + fn) if tp + fn else 0.0
p95 = statistics.quantiles([r[2] for r in rows], n=20)[18]
cost = sum(r[3] for r in rows) / len(rows)
assert precision >= 0.90, f"precision {precision:.3f}"
assert recall >= 0.75, f"recall {recall:.3f}"
assert p95 <= 4.0, f"p95 latency {p95:.2f}s"
assert cost <= 0.05, f"cost per item ${cost:.4f}"
Evidence
- DFPN’s verdict lifecycle (submit, route, analyse, commit, reveal, consensus, challenge) and its storage model are documented on how DFPN works. Six adversary classes and their mitigations are set out in the DFPN threat model, and the C2PA comparison and documentation cover the rest.
- commit-reveal is published on PyPI. Its timing-safe reveal, hash allowlist, and out-of-scope list are stated in the security model and the documentation.
- DataMgmt Node’s encryption (Fernet with PBKDF2 key derivation), P2P layer, and on-chain share records are described in how DataMgmt Node works and its documentation.
- How zkTLS, commit-reveal, and replicated inference fit together, and what each does not prove, is covered in verifiable on-chain AI inference and resolution. The same evidence discipline applied to regulated assets is in tokenizing real-world assets.
- The Decentralized Deepfake Detection Network paper is listed under research. Evidence levels for all three projects are in the maturity register.
- Not yet evidenced: DFPN has no published accuracy benchmark and is not live on mainnet. None of the three projects has a published third-party audit. DataMgmt Node is not a certified compliance system.
What drives the cost
- Test-set construction. Licensing, labelling, and slicing permitted data is often the largest single cost.
- Modalities. Video and audio cost more to evaluate than still images, in compute, vendor fees, and labelling.
- Number of candidates. Each extra vendor or model adds integration work and pass-through fees.
- Evidence requirements. An internal signed log costs little. Public anchoring adds transaction costs, key management, and a privacy review.
- Workflow integration. Wiring results into moderation queues, appeals, or partner data exchanges usually costs more than the evaluation itself.
Limitations and what we won’t do
- We will not quote detection accuracy before the pilot, and no evaluation proves future performance against generators that do not exist yet.
- A detector’s verdict is evidence, not proof. We design the human-review and appeals path, but enforcement decisions stay with you.
- zkTLS proves that a source said something, not that it is true, and oracle agreement is not correctness. The architecture names whose honesty you still rely on.
- DataMgmt Node can support a compliance programme, but it does not make you compliant. Its on-chain records need legal review against your erasure and residency obligations.
- Our assessments are engineering reviews, not independent audits. Where an audit is needed, we scope it and bring in a third party, priced separately.
Next step
Tell us which decisions you need to verify, a sample of the inputs (or a description, if they cannot leave your environment), and who must be able to check the evidence. The assessment returns a test-set plan, a shortlist of candidates, the pre-registered criteria, and a costed pilot scope.