Specialist engineering

Deepfake detection evaluation and verifiable data provenance

“We have to decide whether media and external data are genuine, and share sensitive records with partners, with evidence we can show an auditor later.”

Cryptuon evaluates deepfake detectors and data-attestation options on your permitted test data, then builds provenance pipelines with evidence trails you can audit.

When teams bring us this

  • You must choose a deepfake or synthetic-media detector and have only vendor-reported accuracy to go on
  • Moderation, newsroom, or onboarding decisions now depend on media whose origin you cannot establish
  • A contract or automated workflow needs an external fact (a price, a record, an API response) it can check rather than trust
  • Partners or regulators ask who received shared sensitive data, and your logs can be edited by your own administrators
  • An appeal or audit asked you to reproduce a past verdict, and you could not

Usually owned by

  • Head of trust and safety
  • Platform trust lead
  • Enterprise engineering lead
  • Head of data governance
Scope

What you receive, and how “done” is defined

Deliverables

  • Evaluation plan: permitted test set, slices (modality, compression, generator family), and success criteria written down before testing
  • Side-by-side evaluation of two to four detection or attestation options: precision and recall at threshold, latency percentiles, cost per item
  • Provenance architecture showing where C2PA checks, detection, attestation, and audit anchoring sit in your pipeline, and what is logged where
  • Controlled pilot integrated with your review or data-sharing workflow, producing an evidence record for every decision
  • Pilot report against the agreed criteria, with a go, no-go, or change-option recommendation

Example acceptance criteria

  • Success criteria (threshold, minimum precision and recall, p95 latency, cost per item) are committed to before any option sees the test set
  • Every option is scored on the same permitted test set, and per-item results are delivered so your team can recompute every metric
  • Each pilot decision can be reproduced from its evidence record: input hash, option and model version, score, threshold, and timestamp
  • Audit-trail entries can be verified by a party without write access to your systems (on-chain anchor or signed append-only log, as agreed)

Not included unless scoped

  • Accuracy promises before the pilot; results are reported against the agreed criteria, whatever they turn out to be
  • Moderation policy, takedown decisions, or legal attribution of who created a piece of media
  • Compliance certification (HIPAA, GDPR, SOX) or legal advice on recording personal-data references on a public chain
  • Handling material that is illegal to possess; those pipelines need specialist providers with legal authority
Engagement

How this is bought

Start with a fixed-fee assessment. Every later stage is optional and scoped in writing before it starts.

  1. Step 1 · Assess

    Feasibility or bottleneck assessment

    from $1,500 · 1–2 weeks

    • Current-state analysis against your real workload or codebase
    • Options compared, including ones that do not use Cryptuon technology
    • Risk register and costed, scoped recommendation
  2. Step 2 · Specialist review

    Specialist technical assessment

    from $3,000 · 1–3 weeks

    • Signing, cross-chain, runtime, or production-readiness analysis with a written scope
    • Failure-mode and rehearsal plan for the change you intend to make
    • Evidence pack you can share with your own reviewers
  3. Step 3 · Implement

    Implementation sprint

    from $12,000 · 3–8 weeks

    • The agreed migration or integration, delivered against written acceptance criteria
    • Tests, runbooks, and documentation
    • Handover, or transition into managed operations

Prices exclude independent audits, substantial infrastructure consumption, and legal advice unless written into the scope. All engagement types →

The short answer

Verifying external information means answering one of three questions with evidence: where did this come from (provenance), does this look manipulated (detection), or did this system really say this (data attestation). Cryptuon designs the evaluation first. We agree a permitted test set and fix the success criteria before testing, then score two to four options side by side, including commercial APIs and open standards. Only then do we build a controlled pilot that writes a reproducible evidence record for every decision.

We do not promise detection accuracy in advance. The pilot measures it against the criteria you sign off.

Decision criteria

Which question are you actually answering?

  • Origin: C2PA manifests answer “who captured or edited this”, but only when the producer signed it.
  • Manipulation: detectors give a probability, not proof, and they degrade on generators they were not trained on.
  • External facts: oracle networks, zkTLS, and multi-reporter commit-reveal each prove something different: node agreement, server provenance, and independence respectively.

Many pipelines need two of the three. Choosing the wrong one is the most expensive mistake in this area.

What does each kind of error cost you?

A false positive on a newsroom photo damages credibility. A false negative on an onboarding video lets fraud through. That asymmetry sets the threshold, and the threshold sets precision and recall. We agree the operating point, and decide what goes to a human review queue, before measuring anything.

Can you lawfully use a test set that looks like production?

Vendor benchmarks rarely match your compression, resolution, languages, or generator mix. The evaluation needs labelled media that you have the rights to process and to share with each candidate vendor. Building or licensing that set is often the critical path.

Who must be able to check the evidence, and must it be public?

If only your auditors need to check it, a signed append-only log may be enough. If counterparties or the public must check it without trusting you, anchor hashes on a public chain. Public chains are immutable, though, and even a hash or a pseudonymous identifier can count as personal data. Erasure obligations under GDPR then need a design that keeps personal data off-chain, and that design needs legal review. We flag where the review is needed; we do not give legal advice.

What are the volume and latency budgets?

Synchronous upload screening needs seconds. Batch review of flagged items can take minutes. Multi-operator consensus and on-chain anchoring add latency and per-item cost, so they suit high-stakes decisions better than every upload.

Implementation options

ApproachHow it worksStrengthTradeoffMaturity
DFPN (Cryptuon)Independent operators run their own detectors, commit then reveal results, and a reputation-weighted verdict is anchored on SolanaAuditable verdict from several models; no single vendor controls the scoreAccuracy depends on the operators’ models; adds latency and on-chain costDeclared; no published accuracy benchmark; network not on mainnet
DataMgmt Node (Cryptuon)Encrypts data shared between nodes and records each share as an EVM transaction, checkable via /verify_dataA trail of who shared what, verifiable without access to your systemsSelf-hosted P2P infrastructure; on-chain metadata needs privacy reviewDeclared; not a certified compliance system
C2PA / Content CredentialsSigned manifests attached at capture or edit and verified downstreamOpen standard with broad industry backing; proves origin, not just likelihoodCovers only cooperating producers; manifests can be strippedEstablished standard, adoption still growing
Commercial detection APIsVendor-hosted proprietary models return a manipulation scoreManaged, multimodal, contractual SLAOpaque models; scores not comparable across vendors; per-item pricingProduction
Oracle networks (e.g. Chainlink)A staked node set fetches, aggregates, and posts external data on-chainWidely integrated, long production record for price dataYou trust the node set and its sources; custom data needs custom jobsProduction
zkTLS (e.g. TLSNotary, Reclaim)Proves that a named HTTPS server returned a given responseStrong provenance for web data, no server cooperation neededProves the source said it, not that it is true; tooling is youngEmerging

The delivery sequence

  1. Frame the question. Classify each decision as origin, manipulation, or external fact, and record what an error costs.
  2. Build the permitted test set. Source and label media or records you have the rights to use, sliced by modality, compression, and generator family.
  3. Pre-register. Fix the threshold and the pass criteria, and commit to them before any candidate runs.
  4. Evaluate side by side. Run every candidate on the same set, then publish per-item results and metrics per slice.
  5. Architect. Decide what is checked in what order (typically provenance, then detection, then human review), what is logged, and where the evidence is anchored.
  6. Pilot. Integrate the chosen option with a bounded slice of real workflow and an evidence record for each decision, then report against the criteria.

The pre-registration step uses commit-reveal so that thresholds and labels are sealed while candidates are being scored:

# Seal the pilot's success criteria and label file before any candidate runs.
import hashlib
import json

from commit_reveal import CommitRevealScheme

with open("labels.csv", "rb") as f:
    labels_sha256 = hashlib.sha256(f.read()).hexdigest()

criteria = json.dumps({
    "labels_sha256": labels_sha256,
    "threshold": 0.80,
    "min_precision": 0.90,
    "min_recall": 0.75,
    "max_p95_latency_s": 4.0,
    "max_cost_per_item_usd": 0.05,
}, sort_keys=True)

scheme = CommitRevealScheme()                 # SHA-256 by default
commitment, salt = scheme.commit(criteria)
# Share `commitment` with all parties (or anchor it on-chain). Keep `salt` sealed.

# After scoring: reveal, and anyone can check that nothing moved.
assert scheme.reveal(criteria, salt, commitment)

The acceptance test then runs unchanged against every candidate. In this sketch, detector and permitted_test_set are pytest fixtures that wrap a vendor API, an open model, or a DFPN request; the numbers are placeholders for the criteria you agree.

import statistics
import time


def score_all(detector, items):
    rows = []
    for item in items:
        start = time.perf_counter()
        score, cost_usd = detector.score(item.path)
        rows.append((item.manipulated, score, time.perf_counter() - start, cost_usd))
    return rows


def test_candidate_meets_preregistered_criteria(detector, permitted_test_set):
    rows = score_all(detector, permitted_test_set)
    tp = sum(1 for y, s, _, _ in rows if y and s >= 0.80)
    fp = sum(1 for y, s, _, _ in rows if not y and s >= 0.80)
    fn = sum(1 for y, s, _, _ in rows if y and s < 0.80)
    precision = tp / (tp + fp) if tp + fp else 0.0
    recall = tp / (tp + fn) if tp + fn else 0.0
    p95 = statistics.quantiles([r[2] for r in rows], n=20)[18]
    cost = sum(r[3] for r in rows) / len(rows)

    assert precision >= 0.90, f"precision {precision:.3f}"
    assert recall >= 0.75, f"recall {recall:.3f}"
    assert p95 <= 4.0, f"p95 latency {p95:.2f}s"
    assert cost <= 0.05, f"cost per item ${cost:.4f}"

Evidence

  • DFPN’s verdict lifecycle (submit, route, analyse, commit, reveal, consensus, challenge) and its storage model are documented on how DFPN works. Six adversary classes and their mitigations are set out in the DFPN threat model, and the C2PA comparison and documentation cover the rest.
  • commit-reveal is published on PyPI. Its timing-safe reveal, hash allowlist, and out-of-scope list are stated in the security model and the documentation.
  • DataMgmt Node’s encryption (Fernet with PBKDF2 key derivation), P2P layer, and on-chain share records are described in how DataMgmt Node works and its documentation.
  • How zkTLS, commit-reveal, and replicated inference fit together, and what each does not prove, is covered in verifiable on-chain AI inference and resolution. The same evidence discipline applied to regulated assets is in tokenizing real-world assets.
  • The Decentralized Deepfake Detection Network paper is listed under research. Evidence levels for all three projects are in the maturity register.
  • Not yet evidenced: DFPN has no published accuracy benchmark and is not live on mainnet. None of the three projects has a published third-party audit. DataMgmt Node is not a certified compliance system.

What drives the cost

  • Test-set construction. Licensing, labelling, and slicing permitted data is often the largest single cost.
  • Modalities. Video and audio cost more to evaluate than still images, in compute, vendor fees, and labelling.
  • Number of candidates. Each extra vendor or model adds integration work and pass-through fees.
  • Evidence requirements. An internal signed log costs little. Public anchoring adds transaction costs, key management, and a privacy review.
  • Workflow integration. Wiring results into moderation queues, appeals, or partner data exchanges usually costs more than the evaluation itself.

Limitations and what we won’t do

  • We will not quote detection accuracy before the pilot, and no evaluation proves future performance against generators that do not exist yet.
  • A detector’s verdict is evidence, not proof. We design the human-review and appeals path, but enforcement decisions stay with you.
  • zkTLS proves that a source said something, not that it is true, and oracle agreement is not correctness. The architecture names whose honesty you still rely on.
  • DataMgmt Node can support a compliance programme, but it does not make you compliant. Its on-chain records need legal review against your erasure and residency obligations.
  • Our assessments are engineering reviews, not independent audits. Where an audit is needed, we scope it and bring in a third party, priced separately.

Next step

Tell us which decisions you need to verify, a sample of the inputs (or a description, if they cannot leave your environment), and who must be able to check the evidence. The assessment returns a test-set plan, a shortlist of candidates, the pre-registered criteria, and a costed pilot scope.

Technology

Cryptuon technology we may use

Open-source components we can bring to this work. They are used only where testing on your workload supports them, and evidence levels come from the public maturity register.

Alternatives

Options we’d recommend when they fit better

The assessment compares these on your actual workload. If one of them wins, the recommendation says so.

C2PA / Content Credentials

Your sources (cameras, editing tools, partner publishers) can sign at capture or edit time, and you need origin rather than a manipulation score

Commercial detection APIs (e.g. Reality Defender, Hive)

You want a managed detection service with a vendor contract and SLA, and accept one vendor's proprietary models and scoring

Oracle networks (e.g. Chainlink)

A contract needs widely used external data such as price feeds, and trusting an established decentralised node set is acceptable

zkTLS (e.g. TLSNotary, Reclaim Protocol)

You need to prove that a specific HTTPS endpoint returned a specific response, without that endpoint's cooperation

FAQ

Questions buyers ask

How do we evaluate deepfake detection tools for our platform?

On a test set you are permitted to use that resembles your real traffic, with success criteria fixed in advance: a score threshold, minimum precision and recall at that threshold, p95 latency, and cost per item. Every candidate runs on the same set, and results are broken down by modality, compression, and generator family, because aggregate accuracy hides the slices where detectors fail.

How much does a deepfake detection evaluation or provenance pilot cost?

A scoping assessment starts from $1,500 and a specialist assessment (evaluation design, threat model, provenance architecture) from $3,000. A controlled pilot integrated with your workflow is an implementation sprint, from $12,000. Vendor API fees, GPU time, and on-chain transaction costs during the pilot are pass-through costs, listed separately in the scope.

Do you only use DFPN?

No. DFPN is one option, evaluated alongside commercial detection APIs and open models on the same test set. It has no published accuracy benchmark, and its accuracy depends on the models its operators run. If a commercial API or a C2PA-first pipeline meets your criteria better, the report recommends that.

Can you guarantee a detection accuracy?

No, and you should be wary of anyone who does. Detection is probabilistic, and generators improve faster than any single model updates. We guarantee the method: criteria agreed before testing, the same data for every option, and per-item results you can recompute. The pilot then reports whatever the numbers turn out to be.

Should we use C2PA provenance or deepfake detection?

Usually both. C2PA proves where media came from when the capture device or editor signed it, but most media arrives with no manifest, or with one that was stripped. Detection gives a probabilistic signal for that unsigned majority. The usual design checks provenance first and falls through to detection.

How can a smart contract verify data from an external API?

A contract cannot call an API, so the fact has to be brought on-chain with evidence. An oracle network gives you agreement among a staked node set. zkTLS proves that a named HTTPS server returned a given response. Commit-reveal across independent reporters stops them copying one another. The right choice depends on the data source, latency, and how much a wrong value costs.

Bring us the requirement

Describe the outcome, what is blocking it, and when it must work. We reply with what a scoped assessment would cover and cost.