Sponsored & research

Forecasting and agent-competition application engineering

“We want to build a forecasting or agent-competition application, with question design, resolution, and scoring that hold up to scrutiny — not a demo.”

Cryptuon builds forecasting and agent-competition pilots (question design, resolution, scoring, settlement) for teams or licensed operators. We do not run markets.

When teams bring us this

  • You want an internal forecasting programme (points, no real money) whose questions resolve without argument
  • You are running or sponsoring a competition where AI agents or models are scored against data they must not see
  • A licensed operator wants a new question type, resolution pipeline, or agent interface built against its own rules
  • An earlier pilot stalled on disputed outcomes or ambiguously worded questions
  • An ecosystem programme wants an open reference agent arena or bounty mechanism built and documented

Usually owned by

  • Platform founder or CEO
  • Enterprise innovation or strategy lead
  • Ecosystem or foundation programme lead
Scope

What you receive, and how “done” is defined

Deliverables

  • Question specification: wording rules, named resolution sources, timestamped triggers, fallback, and invalidation clauses
  • Resolution and evaluation framework: how outcomes are decided, challenged, and scored (for example Brier score or hidden-test leaderboards)
  • Settlement design for the chosen model: a points ledger, a licensed operator's systems, or on-chain escrow on testnet
  • Pilot application covering authoring, curation, sealed submission, and scoring, with tests and a runbook
  • Launch-readiness pack listing the regulatory questions your counsel must answer before any real-money or prize use

Example acceptance criteria

  • Every pilot question passes an automated spec check: named source, timestamped trigger, invalidation clause, and no discretionary wording
  • Resolution is reproducible: an independent reviewer re-derives every pilot outcome from the recorded source evidence
  • Sealed submissions cannot be changed after the deadline: every accepted reveal verifies against its commitment, and late or mismatched reveals are rejected in tests
  • Scoring is deterministic: re-running the evaluation on the frozen dataset produces an identical leaderboard

Not included unless scoped

  • Operating a real-money prediction market, sportsbook, or prize game, or holding participant funds
  • Legal or regulatory advice, licence applications, or KYC/AML programme design
  • Market-making, liquidity provision, or funding prize pools
  • Independent security audit of any contract that will hold value (scoped and priced separately)
Engagement

How this is bought

Start with a fixed-fee assessment. Every later stage is optional and scoped in writing before it starts.

  1. Step 1 · Assess

    Feasibility or bottleneck assessment

    from $1,500 · 1–2 weeks

    • Current-state analysis against your real workload or codebase
    • Options compared, including ones that do not use Cryptuon technology
    • Risk register and costed, scoped recommendation
  2. Step 2 · Sponsor

    Sponsored engineering

    from $15,000 · Milestone-based

    • Milestone-based delivery of a capability a buyer or ecosystem explicitly wants funded
    • Reproducible implementation and evaluation report
    • Open-source release where the sponsor agrees

Prices exclude independent audits, substantial infrastructure consumption, and legal advice unless written into the scope. All engagement types →

The short answer

A forecasting or agent-competition application succeeds or fails on three decisions made before any code: how questions are worded, how outcomes are resolved, and what is at stake. Cryptuon designs and builds the pilot around those decisions. That means question specs that resolve mechanically, sealed submissions, deterministic scoring, and a settlement model that fits your legal position. Typical settlement models are a points ledger, a licensed operator’s systems, or testnet escrow.

Cryptuon builds the software. We do not run real-money markets or prize games. If your plan needs a real-money market or prize game, a licensed operator runs it, after legal review.

Decision criteria

The questions that decide cost and approach:

Is real money, or anything of value, at stake?

This is the first fork, and it is a legal question as much as a technical one. Points-only internal forecasting is the simplest case. Prizes, entry fees, or tradeable positions can bring in gambling, derivatives, or financial-promotion rules. For example, the FCA’s cryptoasset financial-promotion regime applies to promotions aimed at UK consumers whatever the firm’s location. Pay-to-play prize mechanics, such as metered entry with a winner-takes-pot payout, need specific legal review in every target jurisdiction before launch. The assessment lists those questions for your counsel. It does not answer them.

Can every question be resolved from a named source?

Questions such as “Will metric X on API Y be at least Z at time T?” resolve mechanically. Questions such as “Will the launch be considered successful?” do not. Mechanical questions can resolve from recorded evidence, and in future from zkTLS attestation where a source supports it. Judgement questions need a bonded challenge process or a named panel. The share of judgement questions in your pipeline largely decides how much resolution infrastructure you need.

Who scores submissions, and can anyone see the answers early?

For agent and model competitions, the main risks are leakage and copying. Leakage means the evaluation set becomes reachable. Copying means a late entrant reads earlier answers. A hidden test split, sealed commit-then-reveal submissions, and a deterministic scorer deal with both. Deterministic scoring also makes every disputed result re-runnable.

Does settlement need to be on-chain?

Often it does not. A points ledger in an ordinary database is enough for an internal programme, and it avoids custody questions entirely. On-chain escrow is useful when participants should not have to trust the organiser to pay out. In that case it should start on testnet and run with real value only under a licensed operator, after an independent audit.

Who curates questions, and how many per week?

AI drafting can produce candidate questions quickly, but a human still has to approve wording before anything opens. The curator workflow, its review time, and its audit trail are a recurring cost, so they belong in the plan.

Implementation options

ApproachHow it worksStrengthTradeoffMaturity
Cryptuon scoped pilotQuestion specs, resolution, and sealed scoring built for your case, reusing Mentat, Sarpoy, and commit-reveal components only where tests support themFits your questions, your settlement model, and your legal constraints; you own the codeYou fund a build; on-chain parts need an independent audit before holding valueMentat and Sarpoy declared (source on GitHub, not on mainnet); commit-reveal reproduced (PyPI); no third-party audit published
UMA optimistic oracleA proposer posts a bonded answer; if nobody disputes within the challenge window it stands, otherwise UMA token holders voteHandles judgement questions on-chain; widely integratedDisputed outcomes take longer; the outcome of an escalation depends on how token holders voteEstablished, in production use
Regulated venue’s APIYour application trades or lists through a licensed exchange’s API under its rulebookThe venue carries the regulatory permissions for real-money contractsLimited to the venue’s jurisdictions, contract types, and termsEstablished
Play-money or reputation forecasting platformHosted questions, points, and leaderboardsFastest route for internal forecasting; no custodyLittle control over resolution logic, data, or agent interfacesEstablished
Kaggle-style competitionModels are scored against public and private leaderboard splits on a hosted platformMature anti-leakage practice and a large participant poolBuilt for offline datasets; no live forecasting or on-chain settlementEstablished

Many good pilots combine options. One common shape is internal forecasting with points, mechanical questions resolved from recorded API evidence, and the few judgement questions sent to a published panel. Prizes come later, and only if counsel clears them.

The delivery sequence

  1. Assess. Review the intended audience, what is at stake, the target jurisdictions, and a sample of 20–50 real questions or tasks.
  2. Specify. Set wording rules and a canonical question schema, and record a resolution source and fallback for every question type.
  3. Choose resolution. Classify each question type as mechanical (recorded evidence, or zkTLS where supported), optimistic oracle, or panel.
  4. Build the pilot. Deliver authoring and curation, sealed submissions, the scorer, and settlement (points ledger, operator integration, or testnet escrow).
  5. Rehearse. Run a closed round end to end, then have an independent reviewer re-derive every outcome and score.
  6. Hand over. Deliver the runbook, curator guide, and the launch-readiness pack with open legal questions.

The spec check below becomes an acceptance test. Its schema is illustrative and modelled on the canonical market schema Mentat publishes:

# Acceptance test: every pilot question must be mechanically resolvable.
# `pilot_questions` is a pytest fixture loading the curated question set.
import re
from datetime import datetime

REQUIRED = ("question", "resolution_criteria", "source_allowlist", "trigger", "invalidation")
DISCRETIONARY = re.compile(
    r"\b(at the discretion of|as determined by|reasonably|substantially|widely reported)\b", re.I
)

def check_spec(spec: dict) -> list[str]:
    errors = [f"missing {key}" for key in REQUIRED if not spec.get(key)]
    trigger_at = (spec.get("trigger") or {}).get("at", "")
    try:
        datetime.fromisoformat(trigger_at.replace("Z", "+00:00"))
    except ValueError:
        errors.append(f"trigger time is not ISO 8601: {trigger_at!r}")
    if DISCRETIONARY.search(spec.get("resolution_criteria", "")):
        errors.append("resolution criteria contain discretionary wording")
    return errors

def test_every_pilot_question_is_resolvable(pilot_questions):
    for spec in pilot_questions:
        assert check_spec(spec) == [], spec["question"]

Sealed submissions use the published commit-reveal API. In this example, ledger stands for your pilot’s own store:

from commit_reveal import CommitRevealScheme

scheme = CommitRevealScheme()  # SHA-256 by default

# Before the deadline: the agent publishes only the commitment.
forecast = "q-0142:YES:0.71"
commitment, salt = scheme.commit(forecast)
ledger.record_commitment(agent_id="agent-7", commitment=commitment)

# After the deadline: accept a reveal only if it matches what was committed.
def accept_reveal(agent_id: str, value: str, salt) -> bool:
    committed = ledger.commitment_for(agent_id)
    return committed is not None and scheme.reveal(value, salt, committed)

Forecast scoring is kept deterministic and simple enough to re-run during a dispute:

def brier_score(forecasts: list[float], outcomes: list[int]) -> float:
    """Mean squared error of probability forecasts; 0 is perfect, lower is better."""
    if not forecasts or len(forecasts) != len(outcomes):
        raise ValueError("forecasts and outcomes must be non-empty and equal length")
    return sum((p - o) ** 2 for p, o in zip(forecasts, outcomes)) / len(outcomes)

Evidence

  • Mentat publishes its pipeline (Scout, Draft, and Validator agents, a curator console, and Anchor factory and settlement programs) and a canonical market schema with source allowlist, trigger, and invalidation fields in how it works. Its FAQ states the status plainly: it is not on mainnet, the Solana programs are in development, and the zkTLS proof service is a later milestone. Documentation: docs.cryptuon.com/mentat.
  • Sarpoy documents a four-instruction Anchor lifecycle with a treasury PDA that only the program can move, in how it works. Note that answer verification runs in the API or a creator-configured oracle before the program pays out, so the verifier is a trust point. Documentation: docs.cryptuon.com/sarpoy.
  • commit-reveal is on PyPI with an explicit threat model. Comparison is timing-safe, but the secp256k1 arithmetic used by the optional proofs is not constant-time. Documentation: docs.cryptuon.com/commit-reveal.
  • The resolution design space is covered by zkTLS, commit-reveal, and replicated inference, and is compared in verifiable on-chain AI inference and resolution.
  • Evidence levels and limitations for each project are in the maturity register.
  • Not yet evidenced: no Cryptuon forecasting or competition system has run in production or on mainnet. No third-party audit has been published for Mentat, Sarpoy, or commit-reveal. zkTLS resolution has not been demonstrated end to end.

What drives the cost

  • What is at stake. Points-only pilots avoid custody, escrow audits, and most of the legal workstream.
  • Number of question types. Each new resolution source needs its own adapter, fallback rule, and evidence capture.
  • Share of judgement questions. These need a challenge process or a panel, plus its tooling.
  • Settlement model. On-chain escrow adds contract work, testnet rehearsal, and an independent audit before real value.
  • Integration with an operator. Building to a licensed operator’s rulebook, APIs, and compliance controls adds review cycles.
  • Curation volume. The number of questions per week sets the size of the curator workflow and its audit trail.

Limitations and what we won’t do

  • We do not operate real-money prediction markets, sportsbooks, or prize games, hold participant funds, or seed pools.
  • We do not give legal advice. Legal review of the specific activity and jurisdictions is required before launch, and the pilot is built so it can stay points-only if counsel says no.
  • Mentat and Sarpoy are reference designs at the declared evidence level. They are not production systems, and any on-chain component needs an independent audit before it holds value.
  • zkTLS resolution depends on the source supporting it. Many official sources will still need recorded evidence, an optimistic oracle, or a panel.
  • If a hosted forecasting platform or a Kaggle-style competition meets the need, the assessment recommends it over a build.

Related work: agent payments and payouts, data verification and provenance, and how we structure engagements.

Next step

Send us the intended audience, what participants stand to win or lose, target jurisdictions, and 20–50 example questions or tasks through /start. The assessment returns a question specification, a recommended resolution and settlement model, and a costed pilot plan. It also lists the legal questions your counsel should answer first.

Technology

Cryptuon technology we may use

Open-source components we can bring to this work. They are used only where testing on your workload supports them, and evidence levels come from the public maturity register.

Alternatives

Options we’d recommend when they fit better

The assessment compares these on your actual workload. If one of them wins, the recommendation says so.

UMA optimistic oracle

You need on-chain resolution of questions that no single API answers, and can accept a bonded challenge window with escalation to a UMA token-holder vote

Build on a regulated venue's API

You want real-money event contracts and a regulated exchange (for example Kalshi, a CFTC-designated contract market) already lists or can list the contract

Hosted play-money or reputation forecasting platform

An internal programme only needs questions, points, and leaderboards, and an existing SaaS or community platform is faster than a build

Kaggle-style evaluation competition

You are scoring models against a hidden test set and want an established host with a large participant pool

FAQ

Questions buyers ask

How much does it cost to build a forecasting or agent-competition pilot?

Most engagements start with a fixed-fee assessment (from $1,500) that settles the question design, resolution model, and settlement approach. A scoped pilot is delivered as milestone-based sponsored engineering, from $15,000. The main cost drivers are whether anything of value is at stake, how many question types need their own resolution sources, and whether settlement is on-chain.

Can Cryptuon run our prediction market or prize competition for us?

No. Cryptuon builds software and pilots: internal forecasting with points and no real money, evaluation competitions, testnet prototypes, or systems for a licensed operator who runs them under its own permissions. We do not operate real-money markets, hold participant funds, or seed prize pools.

Is it legal to run a prediction market or a paid agent competition?

It depends on the activity and on where participants are. Real-money prediction markets and prize games can fall under gambling, derivatives, or financial-promotion rules. For example, the UK FCA's cryptoasset financial-promotion regime applies to marketing aimed at UK consumers wherever the firm is based. Legal review of the specific activity and every target jurisdiction is required before launch. We flag the questions; we do not give legal advice.

Do you only use Mentat?

No. Mentat is a reference design, and it is not live on mainnet. Its zkTLS proof service is a roadmap item. Many pilots are better served by a points ledger and recorded API evidence, by the UMA optimistic oracle, or by building on a regulated venue's API. The assessment compares these against your questions.

How do you resolve forecasting questions without a trusted arbiter?

Mostly by design at drafting time. Every question names its source, trigger time, and invalidation rule, so resolution is mechanical. Mechanical questions resolve from recorded source evidence that anyone can check. Questions that need judgement go to a bonded optimistic oracle such as UMA, or to a named panel with a published process. zkTLS attestation is used only where a source supports it.

How do you stop AI agents from copying each other or gaming a competition?

Submissions are committed before the deadline and revealed after it, so nobody can copy an answer early. Evaluation data stays hidden behind a separate private split. Rate limits or rising per-query costs make brute-force probing expensive. Because a deterministic scorer gets the same result on every run, a disputed result can be checked by running it again.

Bring us the requirement

Describe the outcome, what is blocking it, and when it must work. We reply with what a scoped assessment would cover and cost.