The short answer
A forecasting or agent-competition application succeeds or fails on three decisions made before any code: how questions are worded, how outcomes are resolved, and what is at stake. Cryptuon designs and builds the pilot around those decisions. That means question specs that resolve mechanically, sealed submissions, deterministic scoring, and a settlement model that fits your legal position. Typical settlement models are a points ledger, a licensed operator’s systems, or testnet escrow.
Cryptuon builds the software. We do not run real-money markets or prize games. If your plan needs a real-money market or prize game, a licensed operator runs it, after legal review.
Decision criteria
The questions that decide cost and approach:
Is real money, or anything of value, at stake?
This is the first fork, and it is a legal question as much as a technical one. Points-only internal forecasting is the simplest case. Prizes, entry fees, or tradeable positions can bring in gambling, derivatives, or financial-promotion rules. For example, the FCA’s cryptoasset financial-promotion regime applies to promotions aimed at UK consumers whatever the firm’s location. Pay-to-play prize mechanics, such as metered entry with a winner-takes-pot payout, need specific legal review in every target jurisdiction before launch. The assessment lists those questions for your counsel. It does not answer them.
Can every question be resolved from a named source?
Questions such as “Will metric X on API Y be at least Z at time T?” resolve mechanically. Questions such as “Will the launch be considered successful?” do not. Mechanical questions can resolve from recorded evidence, and in future from zkTLS attestation where a source supports it. Judgement questions need a bonded challenge process or a named panel. The share of judgement questions in your pipeline largely decides how much resolution infrastructure you need.
Who scores submissions, and can anyone see the answers early?
For agent and model competitions, the main risks are leakage and copying. Leakage means the evaluation set becomes reachable. Copying means a late entrant reads earlier answers. A hidden test split, sealed commit-then-reveal submissions, and a deterministic scorer deal with both. Deterministic scoring also makes every disputed result re-runnable.
Does settlement need to be on-chain?
Often it does not. A points ledger in an ordinary database is enough for an internal programme, and it avoids custody questions entirely. On-chain escrow is useful when participants should not have to trust the organiser to pay out. In that case it should start on testnet and run with real value only under a licensed operator, after an independent audit.
Who curates questions, and how many per week?
AI drafting can produce candidate questions quickly, but a human still has to approve wording before anything opens. The curator workflow, its review time, and its audit trail are a recurring cost, so they belong in the plan.
Implementation options
| Approach | How it works | Strength | Tradeoff | Maturity |
|---|---|---|---|---|
| Cryptuon scoped pilot | Question specs, resolution, and sealed scoring built for your case, reusing Mentat, Sarpoy, and commit-reveal components only where tests support them | Fits your questions, your settlement model, and your legal constraints; you own the code | You fund a build; on-chain parts need an independent audit before holding value | Mentat and Sarpoy declared (source on GitHub, not on mainnet); commit-reveal reproduced (PyPI); no third-party audit published |
| UMA optimistic oracle | A proposer posts a bonded answer; if nobody disputes within the challenge window it stands, otherwise UMA token holders vote | Handles judgement questions on-chain; widely integrated | Disputed outcomes take longer; the outcome of an escalation depends on how token holders vote | Established, in production use |
| Regulated venue’s API | Your application trades or lists through a licensed exchange’s API under its rulebook | The venue carries the regulatory permissions for real-money contracts | Limited to the venue’s jurisdictions, contract types, and terms | Established |
| Play-money or reputation forecasting platform | Hosted questions, points, and leaderboards | Fastest route for internal forecasting; no custody | Little control over resolution logic, data, or agent interfaces | Established |
| Kaggle-style competition | Models are scored against public and private leaderboard splits on a hosted platform | Mature anti-leakage practice and a large participant pool | Built for offline datasets; no live forecasting or on-chain settlement | Established |
Many good pilots combine options. One common shape is internal forecasting with points, mechanical questions resolved from recorded API evidence, and the few judgement questions sent to a published panel. Prizes come later, and only if counsel clears them.
The delivery sequence
- Assess. Review the intended audience, what is at stake, the target jurisdictions, and a sample of 20–50 real questions or tasks.
- Specify. Set wording rules and a canonical question schema, and record a resolution source and fallback for every question type.
- Choose resolution. Classify each question type as mechanical (recorded evidence, or zkTLS where supported), optimistic oracle, or panel.
- Build the pilot. Deliver authoring and curation, sealed submissions, the scorer, and settlement (points ledger, operator integration, or testnet escrow).
- Rehearse. Run a closed round end to end, then have an independent reviewer re-derive every outcome and score.
- Hand over. Deliver the runbook, curator guide, and the launch-readiness pack with open legal questions.
The spec check below becomes an acceptance test. Its schema is illustrative and modelled on the canonical market schema Mentat publishes:
# Acceptance test: every pilot question must be mechanically resolvable.
# `pilot_questions` is a pytest fixture loading the curated question set.
import re
from datetime import datetime
REQUIRED = ("question", "resolution_criteria", "source_allowlist", "trigger", "invalidation")
DISCRETIONARY = re.compile(
r"\b(at the discretion of|as determined by|reasonably|substantially|widely reported)\b", re.I
)
def check_spec(spec: dict) -> list[str]:
errors = [f"missing {key}" for key in REQUIRED if not spec.get(key)]
trigger_at = (spec.get("trigger") or {}).get("at", "")
try:
datetime.fromisoformat(trigger_at.replace("Z", "+00:00"))
except ValueError:
errors.append(f"trigger time is not ISO 8601: {trigger_at!r}")
if DISCRETIONARY.search(spec.get("resolution_criteria", "")):
errors.append("resolution criteria contain discretionary wording")
return errors
def test_every_pilot_question_is_resolvable(pilot_questions):
for spec in pilot_questions:
assert check_spec(spec) == [], spec["question"]
Sealed submissions use the published commit-reveal API. In this example, ledger stands for your pilot’s own store:
from commit_reveal import CommitRevealScheme
scheme = CommitRevealScheme() # SHA-256 by default
# Before the deadline: the agent publishes only the commitment.
forecast = "q-0142:YES:0.71"
commitment, salt = scheme.commit(forecast)
ledger.record_commitment(agent_id="agent-7", commitment=commitment)
# After the deadline: accept a reveal only if it matches what was committed.
def accept_reveal(agent_id: str, value: str, salt) -> bool:
committed = ledger.commitment_for(agent_id)
return committed is not None and scheme.reveal(value, salt, committed)
Forecast scoring is kept deterministic and simple enough to re-run during a dispute:
def brier_score(forecasts: list[float], outcomes: list[int]) -> float:
"""Mean squared error of probability forecasts; 0 is perfect, lower is better."""
if not forecasts or len(forecasts) != len(outcomes):
raise ValueError("forecasts and outcomes must be non-empty and equal length")
return sum((p - o) ** 2 for p, o in zip(forecasts, outcomes)) / len(outcomes)
Evidence
- Mentat publishes its pipeline (Scout, Draft, and Validator agents, a curator console, and Anchor factory and settlement programs) and a canonical market schema with source allowlist, trigger, and invalidation fields in how it works. Its FAQ states the status plainly: it is not on mainnet, the Solana programs are in development, and the zkTLS proof service is a later milestone. Documentation: docs.cryptuon.com/mentat.
- Sarpoy documents a four-instruction Anchor lifecycle with a treasury PDA that only the program can move, in how it works. Note that answer verification runs in the API or a creator-configured oracle before the program pays out, so the verifier is a trust point. Documentation: docs.cryptuon.com/sarpoy.
- commit-reveal is on PyPI with an explicit threat model. Comparison is timing-safe, but the secp256k1 arithmetic used by the optional proofs is not constant-time. Documentation: docs.cryptuon.com/commit-reveal.
- The resolution design space is covered by zkTLS, commit-reveal, and replicated inference, and is compared in verifiable on-chain AI inference and resolution.
- Evidence levels and limitations for each project are in the maturity register.
- Not yet evidenced: no Cryptuon forecasting or competition system has run in production or on mainnet. No third-party audit has been published for Mentat, Sarpoy, or commit-reveal. zkTLS resolution has not been demonstrated end to end.
What drives the cost
- What is at stake. Points-only pilots avoid custody, escrow audits, and most of the legal workstream.
- Number of question types. Each new resolution source needs its own adapter, fallback rule, and evidence capture.
- Share of judgement questions. These need a challenge process or a panel, plus its tooling.
- Settlement model. On-chain escrow adds contract work, testnet rehearsal, and an independent audit before real value.
- Integration with an operator. Building to a licensed operator’s rulebook, APIs, and compliance controls adds review cycles.
- Curation volume. The number of questions per week sets the size of the curator workflow and its audit trail.
Limitations and what we won’t do
- We do not operate real-money prediction markets, sportsbooks, or prize games, hold participant funds, or seed pools.
- We do not give legal advice. Legal review of the specific activity and jurisdictions is required before launch, and the pilot is built so it can stay points-only if counsel says no.
- Mentat and Sarpoy are reference designs at the declared evidence level. They are not production systems, and any on-chain component needs an independent audit before it holds value.
- zkTLS resolution depends on the source supporting it. Many official sources will still need recorded evidence, an optimistic oracle, or a panel.
- If a hosted forecasting platform or a Kaggle-style competition meets the need, the assessment recommends it over a build.
Related work: agent payments and payouts, data verification and provenance, and how we structure engagements.
Next step
Send us the intended audience, what participants stand to win or lose, target jurisdictions, and 20–50 example questions or tasks through /start. The assessment returns a question specification, a recommended resolution and settlement model, and a costed pilot plan. It also lists the legal questions your counsel should answer first.