Start with a question. Specialist agents plan, experiment and argue.
Only an exact verifier decides what counts.

Recorded mode replays a real run with no API key needed. Live mode starts new agents on the server.

Hackathon Sponsors

Every result,
verified.

Each card is a claim an agent proposed and the verifier accepted, with an exact certificate. Press play to replay the run that produced it.

Watch it
think.

A lead agent splits the question. The planner weighs options by expected gain and verifier cost. The researcher runs the experiment. The verifier issues exact certificates, the red team tries to break them, and surprises reopen the plan.

Open the lab →
record.jsonl · replayrunning

    Proof, not
    persuasion.

    A language model can sound right and be wrong. Here, nothing is accepted on the agent's word. A claim is either backed by an exact certificate (rational arithmetic, symbolic positivity, Lean) or it stays a hypothesis.

    Rules no agent
    can break.

    One gatekeeper

    Only the researcher may call the verifier, and only the verifier can mark a claim as true. Agents generate; code decides.

    Policies in the loop

    Omnigent guardrails block data leaks and resubmissions, and pause for a human before anything is published.

    Built-in opposition

    A red-team agent attacks every result. Contradictions and surprises reopen the plan instead of being ignored.

    Questions

    Ask something
    worth proving.

    Open the lab