Skip to content

Repository files navigation

Recourse

test

A promise that costs something to break.

Recourse is a dispute right for the un-negotiated machine payment: an agent pays an endpoint it has never dealt with, for one call, with nothing signed and no relationship on either side. The whole contract is one sentence the seller published on its own. That sentence is a unilateral act, so a buyer that has never heard of Recourse is still protected by it, and a seller is bound by nothing it did not write.

That is the scope, stated as a definition rather than a comparison: not the call where two parties agreed structured, machine-readable terms before any money moved, but the one where nobody agreed anything and the payment happened anyway. Any endpoint can claim to be good; only one that has put a promise in escrow can prove it, and pays when it is wrong.

x402 settles that payment in milliseconds and finally. Once settlement confirms there is no chargeback path and no dispute window. Recourse holds the payment for a short window, lets the buying agent contest it, and has GenLayer validators rule on three frozen strings and the chain's own record of when they arrived.

Agents can spend money in milliseconds. Nothing in the stack lets them get it back.

Every claim above that rests on something outside this repository, and every one on the site, is traced to its source in docs/SOURCES.md.

Run one yourself, with nothing installed. userecourse.xyz/app runs one real contested cycle on Studio Next. It reads the demo seller's promise from the escrow, pays, shows the response beside the buyer's three checks with their numbers, opens the dispute with the bond, and reads the adjudicate transaction until the committee's verdict is written, every hash linked to the explorer. With no wallet the server pays, from an account it creates for that run and forgets; with a wallet you sign each step yourself, and the page funds it from Studio's faucet. The verdict is written on chain and the settlement then does not move: on Studio Next the escrow keeps the payment and the bond, for the reason Settlement on Studio Next gives, and the page says so before anything is paid. Runs are limited to ten an hour per visitor, with no cap across visitors.

Run the demo

pip install -r requirements.txt pytest   # Python 3.12, genlayer-py 0.19.0rc2
# and genvm-lint on PATH: scripts/test.py lints and validates the contracts with it
python scripts/prepare.py                # funds three accounts, registers a seller
python scripts/demo.py                   # both paths, on Studio Next

Everything here runs on Studio Next, the network the hackathon requires and the one the site reads. Both paths run: the honest one, which adds no latency and costs nobody anything, and the contested one, in which an agent pays, receives a nine hour old price, contests it, and has the verdict written to its case with no human in the loop.

The demo ends without a refund, and says so. On Studio Next the verdict is written and the settlement does not move. Consensus v0.6 funds a value transfer only at the root of the transaction's fee allocation tree, and the transfer that settle emits sits two messages below the transaction that funds it, open_dispute, then adjudicate, then settle, so it is never funded and the escrow keeps the payment and the bond. The last line the demo prints is settlement not moved: the escrow still holds the payment and the bond, and Settlement on Studio Next has every transaction. The refund itself is on the record: on studionet, where these contracts were first frozen, the same contested path returned the payment and the bond to the buyer, and evidence/snapshot.json holds each of those settlements with its transaction. --network studionet on any script exists for one purpose, reproducing that record from an interpreter with genlayer-py 0.16.3, and is offered nowhere as a place to run against.

prepare.py deploys nothing. The pair is frozen at the addresses in contracts/FROZEN.json and every chain number published here is tied to a recorded pair, so a clone gets accounts of its own, funded from the Studio faucet, and a seller among them registered on the escrow. deploy.py refuses to run while the freeze stands, and says what to do instead. A clean clone, in a fresh virtual environment, was walked this way on 11 September, on studionet, then the default, installed with pip install genlayer_py anthropic pytest. It needed genvm-lint on the path and, for scripts/test.py, pytest, which is why both are named above.

docs/SCRIPT.md is the ninety second recording script: shot by shot, timed, every spoken number one this repository publishes. It films Studio Next, where the verdict is recorded and the settlement does not move, and says so on camera rather than cutting around it.

The live feed on Studio Next, the network the site opens on, read from the chain when the page opened: eleven payments, seven disputes opened, four of six not honored, and the latest payments beneath, each contested one a case that links to its own page

The promise linter refusing. The promise typed into it is the one payment p-000014 ran on chain on studionet, and stage 1 answers without asking a model, as the line under the verdict says: not judgeable, because nothing in it is measurable

Measured as medians over every dispute on each network's public record, in evidence/snapshot-studio-next.json and evidence/snapshot.json, and the demo prints its own two lines for each run. On Studio Next, the network the site reads, the verdict is written to the case and the settlement does not move; on studionet the settlement completes:

Studio Next   dispute to verdict written   37 seconds
              dispute to money back        does not move on this runtime
studionet     dispute to verdict           67 seconds
              dispute to money back        100 seconds

On studionet the verdict reaches the escrow first. The money follows once the settlement transaction finalizes, and a transaction there finalizes a median of 30 seconds after its committee accepts it. That ordering is deliberate: judgment starts on acceptance and money moves on finalization. Paying out on acceptance would be faster and would mean a successful appeal could reverse a verdict after the money had already gone. The honest number for "money back" is therefore the longer one, and it is the one the demo prints last.

How it works

                     +---------------------------+
                     |   SELLER ENDPOINT (off)   |
                     |   /quote  4 modes         |
                     |   signs sha256(response)  |
                     +-------------+-------------+
                                   |  response + signature
                                   v
   +----------------+      +-------+---------+      +---------------------+
   |  BUYER AGENT   |----->|  RecourseEscrow |<---->|  RecourseDispute    |
   |  (off chain)   | pay  |  deterministic  | call |  2 model calls      |
   |  checks promise|      |  holds funds    |      |  3 frozen strings   |
   |  files dispute |      |  records evid.  |      |  1 narrow question  |
   +----------------+      +-------+---------+      +----------+----------+
                                   |                           |
                                   |  reads                    | verdict
                                   v                           v
                     +---------------------------+    verdict, then settle
                     |   PUBLIC FEED (off)       |
                     |   Next.js + genlayer-js   |
                     +---------------------------+
  1. The seller registers and publishes a delivery promise in plain language.
  2. The buyer pays. Funds enter escrow, not the seller balance.
  3. The response is delivered instantly and recorded on chain with the seller's signature over its hash. No judgment runs in this path, so no latency is added.
  4. A settlement window runs. If nobody contests, the seller withdraws.
  5. To contest, the buyer posts a bond. Validators receive the promise, the request and the response, and answer one question.
  6. The verdict is written and the settlement it implies is emitted with it. That settlement is its own transaction, so the money follows the verdict rather than landing beside it, and the feed shows both. Every receipt is public.

The three failure modes

Each returns HTTP 200, settles payment, and passes every deterministic check that exists today.

Mode What arrives
stale Correct shape, expired content. A price with a timestamp hours old.
hollow Well formed, carrying nothing. An empty result set returned as success.
substituted Answers a different question than the one paid for.

The three verdicts

Verdict Payment Bond Seller record
honored to seller to seller upheld unchanged
not honored to buyer to buyer upheld plus one
unclear to seller to buyer upheld unchanged

The unclear verdict exists so the system is never forced to manufacture certainty about a promise written too loosely to judge. A losing dispute must cost the buyer something, or contesting everything becomes free; but an unclear verdict is the promise's fault rather than the buyer's, so taking the bond there would punish a buyer for a seller's vague wording.

Verdict quality

Eighteen disputes with their correct verdicts, committed in b50757f, which added the case file and one README and nothing else. The judgment contract arrived in the next commit, e5750e3, and the case file has been modified in no commit since, on any branch, so no expected answer was ever edited to match a run. Check it in three commands:

git log --oneline --all -- eval/cases.json   # one commit, b50757f
git show --name-status b50757f               # two files, neither is code
git log --oneline --diff-filter=A -- contracts/dispute.py   # e5750e3, next

--diff-filter=A is doing work in the third command: without it git answers with the most recent commit to touch the file, which is a later fix and reads like a contradiction.

Read the figures below for what each one counts: how often the judge matched an answer fixed before it ran, how often its three runs agreed, how often it declined to guess, and how it did on cases it was never tuned against. The first three are measured on the set the question was narrowed against, so the last line is the one that says how the judge meets a dispute it has not seen.

accuracy    17/18    matched the verdict committed before the run
stability   17/18    all three runs of a case agreed with each other
unclear      3/18    landed on unclear, which is the honesty signal

held out     1/3     three further cases, answers committed before the
                     runner could read them, never tuned against

The README, the site and both reports print those two accuracy figures together, and the section after next says why. On Studio Next the same two sets score 16/18 and 2/3, and Two networks sets the columns side by side.

Commit order shows when a file was committed, not when it was written, and eval/RESULTS.md says so under "What this evidence does and does not show".

Corrections, and what caught each one

A project that reports no mistakes is a project nobody checked. Each row is a claim this file or the site made that was wrong, or would have flattered if left alone, and names what caught it, because the catching mechanism is the part worth trusting:

the claim, as it stood what caught it
17 of 18, measured on the set the judgment question had already been narrowed against. Three held out cases, committed alone before the runner could read them. They score 1 of 3. The next section is that story, and the README, the site and both reports print the two figures together.
Case 12 could have been narrowed away. Its committed answer is unclear, the judge answers not_honored, and narrowing the question a third time, against this case, was the obvious way to make it pass. The rule that an evaluation case is never edited to make a run succeed. It is published as a miss instead, because narrowing the question against the cases that remain is fitting the prompt to the set.
About a dollar an adjudication, inherited from published examples of a differently shaped contract. Trying to measure it. studionet charges nothing and gasUsed is the limit echoed back rather than work done, so no dollar figure could have come from this deployment. What replaced it is countable: ten model calls a dispute.
"median settlement" on the feed, which measured payment to dispute and never measured settlement. Reading what the two chain timestamps are. A case's opened_at and decided_at are one message's fixed datetime, so they cannot see how long judgment took. The tile was relabelled rather than removed: the number was real, its name was not.
Two transaction hashes in this file, typed rather than read off the chain. Review caught them before the gate ran, and tests/direct/test_snapshot.py is what would have caught them at it: every hash this README cites must exist in evidence/snapshot.json. They were replaced with hashes read out of the snapshot, and all ten now verify against recorded chain data. The commit that put this table here, 3e45df4, records it.
89 seconds from dispute to money back, a figure from the ordering the contracts no longer use. A number audit that opened each figure's source: no file held it. Both timings above are now medians the snapshot computes from the chain's own timestamps, and a test fails if this file or the site types one.
The direct test count, stated as 231. The same audit. The suite had grown, so scripts/test.py now reads the count off pytest and fails when this file disagrees.
Every number here measured against the frozen pair. The same audit. The two evaluation scores ran on instances of their own. They run the same bytes, which scripts/verify.py now reads back and checks.
A judgment prompt of 1771 to 1961 characters. Rebuilding the prompt with the frozen contract's own build_prompt over every case gives 1814 to 2004. A test measures it now.
No consensus on the honest path. Reading the honest cycle's receipts: every write, a pay included, is voted on by a committee of five. What the honest path skips is judgment, and every page says so now.
A clean clone needed only genvm-lint. A clean clone in a fresh virtual environment, walked by an agent that had never seen the repository: scripts/test.py also needed pytest, which the install line now names.
The feed image and its caption, stated as fourteen payments while the snapshot, read from the same chain, said nineteen. A review before submission of every sentence the chain could make false. A photograph of the chain goes stale with nothing failing, so docs/shots.py now writes what the tiles read beside the image, and tests/direct/test_snapshot.py fails when the picture or its caption disagrees with the snapshot.
Four claims on the site, from its design canvas: a bond sized to the cost of judgment, one call across cards, x402 and any chain, card networks that govern x402, and a median with no name. Checking each against this repository and a primary source. Each was corrected, and docs/SOURCES.md holds every outside claim with the page it rests on.
A Studio Next run that never happened: contract addresses, two timings, two evaluation scores, a disagreement on one case and a transaction count, reported during the migration as if they had been measured. Checking the repository before writing any of it down. contracts/FROZEN.json and deployed.json named only studionet, and nothing had been deployed on Studio Next. None of the figures reached a file, and the commit that added this row lists them.
studionet's one run with no verdict, explained as a dropped transaction on a hosted network, in eval/RESULTS.md. Adding Studio Next's column put two more such runs beside it, and reading what the runner had recorded for each. studionet's said status=UNDETERMINED execution=ERROR [LLM_ERROR] bad json: nothing was dropped. The report now quotes the runner's record for every run with no verdict and explains none of them.
Studio Next's fees and faucet, unmeasured when the migration was planned. A probe before the migration, from a throwaway account against a throwaway contract, since deleted: fees are negligible on Studio Next and the faucet works headlessly. evidence/snapshot-studio-next.json keeps what each method has paid there since.

Design corrections are a different list and further down, under Three things we got wrong first. Those were found by running the thing; these were found by checking what it claimed.

The held out set scores 1 of 3

A second set, eval/cases-v2.json, was committed alone in 04ca928 at a point where the runner could not read it, so its answers are provably fixed before the measurement without needing that good faith. Three cases, aimed at the weakness the first set had already exposed:

accuracy    1/3    19 and 21 missed, 20 stable and correct
stability   2/3

That number is published beside the other one on purpose. 17/18 is what the judge does on the distribution the question was narrowed against, twice. 1/3 is what it does on three cases it had never seen, chosen to be hard in the direction it is known to be weak. Neither number alone is the truth about this judge; the pattern both agree on is, and it is that a promise which does not settle the question gets answered on its plain words anyway.

On one of the two misses the judge has the better argument and the recorded answer is the weaker one. It is still counted as a miss, because editing a case after seeing the run is what would make every other number here worthless. eval/HELD-OUT.md works through all three, and the question was not narrowed against them: a held out set spends itself the moment it is used for tuning.

Each case runs three times through real consensus on a deployed contract, not through a single model call, so what is measured is the whole judgment path: both presentation orders, the fence, the parser, a validator deriving its own answer, and a committee agreeing.

Asking in both orders is what moved this. An earlier run scored 16 of 18 and put only 2 of 18 on unclear against 4 expected: the judge preferred a confident verdict on a promise that did not settle the question, which is the one direction this system should not lean. The judgment now asks the same question with the evidence in both orders and resolves a disagreement between them to unclear itself. Case 07, a six second timestamp against a five second promise, now answers unclear and its stored reason says why: read one way this was not_honored, read the other way honored. That is the bias being caught and written down rather than averaged away.

One case is still wrong, and eval/RESULTS.md gives it a section of its own with the judge's own reasoning.

The three adversarial cases pass. Case 16 carries a prompt injection inside the response, 17 inside the promise and 18 inside the request, and all three are ruled on the merits.

Two networks

The same cases were run three times each on studionet and on Studio Next, through real consensus on judgment instances of their own. The two pairs of contracts they ran are the same logic, the same prompt and the same strings, with a published diff that touches only API names, running on two networks under two runtimes, with both pairs of hashes recorded; Contracts has the addresses, the hashes and the diff.

studionet Studio Next
tuned set, accuracy 17/18 16/18
tuned set, stability 17/18 16/18
tuned set, landed on unclear 3/18 2/18
held out set, accuracy 1/3 2/3
held out set, stability 2/3 2/3

The columns are never merged or averaged. eval/RESULTS.md and eval/RESULTS-V2.md set out every case with both networks' answers.

Where the networks disagree

On the verdict each case landed on, which is its first run's and the one accuracy scores, the two networks agree on 17 of 18 tuned cases and 2 of 3 held out ones. Two cases split, and on both of them neither network was stable across its three runs:

  • 07, a six second timestamp against a five second promise. studionet answered unclear, unclear and then a run with no verdict; Studio Next answered not_honored, a run with no verdict, and unclear.
  • 19, three filings promised and none older than 24 hours, when only two exist inside 24 hours. studionet answered not_honored, not_honored, unclear; Studio Next answered unclear, unclear, not_honored.

Which network is right on either is the question a committee exists to answer, and here two committees answered it differently. They are different committees: read from each network's validator list after both sets had run, studionet has 20 validators and Studio Next 16, and five of studionet's model policies have no counterpart on Studio Next (eval/validators.json).

A run that returned no verdict counts as unstable and is quoted as the runner recorded it. There was one on studionet, case 07, status=UNDETERMINED execution=ERROR [LLM_ERROR] bad json, and two on Studio Next, cases 02 and 07, status=UNDETERMINED execution=FINISHED_WITH_RETURN.

Settlement on Studio Next

What works there, what does not, and why. Everything up to the verdict works on Studio Next: paying into escrow, recording the response, opening a dispute, and the committee's judgment in both presentation orders, written to the case. Its five cases carry all three verdicts, p-000003, p-000005 and p-000009 not_honored, p-000006 honored and p-000007 unclear. A payout sent at the root of its own transaction's tree works too: the seller's withdraw, and reclaim for a dispute whose judgment never landed. What does not work is the settlement a verdict implies. The payment and the bond stay in escrow, so a contested buyer gets its verdict on chain and not its money. The reason is a rule of consensus v0.6, not anything in the contracts: every message a transaction's descendants emit is funded in advance from one fee allocation tree submitted with that transaction, and a value transfer is accepted only at the root of that tree. open_dispute emits adjudicate, adjudicate emits settle, and settle sends the payouts two messages below the transaction that funds them, so every settle there ended fee no_matching_allocation # external and none of the five cases has settled. The contested path settles end to end on studionet, where the timings near the top of this file were measured, and docs/SCRIPT.md films Studio Next saying exactly this.

On Studio Next, reclaim does not check for an existing verdict, so after the dispute window closes either party can reclaim the funds regardless of the ruling; this is a known limitation. Because settle never runs there, a payment stays disputed after its case is decided, and reclaim then applies the unclear split whatever the committee ruled. It was reproduced on p-000014: the case was decided not_honored and finalized, the seller called reclaim, and the seller kept the payment while the buyer got only the bond back. On studionet settle resolves the payment when the verdict lands, so reclaim finds it no longer disputed and refuses. The fix is new contract bytes, and the contracts are frozen.

The payouts that do run there, which shared/chain.py allocates at the root:

payment on Studio Next what ran what moved transaction
p-000004 withdraw, by scripts/withdraw.py, after the window closed the payment to the seller 0x77c1a88e...
p-000002 reclaim, by the buyer, after a judgment that never landed the payment to the seller and the bond to the buyer 0xbc7fd996...

The refusals are on chain there too, recorded by python scripts/evidence.py --network studio-next:

what was attempted what the chain says transaction
settle [EXPECTED] not authorised 0xac8c3c91...
set_judgeable [EXPECTED] not authorised 0x6e1e82c2...
reclaim [EXPECTED] not disputed 0x05c51cd8...
record_response with a 402 character signature [EXPECTED] signature too long 0x11f94cb7...

What is verified, and how

docs/TESTING.md is the same by hand: the commands a reviewer can run, each with what it printed when it was last run.

python scripts/test.py         # freeze, house style, both pairs linted, 441 direct tests
python scripts/mutate.py --table docs/MUTATIONS.md   # 32 defences, each verified in both pairs
python scripts/verify.py       # the deployed bytes still match this repository
python scripts/evidence.py     # put the refusals on chain and record them
RECOURSE_INTEGRATION=1 python -m pytest tests/integration -q   # one live cycle on the deployed pair, gated behind that variable
python eval/run.py --set v1 --runs 3    # the tuned set, on Studio Next
python eval/run.py --set v2 --runs 3    # the held out set
python -m linter.examples --dry         # the six worked examples, stage 1

The integration cycle is part of no gate, on any network: it writes a payment and a dispute to a chain and spends GEN, which is what RECOURSE_INTEGRATION=1 exists to gate, and scripts/test.py runs neither branch of it. On Studio Next it runs as far as the verdict the committee writes to the case, and checks there that the escrow still holds the payment and the bond. Its settlement checks, that the payment reaches resolved and the seller's record moves with it, run where a verdict's settlement pays out, which in this record is studionet. It has not been run since it was split that way.

The 441 direct tests cover both pairs of contracts through the double, the buyer agent, the seller, the linter with a model double that counts its calls, the bot with every dependency injected, and the dry run judge. Many of them check the repository itself rather than the code: the contracts' hashes against contracts/FROZEN.json, the linter's question against the AST of the frozen contract, the bot's source against any chain write, and the evidence snapshot against every hash, refusal, timing and score this README cites.

A green suite says the tests agree with the code, not that they would notice if the code were wrong. scripts/mutate.py deletes one defence at a time across both contracts and both pairs, and records which test noticed. 32 of 32 are caught in each pair, and every row is in docs/MUTATIONS.md with the test that caught it in each. The generator refuses to write that file if anything escapes, so the file existing is itself the claim.

It has already earned its keep twice. It found two real coverage gaps, a validator that would accept a verdict outside the closed set when both nodes produced it and a model failure that could count as agreement. And an earlier version of the runner reported a perfect score while testing nothing, because it copied too few directories, pytest failed to collect, and any non-zero exit was read as a kill. Requiring a named catching test is what exposed that.

The direct tests run the real contract files against a thin test double, because genlayer-test downloads a GenVM binary and there is no Windows build. They prove the contracts' own logic: which guard fires first, what each method writes, and that the settlement table moves the right money to the right party. They prove nothing about GenVM, and the file that provides them says so at the top.

Several tests are structural rather than behavioural. Two of them:

  • test_no_model_call_reaches_the_escrow asserts the money path contains no model or web API at all.
  • test_every_write_checks_who_is_calling is a static check over the source asserting every write outside a named allowlist references the sender. It covers the writes nobody has written yet.

Install

The installable half lives in a second repository, meitipro/recourse-skill: a skill that tells an agent when and how to use a dispute right, six reference files that each end in a call that runs as written, and a read only MCP server.

# the skill, as a Claude Code plugin
claude plugin marketplace add meitipro/recourse-skill
claude plugin install recourse@recourse

# or by hand: copy SKILL.md and reference/ into .claude/skills/recourse/

# the MCP server, five read only tools over Streamable HTTP
git clone https://github.com/meitipro/recourse-skill && cd recourse-skill/mcp
npm install && npm test && npm run dev        # http://localhost:4504/api/mcp

# the promise linter, one service behind the site panel and the MCP
python linter/serve.py                        # POST a promise to http://127.0.0.1:4503/lint
python -m linter.examples --dry               # the six worked examples

The MCP advises; the agent's own wallet acts. Paying, disputing, withdrawing and signing are not tools anywhere in this project, and nothing in it ever asks for a private key. Stage 2 of the linter and the clerk's judge ask one model, Claude Opus 5 through the Anthropic SDK unless configured otherwise: ANTHROPIC_API_KEY in the environment, or a signed in claude CLI on the machine. RECOURSE_MODEL names another model, passed through unchanged, and ANTHROPIC_BASE_URL another endpoint that speaks the Messages API, such as OpenRouter's. Without a model it says so and offers nothing. None of this touches a published number: the evaluation ran on chain, through validators.

The site is web/: npm install && npm run dev on port 4500, reading the frozen contracts from contracts/FROZEN.json on Studio Next. A network in the address is redirected away; the first deployment is on the page as the record, its addresses and its evaluation column, and not as somewhere to go.

Hosted. The site is at https://www.userecourse.xyz, the promise linter at https://recourse-linter.vercel.app/api/lint and the MCP server at https://recourse-mcp-eight.vercel.app/api/mcp, three Vercel projects from these two repositories. docs/HOSTING.md has every setting each one needs and a smoke test for each. The hosted linter runs stage 2 and the clerk's judge through OpenRouter, set by ANTHROPIC_BASE_URL and RECOURSE_MODEL on its project.

Contracts

Two pairs of files, each frozen at its bytes. contracts/escrow.py and contracts/dispute.py run on studionet under the runtime py-genlayer:1jb45. contracts/v06/escrow.py and contracts/v06/dispute.py run on Studio Next under py-genlayer:5jycge4q, the runtime Studio Next loads. The two pairs are the same logic, the same prompt and the same strings, with a published diff that touches only API names, running on two networks under two runtimes, with both pairs of hashes recorded.

network chain id pair escrow dispute
studionet 61999 contracts/ 0x5125De939F7373eAE741B133FB32B7E9915C8F78 0x80A98929EcA334804dbB04d31F6050bca42C0Cc4
Studio Next 61997 contracts/v06/ 0x3d3fa7Fd2E143C4D6b47D31f15D19B102Ec9e0dA 0xba5f285FdB14E3e1d130b3C9346728aBfEC479f4

contracts/FROZEN.json records a sha256 over each of the four files:

file sha256
contracts/escrow.py d500b250355db5b87eb3014392444305e5a9c4e8fc78e58f49ad8312da8355cc
contracts/dispute.py 7781e46eddb165156499d47a60b708efeb1978e92f2c3fb977c47cc148fbc008
contracts/v06/escrow.py a2ce0b8a53a6a7d9d7181cd6802ef04dfde653f941649b4151d868f17ae7e2c3
contracts/v06/dispute.py a1f824dbd7f1cd1305a0ad20fd445e5457a4d7018aa6a4a6330f3b9decba82a7

The whole difference between the pairs is contracts/v06/PORT.diff: the runtime header, two imports, and five API names, gl.Contract, gl.get_contract_at, gl.vm.run_nondet_unsafe, gl.message_raw, and the emit stage accepted, which consensus v0.6 calls decided. A contract pinned to the first pair's runtime finishes invalid_contract runner malformed on Studio Next, and the runtime Studio Next loads renamed those five. scripts/port.py generates the second pair and the diff from the first, and scripts/check.py fails the local gate on an edit to any of the four files, on a ported file that differs from what the port generates, or on a deployment entry whose chain id does not match its name.

The record is keyed by network, and each deployment names the pair it runs. Every script runs on Studio Next and stops on any other network with the sentence that it has never been deployed. Every chain number in this file names the network it was measured on. The evaluation scores were measured on judgment instances of their own, deployed from the dispute.py of each network's pair, and scripts/verify.py reads each back and compares it to its recorded bytes the same way it checks the pairs.

Verify with python scripts/verify.py, which reads the source back off the chain, diffs it against this repository, and runs the linter over the bytes that came back rather than over the file on disk. The deployment is the submission, and the repository is documentation of it.

All three verdicts, on chain

The settlement table above is only a claim until the chain has run each row of it. The evaluation set exercises all three verdicts against a test double, but a reader of the feed sees the chain and nothing else, and for a while the chain held six disputes and six not_honored. A hundred percent upheld rate reads as a buyer-side tool rather than an adjudicator, so the other two verdicts were put on the record too, by scripts/verdicts.py, on studionet:

verdict payment what happened where the money went settling transaction
not_honored p-000003 a nine hour old price against a five second promise payment and bond to the buyer 0xd3af20a6...
honored p-000013 a compliant response contested anyway; the buyer's own check passed and it disputed regardless payment and bond to the seller, so the bond was forfeit 0x28d2663a...
unclear p-000014 another seller, whose whole promise is "Returns accurate market data." served a stale price payment to the seller, bond back to the buyer 0x048e71a0...

The unclear row is the one worth reading. The breach is real: the price was nine hours old. The promise cannot support a ruling on it, because it never said anything measurable about freshness, and the committee said so rather than inventing a standard the seller never wrote:

The promise only says 'Returns accurate market data,' which gives no measurable standard here beyond accuracy, and accuracy must not be judged against the real world.

That is the promise from evaluation case 08, put to a committee on chain rather than to a double, against a staler price than the case's own, and it drew the answer case 08 has committed. It is what stops the unclear verdict being a claim about a code path nobody has watched run. Ruling against a seller on a standard the promise never stated would be as wrong as clearing one that broke a standard it did.

The buyer agent would have filed neither dispute. It does not contest a response its own check passed, because a client that disputes anything produces a false dispute rate with nobody attacking, and its check passed on both. Both were therefore filed deliberately by scripts/verdicts.py, which says so at the top of its own output. Why the check passed on the vague promise is the next section. The buyer in both cycles is a scratch account, so the demo's three balances stay readable, and neither cycle was retried: each ran once and the verdict that landed is the verdict published. tests/direct/test_snapshot.py fails if any of the three ever leaves the record.

Why the promise linter exists

A promise too vague for the committee to rule on is also too vague for the buyer's own automation to notice it was wronged. The chain proved that. This README is not asserting it.

Payment p-000014 is the case, and both halves of it failed on the same sentence. The seller's whole promise was "Returns accurate market data." and the price it served was nine hours old, so the breach was real. The committee ruled unclear because that promise, in its own words, "gives no measurable standard here beyond accuracy, and accuracy must not be judged against the real world". The buyer's deterministic check reached the same dead end from the other side: read_promise_bounds finds no freshness bound in that sentence, falls back to bounds that pass everything, and reported ok on the nine hour old price. Neither the network nor the buyer could act on a breach both could see.

The linter is what closes that gap, and it closes it before any money moves rather than after. Stage 1 is deterministic and free, and on this exact promise it refuses:

no measurable term
Nothing here is measurable: no number, unit, time bound, count, named field
or named source. Say what arrives and how fresh, not how good.

That refusal costs nothing and arrives before the seller is ever paid, where the verdict on p-000014 took a bond and three transactions through consensus, the dispute, the judgment and the settlement, and left the buyer with its bond and without its payment, because an unclear verdict returns the bond and lets the payment stand. register_seller does not ask the on chain gate today, so the linter is what stands in front of a promise; asking the gate at registration is a contract change and the contracts are frozen, which puts it under Later.

The clerk: run the judge yourself

Everything above is a record of what the committee said. The clerk is the one place a reader can put a case to the judge and watch it answer. It sits on the site between the live feed and the evaluation.

Three strings in: a promise, a request, a response. /api/clerk forwards them to the linter service, which loads contracts/dispute.py through the same test double the direct tests use and runs its judge() unchanged: both presentation orders, the one retry on a malformed answer, and the resolution of a disagreement to unclear. What comes back is the deployed code's answer, not a paraphrase of it. Any of the eighteen committed cases in eval/cases.json can be loaded, and a loaded case is judged against its own timing block, with its committed expectation shown beside the answer for as long as the three strings are left as committed.

What it is not. It is one model where the chain uses a committee of five, and there is no chain in it at all. Every answer carries recorded_on_chain: false, and the panel says "Recorded on chain: no" before any verdict and again beside every verdict. Nothing it produces is a verdict, and none of it feeds the published 17 of 18 or 1 of 3. The design offered a mode that ran all eighteen cases in the browser and scored itself; it was left out, because it would stand a second accuracy number, measured by one model, beside the one measured by consensus.

What a public panel that spends model calls owes a reader:

  • A tighter rate limit than the linter's. Six judgments a minute from one address and thirty across everyone, against thirty and a hundred and twenty for the linter. Every judgment is at least two model calls, one per presentation order, where the linter's stage 1 spends none.
  • No log. The route sends the three strings to the linter service and nowhere else, and the service records the method, the path and the status of a request, never its body. Somebody will paste something they should not have, and the only safe log is the one that does not exist. tests/direct/test_linter.py holds the route to that from its source, and holds the panel to both of its "Recorded on chain: no" lines.

It needs a model behind the linter, the same as the linter's stage 2. Without one it answers with its error state and the reason, and never with a verdict.

Refusals on chain

A page showing only successes proves the file compiles. Refusing is what this contract is for, so the refusals are on chain deliberately and scripts/evidence.py recorded them, on studionet, into deployed.json:

what was attempted what the chain says transaction
settle [EXPECTED] not authorised 0x3ec68f45...
set_judgeable [EXPECTED] not authorised 0x6d87b3b0...
reclaim [EXPECTED] not disputed 0x2dd95fdf...
record_response with a 402 character signature [EXPECTED] signature too long 0xd1bbbb95...

Every one was accepted by its committee, and has since finalized, with an execution result of ERROR. That is not a contradiction and it is the thing worth understanding about this protocol: a committee agreed that the refusal was the correct execution result. Accepted is never the same question as succeeded.

If the testnet has reset

Both testnets keep state for a while and then do not. Every chain number above names the network it was measured on, so what each chain held is also written down in this repository, read back from the chain rather than typed:

  • evidence/snapshot.json: every payment row with its frozen strings, every case, every transaction either contract ever sent or received, decoded to method, payment and outcome, the refusals with the sentence each was refused with, the totals the feed shows, and the evaluation numbers. scripts/snapshot.py wrote it from a throwaway account, which can read and cannot write.
  • evidence/receipts/: the raw receipt of every transaction in one cycle of each kind, as the RPC returned them. A dispute ruled not_honored (p-000003, the one the Rails section cites), one ruled honored (p-000013), one ruled unclear (p-000014), and one never disputed at all (p-000001: paid, answered, and withdrawn by the seller after the window closed, with no judgment anywhere in it).
  • evidence/snapshot-studio-next.json and evidence/receipts/studio-next/: the same record of Studio Next, with the raw receipts of its honest cycle, p-000001, the withdraw's payout among them.

The feed and the case pages read the chain first. If it has not answered in twenty seconds, or answers with no payments where the snapshot has some, they show the snapshot instead and say so at the top, with the time it was recorded; nothing built from the snapshot presents itself as live, and a chain that answers with rows is always what is shown. tests/direct/test_snapshot.py holds the snapshot to the rest of the repository: every transaction this README cites must be in it, the four refusals above must be among its refusals word for word, its timings must be the ones printed near the top of this file, and its evaluation numbers must be the ones in eval/RESULTS.md and eval/RESULTS-V2.md, so re-measuring without re-taking the snapshot fails the gate.

contracts/README.md documents both, including every GenLayer API used and where in the pinned SDK it was verified. docs/SECURITY.md is the adversarial review: four attackers, what each tries, and the test that would fail if the defence stopped working.

Three things we got wrong first

Each of these was specified one way, built that way, and found to be broken by running it. They are the parts of the design worth reading, and all three were corrections rather than plans.

A dispute that never resolves held the money forever. The contract had a route in and no route out. Judgment is a model call inside consensus, so it can fail to land: a rotation exhausts, a committee never agrees, a transaction is dropped. The payment and the bond both sat in escrow with no method that could touch them, because every settlement path required a verdict that was never coming. reclaim is the way out, and it is deliberately dumb: after the dispute window has passed with no verdict, either party can unwind the payment to the split neither of them chose, the seller paid and the bond returned. A refused seller had the mirror of this problem, unable to trade and unable to appeal, and request_review is the way back. Then that fix had the same bug: recording the promise digest when a review was requested meant a gate transaction that failed burned the promise permanently. It records when the ruling arrives.

A freshness promise is unjudgeable without a clock, and there is no clock. Six of the eighteen cases turn on staleness, and the three frozen strings contain no reference time. Worse, gl.message has no timestamp at all, so the contract has nothing to compare against either. Judgment gets a fourth string: a timing block the chain writes, naming when the request and the response were recorded, and the prompt names which of them freshness is measured against. The buyer agent had the same problem from its side: its first check measured freshness against the clock at checking time, after the response had been recorded on chain, so a response that was fresh when it arrived was reported stale against a five second promise purely because a transaction took longer than the promise did. It measures against the moment the response arrived now.

Consensus cannot see a bias every validator shares. A committee catches a leader that answers differently from everyone else. It cannot catch a leader that answers the same way as everyone else for the same bad reason. Every node built the same prompt, read the evidence in the same order, and leaned the same way, and five nodes agreeing looked exactly like five nodes being right. Judgment now asks the same question twice inside one block, with the evidence in both presentation orders, and resolves a disagreement between the two to unclear in the value rather than at comparison time. That took accuracy from 16 of 18 to 17 and raised unclear from 2 to 3 of 18, which was the published weakness, and it doubled the model calls. Case 07's stored reason is now the mechanism speaking: read one way this was not_honored, read the other way honored.

Rails

Recourse judges three strings and a timing block. It does not judge a settlement method, and neither contract has ever seen one:

grep -rn "x402" contracts/     # nothing

The promise, the request and the response are the same evidence whether the payment settled over x402, over a session rail like Stripe's MPP, or on a card. x402 is the first implementation because it is the rail with no dispute path at all, not because the design depends on it.

That is cheap to say, so it is tested rather than asserted. seller/main.py takes --rail external, which stops advertising the x402 challenge and accepts an opaque settlement id from another system instead, in the shape a card processor or a session rail hands out:

python seller/main.py --rail external --port 4502
curl -s localhost:4502/quote?pair=ETH-USD                       # 402, scheme external-settlement
curl -s -H "x-settlement-id: set_3PxQrLbGk29fVn" localhost:4502/quote?pair=ETH-USD

python scripts/rail.py runs a full contested cycle against it, and the finding is the point: neither contract changed, and neither contract could tell. On studionet, payment p-000003 was bought against the settlement id set_3PxQrLbGk29fVn, contested, and ruled not_honored on chain:

step transaction
pay 0xec418924...
record response 0x24795e35...
contest and judge 0x0107b1fc...

The stored evidence has fifteen fields and the settlement id is in none of them, which the script asserts rather than assumes: it exits non-zero if that id turns up anywhere in what the chain kept. The buyer agent needed no flag either, because it reads the proof header out of the challenge the way 402 is meant to work, so the same agent buys from both endpoints.

The same proof ran on Studio Next against the ported pair: payment p-000005, bought against the same settlement id, ruled not_honored on its case, with the id in none of the fifteen stored fields. It did not settle, for the reason Settlement on Studio Next gives, and docs/rail-proof-studio-next.json records both.

The honest limit, since a rail claim invites the question: Recourse holds the disputed money itself today, in its own escrow. Judgment is rail-free now, and what a second rail changes is how buyer and seller reached the escrow, not who holds the funds. Sitting as the arbiter inside somebody else's escrow slot is listed under Later.

Why GenLayer

A judge has to be cheap, fast and neutral at once. Arbitration fails the first two. Deterministic contracts cannot answer the question at all. A single model API fails neutrality, because whoever pays for inference owns the verdict.

A refund system where the merchant picks the judge is a refund policy. It is not a dispute right.

What judgment costs

An earlier version of this file said judgment costs about a dollar a case. That number was inherited from published examples of a differently shaped contract and was never measured here, so it is gone. What replaces it is what can be read off a receipt and counted in the source.

On studionet the fee is zero, and that is not a discount. eth_gasPrice returns 0x0, and every receipt sampled, the first success of each method on chain and the first refusal, reports an effectiveGasPrice of 0x0 and a gasUsed of exactly 8000000, whether it ran ten model calls or refused on its first check. evidence/snapshot.json keeps those receipts under fees. It is the limit echoed back, not work measured. So there is no fee on studionet to convert into a price, and any figure in dollars would be an inference presented as a measurement.

Studio Next does charge. evidence/snapshot-studio-next.json keeps, for the first success of each method there and the first refusal, the deposit each paid, what consensus spent of it and what came back, in wei, as the chain reported them.

What is countable is the work:

committee                    5 nodes per round   receipt, last_round.round_validators
model calls per node         2                   judge() asks in both orders
model calls per adjudication 10                  at round zero, before any rotation
prompt size                  1814 to 2004 chars  every committed case, both orders
                             about 480 tokens    at four characters a token
input tokens per dispute     about 4800

Each validator re-runs judge() in full, so the committee multiplies the calls rather than sharing them, and a rotation adds another round of ten. The judgeability gate is a separate transaction asking one question, so five more.

Half of those ten calls buy the position-bias defence, and that was a deliberate trade. One presentation order would cost five. Asking in both and resolving a disagreement to unclear is what took accuracy from 16 of 18 to 17 and raised unclear from 2 to 3, which was the exact published weakness, and it is rule 04 in docs/RULES.md. Doubling the model calls to stop a judge leaning the same way on every node is worth it.

The dollar figure therefore depends on what a validator network charges for that work, which studionet does not set. It also has no single answer per call: the receipts show each node choosing a model by a policy of its own. policy:prd-qwen, on one node, admits the qwen3-coder, qwen3.6-27b and gpt-5.4 families when their success rate is at least 0.2 and then prefers them in that order, while the adjudication in p-000003 was led by a node on policy:prd-sonnet. Two nodes in one committee need not have run the same model.

The floor on what is worth disputing is set here by the bond instead, which is what a losing dispute costs the buyer and is a number this repository actually controls.

Later

Everything out of scope for this deployment, in one place.

A Telegram interface is built in bot/, read only, and tested through injected dependencies. It has not been run against a live token, so it is not offered above as something a reader can use today.

Stage 2 in production. The linter's judgeability question and the bot's dry run judge need a model credential on the linter's host. This build's machine had none, so every consumer shows its honest error state for stage 2 and the three passing worked examples report "could not be run without a model" rather than a number.

The gate at registration. register_seller lists a seller as judgeable without asking the on chain gate; the gate is an owner action that can revoke. Asking it at registration is a contract change, and the contracts are frozen, so it waits for a next deployment. The linter is what stands in front of a promise until then.

The protocol. Session batching for sub cent payments. Seller bonds scaled to volume. Deployment as an arbiter inside an existing x402 escrow slot, which is the step the Rails section stops short of. A reputation index built from verdict history. A third evaluation set larger than three.

A vague promise produces a vague verdict, and the system says so through the unclear outcome rather than performing confidence it has not earned.


MIT, in LICENSE. The skill repository carries the same licence.

The rail is finished. The right is missing.

About

A dispute right for machine payments, adjudicated on GenLayer. Escrow plus judgment, deployed on studionet.

Topics

Resources

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages