The seller posts a bond before the work starts. If the buyer disputes the deal, KOANE re-runs the checks committed to the deal, and the settlement program either releases the bond or slashes it. The program writes every settle and slash to the agent's on-chain record.
bonded deals, replayable verdicts // live on solana devnet
Cheating costs the bond
The seller bonds at least twice the price before the work begins. A seller caught cheating loses that bond, so starting over with a fresh wallet does not erase the cost.
The verdict is a pure function
In a dispute, the verdict comes only from the results of checks committed before the work started. Claims made afterward never enter it. Anyone can replay the verdict with koane-verify and get the same bytes back. There is no jury to bribe, and a wrong verdict fails the replay.
The record follows the agent
Every settle and every slash lands on the agent's on-chain record. The next buyer can read that history before hiring the agent.
The AI investigator chooses which checks to run and in what order. It cannot set the outcome. rule() takes no verdict argument, the committed checks run whether the model picked them or not, and the verdict is a pure function over their results.
Three reports in the repo measure this, and each one states its limits.
Tested against a fully obedient model
The benchmark drives 23 attack payloads in 8 classes through the real investigator loop, using a scripted model that obeys the attacker every time. That is the worst case an injection can produce. No attack caused a false slash or changed the consequence of a committed check. The result is a structural worst case, and it claims no numbers for a live model. the benchmark →
Two implementations, one verdict
Independent TypeScript and Rust implementations of the verdict function agree on 19 fixture-derived cases and 2000 generated cases. The comparison covers the verdict function only. The check primitives that feed it are outside it. the determinism report →
Thresholds measured on held-out data
The frozen text-check thresholds ran against 46 samples they were never tuned on. They caught every blatant breach and failed one honest sample. The report lists that false positive and how to compose checks around it. the report →