Defective runs caught
Full gate sensitivity on the 54-run corpus — vs 64% (7/11) for evidence-only review. Wilson 95% CI [74%, 100%]; bounded-sample caveat disclosed
Loading project...

Sole Designer, Builder & Operator
Solo (AI sub-agents for research, implementation and review)
3-day sprint: first commit Sep 29 → live-corpus measurement Oct 1 → submitted
A submission to the Nebius × NVIDIA Global AI Hackathon (Coding & Agentic Engineering track). CI already distrusts artifacts — but nothing distrusts an agent's *report*. When a coding agent says 'fixed it, tests pass,' runanchor doesn't take the account at face value: every run gets a receipt anchored to provider-issued execution records (Nebius Sandboxes operation UUID + image IDs), is replayed from the recorded environment, checked by a hidden oracle suite the agent never saw, and only adopted after human or measured-machine approval. Rejected and mismatched runs stay on an append-only hash-chained ledger. Measured on a 54-run live corpus: the full gate caught 100% of defective runs (11/11) at 98% specificity — where an evidence-only review let 4 defective runs slip through (64%). Public repo, MIT licensed; the measurement procedure is committed for reproduction.
Coding agents write, run, and test code autonomously — but when one reports 'done, tests pass,' nothing checks the report itself. CI distrusts artifacts; nothing distrusts the agent's account of what happened. So teams pay for the gap the expensive way: a human re-runs everything by hand. A green-looking log can hide three different failures — the run that never actually ran, the run whose tests passed on a hardcoded answer, and the run whose produced state is simply wrong.
A verification and approval gate between the agent's claim and adoption. Every run inside a Nebius Sandboxes environment gets a receipt anchored to records the agent cannot fabricate — the provider's operation UUID and image IDs. Verification runs on two axes: replay forks the recorded start image and reruns the recorded command, comparing exit code and normalized output fingerprints; a hidden oracle forks the produced result image and runs a test suite the agent never saw — catching hardcoded answers and latent spec violations that a green log cannot show. Runs are adopted only after a human or a measured machine judge approves; rejected, mismatched, and unresolved runs stay on an append-only hash-chained ledger — evidence that never silently disappears.
Coding agents write, run, and test code autonomously — but when one reports 'done, tests pass,' nothing checks the report itself. CI distrusts artifacts; nothing distrusts the agent's account of what happened. So teams pay for the gap the expensive way: a human re-runs everything by hand. A green-looking log can hide three different failures — the run that never actually ran, the run whose tests passed on a hardcoded answer, and the run whose produced state is simply wrong.
A verification and approval gate between the agent's claim and adoption. Every run inside a Nebius Sandboxes environment gets a receipt anchored to records the agent cannot fabricate — the provider's operation UUID and image IDs. Verification runs on two axes: replay forks the recorded start image and reruns the recorded command, comparing exit code and normalized output fingerprints; a hidden oracle forks the produced result image and runs a test suite the agent never saw — catching hardcoded answers and latent spec violations that a green log cannot show. Runs are adopted only after a human or a measured machine judge approves; rejected, mismatched, and unresolved runs stay on an append-only hash-chained ledger — evidence that never silently disappears.
Publish only what you measured — and make the measurement reproducible. The 100% catch claim is bounded and disclosed as such: 11 defective items is a small sample, so the number ships with its Wilson confidence interval [74%, 100%] and the caveat that the two judgment layers are not independent detectors (the gate judge is shown the oracle outcome). Ground truth is executable, not labels: the oracle's exit code on the produced state, which is why the corpus also caught two 'clean' items where the agent's honest fix simply failed. The one false reject is disclosed too — a run whose produced state passed the oracle but whose receipt ends on a permission error, so the suite was never demonstrably green. The full procedure — corpus, judge prompt, command sequence — is committed to the repo: you verify our numbers, not trust them.
Each receipt references ConTree operation UUIDs and image IDs issued by Nebius itself — attestations the agent cannot fabricate. 'Demo' receipts from fixtures are structurally marked and can never masquerade as live runs.
Replay catches 'reported but never run' — the recorded command is re-executed from the recorded start image and compared by exit code and output fingerprint. The hidden oracle catches 'green but wrong' — a suite the agent never saw runs against the produced result image. In the benchmark, verification turned 4 baseline misses into catches and flipped 3 borderline rejects into adopts.
Rejected, mismatched, and unresolved runs remain on a hash-chained ledger — the run history can't be quietly rewritten. `runanchor check --ledger` verifies chain integrity: the live benchmark produced 455 anchored snapshots with an intact chain.
Built in a 3-day hackathon sprint and measured before it was claimed. The benchmark ran the same 54-run series through two layers — an evidence-only judge (the baseline) and the full gate (evidence + replay + oracle) — against a corpus of 39 seeded-trap cases and 15 clean controls where the ground truth is the oracle's exit code, not the planted label. The gate caught all 11 defective runs at 98% specificity; the evidence-only baseline let four defective runs through (including seeded hardcode and spec-contradiction traps). Honest caveats are published in the README rather than rounded away: the small denominator, the non-independence of the two layers, and the one false reject. Total inference cost measured on the console: ~$0.30; sandbox operations were $0 during the provider beta. A technical article covering the measurement method was published on Zenn the same week.
Full gate sensitivity on the 54-run corpus — vs 64% (7/11) for evidence-only review. Wilson 95% CI [74%, 100%]; bounded-sample caveat disclosed
One false reject: produced state passed the oracle, but the receipt ends on a permission error — the run was never demonstrably green, so the boundary is documented
Anchored run/replay/oracle snapshots across the live benchmark — hash chain verified intact via `runanchor check`
Total project inference through 2026-10-01 (console-measured); sandbox ops $0 during Nebius beta — full corpus + judge prompt + commands committed for reproduction

Under the hood — every run becomes a provider-anchored receipt, verified by replay and a hidden oracle, then recorded to an append-only hash-chained ledger (pending → adopted/rejected; mismatches are never deleted)