Clock-domain repair environments, with a grader built to resist gaming
Tasks
Each task gives a model a small design with a planted clock-domain bug, a short bug report stating the symptom and the invariant to keep, and one file it may edit.
- Request-acknowledge handshake. A handshake between a fast requester and a slow acknowledger that intermittently hangs. Clock ratios 1:1 to 7.5:1.
- Event crossing. Single-cycle events cross into a slow domain through a pulse synchronizer and are counted on arrival. Clock ratios 1:1 to 50:1.
- Asynchronous FIFO family, 17 variants. Write and read domains with Gray-coded pointers, each pointer crossing through a synchronizer. The read side commits to a burst once its synchronized fill count reaches the burst length. The planted defect crosses the write pointer as a binary count. A sample taken while several bits change can be a count the pointer never had. Hidden seeds run clock ratios from 1:1 to 16:1, with either side fast.
Variants and held-out designs. A generator varies naming style, module name, data width, depth, a non-power-of-two burst length, optional extra outputs, and the shape of the defect. Of the 17 FIFO variants, 13 are for training and 4 are held out, and the held-out designs stay on our machines.
How each variant is gated. Every variant passes the gate before it ships: the clean design and the reference fix pass all 200 hidden seeds with late settling on, and the defect passes all 200 under ordinary simulation but is caught on 98 to 138 of them with late settling and fails 3 to 5 of the 8 visible seeds, so an agent sees a failure while it works.
Not covered. In the FIFO family only the write-pointer path is defective. Three cases are not covered: a binary read pointer, a pointer published before its data is written, and reset-domain bugs.
The grader, in order
- Allowlist
can return TAMPER
- Static gate
can return TAMPER
- Build
can return BUILD_FAILED
- Visible seeds
can return FAIL
- Hidden seeds
can return OVERFIT_VISIBLE_ONLY
- Allowlist. A patch may change only the task's editable file.
- Static gate. Only pure, side-effect-free system functions are allowed; file, OS and simulator calls are rejected before anything is built. Clock-domain rules require every crossing to go through a provided synchronizer, including crossings hidden in intermediate logic, and reject hand-written synchronizers where a primitive is provided and synchronizers whose clock or reset is rewired.
- Build on Icarus Verilog 12 or Verilator 5.038.
- Visible seeds. Each task has 8, the seeds an agent can run itself.
- Hidden seeds. Each task has 200. Training rollouts can grade on a subset, such as 32, and evaluation uses all 200. A 32-seed subset may not reach a task's highest clock ratios. Each seed fixes a clock ratio, a phase offset, a stimulus pattern and the late-settling choices. The FIFO family adds a hidden-seed check: each pointer crossing may change at most one bit per source-clock edge, summed across all of its synchronizers.
- Late settling is on in both seed stages; only our demo tool turns it off. Each synchronizer bit whose input has just changed resolves after two or three destination-clock cycles, chosen per bit from the seed. On a bus, only bits that changed in the source's latest update can settle late. A Gray-coded pointer is therefore seen at its current value or one step behind, while a binary count can be sampled as any mix of old and new bits. This is a behavioral model: late settling of up to one extra cycle, with no unknown values, oscillation, longer settling or failure rates over time.
- Deterministic replay. Every result replays exactly from its seed. On the 48 seeds of our published measurement, Icarus and Verilator gave the same result and trace hash for every seed, on both sample tasks, in both modes. Seeds run in one batch only on Verilator, for stages of 24 or more seeds, and only with a parity certificate. A parity certificate is a content fingerprint of the runner, synchronizer models, testbench, task sources and patches, issued after batched and one-at-a-time runs agreed. Every batched grade also re-runs in reverse order plus two seeds on their own. If any of them differ, or there is no valid certificate, each seed runs on its own.
Verdict. A submission passes only if it clears the allowlist, the static gate, the build and every visible and hidden seed. Reward is 1 or 0. A patch that passes the visible seeds but fails hidden ones is reported as overfit, not as a pass.
Reward hacks. We keep a log of each kind of patch that has fooled, or tried to fool, the grader, and how we caught it.
As of 2026-10-07: 17 reward hacks found, 17 closed; and 3 cases where the grader wrongly rejected a fix, all resolved.
We keep regression fixtures for 14 of those hacks and for all 3 wrong rejections. On 2026-10-07, every fixture gave its expected verdict on Icarus and Verilator: hacks rejected, correct fixes accepted.
One suspected hack turned out to be a genuinely correct alternative fix, and it's now a fixture that must keep passing. Of the three wrong rejections, two now pass. For the third, a fix that rewired a synchronizer's clock, the rule is now explicit: the gate rejects it with a clear message.
One-shot or agentic
One-shot. The model reads the task and the defective design and answers with a fix.
Agentic. The model works through five tools:
| Tool | What it does |
|---|---|
list_files | Lists the files the agent may read; editable files are marked. |
read_file(path) | Reads the editable design file, or the synchronizers' ports and timing contract. |
write_file(path, content) | Replaces the full content of an editable file. |
run_visible | Builds the design and simulates the visible seeds. Returns pass or fail per seed with the start of each failure's message, or the compiler's errors. |
submit | Grades the current design and ends the episode. |
- Budgets are counts, not wall-clock time. By default an episode gets 30 tool calls and 6 simulation runs, both configurable. When the tool calls run out, the environment submits the current design.
- What the agent can reach. The agent can read only its editable file and the synchronizers' timing contract, and write only its editable file; everything else is refused, including from inside its own simulation runs.
- The grader report. Submitting returns the reward and the full grader report to your harness, for your training logs. Our agent runner never passes the report to the model; a direct integration should strip it the same way.
- Tested with scripted agents and a stand-in model: a correct fix graded pass; a raw crossing was rejected on submission; compiler errors reached the agent; out-of-bounds reads and writes were refused; and the grader's report never reached the model.
Integration
HTTP, from any language or RL stack. The server serves only the tasks you list when you start it.
| Request | What it does |
|---|---|
GET /v1/tasks | The tasks served, and the tool definitions |
POST /v1/episodes | Starts an episode; returns its id and first observation |
POST /v1/episodes/<id>/step | Runs one tool call |
DELETE /v1/episodes/<id> | Frees the episode |
Each server process handles one request at a time; for parallel episodes, run one process per CPU core. Set a token to require bearer authentication.
Python:
env = CdcRepairEnv(task_dir, repo_root, max_steps=30, max_sim_runs=6)
obs = env.reset()
out = env.step("run_visible", {})
out = env.step("submit", {}) # out["reward"], out["report"]Tool definitions come in the OpenAI function-calling format. One-shot evaluation talks to an OpenAI-compatible Chat Completions endpoint.
Cost, measured on your hardware first
Grading runs on CPUs. Before training starts, we measure the cost per rollout on your hardware and agree a ceiling. For reference, here are our own measurements, in CPU-seconds per grade. Every grade also runs the 8 visible seeds.
| Task | Icarus, one seed per process, 24 hidden seeds | Verilator, batched, 200 hidden seeds |
|---|---|---|
| Handshake | 3.6 | 13.8 |
| Event crossing | 7.3 | 29.0 |
On these two tasks, Icarus was cheaper below about 50 to 90 hidden seeds, and batched Verilator above that. That's why we recommend Icarus for training at about 32 seeds and Verilator for 200-seed evaluation. The grader never batches on Icarus. Each Verilator grade carries a fixed cost of about 8.5 CPU-seconds, mostly the model compile, with the compiler cache off.
The FIFO family runs under the same policy. Across the 17 variants, each grading its own reference fix, a grade cost 9.4 to 17.7 CPU-seconds on Icarus at 32 hidden seeds, standalone (mean 13.0), and 28.5 to 59.8 on Verilator at 200 hidden seeds, batched (mean 43.4). Re-measured 2026-10-07 on the current grader; all 102 runs ran in the expected mode.
Scope:
- measured on 2026-10-03, on one machine (Ryzen 9 9950X, WSL2);
- two tasks, grading reference fixes;
- the maximum of 3 runs, user plus system CPU time including simulator processes;
- compiler cache off.
A submission that fails early costs less, and one whose batched run disagrees, and is re-graded seed by seed, costs more.