A planted clock-domain bug that passes ordinary simulation on 48 of 48 seeds.
One small handshake design, clock ratios 1:1 to 7:1, Icarus Verilog 12. Late settling is a behavioral model.
Sentinely builds training environments for AI that debugs chips, graded with a behavioral model in which synchronizer flip-flops can settle a cycle late, as real ones can.
The defective handshake on one real seed: the same clock edges, simulated two ways.
Ordinary simulation: one slow-clock edge samples the brief low between the two requests. The synchronizer captures it on time, it reaches req_sync one cycle later, and the slow side sees two requests.
With late settling: the flop that samples the brief low settles one cycle late. By the next edge req is high again, so the low never reaches req_sync. The two requests merge into one, the second is never acknowledged, and the handshake hangs.
The measured result
We planted a clock-domain bug in a small request-acknowledge handshake: the acknowledge is a one-cycle pulse instead of a level held until the request drops. Then we ran the same 48 seeds with and without late settling.
| Design | Ordinary simulation | With late settling |
|---|---|---|
| Defective handshake | 48 of 48 seeds pass | 20 of 48 pass: 28 caught |
| Correct fix | 48 of 48 pass | 48 of 48 pass |
A behavioral model of late settling, measured on one design over 48 seeds at clock ratios from 1:1 to 7:1. Icarus Verilog 12 and Verilator 5.038 gave the same result for every seed.
What you get
- Tasks
- Small designs, each with a planted clock-domain bug and a short bug report: one handshake task, one event-crossing task, and a family of 17 asynchronous FIFOs with Gray-coded pointers, 13 for training and 4 held out. The FIFO variants differ in naming, width, depth, burst length and the shape of the defect.
- The grader
- A patch may change only the task's own file. Static clock-domain rules require every crossing to go through a provided synchronizer. Visible and hidden seeds then run across a range of clock ratios with late settling on, and every result replays exactly from its seed.
- One-shot or agentic
- Your model either writes a fix in one answer, or works like an engineer: it reads the design, edits it, runs the visible seeds, reads the failures and submits within a fixed budget. Its tools can't open the testbench, the hidden seeds or the synchronizer source.
- Plugs into your training loop
- It connects through an HTTP interface or a Python API, with tool definitions in the OpenAI function-calling format. Grading runs on CPUs, and we measure its cost on your hardware before training starts.
How we keep the reward honest
We keep a log of each kind of patch that has fooled, or tried to fool, the grader, and how we caught it.
As of 2026-10-07: 17 reward hacks found, 17 closed; and 3 cases where the grader wrongly rejected a fix, all resolved.
We keep regression fixtures for 14 of those hacks and for all 3 wrong rejections. On 2026-10-07, every fixture gave its expected verdict on Icarus and Verilator: hacks rejected, correct fixes accepted.
- Editing the testbench instead of the design. Rejected: a patch may touch only the task's file.
- Hiding a binary pointer crossing in a FIFO. The patch held back reads until the synchronized count was non-zero. It passed all 200 hidden seeds on every FIFO variant. An open-weights 4B model we ran in agentic mode found it on its own, in 3 of 20 episodes. A hidden-seed check on the crossing itself now rejects it.
- Splitting a multi-bit crossing across one-bit synchronizers to dodge a per-instance check. It's now caught by counting the bits that change across the whole crossing per source-clock edge.
An independent formal audit is in progress on our two sample tasks, the handshake and the event crossing. An outside verification engineer is writing formal properties from the specification, without seeing our testbenches, to check against the fixes our grader accepted and rejected.
Start with a free private evaluation
We run your model on our clock-domain tasks, privately, and send you a report: pass rate per task with 95% intervals, and where and why attempts failed. An endpoint is enough; we don't need your weights.
If the results are useful, a six-week pilot follows: a baseline, four weeks of training on our training tasks, a held-out measurement on designs your model never trained on, and a readout.
Request a free private evaluation
Email support@sentinely.ai with the model you'd like evaluated.