Skip to content
SentinelyRequest a free evaluation

A planted clock-domain bug that passes ordinary simulation on 48 of 48 seeds.

One small handshake design, clock ratios 1:1 to 7:1, Icarus Verilog 12. Late settling is a behavioral model.

Sentinely builds training environments for AI that debugs chips, graded with a behavioral model in which synchronizer flip-flops can settle a cycle late, as real ones can.

The defective handshake on one real seed: the same clock edges, simulated two ways.

Ordinary simulation: one slow-clock edge samples the brief low between the two requests. The synchronizer captures it on time, it reaches req_sync one cycle later, and the slow side sees two requests.

Figure 1. Behavioral model of late settling, up to one extra cycle. Seed 1027 of the 48; the slow clock's period is 7 times the fast clock's.

The measured result

We planted a clock-domain bug in a small request-acknowledge handshake: the acknowledge is a one-cycle pulse instead of a level held until the request drops. Then we ran the same 48 seeds with and without late settling.

Table 1. The same 48 seeds with and without late settling.
DesignOrdinary simulationWith late settling
Defective handshake48 of 48 seeds pass20 of 48 pass: 28 caught
Correct fix48 of 48 pass48 of 48 pass

A behavioral model of late settling, measured on one design over 48 seeds at clock ratios from 1:1 to 7:1. Icarus Verilog 12 and Verilator 5.038 gave the same result for every seed.

Read the finding

What you get

Tasks
Small designs, each with a planted clock-domain bug and a short bug report: one handshake task, one event-crossing task, and a family of 17 asynchronous FIFOs with Gray-coded pointers, 13 for training and 4 held out. The FIFO variants differ in naming, width, depth, burst length and the shape of the defect.
The grader
A patch may change only the task's own file. Static clock-domain rules require every crossing to go through a provided synchronizer. Visible and hidden seeds then run across a range of clock ratios with late settling on, and every result replays exactly from its seed.
One-shot or agentic
Your model either writes a fix in one answer, or works like an engineer: it reads the design, edits it, runs the visible seeds, reads the failures and submits within a fixed budget. Its tools can't open the testbench, the hidden seeds or the synchronizer source.
Plugs into your training loop
It connects through an HTTP interface or a Python API, with tool definitions in the OpenAI function-calling format. Grading runs on CPUs, and we measure its cost on your hardware before training starts.

How we keep the reward honest

We keep a log of each kind of patch that has fooled, or tried to fool, the grader, and how we caught it.

As of 2026-10-07: 17 reward hacks found, 17 closed; and 3 cases where the grader wrongly rejected a fix, all resolved.

We keep regression fixtures for 14 of those hacks and for all 3 wrong rejections. On 2026-10-07, every fixture gave its expected verdict on Icarus and Verilator: hacks rejected, correct fixes accepted.

  • Editing the testbench instead of the design. Rejected: a patch may touch only the task's file.
  • Hiding a binary pointer crossing in a FIFO. The patch held back reads until the synchronized count was non-zero. It passed all 200 hidden seeds on every FIFO variant. An open-weights 4B model we ran in agentic mode found it on its own, in 3 of 20 episodes. A hidden-seed check on the crossing itself now rejects it.
  • Splitting a multi-bit crossing across one-bit synchronizers to dodge a per-instance check. It's now caught by counting the bits that change across the whole crossing per source-clock edge.

An independent formal audit is in progress on our two sample tasks, the handshake and the event crossing. An outside verification engineer is writing formal properties from the specification, without seeing our testbenches, to check against the fixes our grader accepted and rejected.

Start with a free private evaluation

We run your model on our clock-domain tasks, privately, and send you a report: pass rate per task with 95% intervals, and where and why attempts failed. An endpoint is enough; we don't need your weights.

If the results are useful, a six-week pilot follows: a baseline, four weeks of training on our training tasks, a held-out measurement on designs your model never trained on, and a readout.

How the pilot works

Request a free private evaluation

Email support@sentinely.ai with the model you'd like evaluated.