Start with a free private evaluation
The free private evaluation
We run your model on our two sample tasks, the handshake and the event crossing, and on a few of our FIFO training variants, never on held-out designs, and send you a private report.
What we need
- An OpenAI-compatible endpoint: base URL, API key and model name. Or open weights, which we serve ourselves with vLLM.
- Your rate limits, so the run doesn't trip them.
- Agreement on data handling (see the FAQ).
- For an agentic evaluation, an endpoint that supports tool calling.
What you get back
- Pass rate per task with 95% intervals.
- Where attempts went, with what each outcome means. For example: passed every seed; passed visible but failed hidden seeds; failed simulation; rejected for an unsynchronized crossing; edited files outside the task; or no usable patch.
- Tokens per attempt, when your endpoint reports them.
- One-shot and agentic results reported separately.
The six-week pilot
| When | What happens |
|---|---|
| Week 0 | Kickoff. We measure grading cost on your hardware, run a baseline evaluation of your model, and agree the success criterion with you. |
| Weeks 1 to 4 | You train on our 15 training tasks, with the grader running on your machines or ours. A weekly 30-minute check-in. |
| Week 5 | Held-out measurement of your trained model on a fresh set of held-out designs generated for you, with the same grader, seeds and intervals as the baseline. |
| Week 6 | Readout: what changed, by how much, and where the model still fails. A decision on a license. |
What we need from you: a technical owner, access to your model (an API endpoint or weights), and the compute for your training runs.
Held-out designs, made for you. For the held-out measurement, we generate a fresh set of held-out designs for each customer with our variant generator, and never reuse it with another customer. The pilot agreement forbids training on it. That keeps the measurement honest even when we reach your model through an endpoint, so weights stay optional.
FAQ
- What do you need from us?
- For the evaluation: an OpenAI-compatible endpoint or open weights, your rate limits, and agreement on data handling. For a pilot: a technical owner, model access and the compute for training.
- What happens to our model's outputs?
- We use them only to prepare your report. We never share them and never use them for training. We delete them, with your key, within 30 days.
- Do you need our weights?
- No. An endpoint works, including for the held-out measurement: its designs are generated fresh for you, never reused with another customer, and the pilot agreement forbids training on them.
- Where does grading run?
- For the evaluation, on our machines; you only give us an endpoint. In a pilot, grading for training runs on your machines or ours, and held-out grading always runs on ours.
- Which simulators?
- Icarus Verilog 12 and Verilator 5.038. On our handshake and event tasks, Icarus measured cheaper at training scale and Verilator at 200-seed evaluation, so we use each where it's cheaper.
- What does it cost?
- The evaluation is free. The pilot is a fixed fee, scoped with you after the evaluation.