IntellectEU has released a test-generation benchmark for Daml, giving Canton developers a way to check whether an AI agent’s tests can catch faults in smart-contract code.
The public release contains 34 tasks drawn from six existing code repositories, including Canton, Splice and Daml Finance. Instead of asking the agent to build an application, each task provides working code and an emptied test file. The agent must write the tests.
Those tests face two checks. First, they must work against the correct implementation. They are then run against altered versions containing deliberately introduced faults. A successful test catches a fault by failing when the code is wrong.
The public task set contains 159 such faults. Some are based on real bugs from the source repositories’ history; others were created and reviewed to represent plausible developer mistakes.
In the published baseline, OpenAI’s gpt-6-sol wrote tests that caught 151 of the 159 faults, or 95%. All 34 public tasks produced test files that passed against the correct code. The agent used the benchmark’s default setup, with a ten-minute limit per task and no extra skills or prompt guidance.
The release includes the results behind those totals. A dashboard shows the agent’s steps, the code it produced and how its tests performed against each fault. Cost and runtime are recorded too, allowing developers to compare performance alongside the resources used.
Teams can add tasks from their own Daml projects, including private codebases, without changing the benchmark’s core machinery.
The work forms part of the Canton Development Fund’s Daml Code Assistant project. Its latest submission covers the test-generation benchmark deliverable. The public repository contains the evaluation code, dashboard and baseline records.



