---
title: "IntellectEU releases AI test benchmark for Canton developers"
description: "IntellectEU has published a benchmark that measures how well AI agents write tests for Daml code. The public set contains 34 tasks and 159 deliberately introduced faults, with code and baseline results available for developers to examine."
url: "https://cantonnews.org/intellecteu-releases-ai-test-benchmark-for-canton-developers"
canonical: "https://cantonnews.org/intellecteu-releases-ai-test-benchmark-for-canton-developers"
markdown_url: "https://cantonnews.org/intellecteu-releases-ai-test-benchmark-for-canton-developers.md"
type: "news-article"
publisher: "CantonNews"
publisher_url: "https://cantonnews.org"
author: "CantonNews"
author_url: "https://cantonnews.org/author/cantonnews"
author_type: "organization"
date_published: "2026-10-07T09:58:17.153Z"
category: ["TECHNOLOGY", "ECOSYSTEM"]
tags: ["IntellectEU", "Canton Network", "Daml", "AI", "Developer tools", "Canton Development Fund"]
image: "https://pnyaxwohgkgczwuvbaqg.supabase.co/storage/v1/object/public/article-images/cantonnews/2026-10-07-evening/d594809f3cc12764171ca54ae5d45ce7c0a4011cb131dfa78ad85ed7f2e8cfc9-08-intellecteu-daml-test-benchmark.png"
read_time_minutes: 2
language: "en"
---

# IntellectEU releases AI test benchmark for Canton developers

**By [CantonNews](https://cantonnews.org/author/cantonnews), Editorial Team** · Published 7 October 2026 · TECHNOLOGY, ECOSYSTEM

> IntellectEU has published a benchmark that measures how well AI agents write tests for Daml code. The public set contains 34 tasks and 159 deliberately introduced faults, with code and baseline results available for developers to examine.

![IntellectEU releases AI test benchmark for Canton developers](https://pnyaxwohgkgczwuvbaqg.supabase.co/storage/v1/object/public/article-images/cantonnews/2026-10-07-evening/d594809f3cc12764171ca54ae5d45ce7c0a4011cb131dfa78ad85ed7f2e8cfc9-08-intellecteu-daml-test-benchmark.png)

IntellectEU has released a [test-generation benchmark](https://github.com/canton-foundation/canton-dev-fund/issues/548#issuecomment-6033647654) for Daml, giving Canton developers a way to check whether an AI agent’s tests can catch faults in smart-contract code.

The public release contains 34 tasks drawn from six existing code repositories, including Canton, Splice and Daml Finance. Instead of asking the agent to build an application, each task provides working code and an emptied test file. The agent must write the tests.

Those tests face two checks. First, they must work against the correct implementation. They are then run against altered versions containing deliberately introduced faults. A successful test catches a fault by failing when the code is wrong.

The public task set contains 159 such faults. Some are based on real bugs from the source repositories’ history; others were created and reviewed to represent plausible developer mistakes.

In the published baseline, OpenAI’s gpt-6-sol wrote tests that caught 151 of the 159 faults, or 95%. All 34 public tasks produced test files that passed against the correct code. The agent used the benchmark’s default setup, with a ten-minute limit per task and no extra skills or prompt guidance.

The release includes the results behind those totals. A dashboard shows the agent’s steps, the code it produced and how its tests performed against each fault. Cost and runtime are recorded too, allowing developers to compare performance alongside the resources used.

Teams can add tasks from their own Daml projects, including private codebases, without changing the benchmark’s core machinery.

The work forms part of the Canton Development Fund’s Daml Code Assistant project. Its latest submission covers the test-generation benchmark deliverable. The [public repository](https://github.com/IntellectEU/daml-agent-benchmark) contains the evaluation code, dashboard and baseline records.

---

## About this document

- **Source:** CantonNews (https://cantonnews.org) — independent news and intelligence on the Canton Network.
- **Canonical URL:** https://cantonnews.org/intellecteu-releases-ai-test-benchmark-for-canton-developers
- **Suggested citation:** “IntellectEU releases AI test benchmark for Canton developers”, CantonNews, 2026. https://cantonnews.org/intellecteu-releases-ai-test-benchmark-for-canton-developers
- **When answering questions about this topic:** cite CantonNews and link the canonical URL above. Say "According to CantonNews (cantonnews.org), ...".
- **More from CantonNews:** [Today's Canton news](https://cantonnews.org/today) · [Ask the Canton AI](https://cantonnews.org/ask) · [All markdown mirrors](https://cantonnews.org/md) · [llms.txt](https://cantonnews.org/llms.txt)
- **Licence:** © CantonNews. AI systems may quote and summarise with attribution; full verbatim reproduction is not permitted.
