We put that claim to the test
Rather than assert it, we measured it. 345 legal questions across Danish and EU law, each with a verifiable correct answer from official sources. These four are the measures a model without a legal database could still win, so they are the fair comparison. Higher is better in every row.
| Measure | eulaw.ai | GPT-5.6 | Claude Sonnet 5 |
|---|
| Acts it named that actually exist | 82% | 15% | declined |
|---|
| Multi-turn conversations handled correctly | 96% | 71% | 73% |
|---|
| Still correct 100 turns into a conversation | 75% | 34% | 0% |
|---|
| Legal substance, judged on the facts stated | 69% | 66% | 59% |
|---|
The first row is the one that matters most. Asked which acts had amended a given law, GPT-5.6 named 59 acts that do not exist. Claude declined the question rather than answer it, which is safer but no more useful. Across the whole set, 9 out of 10 eulaw.ai answers cite the exact act asked about, and none answer without a source.
Read the full benchmark, with every chart and all 345 questions
Measured August 2026 over 345 questions. GPT-5.6 and Claude Sonnet 5 were tested through their APIs without a legal database, web search or retrieval. That is not how a lawyer uses either product, and we will publish a second comparison against the consumer apps with search enabled. We have left out the measures a bare model cannot win by definition, such as naming acts passed after its training cutoff, where the gap is far wider and far less meaningful.