We recently had a great chat with a developer building on the Canton Network, and at some point they shared an audit report that a well-known security firm had done for a DAML project. Afterwards we found a second, public DAML audit by the same firm.

Security people have a strange idea of fun. Ours was: what happens if we point our own DAML security tool at the same code and compare notes?

This post covers how that went, what the numbers do and don't mean, and, probably more useful if you build on Canton, what these findings say about where DAML bugs like to hide.

The setup

Two codebases, both audited by the same firm:

  • Codebase A: 19 files, 5 issues in the audit report
  • Codebase B: 34 core in-scope files, 7 issues in the audit report

We ran our tool (internally we call it damlsec) on the in-scope code of each one, then went through its output and compared it with the reports. We matched on root cause, not on keywords. If the tool found something nearby but not the actual issue, we counted it as a miss.

We're not naming the projects or the firm. This is about the tool, not about them.

A quick word on how the tool works: it generates a long list of candidate issues, checks each one, and throws out anything that doesn't hold up or can't cause material harm. On these two codebases it went from 97 candidates to 26 findings, and from 92 candidates to 42.

The results

10 / 12
audit issues found by the tool
10 / 11
leaving out the naming issue with no security impact
189 → 68
candidates validated down to reported findings
Codebase ACodebase BTotal
Issues in the audit report5712
Found by our tool3710
Candidates before validation9792189
Findings the tool reported264268
From candidates to findings to matches
Count per codebase, on one shared scale
Candidates, reported findings and audit issues found, per codebase Codebase A: 97 candidates, 26 reported findings, 3 of 5 audit issues found. Codebase B: 92 candidates, 42 reported findings, 7 of 7 audit issues found. 0 25 50 75 100 Codebase A Codebase A: 97 candidates 97 candidates Codebase A: 26 reported 26 reported Codebase A: 3 of 5 audit issues 3 of 5 audit issues Codebase B Codebase B: 92 candidates 92 candidates Codebase B: 42 reported 42 reported Codebase B: 7 of 7 audit issues 7 of 7 audit issues
Validation threw out roughly two thirds of the candidates. The audit issues it found sit inside the reported findings.

And issue by issue:

Issue in the audit reportDid our tool find it?
Confirmations could be created without any permission checkYes, with the same root cause and fix
A multi-party approval could be configured as 1-of-1Yes, and it went further: the threshold could also be set to 0, so actions could execute with no confirmations at all
Audit trail records lacked contextYes
Expired confirmations could still be usedNear miss (more on that below)
A naming issue in an identifier fieldNo, on purpose (also below)
Votes could be changed at any time, with no cooldownYes, flagged independently by several of the tool's agents
An election could be won with half the quorum instead of a majorityYes, exact match
Validator rewards could go to a third partyYes, same root cause: the reward coupon wasn't tied to the validator
Expiry was checked on some paths but not othersYes, caught by a check that looks specifically for this kind of asymmetry
A stepped fee schedule charged the wrong feeYes, plus a second bug in the same fee logic
A creation fee was applied inconsistently on burnYes, at the same place in the code
A secret was stored without being hashedYes, exact match

That's 10 of 12. Leave out the naming issue, which has no security impact, and it's 10 of 11.

To be clear, this isn't a dig at the firm. Their reports were solid and all twelve issues were real, which is exactly what made them a good benchmark.

The two misses

The first was a naming issue: an identifier field whose name could have been better. Fair point for code quality, but nobody loses money over it. The tool rated it informational, and by design it filters out anything below a material-harm threshold. We're not going to lose sleep over a field name.

The second was closer. The report pointed out that expired confirmations could still be used. Our tool flagged related expiry problems in the same code, but treated this exact one as informational, so it didn't make the final list. It's a low-impact issue, but a miss is a miss, so we counted it as one.

How to read these numbers

We're happy with this result, and we also know how easy it is to oversell one. So, some honest caveats:

  • Twelve issues across two codebases is a small sample. It's an encouraging result, not a benchmark study.
  • The tool also reported 50+ issues that aren't in either report. We're validating them now. Some will be real. Some will be the tool being dramatic.
  • The matches weren't all at the top of the list. One was the tool's #2 finding, another was #28. A match at #28 is still a match, but only if someone reads all the way to #28.

That last point is why we see the tool as the start of an audit, not a replacement for one. (We wrote more about what AI tools can and can't do for security.)

Where DAML bugs hide

DAML gets a lot right by design. Authorization is built into the language, and a contract can't be created without the consent of its signatories. But these twelve findings show that a good model doesn't stop you from writing the wrong rules. If you build on Canton, these are the questions worth asking about your own code:

QuestionWhat to check
Who can create this contract?If a template's only signatory is the party creating it, anyone can create one, signed by themselves. That's fine, unless another workflow treats the contract's existence as an approval. Make sure that workflow also checks who signed it.
Does the twin choice check it too?Expiry, authorization and fee checks tend to get added to one choice and forgotten in another that does the same job. Example below.
What happens with a threshold of 0?Put configuration limits in the template's ensure clause, like ensure threshold > 0 && threshold <= length approvers, or a higher minimum if the whole point is multi-party approval. Then a bad configuration can't even be created.
What happens at exactly half?Governance and fee logic break at the boundaries. Test exactly half the votes, one vote over half, and the first unit of every fee tier.
Who can read this field?Every stakeholder on a contract can see all of its data. If a value is supposed to stay secret, like the first half of a commit-reveal, store a hash and reveal the value only when it's needed. “Secret” is a strong word for something every observer can read.

The twin-choice problem is the easiest one to miss, because each choice looks fine on its own:

-- Inside a Voucher template with fields holder and expiresAt
choice Redeem : ()
  controller holder
  do
    now <- getTime
    assertMsg "voucher expired" (now < expiresAt)
    pure () -- pay out the voucher

choice RedeemForCredit : ()
  controller holder
  do
    pure () -- also pays out, but nobody checks the expiry

Whenever you add a check to one choice, search for every other choice that consumes the same contract or has the same effect, and ask whether it needs the check too.

How we use it

The tool is the first pass in our DAML audits. It goes through the whole codebase quickly and hands our researchers a ranked list of places to dig. Then the researchers do the part a tool can't: confirm what's real, work out how bad it is, and go looking for the bugs no tool knows how to look for.

Audits like these cost real money. Letting a tool cover the known patterns means the hours you pay for go into the parts of your protocol that only a human will catch.

Building on Canton or DAML? We'd love to help. Get in touch and tell us what you're working on, or see our DAML audit service.