In August we tested our DAML tool against two real audits. The obvious next question was whether our Solidity tool holds up the same way, so we gave it a harder exam: a public Sherlock contest, where a crowd of researchers had already been through the code and the judges had already decided what counts.

The codebase was Ammplify, Sherlock contest #1054 from September 2025. It's a liquidity layer on top of Uniswap V3: makers deposit liquidity, takers borrow it, and everything is tracked in a segment tree of nodes with its own fee accounting. Lots of math, lots of state, and lots of places for the numbers to disagree. A good test.

The setup

  • Scope: 39 Solidity files, about 3,900 nSLOC: 8 contracts, 26 libraries and 10 interfaces
  • Ground truth: the 36 issues the contest judges marked valid, 13 High and 23 Medium, one per issue family, with their final severities
  • Tool: our Solidity security tool (internally, soliditysec) version 1.9.0, running three passes

Matching was strict. Four separate matchers compared the tool's report with each judged issue, and a finding only counted if it had the same root cause and a similar impact. Something in the right neighbourhood but describing a different bug counted as a miss.

The results

33 / 36
valid issues caught, same root cause and impact
11 / 13
High-severity issues caught (plus one partial)
22 / 23
Medium-severity issues caught (plus one partial)
33 caught, 2 partial, 1 missed
One square per valid issue, grouped by the judges' final severity
High · 13 issues
Medium · 23 issues
Caught, same severity Caught, rated differently Partial Missed
“Rated differently” means the tool found the bug but scored it higher or lower than the judges did.

Finding the bug is one half of the job. Rating it is the other, and that's where the tool was weaker:

Judges' severityTotalTool: HighTool: MediumTool: LowPartialMissed
High1374011
Medium23414410
Total361118421

Of the 11 Highs it caught, it rated 7 as High and 4 as Medium. Mediums mostly landed on Medium, with 4 pushed up to High and 4 down to Low. Count it up and the tool rated 8 issues lower than the judges and 4 higher. When it gets a severity wrong, it's more often too cautious than too harsh.

What it caught

A few of the matches, to give a sense of the kind of bug this is. None of these is a missing modifier or a textbook reentrancy:

Judged issueSeverityWhat the tool said
Borrow fee uses an annual rate as a per-second rateHighThe fee curve is applied per second with no 365-day divisor, so borrow fees are inflated about 31.5 million times
Taker collateral can be stolen by re-entering through the payment callbackHighPayments are priced before the payer's callback, the Uniswap mint runs after it, and a payer that moves the price in between makes the shared balance pay the difference
Attackers can drain the protocol with a fake poolHighPool addresses are never checked against the Uniswap factory, so a fake pool can report whatever it likes and get paid from everyone's tokens
Uninitialized Uniswap ticks break the inside-fee mathHighA checked subtraction on modular fee growth either reverts forever, freezing funds, or credits fees nobody earned
The settle walk and the modify walk take different routesHighOne walk treats the upper bound as inclusive and the other doesn't, so some changed nodes are never synced to Uniswap
Utilisation wraps to zero at 100%MediumNarrowing to uint64 turns 100% utilisation into 0, so takers pay the minimum rate exactly when the pool is fully borrowed

The three it didn't fully get

Missed: takers overcharged when a range spans several nodes (High). Ammplify values each node of the tree at its own midpoint price and adds the results up. One taker range split across several nodes therefore gets charged very differently from the same range sitting in a single node. The tool found other problems in the same pricing code, but not this one. The bug is that the pieces don't add up to the whole, and nothing in the tool was asking that question.

Partial: first-depositor share inflation (High). The tool found the right root cause: per-node compounding shares with no minimum and no virtual offset. But it rated the impact as dust, up to one share's worth. The judges' version chains it with a donation and a shrink to a single share, which wipes out the victim's whole deposit. The tool saw the bug and underestimated what it could do.

Partial: tokens that return false on success (Medium). The contest README listed “missing return values” as a supported token behaviour. The tool flagged the failing transfer helper, but only on one function, not across every path that uses it.

How to read these numbers

We're proud of 33 out of 36. We also know how easy a number like that is to oversell, so:

  • The report is big. It holds 325 findings: 26 High, 87 Medium, 192 Low and 20 Informational. Many of them are several angles on the same bug. Some judged issues matched four or five entries, and plenty of the rest won't survive a human review. A report like this is a map for researchers, not something to hand straight to a client.
  • Nothing was executed. No proof-of-concept tests were written or run in this benchmark. Every finding is a trace through the source code. Proving the serious ones is still a human's job.
  • One codebase is one data point. Ammplify suits a tool that reads math carefully. A protocol whose bugs live in off-chain assumptions or cross-protocol economics might score differently.

The tool finds the bugs. Researchers decide which ones are real, how bad they are, and how to prove it. That's the split we use on every audit.

What we changed

Every miss and every wrong severity became a rule change in version 1.10.0, written for the type of bug rather than for Ammplify:

GapWhat the tool now asks
Pieces that don't add up to the wholeWhen a protocol splits something into pieces (tree nodes, buckets, tranches, epochs) and values each one separately, does the sum equal the value of the whole?
Hidden share poolsFind every per-key share pool, not just the main vault, and chain enter-first, shrink-to-one-share and donate before calling a rounding loss dust
Declared token behavioursFor every token behaviour the docs say is supported, check every call site and every transfer helper, not just the first one
Severity: masked by another bugRate each bug as if every other reported bug had been fixed the way the code intends
Severity: attacker actions counted as luckA swap, flash loan or front-run the attacker performs is part of the attack, not an external condition that lowers severity
Severity: functions that can never workOwnership hand-overs that always revert and ignored recipient parameters are broken functionality, not Low

One caveat: these rules came from this benchmark, so Ammplify can't measure them. The real test is the next codebase the tool has never seen.

How we use it

On a Solidity audit the tool runs first. It reads every function, builds a map of where the money moves, and hands our researchers a ranked list of places to dig. Then the researchers do what it can't: throw out the noise, prove what's real, get the severity right, and go after the bugs no rule describes yet.

You still need people for an audit. The tool means less of their time goes on the patterns we already know about. (More on that split in AI Writes Better Smart Contracts Now. You Still Need an Audit.)

Shipping Solidity soon? See our Solidity audit service or tell us what you're building.