Loading ForensicBlock
Preparing your blockchain forensics platform...
Preparing your blockchain forensics platform...
Daubert factor 2 — known or potential rate of error
In United States v. Sterlingov, the defense attacked commercial tracing software for lacking a published error rate. It is the reliability factor this industry does not answer. Here is ours, with the ground truth it was measured against and the interval that bounds it.
Read the interval, not the point. 0 of 46 cases produced a value other than the one its primary source establishes. A corpus of this size cannot establish a rate below 7.7%, and we do not claim one. The interval narrows as the corpus grows — that is the incentive, and every case added is published here.
Reported separately, never blended into one number — a small corpus on one engine cannot be hidden behind a large one on another.
| Engine | Cases | Errors | 95% CI on the error rate |
|---|---|---|---|
| btc.coinjoin-structure | 2 | 0 | [0.0%, 65.8%] |
| btc.utxo-edges | 1 | 0 | [0.0%, 79.3%] |
| tokens.classify-asset | 4 | 0 | [0.0%, 49.0%] |
| units.base-to-native | 23 | 0 | [0.0%, 14.3%] |
| units.denomination-from-symbol | 14 | 0 | [0.0%, 21.5%] |
| units.normalize-native | 2 | 0 | [0.0%, 65.8%] |
31 cases require an engine to produce a recorded value. 15 are negative controls requiring a guard not to fire — a guard that deletes real value is harder to notice than one that inflates it, so both directions are measured.
This measures denomination, unit conversion, value apportionment and asset classification — the deterministic engines where every court-fatal defect this platform has shipped actually occurred. It establishes nothing about:
Attribution coverage is our binding constraint: of 4,586 distinct counterparties across our 30 most recent completed investigations, 47 resolve in the entity catalog — one percent. So we asked whether the chain itself could fill the gap: infrastructure and personal wallets have different transaction shapes, and shape needs no catalog. We built the classifier and measured it against catalog-labeled addresses.
Behavioral role inference does not substitute for attribution data at our current observation depth. Only 0.8% of catalog-labeled service addresses carry enough observed transaction shape in our warehouse to classify at all, and on that observable subset the classifier recognises 33.3% (95% CI 24.2%–43.9%). The limit is observation, not the rule set: most addresses are seen only through the slice one trace touched. Deeper per-address history or more labeled ground truth would change this; a better classifier would not. We publish this because a negative result about our own approach is worth exactly as much as a positive one, and a buyer deciding whether to trust derived attribution is entitled to the number.
What this figure is not. It measures whether transaction shape RECOGNISES a known service — not whether a label naming a party is correct. Label correctness remains unmeasured and is listed above as such. Recognising a role and verifying a name are different questions, and the measured one does not cover the unmeasured one.
A confidence score shown anywhere in our product is not an error rate. It states the strength of the inputs behind a specific conclusion under a disclosed rule, and it has not been calibrated against ground truth. Offering one as the other is precisely what this page exists to stop doing.
The methodology this was measured under · Verify a sealed report