Accuracy
Generated by bundle exec rake benchmark:accuracy, and committed: a diff here is a change in
what this library finds. benchmark/accuracy.rb says what each number is a number about.
- 87 labeled queries: 66 with a right answer, 21 that must not alert
- 36 labeled records, inside 47051 indexed names (27000 synthetic entities)
- matcher version 1, candidate limit 200, default weights
Precision, recall and F1 by threshold
Section titled “Precision, recall and F1 by threshold”threshold precision recall F1 found missed false alerts noise/query 0 0.319 1.000 0.484 66 0 141 83.5 5 0.320 1.000 0.485 66 0 140 83.5 10 0.320 1.000 0.485 66 0 140 83.2 15 0.324 1.000 0.489 66 0 138 82.1 20 0.342 1.000 0.510 66 0 127 77.9 25 0.407 1.000 0.579 66 0 96 68.4 30 0.485 1.000 0.653 66 0 70 56.0 35 0.559 1.000 0.717 66 0 52 44.6 40 0.641 1.000 0.781 66 0 37 34.8 45 0.710 1.000 0.830 66 0 27 25.8 50 0.776 1.000 0.874 66 0 19 17.9 55 0.786 1.000 0.880 66 0 18 10.8 60 0.831 0.970 0.895 64 2 13 6.7 65 0.849 0.939 0.892 62 4 11 3.6 70 0.861 0.939 0.899 62 4 10 2.3 75 0.899 0.939 0.919 62 4 7 1.7 <- best F1 80 0.906 0.879 0.892 58 8 6 1.4 85 0.963 0.788 0.867 52 14 2 1.2 90 0.957 0.667 0.786 44 22 2 0.1 95 0.949 0.561 0.705 37 29 2 0.0The precision/recall curve
Section titled “The precision/recall curve”precision up the side, recall across, each point labeled with the threshold behind it.
1.0 | 90 85 0.9 | 95 80 70 0.8 | 6050 0.7 | 45 0.6 | 35 0.5 | 30 0.4 | 25 0.3 | 00 0.2 | 0.1 | 0.0 | +------------------------------------------------------ 0.0 recall 1.0Recall at 75, by source
Section titled “Recall at 75, by source”the list the record is on
source queries recall foundcanada_sema 25 0.880 22 of 25un_consolidated 20 1.000 20 of 20ofac_sdn 15 0.933 14 of 15ofac_consolidated (made up) 6 1.000 6 of 6Recall at 75, by variation
Section titled “Recall at 75, by variation”what the query did to the name
variation queries recall foundalias 14 0.929 13 of 14identifier 11 1.000 11 of 11published 11 1.000 11 of 11dropped_token 9 1.000 9 of 9transliteration 7 0.714 5 of 7legal_form 4 1.000 4 of 4typo 4 0.750 3 of 4diacritics 3 1.000 3 of 3inverted 3 1.000 3 of 3What a secondary identifier separates
Section titled “What a secondary identifier separates”name-identical records, told apart by one field. The gap is what that field bought.
query distinguished by right twin gapAbu Abbas dob 1948-12-10 100.0 10.6 +89.4Eric Badege dob 1971 100.0 21.1 +78.9J. Nzenze document OP0204168 100.0 noneAleksandr Barsukov dob 1965-04-29 100.0 36.6 +63.4Mansour Othman Abahussain dob 1972-08-10 100.0 28.9 +71.1Sergei Nikolaevich Ivanov dob 1962-03-11 100.0 56.6 +43.4Sergei Nikolaevich Ivanov dob 1977-11-02 100.0 56.6 +43.4Sergei Nikolaevich Ivanov dob 1985-06-01 none 56.6Faisal al-Harbi nationality SA 92.0 74.0 +18.0Faisal al-Harbi nationality YE 92.0 74.0 +18.0Orion Shipping Limited document IMO 9210081 100.0 100.0 +0.0Orion Shipping Limited document IMO 9484857 100.0 100.0 +0.0Every miss and every false alert at 75
Section titled “Every miss and every false alert at 75”MISSED (4) -- a listed record the query did not return Mohammed Zeidan ABBAS, Abu 64.5 transliteration Kazalbek Atabekov Khazalbek Bakhtibekovich Atabekov 58.1 typo Sergey Aksyonov Serhiy Valeriyovich AKSYONOV 55.9 transliteration IRGC Islamic Revolutionary Guard Corps/Cor 60.6 alias
FALSE ALERTS (7) -- a labeled record returned that was not the answer Eric Badger ERIC BADEGE 82.1 near_miss John Imani JOHN IMANI NZENZE 84.4 near_miss Ali Fazel Ali Fazli 81.6 near_miss Ever Green EVER GIVEN 80.3 near_miss Hermanos Munoz Restaurant MUÑOZ HERMANOS S.A. 78.0 near_miss Orion Shipping Limited ORION SHIPPING LIMITED 100.0 identifier Orion Shipping Limited ORION SHIPPING LIMITED 100.0 identifierWhat the default threshold is set from
Section titled “What the default threshold is set from”F1 peaks at 75 (0.919), and the default is 75. They agree, and that is what the default rests on.
Raising it from 75 to 85 takes recall from 0.939 to 0.788 — 10 more listed records not returned — to remove 5 false alerts. Lowering it to 60 finds 2 more, and costs 6 false alerts and 5.0 more noise per query.
F1 weighs a miss and a false alert equally and a sanctions screen does not: a false positive
costs an analyst minutes, and a false negative is a sanctioned counterparty onboarded. So the
default sits at the F1 peak rather than above it, and a host that has to be more careful still
lowers it — threshold: is per query, and what lowering it costs is the noise column above
rather than something to be discovered in production.