How matching works
This page is for someone who has already screened a name and now has to defend the result, or tune it. It does not tell you what to type — where you want that, Get started and Guides do. It explains why the number came out the way it did.
Every hosted competitor in this market reports a score and a vague summary of which fields contributed. This library reports reasons that sum to the score exactly, stamped with the weights and the matcher version that produced them. That is the most defensible claim it makes, and this page is the argument for it.
The pipeline, end to end
Section titled “The pipeline, end to end”A screening call is five stages, each narrowing or explaining what the one before it did. Fold the query once. Retrieve the names worth comparing. Score each candidate against four string algorithms and a phonetic pass. Adjust for what the record says beyond the name. Report the reasons — which are the score, not a commentary on it.
Normalizer → Index → NameScore → Adjustments → Reason. Stage numbers below match the section headings on this page.
Normalizing a name
Section titled “Normalizing a name”Screening compares folded strings, never published ones. ActiveSanction::Normalizer
is stage one, and the only place in the library a name is folded at all:
require "active_sanction"
ActiveSanction::Normalizer.call("O'Brien, Seán").value# => "o brien sean"Five stages, in order: Unicode NFKD, strip the combining marks it separated,
casefold, punctuation to spaces, collapse whitespace — plus one table for the
Latin letters decomposition cannot reach, so Bjørn still meets Bjorn and
the same goes for ł, đ, þ, æ, ı and ə. Told what kind of entity a
name belongs to, a sixth stage drops the tokens that identify nothing rather
than which entity: legal forms and function words from an organization,
honorifics from an individual — “Public Joint Stock Company Gazprom”, “PJSC
Gazprom” and “Gazprom” all fold to gazprom. A fixed list of particles —
bin, ibn, al, von, de, and a dozen more — is never stripped,
because “Osama bin Laden” without bin is not a shorter name, it is a
different one.
There is one code path, and that is the point. A matcher whose index and
query fold differently does not fail loudly; it silently stops matching, on
exactly the records the difference touches. Both sides of every comparison
call the same Normalizer.call, so a stoplist entry added today changes what
matches today, everywhere, rather than on one side of a comparison and not
the other.
Be explicit about the limit. Non-Latin script is not transliterated.
Путин does not fold to putin — Cyrillic, Arabic, Han, Kana and Hangul
come out casefolded and stripped of marks, in their own script, so a
Cyrillic name matches a Cyrillic query and nothing else. What makes that
survivable is that these publishers ship a romanized name as an additional
alias rather than instead of the original — the UN’s ORIGINAL_SCRIPT
entries and Canada’s Cyrillic ones sit on records that carry a Latin
spelling too, which is the one an English-language query finds. Guessing at
a romanization costs precision everywhere it is tried, for every script, and
buys recall on the one record a publisher has not already covered — see
the case one name transliterated two ways does not solve,
below.
Candidate retrieval
Section titled “Candidate retrieval”The corpus is roughly 46,000 searchable name strings. Scoring all of them against a query costs a few hundred milliseconds per call, which is fine once and hopeless for a service — so an inverted index narrows the field first, the way a search engine does, rather than scanning.
A folded name is described by three feature spaces — its tokens, its
character trigrams, and its Double Metaphone keys — and any indexed name
sharing a single feature with the query is a candidate. Trigrams catch a
typo a token match would miss; phonetic keys catch a name filed under a
different spelling of itself, which is the retrieval half of the
transliteration limit below. Candidates
are then ranked by a cosine over those three spaces, each feature weighted
by how rare it is — a shared mohammed says almost nothing on a corpus
where thousands of names carry it; a shared surname held by two says nearly
everything — divided by how much of each name the query actually accounts
for, so a short name matched completely outranks a long one matched partly.
Recall is this stage’s whole job, and that is why the cap exists at all.
A name this stage does not retrieve is never compared to anything by the
stages after it — it cannot score badly, it does not appear, and there is no
signal to a caller that it happened. So retrieval is generous on purpose,
capped at 200 candidates by default, and the number is set from a measured
recall curve rather than a round one: it is where recall against a
deliberately damaged query set stops improving and latency keeps costing.
ActiveSanction.configure { |c| c.candidate_limit = 500 } raises it for a
caller who has the budget and wants the margin.
The blend, and what it cannot fix
Section titled “The blend, and what it cannot fix”ActiveSanction::Scorer::NameScore is stage three: four string algorithms
and a phonetic pass, blended into one similarity on the 0..100 scale a score
is read on.
| Component | Default weight | What it catches |
|---|---|---|
token_set | 0.45 | Every word of the shorter name appearing in the longer one — the shape of nearly every honest partial query. |
token_sort | 0.25 | The same words in a different order — a record published surname-first against a query typed given-name-first. |
jaro_winkler | 0.15 | A shared prefix and a small edit distance — the cheapest of the five to compute, and the first one measured. |
levenshtein | 0.1 | Raw character edits, no rearrangement forgiven — the brake that keeps kim jong un apart from kim yong chol, where the token ratios cannot tell the two apart. |
phonetic | 0.05 | Shared Double Metaphone keys — real evidence, and weak evidence, since HSN is the key for both HUSSEIN and HASSAN. |
The five are shares of one weighted mean, held to summing to 1.0. Every
number lives in ActiveSanction::Scorer::Weights, with the reason it is
what it is, and a host can change any of it:
ActiveSanction.configure { |c| c.scorer_weights = { dob_conflict: -20.0 } }Why a mean, and not the weighted maximum the well-known Python ratio
uses. token_set returns 1.0 whenever one name’s words are a subset of
the other’s, so a maximum would score the query Mohammed against
MOHAMMED AL-ZAWAHIRI in the nineties. On a corpus where a quarter of the
individuals share a handful of given names, that is not tolerance, it is an
alert queue nobody can work through. A mean puts the same pair in the high
seventies — still high, because the caller’s whole query really is on the
record, which is the honest answer — and what pulls it apart from a true
match is not the name at all. It is the identifiers, below.
One name transliterated two different ways is the case this stage does
not solve. QADHAFI, Muammar against Muammar Gaddafi scores 58.8, and
raising the phonetic share does not fix it — pushing it from 0.05 to 0.15
moves the pair to 66.8, still under any threshold worth setting, while
lifting every common-name near-miss by the same few points. What actually
covers the case is upstream: these lists publish the variants themselves —
OFAC’s Qadhafi record carries QADHAFI, QADAFI, GADAFI and KADAFI
among others — the index keys on Double Metaphone so
a query for one spelling retrieves a record filed under another, and the
scorer takes the maximum over an entity’s names, so the query is scored
against the alias it is actually a spelling of. The residue — a record
carrying one spelling and one only, queried with a different one — is a
real recall limitation, and the honest mitigation is the identifier fields
rather than a bigger number in the blend.
Adjustments, and why absence is never conflict
Section titled “Adjustments, and why absence is never conflict”Name similarity alone puts thousands of people on a list of a few hundred —
a quarter of the individuals on these lists share a handful of given names,
and every one of them scores in the seventies against every other. The
passport number, the date of birth and the nationality are the corrective,
and they are the fields a compliance officer already has in a customer
record. ActiveSanction::Scorer::Adjustments is stage four: points added to
the name score, strongest evidence first.
| Adjustment | Points |
|---|---|
| Passport / national ID exact match | +40 |
| Date of birth, exact full date | +15 |
| Date of birth, year-only overlap | +6 |
| Date of birth, genuine conflict | -35 |
| Nationality agreement | +6 |
| Nationality conflict | -12 |
| Entity type mismatch | filtered out entirely, at any name similarity |
A UN alias graded QUALITY=Low by the Committee itself costs
10 points before the maximum over an
entity’s names is taken — not after — so a good name scoring 85 beats a low-quality one
scoring 90, rather than the other way around.
Every adjustment above fires only when both sides carry the field, and that is the rule everything here obeys. Most records lack most identifiers: Canada publishes no aliases and frequently no date of birth, OFAC’s dates are prose in a remarks field, and the UN grades what it has and says nothing about what it does not. A missing field produces no reason at all — not a small penalty. Treating absence as disagreement would systematically under-score the jurisdictions that publish least, Canada above all, and would hide real hits below the threshold. That is a compliance argument, not a coding-style one.
The practical consequence follows directly: a subject carrying a passport number needs 40 points less name similarity than one carrying nothing, so a query passing only a name leaves most of the library’s discrimination unused — and it is discrimination against the false positives, not against the hits.
Reading a score
Section titled “Reading a score”The reasons sum to the score, rounded once — there is no arithmetic anywhere in this library that can move one without the other:
result.score# => 82.0result.explanation.map(&:to_s)# => ["+42.0 name: matched primary name \"NTAGANDA, Bosco\"",# "+40.0 identifier: passport 750123456 matches"]Do not band by score alone. A 97 that is all name similarity and an 82 with a matching passport number are different findings, and the number does not distinguish them. A score that merely travelled beside the reasons that produced it could come apart from them in a later release and be quietly wrong for a year — so it never travels beside them, it is computed from them, on every read.
Why the threshold defaults to 75
Section titled “Why the threshold defaults to 75”The default was a guess until rake benchmark:accuracy measured it, against
87 labeled queries: real published records queried the way a customer
record spells them, plus the common names and near misses that must not
alert.
threshold precision recall F1 found missed false alerts 60 0.831 0.970 0.895 64 2 13 75 0.899 0.939 0.919 62 4 7 <- best F1 85 0.963 0.788 0.867 52 14 2F1 peaks at exactly the number this library ships, and that agreement is
the whole argument for the default — but it is worth being clear about what
F1 cannot argue. It weighs a miss and a false alert equally, and a
sanctions screen does not: a false positive costs an analyst minutes, and a
false negative is a sanctioned counterparty onboarded. So 75 is a floor
to tune down from, not a ceiling, and threshold: is per query for a
host that has to be more careful still.
Recall is not uniform across the lists, and an averaged figure would hide it — Canada’s 0.880 against the UN’s 1.000 is not the matching’s fault; the list publishes fewer aliases and fewer dates of birth, so there is less to match against.
benchmark/results/accuracy.md
is regenerated by rake benchmark:accuracy and committed. It names every
miss and every false alert at the default threshold, breaks recall down by
what the query did to the name, and is the one place these figures are
rendered — this page argues about what they mean and links there rather
than restating them, so the numbers cannot drift into disagreement with
themselves.
Reproducibility
Section titled “Reproducibility”Four fields make a past screening decision re-derivable, and each is a way the same query could score differently today:
| Field | What it pins down |
|---|---|
snapshot_id | the checksum of the exact list version that answered. Publishers overwrite their files in place, so “the OFAC list” is not a citable thing; a checksum is |
matcher_version | which matching pipeline scored it |
weights | what each signal was worth. A host that retunes dob_conflict changes what every past decision would score today |
query | what was screened, and under what threshold. A hit at 78 means one thing under a threshold of 75, and could not have existed under 85 |
matcher_version (currently "{weightsData.matcher_version}") is
deliberately not the gem version. The gem version moves for a new source
adapter, a storage fix, or a documentation release — none of which change
what a name scores. It is bumped only when a change to the normalizer, the
index, the similarity algorithms or the scorer could move a score,
because an auditor asking “would this screening come out the same today?”
needs the answer to that specific question, not to several others that
happen to share a release number.
Together with the snapshot checksum and the weights, it is what turns a
MatchResult sitting in an audit record into something a compliance
officer can re-derive months later, rather than something they have to
trust.