Class: ActiveSanction::Normalizer::Form
- Inherits:
-
Object
- Object
- ActiveSanction::Normalizer::Form
- Extended by:
- T::Sig
- Defined in:
- lib/active_sanction/normalizer/form.rb
Overview
A name in both of the forms a screening decision needs: the string the publisher wrote, and the folded string a comparison actually runs against.
form = ActiveSanction::Normalizer.call("O'Brien, Seán")
form.original # => "O'Brien, Seán"
form.value # => "o brien sean"
form.tokens # => ["o", "brien", "sean"]
org = ActiveSanction::Normalizer.call("Rosneft Oil Company", type: :organization)
org.value # => "rosneft oil"
org.type # => :organization
Both halves travel together because both are needed at different ends of
the same query. The scorers (#28, #29) compare value; the index (#31)
keys on tokens; and what a compliance user reads in a hit is
original, in the government's own capitalization and punctuation. A
report that quotes the folded string instead is quoting this library
rather than the list, which is not something anyone can take to an
examiner.
Instances are frozen on construction and compare by value.
The pipeline
Five stages, applied in this order to indexed names and query names alike -- see Normalizer for why that sameness is the whole point -- and a sixth that runs only for a caller who said what kind of entity the name belongs to:
-
Unicode NFKD. Decomposes
éintoe+ combining acute, and folds the compatibility forms a publisher's export tooling emits: full-widthABCbecomesABC,ⅣbecomesIV,①becomes1, and the no-break spaces scattered through the delimited lists become ordinary ones. -
Strip combining marks. What stage 1 separated is dropped, so
BélarusisbelarusandMüllerismuller. Arabic gets the same treatment for free and wants it: the harakat are optional in writing, so one publisher'sمُحَمَّدand another'sمحمدhave to fold together, andأdecomposes to a bare alef rather than staying a third spelling of the same first letter. -
Casefold, Unicode-aware:
String#downcase(:fold)rather thanString#downcase, which is what turnsStraßeintostrasseinstead of leaving aßthat no query will ever be typed with. 3b. Transliterate the letters NFKD cannot help with -- see TRANSLITERATIONS. -
Punctuation to spaces, not to nothing:
Al-Qaidaisal qaidaandO'Brieniso brien. Splitting is the conservative direction. A hyphen and a space are written interchangeably across these lists, so joiningAl-Qaidaintoalqaidawould make it unreachable from theal qaidaa caller types, while splitting it leaves both sides as the same two tokens for the token ratios (#29) to work on. -
Collapse whitespace and strip, which is what
tokensis: the folded string split on whitespace, withvalueits single-spaced join. -
Drop the tokens that carry no identifying information, given a Stoplist:
LTDandCOMPANYfrom an organization,SHAYKHfrom a person, nothing at all from either without one. Stage 6 is the only one that depends on something outside the string, which is why it arrives as an argument -- see Dictionary for what is on the lists and for the particles they may never touch.A name that folds away entirely keeps its unstripped tokens. An organization called "The Company" is a poor name to screen on and a worse one to index as the empty string, which matches everything or nothing depending on which scorer sees it first.
What it deliberately does not do
Non-Latin script is not transliterated. Cyrillic, Arabic, Han, Kana
and Hangul come out of here casefolded and stripped of marks, in their
own script. Путин does not become putin, so a Cyrillic name matches a
Cyrillic query and nothing else. That is a real recall limitation, and it
is stated rather than papered over.
What makes it survivable is that these lists publish a non-Latin name as an additional variant rather than instead of a Latin one -- the UN's ORIGINAL_SCRIPT aliases and Canada's Cyrillic ones both sit on records that carry a romanized name too, which is the one an English-language query finds. Romanization itself is a per-script problem with several competing standards for Cyrillic alone, and guessing at it costs precision everywhere, so v1 does not. Double Metaphone (#30) covers the case this actually leaves open, which is one name romanized two ways.
One consequence worth knowing: NFKD decomposes Hangul syllables into
jamo, so 김정은 folds to its letters rather than its syllable blocks.
Nothing downstream cares -- both sides of a comparison are folded the
same way -- but the value is not the string a Korean reader would type.
Instance Attribute Summary collapse
-
#original ⇒ String
readonly
The publisher's string, untouched.
-
#tokens ⇒ Array<String>
readonly
valuesplit on whitespace. -
#type ⇒ Symbol?
readonly
The entity type this name was folded for, or nil when the caller did not say.
-
#value ⇒ String
readonly
The folded form: lowercase, unmarked, punctuation-free, single-spaced.
Instance Method Summary collapse
-
#==(other) ⇒ Boolean
(also: #eql?)
Class is part of the comparison to keep #== and #hash agreeing, which is what Hash and Set rely on -- and the index is built out of both.
-
#empty? ⇒ Boolean
True when nothing survived the fold: a name of
"---", of punctuation, of an emoji, or of whitespace alone. - #hash ⇒ Integer
-
#initialize(original, stoplist: nil) ⇒ void
constructor
Untyped for the reason the rest of the model is: what arrives here is a publisher's text as whatever the parser made of it.
- #inspect ⇒ String
- #to_s ⇒ String
Constructor Details
#initialize(original, stoplist: nil) ⇒ void
Untyped for the reason the rest of the model is: what arrives here is a
publisher's text as whatever the parser made of it. Anything that
responds to to_s works, which includes Name -- Name#to_s is its
value -- so an indexer can hand over the object it already has.
stoplist is stage 6, and comes from a Dictionary rather than from
here: which tokens carry no information is a property of the entity
type and of the lists in force, neither of which the string knows.
Building a Form directly skips the cache; Normalizer.call is the entry point everything in the library goes through, and the one that resolves a type into the stoplist for it.
184 185 186 187 188 189 190 |
# File 'lib/active_sanction/normalizer/form.rb', line 184 def initialize(original, stoplist: nil) @original = T.let(-original.to_s, String) @type = T.let(stoplist&.type, T.nilable(Symbol)) @tokens = T.let(fold(@original, stoplist), T::Array[String]) @value = T.let(-@tokens.join(" "), String) freeze end |
Instance Attribute Details
#original ⇒ String (readonly)
The publisher's string, untouched. This is what a hit is reported in.
150 151 152 |
# File 'lib/active_sanction/normalizer/form.rb', line 150 def original @original end |
#tokens ⇒ Array<String> (readonly)
value split on whitespace. Frozen, and the array the token ratios
(#29) and the inverted index (#31) read rather than splitting again per
comparison.
162 163 164 |
# File 'lib/active_sanction/normalizer/form.rb', line 162 def tokens @tokens end |
#type ⇒ Symbol? (readonly)
The entity type this name was folded for, or nil when the caller did
not say. It is what decides stage 6, and it travels with the form
because two folds of the same string under different types are two
different answers -- which is also why value is part of #==.
169 170 171 |
# File 'lib/active_sanction/normalizer/form.rb', line 169 def type @type end |
#value ⇒ String (readonly)
The folded form: lowercase, unmarked, punctuation-free, single-spaced. Empty when the original carried nothing a comparison can use -- see #empty?.
156 157 158 |
# File 'lib/active_sanction/normalizer/form.rb', line 156 def value @value end |
Instance Method Details
#==(other) ⇒ Boolean Also known as: eql?
Class is part of the comparison to keep #== and #hash agreeing, which is what Hash and Set rely on -- and the index is built out of both.
Two forms are equal when they came from the same original and folded
to the same value. The second half is not redundant now that stage 6
exists: value is a pure function of the original and the stoplist,
so "Rosneft Oil Company" folded as an organization and the same string
folded as nothing in particular are two different answers rather than
one.
Note that this makes two differently-written names that fold to the same string unequal as forms while comparing as identical for matching, which is the distinction the whole pipeline rests on.
218 219 220 221 222 |
# File 'lib/active_sanction/normalizer/form.rb', line 218 def ==(other) return false unless other.instance_of?(self.class) original == other.original && value == other.value end |
#empty? ⇒ Boolean
True when nothing survived the fold: a name of "---", of punctuation,
of an emoji, or of whitespace alone. It happens in real data, and it
matters because such a name cannot be indexed and cannot be scored --
every comparison against it is meaningless rather than merely bad. The
index (#31) skips these; the alternative is a record that matches
everything or nothing depending on which scorer sees it first.
199 |
# File 'lib/active_sanction/normalizer/form.rb', line 199 def empty? = value.empty? |
#hash ⇒ Integer
226 227 228 |
# File 'lib/active_sanction/normalizer/form.rb', line 226 def hash [self.class, original, value].hash end |
#inspect ⇒ String
231 232 233 |
# File 'lib/active_sanction/normalizer/form.rb', line 231 def inspect "#<#{self.class} #{original.inspect} => #{value.inspect}>" end |
#to_s ⇒ String
202 |
# File 'lib/active_sanction/normalizer/form.rb', line 202 def to_s = value |