Class: ActiveSanction::Normalizer::Form

Inherits:
Object
  • Object
show all
Extended by:
T::Sig
Defined in:
lib/active_sanction/normalizer/form.rb

Overview

A name in both of the forms a screening decision needs: the string the publisher wrote, and the folded string a comparison actually runs against.

form = ActiveSanction::Normalizer.call("O'Brien, Seán")
form.original   # => "O'Brien, Seán"
form.value      # => "o brien sean"
form.tokens     # => ["o", "brien", "sean"]

org = ActiveSanction::Normalizer.call("Rosneft Oil Company", type: :organization)
org.value       # => "rosneft oil"
org.type        # => :organization

Both halves travel together because both are needed at different ends of the same query. The scorers (#28, #29) compare value; the index (#31) keys on tokens; and what a compliance user reads in a hit is original, in the government's own capitalization and punctuation. A report that quotes the folded string instead is quoting this library rather than the list, which is not something anyone can take to an examiner.

Instances are frozen on construction and compare by value.

The pipeline

Five stages, applied in this order to indexed names and query names alike -- see Normalizer for why that sameness is the whole point -- and a sixth that runs only for a caller who said what kind of entity the name belongs to:

  1. Unicode NFKD. Decomposes é into e + combining acute, and folds the compatibility forms a publisher's export tooling emits: full-width ABC becomes ABC, Ⅳ becomes IV, ① becomes 1, and the no-break spaces scattered through the delimited lists become ordinary ones.

  2. Strip combining marks. What stage 1 separated is dropped, so Bélarus is belarus and Müller is muller. Arabic gets the same treatment for free and wants it: the harakat are optional in writing, so one publisher's مُحَمَّد and another's محمد have to fold together, and أ decomposes to a bare alef rather than staying a third spelling of the same first letter.

  3. Casefold, Unicode-aware: String#downcase(:fold) rather than String#downcase, which is what turns Straße into strasse instead of leaving a ß that no query will ever be typed with. 3b. Transliterate the letters NFKD cannot help with -- see TRANSLITERATIONS.

  4. Punctuation to spaces, not to nothing: Al-Qaida is al qaida and O'Brien is o brien. Splitting is the conservative direction. A hyphen and a space are written interchangeably across these lists, so joining Al-Qaida into alqaida would make it unreachable from the al qaida a caller types, while splitting it leaves both sides as the same two tokens for the token ratios (#29) to work on.

  5. Collapse whitespace and strip, which is what tokens is: the folded string split on whitespace, with value its single-spaced join.

  6. Drop the tokens that carry no identifying information, given a Stoplist: LTD and COMPANY from an organization, SHAYKH from a person, nothing at all from either without one. Stage 6 is the only one that depends on something outside the string, which is why it arrives as an argument -- see Dictionary for what is on the lists and for the particles they may never touch.

    A name that folds away entirely keeps its unstripped tokens. An organization called "The Company" is a poor name to screen on and a worse one to index as the empty string, which matches everything or nothing depending on which scorer sees it first.

What it deliberately does not do

Non-Latin script is not transliterated. Cyrillic, Arabic, Han, Kana and Hangul come out of here casefolded and stripped of marks, in their own script. Путин does not become putin, so a Cyrillic name matches a Cyrillic query and nothing else. That is a real recall limitation, and it is stated rather than papered over.

What makes it survivable is that these lists publish a non-Latin name as an additional variant rather than instead of a Latin one -- the UN's ORIGINAL_SCRIPT aliases and Canada's Cyrillic ones both sit on records that carry a romanized name too, which is the one an English-language query finds. Romanization itself is a per-script problem with several competing standards for Cyrillic alone, and guessing at it costs precision everywhere, so v1 does not. Double Metaphone (#30) covers the case this actually leaves open, which is one name romanized two ways.

One consequence worth knowing: NFKD decomposes Hangul syllables into jamo, so 김정은 folds to its letters rather than its syllable blocks. Nothing downstream cares -- both sides of a comparison are folded the same way -- but the value is not the string a Korean reader would type.

Instance Attribute Summary collapse

Instance Method Summary collapse

Constructor Details

#initialize(original, stoplist: nil) ⇒ void

Untyped for the reason the rest of the model is: what arrives here is a publisher's text as whatever the parser made of it. Anything that responds to to_s works, which includes Name -- Name#to_s is its value -- so an indexer can hand over the object it already has.

stoplist is stage 6, and comes from a Dictionary rather than from here: which tokens carry no information is a property of the entity type and of the lists in force, neither of which the string knows.

Building a Form directly skips the cache; Normalizer.call is the entry point everything in the library goes through, and the one that resolves a type into the stoplist for it.

Parameters:

  • original (T.untyped)
  • stoplist (Dictionary::Stoplist, nil) (defaults to: nil)


184
185
186
187
188
189
190
# File 'lib/active_sanction/normalizer/form.rb', line 184

def initialize(original, stoplist: nil)
  @original = T.let(-original.to_s, String)
  @type = T.let(stoplist&.type, T.nilable(Symbol))
  @tokens = T.let(fold(@original, stoplist), T::Array[String])
  @value = T.let(-@tokens.join(" "), String)
  freeze
end

Instance Attribute Details

#original ⇒ String (readonly)

The publisher's string, untouched. This is what a hit is reported in.

Returns:

  • (String)


150
151
152
# File 'lib/active_sanction/normalizer/form.rb', line 150

def original
  @original
end

#tokens ⇒ Array<String> (readonly)

value split on whitespace. Frozen, and the array the token ratios (#29) and the inverted index (#31) read rather than splitting again per comparison.

Returns:

  • (Array<String>)


162
163
164
# File 'lib/active_sanction/normalizer/form.rb', line 162

def tokens
  @tokens
end

#type ⇒ Symbol? (readonly)

The entity type this name was folded for, or nil when the caller did not say. It is what decides stage 6, and it travels with the form because two folds of the same string under different types are two different answers -- which is also why value is part of #==.

Returns:

  • (Symbol, nil)


169
170
171
# File 'lib/active_sanction/normalizer/form.rb', line 169

def type
  @type
end

#value ⇒ String (readonly)

The folded form: lowercase, unmarked, punctuation-free, single-spaced. Empty when the original carried nothing a comparison can use -- see #empty?.

Returns:

  • (String)


156
157
158
# File 'lib/active_sanction/normalizer/form.rb', line 156

def value
  @value
end

Instance Method Details

#==(other) ⇒ Boolean Also known as: eql?

Class is part of the comparison to keep #== and #hash agreeing, which is what Hash and Set rely on -- and the index is built out of both.

Two forms are equal when they came from the same original and folded to the same value. The second half is not redundant now that stage 6 exists: value is a pure function of the original and the stoplist, so "Rosneft Oil Company" folded as an organization and the same string folded as nothing in particular are two different answers rather than one.

Note that this makes two differently-written names that fold to the same string unequal as forms while comparing as identical for matching, which is the distinction the whole pipeline rests on.

Parameters:

  • other (T.untyped)

Returns:

  • (Boolean)


218
219
220
221
222
# File 'lib/active_sanction/normalizer/form.rb', line 218

def ==(other)
  return false unless other.instance_of?(self.class)

  original == other.original && value == other.value
end

#empty? ⇒ Boolean

True when nothing survived the fold: a name of "---", of punctuation, of an emoji, or of whitespace alone. It happens in real data, and it matters because such a name cannot be indexed and cannot be scored -- every comparison against it is meaningless rather than merely bad. The index (#31) skips these; the alternative is a record that matches everything or nothing depending on which scorer sees it first.

Returns:

  • (Boolean)


199
# File 'lib/active_sanction/normalizer/form.rb', line 199

def empty? = value.empty?

#hash ⇒ Integer

Returns:

  • (Integer)


226
227
228
# File 'lib/active_sanction/normalizer/form.rb', line 226

def hash
  [self.class, original, value].hash
end

#inspect ⇒ String

Returns:

  • (String)


231
232
233
# File 'lib/active_sanction/normalizer/form.rb', line 231

def inspect
  "#<#{self.class} #{original.inspect} => #{value.inspect}>"
end

#to_s ⇒ String

Returns:

  • (String)


202
# File 'lib/active_sanction/normalizer/form.rb', line 202

def to_s = value