Class: ActiveSanction::Matcher
- Inherits:
-
Object
- Object
- ActiveSanction::Matcher
- Extended by:
- T::Sig
- Defined in:
- lib/active_sanction/matcher.rb
Overview
The screening call, and the object a server holds.
matcher = ActiveSanction::Matcher.build(store)
results = matcher.screen(
name: "Bosco Ntaganda",
type: :individual,
date_of_birth: "1973",
countries: %w[CD],
sources: %i[ofac_sdn un_consolidated],
threshold: 75,
limit: 10
)
results.first.score # => 100.0
results.first.snapshot_id # => "sha256:9f86d081884c7d65..."
Stage five, and the only one with nothing after it. It runs the pipeline the other four stages are: fold the query once (Normalizer), retrieve the names worth comparing (Index), score each of them with reasons (Scorer), then filter, sort, cap and stamp. Nothing here decides whether two names are the same person; what it decides is what a caller is handed and what a decision can be defended with.
It is built once and then only read
A Matcher holds an index, the checksum of every list in it, the weights it
scores with and the candidate cap it retrieves with. All of it is fixed at
construction and the object is frozen, so screen allocates locals and
touches nothing shared. A web process builds one at boot and screens from
every thread without a lock:
MATCHER = ActiveSanction::Matcher.build(store) # in an initializer
MATCHER.screen(name: params[:name]) # in a request
Nothing on the query path reads configuration. That is a stronger statement than thread safety and it is the one that matters for an audit: a threshold, a weight or a candidate cap changed halfway through a batch cannot produce a run that is half one set of numbers and half another, because the numbers were read once -- into the Query, and into this.
A sync does not update a matcher
It builds a new one, and the application swaps its reference:
MATCHER = ActiveSanction::Matcher.build(store) # after a sync
A plain reassignment is enough on CRuby, where a reference assignment is
atomic; Concurrent::AtomicReference is the portable spelling. Requests
in flight keep the matcher they started with and finish against one
consistent list version, which is what makes their results re-derivable --
a matcher that mutated underneath a query would produce a result no
snapshot checksum explains. See Index, which is immutable for this reason.
An empty matcher is refused rather than built
Screening against a list that is not there returns a clean report, and a clean report is the most expensive thing this library can get wrong. So a store with nothing in it raises NotSynced at build, a named source that has never been synced raises Storage::MissingSnapshot, and a query naming a source this matcher does not hold raises rather than quietly covering two of the three lists it was asked for.
Where the backend seam goes
This is the Local backend's implementation (#56): Backend::Local#screen
is this call, and a hosted backend answers the same query with the same
MatchResults against data somebody else keeps fresh. Which one answered is
on every result. What a server holds is a Client (#55) rather than one of
these directly, because a client is what pairs an index with the
configuration it was built under; this stays the object to build by hand
when a caller already has an index -- a spec, or a process screening one
name against several list versions of the same store.
Defined Under Namespace
Classes: NotSynced
Instance Attribute Summary collapse
- #backend ⇒ Symbol readonly
-
#candidate_limit ⇒ Integer
readonly
How many names the index hands the scorer per query.
- #index ⇒ Index readonly
-
#instrumenter ⇒ T.untyped
readonly
Where the
:screenevent goes, or nil for nothing listening. -
#snapshots ⇒ Hash{Symbol => String}
readonly
Which lists this matcher holds, and the checksum of each.
-
#verified ⇒ Array<Symbol>
readonly
The lists in here that arrived cryptographically attested -- read from a signed bundle (#57) that verified under a key this installation supplied -- sorted.
-
#weights ⇒ Scorer::Weights
readonly
What each signal was worth when this matcher was built, and what every result it produces records.
Class Method Summary collapse
-
.build(store = nil, sources: nil, weights: nil, candidate_limit: nil, backend: MatchResult::DEFAULT_BACKEND, instrumenter: nil) ⇒ Matcher
A matcher over what a store holds, or over the lists named:.
Instance Method Summary collapse
-
#initialize(index:, snapshots:, weights: nil, candidate_limit: nil, backend: MatchResult::DEFAULT_BACKEND, verified: nil, instrumenter: nil) ⇒ void
constructor
Built by .build, which is what a caller almost always wants.
- #inspect ⇒ String
-
#screen(query = nil, **overrides) ⇒ Array<MatchResult>
The hits, highest score first:.
-
#screen_all(queries, **overrides) ⇒ Array<Array<MatchResult>>
A book of names against one list version:.
-
#size ⇒ Integer
How many names are screened against.
-
#snapshot_id(source) ⇒ String?
The checksum of the list version this matcher holds for a source, or nil for one it does not.
-
#sources ⇒ Array<Symbol>
The lists this matcher screens against, sorted.
-
#verified?(source) ⇒ Boolean
Whether the list this matcher holds for a source was attested.
Constructor Details
#initialize(index:, snapshots:, weights: nil, candidate_limit: nil, backend: MatchResult::DEFAULT_BACKEND, verified: nil, instrumenter: nil) ⇒ void
Built by .build, which is what a caller almost always wants. Taken directly by a caller that already has an index -- a process screening one name against several list versions, or a spec.
248 249 250 251 252 253 254 255 256 257 258 259 260 |
# File 'lib/active_sanction/matcher.rb', line 248 def initialize(index:, snapshots:, weights: nil, candidate_limit: nil, backend: MatchResult::DEFAULT_BACKEND, verified: nil, instrumenter: nil) @index = index @snapshots = T.let(snapshots!(snapshots), T::Hash[Symbol, String]) @verified = T.let(verified!(verified), T::Array[Symbol]) raise NotSynced, "the lists given hold no names to screen against" if index.empty? @weights = T.let(Scorer::Weights.build(weights), Scorer::Weights) @candidate_limit = T.let(candidate_limit!(candidate_limit), Integer) @backend = T.let(backend.to_s.to_sym, Symbol) @instrumenter = T.let(instrumenter.nil? ? ActiveSanction.config.instrumenter : instrumenter, T.untyped) freeze end |
Instance Attribute Details
#backend ⇒ Symbol (readonly)
133 134 135 |
# File 'lib/active_sanction/matcher.rb', line 133 def backend @backend end |
#candidate_limit ⇒ Integer (readonly)
How many names the index hands the scorer per query. See
Configuration::DEFAULT_CANDIDATE_LIMIT -- and note that it bounds a
query's limit: in practice, since a result cannot be returned for a
name that was never retrieved.
130 131 132 |
# File 'lib/active_sanction/matcher.rb', line 130 def candidate_limit @candidate_limit end |
#index ⇒ Index (readonly)
118 119 120 |
# File 'lib/active_sanction/matcher.rb', line 118 def index @index end |
#instrumenter ⇒ T.untyped (readonly)
Where the :screen event goes, or nil for nothing listening. Read once
at construction and frozen with everything else here, which is the rule
the class comment states for the whole query path: a subscriber swapped
halfway through a batch cannot make half of it instrumented. See
Instrumentation.
141 142 143 |
# File 'lib/active_sanction/matcher.rb', line 141 def instrumenter @instrumenter end |
#snapshots ⇒ Hash{Symbol => String} (readonly)
Which lists this matcher holds, and the checksum of each. The stamp on every result comes from here.
103 104 105 |
# File 'lib/active_sanction/matcher.rb', line 103 def snapshots @snapshots end |
#verified ⇒ Array<Symbol> (readonly)
The lists in here that arrived cryptographically attested -- read from a signed bundle (#57) that verified under a key this installation supplied -- sorted. Usually empty, because a list this installation fetched and parsed itself is not attested by anybody.
Kept beside snapshots rather than folded into it because it is a fact
about a different thing: a checksum says which list version answered, and
this says who vouched for it. Every result the matcher produces carries
both. See MatchResult#verified?.
115 116 117 |
# File 'lib/active_sanction/matcher.rb', line 115 def verified @verified end |
#weights ⇒ Scorer::Weights (readonly)
What each signal was worth when this matcher was built, and what every result it produces records.
123 124 125 |
# File 'lib/active_sanction/matcher.rb', line 123 def weights @weights end |
Class Method Details
.build(store = nil, sources: nil, weights: nil, candidate_limit: nil, backend: MatchResult::DEFAULT_BACKEND, instrumenter: nil) ⇒ Matcher
A matcher over what a store holds, or over the lists named:
ActiveSanction::Matcher.build # the configured store
ActiveSanction::Matcher.build(store)
ActiveSanction::Matcher.build(store, sources: %i[ofac_sdn])
Snapshots are read one at a time and each is released before the next is opened, so building never holds every list in memory at once -- and each list's checksum is taken from the very snapshot that was indexed, rather than read separately afterwards, where a concurrent sync could put a stamp on results the list no longer explains.
sources: nil means whatever is stored. Naming a list that has never
been synced raises instead: a run that quietly covers two of the three
lists an application configured is indistinguishable from one that
covers all three, and both report the name clear.
166 167 168 169 170 171 172 173 174 175 |
# File 'lib/active_sanction/matcher.rb', line 166 def build(store = nil, sources: nil, weights: nil, candidate_limit: nil, backend: MatchResult::DEFAULT_BACKEND, instrumenter: nil) store ||= ActiveSanction.config.storage listening = instrumenter.nil? ? ActiveSanction.config.instrumenter : instrumenter built = Instrumentation.instrument(listening, :"index.build", { store: store.class.name }) do |event| index_over(store, sources, event) end new(index: built.fetch(:index), snapshots: built.fetch(:checksums), verified: built.fetch(:attested), weights: weights, candidate_limit: candidate_limit, backend: backend, instrumenter: listening) end |
Instance Method Details
#inspect ⇒ String
334 |
# File 'lib/active_sanction/matcher.rb', line 334 def inspect = "#<#{self.class} #{size} names from #{sources.join(", ")}>" |
#screen(query = nil, **overrides) ⇒ Array<MatchResult>
The hits, highest score first:
matcher.screen(name: "Bosco Ntaganda", threshold: 75)
matcher.screen("Bosco Ntaganda") # a name and nothing else
matcher.screen(query, limit: 25) # a Query, with one option changed
An empty array is a real answer and the common one -- most customers are not on a sanctions list. It is not the same answer as an exception, and everything that could make it a lie rather than a fact raises instead: see the note on an empty matcher above.
One result per entity, not per name
An entity is retrieved once for every one of its names the query looks like, and its score is the best of those names (see Scorer). So each entity is scored once and reported once, in the alias that won.
The order is re-derivable
Score descending, and equal scores by list and then entity id. Ties are not a corner case on this corpus -- a query matching two records of the same name scores them identically -- and which one is listed first has to be the same answer in a year's time.
286 287 288 |
# File 'lib/active_sanction/matcher.rb', line 286 def screen(query = nil, **overrides) run(Query.build(query, **overrides), Time.now.utc) end |
#screen_all(queries, **overrides) ⇒ Array<Array<MatchResult>>
A book of names against one list version:
matcher.screen_all(["Bosco Ntaganda", "Gazprom"], threshold: 80)
matcher.screen_all(customers.map { |c| { name: c.name, dob: c.born_on } })
Index-aligned: the nth element is the nth query's results, and it is an empty array for a name that hit nothing. Deliberately not keyed by name -- a batch of customers contains the same name twice often enough, and a Hash would silently screen one of them and report both.
Every result in the batch carries one screened_at, because a batch is
one screening run: a rescreening of a customer book against a new list
version is a single event in an audit trail, not ten thousand of them a
microsecond apart.
305 306 307 308 309 310 |
# File 'lib/active_sanction/matcher.rb', line 305 def screen_all(queries, **overrides) raise QueryError, "screen_all takes an Array of queries, got #{queries.class}" unless queries.is_a?(Array) screened_at = Time.now.utc queries.map { |query| run(Query.build(query, **overrides), screened_at) } end |
#size ⇒ Integer
How many names are screened against. Names rather than entities -- see Index#size.
331 |
# File 'lib/active_sanction/matcher.rb', line 331 def size = index.size |
#snapshot_id(source) ⇒ String?
The checksum of the list version this matcher holds for a source, or nil for one it does not. Named for the Backend contract (#56), where every backend has to be able to answer it or reproducibility breaks at the seam.
321 |
# File 'lib/active_sanction/matcher.rb', line 321 def snapshot_id(source) = snapshots[Sources::Definition.key!(source)] |
#sources ⇒ Array<Symbol>
The lists this matcher screens against, sorted.
314 |
# File 'lib/active_sanction/matcher.rb', line 314 def sources = snapshots.keys.sort |
#verified?(source) ⇒ Boolean
Whether the list this matcher holds for a source was attested. What every result off that list records.
326 |
# File 'lib/active_sanction/matcher.rb', line 326 def verified?(source) = verified.include?(Sources::Definition.key!(source)) |