Detect when a publisher changes its format
Notice when a government publisher has changed what its file means, not just whether the file still parses.
The dangerous change is the one where the file still parses cleanly and means something different. 19,321 entities carrying zero passport numbers looks exactly as healthy as 19,321 carrying 23,429 if the only thing anyone counts is records — and screening a passport number against the first returns a clean result for somebody who is actually on the list. A sync would not notice that for months; the doctor exists to.
Run it
Section titled “Run it”report = ActiveSanction.doctor # every configured sourcereport = ActiveSanction.doctor(:ofac_sdn) # onereport = ActiveSanction.doctor(tolerance: 0.05) # report smaller movements
report.ok? # => falsereport.findings # => [Doctor::Finding, ...]report[:ofac_sdn].severity # => :warnreport[:ofac_sdn].profile.fill # => { dates_of_birth: 0.12, identifiers: 0.34, ... }exit report.exit_code # 1 on any error; exit_code(on: :warn) for a build that should fail on warnings too3 sources in 41.07s: 2 with findings, 1 unreadableofac_sdn WARN 3 findings warn remarks coverage 71.4% (was 97.3%): "Passport No. #####" x 1,880 unrecognized warn individuals with a date of birth 12% (was 61%) of 11,704 info unknown SDN_Type "syndicate"; treated as an organization (41 rows)un_consolidated OKeu_fsf ERROR 1 finding error could not be read: ActiveSanction::HttpClient::TimeoutError: execution expiredDoctor::Report#to_h round-trips through JSON, the same as a sync report —
alert on a finding through whatever your job runner already uses rather than
scraping console output.
Run it nightly, on its own schedule
Section titled “Run it nightly, on its own schedule”A doctor invoked by hand only confirms a regression somebody already
suspected. The value is in catching one nobody suspected, which means
something has to run it when nobody is looking:
task doctor_sanctions: :environment do report = ActiveSanction.doctor warn report.to_s exit report.exit_code(on: :warn)endThe doctor never writes anything — no snapshot, no payload cache, no
conditional-GET validators — so it always diagnoses bytes the publisher is
serving right now, and a doctor run never makes a later sync think an
unchanged list is unchanged when the doctor already fetched it. That costs a
full download of every list on every run; run it on its own cadence rather
than folding it into every sync!.
Compare against yesterday’s run, not just the stored snapshot
Section titled “Compare against yesterday’s run, not just the stored snapshot”The doctor’s default baseline is the last stored snapshot, which is free and never goes stale on its own. Two things it measures only exist while a parse is actually running — warning counts and free-text coverage — so a job that wants those tracked week over week keeps its own report:
yesterday = JSON.parse(File.read("doctor.json"))report = ActiveSanction.doctor(baseline: ActiveSanction::Doctor::Report.from_h(yesterday))File.write("doctor.json", JSON.generate(report.to_h))Read warn versus error correctly
Section titled “Read warn versus error correctly”It is not about the size of the movement — it is whether the reading can
be explained by the list changing rather than the file changing. A third
of the records disappearing is a warn: a delisting wave looks exactly like
a truncated download, and deciding automatically that it was the delisting
is how a compliance tool ends up quietly screening against a list it threw
half away. A column that used to hold numbers and now holds company names is
an error, and so is every record on a list losing a field all of them used
to carry — nothing a government does to its own list produces either of
those.
Fill rate is the specific check that catches a clean parse of a changed file, because record counts do not move when a publisher renames or reorders a column but the share of records carrying each field does. A date of birth is measured over individuals alone — an organization never has one — and a swapped-but-not-inserted column is invisible to a declared width, which is why declaring column values, not just column count, matters for a headerless file:
class Ofac < ActiveSanction::Sources::Base floor :remarks_coverage, 0.90endA floor is a coarse backstop for the run that has nothing to compare
against — a first sync, a new source, a store that was cleared — not a
substitute for the snapshot-to-snapshot comparison, which catches drift a
fixed number wide enough to survive years of a growing list cannot.
The upstream canary: the same idea, run on a schedule against real endpoints
Section titled “The upstream canary: the same idea, run on a schedule against real endpoints”Where the doctor is something you run and read, the canary is the same
comparison run automatically: it fetches every registered source on
weekdays and holds each against .github/baselines/<key>.json, opening an
issue when a number moves outside its tolerance.
$ bundle exec rake canary # what every source measures right now$ CANARY_SOURCES=ofac_sdn bundle exec rake canary:refresh # accept the current numbers as the new baselineIt is deliberately excluded from the normal test suite and from CI’s default
run — it reaches seven government endpoints, and a red build should mean
this gem’s own code broke rather than that a publisher rate-limited a runner.
.github/workflows/canary.yml is what runs it on a schedule, and
.github/baselines/README.md
documents the baseline format and the per-key tolerances. When adding a new
source, commit its baseline in the same pull request as the adapter — see
Add a sanctions source.
Nothing is reported until two consecutive runs agree. Government endpoints 403 a non-browser user agent and block cloud IP ranges, and a canary that cried wolf on one bad afternoon would be muted just as fast as a red badge — so a fetch that failed and a file that parsed into something different are kept as separate findings, each run’s report is saved as a workflow artifact, and only a finding both of two consecutive runs made is opened as an issue. A clean run opens a rolling pull request keeping the baselines current instead, so a number moving there still has a date, an author and a review, without anybody hand-editing the JSON.