# Coverage due-diligence protocol

Created 2026-08-16. This file defines what must be done before the project says
it has made a genuinely comprehensive search for land-taking and removal events.
It does not promise that a finite historical register can prove that no event
has ever been missed.

## Scope to freeze before searching

- Temporal range: approximately 6500 BCE to the present, matching the atlas.
- Inclusion: historically identifiable direct land taking **or forced removal**,
  including settlement, conquest, demographic remaking, expulsion, blocked
  return, and land clearance for military, development, extraction, or
  conservation use.
- Distinction: record whether newcomers conquered and stayed, residents were
  removed without replacement, or land was cleared for a non-settlement use.
- Event unit: split when actors, affected population, mechanism, geography, or
  chronology changes materially.
- Evidence: retain the audit's source and confidence requirements. Record
  uncertainty rather than filling gaps by analogy.
- Source-point preservation: save every precise fact, contradiction, and
  negative finding with its source and page/section-level locator when
  available. Preserve the claim-to-source relationship in an event review,
  source brief, candidate ledger, or search log; do not leave it only in chat,
  an open browser tab, or an unindexed download.
- Editorial objective: demonstrate documentable patterns across millennia.
  Ancient and premodern searches are therefore required coverage work, not an
  optional appendix; no event is added merely to satisfy an era quota.

## Release and correction policy

The atlas is allowed to publish a strong provisional register before every
possible expert has reviewed it. Publication should expose the event definitions,
sources, confidence, known blind spots, and a practical correction channel.
Public criticism and specialist feedback are expected discovery mechanisms.

The correction channel must be available inside the atlas. A case-scoped report
should automatically identify the case and build; a general report should allow
whole-atlas criticism and missing-event proposals. Substantive reports must enter
the same candidate/correction ledger used by the audit rather than disappear into
an untracked inbox.

This iterative approach does not lower the evidence threshold or excuse an
uncrosswalked source universe. It changes outside review from a pre-publication
perfection gate into a versioned post-publication process: publish a defensible
version, log challenges, correct it, and rerun the affected coverage checks.

## Benchmark indexes to crosswalk

There is no single index with this scope. The existing register uses individual
UNHCR and IDMC sources, but as of 2026-08-16 it has not recorded a systematic
record-by-record crosswalk against the following benchmark datasets:

- UNHCR Refugee Data Finder, country/year refugee and related population data
  from 1951 onward:
  https://popstats.unhcr.org/refugee-statistics/methodology/data-content/
- IDMC Global Internal Displacement Database, event/trigger records from 2008
  onward:
  https://www.internal-displacement.org/database/displacement-data/
- Uppsala Conflict Data Program event data, organized violence from 1989 onward:
  https://ucdp.uu.se/downloads/index.html
- Government-Sponsored Mass Expulsion dataset, 139 cross-border mass-expulsion
  episodes from 1900–2020:
  https://www.belfercenter.org/publication/introducing-government-sponsored-mass-expulsion-dataset
- Ethnicity of Refugees dataset, refugee-stock composition from 1975 onward:
  https://icr.ethz.ch/data/epr/er/beta.html
- Oregon State University Dam Impacts Database, including displacement and
  resettlement fields for more than 500 selected dams:
  https://did.oregonstate.edu/
- Environmental Justice Atlas, a growing global catalog of environmental
  conflicts and community removals:
  https://ejatlas.org/
- Land Matrix, large-scale land acquisitions and their effects:
  https://landmatrix.org/
- RAMÛ Records of Ancient Migration, source-level imperial coerced migration
  from the Neo-Assyrian through Achaemenid periods:
  https://researchportal.helsinki.fi/en/projects/records-of-ancient-migration-locating-the-relocated-under-empire/

Add specialist regional or mechanism-specific datasets when discovered. These
indexes are candidate generators, not authorities that override event-level
historical research.

## Required procedure

1. **Crosswalk every benchmark record.** Give each a disposition: existing exact
   match, merged duplicate, new candidate, below threshold, out of temporal
   scope, unsupported, or unresolved. Never silently ignore a record.
2. **Build an era × region × search-mechanism review matrix.** At minimum divide the
   project into prehistoric, ancient, medieval, early modern, nineteenth
   century, 1900–1945, 1945–1989, and 1990–present cells. Record the sources and
   search terms used for each cell.
3. **Run specialist older-era sweeps.** Commission separate reviews of ancient
   imperial relocations, Greco-Roman settlement and deportation, South and East
   Asian imperial population policy, medieval religious and dynastic
   expulsions, early-modern confessional transfers, and precolonial African,
   American, and Oceanian cases. Modern humanitarian indexes cannot cover these.
4. **Run search-mechanism sweeps.** Independently review settler replacement,
   state-directed demographic remaking, expulsions and blocked return,
   conflict flight, population exchange, coerced labor movement, dams and
   infrastructure, military clearance, conservation removal, and extraction or
   plantation land seizure. These ten terms are discovery prompts only. They
   are not a public taxonomy or coded event field, and they do not revive the
   normalized mechanism vocabulary that failed its reliability retest.
5. **Keep a candidate ledger.** Every proposed event receives its sources,
   closest existing event, inclusion reasoning, rejection reasoning when
   applicable, and reviewer/date. Preserve negative findings. For older-era
   work, keep the source/search disposition log separately from the candidate
   list so a search that finds only an existing case or an access blocker is not
   silently lost. When a source supplies a precise date, count, place, route,
   artifact, legal act, or evidentiary limitation, save that point and its
   locator even if it is too detailed for the eventual atlas prose.
6. **Red-team the register independently.** Search specifically for omissions,
   bundled events, Eurocentric and modern-era bias, and events hidden by modern
   country containers. A verifier should not simply repeat the first searcher's
   source list.
7. **Publish an auditable provisional register and solicit correction.** Make
   sources, event boundaries, confidence, exclusions, and known blind spots
   visible. Invite public and specialist challenges, prioritize weak cells, and
   disposition every substantive correction in the candidate ledger.
8. **Repeat the implicated cause class after corrections.** Cause-code each
   omission and rescan the discovery method or scope that missed it, while
   retaining periodic cross-class checks. Do not restart every historical
   search indefinitely.

## Relationship to the statistical second pass

Ancient and premodern coverage work is a required corpus-preparation stage for
the queued second classification pass. The sequence is:

1. Run a declared older-era crosswalk and specialist search, beginning with
   RAMÛ and the regional sweeps above.
2. Apply the normal inclusion, event-splitting, sourcing, and confidence rules.
3. Code every accepted addition with the same eight-trait codebook and preserve
   unknowns where the evidence cannot support a value.
4. Freeze and identify that expanded corpus version.
5. Rerun the classification tests and compare them directly with the original
   407-event result, including era sensitivity and documentation-density effects.

This does not require delaying analysis until historical completeness can be
claimed. It requires a substantial, auditable older-era search before the next
corpus is frozen; later corrections can produce later versioned reruns. No era
quota and no weaker ancient evidence threshold should be used to manufacture
balance.

### Current progress, 2026-08-20

The first iteration of that sequence is complete, but the due-diligence
protocol is not:

- the repaired 407-event corpus is frozen as `analysis/second_pass/baseline_407/`;
- 84 bounded ancient/premodern cases across nine source tranches passed source
  briefs, all 60 coding judgments, an expanded 490-event rerun, and manual
  family review;
- the c. 900–612 BCE Neo-Assyrian umbrella has been source-reviewed and replaced
  without parent/child coexistence. Eleven later public-RINAP children retire
  the parent, while a completed ruler-by-ruler official-inscription review adds
  44 separately bounded c. 900–745 BCE episodes. The active arithmetic is
  `407 + 84 gross additions - 1 retired umbrella = 490`;
- the early Neo-Assyrian source ledger preserves 75 explicit source records.
  Forty-two advance dispositions become 44 non-overlapping event units after
  documented splits: 16 under Ashurnasirpal II, 24 under Shalmaneser III, and
  four under Šamši-Adad V. Holds, army-only taking, cumulative duplicates,
  ambiguous terminology, and source-poor years remain visible rather than being
  silently promoted;
- the eight-trait structure remains stable, while thirteen of the 84 additions
  surface as current nearest-prototype hybrids;
- the free automatic k=6 neighborhoods remain near-degenerate. The 446-to-490
  test moves 71 of 446 shared cases in independent fits and 20 under continuity
  initialization; carrying the prior partition costs only 0.9827% more
  within-cluster error. These clusters are diagnostic and cannot define the
  public taxonomy;
- a new source-resolution sensitivity test shows that much of that apparent
  movement is produced by archive density. Replacing the 44 early Neo-Assyrian
  records with three ruler medians, nine evidence-pattern medians, or thirteen
  ruler-by-pattern medians reduces shared-case movement to 12, six, or seven.
  Future tranches dominated by one archive must receive the same diagnostic;
- the era-balanced test shows that state orchestration, legal machinery, and
  coerced labor redistribute some loadings within a stable combined block,
  reinforcing the need for broader older-era coverage before label freeze;
- the Greco-Roman balancing ledger now preserves **407 explicit source
  dispositions**, including 97 direct Greek/Macedonian/Hellenistic prospective
  event units. Its separate 67-row Roman Republican colony catalogue has
  completed northern, central/southern, and remaining named-site priority
  reviews. Five northern event
  units advance behind broad-parent supersession, while Luna is a probable
  same-event Apuan land-use phase. Nine further catalogue rows yield ten
  central/southern event units because Fregellae's Roman foundation, intervening
  Samnite capture, and later Roman destruction require separate boundaries.
  Cales and Ariminum add two more event units; Luca is a negative control and
  Dertona remains an evidence hold. The corrected 28-row Sullan benchmark adds
  eight direct prospective units, including one medium-confidence Alban Hills
  regional unit resolved through a source-critical land-surveyor review,
  while preserving colony-only sites as holds and the national program as a
  parent control. The separate 23-row Sullan punishment sibling review adds
  seven prospective units, preserves five holds and four support or overlap
  records, and rejects six booty or spoil leads without a resident-land
  outcome. The first eighteen-row standard-tier foundation batch adds eleven
  prospective 329-241 BCE units, four same-event or origin-land supports, and
  three evidence holds. It groups linked Auruncian and Aequian sequences and
  prevents Firmum from duplicating the existing Picentine transfer. The second
  eighteen-row checkpoint covers all seventeen remaining standard rows plus the
  Castrum control: thirteen prospective 194-122 BCE units, three support nodes,
  one standard evidence hold, and one related textual-identity hold. It groups
  the four Campanian coastal nodes and keeps different affected communities and
  source lands separate. A subsequent twelve-row pre-Gracchan viritane-and-
  controls benchmark advances eight prospective units, preserves one location
  and overlap hold, rejects two false or non-qualifying event leads, and kept
  Labici as a qualifying scope control that opened a separate pre-338 BCE
  foundation census. It also identifies Capua's 211-210 BCE removal, confiscation, and
  civic dismantling without inventing an incoming settler population, and
  treats the 232 BCE *lex Flaminia* as a later settlement phase rather than a
  duplicate Senone expulsion. The subsequent nine-row Gracchan review advances
  five bounded regional land-division implementations, preserves Polla and
  Calabria/Bruttium as explicit source-resolution holds, and keeps both the
  national program and the 111 BCE regularization out of the event count. The
  subsequent twelve-row late-Republic residual review advances four prospective
  units, preserves four source or implementation holds, and records four
  rejected, annulled, or unlocated-program controls. It separates the 146 BCE
  destruction of Carthage from Junonia's later allotments, records Narbo as
  land taking with Indigenous continuity, and groups the disputed Marian
  African evidence into one medium-confidence program rather than one event per
  late title. The following two-row Salassian review advances one 25–24 BCE
  mass-enslavement, best-land-transfer, and Augusta Praetoria event and
  preserves the early Salassian *incolae* dedication as same-event survivor-
  continuity evidence. The completed thirty-three-row pre-338 BCE foundation
  census then advances eighteen bounded event units and preserves fifteen
  uncertain foundations, same-site supports, failed or unidentifiable
  proposals, and extension controls. Labici is updated in place, and the
  review freezes proposal/implementation, parent/child, sequence, and paired-
  versus-local sensitivities. All 22 priority catalogue rows remain
  dispositioned. A final row-level residual audit gives all 67 canonical parent
  components and all 96 Roman coding units explicit dispositions. It freezes
  complete retirement of both broad Roman parents when children activate,
  without using eleven evidence holds as a fictional residual event. The queue
  is 95 prospective events plus one count-neutral Apuan/Luna recode; **all
  ninety-six Roman coding units** now have a validated combined source-brief
  row and a frozen codebook-1.1 Roman coding protocol, but remain inactive or
  unimplemented behind the sixty-dimension coding and sensitivity gates. The
  546 parent-only and 583 full-candidate
  counts are projections only and do not alter the active 490-event matrix;
- a separate 35-row Hellenistic royal-transfer inventory now fully dispositions
  the first western/Macedonian synoikism and relocation benchmark. All records
  are source-frozen: twenty event-unit advances, eleven controls, and four
  evidence holds closed under the declared search. No Cohen catalogue lead or
  named first-pass gate remains open. The review requires implemented
  removal, clearance, land transfer, or bounded successor settlement rather
  than treating royal foundation or political union as proof of displacement.
- a separate 19-row eastern Hellenistic priority inventory now preserves nine
  prospective event units, five negative controls, and five evidence holds.
  It tests destruction and removal outside foundation catalogues, counts
  Ptolemy's 312 BCE four-city razing as one linked operation, separates the
  reversed Jerusalem Akra removals, and rejects generic Ptolemaic cleruchies
  and a mass-Babylonian-transfer claim as event units;
- a follow-on 25-row Cohen-2013 eastern catalogue-trigger inventory preserves
  five prospective event units, twelve evidence or scope holds, and eight
  negative or overlap controls. It completes candidate disposition across the
  catalogue's geographic divisions, corrects Apameia on the Silhu to December
  140 BCE, and keeps Ai Khanoum inside one broad Bactrian sequence rather than
  double-counting it. This closes the declared catalogue trigger pass but does
  not close the non-catalogue search;
- the first full Alexander Indian city-treatment sibling pass adds a separate
  27-row source-frozen inventory: eleven prospective event units, one
  Peucelaotis identity/resident-outcome hold, and fifteen negative, nested, or
  same-event controls. Connected Assacenian and Mallian narratives are grouped
  rather than weighted one case per city; peaceful submissions remain visible;
  and Rhambakia advances only inside the wider Oreitai conquest-and-colony
  sequence. These units are not coded or active;
- the first bounded named Ptolemaic Egypt/Cyrenaica pass adds a separate 25-row
  source-frozen inventory: four prospective event units, four evidence or
  overlap holds, and seventeen negative, overlap, nested, or parent controls.
  Alexandria's 145–124 BCE evidence is one repeated sequence; the
  Barke/Ptolemais transfer is resolved as a non-event control; and ordinary
  foundations, reclaimed-land allotments, and the Great Revolt wrapper do not
  become cases without bounded resident–land outcomes. These new units are not
  coded or active;
- the bounded Hasmonean regional sibling pass adds a separate 32-row
  source-frozen inventory: thirteen prospective event units—twelve newly
  identified plus the already counted Akra reversal—five evidence holds, and
  fourteen negative, nested, military-only, or parent controls. Three
  identifiers already existed, so 29 rows enter the shared ledger. Joppa's
  164/3 BCE resident killing, 147 BCE negotiated capture, and 143 BCE expulsion
  remain distinct; Simon's Joppa, Gezer, and Akra operations remain separate;
  and source-dense Gilead and Jannaeus sequences are grouped to avoid archive-
  density inflation. Nothing is coded or active, and the 490-event matrix is
  unchanged;
- the completed cross-inventory Hellenistic overlap audit gives all 97 direct
  Greek/Macedonian/Hellenistic advances a row-level shared-place, chronology,
  active-parent, and event-unit disposition. It finds no duplicate unit, no
  active-event supersession, and no parent/child coexistence. Seven rows can
  coexist with four active archaic Greek colonization parents only after
  explicit scope narrowing. Forty-two rows fall into twelve compact ruler,
  campaign, or shared-source groups that must receive both one-group-median and
  group-by-outcome-pattern sensitivity tests after coding. Four
  `advance_pending_*` Greek rows remain outside the freeze.

These are nine coded and modeled source tranches plus western and eastern
pre-coding Hellenistic source checkpoints, not saturation. RAMÛ remains blocked
at record level, the other declared benchmark indexes have not all been
crosswalked, and later Sasanian and other regional source series still require
event splitting. The Neo-Assyrian official-inscription series is now bounded,
coded, modeled, and manually reviewed for both the later replacement children
and the admitted early tranche, but it is not a benchmark-complete crosswalk:
specialist and non-royal evidence can still refine the record, and neither the
public children nor the official corpus substitutes for RAMÛ's 4,327-record
disposition. The Mongol parent is partially dispositioned after nine
western-campaign additions; European captives and later workshop transfers
remain explicit community-unit and overlap work. The New Kingdom source series is only partially
dispositioned: two Levantine cases are coded, Konosso remains held against the
broad Nubia wrapper, and the reported 74-case record table is unavailable.
Three non-admitted
Achaemenid children and two early Sasanian reports now have documented
evidence-limited holds rather than open reviews; they reopen only if RAMÛ or a
new source supplies the missing event or coercion evidence. The first manual
scope tranche moved all 26 formerly open removal-family rows to presumptive
inclusion under the forced-removal rule. The second manual tranche advanced 16
more and recorded nine explicit scope/evidence/event-unit holds. No omission
remains in an undefined open bucket, and no pair of successive independent
zero-new-event scans has occurred. The project therefore may describe a
validated first expansion, but not completed comprehensive due diligence.

## Declared-scope coverage-report evidence

The project may publish a finite coverage report for a declared scope when all
of the following are true:

- 100% of records from the declared benchmark datasets have a documented
  disposition.
- Every applicable era × controlled-region × search-mechanism cell has at least
  two methodologically distinct search channels, or an explicit non-applicable
  or evidence/access-limited disposition. Any unsearched cell is listed.
- Every included event has stable provenance, source minimums, and an event-unit
  review; every exclusion has a reason.
- Cross-dataset duplicates and umbrella events are resolved explicitly.
- Every located candidate has a current disposition, including explicit holds;
  search leads without real source locators remain separately identified.
- Channel yields and overlaps, valid benchmark-disposition fractions, and known
  unsearched cells are published. No global coverage percentage is reported
  without a specified and validated estimator.
- The remaining uncertainties and known blind spots are published alongside the
  coverage claim.

The defensible wording is then: **this coverage report states the declared
scope, where and how the atlas searched, what it found, every current candidate
disposition, and the remaining limitations.** It closes that declared search
program, not history. Later evidence may reopen a cell. Do not claim saturation,
historical completeness, or a mathematically complete census.
