# Prospectively frozen exploratory audit: kinship and chronology anomalies

**Status:** Frozen before inspection or extraction of any additional candidate outcomes
**Freeze date:** 2026-08-10 (Australia/Adelaide)
**Characterisation:** Prospectively frozen exploratory audit, not an independent preregistered confirmation

## 1. Question and permitted inference

The principal question is whether public ancient-DNA data contain genetically related individuals whose independently dated chronologies remain unusually difficult to reconcile under extreme but ordinary-human generation and lifespan limits after source-level validation.

The conditional hypothesis is that a small family arriving near Cudi Dağı around 2900 BCE may initially have had approximately 79–135-year fathering intervals, delayed maturation and long lifespans; lineage growth and outbreeding may have made descendants more visible only after dispersal from Shinar during approximately 2370–2033 BCE. The analysis tests for a predicted *class of discrepancy*. It cannot identify a person or lineage as Noahic, prove Babel, distinguish institutional diffusion from biological migration, or estimate Noah's total genetic contribution.

A null result is acceptable. Tatika-era remains are local baselines, not presumptive descendants.

## 2. Prior-information boundary

The materials and observations in `prior-observations.md` were read before this plan. No additional AADR family-relation rows, ancIBD pair/segment outcomes, candidate source supplements, or anomaly-search outcomes were opened or extracted before the freeze. AADR's header and the metadata-audit README/file inventory were inspected to establish available fields and infrastructure only.

## 3. Frozen chronology and geography

### Chronology

- Proposed arrival/local-baseline anchor: 2900 BCE near Cudi Dağı.
- Gathering at Shinar: approximately 2800 BCE.
- Dispersal onset: 2370 BCE.
- Proposed route arrivals: 2250–2033 BCE.
- Peleg/Babel completion anchor: approximately 2033 BCE.
- **Primary analysis window:** 2200–1000 BCE.
- **Sensitivity A:** 2370–1000 BCE (dispersal-onset inclusive).
- **Sensitivity B:** 2033–1000 BCE (post-completion).
- **Local temporal control:** 3500–2371 BCE in the same target regions.

Selection into a window is based on at least 0.50 probability mass of an individual's calibrated distribution within the window. The pair enters when both individuals satisfy the rule. Sensitivities use any-overlap and posterior-mean inclusion, reported separately. BCE years are stored as signed astronomical/calendar values consistently; midpoint years are never treated as observed death dates.

### Target regions

The frozen post-Babel target comprises these website-derived route regions: Armenia/southern Caucasus (Togarmah), Anatolia and the Aegean (Lud/Javan), northern Mesopotamia and Syria (Asshur/Aram), Iran including Media and Elam, the Levant (Canaan), Egypt (Mizraim), and the Upper Nile/Nubia (Cush). Operationally, records in Armenia, Azerbaijan, Georgia, Turkey, Greece, Cyprus, Iraq, Syria, Iran, Lebanon, Israel, Palestine, Jordan, Egypt, Sudan, Ethiopia and Libya are discovered, then assigned by coordinates and site context to one of the listed regions. Border cases remain in the table with an explicit classification note.

These routes imply possible founder migration or contact, not population replacement. They are not assumed to require a detectable ancestry pulse. Canaan's conflicting Sumer/Egypt website transmission labels are pooled as Levant target status rather than resolved after results. The Cain/Zagros 5000/4321 BCE material is outside this audit's chronology.

### Geographic controls

1. **Off-route West Eurasian controls:** directly dated ancient relatives from Europe outside Greece/Cyprus, the steppe north of the Caucasus, Central Asia east of Iran, and South Asia, within the same primary or sensitivity window.
2. **Earlier local controls:** target-region relative pairs in 3500–2371 BCE.

Pairs from the Americas, Oceania, eastern Asia, and modern populations are excluded from primary controls. They may appear only in methodological validation.

## 4. Sources and acquisition

Discovery sources are AADR v66.p1 family-relation annotations and the public ancIBD 2024 dataset (Ringbauer et al., DOI 10.1038/s41588-023-01582-w; Zenodo 10.5281/zenodo.8417049). AADR and ancIBD are discovery aids. Every screened candidate must be traced to the earliest primary publication, supplementary table, direct-date record, and relationship method.

The audit will retain all discoverable first- and second-degree pairs and all published/reliable third- through sixth-degree pairs in scope, without requiring date overlap. It will also retain lower-quality and contextual-date cases in separate descriptive tiers. A source manifest will record URL, access date, byte size and SHA-256 where possible.

## 5. Unit of analysis and deduplication

The basic row is an unordered pair of archaeological individuals. Aliases and multiple AADR representations are collapsed to archaeological person IDs. Multiple relationship calls for one pair occupy one row with method-specific fields. Multiple pairs within the same connected pedigree/network are not independent positive events; the independent unit for the replicated-pattern rule is a connected network/site-publication cluster.

## 6. Relationship and data-quality tiers

### Relationship classes

- Resolved parent–offspring.
- Resolved full siblings.
- Unresolved first degree.
- Resolved grandparent–grandchild.
- Resolved aunt/uncle–niece/nephew.
- Unresolved second degree.
- Resolved first cousins.
- Unresolved third degree.
- Published/reliable fourth–sixth degree, retained for IBD/network analysis only.

### Genetic quality

- **G1 (high):** both individuals have at least 400,000 covered 1240K SNPs (or published equivalent high-confidence eligibility), no AADR CRITICAL/FAIL warning, and a relationship supported in a primary paper by a quantitative method or resolved pedigree.
- **G2 (moderate):** both have at least 100,000 covered 1240K SNPs, no CRITICAL/FAIL warning, and a published or AADR-annotated relationship that survives source checking.
- **G3 (exploratory):** below G2, quality warning, unresolved aliasing, or discovery annotation not independently confirmed.

For ancIBD, the published high-confidence gate is used: at least 70% of chromosome-3 1240K SNPs with maximum imputed genotype probability above 0.99; segments longer than 8 cM are retained, while close-degree classification uses segments longer than 12 cM. Exact topology beyond third degree is not inferred from a degree label alone.

### Dating quality

- **D1 (high):** both individuals directly dated, raw conventional radiocarbon ages/errors and laboratory codes available, dated elements identified, and calibration/reservoir treatment auditable.
- **D2 (moderate):** both directly dated but one raw input, element, or reservoir detail is incomplete; a full published calibrated distribution or defensible reconstruction is available.
- **D3 (exploratory):** contextual/associated/modelled date, only a broad range, mixed-individual context, or unresolved element identity.

**Primary set:** G1 + D1, relationship classes with a defensible topology-specific ceiling (resolved parent–offspring, full siblings, unresolved first degree, resolved grandparent–grandchild, resolved aunt/uncle–niece/nephew, unresolved second degree, resolved first cousins).
**Expanded sensitivity:** G1/G2 + D1/D2.
All other pairs remain in the complete table but cannot drive the conclusion.

## 7. Ordinary-human comparison

The hard limits come from Sedig et al., *Journal of Archaeological Science* 133 (2021) 105452, DOI 10.1016/j.jas.2021.105452:

| Topology | Frozen maximum absolute death-date separation |
|---|---:|
| Parent–offspring | 100 years |
| Full siblings | 135 years |
| Unresolved first degree | 135 years |
| Grandparent–grandchild | 180 years |
| First cousins | 195 years |
| Aunt/uncle–niece/nephew | 210 years |
| Unresolved second degree | 210 years |

The paper's sibling worked example totals 130 years while its formal table says 135; 135 is frozen as the more conservative value. It supplies no defensible generic hard maximum for unresolved third- through sixth-degree pairs. Such pairs are analysed descriptively through topology mixtures, calendar time per inferred meiosis and multi-person networks, and cannot alone be stringent anomalies.

An empirical sensitivity uses Sedig et al.'s historical/genealogical absolute separation summaries: parent–offspring 28.84±18.94; siblings 26.33±22.41; grandparent–grandchild 35.00±31.93; pooled other second/third degree 34.94±25.60 years. These are comparison distributions, not ancient-population hard limits.

## 8. Calendar distributions and residuals

Raw radiocarbon measurements will be recalibrated on a one-year grid with IntCal20 for terrestrial samples, Marine20 or a published mixed curve only when source evidence requires it. Published OxCal posterior distributions are preferred when downloadable. Calibration uncertainty is retained; broad and multimodal distributions are not collapsed.

For each pair with distributions `p1(t)` and `p2(t)`, calculate the full convolution of absolute separation and `P(|T1-T2| > L)`, where `L` is the frozen topology-specific limit. Also report median separation, 95.4% interval, and minimum distribution-supported separation. Calendar distributions are independent for screening unless a published laboratory/context model supplies covariance; shared lab effects are assessed in sensitivity.

### Tissues, age at death and reservoir effects

- Direct dates are measurements on tissues, not automatically exact death clocks.
- When a tooth/enamel/dentine formation offset and morphological age-at-death range are available, transform to a death-date distribution using the published offset/range.
- When a potentially early-forming tissue lacks an auditable offset, the pair cannot be a stringent survivor; sensitivity broadens the date by the widest source-supported formation-to-death interval.
- Published reservoir corrections are used. Uncorrected marine/freshwater dietary risk, calibration-curve ambiguity, or large isotope uncertainty prevents stringent survival unless every plausible correction leaves the flag intact.
- Contextual and direct dates are never pooled as equivalent.

## 9. Frozen flagging and multiple-comparison rules

- **Source-investigation screen:** `P(exceeds hard limit) >= 0.20`, or a published author flags incompatibility/mismatch even when a complete distribution cannot be reconstructed.
- **Unresolved stringent discrepancy:** G1+D1, source-validated topology, `P(exceeds) >= 0.95`, robust under all prespecified calibration/reservoir/tissue sensitivities, with no plausible identity, commingling, intrusion, reuse, catalogue or lab-error resolution.
- **Multiplicity-surviving discrepancy:** additionally `m * (1 - P(exceeds)) <= 0.05`, where `m` is the number of topology-eligible primary pairs. This Bonferroni-like posterior-error screen is descriptive and is not labelled a frequentist p-value.

Every screened case enters `candidate-resolution.csv`, including disproved candidates and negative controls. No candidate is silently removed.

## 10. Matched-control construction

Controls are selected before anomaly scores are viewed. For each target pair, select up to three off-route pairs and up to three earlier-local pairs, exact-matching relationship topology/class and dating tier where possible. Nearest-neighbour distance uses prespecified covariates only: pair mean calendar period, log minimum SNP coverage, burial/context class, direct-date completeness, and publication/laboratory identity indicators. Outcome residuals, calibrated pair separation and anomaly flags are forbidden matching variables. Matching is without replacement where feasible; shortages and fallback-with-replacement are logged.

Target enrichment is tested at the independent network/site-publication level by a one-sided permutation/Fisher analysis and reported with an odds ratio and interval. A positive pattern requires both `p <= 0.05` and odds ratio at least 2.0; otherwise enrichment is not established. Raw rates and uncertainty are reported even when testing is underpowered.

## 11. IBD and conditional long-generation analyses

Published ancIBD segments and pair summaries will be used to reconstruct connected networks. Total IBD, count of >12 cM segments, maximum segment, inferred degree, calendar separation and source geography are reported. An isolated long segment is not an independent positive. Endogamy, ROH, imputation quality and relationship-class overlap are explicit confounders.

Where topology is resolved and data support it, compare:

1. the ordinary empirical separation distributions above; and
2. a conditional model using the pre-disclosed Genesis fathering intervals (79–135 years) and declining lifespans from Shem through Terah.

The conditional model is a likelihood/simulation comparison, not a lineage identifier. It will not be fitted if fewer than five primary pairs or one informative network are available; that limitation will be stated rather than filled with speculative estimates.

## 12. Conventional anomaly vocabulary and secondary evidence

For every screened pair, site and primary paper, search original publications and osteological reports for the vocabulary listed in the user brief: radiocarbon outlier/reservoir/calibration problems; intrusion, reopening, reuse, commingling or sample mismatch; discordant dental/skeletal age; delayed or unfused epiphyses; endocrine/developmental delay; cementum anomalies; and unexpected adult wear/degeneration in subadult classifications. Record raw observation and authors' explanation separately.

Multi-tissue cases can support but not create a positive pattern unless identity is established (ideally DNA), at least two direct tissue dates exist, preservation and isotopes are reported, and mixing/reservoir/tissue-turnover explanations are tested. Cementum or morphology alone is insufficient.

## 13. Positive-result rule and outcome classes

A **conditionally positive** result requires all of:

1. at least three multiplicity-surviving high-quality discrepancies;
2. at least two independent connected networks/sites and two publications;
3. survival of original-source, identity, calibration, tissue and reservoir audits;
4. target enrichment versus frozen matched controls (one-sided `p <= 0.05`, OR >= 2.0);
5. no single pedigree/network counted more than once toward the three-case minimum; and
6. at least one case supported by an independent chronological or developmental measurement.

The rule will not be weakened. Final classes:

- **Null:** no high-quality discrepancy survives.
- **Candidate-generating:** unresolved cases merit redating/reassessment but no replicated enriched pattern survives.
- **Conditionally positive:** the complete rule is met, without identifying a biblical lineage.
- **Artefact-positive:** automated flags are explained by metadata, identity, dating or method artefacts.

## 14. Sensitivities

Prespecified sensitivities are: chronology windows A/B; probability-mass versus any-overlap window selection; G1-only versus G1+G2; D1-only versus D1+D2; formal 135-year sibling limit versus the prose 130-year value; relationship-resolved only; IntCal20 versus published posterior; laboratory systematic shifts of ±20 and ±40 years; published reservoir correction extremes; tissue-formation offsets; shared-context covariance; target subregions separately; controls with/without replacement; and network-level rather than pair-level counting.

## 15. Null upper bound

If no qualifying pattern survives, estimate a 95% upper bound on the fraction `f` of sampled, high-quality topology-eligible relationships that carry a *detectable chronology discrepancy*. Pair-specific detection probabilities `d_i` will be estimated by simulation using each pair's calibrated uncertainty and two frozen injected effects: true separation 30 years beyond its hard limit and 100 years beyond it. Solve `product(1 - f*d_i) = 0.05`; report both scenarios and effective sample size `sum(d_i)`. If the observed surviving count is nonzero, use the corresponding Poisson-binomial likelihood.

This upper bound constrains only detectable discrepancies among sampled high-quality public relationships under the stated effect and measurement model. It does not bound Noahic ancestry, lineage frequency in unsampled populations, longevity without a chronology discrepancy, or the hypothesis in regions with no usable pairs.

## 16. Reproducibility and deviations

Scripts, raw-source hashes, intermediate tables, random seed (`20260810`), software versions and executed commands will be retained. Initial plan files are SHA-256 hashed and committed before candidate extraction. Any unavoidable change is appended prospectively to `deviations-log.md` with date, reason, effect and whether outcomes had been inspected; the frozen rule is never edited in place.
