# 8. Macro-Culture Historical Research SOP

**Status:** Active working procedure  
**Version:** 1.0  
**Date:** 2026-08-06  
**Applies to:** The 24 × 24 Macro-Culture sector model and its Great Lakes pilot  
**Primary objective:** Build historically grounded sector records by working from modern structured data and historical archival evidence inward toward a documented overlap zone.

## 1. Purpose

This SOP defines how Macro-Culture research will be conducted for all 576 sectors and how the first nine-sector Great Lakes pilot will be assembled.

The procedure is designed to preserve four distinctions:

1. **Observation** — what a source directly reports.
2. **Interpretation** — what the observation may mean in the sector context.
3. **Derivation** — what a transparent formula calculates from observations.
4. **Classification** — the provisional quadrant or model judgment assigned afterward.

No historical description should be silently converted into a modern-looking number. No model output should be presented as an archival fact.

## 2. Fixed project decisions

### 2.1 Spatial frame

The canonical grid is fixed at **24 × 24**:

- 24 latitude bands, each 7.5° high;
- 24 longitude bands, each 15° wide;
- 576 total sectors;
- stable sector identifiers and boundaries across all periods.

Land, ocean, and mixed cells remain in the grid. Each record must state its land/water composition for the relevant period and source resolution.

The previously encountered 24 × 48 visualization is not the canonical research grid. Any data or visualization using another geometry must be explicitly marked as a derivative, crosswalk, or unresolved legacy artifact.

### 2.2 Evidence direction

Research proceeds from both ends inward:

- **Modern end:** standardized, machine-readable, spatially comparable datasets.
- **Historical end:** dated archival sources, maps, reports, censuses, newspapers, directories, and institutional records.
- **Center:** periods and variables where the two evidence systems can be compared without pretending they are identical.

### 2.3 First study area

The first complete pilot is the existing nine-sector Great Lakes octagon centered on `N05_E18`.

Global expansion begins only after the pilot demonstrates that the template, provenance rules, period crosswalks, and formulas can be applied consistently.

## 3. Research unit: the sector-period record

The basic unit is not a timeless sector. It is a **sector during a defined time band**.

```text
sector_period_record
├── sector_id
├── period_id
├── geometry
├── land_water_composition
├── climate_profile
├── water_profile
├── energy_profile
├── population_settlement_profile
├── infrastructure_connectivity
├── historical_institutional_profile
├── neighboring_sector_relationships
├── raw_observations
├── derived_constraints
├── adaptive_capacity_indicators
├── provisional_quadrant
├── confidence
├── missingness
└── source_provenance
```

A sector-level summary may be generated later, but it must retain links to the underlying period records.

## 4. Standard time bands

Use broad bands first. Do not imply annual precision when the sources do not support it.

| Period | Purpose | Typical evidence |
|---|---|---|
| 2000–present | Modern baseline | Reanalysis, satellite, gridded population, current infrastructure and energy datasets |
| 1950–2000 | Recent historical bridge | Statistical yearbooks, census series, industrial and energy statistics, aerial/satellite records, digitized technical reports |
| 1800–1950 | Industrial transition | Censuses, agricultural and geological surveys, railroad and port records, directories, maps, municipal reports |
| 1500–1800 | Early modern reconstruction | Gazetteers, land and tax records, travel accounts, historical maps, trade and settlement records |
| Before 1500 | Selective deep history | Archaeology, paleoenvironmental reconstructions, historical geography, surviving primary sources |

These bands are a starting crosswalk, not a claim that every source fits neatly inside them. A source may receive a narrower `date_start` and `date_end` within a band.

## 5. Source hierarchy

### Tier A — structured primary or authoritative data

Use first when available:

- official statistical agencies and censuses;
- national geological, hydrological, meteorological, and mapping agencies;
- intergovernmental datasets with documented methods;
- peer-reviewed datasets with clear provenance;
- original historical surveys and administrative records.

### Tier B — institutional and technical records

Use for sector history and calibration:

- engineering and geological reports;
- railroad, canal, port, utility, and industrial records;
- agricultural experiment-station reports;
- municipal and regional planning documents;
- historical atlases and gazetteers;
- university digital collections.

### Tier C — Internet Archive and other digitized repositories

Use Archive.org for discovery, OCR, scans, maps, newspapers, directories, government documents, and historical books. Record the item identifier, file name, page or image number, date, creator, collection, and access URL.

### Tier D — secondary synthesis

Use scholarly books, review articles, local histories, and reputable reference works to locate evidence and explain context. Do not let a secondary synthesis replace a primary source when the claim is central to a score or causal interpretation.

### Tier E — exploratory or uncited material

Use only as a lead. It cannot support a final observation, score, or conclusion until independently verified.

## 6. Modern-end procedure

For each pilot sector and period:

1. Load the fixed 24 × 24 geometry.
2. Extract climate and water variables from the selected gridded datasets.
3. Extract population and settlement variables from the selected population and built-environment datasets.
4. Extract energy and electricity variables from authoritative energy and access datasets.
5. Extract roads, rail, ports, waterways, grid lines, and other connectivity layers.
6. Preserve the original source units and spatial resolution.
7. Aggregate to the sector using a documented rule: area-weighted mean, sum, median, share, count, nearest feature, or categorical majority.
8. Store the dataset version, retrieval date, variable name, unit, aggregation rule, and uncertainty.
9. Flag modeled, interpolated, reanalysis, and observed values separately.
10. Do not assign a quadrant during ingestion.

The modern baseline should be reproducible from the source files and extraction script or notebook.

## 7. Historical-end procedure

### 7.1 Define the search envelope

Before searching Archive.org, create a source brief for the sector containing:

- sector ID and coordinates;
- modern place names;
- historical place names and spelling variants;
- rivers, lakes, ports, cities, indigenous territories, and administrative units;
- relevant industries and commodities;
- target period;
- target variables;
- neighboring sectors and likely exchange routes.

### 7.2 Search in source families

Search each sector through multiple families rather than one broad query:

```text
place + census / population / directory
place + climate / weather / drought / flood
place + river / watershed / lake / irrigation
place + coal / timber / oil / electricity / utility
place + railroad / canal / port / shipping / highway
place + industry / manufacturing / agriculture / mining
place + law / territory / land / government / institution
place + neighboring region / trade / migration / commodity
```

Repeat searches with historical names and nearby administrative units.

### 7.3 Inspect and extract

For each candidate item:

1. Read the catalog metadata.
2. Check date, creator, geographic scope, and collection.
3. Inspect available OCR and file derivatives.
4. Search OCR for target terms.
5. Read the surrounding passage, not just the matching line.
6. Inspect the original page or map when OCR is ambiguous.
7. Extract a bounded observation with page or image reference.
8. Record whether the observation is measured, reported, estimated, or interpretive.
9. Preserve the original wording when it carries uncertainty or bias.
10. Download only the files needed for verification and reproducibility.

OCR is a locator and transcription aid, not automatically authoritative evidence. Tables, maps, numbers, names, and negations require visual verification when possible.

## 8. The bridge toward the center

The bridge aligns modern and historical evidence without erasing their differences.

### 8.1 Geography crosswalk

For every historical source, record:

- historical place name;
- source boundary or geographic extent;
- modern equivalent, if known;
- relationship to the fixed 24 × 24 cell;
- whether the source covers one cell, several cells, or only a point;
- boundary uncertainty.

If a source crosses sector boundaries, do not assign it to one cell without noting the cross-boundary issue.

### 8.2 Variable crosswalk

For every variable, record whether the historical and modern measures are:

- **directly comparable**;
- **related but not equivalent**;
- **qualitative only**;
- **not comparable**.

Examples:

```text
Modern: built-up area percentage
Historical: description of dense settlement
Status: related but not equivalent

Modern: electricity-access rate
Historical: utility service area or household lighting source
Status: related but not equivalent

Modern: annual precipitation
Historical: weather-station series
Status: potentially directly comparable after unit and period checks
```

### 8.3 Period anchors

Anchor the historical timeline to events that can be independently dated:

- census years;
- railway or canal openings;
- port construction;
- major dams or power stations;
- mine or factory openings and closures;
- boundary changes;
- severe floods, droughts, fires, epidemics, or wars;
- major migration waves;
- energy-regime transitions.

Anchors organize the timeline; they do not prove causation.

## 9. Observation and provenance standard

Every extracted observation must use a record like this:

```json
{
  "sector_id": "N05_E18",
  "period_id": "1850-1900",
  "variable": "infrastructure_connectivity",
  "subvariable": "rail_connection",
  "value": "Rail connection to Chicago and eastern markets reported",
  "value_type": "qualitative_observation",
  "date_start": "1850",
  "date_end": "1900",
  "geographic_scope": "historical place or source boundary",
  "comparability": "related_not_equivalent",
  "source_tier": "B",
  "source": {
    "repository": "Internet Archive",
    "identifier": "archive_item_identifier",
    "file": "file_name",
    "page_or_image": "page_or_image_reference",
    "title": "source_title",
    "creator": "source_creator",
    "publication_date": "source_date",
    "url": "https://archive.org/details/..."
  },
  "extraction_method": "OCR_checked_against_scan",
  "source_language_or_bias": "description",
  "confidence": "medium",
  "notes": "Boundary and terminology require crosswalk review"
}
```

The exact schema may evolve, but these provenance fields are mandatory for research-grade observations.

## 10. Confidence and missingness

Confidence is assigned to the observation, not to the whole theory.

### 10.1 Confidence dimensions

Score or describe separately:

- source authority;
- geographic fit;
- temporal fit;
- measurement quality;
- extraction quality;
- independent corroboration;
- comparability to other periods;
- known source bias.

A single summary label may be used only alongside these dimensions:

- **high:** authoritative, well-matched, clearly measured, and corroborated;
- **medium:** useful evidence with a limitation, indirect measure, or partial corroboration;
- **low:** plausible lead, weak geographic or temporal fit, substantial uncertainty;
- **unresolved:** insufficient evidence to support an observation.

### 10.2 Missingness codes

Use explicit codes rather than blank cells:

- `not_searched`;
- `searched_no_result`;
- `source_exists_not_accessible`;
- `source_exists_not_digitized`;
- `historical_measure_not_available`;
- `geographic_match_uncertain`;
- `period_match_uncertain`;
- `ocr_or_scan_unreadable`;
- `conflicting_sources`;
- `not_applicable`;
- `ocean_or_nonsettled_cell`.

Archival silence is not evidence that an event or condition did not exist.

## 11. Derived constraints and adaptive capacity

Do not calculate the model layer until the raw observation table is complete enough to inspect.

### 11.1 Derived constraints

Possible components include:

- climate stress;
- water stress;
- energy-access or energy-regime constraint;
- transport and connectivity constraint;
- hazard exposure;
- resource concentration or dependency.

Each component must state its formula, normalization range, weights, and missing-data behavior.

### 11.2 Adaptive capacity

Possible indicators include:

- infrastructure redundancy;
- energy diversity;
- connectivity to alternatives;
- institutional capacity;
- documented repair or adaptation;
- population mobility options;
- technological substitution;
- evidence of institutional learning.

Adaptive capacity must not be inferred merely from wealth, survival, expansion, or present-day success. Those outcomes require an explicit mechanism and evidence.

### 11.3 Formula rule

The formula should remain fixed across sectors during a validation round. If the formula changes, rerun prior sectors and record the version change.

## 12. Provisional quadrant assignment

Quadrants are model outputs, not cultural essences:

- high constraint / high comprehension;
- high constraint / low comprehension;
- low constraint / low comprehension;
- low constraint / high comprehension.

A provisional quadrant may be assigned only when:

1. the relevant raw variables are present or their missingness is explicit;
2. the formula version is recorded;
3. the evidence quality is sufficient for a provisional judgment;
4. alternative explanations are listed;
5. the confidence level is visible.

If these conditions are not met, use `unresolved` rather than forcing a quadrant.

## 13. Neighbor and matched-sector analysis

After individual sector records are assembled, compare:

- adjacent sectors;
- latitudinally similar sectors;
- longitudinally connected sectors;
- resource-complementary sectors;
- sectors with similar physical conditions but different institutional histories;
- sectors with different physical conditions but similar institutions or infrastructure.

Neighbor relationships must be directional where possible. Record whether a relationship is based on observed trade, migration, transport, political control, ecological dependence, or only geographic proximity.

## 14. Validation sequence

### Phase 1 — Great Lakes pilot

Complete the nine sectors with the same schema and source rules.

### Phase 2 — Modern/historical overlap

Test the 1950–2000 bridge first, where structured statistics and archival material are both relatively available.

### Phase 3 — Industrial transition

Extend selected variables into 1800–1950 and identify which measures can be reconstructed reliably.

### Phase 4 — Matched comparison

Select comparison sectors before looking at outcomes. Do not choose only cases that appear to confirm the thesis.

### Phase 5 — Formula lock

Freeze the first validated formula version and apply it to a larger sector sample.

### Phase 6 — Global expansion

Apply the template to all 576 sectors only after the pilot passes the following checks:

- geometry is stable;
- source provenance is complete enough to audit;
- missingness is visible;
- variable crosswalks are explicit;
- independent comparisons are possible;
- counterexamples have been documented;
- the formula does not require hidden judgments.

## 15. Falsification and stop conditions

Research must actively search for cases that weaken the model.

Pause expansion and review the method if:

- changing the geographic boundary changes the result without a principled reason;
- the model only works for cases used to create it;
- the same observed outcome can be explained equally well by institutions, technology, colonial power, class, or political history alone;
- historical sources contradict the assigned quadrant;
- missingness is concentrated in particular regions or populations;
- the model treats a population as possessing a fixed cultural essence;
- qualitative archival evidence is being used to manufacture precise scores;
- a visually complete map conceals incomplete data.

The desired outcome is not confirmation of every prediction. It is a more useful comparative explanation with visible limits.

## 16. Research deliverables

Each completed pilot sector should produce:

1. `raw-observations.json` — extracted facts and measurements;
2. `historical-timeline.md` — dated events and transitions with citations;
3. `variable-crosswalk.md` — modern/historical comparability notes;
4. `derived-profile.json` — calculated indicators and formula version;
5. `provenance.md` — complete source register;
6. `missingness.md` — unresolved gaps and access limitations;
7. `review-notes.md` — disagreements, alternative explanations, and falsification checks.

The project-level deliverable should include:

- the 24 × 24 grid definition;
- the source registry;
- the formula specification;
- the confidence rubric;
- the missingness report;
- pilot comparison results;
- a list of revisions made to the SOP.

## 17. Working rule

> **Modern datasets establish comparable measurements. Historical sources establish temporal depth and institutional context. The bridge records where they converge, where they disagree, and where comparison is not justified.**

Macro-Culture remains a comparative systems lens. It does not claim that geography mechanically determines culture, and it must not be used to assign fixed cultural traits to populations based on location alone.
