# Chapter refresh record: `poverty`

Human-readable ops checklist for regenerating the **poverty and wealth** chapter survey spine that feeds [`book/poverty.md`](../../../book/poverty.md). Feeds [**Task 44**](../../../refresh_manifest_plan.md) / future `analysis_manifest.yaml` and **Task 127**.

> *If a new GSS or ANES extract were released tomorrow, what would we need to regenerate this chapter’s attitude opening and exploratory candidate stack?*

**Naming:** file stem matches the book chapter (`poverty`). Variable lists = **EDA strong candidates** ([`eda_poverty.md`](../../eda_poverty.md)) plus wording companion `natfarey` — not the full secondary inventory ([`inventories/poverty_inventory.md`](../../inventories/poverty_inventory.md)). **One mixed list** (GSS + ANES); Slurm / local runners auto-detect survey from the variable name (`VCF*` → ANES).

**Status (2026-09-08):** Strong-candidate stack fit on Unity (4000/4000 @ 0.99); **`VCF0870` dropped** from active lists (sociotropic mood not useful here). Soft-gate follow-ups remain (esp. `natfare`, `VCF0839`). Objective Q1–Q4 series stay outside this Bayesian refresh path.

---

## Chapter metadata

| Field | Value |
|-------|--------|
| **Chapter / record id** | `poverty` |
| **Scope** | EDA strong candidates + `natfarey` (GSS + ANES), minus dropped mood item `VCF0870`. Featured opening: `getahead` / `natfare` / `dwelown`. |
| **Data version** | `gss2024_r3a` + current ANES CDF extract |
| **GSS extract** | `notebooks/gss_extract_2024_r3a.hdf` (`GSS_DEFAULT_HDF`) |
| **ANES extract** | `notebooks/anes_extract_anes_timeseries_cdf_stata_20260205.hdf` (`ANES_DEFAULT_HDF`) |
| **Survey-year mapping** | `notebooks/survey_year_mapping.csv` |
| **Refresh dates** | **2026-09-07** — R3a + EDA; local scouts. **2026-09-08** — Unity batch 1+2; mixed lists |
| **Conda env** | `CultureWar` |
| **Book consumer** | [`book/poverty.md`](../../../book/poverty.md) (**Task 127**); results digest [`jb/poverty_results.md`](../../../jb/poverty_results.md) (**Task 130**) |
| **EDA** | [`notebooks/eda_poverty.md`](../../eda_poverty.md) |
| **Related tasks** | **Task 127**, **44**/**45**, **86**/**87**, **84**/**126**, **128** |
| **Variable lists** | Full: [`slurm/variables_poverty.txt`](../../../slurm/variables_poverty.txt) (**20**). Batch 1: [`variables_poverty_batch1.txt`](../../../slurm/variables_poverty_batch1.txt) (4). Batch 2: [`variables_poverty_batch2.txt`](../../../slurm/variables_poverty_batch2.txt) (**16**) |
| **Unity driver** | [`slurm/survey_pipeline_array.sh`](../../../slurm/survey_pipeline_array.sh) — default **auto** data source |

**Sampling convention:**

| Tier | Where | `tune` / `draws` | `target_accept` | Role |
|------|-------|------------------|-----------------|------|
| **Local scout** | workstation | **2000 / 2000** | **0.95** | Smoke / metadata; `FORCE_RUN=1` |
| **Unity production** | cluster | **4000 / 4000** | **0.99** | Print-quality idata |
| Decade PC | either | — | — | Optional (`--pc-diag` on `run_survey_pipeline.sh`) |

**Artifact stems:** `rw2_overdisp_weighted_mean_binoround_ncrw2_nceps` (time_model); `double_cumsum_Exponential_c0.125_t0.125_weighted_mean_overdisp_initpriors_nc` (model10 / trajectories).

---

## Variables

| Variable | Survey | y=1 coding | Role |
|----------|--------|------------|------|
| ★ `getahead` | GSS | Hard work most important **[1]** | Featured #1 |
| ★ `natfare` | GSS | Welfare too much **[3]** | Featured #2 |
| ★ `dwelown` | GSS | Own or buying **[1]** | Featured #3 (descriptive) |
| `natfarey` | GSS | Assistance to poor too much **[3]** | Wording companion (do **not** pool) |
| `eqwlth` | GSS | Oppose reducing income differences **[5,6,7]** | Inequality |
| `satfin` | GSS | Pretty well satisfied **[1]** | Descriptive well-being |
| `finrela` | GSS | Below / far below average **[1,2]** | Descriptive relative standing |
| `class` | GSS | Working class **[2]** | Descriptive class ID |
| `goodlife` | GSS | Agree can improve living standard **[1,2]** | Descriptive outlook |
| `parsol` | GSS | Better than parents **[1,2]** | Descriptive generational |
| `VCF0809` | ANES | Let each person get ahead **[5,6,7]** | Redistribution / safety net |
| `VCF0886` | ANES | Decrease aid to the poor **[3]** | Near `natfarey` wording |
| `VCF0839` | ANES | Fewer services **[1,2,3]** | Services–spending |
| `VCF0806` | ANES | Private health insurance **[5,6,7]** | Medical COL |
| `VCF0148` | ANES | Lower / working class **[0–3]** | Class ID (descriptive; ends 2016) |
| `VCF9013` | ANES | Disagree equal-opportunity mandate **[4,5]** | Anti-egalitarian |
| `VCF9015` | ANES | Disagree unequal chance is big problem **[4,5]** | Anti-egalitarian (ends 2012) |
| `VCF9016` | ANES | Agree OK if unequal chance **[1,2]** | Anti-egalitarian |
| `VCF9017` | ANES | Agree worry less about equality **[1,2]** | Anti-egalitarian |
| `VCF0220` | ANES | Cold thermometer **[0–49]** | Cool toward people on welfare |

**Dropped from active spine:** `VCF0870` (economy past year) — sociotropic mood; not useful for this chapter’s attitude exploration.

**Valence:** Policy / spending / egalitarianism → **conservative y=1**. Descriptive items (`satfin`, `finrela`, `class`, `goodlife`, `parsol`, `dwelown`, `VCF0148`) are substantive, not partisan. Metadata for active list vars is in `utils.py`.

**Not in the list:** secondary / hand-offs (`natrace`→race, `nataid`→nationality, income covariates, thin ANES COL items, `VCF0870`, …); external Q1–Q4 series. `dwelown` may need an age–cohort model (**Task 132**) rather than over-reading period–cohort fans.

---

## Analysis workflow

National stack only for v1 (no sex compose). Per variable: time_model → model10 → cohort_trajectories.

### Local scout (2000 / 2000 @ 0.95)

One list = one extract (**Task 149** — no `--data-source auto`):

```bash
cd /home/downey/CultureWar
conda activate CultureWar
FORCE_RUN=1 ./scripts/run_survey_pipeline.sh \
  --extract gss_extract_2024_r3a.hdf \
  $(grep -vE '^\s*(#|$)' slurm/variables_poverty_gss.txt)
FORCE_RUN=1 ./scripts/run_survey_pipeline.sh \
  --extract anes_extract_anes_timeseries_cdf_stata_20260205.hdf \
  $(grep -vE '^\s*(#|$)' slurm/variables_poverty_anes.txt)
```

### Unity (4000 / 4000 @ 0.99)

```bash
# Batch 1 — featured four (GSS)
N=$(grep -cvE '^\s*(#|$)' slurm/variables_poverty_batch1.txt)
sbatch --array=0-$((N-1))%4 --export=ALL,CW_TARGET_ACCEPT=0.99 \
  slurm/survey_pipeline_array.sh slurm/variables_poverty_batch1.txt gss_extract_2024_r3a.hdf

# Batch 2 — GSS explorers
N=$(grep -cvE '^\s*(#|$)' slurm/variables_poverty_batch2_gss.txt)
sbatch --array=0-$((N-1))%4 --export=ALL,CW_TARGET_ACCEPT=0.99 \
  slurm/survey_pipeline_array.sh slurm/variables_poverty_batch2_gss.txt gss_extract_2024_r3a.hdf

# Batch 2 — ANES CDF
N=$(grep -cvE '^\s*(#|$)' slurm/variables_poverty_batch2_anes.txt)
sbatch --array=0-$((N-1))%4 --export=ALL,CW_TARGET_ACCEPT=0.99 \
  slurm/survey_pipeline_array.sh slurm/variables_poverty_batch2_anes.txt \
  anes_extract_anes_timeseries_cdf_stata_20260205.hdf
```

Full catalogs: `variables_poverty_gss.txt` / `variables_poverty_anes.txt`. Mixed `variables_poverty.txt` is summarizer-only — do not sbatch it.

### Post-pull summary

```bash
python scripts/summarize_survey_pipeline_runs.py \
  --list slurm/variables_poverty.txt \
  --job-id 64093732 --job-id 64094131 --job-id 64094145 \
  --out-csv notebooks/tables/poverty_unity_sampling_summary.csv
```

CSV: [`notebooks/tables/poverty_unity_sampling_summary.csv`](../../../notebooks/tables/poverty_unity_sampling_summary.csv).

---

## Unity history (2026-09-08)

| Wave | Job | List (as run) | Result |
|------|-----|---------------|--------|
| Batch 1 | **64093732** | featured 4 GSS | All `SURVEY_PIPELINE_OK`; **`natfare`** soft-fail |
| Batch 2a | **64094131** | 6 GSS explorers | All OK; `goodlife` mild R̂ soft-fail |
| Batch 2b | **64094145** | 11 ANES | All OK; **`VCF0839`** soft-fail; several mild divergences |

Going forward, batch 2 is the single mixed file `variables_poverty_batch2.txt` (**16** vars after dropping `VCF0870`).

---

## Sampling notes (Unity model10, 4000/4000 @ 0.99)

| Variable | Notes |
|----------|-------|
| `getahead`, `dwelown`, `natfarey`, `eqwlth`, `class`, `parsol`, most ANES | Clean or mild div only |
| **`natfare`** | Hard soft-fail — R̂ ~1.18, ESS ~16/8 (retry before print) |
| **`VCF0839`** | Soft-fail — R̂ ~1.07, ESS bulk ~47 |
| `goodlife` | Soft-fail — R̂ 1.032 (barely over gate) |
| Several ANES | 1–3 divergences (soft-fail on div=0 gate) |

Full table: re-run `summarize_survey_pipeline_runs.py` or see the CSV.

---

## Checklist

- [x] R3a extract + EDA + metadata for all 21
- [x] Unified `variables_poverty.txt` (+ batch1 / batch2); auto data-source in array / local runner
- [x] Unity batch 1 + batch 2 (historically split GSS/ANES jobs; lists now mixed)
- [x] Pull + [`summarize_survey_pipeline_runs.py`](../../../scripts/summarize_survey_pipeline_runs.py)
- [ ] Retry soft-gate fails (`natfare`, `VCF0839`, …)
- [ ] Explore figs; embed featured opening in [`book/poverty.md`](../../../book/poverty.md)
- [ ] Optional decade PC; Task **128** splash; external Q1–Q4

---

## Related

- Chapter stub: [`book/poverty.md`](../../../book/poverty.md)
- Topics index: [`README.md`](README.md)
- Gender template: [`gender_refresh.md`](gender_refresh.md)
- Unity ops: [`unity_workflow.md`](../../../unity_workflow.md)
- Runners: [`scripts/run_survey_pipeline.sh`](../../../scripts/run_survey_pipeline.sh), [`slurm/survey_pipeline_array.sh`](../../../slurm/survey_pipeline_array.sh)
- Summary: [`scripts/summarize_survey_pipeline_runs.py`](../../../scripts/summarize_survey_pipeline_runs.py)
- Auto source helper: `infer_data_source_for_variable` in [`notebooks/survey_data.py`](../../survey_data.py)
