Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
1196449
Add new Commerce EDA Poverty import
shvngisingh Aug 24, 2026
4b0b7da
Remove committed test fixture files and generate them dynamically in …
shvngisingh Aug 25, 2026
54aa562
Address code review findings for Commerce EDA Poverty import
shvngisingh Sep 10, 2026
0b2045e
Address review findings: add counters to source_files, expand test co…
shvngisingh Sep 10, 2026
d4a33e5
Restore output and counters directory structure matching usaspending,…
shvngisingh Sep 10, 2026
a6f21ff
Configure 100-row sample for Poverty_cleaned.csv and Poverty_output.csv
shvngisingh Sep 10, 2026
e9bd651
Address code review: process full 3,232 counties, restore place valid…
shvngisingh Sep 10, 2026
371d754
docs: add post-mortem document for Commerce EDA Poverty import
shvngisingh Sep 10, 2026
0af0843
Address PR #2178 review comments: split 2020/2021 poverty rates, untr…
shvngisingh Sep 15, 2026
5cd9e06
Remove gitignore, postmortem.md, and old fixtures; add 200-line test_…
shvngisingh Sep 15, 2026
bdaa18c
Comply with review checklist: replace logging.fatal+raise with loggin…
shvngisingh Sep 15, 2026
ff9589c
Address code review findings: remove python3 script prefix, add cron_…
shvngisingh Sep 15, 2026
83ed2b2
Address Saanika review comments: add expected test output fixture and…
shvngisingh Sep 18, 2026
ced4ec2
Create poverty download script and decouple process from GCS
shvngisingh Sep 22, 2026
54d2882
Update download and preprocessing scripts to source EDA Persistent Po…
shvngisingh Sep 23, 2026
69f5e36
Address linting in download_poverty.py and process_poverty.py
shvngisingh Sep 23, 2026
002e095
Fix CodeQL incomplete URL substring sanitization in download_poverty_…
shvngisingh Sep 23, 2026
44b4e1d
Address review findings from report 5996883579895808 for Commerce EDA…
shvngisingh Sep 24, 2026
32a37ba
Remove golden check from validation_config.json, manifest.json, and g…
shvngisingh Sep 24, 2026
a1437fd
Remove redundant download.py alias to keep single downloading and pro…
shvngisingh Sep 24, 2026
2701c11
Address polish findings [P3-01], [P3-02], and [P3-03] for Commerce ED…
shvngisingh Sep 24, 2026
b042cbe
Remove unused io import in process_poverty_test.py
shvngisingh Sep 24, 2026
8c783f2
Address review findings for Commerce EDA Poverty dataset import
shvngisingh Sep 29, 2026
238916a
Make poverty_pvmap.csv generic with long-format (GEOID, year, poverty…
shvngisingh Sep 30, 2026
d2c6faa
Address review findings: harden download fallback, robust header dete…
shvngisingh Sep 30, 2026
3bd81c4
Address CodeQL alert: replace backtracking regex with linear findall …
shvngisingh Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion import-automation/executor/requirements.txt
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ chardet
chromedriver_py
croniter
dataclasses
datacommons
datacommons==1.4.3
datacommons_client
db-dtypes
duckdb
Expand Down
178 changes: 178 additions & 0 deletions statvar_imports/commerce_eda_poverty/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,178 @@
# Commerce EDA Poverty: Persistent Poverty County Status and Rates

Author: Shivangi Singh
Date: *September 2026*

| Parameter | Details |
| :--- | :--- |
| Import Type | **Automated Import** (raw dataset downloaded directly from source website or local input file, and processed locally) |
| Link to dataset preview or raw data | [EDA Persistent Poverty Counties](https://www.eda.gov/performance/resources/persistent-poverty-counties) |
| Direct Workbook URL | [`EDA_FY23_PPCs.xlsx`](https://www.eda.gov/sites/default/files/2023-03/EDA_FY23_PPCs.xlsx) |
| Archive Mirror URL | [`EDA_FY23_PPCs.xlsx` (Wayback Machine)](https://web.archive.org/web/20250308204521if_/https://www.eda.gov/sites/default/files/2023-03/EDA_FY23_PPCs.xlsx) |
| Place types covered | U.S. Counties, County Equivalents, and Island Territories (`County` / `AdministrativeArea1`) |
| Place ID resolution | `country/USA` FIPS (5-digit county and county-equivalent FIPS codes, `geoId/XXXXX`) |
| Date range covered | 1990, 2000, 2020, 2021 (1990 Decennial Census, 2000 Decennial Census, 2020 Island Area Decennial Census, 2017–2021 ACS 5-Year) |
| Statistical Variables | `Count_Person_BelowPovertyLevelInThePast12Months_AsFractionOf_Count_Person` |
| Unit / Scaling | `Percent` / `100` |
| Refresh Cycle | Annual automated check (`30 05 1 1 *`) aligned with EDA/Census releases |

---

## Overview

This automated dataset import fetches and processes historical and recent county-level poverty percentage rates published by the U.S. Economic Development Administration (EDA) (`https://www.eda.gov/performance/resources/persistent-poverty-counties`) for Persistent Poverty Counties (PPCs) and all benchmarked U.S. counties.

The download script (`download_poverty.py`) downloads the official `EDA_FY23_PPCs.xlsx` workbook directly from EDA (with automatic fallback to the Wayback Machine archive mirror if Cloudflare bot detection blocks automated requests, or from a local file via `--input_file`) into `input_files/EDA_FY23_PPCs.xlsx`. It extracts the `Underlying_Data` sheet (3,241 county rows) into `input_files/Poverty.csv` and `output/Poverty_original.csv`.

The preprocessing script (`process_poverty.py`) reads the downloaded source file locally, cleans and standardizes 5-digit FIPS codes and poverty percentages into normalized `(GEOID, year, poverty_rate)` records in `output/Poverty_cleaned.csv`, and feeds the cleaned data into `stat_var_processor.py`. Neither script relies on Google Cloud Storage (GCS) staging.

The dataset benchmarks poverty rates across statutory periods:
- **1990**: 1990 Decennial Census (`year=1990`)
- **2000**: 2000 Decennial Census (`year=2000`)
- **2020**: 2020 Island Areas Decennial Census (`year=2020` for territory county equivalents `60`, `66`, `69`, `78`)
- **2021**: 2021 SAIPE / 2017–2021 ACS 5-Year Estimates (`year=2021` for all 50 states, DC, and Puerto Rico)

It covers 3,232 U.S. counties, county equivalents, and island territories. Island territories in the EDA dataset use 5-digit county-equivalent FIPS codes (e.g., 60010 for Eastern District, AS; 66010 for Guam; 69085 for Northern Islands, MP; 78010 for St. Croix, VI).

> [!NOTE]
> Direct automated HTTP requests to `eda.gov` may encounter Cloudflare bot protection (HTTP 403 Forbidden). `download_poverty.py` automatically falls back to an archive mirror of the official FY23 workbook. For manual/semi-automated refresh when upstream releases a new workbook, operators can download via a browser and provide it locally via `--input_file`.

---

## Statistical Variable

```mcf
Node: dcid:Count_Person_BelowPovertyLevelInThePast12Months_AsFractionOf_Count_Person
typeOf: dcid:StatisticalVariable
name: "Population: Below Poverty Level in The Past 12 Months (Per Capita)"
populationType: dcid:Person
measuredProperty: dcid:count
statType: dcid:measuredValue
measurementDenominator: dcid:Count_Person
povertyStatus: dcid:BelowPovertyLevelInThePast12Months
```

---

## Working Directory Context

Commands in this workflow depend on the working directory:
- **Module directory (`statvar_imports/commerce_eda_poverty/`)**: Execute download, preprocessing, and pipeline scripts (`download_poverty.py`, `process_poverty.py`, `stat_var_processor.py`).
- **Repository root (`data/`)**: Execute validation runner, test scripts (`./run_tests.sh`), and unittest module invocations.

---

## Pipeline Execution (Automated Import)

### Prerequisites
Ensure Python dependencies are available:
```bash
pip install pandas openpyxl requests absl-py duckdb
```

For running `stat_var_processor.py` locally, authenticate Google Cloud Application Default
Credentials to allow reading the central schema definitions from Google Cloud Storage:
```bash
gcloud auth application-default login
```
*(Required to read `--existing_statvar_mcf=gs://unresolved_mcf/scripts/statvar/stat_vars.mcf`).*

### 1. Download Source Dataset (`download_poverty.py`)
Run from `statvar_imports/commerce_eda_poverty/`. Downloads the official `EDA_FY23_PPCs.xlsx` workbook directly from EDA (or archive mirror / local input file) and extracts the `Underlying_Data` sheet to `input_files/Poverty.csv` (also staging `output/Poverty_original.csv`):
```bash
cd statvar_imports/commerce_eda_poverty
python3 download_poverty.py
```

To ingest a local copy directly without downloading:
```bash
python3 download_poverty.py --input_file=input_files/EDA_FY23_PPCs.xlsx
```

### 2. Preprocess Dataset (`process_poverty.py`)
Run from `statvar_imports/commerce_eda_poverty/`. Ingests the downloaded source file locally, standardizes 5-digit county and county-equivalent island territory FIPS codes, partitions the recent rates into 2020 (island territories) and 2021 (states, DC, PR), validates survey year bounds [2020..current year], enforces percentage value bounds $[0.0, 100.0]$, and atomically outputs `output/Poverty_cleaned.csv`:
```bash
python3 process_poverty.py
```

### 3. Generate Data Commons Observations (`stat_var_processor.py`)
Run from `statvar_imports/commerce_eda_poverty/`. Matches the scripts configured in `manifest.json`:
```bash
python3 ../../tools/statvar_importer/stat_var_processor.py \
--input_data=output/Poverty_cleaned.csv \
--pv_map=poverty_pvmap.csv \
--config_file=poverty_metadata.csv \
--output_path=output/Poverty_output \
--existing_statvar_mcf=gs://unresolved_mcf/scripts/statvar/stat_vars.mcf \
--output_counters=counters/Poverty_counters.csv
```

---

## Validation

Validate generated outputs using the Import Validation Framework and `validation_config.json`. Run from the repository root `data/`:
```bash
python3 -m tools.import_validation.runner \
--validation_config=statvar_imports/commerce_eda_poverty/validation_config.json \
--stats_summary=statvar_imports/commerce_eda_poverty/dc_generated/summary_report.csv \
--differ_output=statvar_imports/commerce_eda_poverty/dc_generated \
--lint_report=statvar_imports/commerce_eda_poverty/dc_generated/report.json \
--validation_output=statvar_imports/commerce_eda_poverty/dc_generated/validation_report.json
```

Validation rules configured in `validation_config.json`:
1. `check_percent_min_value`: Asserts poverty rate values are $\ge 0.0\%$.
2. `check_percent_max_value`: Asserts poverty rate values are $\le 100.0\%$.
3. `check_num_places_count`: Asserts total places count is between 3,100 and 3,250.
4. `check_num_observations_count`: Asserts total observation count is between 9,000 and 10,000.
5. `check_max_date_consistent`: Asserts MaxDate is consistent across StatVars.
6. `check_date_span_sql`: Asserts `TRY_CAST(MinDate AS INT) = 1990 AND TRY_CAST(MaxDate AS INT) >= 2021`.
7. `check_missing_refs_count`: Asserts zero unresolved entity or schema references.
8. `check_lint_error_count`: Asserts zero lint errors.
*(Note: `check_deleted_records_percent` is inherited from the base validation configuration with threshold 0).*

---

## Troubleshooting & Operational Runbook

### 1. Cloudflare Bot Detection / HTTP 403 Forbidden
- **Symptom:** `download_poverty.py` receives HTTP 403 when requesting EDA URLs.
- **Automatic Mitigation:** The script automatically catches non-200 responses and falls back to
the Wayback Machine archive mirror (`EDA_PPC_MIRROR_URL`).
- **Manual Workaround:** Download the workbook manually via a browser from the official
[EDA PPC page](https://www.eda.gov/performance/resources/persistent-poverty-counties) and
pass it via `--input_file`:
```bash
python3 download_poverty.py --input_file=input_files/EDA_FY23_PPCs.xlsx
```

### 2. Survey Year Mismatch / Upstream Schema Changes
- **Symptom:** `process_poverty.py` raises `ValueError: Unexpected survey year...`.
- **Cause:** EDA periodically releases updated PPC workbooks baselining newer Census/SAIPE
estimates (e.g. transitioning from SAIPE 2021 to 2022 or 2023).
- **Remediation:**
1. Inspect the new column headers and update the date mapping logic in `process_poverty.py`.
2. Update `poverty_metadata.csv` and `poverty_pvmap.csv` if new column names are introduced.
3. Update `validation_config.json` date span rules to accommodate the new maximum date.

### 3. County Count Threshold Failures
- **Symptom:** Validation fails on `NUM_PLACES_COUNT` outside `[3100, 3250]` or `process_poverty.py`
raises `Cleaned county count below minimum threshold`.
- **Cause:** Upstream sheet layout changes (e.g., altered sheet name, modified header row
offset, or unexpected GEOID formatting).
- **Remediation:** Check whether EDA modified the sheet structure or header rows. Re-run
preprocessing and inspect `output/Poverty_cleaned.csv`.

---

## Testing

Run unit tests verifying workbook download and mirror failover, local file ingestion, GEOID standardization, and value sanitation from the repository root `data/`:
```bash
python3 -m unittest discover -s statvar_imports/commerce_eda_poverty -p "*test*.py"
```
Or via the test runner script:
```bash
./run_tests.sh -p statvar_imports/commerce_eda_poverty
```
Loading
Loading