diff --git a/import-automation/executor/requirements.txt b/import-automation/executor/requirements.txt
index 0231dab808..ae652381fc 100644
--- a/import-automation/executor/requirements.txt
+++ b/import-automation/executor/requirements.txt
@@ -7,7 +7,7 @@ chardet
chromedriver_py
croniter
dataclasses
-datacommons
+datacommons==1.4.3
datacommons_client
db-dtypes
duckdb
diff --git a/statvar_imports/commerce_eda_poverty/README.md b/statvar_imports/commerce_eda_poverty/README.md
new file mode 100644
index 0000000000..538f8f15b3
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/README.md
@@ -0,0 +1,178 @@
+# Commerce EDA Poverty: Persistent Poverty County Status and Rates
+
+Author: Shivangi Singh
+Date: *September 2026*
+
+| Parameter | Details |
+| :--- | :--- |
+| Import Type | **Automated Import** (raw dataset downloaded directly from source website or local input file, and processed locally) |
+| Link to dataset preview or raw data | [EDA Persistent Poverty Counties](https://www.eda.gov/performance/resources/persistent-poverty-counties) |
+| Direct Workbook URL | [`EDA_FY23_PPCs.xlsx`](https://www.eda.gov/sites/default/files/2023-03/EDA_FY23_PPCs.xlsx) |
+| Archive Mirror URL | [`EDA_FY23_PPCs.xlsx` (Wayback Machine)](https://web.archive.org/web/20250308204521if_/https://www.eda.gov/sites/default/files/2023-03/EDA_FY23_PPCs.xlsx) |
+| Place types covered | U.S. Counties, County Equivalents, and Island Territories (`County` / `AdministrativeArea1`) |
+| Place ID resolution | `country/USA` FIPS (5-digit county and county-equivalent FIPS codes, `geoId/XXXXX`) |
+| Date range covered | 1990, 2000, 2020, 2021 (1990 Decennial Census, 2000 Decennial Census, 2020 Island Area Decennial Census, 2017–2021 ACS 5-Year) |
+| Statistical Variables | `Count_Person_BelowPovertyLevelInThePast12Months_AsFractionOf_Count_Person` |
+| Unit / Scaling | `Percent` / `100` |
+| Refresh Cycle | Annual automated check (`30 05 1 1 *`) aligned with EDA/Census releases |
+
+---
+
+## Overview
+
+This automated dataset import fetches and processes historical and recent county-level poverty percentage rates published by the U.S. Economic Development Administration (EDA) (`https://www.eda.gov/performance/resources/persistent-poverty-counties`) for Persistent Poverty Counties (PPCs) and all benchmarked U.S. counties.
+
+The download script (`download_poverty.py`) downloads the official `EDA_FY23_PPCs.xlsx` workbook directly from EDA (with automatic fallback to the Wayback Machine archive mirror if Cloudflare bot detection blocks automated requests, or from a local file via `--input_file`) into `input_files/EDA_FY23_PPCs.xlsx`. It extracts the `Underlying_Data` sheet (3,241 county rows) into `input_files/Poverty.csv` and `output/Poverty_original.csv`.
+
+The preprocessing script (`process_poverty.py`) reads the downloaded source file locally, cleans and standardizes 5-digit FIPS codes and poverty percentages into normalized `(GEOID, year, poverty_rate)` records in `output/Poverty_cleaned.csv`, and feeds the cleaned data into `stat_var_processor.py`. Neither script relies on Google Cloud Storage (GCS) staging.
+
+The dataset benchmarks poverty rates across statutory periods:
+- **1990**: 1990 Decennial Census (`year=1990`)
+- **2000**: 2000 Decennial Census (`year=2000`)
+- **2020**: 2020 Island Areas Decennial Census (`year=2020` for territory county equivalents `60`, `66`, `69`, `78`)
+- **2021**: 2021 SAIPE / 2017–2021 ACS 5-Year Estimates (`year=2021` for all 50 states, DC, and Puerto Rico)
+
+It covers 3,232 U.S. counties, county equivalents, and island territories. Island territories in the EDA dataset use 5-digit county-equivalent FIPS codes (e.g., 60010 for Eastern District, AS; 66010 for Guam; 69085 for Northern Islands, MP; 78010 for St. Croix, VI).
+
+> [!NOTE]
+> Direct automated HTTP requests to `eda.gov` may encounter Cloudflare bot protection (HTTP 403 Forbidden). `download_poverty.py` automatically falls back to an archive mirror of the official FY23 workbook. For manual/semi-automated refresh when upstream releases a new workbook, operators can download via a browser and provide it locally via `--input_file`.
+
+---
+
+## Statistical Variable
+
+```mcf
+Node: dcid:Count_Person_BelowPovertyLevelInThePast12Months_AsFractionOf_Count_Person
+typeOf: dcid:StatisticalVariable
+name: "Population: Below Poverty Level in The Past 12 Months (Per Capita)"
+populationType: dcid:Person
+measuredProperty: dcid:count
+statType: dcid:measuredValue
+measurementDenominator: dcid:Count_Person
+povertyStatus: dcid:BelowPovertyLevelInThePast12Months
+```
+
+---
+
+## Working Directory Context
+
+Commands in this workflow depend on the working directory:
+- **Module directory (`statvar_imports/commerce_eda_poverty/`)**: Execute download, preprocessing, and pipeline scripts (`download_poverty.py`, `process_poverty.py`, `stat_var_processor.py`).
+- **Repository root (`data/`)**: Execute validation runner, test scripts (`./run_tests.sh`), and unittest module invocations.
+
+---
+
+## Pipeline Execution (Automated Import)
+
+### Prerequisites
+Ensure Python dependencies are available:
+```bash
+pip install pandas openpyxl requests absl-py duckdb
+```
+
+For running `stat_var_processor.py` locally, authenticate Google Cloud Application Default
+Credentials to allow reading the central schema definitions from Google Cloud Storage:
+```bash
+gcloud auth application-default login
+```
+*(Required to read `--existing_statvar_mcf=gs://unresolved_mcf/scripts/statvar/stat_vars.mcf`).*
+
+### 1. Download Source Dataset (`download_poverty.py`)
+Run from `statvar_imports/commerce_eda_poverty/`. Downloads the official `EDA_FY23_PPCs.xlsx` workbook directly from EDA (or archive mirror / local input file) and extracts the `Underlying_Data` sheet to `input_files/Poverty.csv` (also staging `output/Poverty_original.csv`):
+```bash
+cd statvar_imports/commerce_eda_poverty
+python3 download_poverty.py
+```
+
+To ingest a local copy directly without downloading:
+```bash
+python3 download_poverty.py --input_file=input_files/EDA_FY23_PPCs.xlsx
+```
+
+### 2. Preprocess Dataset (`process_poverty.py`)
+Run from `statvar_imports/commerce_eda_poverty/`. Ingests the downloaded source file locally, standardizes 5-digit county and county-equivalent island territory FIPS codes, partitions the recent rates into 2020 (island territories) and 2021 (states, DC, PR), validates survey year bounds [2020..current year], enforces percentage value bounds $[0.0, 100.0]$, and atomically outputs `output/Poverty_cleaned.csv`:
+```bash
+python3 process_poverty.py
+```
+
+### 3. Generate Data Commons Observations (`stat_var_processor.py`)
+Run from `statvar_imports/commerce_eda_poverty/`. Matches the scripts configured in `manifest.json`:
+```bash
+python3 ../../tools/statvar_importer/stat_var_processor.py \
+ --input_data=output/Poverty_cleaned.csv \
+ --pv_map=poverty_pvmap.csv \
+ --config_file=poverty_metadata.csv \
+ --output_path=output/Poverty_output \
+ --existing_statvar_mcf=gs://unresolved_mcf/scripts/statvar/stat_vars.mcf \
+ --output_counters=counters/Poverty_counters.csv
+```
+
+---
+
+## Validation
+
+Validate generated outputs using the Import Validation Framework and `validation_config.json`. Run from the repository root `data/`:
+```bash
+python3 -m tools.import_validation.runner \
+ --validation_config=statvar_imports/commerce_eda_poverty/validation_config.json \
+ --stats_summary=statvar_imports/commerce_eda_poverty/dc_generated/summary_report.csv \
+ --differ_output=statvar_imports/commerce_eda_poverty/dc_generated \
+ --lint_report=statvar_imports/commerce_eda_poverty/dc_generated/report.json \
+ --validation_output=statvar_imports/commerce_eda_poverty/dc_generated/validation_report.json
+```
+
+Validation rules configured in `validation_config.json`:
+1. `check_percent_min_value`: Asserts poverty rate values are $\ge 0.0\%$.
+2. `check_percent_max_value`: Asserts poverty rate values are $\le 100.0\%$.
+3. `check_num_places_count`: Asserts total places count is between 3,100 and 3,250.
+4. `check_num_observations_count`: Asserts total observation count is between 9,000 and 10,000.
+5. `check_max_date_consistent`: Asserts MaxDate is consistent across StatVars.
+6. `check_date_span_sql`: Asserts `TRY_CAST(MinDate AS INT) = 1990 AND TRY_CAST(MaxDate AS INT) >= 2021`.
+7. `check_missing_refs_count`: Asserts zero unresolved entity or schema references.
+8. `check_lint_error_count`: Asserts zero lint errors.
+*(Note: `check_deleted_records_percent` is inherited from the base validation configuration with threshold 0).*
+
+---
+
+## Troubleshooting & Operational Runbook
+
+### 1. Cloudflare Bot Detection / HTTP 403 Forbidden
+- **Symptom:** `download_poverty.py` receives HTTP 403 when requesting EDA URLs.
+- **Automatic Mitigation:** The script automatically catches non-200 responses and falls back to
+ the Wayback Machine archive mirror (`EDA_PPC_MIRROR_URL`).
+- **Manual Workaround:** Download the workbook manually via a browser from the official
+ [EDA PPC page](https://www.eda.gov/performance/resources/persistent-poverty-counties) and
+ pass it via `--input_file`:
+ ```bash
+ python3 download_poverty.py --input_file=input_files/EDA_FY23_PPCs.xlsx
+ ```
+
+### 2. Survey Year Mismatch / Upstream Schema Changes
+- **Symptom:** `process_poverty.py` raises `ValueError: Unexpected survey year...`.
+- **Cause:** EDA periodically releases updated PPC workbooks baselining newer Census/SAIPE
+ estimates (e.g. transitioning from SAIPE 2021 to 2022 or 2023).
+- **Remediation:**
+ 1. Inspect the new column headers and update the date mapping logic in `process_poverty.py`.
+ 2. Update `poverty_metadata.csv` and `poverty_pvmap.csv` if new column names are introduced.
+ 3. Update `validation_config.json` date span rules to accommodate the new maximum date.
+
+### 3. County Count Threshold Failures
+- **Symptom:** Validation fails on `NUM_PLACES_COUNT` outside `[3100, 3250]` or `process_poverty.py`
+ raises `Cleaned county count below minimum threshold`.
+- **Cause:** Upstream sheet layout changes (e.g., altered sheet name, modified header row
+ offset, or unexpected GEOID formatting).
+- **Remediation:** Check whether EDA modified the sheet structure or header rows. Re-run
+ preprocessing and inspect `output/Poverty_cleaned.csv`.
+
+---
+
+## Testing
+
+Run unit tests verifying workbook download and mirror failover, local file ingestion, GEOID standardization, and value sanitation from the repository root `data/`:
+```bash
+python3 -m unittest discover -s statvar_imports/commerce_eda_poverty -p "*test*.py"
+```
+Or via the test runner script:
+```bash
+./run_tests.sh -p statvar_imports/commerce_eda_poverty
+```
diff --git a/statvar_imports/commerce_eda_poverty/download_poverty.py b/statvar_imports/commerce_eda_poverty/download_poverty.py
new file mode 100644
index 0000000000..181be21412
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/download_poverty.py
@@ -0,0 +1,392 @@
+# Copyright 2026 Google LLC
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+# https://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+
+"""Download script for Commerce EDA Persistent Poverty Counties (PPC) dataset.
+
+This script fetches the official Persistent Poverty Counties dataset from the
+U.S. Economic Development Administration (EDA) / Department of Commerce:
+ https://www.eda.gov/performance/tools/ (EDA_FY23_PPCs.xlsx)
+It downloads the official Excel workbook (with automatic mirror failover),
+supports ingesting directly from an existing input file, extracts the underlying
+county-level poverty data table (3,241 places across 1990, 2000, and 2020/2021),
+and stages the raw CSV and workbook under input_files/ and output/ for preprocessing.
+"""
+
+import csv
+import io
+import os
+import shutil
+import tempfile
+
+from absl import app, flags, logging
+import openpyxl
+import requests
+from requests.adapters import HTTPAdapter
+from urllib3.util import Retry
+
+MODULE_DIR = os.path.dirname(os.path.abspath(__file__))
+
+EDA_PPC_XLSX_URL = (
+ "https://www.eda.gov/sites/default/files/2023-03/EDA_FY23_PPCs.xlsx"
+)
+EDA_PPC_MIRROR_URL = (
+ "https://web.archive.org/web/20250308204521if_/"
+ "https://www.eda.gov/sites/default/files/2023-03/EDA_FY23_PPCs.xlsx"
+)
+
+DEFAULT_OUTPUT_DIR = os.path.join(MODULE_DIR, "input_files")
+DEFAULT_OUTPUT_XLSX = os.path.join(DEFAULT_OUTPUT_DIR, "EDA_FY23_PPCs.xlsx")
+DEFAULT_INPUT_CSV = os.path.join(DEFAULT_OUTPUT_DIR, "Poverty.csv")
+DEFAULT_RAW_CSV_FILE = os.path.join(MODULE_DIR, "output", "Poverty_original.csv")
+
+HTTP_HEADERS = {
+ "User-Agent": (
+ "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
+ "(KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
+ ),
+ "Accept": "*/*",
+}
+
+FLAGS = flags.FLAGS
+flags.DEFINE_string(
+ "source_url",
+ EDA_PPC_XLSX_URL,
+ "Primary URL to download the Persistent Poverty Counties Excel workbook.",
+)
+flags.DEFINE_string(
+ "mirror_url",
+ EDA_PPC_MIRROR_URL,
+ "Fallback mirror URL to download the Persistent Poverty Counties Excel workbook.",
+)
+flags.DEFINE_string(
+ "input_file",
+ None,
+ "Optional path to a local input file (.xlsx or .csv) to use instead of downloading.",
+)
+flags.DEFINE_string(
+ "output_xlsx_path",
+ DEFAULT_OUTPUT_XLSX,
+ "Destination path to save the downloaded source Excel workbook.",
+)
+flags.DEFINE_string(
+ "output_csv_path",
+ DEFAULT_INPUT_CSV,
+ "Destination path to save the extracted input CSV file.",
+)
+flags.DEFINE_string(
+ "raw_csv_path",
+ DEFAULT_RAW_CSV_FILE,
+ "Optional path to save an extracted raw CSV copy of the workbook.",
+)
+flags.DEFINE_integer(
+ "max_retries",
+ 3,
+ "Maximum number of download retry attempts per URL.",
+)
+flags.DEFINE_integer(
+ "timeout",
+ 60,
+ "HTTP request timeout in seconds.",
+)
+
+
+def download_file(
+ download_url,
+ output_path,
+ session=None,
+ max_retries=3,
+ backoff_factor=1.5,
+ timeout=60,
+):
+ """Downloads a file from download_url and saves it atomically to output_path."""
+ logging.info("Downloading file from: %s", download_url)
+ dst_dir = os.path.dirname(os.path.abspath(output_path))
+ os.makedirs(dst_dir, exist_ok=True)
+
+ owns_session = session is None
+ active_session = requests.Session() if owns_session else session
+ temp_path = None
+
+ try:
+ adapter = HTTPAdapter(
+ max_retries=Retry(
+ total=max_retries,
+ backoff_factor=backoff_factor,
+ status_forcelist=[429, 500, 502, 503, 504],
+ allowed_methods=["GET"],
+ raise_on_status=True,
+ )
+ )
+ if hasattr(active_session, "mount") and callable(active_session.mount):
+ active_session.mount("https://", adapter)
+ active_session.mount("http://", adapter)
+
+ try:
+ response = active_session.get(
+ download_url, headers=HTTP_HEADERS, timeout=timeout
+ )
+ if response.status_code == 404:
+ logging.warning(
+ "HTTP 404 received for %s; failing fast without retry.",
+ download_url,
+ )
+ raise requests.HTTPError(
+ f"404 Client Error: Not Found for url: {download_url}",
+ response=response,
+ )
+ response.raise_for_status()
+ content = response.content
+ if not content:
+ raise RuntimeError(
+ f"Empty response body received from {download_url}"
+ )
+ with tempfile.NamedTemporaryFile(
+ "wb", dir=dst_dir, delete=False, suffix=".tmp"
+ ) as tmp_file:
+ temp_path = tmp_file.name
+ tmp_file.write(content)
+ except requests.RequestException as e:
+ logging.error("Failed to download %s: %s", download_url, e)
+ raise RuntimeError(f"Failed to download {download_url}: {e}") from e
+
+ os.replace(temp_path, output_path)
+ temp_path = None
+ logging.info(
+ "Download completed successfully. Saved %d bytes to %s",
+ len(content),
+ output_path,
+ )
+ return content
+ finally:
+ if temp_path and os.path.exists(temp_path):
+ try:
+ os.unlink(temp_path)
+ except OSError:
+ pass
+ if owns_session:
+ active_session.close()
+
+
+def extract_sheet_to_csv(
+ excel_source,
+ csv_output_path,
+ target_sheet_name="Underlying_Data",
+):
+ """Extracts the underlying data worksheet from Excel workbook to a raw CSV."""
+ if isinstance(excel_source, bytes):
+ wb = openpyxl.load_workbook(io.BytesIO(excel_source), data_only=True)
+ else:
+ wb = openpyxl.load_workbook(excel_source, data_only=True)
+
+ try:
+ sheet_names = wb.sheetnames
+ selected_sheet = None
+ if target_sheet_name in sheet_names:
+ selected_sheet = target_sheet_name
+ else:
+ for name in sheet_names:
+ if any(
+ k in name.lower()
+ for k in ["underlying", "poverty", "data", "ppc"]
+ ):
+ selected_sheet = name
+ break
+ if not selected_sheet:
+ selected_sheet = sheet_names[0]
+ logging.warning(
+ "Worksheet '%s' not found in %s; falling back to '%s'.",
+ target_sheet_name,
+ sheet_names,
+ selected_sheet,
+ )
+
+ logging.info("Extracting sheet '%s' from Excel workbook...", selected_sheet)
+ ws = wb[selected_sheet]
+
+ dst_dir = os.path.dirname(os.path.abspath(csv_output_path))
+ os.makedirs(dst_dir, exist_ok=True)
+
+ row_count = 0
+ tmp_path = None
+ try:
+ with tempfile.NamedTemporaryFile(
+ "w",
+ dir=dst_dir,
+ delete=False,
+ suffix=".tmp",
+ encoding="utf-8",
+ newline="",
+ ) as tmp:
+ tmp_path = tmp.name
+ writer = csv.writer(tmp)
+ for row in ws.iter_rows(values_only=True):
+ if not any(row):
+ continue
+ writer.writerow([("" if c is None else str(c)) for c in row])
+ row_count += 1
+
+ os.replace(tmp_path, csv_output_path)
+ tmp_path = None
+ finally:
+ if tmp_path and os.path.exists(tmp_path):
+ try:
+ os.unlink(tmp_path)
+ except OSError:
+ pass
+
+ logging.info(
+ "Extracted %d rows from sheet '%s' to %s",
+ row_count,
+ selected_sheet,
+ csv_output_path,
+ )
+ return csv_output_path
+ finally:
+ wb.close()
+
+
+def copy_file_atomically(src_path, dst_path):
+ """Copies src_path to dst_path atomically."""
+ dst_dir = os.path.dirname(os.path.abspath(dst_path))
+ os.makedirs(dst_dir, exist_ok=True)
+ tmp_path = None
+ try:
+ with tempfile.NamedTemporaryFile(
+ "wb", dir=dst_dir, delete=False, suffix=".tmp"
+ ) as tmp:
+ tmp_path = tmp.name
+ with open(src_path, "rb") as fsrc:
+ shutil.copyfileobj(fsrc, tmp)
+ os.replace(tmp_path, dst_path)
+ tmp_path = None
+ finally:
+ if tmp_path and os.path.exists(tmp_path):
+ try:
+ os.unlink(tmp_path)
+ except OSError:
+ pass
+ logging.info("Copied %s to %s", src_path, dst_path)
+
+
+def download_poverty_dataset(
+ source_url=EDA_PPC_XLSX_URL,
+ mirror_url=EDA_PPC_MIRROR_URL,
+ input_file=None,
+ output_xlsx_path=DEFAULT_OUTPUT_XLSX,
+ output_csv_path=DEFAULT_INPUT_CSV,
+ raw_csv_path=DEFAULT_RAW_CSV_FILE,
+ max_retries=3,
+ timeout=60,
+):
+ """Main workflow to download official EDA PPC workbook or ingest input file."""
+ # Case 1: Local input file specified
+ if input_file:
+ if not os.path.exists(input_file) or os.path.getsize(input_file) == 0:
+ raise FileNotFoundError(f"Input file not found or empty: {input_file}")
+ logging.info("Using provided local input file: %s", input_file)
+ if input_file.lower().endswith((".xlsx", ".xls")):
+ copy_file_atomically(input_file, output_xlsx_path)
+ extract_sheet_to_csv(input_file, output_csv_path)
+ else:
+ copy_file_atomically(input_file, output_csv_path)
+ if raw_csv_path:
+ copy_file_atomically(output_csv_path, raw_csv_path)
+ return output_csv_path
+
+ # Clean up existing target files before a fresh download to avoid reusing stale files
+ for path_to_clean in [output_xlsx_path, output_csv_path, raw_csv_path]:
+ if path_to_clean and os.path.exists(path_to_clean):
+ try:
+ os.remove(path_to_clean)
+ logging.info(
+ "Cleaned up existing target file before fresh download: %s",
+ path_to_clean,
+ )
+ except OSError as e:
+ logging.warning(
+ "Could not remove existing file %s: %s", path_to_clean, e
+ )
+
+ # Case 2: Download from web with primary and mirror fallback
+ with requests.Session() as session:
+ content = None
+ urls_to_try = []
+ if source_url:
+ urls_to_try.append(source_url)
+ if mirror_url and mirror_url != source_url:
+ urls_to_try.append(mirror_url)
+
+ last_err = None
+ for url in urls_to_try:
+ try:
+ content = download_file(
+ download_url=url,
+ output_path=output_xlsx_path,
+ session=session,
+ max_retries=max_retries,
+ timeout=timeout,
+ )
+ if not content.startswith(b"PK\x03\x04"):
+ raise ValueError(
+ f"Downloaded content from {url} is not a valid ZIP/XLSX archive."
+ )
+ # Extract Underlying_Data sheet to output_csv_path
+ extract_sheet_to_csv(content, output_csv_path)
+ logging.info(
+ "Successfully downloaded workbook and extracted sheet from %s", url
+ )
+ break
+ except Exception as e:
+ last_err = e
+ logging.warning("Failed to download or process from %s: %s", url, e)
+ if os.path.exists(output_xlsx_path):
+ try:
+ os.remove(output_xlsx_path)
+ except OSError:
+ pass
+ if os.path.exists(output_csv_path):
+ try:
+ os.remove(output_csv_path)
+ except OSError:
+ pass
+ content = None
+
+ if not content:
+ raise RuntimeError(
+ f"Failed to acquire dataset from all URLs: {urls_to_try}. Last error: {last_err}"
+ ) from last_err
+
+ if raw_csv_path:
+ copy_file_atomically(output_csv_path, raw_csv_path)
+
+ return output_csv_path
+
+
+def main(argv):
+ """Main entrypoint for downloading the poverty dataset."""
+ del argv # Unused
+ download_poverty_dataset(
+ source_url=FLAGS.source_url,
+ mirror_url=FLAGS.mirror_url,
+ input_file=FLAGS.input_file,
+ output_xlsx_path=FLAGS.output_xlsx_path,
+ output_csv_path=FLAGS.output_csv_path,
+ raw_csv_path=FLAGS.raw_csv_path,
+ max_retries=FLAGS.max_retries,
+ timeout=FLAGS.timeout,
+ )
+
+
+if __name__ == "__main__":
+ app.run(main)
diff --git a/statvar_imports/commerce_eda_poverty/download_poverty_test.py b/statvar_imports/commerce_eda_poverty/download_poverty_test.py
new file mode 100644
index 0000000000..b86995b3f0
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/download_poverty_test.py
@@ -0,0 +1,459 @@
+# Copyright 2026 Google LLC
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+# https://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+
+"""Unit tests for Commerce EDA Persistent Poverty Counties download script."""
+
+import io
+import os
+import sys
+import tempfile
+import unittest
+from unittest import mock
+from urllib import parse
+
+import openpyxl
+import pandas as pd
+import requests
+from requests.adapters import HTTPAdapter
+
+MODULE_DIR = os.path.dirname(os.path.abspath(__file__))
+PROJECT_ROOT = os.path.abspath(os.path.join(MODULE_DIR, "..", ".."))
+sys.path.insert(0, PROJECT_ROOT)
+
+from statvar_imports.commerce_eda_poverty.download_poverty import (
+ EDA_PPC_MIRROR_URL,
+ EDA_PPC_XLSX_URL,
+ download_file,
+ download_poverty_dataset,
+ extract_sheet_to_csv,
+)
+
+
+def _create_mock_eda_workbook(filepath=None):
+ """Creates a mock Excel workbook containing the Underlying_Data worksheet."""
+ wb = openpyxl.Workbook()
+ try:
+ # Sheet 1: Readme
+ ws_readme = wb.active
+ ws_readme.title = "EDA Read Me"
+ ws_readme.append(["FY2023 PERSISTENT POVERTY COUNTIES (PPCs)", ""])
+
+ # Sheet 2: Underlying_Data
+ ws_data = wb.create_sheet(title="Underlying_Data")
+ ws_data.append([
+ "Table. FY2023 Persistent Poverty County Status - as of Data Year 2021",
+ "", "", "", "", "", "", ""
+ ])
+ ws_data.append([
+ "Identifing Information", "", "Census Bureau Data", "", "", "",
+ "FY23 Persistent Poverty", "Census GEO PPC Code"
+ ])
+ ws_data.append([
+ "Name",
+ "GEOID",
+ "1990 Decennial Census, % in Poverty",
+ "2000 Decennial Census, % in Poverty",
+ "Most Recent Estimate, % in Poverty* ",
+ "Data Source―Most Recent Estimate",
+ "",
+ "",
+ ])
+ ws_data.append(["Autauga County, AL", "01001", 15.7, 10.9, 13.3, "SAIPE, 2021", "No", 1])
+ ws_data.append(["Barbour County, AL", "01005", 25.2, 26.8, 29.0, "SAIPE, 2021", "Yes", 2])
+ ws_data.append([
+ "Eastern District, AS", "60010", 56.0, 58.6, 52.2, "Decennial Census, 2020", "Yes", 2
+ ])
+
+ if filepath:
+ wb.save(filepath)
+ return filepath
+
+ bio = io.BytesIO()
+ wb.save(bio)
+ return bio.getvalue()
+ finally:
+ wb.close()
+
+
+class TestDownloadPoverty(unittest.TestCase):
+
+ def test_download_file_success(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ out_file = os.path.join(tmpdir, "test.xlsx")
+ test_content = b"PK\x03\x04test_content"
+
+ mock_resp = mock.MagicMock()
+ mock_resp.__enter__.return_value = mock_resp
+ mock_resp.content = test_content
+ mock_resp.status_code = 200
+
+ mock_session = mock.MagicMock()
+ mock_session.get.return_value = mock_resp
+
+ content = download_file(
+ "https://example.gov/EDA_FY23_PPCs.xlsx",
+ out_file,
+ session=mock_session,
+ max_retries=1,
+ )
+
+ self.assertEqual(content, test_content)
+ self.assertEqual(mock_session.get.call_count, 1)
+ self.assertTrue(os.path.exists(out_file))
+ with open(out_file, "rb") as f:
+ self.assertEqual(f.read(), test_content)
+
+ def test_download_file_retry_and_succeed(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ out_file = os.path.join(tmpdir, "retry.xlsx")
+ test_content = b"ExcelData"
+
+ resp_503 = mock.MagicMock(
+ status=503, reason="Service Unavailable", msg=None, headers={}
+ )
+ resp_503.getheaders.return_value = []
+ resp_503.get_redirect_location.return_value = None
+ resp_503.isclosed.return_value = True
+
+ resp_200 = mock.MagicMock(status=200, reason="OK", msg=None, headers={})
+ resp_200.getheaders.return_value = []
+ resp_200.get_redirect_location.return_value = None
+ resp_200.isclosed.return_value = True
+ resp_200.data = test_content
+ resp_200.read.return_value = test_content
+ resp_200.stream.return_value = [test_content]
+
+ with mock.patch(
+ "urllib3.connectionpool.HTTPConnectionPool._make_request",
+ side_effect=[resp_503, resp_200],
+ ) as mock_make_request:
+ content = download_file(
+ "https://example.gov/EDA_FY23_PPCs.xlsx",
+ out_file,
+ max_retries=3,
+ backoff_factor=0.01,
+ )
+
+ self.assertEqual(mock_make_request.call_count, 2)
+ self.assertEqual(content, test_content)
+ self.assertTrue(os.path.exists(out_file))
+ with open(out_file, "rb") as f:
+ self.assertEqual(f.read(), test_content)
+
+ def test_download_file_fails_after_retries(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ out_file = os.path.join(tmpdir, "fail.xlsx")
+
+ resp_503 = mock.MagicMock(
+ status=503, reason="Service Unavailable", msg=None, headers={}
+ )
+ resp_503.getheaders.return_value = []
+ resp_503.get_redirect_location.return_value = None
+ resp_503.isclosed.return_value = True
+
+ with mock.patch(
+ "urllib3.connectionpool.HTTPConnectionPool._make_request",
+ side_effect=[resp_503, resp_503, resp_503],
+ ) as mock_make_request:
+ with self.assertRaises(RuntimeError):
+ download_file(
+ "https://example.gov/EDA_FY23_PPCs.xlsx",
+ out_file,
+ max_retries=2,
+ backoff_factor=0.01,
+ )
+
+ self.assertEqual(mock_make_request.call_count, 3)
+ self.assertFalse(os.path.exists(out_file))
+
+ def test_download_file_empty_body_raises(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ out_file = os.path.join(tmpdir, "empty.xlsx")
+
+ mock_resp = mock.MagicMock()
+ mock_resp.__enter__.return_value = mock_resp
+ mock_resp.content = b""
+ mock_resp.status_code = 200
+
+ mock_session = mock.MagicMock()
+ mock_session.get.return_value = mock_resp
+
+ with self.assertRaises(RuntimeError):
+ download_file(
+ "https://example.gov/EDA_FY23_PPCs.xlsx",
+ out_file,
+ session=mock_session,
+ max_retries=1,
+ )
+
+ self.assertFalse(os.path.exists(out_file))
+
+ def test_download_file_mounts_adapter_on_session(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ out_file = os.path.join(tmpdir, "mount.xlsx")
+ mock_session = mock.MagicMock()
+ mock_resp = mock.MagicMock(status_code=200, content=b"data")
+ mock_resp.__enter__.return_value = mock_resp
+ mock_session.get.return_value = mock_resp
+
+ download_file(
+ "https://example.gov/mount.xlsx",
+ out_file,
+ session=mock_session,
+ max_retries=2,
+ )
+
+ self.assertEqual(mock_session.mount.call_count, 2)
+ mounted = {
+ call[0][0]: call[0][1]
+ for call in mock_session.mount.call_args_list
+ }
+ self.assertIn("https://", mounted)
+ self.assertIn("http://", mounted)
+ self.assertIsInstance(mounted["https://"], HTTPAdapter)
+ self.assertIsInstance(mounted["http://"], HTTPAdapter)
+ self.assertEqual(mounted["https://"].max_retries.total, 2)
+ self.assertEqual(mounted["http://"].max_retries.total, 2)
+
+ def test_download_file_fast_fails_on_404(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ out_file = os.path.join(tmpdir, "404.xlsx")
+ mock_session = mock.MagicMock()
+ mock_resp = mock.MagicMock(status_code=404)
+ mock_resp.__enter__.return_value = mock_resp
+ mock_session.get.return_value = mock_resp
+
+ with self.assertRaises(RuntimeError) as ctx:
+ download_file(
+ "https://example.gov/404.xlsx",
+ out_file,
+ session=mock_session,
+ max_retries=3,
+ )
+
+ self.assertIn("404", str(ctx.exception))
+ self.assertEqual(mock_session.get.call_count, 1)
+ self.assertFalse(os.path.exists(out_file))
+
+ def test_download_file_session_without_mount_supported(self):
+ class DuckSession:
+
+ def __init__(self):
+ self.call_count = 0
+
+ def get(self, url, headers=None, timeout=None, **kwargs):
+ self.call_count += 1
+ resp = mock.MagicMock()
+ resp.__enter__.return_value = resp
+ resp.status_code = 200
+ resp.content = b"duck_data"
+ return resp
+
+ with tempfile.TemporaryDirectory() as tmpdir:
+ out_file = os.path.join(tmpdir, "duck.xlsx")
+ duck_session = DuckSession()
+ content = download_file(
+ "https://example.gov/duck.xlsx",
+ out_file,
+ session=duck_session,
+ )
+
+ self.assertEqual(content, b"duck_data")
+ self.assertEqual(duck_session.call_count, 1)
+ self.assertTrue(os.path.exists(out_file))
+ with open(out_file, "rb") as f:
+ self.assertEqual(f.read(), b"duck_data")
+
+ def test_extract_sheet_to_csv_from_bytes(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ csv_path = os.path.join(tmpdir, "extracted.csv")
+ excel_bytes = _create_mock_eda_workbook()
+
+ extract_sheet_to_csv(excel_bytes, csv_path, target_sheet_name="Underlying_Data")
+ self.assertTrue(os.path.exists(csv_path))
+
+ df = pd.read_csv(csv_path, skiprows=2, dtype=str)
+ self.assertEqual(len(df), 3)
+ self.assertIn("GEOID", df.columns)
+ self.assertEqual(list(df["GEOID"]), ["01001", "01005", "60010"])
+
+ def test_extract_sheet_to_csv_from_file_path(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ xlsx_path = os.path.join(tmpdir, "EDA_FY23_PPCs.xlsx")
+ csv_path = os.path.join(tmpdir, "Poverty.csv")
+ _create_mock_eda_workbook(xlsx_path)
+
+ extract_sheet_to_csv(xlsx_path, csv_path)
+ self.assertTrue(os.path.exists(csv_path))
+
+ df = pd.read_csv(csv_path, skiprows=2, dtype=str)
+ self.assertEqual(len(df), 3)
+ self.assertIn("GEOID", df.columns)
+
+ def test_download_poverty_dataset_with_input_file_csv(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ src_csv = os.path.join(tmpdir, "source_input.csv")
+ dst_xlsx = os.path.join(tmpdir, "out.xlsx")
+ dst_csv = os.path.join(tmpdir, "out.csv")
+ raw_csv = os.path.join(tmpdir, "raw.csv")
+
+ with open(src_csv, "w", encoding="utf-8") as f:
+ f.write("Line 1\nLine 2\nName,GEOID\nCounty A,01001\n")
+
+ res = download_poverty_dataset(
+ input_file=src_csv,
+ output_xlsx_path=dst_xlsx,
+ output_csv_path=dst_csv,
+ raw_csv_path=raw_csv,
+ )
+
+ self.assertEqual(res, dst_csv)
+ self.assertTrue(os.path.exists(dst_csv))
+ self.assertTrue(os.path.exists(raw_csv))
+ with open(dst_csv, "r", encoding="utf-8") as f:
+ content = f.read()
+ self.assertIn("County A,01001", content)
+
+ def test_download_poverty_dataset_with_input_file_xlsx(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ src_xlsx = os.path.join(tmpdir, "source_input.xlsx")
+ dst_xlsx = os.path.join(tmpdir, "EDA_FY23_PPCs.xlsx")
+ dst_csv = os.path.join(tmpdir, "Poverty.csv")
+ raw_csv = os.path.join(tmpdir, "Poverty_original.csv")
+
+ _create_mock_eda_workbook(src_xlsx)
+
+ res = download_poverty_dataset(
+ input_file=src_xlsx,
+ output_xlsx_path=dst_xlsx,
+ output_csv_path=dst_csv,
+ raw_csv_path=raw_csv,
+ )
+
+ self.assertEqual(res, dst_csv)
+ self.assertTrue(os.path.exists(dst_xlsx))
+ self.assertTrue(os.path.exists(dst_csv))
+ self.assertTrue(os.path.exists(raw_csv))
+
+ df = pd.read_csv(dst_csv, skiprows=2, dtype=str)
+ self.assertEqual(len(df), 3)
+
+ def test_download_poverty_dataset_primary_fail_mirror_succeed(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ dst_xlsx = os.path.join(tmpdir, "EDA_FY23_PPCs.xlsx")
+ dst_csv = os.path.join(tmpdir, "Poverty.csv")
+ raw_csv = os.path.join(tmpdir, "Poverty_original.csv")
+
+ excel_bytes = _create_mock_eda_workbook()
+
+ def mock_get(url, **kwargs):
+ resp = mock.MagicMock()
+ resp.__enter__.return_value = resp
+ if parse.urlparse(url).netloc == "www.eda.gov":
+ resp.status_code = 403
+ resp.raise_for_status.side_effect = requests.HTTPError("403 Forbidden")
+ return resp
+ resp.status_code = 200
+ resp.content = excel_bytes
+ return resp
+
+ with mock.patch("requests.Session.get", side_effect=mock_get):
+ res = download_poverty_dataset(
+ source_url=EDA_PPC_XLSX_URL,
+ mirror_url=EDA_PPC_MIRROR_URL,
+ output_xlsx_path=dst_xlsx,
+ output_csv_path=dst_csv,
+ raw_csv_path=raw_csv,
+ max_retries=1,
+ )
+
+ self.assertEqual(res, dst_csv)
+ self.assertTrue(os.path.exists(dst_xlsx))
+ self.assertTrue(os.path.exists(dst_csv))
+ df = pd.read_csv(dst_csv, skiprows=2, dtype=str)
+ self.assertEqual(len(df), 3)
+
+ def test_download_poverty_dataset_missing_input_file_raises(self):
+ with self.assertRaises(FileNotFoundError):
+ download_poverty_dataset(input_file="/nonexistent/path/Poverty.csv")
+
+ def test_download_poverty_dataset_all_urls_fail_raises_and_cleans_stale_files(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ dst_xlsx = os.path.join(tmpdir, "EDA_FY23_PPCs.xlsx")
+ dst_csv = os.path.join(tmpdir, "Poverty.csv")
+ dst_raw_csv = os.path.join(tmpdir, "Poverty_original.csv")
+ for p in [dst_xlsx, dst_csv, dst_raw_csv]:
+ with open(p, "w", encoding="utf-8") as f:
+ f.write("stale_data")
+
+ def mock_get(url, **kwargs):
+ resp = mock.MagicMock()
+ resp.status_code = 500
+ resp.raise_for_status.side_effect = requests.HTTPError("500 Server Error")
+ return resp
+
+ with mock.patch("requests.Session.get", side_effect=mock_get):
+ with self.assertRaises(RuntimeError) as ctx:
+ download_poverty_dataset(
+ source_url="https://example.gov/p1.xlsx",
+ mirror_url="https://example.gov/m1.xlsx",
+ output_xlsx_path=dst_xlsx,
+ output_csv_path=dst_csv,
+ raw_csv_path=dst_raw_csv,
+ max_retries=1,
+ )
+
+ self.assertIn("Failed to acquire dataset from all URLs", str(ctx.exception))
+ self.assertFalse(os.path.exists(dst_xlsx))
+ self.assertFalse(os.path.exists(dst_csv))
+ self.assertFalse(os.path.exists(dst_raw_csv))
+
+ def test_download_poverty_dataset_primary_html_challenge_mirror_succeed(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ dst_xlsx = os.path.join(tmpdir, "EDA_FY23_PPCs.xlsx")
+ dst_csv = os.path.join(tmpdir, "Poverty.csv")
+ raw_csv = os.path.join(tmpdir, "Poverty_original.csv")
+
+ excel_bytes = _create_mock_eda_workbook()
+
+ def mock_get(url, **kwargs):
+ resp = mock.MagicMock()
+ resp.__enter__.return_value = resp
+ resp.status_code = 200
+ if parse.urlparse(url).netloc == "www.eda.gov":
+ # Cloudflare challenge served with HTTP 200 HTML body
+ resp.content = b"
Just a moment..."
+ return resp
+ resp.content = excel_bytes
+ return resp
+
+ with mock.patch("requests.Session.get", side_effect=mock_get):
+ res = download_poverty_dataset(
+ source_url=EDA_PPC_XLSX_URL,
+ mirror_url=EDA_PPC_MIRROR_URL,
+ output_xlsx_path=dst_xlsx,
+ output_csv_path=dst_csv,
+ raw_csv_path=raw_csv,
+ max_retries=1,
+ )
+
+ self.assertEqual(res, dst_csv)
+ self.assertTrue(os.path.exists(dst_xlsx))
+ self.assertTrue(os.path.exists(dst_csv))
+ df = pd.read_csv(dst_csv, skiprows=2, dtype=str)
+ self.assertEqual(len(df), 3)
+
+
+if __name__ == "__main__":
+ unittest.main()
diff --git a/statvar_imports/commerce_eda_poverty/manifest.json b/statvar_imports/commerce_eda_poverty/manifest.json
new file mode 100644
index 0000000000..1205c6d5bc
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/manifest.json
@@ -0,0 +1,39 @@
+{
+ "import_specifications": [
+ {
+ "import_name": "Commerce_EDA_Poverty",
+ "curator_emails": [
+ "support@datacommons.org"
+ ],
+ "provenance_url": "https://www.eda.gov/performance/resources/persistent-poverty-counties",
+ "provenance_description": "U.S. Economic Development Administration (EDA) - Persistent Poverty Counties.",
+ "cron_schedule": "30 05 1 1 *",
+ "scripts": [
+ "download_poverty.py",
+ "process_poverty.py",
+ "../../tools/statvar_importer/stat_var_processor.py --input_data=output/Poverty_cleaned.csv --pv_map=poverty_pvmap.csv --config_file=poverty_metadata.csv --output_path=output/Poverty_output --existing_statvar_mcf=gs://unresolved_mcf/scripts/statvar/stat_vars.mcf --output_counters=counters/Poverty_counters.csv"
+ ],
+ "import_inputs": [
+ {
+ "template_mcf": "output/Poverty_output.tmcf",
+ "cleaned_csv": "output/Poverty_output.csv",
+ "node_mcf": "output/*.mcf"
+ }
+ ],
+ "validation_config_file": "validation_config.json",
+ "source_files": [
+ "download_poverty.py",
+ "process_poverty.py",
+ "input_files/*",
+ "output/*.mcf",
+ "output/Poverty_cleaned.csv",
+ "output/Poverty_original.csv",
+ "counters/*",
+ "validation_config.json",
+ "poverty_metadata.csv",
+ "poverty_pvmap.csv",
+ "manifest.json"
+ ]
+ }
+ ]
+}
diff --git a/statvar_imports/commerce_eda_poverty/poverty_metadata.csv b/statvar_imports/commerce_eda_poverty/poverty_metadata.csv
new file mode 100644
index 0000000000..20e3518262
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/poverty_metadata.csv
@@ -0,0 +1,13 @@
+parameter,value
+url,https://www.eda.gov/performance/resources/persistent-poverty-counties
+description,Persistent Poverty County Status (sourced from U.S. Economic Development Administration)
+#place_type,County
+#places_within,country/USA
+#start_date,1990
+#end_date,2021
+release_frequency,P1Y
+process,
+comments,
+output_columns,"observationAbout,observationDate,variableMeasured,value,unit,scalingFactor"
+header_rows,1
+output_only_new_statvars,True
diff --git a/statvar_imports/commerce_eda_poverty/poverty_pvmap.csv b/statvar_imports/commerce_eda_poverty/poverty_pvmap.csv
new file mode 100644
index 0000000000..cc69d3bcf2
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/poverty_pvmap.csv
@@ -0,0 +1,4 @@
+key,prop,value,p1,v1,p2,v2,p3,v3
+GEOID,observationAbout,geoId/{Data},,,,,,
+year,observationDate,{Data},,,,,,
+poverty_rate,unit,Percent,scalingFactor,100,value,{Number},variableMeasured,dcs:Count_Person_BelowPovertyLevelInThePast12Months_AsFractionOf_Count_Person
diff --git a/statvar_imports/commerce_eda_poverty/process_poverty.py b/statvar_imports/commerce_eda_poverty/process_poverty.py
new file mode 100644
index 0000000000..fe74345c11
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/process_poverty.py
@@ -0,0 +1,387 @@
+# Copyright 2026 Google LLC
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+# https://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+
+"""Preprocessing script for Commerce EDA Persistent Poverty Counties dataset.
+
+This script ingests the raw Persistent Poverty Counties dataset (downloaded
+from U.S. Economic Development Administration (EDA) / Department of Commerce),
+standardizes geographic identifiers (5-digit county FIPS codes across 50 US
+states, DC, and Puerto Rico, and 5-digit island territory codes), validates
+poverty percentage rates across 1990, 2000, 2020, and 2021, and generates the
+normalized cleaned CSV for stat_var_processor.py.
+"""
+
+import datetime
+import os
+import re
+import tempfile
+
+from absl import app, flags, logging
+import openpyxl
+import pandas as pd
+
+MODULE_DIR = os.path.dirname(os.path.abspath(__file__))
+
+DEFAULT_SOURCE_CSV = os.path.join(MODULE_DIR, "input_files", "Poverty.csv")
+DEFAULT_SOURCE_XLSX = os.path.join(MODULE_DIR, "input_files", "EDA_FY23_PPCs.xlsx")
+DEFAULT_RAW_CSV = os.path.join(MODULE_DIR, "output", "Poverty_original.csv")
+CLEANED_CSV = os.path.join(MODULE_DIR, "output", "Poverty_cleaned.csv")
+
+FLAGS = flags.FLAGS
+flags.DEFINE_string(
+ "source_path",
+ None,
+ "Path to source file (.csv or .xlsx). If not specified, automatically "
+ "resolves from input_files/Poverty.csv, input_files/EDA_FY23_PPCs.xlsx, "
+ "or output/Poverty_original.csv.",
+)
+flags.DEFINE_string("cleaned_csv_path", CLEANED_CSV, "Path to save cleaned output CSV.")
+flags.DEFINE_integer("min_county_count", 3000, "Minimum number of valid counties expected.")
+flags.DEFINE_integer(
+ "min_survey_year", 2020, "Minimum valid survey year for most recent estimate."
+)
+flags.DEFINE_integer(
+ "max_survey_year", None, "Maximum valid survey year (defaults to current year)."
+)
+
+# Valid 2-digit US State and Territory FIPS codes
+VALID_STATE_FIPS = {
+ # 50 States + DC
+ "01", "02", "04", "05", "06", "08", "09", "10", "11", "12", "13", "15",
+ "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27",
+ "28", "29", "30", "31", "32", "33", "34", "35", "36", "37", "38", "39",
+ "40", "41", "42", "44", "45", "46", "47", "48", "49", "50", "51", "53",
+ "54", "55", "56",
+ # Territories: American Samoa, Guam, Northern Mariana Islands, Puerto Rico, Virgin Islands
+ "60", "66", "69", "72", "78",
+}
+
+# Island territories where most recent estimate is from 2020 Decennial Census
+ISLAND_TERRITORY_FIPS = {"60", "66", "69", "78"}
+
+COLUMN_RENAME_MAP = {
+ "GEOID": "GEOID",
+ "1990 Decennial Census, % in Poverty": "poverty_rate_1990",
+ "2000 Decennial Census, % in Poverty": "poverty_rate_2000",
+ "Most Recent Estimate, % in Poverty*": "poverty_rate_recent",
+ "Most Recent Estimate, % in Poverty": "poverty_rate_recent",
+}
+
+
+def clean_geoid(val):
+ """Standardizes GEOIDs to 5 digits and validates against US state FIPS.
+
+ Rejects state summary entries of the form XX000.
+ """
+ if pd.isna(val):
+ return None
+ s = str(val).strip()
+ match = re.match(r"^(\d+)(?:\.0+)?$", s)
+ if not match:
+ return None
+ digits = match.group(1)
+ if len(digits) == 4:
+ digits = digits.zfill(5)
+ if len(digits) == 5 and digits[:2] in VALID_STATE_FIPS and digits[2:] != "000":
+ return digits
+ return None
+
+
+def _extract_dataframe_from_excel(excel_path):
+ """Extracts Underlying_Data sheet from an Excel workbook into a pandas DataFrame."""
+ wb = openpyxl.load_workbook(excel_path, data_only=True)
+ try:
+ sheet_names = wb.sheetnames
+ target_sheet = None
+ if "Underlying_Data" in sheet_names:
+ target_sheet = "Underlying_Data"
+ else:
+ for name in sheet_names:
+ if any(
+ k in name.lower()
+ for k in ["underlying", "poverty", "data", "ppc"]
+ ):
+ target_sheet = name
+ break
+ if not target_sheet:
+ target_sheet = sheet_names[0]
+ logging.warning(
+ "Worksheet 'Underlying_Data' not found in %s; falling back to '%s'.",
+ sheet_names,
+ target_sheet,
+ )
+ ws = wb[target_sheet]
+ rows = list(ws.iter_rows(values_only=True))
+ if not rows:
+ raise ValueError(f"Excel sheet '{target_sheet}' is empty.")
+
+ # Find row with GEOID header
+ header_idx = None
+ for idx, r in enumerate(rows[:10]):
+ row_str = [str(c).strip() for c in r if c is not None]
+ if any(c.upper() == "GEOID" for c in row_str):
+ header_idx = idx
+ break
+
+ if header_idx is None:
+ header_idx = 2 if len(rows) > 2 else 0
+
+ headers = [("" if c is None else str(c).strip()) for c in rows[header_idx]]
+ data_rows = []
+ for r in rows[header_idx + 1:]:
+ if not any(r):
+ continue
+ data_rows.append([("" if c is None else str(c).strip()) for c in r[:len(headers)]])
+
+ return pd.DataFrame(data_rows, columns=headers)
+ finally:
+ wb.close()
+
+
+def resolve_source_file_path(requested_path=None):
+ """Finds the raw source file, raising FileNotFoundError if missing."""
+ if requested_path:
+ if os.path.exists(requested_path) and os.path.getsize(requested_path) > 0:
+ return requested_path
+ raise FileNotFoundError(f"Specified source file not found or empty: {requested_path}")
+
+ candidates = [
+ DEFAULT_SOURCE_CSV,
+ DEFAULT_SOURCE_XLSX,
+ DEFAULT_RAW_CSV,
+ ]
+ for candidate in candidates:
+ if os.path.exists(candidate) and os.path.getsize(candidate) > 0:
+ return candidate
+
+ raise FileNotFoundError(
+ "No downloaded source file found. Please run download_poverty.py first to "
+ f"fetch the dataset, or specify --source_path. Checked: {candidates}"
+ )
+
+
+def preprocess_poverty(
+ src_path=DEFAULT_SOURCE_CSV,
+ dst_path=CLEANED_CSV,
+ min_county_count=3000,
+ min_survey_year=2020,
+ max_survey_year=None,
+):
+ """Preprocesses the raw Poverty dataset into cleaned format with normalized columns."""
+ logging.info("Preprocessing source Poverty dataset from %s...", src_path)
+ if max_survey_year is None:
+ max_survey_year = datetime.date.today().year
+ if not os.path.exists(src_path) or os.path.getsize(src_path) == 0:
+ logging.error("Source file does not exist or is empty: %s", src_path)
+ raise ValueError(f"Source file does not exist or is empty: {src_path}")
+
+ if src_path.lower().endswith((".xlsx", ".xls")):
+ df = _extract_dataframe_from_excel(src_path)
+ else:
+ # Detect header row index by scanning first lines for 'GEOID'
+ skip = 0
+ with open(src_path, "r", encoding="utf-8", errors="ignore") as f:
+ for idx in range(10):
+ line = f.readline()
+ if not line:
+ break
+ if "GEOID" in line.upper():
+ skip = idx
+ break
+ df = pd.read_csv(src_path, skiprows=skip, dtype=str)
+
+ if df.empty:
+ logging.error("Source dataset is empty: %s", src_path)
+ raise ValueError(f"Source dataset is empty: {src_path}")
+
+ # Strip column headers to avoid fragile whitespace issues
+ df.columns = df.columns.str.strip()
+
+ # Check if at least GEOID and the poverty columns are found
+ if "GEOID" not in df.columns:
+ logging.error("Missing required column 'GEOID' in source dataset")
+ raise ValueError("Missing required column 'GEOID' in source dataset")
+
+ # Rename columns to standard names (including any future Decennial Census columns)
+ rename_dict = {}
+ for col in df.columns:
+ clean_col = col.rstrip("*").strip()
+ matched = False
+ for k, v in COLUMN_RENAME_MAP.items():
+ if col == k or clean_col == k.rstrip("*").strip():
+ rename_dict[col] = v
+ matched = True
+ break
+ if not matched:
+ hist_match = re.match(
+ r"^(\d{4})\b.*%\s*in\s*Poverty", clean_col, flags=re.IGNORECASE
+ )
+ if hist_match:
+ rename_dict[col] = f"poverty_rate_{hist_match.group(1)}"
+
+ # Guard against silent year corruption if Data Source column is present
+ data_source_cols = [c for c in df.columns if "data source" in c.lower()]
+ df = df.rename(columns=rename_dict)
+
+ required_cols = [
+ "GEOID", "poverty_rate_1990", "poverty_rate_2000", "poverty_rate_recent"
+ ]
+ missing = [c for c in required_cols if c not in df.columns]
+ if missing:
+ logging.error("Missing required columns in source dataset: %s", missing)
+ raise ValueError(f"Missing required columns in source dataset: {missing}")
+
+ # Standardize and validate GEOIDs
+ df["GEOID"] = df["GEOID"].apply(clean_geoid)
+ df = df.dropna(subset=["GEOID"])
+
+ is_territory = df["GEOID"].str[:2].isin(ISLAND_TERRITORY_FIPS)
+ default_years = pd.Series("2021", index=df.index).where(~is_territory, "2020")
+
+ if data_source_cols:
+ ds_col = data_source_cols[0]
+ parsed_years = []
+ for idx, row in df.iterrows():
+ val = str(row.get(ds_col, "")).strip()
+ if val and val != "nan":
+ matched_years = re.findall(r"\b(19\d\d|20\d\d)\b", val)
+ if not matched_years:
+ error_msg = (
+ f"Unrecognized survey year in {ds_col} ('{val}') for "
+ f"GEOID {row['GEOID']}"
+ )
+ logging.error(error_msg)
+ raise ValueError(error_msg)
+ yr = int(matched_years[-1])
+ if not (min_survey_year <= yr <= max_survey_year):
+ error_msg = (
+ f"Unexpected survey year {yr} in {ds_col} for GEOID"
+ f" {row['GEOID']} (expected between {min_survey_year} and "
+ f"{max_survey_year})"
+ )
+ logging.error(error_msg)
+ raise ValueError(error_msg)
+ parsed_years.append(str(yr))
+ else:
+ parsed_years.append(default_years.loc[idx])
+ recent_years = pd.Series(parsed_years, index=df.index)
+ else:
+ recent_years = default_years
+
+ # Identify all historical year columns dynamically
+ historical_year_cols = sorted(
+ [
+ (re.match(r"^poverty_rate_(\d{4})$", c).group(1), c)
+ for c in df.columns
+ if re.match(r"^poverty_rate_(\d{4})$", c)
+ ],
+ key=lambda x: int(x[0]),
+ )
+
+ # Coerce and validate poverty values within [0.0, 100.0]
+ raw_poverty_cols = [c for _, c in historical_year_cols] + ["poverty_rate_recent"]
+ for col in raw_poverty_cols:
+ df[col] = pd.to_numeric(df[col].astype(str).str.strip(), errors="coerce")
+ invalid_mask = df[col].notna() & ((df[col] < 0.0) | (df[col] > 100.0))
+ if invalid_mask.any():
+ logging.warning(
+ "Found %d out-of-bounds values in %s; setting to NaN",
+ invalid_mask.sum(),
+ col,
+ )
+ df.loc[invalid_mask, col] = None
+
+ # Keep counties that have at least one valid poverty rate observation
+ df = df.dropna(subset=raw_poverty_cols, how="all")
+
+ # Verify sanity threshold on valid county count
+ if len(df) < min_county_count:
+ logging.error(
+ "Sanity check failed: Expected at least %d counties, but found %d.",
+ min_county_count,
+ len(df),
+ )
+ raise ValueError(
+ f"Sanity check failed: Expected at least {min_county_count} "
+ f"counties, but found {len(df)}."
+ )
+
+ # Build generic long-format (GEOID, year, poverty_rate) records
+ records = []
+ for idx, row in df.iterrows():
+ geoid = row["GEOID"]
+ row_obs = {}
+ for yr_str, col_name in historical_year_cols:
+ val = row[col_name]
+ if pd.notna(val):
+ row_obs[yr_str] = float(val)
+ recent_val = row["poverty_rate_recent"]
+ if pd.notna(recent_val):
+ row_obs[str(recent_years.loc[idx])] = float(recent_val)
+ for yr_str in sorted(row_obs.keys(), key=int):
+ records.append(
+ {
+ "GEOID": geoid,
+ "year": yr_str,
+ "poverty_rate": row_obs[yr_str],
+ }
+ )
+
+ df = pd.DataFrame(records, columns=["GEOID", "year", "poverty_rate"])
+
+ # Atomic write to destination file
+ dst_dir = os.path.dirname(os.path.abspath(dst_path))
+ os.makedirs(dst_dir, exist_ok=True)
+ temp_path = None
+ try:
+ with tempfile.NamedTemporaryFile(
+ "w",
+ dir=dst_dir,
+ delete=False,
+ suffix=".tmp",
+ encoding="utf-8",
+ newline="",
+ ) as tmp_file:
+ temp_path = tmp_file.name
+ df.to_csv(tmp_file, index=False)
+
+ os.replace(temp_path, dst_path)
+ temp_path = None
+ finally:
+ if temp_path and os.path.exists(temp_path):
+ try:
+ os.unlink(temp_path)
+ except OSError:
+ pass
+
+ logging.info("Poverty dataset cleaned and saved successfully to %s!", dst_path)
+ logging.info("Shape: %s", df.shape)
+ return dst_path
+
+
+def main(argv):
+ """Main entrypoint for preprocessing the poverty dataset."""
+ del argv # Unused
+ source_path = resolve_source_file_path(FLAGS.source_path)
+ preprocess_poverty(
+ src_path=source_path,
+ dst_path=FLAGS.cleaned_csv_path,
+ min_county_count=FLAGS.min_county_count,
+ min_survey_year=FLAGS.min_survey_year,
+ max_survey_year=FLAGS.max_survey_year,
+ )
+
+
+if __name__ == "__main__":
+ app.run(main)
diff --git a/statvar_imports/commerce_eda_poverty/process_poverty_test.py b/statvar_imports/commerce_eda_poverty/process_poverty_test.py
new file mode 100644
index 0000000000..6818c23b18
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/process_poverty_test.py
@@ -0,0 +1,373 @@
+# Copyright 2026 Google LLC
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+# https://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+
+"""Unit tests for process_poverty.py."""
+
+import os
+import sys
+import tempfile
+import unittest
+
+import openpyxl
+import pandas as pd
+
+MODULE_DIR = os.path.dirname(os.path.abspath(__file__))
+PROJECT_ROOT = os.path.abspath(os.path.join(MODULE_DIR, "..", ".."))
+sys.path.insert(0, PROJECT_ROOT)
+
+from statvar_imports.commerce_eda_poverty.process_poverty import (
+ clean_geoid,
+ preprocess_poverty,
+ resolve_source_file_path,
+)
+
+
+def _create_mock_excel_file(filepath, rows, description_row=True):
+ """Creates a .xlsx workbook mimicking the official EDA Persistent Poverty Counties workbook."""
+ wb = openpyxl.Workbook()
+ ws = wb.active
+ try:
+ ws.title = "Underlying_Data"
+ if description_row:
+ ws.append([
+ "Table. FY2023 Persistent Poverty County Status - as of Data Year 2021",
+ "", "", "", "", ""
+ ])
+ ws.append(["Identifing Information", "", "Census Bureau Data", "", "", ""])
+ ws.append([
+ "Name",
+ "GEOID",
+ "1990 Decennial Census, % in Poverty",
+ "2000 Decennial Census, % in Poverty",
+ "Most Recent Estimate, % in Poverty* ",
+ "Data Source―Most Recent Estimate",
+ ])
+ for r in rows:
+ ws.append(r)
+ wb.save(filepath)
+ finally:
+ wb.close()
+
+
+class TestProcessPoverty(unittest.TestCase):
+
+ def test_clean_geoid(self):
+ # 5-digit strings
+ self.assertEqual(clean_geoid("01001"), "01001")
+ self.assertEqual(clean_geoid("72143"), "72143")
+ self.assertEqual(clean_geoid("78010"), "78010")
+
+ # 4-digit zero-padding (e.g. Alabama 01005, Alaska 02090)
+ self.assertEqual(clean_geoid("1005"), "01005")
+ self.assertEqual(clean_geoid("2090"), "02090")
+
+ # Float strings
+ self.assertEqual(clean_geoid("1001.0"), "01001")
+ self.assertEqual(clean_geoid("01001.0"), "01001")
+ self.assertEqual(clean_geoid("1001.00"), "01001")
+
+ # Numeric floats and ints
+ self.assertEqual(clean_geoid(1001.0), "01001")
+ self.assertEqual(clean_geoid(2090), "02090")
+
+ # State summary FIPS ending in 000 must return None
+ self.assertIsNone(clean_geoid("01000"))
+ self.assertIsNone(clean_geoid("1000"))
+ self.assertIsNone(clean_geoid("72000"))
+
+ # Invalid cases returning None
+ self.assertIsNone(clean_geoid("1001.5"))
+ self.assertIsNone(clean_geoid("-1001.0"))
+ self.assertIsNone(clean_geoid(-1001.0))
+ self.assertIsNone(clean_geoid("abc"))
+ self.assertIsNone(clean_geoid(""))
+ self.assertIsNone(clean_geoid(None))
+ self.assertIsNone(clean_geoid("00100"))
+ self.assertIsNone(clean_geoid("0100"))
+ self.assertIsNone(clean_geoid("99001"))
+ self.assertIsNone(clean_geoid("123456"))
+
+ def test_resolve_source_file_path_specified(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ sample_file = os.path.join(tmpdir, "custom.csv")
+ with open(sample_file, "w") as f:
+ f.write("a,b\n1,2\n")
+ self.assertEqual(resolve_source_file_path(sample_file), sample_file)
+
+ def test_resolve_source_file_path_missing_raises(self):
+ with self.assertRaises(FileNotFoundError):
+ resolve_source_file_path("/nonexistent/file.csv")
+
+ def test_preprocess_poverty_missing_file(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ missing_path = os.path.join(tmpdir, "nonexistent.csv")
+ dst_path = os.path.join(tmpdir, "output.csv")
+ with self.assertRaises(ValueError):
+ preprocess_poverty(src_path=missing_path, dst_path=dst_path)
+
+ def test_preprocess_poverty_empty_file(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ empty_path = os.path.join(tmpdir, "empty.csv")
+ with open(empty_path, "w") as f:
+ pass
+ dst_path = os.path.join(tmpdir, "output.csv")
+ with self.assertRaises(ValueError):
+ preprocess_poverty(src_path=empty_path, dst_path=dst_path)
+
+ def test_preprocess_poverty_header_only_empty(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ header_only = os.path.join(tmpdir, "header_only.csv")
+ with open(header_only, "w") as f:
+ f.write(
+ "Header 1\nHeader 2\n"
+ "Name,GEOID,\"1990 Decennial Census, % in Poverty\","
+ "\"2000 Decennial Census, % in Poverty\","
+ "\"Most Recent Estimate, % in Poverty*\"\n"
+ )
+ dst_path = os.path.join(tmpdir, "output.csv")
+ with self.assertRaises(ValueError):
+ preprocess_poverty(src_path=header_only, dst_path=dst_path)
+
+ def test_preprocess_poverty_missing_columns(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ bad_csv = os.path.join(tmpdir, "bad.csv")
+ with open(bad_csv, "w") as f:
+ f.write("Line 1\nLine 2\nGEOID,OtherCol\n01001,10.0\n")
+ dst_path = os.path.join(tmpdir, "output.csv")
+ with self.assertRaises(ValueError):
+ preprocess_poverty(src_path=bad_csv, dst_path=dst_path)
+
+ def test_preprocess_poverty_unexpected_survey_year(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ for bad_source in ["SAIPE, 1999", "SAIPE, 2099"]:
+ future_csv = os.path.join(tmpdir, "future.csv")
+ with open(future_csv, "w") as f:
+ f.write(
+ "Header 1\nHeader 2\n"
+ "Name,GEOID,\"1990 Decennial Census, % in Poverty\","
+ "\"2000 Decennial Census, % in Poverty\","
+ "\"Most Recent Estimate, % in Poverty*\","
+ "\"Data Source―Most Recent Estimate\"\n"
+ f'"Autauga County, AL",01001,15.7,10.9,13.3,"{bad_source}"\n'
+ )
+ dst_path = os.path.join(tmpdir, "output.csv")
+ with self.assertRaises(ValueError) as ctx:
+ preprocess_poverty(
+ src_path=future_csv, dst_path=dst_path, min_county_count=1
+ )
+ self.assertIn("Unexpected survey year", str(ctx.exception))
+
+ malformed_csv = os.path.join(tmpdir, "malformed.csv")
+ with open(malformed_csv, "w") as f:
+ f.write(
+ "Header 1\nHeader 2\n"
+ "Name,GEOID,\"1990 Decennial Census, % in Poverty\","
+ "\"2000 Decennial Census, % in Poverty\","
+ "\"Most Recent Estimate, % in Poverty*\","
+ "\"Data Source―Most Recent Estimate\"\n"
+ '"Autauga County, AL",01001,15.7,10.9,13.3,"SAIPE FY23"\n'
+ )
+ with self.assertRaises(ValueError) as ctx:
+ preprocess_poverty(
+ src_path=malformed_csv, dst_path=dst_path, min_county_count=1
+ )
+ self.assertIn("Unrecognized survey year", str(ctx.exception))
+
+ def test_preprocess_poverty_valid_survey_year_range(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ sample_csv = os.path.join(tmpdir, "valid_year.csv")
+ with open(sample_csv, "w") as f:
+ f.write(
+ "Header 1\nHeader 2\n"
+ "Name,GEOID,\"1990 Decennial Census, % in Poverty\","
+ "\"2000 Decennial Census, % in Poverty\","
+ "\"Most Recent Estimate, % in Poverty*\","
+ "\"Data Source―Most Recent Estimate\"\n"
+ '"Autauga County, AL",01001,15.7,10.9,13.3,"SAIPE, 2023"\n'
+ )
+ dst_path = os.path.join(tmpdir, "output.csv")
+ preprocess_poverty(src_path=sample_csv, dst_path=dst_path, min_county_count=1)
+ self.assertTrue(os.path.exists(dst_path))
+ df = pd.read_csv(dst_path, dtype={"GEOID": str, "year": str})
+ self.assertEqual(list(df.columns), ["GEOID", "year", "poverty_rate"])
+ row_2023 = df[(df["GEOID"] == "01001") & (df["year"] == "2023")]
+ self.assertEqual(len(row_2023), 1)
+ self.assertEqual(row_2023["poverty_rate"].iloc[0], 13.3)
+
+ def test_preprocess_poverty_min_county_count_failure(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ fixture_path = os.path.join(MODULE_DIR, "test_data", "Poverty_input.csv")
+ dst_path = os.path.join(tmpdir, "output.csv")
+ with self.assertRaises(ValueError):
+ preprocess_poverty(src_path=fixture_path, dst_path=dst_path, min_county_count=3000)
+
+ def test_preprocess_poverty_with_test_data(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ fixture_path = os.path.join(MODULE_DIR, "test_data", "Poverty_input.csv")
+ expected_path = os.path.join(MODULE_DIR, "test_data", "Poverty_expected_output.csv")
+ actual_csv = os.path.join(tmpdir, "Poverty_cleaned.csv")
+
+ preprocess_poverty(src_path=fixture_path, dst_path=actual_csv, min_county_count=100)
+
+ self.assertTrue(os.path.exists(actual_csv))
+ df_actual = pd.read_csv(actual_csv, dtype={"GEOID": str, "year": str})
+ df_expected = pd.read_csv(expected_path, dtype={"GEOID": str, "year": str})
+ pd.testing.assert_frame_equal(df_actual, df_expected)
+
+ # 191 valid counties/territories across 1990, 2000, and 2020/2021
+ self.assertEqual(df_actual["GEOID"].nunique(), 191)
+ self.assertEqual(list(df_actual.columns), ["GEOID", "year", "poverty_rate"])
+
+ # Verify 11 island territories (AS, GU, MP, VI) map recent estimate to year 2020
+ territory_rows = df_actual[df_actual["GEOID"].str[:2].isin({"60", "66", "69", "78"})]
+ self.assertEqual(territory_rows["GEOID"].nunique(), 11)
+ self.assertIn("2020", set(territory_rows["year"]))
+ self.assertNotIn("2021", set(territory_rows["year"]))
+
+ # Verify states and PR (180 counties) map recent estimate to year 2021
+ state_rows = df_actual[~df_actual["GEOID"].str[:2].isin({"60", "66", "69", "78"})]
+ self.assertEqual(state_rows["GEOID"].nunique(), 180)
+ self.assertIn("2021", set(state_rows["year"]))
+ self.assertNotIn("2020", set(state_rows["year"]))
+
+ def test_preprocess_poverty_edge_cases(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ raw_csv = os.path.join(tmpdir, "raw.csv")
+ actual_csv = os.path.join(tmpdir, "cleaned.csv")
+
+ raw_content = (
+ "Header 1\n"
+ "Header 2\n"
+ "Name,GEOID,\"1990 Decennial Census, % in Poverty\","
+ "\"2000 Decennial Census, % in Poverty\","
+ "\"Most Recent Estimate, % in Poverty* \"\n"
+ '"Autauga County, AL",01001,15.7,10.9,13.3\n'
+ '"Alabama State Summary",01000,18.0,16.0,15.0\n'
+ '"Yukon-Koyukuk, AK",2090,7.6,7.8,9.6\n'
+ '"Eastern District, AS",60010,25.0,28.0,30.0\n'
+ '"Barbour County, AL",01005.0,25.2,26.8,29.0\n'
+ '"Bibb County, AL",1007.0,21.2,20.6,24.9\n'
+ '"Blount County, AL",01009,-5.0,150.0,14.5\n'
+ '"Bullock County, AL",01011,-10.0,120.0,999.0\n'
+ '"Invalid 1",99001,15.0,15.0,15.0\n'
+ '"Invalid 2",abc,10.0,10.0,10.0\n'
+ '"Invalid 3",0100,5.0,4.2,3.1\n'
+ '"Source footnote",,,,\n'
+ )
+ with open(raw_csv, "w") as f:
+ f.write(raw_content)
+
+ preprocess_poverty(src_path=raw_csv, dst_path=actual_csv, min_county_count=1)
+ df_actual = pd.read_csv(actual_csv, dtype={"GEOID": str, "year": str})
+
+ expected_data = {
+ "GEOID": [
+ "01001", "01001", "01001",
+ "02090", "02090", "02090",
+ "60010", "60010", "60010",
+ "01005", "01005", "01005",
+ "01007", "01007", "01007",
+ "01009",
+ ],
+ "year": [
+ "1990", "2000", "2021",
+ "1990", "2000", "2021",
+ "1990", "2000", "2020",
+ "1990", "2000", "2021",
+ "1990", "2000", "2021",
+ "2021",
+ ],
+ "poverty_rate": [
+ 15.7, 10.9, 13.3,
+ 7.6, 7.8, 9.6,
+ 25.0, 28.0, 30.0,
+ 25.2, 26.8, 29.0,
+ 21.2, 20.6, 24.9,
+ 14.5,
+ ],
+ }
+ df_expected = pd.DataFrame(expected_data)
+ pd.testing.assert_frame_equal(df_actual, df_expected)
+
+ def test_preprocess_poverty_from_excel(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ excel_path = os.path.join(tmpdir, "mock.xlsx")
+ cleaned_csv = os.path.join(tmpdir, "cleaned.csv")
+
+ mock_rows = [
+ ["Autauga County, AL", "01001", 15.7, 10.9, 13.3, "SAIPE, 2021"],
+ ["Eastern District, AS", "60010", 25.0, 28.0, 30.0, "Decennial Census, 2020"],
+ ]
+ _create_mock_excel_file(excel_path, mock_rows)
+
+ preprocess_poverty(src_path=excel_path, dst_path=cleaned_csv, min_county_count=1)
+ self.assertTrue(os.path.exists(cleaned_csv))
+ df = pd.read_csv(cleaned_csv, dtype={"GEOID": str, "year": str})
+ self.assertEqual(len(df), 6)
+ self.assertEqual(
+ df.loc[(df["GEOID"] == "01001") & (df["year"] == "2021"), "poverty_rate"].iloc[0],
+ 13.3,
+ )
+ self.assertEqual(
+ df.loc[(df["GEOID"] == "60010") & (df["year"] == "2020"), "poverty_rate"].iloc[0],
+ 30.0,
+ )
+
+
+ def test_preprocess_poverty_header_detection_with_1990_title_row(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ csv_path = os.path.join(tmpdir, "title_with_1990.csv")
+ cleaned_csv = os.path.join(tmpdir, "cleaned.csv")
+ headers = (
+ 'Name,GEOID,"1990 Decennial Census, % in Poverty",'
+ '"2000 Decennial Census, % in Poverty",'
+ '"Most Recent Estimate, % in Poverty*",'
+ 'Data Source―Most Recent Estimate\n'
+ )
+ with open(csv_path, "w", encoding="utf-8") as f:
+ f.write("Table 1. EDA PPC Report (1990-2021) Summary\n")
+ f.write(headers)
+ f.write('Autauga County, AL,01001,15.7,10.9,13.3,"SAIPE, 2021"\n')
+
+ preprocess_poverty(src_path=csv_path, dst_path=cleaned_csv, min_county_count=1)
+ self.assertTrue(os.path.exists(cleaned_csv))
+ df = pd.read_csv(cleaned_csv, dtype={"GEOID": str, "year": str})
+ self.assertEqual(len(df), 3)
+ self.assertEqual(df["GEOID"].iloc[0], "01001")
+
+ def test_preprocess_poverty_survey_year_with_footnotes_and_citations(self):
+ with tempfile.TemporaryDirectory() as tmpdir:
+ csv_path = os.path.join(tmpdir, "footnotes.csv")
+ cleaned_csv = os.path.join(tmpdir, "cleaned.csv")
+ headers = (
+ 'Name,GEOID,"1990 Decennial Census, % in Poverty",'
+ '"2000 Decennial Census, % in Poverty",'
+ '"Most Recent Estimate, % in Poverty*",'
+ 'Data Source―Most Recent Estimate\n'
+ )
+ with open(csv_path, "w", encoding="utf-8") as f:
+ f.write(headers)
+ f.write('Autauga County, AL,01001,15.7,10.9,13.3,"SAIPE, 2021*"\n')
+ f.write('Barbour County, AL,01005,25.2,26.8,29.0,"SAIPE, 2021 [1]"\n')
+
+ preprocess_poverty(src_path=csv_path, dst_path=cleaned_csv, min_county_count=1)
+ self.assertTrue(os.path.exists(cleaned_csv))
+ df = pd.read_csv(cleaned_csv, dtype={"GEOID": str, "year": str})
+ self.assertEqual(len(df), 6)
+ recent_years = set(df.loc[df["year"] != "1990"].loc[df["year"] != "2000", "year"])
+ self.assertEqual(recent_years, {"2021"})
+
+
+if __name__ == "__main__":
+ unittest.main()
diff --git a/statvar_imports/commerce_eda_poverty/test_data/Poverty_expected_output.csv b/statvar_imports/commerce_eda_poverty/test_data/Poverty_expected_output.csv
new file mode 100644
index 0000000000..4546bafb05
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/test_data/Poverty_expected_output.csv
@@ -0,0 +1,569 @@
+GEOID,year,poverty_rate
+01001,1990,15.7
+01001,2000,10.9
+01001,2021,13.3
+01003,1990,14.3
+01003,2000,10.1
+01003,2021,12.5
+01005,1990,25.2
+01005,2000,26.8
+01005,2021,29.0
+01007,1990,21.2
+01007,2000,20.6
+01007,2021,24.9
+01009,1990,15.3
+01009,2000,11.7
+01009,2021,14.5
+01011,1990,36.5
+01011,2000,33.5
+01011,2021,39.1
+01013,1990,31.5
+01013,2000,24.6
+01013,2021,27.2
+01015,1990,15.7
+01015,2000,16.1
+01015,2021,21.8
+01017,1990,18.8
+01017,2000,17.0
+01017,2021,23.7
+01019,1990,17.6
+01019,2000,15.6
+01019,2021,21.6
+01021,1990,17.1
+01021,2000,15.7
+01021,2021,18.7
+01023,1990,30.2
+01023,2000,24.5
+01023,2021,28.1
+01025,1990,25.9
+01025,2000,22.6
+01025,2021,23.7
+01027,1990,17.4
+01027,2000,17.1
+01027,2021,21.8
+01029,1990,15.3
+01029,2000,13.9
+01029,2021,18.4
+01031,1990,15.5
+01031,2000,14.7
+01031,2021,18.1
+01033,1990,14.6
+01033,2000,14.0
+01033,2021,19.3
+01035,1990,29.7
+01035,2000,26.6
+01035,2021,28.0
+01037,1990,18.2
+01037,2000,14.9
+01037,2021,21.5
+01039,1990,22.0
+01039,2000,18.4
+01039,2021,23.1
+01041,1990,24.3
+01041,2000,22.1
+01041,2021,21.7
+01043,1990,15.3
+01043,2000,13.0
+01043,2021,16.0
+01045,1990,14.8
+01045,2000,15.1
+01045,2021,18.7
+01047,1990,36.2
+01047,2000,31.1
+01047,2021,35.5
+01049,1990,17.4
+01049,2000,15.4
+01049,2021,22.2
+01051,1990,14.5
+01051,2000,10.2
+01051,2021,14.3
+01053,1990,28.1
+01053,2000,20.9
+01053,2021,28.0
+01055,1990,16.5
+01055,2000,15.7
+01055,2021,20.1
+01057,1990,20.3
+01057,2000,17.3
+01057,2021,23.4
+01059,1990,20.7
+01059,2000,18.9
+01059,2021,22.6
+01061,1990,19.5
+01061,2000,19.6
+01061,2021,25.0
+01063,1990,45.6
+01063,2000,34.3
+01063,2021,40.1
+01065,1990,35.6
+01065,2000,26.9
+01065,2021,27.4
+01067,1990,17.4
+01067,2000,19.1
+01067,2021,19.3
+01069,1990,16.5
+01069,2000,15.0
+01069,2021,20.9
+01071,1990,16.6
+01071,2000,13.7
+01071,2021,21.6
+01073,1990,16.0
+01073,2000,14.8
+01073,2021,18.3
+01075,1990,18.0
+01075,2000,16.1
+01075,2021,21.1
+01077,1990,14.9
+01077,2000,14.4
+01077,2021,18.9
+01079,1990,19.8
+01079,2000,15.3
+01079,2021,17.5
+01081,1990,24.9
+01081,2000,21.8
+01081,2021,20.2
+01083,1990,14.0
+01083,2000,12.3
+01083,2021,11.8
+01085,1990,38.6
+01085,2000,31.4
+01085,2021,34.2
+01087,1990,34.5
+01087,2000,32.8
+01087,2021,33.6
+01089,1990,10.9
+01089,2000,10.5
+01089,2021,11.7
+01091,1990,30.0
+01091,2000,25.9
+01091,2021,29.3
+01093,1990,19.1
+01093,2000,15.6
+01093,2021,21.9
+01095,1990,17.2
+01095,2000,14.7
+01095,2021,18.3
+01097,1990,21.4
+01097,2000,18.5
+01097,2021,20.2
+01099,1990,22.7
+01099,2000,21.3
+01099,2021,26.9
+01101,1990,17.9
+01101,2000,17.3
+01101,2021,24.1
+01103,1990,12.0
+01103,2000,12.3
+01103,2021,15.8
+01105,1990,42.6
+01105,2000,35.4
+01105,2021,41.6
+01107,1990,28.9
+01107,2000,24.9
+01107,2021,26.0
+01109,1990,27.2
+01109,2000,23.1
+01109,2021,27.7
+01111,1990,18.9
+01111,2000,17.0
+01111,2021,25.1
+01113,1990,20.4
+01113,2000,19.9
+01113,2021,25.3
+01115,1990,14.8
+01115,2000,12.1
+01115,2021,15.2
+01117,1990,9.2
+01117,2000,6.3
+01117,2021,9.0
+01119,1990,39.7
+01119,2000,38.7
+01119,2021,41.9
+01121,1990,20.2
+01121,2000,17.6
+01121,2021,22.3
+01123,1990,16.0
+01123,2000,16.6
+01123,2021,21.7
+01125,1990,20.1
+01125,2000,17.0
+01125,2021,17.0
+01127,1990,17.3
+01127,2000,16.5
+01127,2021,23.2
+01129,1990,24.8
+01129,2000,18.5
+01129,2021,24.5
+01131,1990,45.2
+01131,2000,39.9
+01131,2021,39.8
+01133,1990,19.8
+01133,2000,17.1
+01133,2021,22.6
+02013,1990,11.9
+02013,2000,21.8
+02013,2021,23.4
+02016,1990,9.0
+02016,2000,11.9
+02016,2021,12.8
+02020,1990,7.1
+02020,2000,7.3
+02020,2021,10.7
+02050,1990,30.0
+02050,2000,20.6
+02050,2021,29.1
+02060,1990,5.1
+02060,2000,9.5
+02060,2021,14.2
+02063,1990,8.9
+02063,2000,9.8
+02063,2021,9.7
+02066,1990,8.9
+02066,2000,9.8
+02066,2021,17.7
+02068,2000,7.9
+02068,2021,9.3
+02070,1990,24.6
+02070,2000,21.4
+02070,2021,26.8
+02090,1990,7.6
+02090,2000,7.8
+02090,2021,9.6
+02100,1990,9.2
+02100,2000,10.7
+02100,2021,13.5
+02105,2021,21.2
+02110,1990,5.6
+02110,2000,6.0
+02110,2021,9.3
+02122,1990,7.7
+02122,2000,10.0
+02122,2021,13.6
+02130,1990,4.2
+02130,2000,6.5
+02130,2021,13.4
+02150,1990,5.5
+02150,2000,6.6
+02150,2021,10.4
+02158,1990,31.0
+02158,2000,26.2
+02158,2021,36.7
+02164,1990,20.0
+02164,2000,18.9
+02164,2021,25.5
+02170,1990,9.4
+02170,2000,11.0
+02170,2021,11.9
+02180,1990,22.4
+02180,2000,17.4
+02180,2021,23.9
+02185,1990,8.7
+02185,2000,9.1
+02185,2021,15.6
+02188,1990,18.5
+02188,2000,17.4
+02188,2021,25.5
+02195,1990,5.7
+02195,2000,7.9
+02195,2021,9.9
+02198,1990,9.1
+02198,2000,12.1
+02198,2021,18.8
+02220,1990,4.8
+02220,2000,7.8
+02220,2021,10.6
+02230,2021,7.6
+02240,1990,14.2
+02240,2000,18.9
+02240,2021,16.0
+02275,1990,5.7
+02275,2000,7.9
+02275,2021,14.1
+02282,1990,8.9
+02282,2000,13.5
+02282,2021,15.3
+02290,1990,26.0
+02290,2000,23.8
+02290,2021,29.1
+04001,1990,47.1
+04001,2000,37.8
+04001,2021,31.8
+04003,1990,20.3
+04003,2000,17.7
+04003,2021,19.4
+04005,1990,23.1
+04005,2000,18.2
+04005,2021,19.1
+04007,1990,18.3
+04007,2000,17.4
+04007,2021,20.4
+04009,1990,26.7
+04009,2000,23.0
+04009,2021,23.9
+04011,1990,12.6
+04011,2000,9.9
+04011,2021,11.8
+04012,1990,28.2
+04012,2000,19.6
+04012,2021,24.8
+04013,1990,12.3
+04013,2000,11.7
+04013,2021,11.7
+04015,1990,14.2
+04015,2000,13.9
+04015,2021,20.0
+04017,1990,34.7
+04017,2000,29.5
+04017,2021,28.0
+04019,1990,17.2
+04019,2000,14.7
+04019,2021,15.7
+04021,1990,23.6
+04021,2000,16.9
+04021,2021,12.2
+04023,1990,26.4
+04023,2000,24.5
+04023,2021,23.9
+04025,1990,13.6
+04025,2000,11.9
+04025,2021,14.2
+04027,1990,19.9
+04027,2000,19.2
+04027,2021,19.7
+05001,1990,20.4
+05001,2000,17.8
+05001,2021,21.8
+05003,1990,20.9
+05003,2000,17.5
+05003,2021,25.2
+05005,1990,16.3
+05005,2000,11.1
+05005,2021,15.7
+05007,1990,9.6
+05007,2000,10.1
+05007,2021,9.2
+05009,1990,13.9
+05009,2000,14.8
+05009,2021,16.0
+05011,1990,24.9
+05011,2000,26.3
+05011,2021,24.3
+05013,1990,15.6
+05013,2000,16.5
+05013,2021,18.2
+05015,1990,15.2
+05015,2000,15.5
+05015,2021,20.0
+05017,1990,40.4
+05017,2000,28.6
+05017,2021,33.8
+05019,1990,23.9
+05019,2000,19.1
+05019,2021,24.2
+05021,1990,21.2
+05021,2000,17.5
+05021,2021,21.9
+05023,1990,17.3
+05023,2000,13.1
+05023,2021,19.5
+05025,1990,19.0
+05025,2000,15.2
+05025,2021,17.5
+05027,1990,24.4
+05027,2000,21.1
+05027,2021,27.6
+05029,1990,16.5
+05029,2000,16.1
+05029,2021,20.3
+05031,1990,17.0
+05031,2000,15.4
+05031,2021,20.6
+05033,1990,16.3
+05033,2000,14.2
+05033,2021,17.6
+05035,1990,27.1
+05035,2000,25.3
+05035,2021,27.0
+05037,1990,25.4
+05037,2000,19.9
+05037,2021,20.6
+05039,1990,22.3
+05039,2000,18.9
+05039,2021,25.0
+05041,1990,34.0
+05041,2000,28.9
+05041,2021,32.8
+05043,1990,24.2
+05043,2000,18.2
+05043,2021,20.7
+05045,1990,13.8
+05045,2000,12.5
+05045,2021,15.4
+05047,1990,20.4
+05047,2000,15.2
+05047,2021,20.8
+05049,1990,26.3
+05049,2000,16.3
+05049,2021,22.0
+05051,1990,18.0
+05051,2000,14.6
+05051,2021,17.0
+05053,1990,14.9
+05053,2000,10.2
+05053,2021,13.8
+05055,1990,17.9
+05055,2000,13.3
+05055,2021,15.7
+05057,1990,22.7
+05057,2000,20.3
+05057,2021,25.4
+05059,1990,18.6
+05059,2000,14.0
+05059,2021,21.0
+05061,1990,18.6
+05061,2000,15.5
+05061,2021,21.6
+05063,1990,17.1
+05063,2000,13.0
+05063,2021,21.4
+05065,1990,21.1
+05065,2000,17.2
+05065,2021,22.4
+05067,1990,26.6
+05067,2000,17.4
+05067,2021,27.8
+05069,1990,23.9
+05069,2000,20.5
+05069,2021,24.7
+05071,1990,20.1
+05071,2000,16.4
+05071,2021,21.2
+05073,1990,34.7
+05073,2000,23.2
+05073,2021,28.6
+05075,1990,25.0
+05075,2000,18.4
+05075,2021,24.2
+05077,1990,47.3
+05077,2000,29.9
+05077,2021,43.6
+05079,1990,26.2
+05079,2000,19.5
+05079,2021,28.3
+05081,1990,19.3
+05081,2000,15.4
+05081,2021,19.6
+05083,1990,19.3
+05083,2000,15.4
+05083,2021,18.8
+05085,1990,14.9
+05085,2000,10.5
+05085,2021,13.0
+05087,1990,20.1
+05087,2000,18.6
+05087,2021,18.9
+05089,1990,18.9
+05089,2000,15.2
+05089,2021,21.3
+05091,1990,22.4
+05091,2000,19.3
+05091,2021,23.4
+05093,1990,26.2
+05093,2000,23.0
+05093,2021,28.1
+05095,1990,35.9
+05095,2000,27.5
+05095,2021,31.3
+05097,1990,23.8
+05097,2000,17.0
+05097,2021,23.9
+05099,1990,20.3
+05099,2000,22.8
+05099,2021,25.9
+05101,1990,29.6
+05101,2000,20.4
+05101,2021,23.7
+05103,1990,21.2
+05103,2000,19.5
+05103,2021,24.4
+05105,1990,20.3
+05105,2000,14.0
+05105,2021,18.6
+05107,1990,43.0
+05107,2000,32.7
+05107,2021,42.1
+05109,1990,17.9
+05109,2000,16.8
+05109,2021,22.2
+05111,1990,25.6
+05111,2000,21.2
+05111,2021,25.9
+05113,1990,18.5
+05113,2000,18.2
+05113,2021,23.4
+05115,1990,15.4
+05115,2000,15.2
+05115,2021,20.7
+05117,1990,22.7
+05117,2000,15.5
+05117,2021,19.2
+60010,1990,56.0
+60010,2000,58.6
+60010,2020,52.2
+60020,1990,79.8
+60020,2000,65.2
+60020,2020,60.5
+60050,1990,59.3
+60050,2000,62.6
+60050,2020,55.7
+66010,1990,15.0
+66010,2000,23.0
+66010,2020,20.2
+69085,1990,72.2
+69085,2000,83.3
+69085,2020,42.9
+69100,1990,51.7
+69100,2000,34.2
+69100,2020,36.4
+69110,1990,51.4
+69110,2000,46.9
+69110,2020,38.2
+69120,1990,51.3
+69120,2000,41.2
+69120,2020,35.5
+72001,1990,81.5
+72001,2000,65.4
+72001,2021,71.3
+72003,1990,69.7
+72003,2000,59.3
+72003,2021,48.4
+72005,1990,65.3
+72005,2000,55.0
+72005,2021,52.6
+72007,1990,62.6
+72007,2000,51.7
+72007,2021,47.1
+72009,1990,60.6
+72009,2000,51.8
+72009,2021,48.3
+72011,1990,61.7
+72011,2000,51.6
+72011,2021,49.1
+72013,1990,63.9
+72013,2000,50.9
+72013,2021,49.4
+72015,1990,70.7
+72015,2000,55.1
+72015,2021,66.3
+72017,1990,64.4
+72017,2000,56.0
+72017,2021,50.4
+78010,1990,33.7
+78010,2000,38.7
+78010,2020,24.9
+78020,1990,15.0
+78020,2000,18.5
+78020,2020,18.9
+78030,1990,21.3
+78030,2000,27.2
+78030,2020,21.2
diff --git a/statvar_imports/commerce_eda_poverty/test_data/Poverty_input.csv b/statvar_imports/commerce_eda_poverty/test_data/Poverty_input.csv
new file mode 100644
index 0000000000..bee739ef7c
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/test_data/Poverty_input.csv
@@ -0,0 +1,200 @@
+Table. FY2023 Persistent Poverty County Status - as of Data Year 2021,,,,,,,,,
+Identifing Information,,Census Bureau Data,,,,FY23 Persistent Poverty,Census GEO PPC Code,Notes,
+Name,GEOID,"1990 Decennial Census, % in Poverty","2000 Decennial Census, % in Poverty","Most Recent Estimate, % in Poverty* ",Data Source―Most Recent Estimate,,,,
+"Autauga County, AL",01001,15.7,10.9,13.3,"SAIPE, 2021",No,1,,
+"Baldwin County, AL",01003,14.3,10.1,12.5,"SAIPE, 2021",No,1,,
+"Barbour County, AL",01005,25.2,26.8,29,"SAIPE, 2021",Yes,2,,
+"Bibb County, AL",01007,21.2,20.6,24.9,"SAIPE, 2021",Yes,2,,
+"Blount County, AL",01009,15.3,11.7,14.5,"SAIPE, 2021",No,1,,
+"Bullock County, AL",01011,36.5,33.5,39.1,"SAIPE, 2021",Yes,2,,
+"Butler County, AL",01013,31.5,24.6,27.2,"SAIPE, 2021",Yes,2,,
+"Calhoun County, AL",01015,15.7,16.1,21.8,"SAIPE, 2021",No,1,,
+"Chambers County, AL",01017,18.8,17,23.7,"SAIPE, 2021",No,1,,
+"Cherokee County, AL",01019,17.6,15.6,21.6,"SAIPE, 2021",No,1,,
+"Chilton County, AL",01021,17.1,15.7,18.7,"SAIPE, 2021",No,1,,
+"Choctaw County, AL",01023,30.2,24.5,28.1,"SAIPE, 2021",Yes,2,,
+"Clarke County, AL",01025,25.9,22.6,23.7,"SAIPE, 2021",Yes,2,,
+"Clay County, AL",01027,17.4,17.1,21.8,"SAIPE, 2021",No,1,,
+"Cleburne County, AL",01029,15.3,13.9,18.4,"SAIPE, 2021",No,1,,
+"Coffee County, AL",01031,15.5,14.7,18.1,"SAIPE, 2021",No,1,,
+"Colbert County, AL",01033,14.6,14,19.3,"SAIPE, 2021",No,1,,
+"Conecuh County, AL",01035,29.7,26.6,28,"SAIPE, 2021",Yes,2,,
+"Coosa County, AL",01037,18.2,14.9,21.5,"SAIPE, 2021",No,1,,
+"Covington County, AL",01039,22,18.4,23.1,"SAIPE, 2021",No,1,,
+"Crenshaw County, AL",01041,24.3,22.1,21.7,"SAIPE, 2021",Yes,2,,
+"Cullman County, AL",01043,15.3,13,16,"SAIPE, 2021",No,1,,
+"Dale County, AL",01045,14.8,15.1,18.7,"SAIPE, 2021",No,1,,
+"Dallas County, AL",01047,36.2,31.1,35.5,"SAIPE, 2021",Yes,2,,
+"DeKalb County, AL",01049,17.4,15.4,22.2,"SAIPE, 2021",No,1,,
+"Elmore County, AL",01051,14.5,10.2,14.3,"SAIPE, 2021",No,1,,
+"Escambia County, AL",01053,28.1,20.9,28,"SAIPE, 2021",Yes,2,,
+"Etowah County, AL",01055,16.5,15.7,20.1,"SAIPE, 2021",No,1,,
+"Fayette County, AL",01057,20.3,17.3,23.4,"SAIPE, 2021",No,1,,
+"Franklin County, AL",01059,20.7,18.9,22.6,"SAIPE, 2021",No,1,,
+"Geneva County, AL",01061,19.5,19.6,25,"SAIPE, 2021",Yes,2,,
+"Greene County, AL",01063,45.6,34.3,40.1,"SAIPE, 2021",Yes,2,,
+"Hale County, AL",01065,35.6,26.9,27.4,"SAIPE, 2021",Yes,2,,
+"Henry County, AL",01067,17.4,19.1,19.3,"SAIPE, 2021",No,1,,
+"Houston County, AL",01069,16.5,15,20.9,"SAIPE, 2021",No,1,,
+"Jackson County, AL",01071,16.6,13.7,21.6,"SAIPE, 2021",No,1,,
+"Jefferson County, AL",01073,16,14.8,18.3,"SAIPE, 2021",No,1,,
+"Lamar County, AL",01075,18,16.1,21.1,"SAIPE, 2021",No,1,,
+"Lauderdale County, AL",01077,14.9,14.4,18.9,"SAIPE, 2021",No,1,,
+"Lawrence County, AL",01079,19.8,15.3,17.5,"SAIPE, 2021",No,1,,
+"Lee County, AL",01081,24.9,21.8,20.2,"SAIPE, 2021",Yes,2,,
+"Limestone County, AL",01083,14,12.3,11.8,"SAIPE, 2021",No,1,,
+"Lowndes County, AL",01085,38.6,31.4,34.2,"SAIPE, 2021",Yes,2,,
+"Macon County, AL",01087,34.5,32.8,33.6,"SAIPE, 2021",Yes,2,,
+"Madison County, AL",01089,10.9,10.5,11.7,"SAIPE, 2021",No,1,,
+"Marengo County, AL",01091,30,25.9,29.3,"SAIPE, 2021",Yes,2,,
+"Marion County, AL",01093,19.1,15.6,21.9,"SAIPE, 2021",No,1,,
+"Marshall County, AL",01095,17.2,14.7,18.3,"SAIPE, 2021",No,1,,
+"Mobile County, AL",01097,21.4,18.5,20.2,"SAIPE, 2021",No,1,,
+"Monroe County, AL",01099,22.7,21.3,26.9,"SAIPE, 2021",Yes,2,,
+"Montgomery County, AL",01101,17.9,17.3,24.1,"SAIPE, 2021",No,1,,
+"Morgan County, AL",01103,12,12.3,15.8,"SAIPE, 2021",No,1,,
+"Perry County, AL",01105,42.6,35.4,41.6,"SAIPE, 2021",Yes,2,,
+"Pickens County, AL",01107,28.9,24.9,26,"SAIPE, 2021",Yes,2,,
+"Pike County, AL",01109,27.2,23.1,27.7,"SAIPE, 2021",Yes,2,,
+"Randolph County, AL",01111,18.9,17,25.1,"SAIPE, 2021",No,1,,
+"Russell County, AL",01113,20.4,19.9,25.3,"SAIPE, 2021",Yes,2,,
+"St. Clair County, AL",01115,14.8,12.1,15.2,"SAIPE, 2021",No,1,,
+"Shelby County, AL",01117,9.2,6.3,9,"SAIPE, 2021",No,1,,
+"Sumter County, AL",01119,39.7,38.7,41.9,"SAIPE, 2021",Yes,2,,
+"Talladega County, AL",01121,20.2,17.6,22.3,"SAIPE, 2021",No,1,,
+"Tallapoosa County, AL",01123,16,16.6,21.7,"SAIPE, 2021",No,1,,
+"Tuscaloosa County, AL",01125,20.1,17,17,"SAIPE, 2021",No,1,,
+"Walker County, AL",01127,17.3,16.5,23.2,"SAIPE, 2021",No,1,,
+"Washington County, AL",01129,24.8,18.5,24.5,"SAIPE, 2021",No,1,,
+"Wilcox County, AL",01131,45.2,39.9,39.8,"SAIPE, 2021",Yes,2,,
+"Winston County, AL",01133,19.8,17.1,22.6,"SAIPE, 2021",No,1,,
+"Aleutians East Borough, AK",02013,11.9,21.8,23.4,"SAIPE, 2021",No,1,,
+"Aleutians West Census Area, AK",02016,9,11.9,12.8,"SAIPE, 2021",No,1,,
+"Anchorage Borough, AK",02020,7.1,7.3,10.7,"SAIPE, 2021",No,1,,
+"Bethel Census Area, AK",02050,30,20.6,29.1,"SAIPE, 2021",Yes,2,,
+"Bristol Bay Borough, AK",02060,5.1,9.5,14.2,"SAIPE, 2021",No,1,,
+"Chugach Census Area, AK9",02063,8.9,9.8,9.7,"SAIPE, 2021",No,1,Came into existence in 2020. Used to be Valdez−Cordova Census Area (02261). Percent in poverty hand entered based on prior geography.,
+"Copper River Census Area, AK9",02066,8.9,9.8,17.7,"SAIPE, 2021",No,1,Came into existence in 2020. Used to be Valdez−Cordova Census Area (02261). Percent in poverty hand entered based on prior geography.,
+"Denali Borough, AK",02068,NA,7.9,9.3,"SAIPE, 2021",NA,0,Did not exist prior to 2000.,
+"Dillingham Census Area, AK",02070,24.6,21.4,26.8,"SAIPE, 2021",Yes,2,,
+"Fairbanks North Star Borough, AK",02090,7.6,7.8,9.6,"SAIPE, 2021",No,1,,
+"Haines Borough, AK",02100,9.2,10.7,13.5,"SAIPE, 2021",No,1,,
+"Hoonah-Angoon Census Area, AK5",02105,NA,NA,21.2,"SAIPE, 2021",NA,0,Came into existence in 2008. Used to be Skagway-Hoonah-Angoon Census Area (02232).,
+"Juneau Borough, AK",02110,5.6,6,9.3,"SAIPE, 2021",No,1,,
+"Kenai Peninsula Borough, AK",02122,7.7,10,13.6,"SAIPE, 2021",No,1,,
+"Ketchikan Gateway Borough, AK",02130,4.2,6.5,13.4,"SAIPE, 2021",No,1,,
+"Kodiak Island Borough, AK",02150,5.5,6.6,10.4,"SAIPE, 2021",No,1,,
+"Kusilvak Census Area, AK6",02158,31,26.2,36.7,"SAIPE, 2021",Yes,2,Came into existence in 2015. Used to be Wade Hampton Census Area (02270). Percent in poverty hand entered based on prior geography.,
+"Lake and Peninsula Borough, AK",02164,20,18.9,25.5,"SAIPE, 2021",No,1,,
+"Matanuska-Susitna Borough, AK",02170,9.4,11,11.9,"SAIPE, 2021",No,1,,
+"Nome Census Area, AK",02180,22.4,17.4,23.9,"SAIPE, 2021",No,1,,
+"North Slope Borough, AK",02185,8.7,9.1,15.6,"SAIPE, 2021",No,1,,
+"Northwest Arctic Borough, AK",02188,18.5,17.4,25.5,"SAIPE, 2021",No,1,,
+"Petersburg Borough, AK7",02195,5.7,7.9,9.9,"SAIPE, 2021",No,1,Came into existence in 2009. Used to be Wrangell−Petersburg Census Area (02280). Percent in poverty hand entered based on prior geography.,
+"Prince of Wales-Hyder Census Area, AK8",02198,9.1,12.1,18.8,"SAIPE, 2021",No,1,Came into existence in 2009. Used to be Prince of Wales Outer Ketchikan Census Area (02201). Percent in poverty hand entered based on prior greography.,
+"Sitka Borough, AK",02220,4.8,7.8,10.6,"SAIPE, 2021",No,1,,
+"Skagway Municipality, AK5",02230,NA,NA,7.6,"SAIPE, 2021",NA,0,Came into existence in 2008. Used to be Skagway-Hoonah-Angoon Census Area (02232).,
+"Southeast Fairbanks Census Area, AK",02240,14.2,18.9,16,"SAIPE, 2021",No,1,,
+"Wrangell City and Borough, AK7",02275,5.7,7.9,14.1,"SAIPE, 2021",No,1,Came into existence in 2009. Used to be Wrangell−Petersburg Census Area (02280). Percent in poverty hand entered based on prior geography.,
+"Yakutat City and Borough, AK5",02282,8.9,13.5,15.3,"SAIPE, 2021",No,1,Came into existence in 1992. Used to be Skagway-Yakutat-Angoon Census Area (02231). Percent in poverty hand entered based on prior geography. ,
+"Yukon-Koyukuk Census Area, AK",02290,26,23.8,29.1,"SAIPE, 2021",Yes,2,,
+"Apache County, AZ",04001,47.1,37.8,31.8,"SAIPE, 2021",Yes,2,,
+"Cochise County, AZ",04003,20.3,17.7,19.4,"SAIPE, 2021",No,1,,
+"Coconino County, AZ",04005,23.1,18.2,19.1,"SAIPE, 2021",No,1,,
+"Gila County, AZ",04007,18.3,17.4,20.4,"SAIPE, 2021",No,1,,
+"Graham County, AZ",04009,26.7,23,23.9,"SAIPE, 2021",Yes,2,,
+"Greenlee County, AZ",04011,12.6,9.9,11.8,"SAIPE, 2021",No,1,,
+"La Paz County, AZ",04012,28.2,19.6,24.8,"SAIPE, 2021",Yes,2,,
+"Maricopa County, AZ",04013,12.3,11.7,11.7,"SAIPE, 2021",No,1,,
+"Mohave County, AZ",04015,14.2,13.9,20,"SAIPE, 2021",No,1,,
+"Navajo County, AZ",04017,34.7,29.5,28,"SAIPE, 2021",Yes,2,,
+"Pima County, AZ",04019,17.2,14.7,15.7,"SAIPE, 2021",No,1,,
+"Pinal County, AZ",04021,23.6,16.9,12.2,"SAIPE, 2021",No,1,,
+"Santa Cruz County, AZ",04023,26.4,24.5,23.9,"SAIPE, 2021",Yes,2,,
+"Yavapai County, AZ",04025,13.6,11.9,14.2,"SAIPE, 2021",No,1,,
+"Yuma County, AZ",04027,19.9,19.2,19.7,"SAIPE, 2021",No,1,,
+"Arkansas County, AR",05001,20.4,17.8,21.8,"SAIPE, 2021",No,1,,
+"Ashley County, AR",05003,20.9,17.5,25.2,"SAIPE, 2021",No,1,,
+"Baxter County, AR",05005,16.3,11.1,15.7,"SAIPE, 2021",No,1,,
+"Benton County, AR",05007,9.6,10.1,9.2,"SAIPE, 2021",No,1,,
+"Boone County, AR",05009,13.9,14.8,16,"SAIPE, 2021",No,1,,
+"Bradley County, AR ",05011,24.9,26.3,24.3,"SAIPE, 2021",Yes,2,,
+"Calhoun County, AR",05013,15.6,16.5,18.2,"SAIPE, 2021",No,1,,
+"Carroll County, AR",05015,15.2,15.5,20,"SAIPE, 2021",No,1,,
+"Chicot County, AR",05017,40.4,28.6,33.8,"SAIPE, 2021",Yes,2,,
+"Clark County, AR",05019,23.9,19.1,24.2,"SAIPE, 2021",No,1,,
+"Clay County, AR",05021,21.2,17.5,21.9,"SAIPE, 2021",No,1,,
+"Cleburne County, AR",05023,17.3,13.1,19.5,"SAIPE, 2021",No,1,,
+"Cleveland County, AR",05025,19,15.2,17.5,"SAIPE, 2021",No,1,,
+"Columbia County, AR",05027,24.4,21.1,27.6,"SAIPE, 2021",Yes,2,,
+"Conway County, AR",05029,16.5,16.1,20.3,"SAIPE, 2021",No,1,,
+"Craighead County, AR",05031,17,15.4,20.6,"SAIPE, 2021",No,1,,
+"Crawford County, AR",05033,16.3,14.2,17.6,"SAIPE, 2021",No,1,,
+"Crittenden County, AR",05035,27.1,25.3,27,"SAIPE, 2021",Yes,2,,
+"Cross County, AR",05037,25.4,19.9,20.6,"SAIPE, 2021",Yes,2,,
+"Dallas County, AR",05039,22.3,18.9,25,"SAIPE, 2021",No,1,,
+"Desha County, AR",05041,34,28.9,32.8,"SAIPE, 2021",Yes,2,,
+"Drew County, AR",05043,24.2,18.2,20.7,"SAIPE, 2021",No,1,,
+"Faulkner County, AR",05045,13.8,12.5,15.4,"SAIPE, 2021",No,1,,
+"Franklin County, AR",05047,20.4,15.2,20.8,"SAIPE, 2021",No,1,,
+"Fulton County, AR",05049,26.3,16.3,22,"SAIPE, 2021",No,1,,
+"Garland County, AR",05051,18,14.6,17,"SAIPE, 2021",No,1,,
+"Grant County, AR",05053,14.9,10.2,13.8,"SAIPE, 2021",No,1,,
+"Greene County, AR",05055,17.9,13.3,15.7,"SAIPE, 2021",No,1,,
+"Hempstead County, AR",05057,22.7,20.3,25.4,"SAIPE, 2021",Yes,2,,
+"Hot Spring County, AR",05059,18.6,14,21,"SAIPE, 2021",No,1,,
+"Howard County, AR",05061,18.6,15.5,21.6,"SAIPE, 2021",No,1,,
+"Independence County, AR",05063,17.1,13,21.4,"SAIPE, 2021",No,1,,
+"Izard County, AR",05065,21.1,17.2,22.4,"SAIPE, 2021",No,1,,
+"Jackson County, AR",05067,26.6,17.4,27.8,"SAIPE, 2021",No,1,,
+"Jefferson County, AR",05069,23.9,20.5,24.7,"SAIPE, 2021",Yes,2,,
+"Johnson County, AR",05071,20.1,16.4,21.2,"SAIPE, 2021",No,1,,
+"Lafayette County, AR",05073,34.7,23.2,28.6,"SAIPE, 2021",Yes,2,,
+"Lawrence County, AR",05075,25,18.4,24.2,"SAIPE, 2021",No,1,,
+"Lee County, AR",05077,47.3,29.9,43.6,"SAIPE, 2021",Yes,2,,
+"Lincoln County, AR",05079,26.2,19.5,28.3,"SAIPE, 2021",Yes,2,,
+"Little River County, AR",05081,19.3,15.4,19.6,"SAIPE, 2021",No,1,,
+"Logan County, AR",05083,19.3,15.4,18.8,"SAIPE, 2021",No,1,,
+"Lonoke County, AR",05085,14.9,10.5,13,"SAIPE, 2021",No,1,,
+"Madison County, AR",05087,20.1,18.6,18.9,"SAIPE, 2021",No,1,,
+"Marion County, AR",05089,18.9,15.2,21.3,"SAIPE, 2021",No,1,,
+"Miller County, AR",05091,22.4,19.3,23.4,"SAIPE, 2021",No,1,,
+"Mississippi County, AR",05093,26.2,23,28.1,"SAIPE, 2021",Yes,2,,
+"Monroe County, AR",05095,35.9,27.5,31.3,"SAIPE, 2021",Yes,2,,
+"Montgomery County, AR",05097,23.8,17,23.9,"SAIPE, 2021",No,1,,
+"Nevada County, AR",05099,20.3,22.8,25.9,"SAIPE, 2021",Yes,2,,
+"Newton County, AR",05101,29.6,20.4,23.7,"SAIPE, 2021",Yes,2,,
+"Ouachita County, AR",05103,21.2,19.5,24.4,"SAIPE, 2021",Yes,2,,
+"Perry County, AR",05105,20.3,14,18.6,"SAIPE, 2021",No,1,,
+"Phillips County, AR",05107,43,32.7,42.1,"SAIPE, 2021",Yes,2,,
+"Pike County, AR",05109,17.9,16.8,22.2,"SAIPE, 2021",No,1,,
+"Poinsett County, AR",05111,25.6,21.2,25.9,"SAIPE, 2021",Yes,2,,
+"Polk County, AR",05113,18.5,18.2,23.4,"SAIPE, 2021",No,1,,
+"Pope County, AR",05115,15.4,15.2,20.7,"SAIPE, 2021",No,1,,
+"Prairie County, AR",05117,22.7,15.5,19.2,"SAIPE, 2021",No,1,,
+"Eastern District, AS",60010,56,58.6,52.2,"Decennial Census, 2020",Yes,2,,
+"Manu'a District, AS",60020,79.8,65.2,60.5,"Decennial Census, 2020",Yes,2,,
+"Western District, AS",60050,59.3,62.6,55.7,"Decennial Census, 2020",Yes,2,,
+"Guam, GU",66010,15,23,20.2,"Decennial Census, 2020",No,1,,
+"Northern Islands Municipality, MP",69085,72.2,83.3,42.9,"Decennial Census, 2020",Yes,2,,
+"Rota Municipality, MP",69100,51.7,34.2,36.4,"Decennial Census, 2020",Yes,2,,
+"Saipan Municipality, MP",69110,51.4,46.9,38.2,"Decennial Census, 2020",Yes,2,,
+"Tinian Municipality, MP",69120,51.3,41.2,35.5,"Decennial Census, 2020",Yes,2,,
+"Adjuntas Municipio, PR",72001,81.5,65.4,71.3,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Aguada Municipio, PR",72003,69.7,59.3,48.4,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Aguadilla Municipio, PR",72005,65.3,55,52.6,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Aguas Buenas Municipio, PR",72007,62.6,51.7,47.1,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Aibonito Municipio, PR",72009,60.6,51.8,48.3,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Añasco Municipio, PR",72011,61.7,51.6,49.1,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Arecibo Municipio, PR",72013,63.9,50.9,49.4,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Arroyo Municipio, PR",72015,70.7,55.1,66.3,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"Barceloneta Municipio, PR",72017,64.4,56,50.4,"ACS 5-YEAR, 2017-2021",Yes,2,,
+"St. Croix Island, VI",78010,33.7,38.7,24.9,"Decennial Census, 2020",Yes,2,,
+"St. John Island, VI",78020,15,18.5,18.9,"Decennial Census, 2020",No,1,,
+"St. Thomas Island, VI",78030,21.3,27.2,21.2,"Decennial Census, 2020",Yes,2,,
+"Source: 1990, Decennial Census, Public Use; 2000, Decennial Census, Public Use; 2021 Small Area Income and Poverty Estimates, Public Use; 2017-2021 5-Year American Community Survey, Public Use; 2020 Decennial Census, Public Use",,,,,,,,,
+"*Most recent estimate, % in poverty -- For continental U.S., it's 2021 SAIPE, which uses the upper bound of the confidence interval (for statistical uncertainty). For Puerto Rico, it's 2021 ACS 5-year, which uses the percentage estimate+MOE. For Island Areas, it's 2020 decennial number.",,,,,,,,,
+"Note: All three data points must have a valid numerical percent in poverty >= 19.5 percent to be considered in persistent poverty. We are using 19.5 because values for percent in poverty are reflected to one decimal point and EDA rounds up to whole integers. Counties without a numerical percent in poverty are coded as NA. These included areas that were not in existence, experienced a boundary change or are uninhabited. See Technical Documentation for more details.",,,,,,,,,
+"For more information on the decennial census, see the technical documentation here: https://www.census.gov/programs-surveys/decennial-census/technical-documentation/complete-technical-documents.html",,,,,,,,,
+"For more information on the American Community Survey and the margins of error, see the technical documentation here: https://www.census.gov/programs-surveys/acs/technical-documentation.html",,,,,,,,,
+"For more information on SAIPE and the estimation of the 90-percent confidence intervals, see the technical documentation here: https://www.census.gov/programs-surveys/saipe/technical-documentation/methodology/counties-states/county-level.html",,,,,,,,,
diff --git a/statvar_imports/commerce_eda_poverty/validation_config.json b/statvar_imports/commerce_eda_poverty/validation_config.json
new file mode 100644
index 0000000000..0bd8655490
--- /dev/null
+++ b/statvar_imports/commerce_eda_poverty/validation_config.json
@@ -0,0 +1,70 @@
+{
+ "schema_version": "1.0",
+ "rules": [
+ {
+ "rule_id": "check_percent_min_value",
+ "validator": "MIN_VALUE_CHECK",
+ "description": "Ensure poverty percentage values are non-negative (>= 0.0%).",
+ "params": {
+ "minimum": 0
+ }
+ },
+ {
+ "rule_id": "check_percent_max_value",
+ "validator": "MAX_VALUE_CHECK",
+ "description": "Ensure poverty percentage values do not exceed 100.0%.",
+ "params": {
+ "maximum": 100
+ }
+ },
+ {
+ "rule_id": "check_num_places_count",
+ "validator": "NUM_PLACES_COUNT",
+ "description": "Verify county coverage is between 3,100 and 3,250 places.",
+ "params": {
+ "minimum": 3100,
+ "maximum": 3250
+ }
+ },
+ {
+ "rule_id": "check_num_observations_count",
+ "validator": "NUM_OBSERVATIONS_CHECK",
+ "description": "Verify total observations count is between 9,000 and 10,000.",
+ "params": {
+ "minimum": 9000,
+ "maximum": 10000
+ }
+ },
+ {
+ "rule_id": "check_max_date_consistent",
+ "validator": "MAX_DATE_CONSISTENT",
+ "description": "Ensure MaxDate is uniform across all StatVars in the import.",
+ "params": {}
+ },
+ {
+ "rule_id": "check_date_span_sql",
+ "validator": "SQL_VALIDATOR",
+ "description": "Verify MinDate is 1990 and MaxDate is >= 2021.",
+ "params": {
+ "query": "SELECT StatVar, MinDate, MaxDate FROM stats",
+ "condition": "TRY_CAST(MinDate AS INT) = 1990 AND TRY_CAST(MaxDate AS INT) >= 2021"
+ }
+ },
+ {
+ "rule_id": "check_missing_refs_count",
+ "validator": "MISSING_REFS_COUNT",
+ "description": "Ensure zero unresolved entity or schema references.",
+ "params": {
+ "threshold": 0
+ }
+ },
+ {
+ "rule_id": "check_lint_error_count",
+ "validator": "LINT_ERROR_COUNT",
+ "description": "Ensure zero MCF/TMCF lint errors in generated output.",
+ "params": {
+ "threshold": 0
+ }
+ }
+ ]
+}