Reusable Python utilities for California economics and geography research projects. Extracted from working research code across multiple projects; designed to be dropped directly into any scripts/ directory or imported as a local module.
All utilities are self-contained single-file modules with no internal dependencies on each other (except geo_crosswalk.py, which imports download_utils.py).
| Module | Description | Download |
|---|---|---|
download_utils.py |
Download ZIP archives and individual files with skip-if-exists logic. Safe to call repeatedly in reproducible pipelines. | download |
census_api.py |
Fetch ACS 5-year estimates via the Census Bureau Data API. Handles batching (50-var limit), GEOID construction, and sentinel value (-666666666) masking. |
download |
geo_crosswalk.py |
Build area-proportional geographic crosswalk tables: ZCTA→tract, county→tract, and precinct→tract (spatial overlay). All crosswalks target 2020 Census tract boundaries. | download |
arcgis_rest.py |
Paginated GeoJSON downloads from ArcGIS REST feature services. Handles server page limits automatically. No third-party dependencies (stdlib urllib only). |
download |
Copy the module(s) you need into your project's scripts/ directory and import normally:
from census_api import fetch_acs_tracts
from geo_crosswalk import build_county_tract, build_zip_tract
from arcgis_rest import paginate_geojson, save_geojson
from download_utils import download_zip, download_fileOr install requests + geopandas + pandas and run any module directly:
pip install requests geopandas pandasFunctions: download_zip, download_file
Dependencies: requests
Downloads a ZIP file from a URL and extracts it, or downloads a single file. Both functions skip the download if the destination already exists — safe for use in idempotent pipelines.
from download_utils import download_zip, download_file
from pathlib import Path
# Extract a Census TIGER shapefile ZIP
shp_dir = download_zip(
url="https://www2.census.gov/geo/tiger/TIGER2020/TRACT/tl_2020_06_tract.zip",
dest_dir=Path("data/raw/shapefiles"),
name="tl_2020_06_tract",
)
# Download a plain-text relationship file
txt = download_file(
url="https://www2.census.gov/geo/docs/maps-data/data/rel2020/zcta520/tab20_zcta520_tract20_natl.txt",
dest_path=Path("data/raw/shapefiles/tab20_zcta520_tract20_natl.txt"),
)Functions: fetch_acs_tracts, fetch_acs_batch, build_geoid, mask_sentinel
Dependencies: requests, pandas
Fetches ACS 5-year estimates for all Census tracts in a state. Handles the Census API's 50-variable-per-request limit by accepting a list of batches. Automatically builds 11-digit GEOIDs and replaces the Census sentinel value (-666666666) with pd.NA.
Get a free API key at api.census.gov/data/key_signup.html and set export CENSUS_API_KEY=your_key.
from census_api import fetch_acs_tracts
# Pull ACS B25038 (tenure by move-in year) for CA tracts — 2020 5-year estimates
BATCHES = [
["B25038_001E", "B25038_002E", "B25038_003E", "B25038_006E", "B25038_010E"],
]
LABELS = {
"B25038_001E": "owner_occ_total",
"B25038_002E": "moved_in_2015plus",
"B25038_003E": "moved_in_2010_2014",
"B25038_006E": "moved_in_2000_2009",
"B25038_010E": "moved_in_pre2000",
}
df = fetch_acs_tracts(year=2020, variable_batches=BATCHES, variable_labels=LABELS)
df.to_csv("data/raw/acs/acs_tenure_ca_2020.csv", index=False)Functions: build_zip_tract, build_county_tract, build_prec_tract, check_weights, download_tiger_tracts
Dependencies: geopandas, pandas, requests, download_utils
All three crosswalks produce a CSV with a source-unit column, a tract_geoid_20 column, and a weight column that sums to 1.0 per source unit.
| Crosswalk | Use case |
|---|---|
build_zip_tract |
Assign ZIP-level data (Zillow ZHVI, CEC ZEV counts) to Census tracts |
build_county_tract |
Broadcast county-panel data (insurance premiums, YCOM beliefs) to all tracts |
build_prec_tract |
Assign precinct-level election results to Census tracts |
from geo_crosswalk import build_county_tract, build_zip_tract
from pathlib import Path
# County → tract (for broadcasting Keys-Mulder insurance panel to tracts)
build_county_tract(
acs_tracts_csv=Path("data/raw/acs/acs_tenure_ca_2020.csv"),
output_path=Path("data/processed/crosswalk_county_tract.csv"),
)
# ZIP → tract (for Zillow ZHVI)
build_zip_tract(
output_path=Path("data/processed/crosswalk_zip_tract.csv"),
shapes_dir=Path("data/raw/shapefiles"),
)Functions: paginate_geojson, get_record_count, save_geojson, load_geojson
Dependencies: stdlib only (urllib, json, time)
Downloads all features from an ArcGIS REST feature service, handling server-side pagination transparently. Useful for California state agency datasets on GIS servers (CalFire, CalEPA, DWR, CDFW).
from arcgis_rest import paginate_geojson, save_geojson
from pathlib import Path
# CalFire Fire Hazard Severity Zones (SRA, 2007 designations)
gj = paginate_geojson(
base_url=(
"https://services.gis.ca.gov/arcgis/rest/services/Environment/"
"Fire_Severity_Zones/MapServer/0/query"
),
out_fields="HAZ_CLASS,SRA",
)
save_geojson(gj, Path("data/raw/calfire_fhsz/fhsz_sra.geojson"))
# Load into GeoPandas
import geopandas as gpd
gdf = gpd.GeoDataFrame.from_features(gj["features"], crs="EPSG:4326")- Prop 13 / Insurance Wedge Paper — uses
census_api,arcgis_rest,geo_crosswalk,download_utils - Hummers or Hybrids Replication — original source of
geo_crosswalkandcensus_apipatterns
MIT