Skip to content

About

Reusable Python utilities for California economics and geography research: Census API, geo crosswalks, ArcGIS REST, file downloads

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

python-geo-utils

Reusable Python utilities for California economics and geography research projects. Extracted from working research code across multiple projects; designed to be dropped directly into any scripts/ directory or imported as a local module.

All utilities are self-contained single-file modules with no internal dependencies on each other (except geo_crosswalk.py, which imports download_utils.py).


Utilities

Module Description Download
download_utils.py Download ZIP archives and individual files with skip-if-exists logic. Safe to call repeatedly in reproducible pipelines. download
census_api.py Fetch ACS 5-year estimates via the Census Bureau Data API. Handles batching (50-var limit), GEOID construction, and sentinel value (-666666666) masking. download
geo_crosswalk.py Build area-proportional geographic crosswalk tables: ZCTA→tract, county→tract, and precinct→tract (spatial overlay). All crosswalks target 2020 Census tract boundaries. download
arcgis_rest.py Paginated GeoJSON downloads from ArcGIS REST feature services. Handles server page limits automatically. No third-party dependencies (stdlib urllib only). download

Quick start

Copy the module(s) you need into your project's scripts/ directory and import normally:

from census_api import fetch_acs_tracts
from geo_crosswalk import build_county_tract, build_zip_tract
from arcgis_rest import paginate_geojson, save_geojson
from download_utils import download_zip, download_file

Or install requests + geopandas + pandas and run any module directly:

pip install requests geopandas pandas

Module details

download_utils.py

Functions: download_zip, download_file

Dependencies: requests

Downloads a ZIP file from a URL and extracts it, or downloads a single file. Both functions skip the download if the destination already exists — safe for use in idempotent pipelines.

from download_utils import download_zip, download_file
from pathlib import Path

# Extract a Census TIGER shapefile ZIP
shp_dir = download_zip(
    url="https://www2.census.gov/geo/tiger/TIGER2020/TRACT/tl_2020_06_tract.zip",
    dest_dir=Path("data/raw/shapefiles"),
    name="tl_2020_06_tract",
)

# Download a plain-text relationship file
txt = download_file(
    url="https://www2.census.gov/geo/docs/maps-data/data/rel2020/zcta520/tab20_zcta520_tract20_natl.txt",
    dest_path=Path("data/raw/shapefiles/tab20_zcta520_tract20_natl.txt"),
)

census_api.py

Functions: fetch_acs_tracts, fetch_acs_batch, build_geoid, mask_sentinel

Dependencies: requests, pandas

Fetches ACS 5-year estimates for all Census tracts in a state. Handles the Census API's 50-variable-per-request limit by accepting a list of batches. Automatically builds 11-digit GEOIDs and replaces the Census sentinel value (-666666666) with pd.NA.

Get a free API key at api.census.gov/data/key_signup.html and set export CENSUS_API_KEY=your_key.

from census_api import fetch_acs_tracts

# Pull ACS B25038 (tenure by move-in year) for CA tracts — 2020 5-year estimates
BATCHES = [
    ["B25038_001E", "B25038_002E", "B25038_003E", "B25038_006E", "B25038_010E"],
]
LABELS = {
    "B25038_001E": "owner_occ_total",
    "B25038_002E": "moved_in_2015plus",
    "B25038_003E": "moved_in_2010_2014",
    "B25038_006E": "moved_in_2000_2009",
    "B25038_010E": "moved_in_pre2000",
}

df = fetch_acs_tracts(year=2020, variable_batches=BATCHES, variable_labels=LABELS)
df.to_csv("data/raw/acs/acs_tenure_ca_2020.csv", index=False)

geo_crosswalk.py

Functions: build_zip_tract, build_county_tract, build_prec_tract, check_weights, download_tiger_tracts

Dependencies: geopandas, pandas, requests, download_utils

All three crosswalks produce a CSV with a source-unit column, a tract_geoid_20 column, and a weight column that sums to 1.0 per source unit.

Crosswalk Use case
build_zip_tract Assign ZIP-level data (Zillow ZHVI, CEC ZEV counts) to Census tracts
build_county_tract Broadcast county-panel data (insurance premiums, YCOM beliefs) to all tracts
build_prec_tract Assign precinct-level election results to Census tracts
from geo_crosswalk import build_county_tract, build_zip_tract
from pathlib import Path

# County → tract (for broadcasting Keys-Mulder insurance panel to tracts)
build_county_tract(
    acs_tracts_csv=Path("data/raw/acs/acs_tenure_ca_2020.csv"),
    output_path=Path("data/processed/crosswalk_county_tract.csv"),
)

# ZIP → tract (for Zillow ZHVI)
build_zip_tract(
    output_path=Path("data/processed/crosswalk_zip_tract.csv"),
    shapes_dir=Path("data/raw/shapefiles"),
)

arcgis_rest.py

Functions: paginate_geojson, get_record_count, save_geojson, load_geojson

Dependencies: stdlib only (urllib, json, time)

Downloads all features from an ArcGIS REST feature service, handling server-side pagination transparently. Useful for California state agency datasets on GIS servers (CalFire, CalEPA, DWR, CDFW).

from arcgis_rest import paginate_geojson, save_geojson
from pathlib import Path

# CalFire Fire Hazard Severity Zones (SRA, 2007 designations)
gj = paginate_geojson(
    base_url=(
        "https://services.gis.ca.gov/arcgis/rest/services/Environment/"
        "Fire_Severity_Zones/MapServer/0/query"
    ),
    out_fields="HAZ_CLASS,SRA",
)
save_geojson(gj, Path("data/raw/calfire_fhsz/fhsz_sra.geojson"))

# Load into GeoPandas
import geopandas as gpd
gdf = gpd.GeoDataFrame.from_features(gj["features"], crs="EPSG:4326")

Projects using these utilities


License

MIT

About

Reusable Python utilities for California economics and geography research: Census API, geo crosswalks, ArcGIS REST, file downloads

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages