← Back to all work
Data Scientist, Machine Learning Engineer • 2025 ACTIVE

Exploring Disparate Impact with Geospatial Analysis and Communication Graphing

This research explores whether law enforcement disproportionately patrols certain communities, combining geospatial analysis and communication graphing to uncover patterns. The work supports investigations into potential violations of the Fourth Amendment (Search and Seizure) and the Fourteenth Amendment (Equal Protection Clause).

Creative Activity Research Day ↗
Technologies & Frameworks
NetworkXGraph toolingPythonVisualizationArcGISYOLOv5
Key System Outcomes
✓ Open-ended✓ Reduced time-to-action

Georgia von Minden presenting at CARD 2025

Research Questions

  • Where does patrolling occur most intensely?
  • How are these activities discussed internally by law enforcement?
  • Do patterns show disparate impact on protected communities?

Disparate impact refers to policies or practices that adversely affect one group more than others, even when not intentionally discriminatory.


Geospatial Analysis

Dataset

  • Unit of analysis: Census Block Groups (n=230)

  • Attributes extracted using ArcGIS Pro:

    • Demographic: Race/Ethnicity, Poverty Rate, Median Income, English Learner Status, Educational Attainment, Unemployment Rate
    • Environmental: Wetland & Watershed Coverage
    • Zoning: Residential Housing Density, Private Area, Population Estimates, Homeowner Rate
  • Targets:

    • All patrol reports
    • Cannabis-flagged patrol reports
      (Derived from PRA requests and internal documents)

Modeling

  • Model type: Random Forest Regression
  • Train/Test split: 184 / 46 block groups
  • Goal: Assess if cannabis-related patrols concentrate in specific communities

Results

All Patrol Reports

Test R²: 0.41
Validation R²: 0.52

FeatureImportance Score
Wetland/Watershed0.1909
Poverty Rate0.1362
Population Estimate0.1048
Bachelor’s Attainment0.1021
Residential Housing Density0.0917

Cannabis-Flagged Reports

Test R²: Invalid
Validation R²: 0.44

FeatureImportance Score
Bachelor’s Attainment0.2825
Hispanic/Latino0.1642
Wetland/Watershed0.1640
English Learner Rate0.1083
Residential Housing Density0.0741

SHAP Explanations (Cannabis-Flagged Reports)

FeatureMean |SHAP|Mean SHAPStdDev
Bachelor’s Attainment0.44240.05450.9268
English Learner Status0.25750.00450.5071
Hispanic/Latino0.21710.03650.4634
Residential Housing Density0.1324-0.00750.2714
Wetland/Watershed0.05100.00120.6608

Conclusion: Patterns suggest patrols may be concentrated in areas with higher Hispanic/Latino and English Learner populations, especially for cannabis-related investigations. However, model performance is moderate—results should guide further research, not legal conclusions.


Communication Graphing

Dataset

  • ~5,000 internal emails from a law enforcement agency over 4 years
  • ~200 employees, multiple teams
  • Collected via PRA request
  • Unstructured and noisy (non-standard formats)

Methodology

  1. Latent Dirichlet Allocation (LDA) for topic discovery
  2. Network graphing to visualize:
    • Who discusses which topics
    • Topic evolution over time
  3. Statistical significance testing on topic frequency over time (α = 0.05)

Example Workflow

  1. Extracted topic: “permit”, “surveillance”, “cannabis”, “dwelling”
  2. Batch group emails by this topic
  3. Identify individuals discussing it
  4. Map topic intensity over time across the org

Results

  • “Hot docs” identified through LDA improve investigator efficiency
  • Communication networks showed:
    • Team- and individual-level differences
    • Noticeable shifts after legal action (e.g., lawsuit filing)
  • Significant trends in cannabis enforcement topics over time

Conclusion: LDA-based communication modeling provides useful prioritization for litigation teams and supports rapid legal review of internal documents.


Future Directions

  • Modeling improvements:
    Use GWR, GAM, and zero-inflated models to improve patrol prediction accuracy.

  • Communication modeling enhancements:
    Explore graph-based deep learning (e.g., GNNs, GANs) to enrich topic extraction and document prioritization.


Acknowledgements

Special thanks to:

  • Robert Clements, USF M.S. in Data Science
  • Dylan Verner-Crist and Ian Duke, ACLU of Northern California

This work was supported by the Master’s in Data Science program at the University of San Francisco.


Contact

Hadley Dixon
📧 hadley.dixon22@gmail.com

Georgia von Minden
📧 georgia.vonminden@gmail.com

← Back to all work Discuss this project →