GIS Data Scientist — transit analytics & geospatial engineering
I turn transit operations problems into data systems: GTFS pipelines, service KPIs, and the spatial models underneath them. Currently at the Volunteer Transportation Center, running the geospatial and data backbone for public transit across five counties in northern New York — the analytics and tools used by planners, dispatchers, and riders.
M.S. Applied Data Science, Clarkson University · B.E. Computer Science, Savitribai Phule Pune University · Selected speaker, MobilityData Conference, Montréal 2024
🌐 arpatil.xyz — full portfolio, project write-ups, and résumé
📍 Potsdam, NY — open to remote and willing to relocate. Currently looking for Transit Data Analyst, GIS Data Analyst, Data Engineer, and Data Scientist roles.
GTFS-RT pipeline for $4/month — Lambda + S3 with 30-day history, powering on-time performance and service analysis across the agency. Unified GTFS static/RT/Flex behind automated QA and standardized the agency's ridership, OTP, coverage, and access/equity KPIs. Cut manual data processing in half. Presented the architecture at MobilityData 2024.
Unified service design across 5 counties — aligned fixed-route, first/last-mile, and zone-based microtransit using demand, travel-time, and OTP analysis. Co-developed the microtransit zones and service windows, and shipped LLM assistants for fixed-route dispatch that improved dispatch efficiency by 30%.
GTFS-RT validation on $8/month — data-quality monitoring for St. Lawrence County Transit's realtime feed: Terraform-provisioned EC2, Dockerized MobilityData validator, cron scheduling aligned to service hours, SNS alerting on critical errors, and S3 report history. → SLC-Transit-GTFS-RT-Validator
Wildlife–vehicle collision hotspots, A2A corridor — joined 1,000+ collision records with traffic volume, speed limits, land cover, water proximity, and slope; kernel-density screening followed by penalized regression to rank predictors. Mitigation layers delivered to 3 counties.
Booking cancellation model at 88.3% accuracy — CatBoost with cross-validated hyperparameter search; feature importance and partial dependence to explain lead time, ADR, and cancellation history. → repo
Carbon emissions risk models — analyzed 10M+ environmental audit records across US/EU/CA/IN and built Random Forest predictors for emissions risk. → repo



