Andreas Züfle

h-index1
3papers
2citations

3 Papers

6.0AIJun 4
An Infectious Disease Spread Simulation Based on Large Language Model Decision Making

Yonchanok Khaokaew, Ruochen Kong, Andreas Zufle et al.

Modelling individual decision-making during infectious disease outbreaks is crucial for understanding behavioural dynamics and informing effective public health interventions. Prior work has shown that large language models can simulate realistic human behaviour by generating agent decisions based on demographic prompts and situational context. We build on this foundation with a spatially grounded, agent-based simulation framework that integrates LLM-generated decisions about self-reported influenza-like illness into a census-based synthetic population of agents. Location is treated as a central feature: agents are assigned to spatial units within cities, capturing the spatial distributions of different demographic groups using real-world census data and enabling geographically diverse behavioural modelling. We implement and compare three decision scenarios, independent reasoning, household influence, and message framing, and simulate self-reporting outcomes in San Francisco and Atlanta. Results reveal that income and education are the dominant drivers of reporting rate variation, with smaller but consistent effects from geography, LLM model choice, and message framing. Our framework generates synthetic data that captures both social and geographic heterogeneity, supporting spatial epidemiological modelling and bias-aware behavioural analysis.

11.9SOC-PHMay 29
SF-LIFE: A Large-Scale Simulated Movement Dataset for the San Francisco Bay Area

Chanuka Algama, Taylor Anderson, Henrique Ferraz de Arruda et al.

We introduce SF-LIFE, a large-scale simulated movement dataset designed to accelerate research in transportation, mobility, and machine learning. The dataset contains 3,024,000,000,000 location records capturing complete, noise-free, multi-modality trajectories of 500,000 simulated agents observed at a 1Hz frequency navigating the San Francisco Bay Area network over a 70-day period. The data captures (1) needs-driven daily agendas of individual agents generated by an agent-based simulation of human patterns of life and (2) detailed kinematic trajectories moving agents across the OpenStreetMap representation of San Francisco using data from 40+ transit agencies across 9 counties. SF-LIFE provides unprecedented scale and detail as trajectories are based on real transit infrastructure using San Francisco General Transit Feed Specification (GTFS) data, having agent movements across multiple modalities, including bus, rail, bike, automobile, and walking. For this high-fidelity simulated representation of San Francisco, we provide (1) the full trajectory data annotated with transportation mode labels, (2) reduced-size versions of the trajectory data with reduced temporal frequency, (3) agent activity information describing the causal activity why an agent visits a place, (4) agent demographic data, and (5) the underlying OSM road network and building data. As the first dataset of its scale and level of detail, SF-LIFE overcomes the privacy, noise, and completeness limitations inherent in real-world tracking data, providing a robust and ethically sourced resource for research in transit optimization, human mobility analysis, and urban computing.

6.2CEJun 20
Simulating Public Transit Fare Policies in NYC: An Efficient, Socioeconomic-Aware Framework

Parker Wischhover, Hossein Amiri, Kiara Ha et al.

Designing equitable and effective public transit fare policies is challenging due to complex interactions among traveler behavior, multimodal networks, and socioeconomic heterogeneity. This paper presents a scalable, data-driven simulation framework for evaluating transit fare policies in New York City (NYC), integrating a synthetic population, agent-based simulation, multimodal travel-time estimation, and fare-sensitive mode choice modeling. We evaluate multiple fare scenarios, including distance-based pricing, fare increases, and fare-free bus policies. Results show that pricing changes modestly affect total ridership but significantly alter modal composition and produce heterogeneous impacts across income groups. In particular, fare-free bus policies generate substantial benefits for lower-income riders by increasing bus usage and reducing fare burden, while introducing trade-offs in revenue. To support city-scale analysis, we introduce a sampling-based approach that reduces computational cost while preserving aggregate accuracy. The proposed framework provides a practical tool for assessing trade-offs between ridership, revenue, and equity, enabling more informed and equitable transit policy design.