Deep Activity Model: A Generative Approach for Human Mobility Pattern Synthesis
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Deep Activity Model: A Generative Deep Learning Approach for Human Mobility Pattern Synthesis".
Jane: The paper was written by Xishun Liao, Qinhua Jiang, Brian Yueshuai He, Yifan Liu, Chenchen Kuai et al. from University of California, Los Angeles and University of Louisville and Texas A&M University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 1: Tom: Welcome back to the show, everybody. Today we're digging into a paper that's got a wonderfully ambitious title: "Deep Activity Model: A Generative Deep Learning Approach for Human Mobility Pattern Synthesis." Jane, I have to say, just reading that title got me excited, because it's basically saying we can teach a computer to generate realistic daily schedules for millions of people.
Jane: And that's a bigger deal than it sounds, Tom. For decades, transportation planners have built these massive models called Activity-Based Models, or ABMs, to predict how people move around a city. But those models are incredibly expensive to build. They need years of survey data, tons of calibration, and they're so rigid that you can't easily move them to a new city.
Tom: Right, and this paper from the UCLA Mobility Lab is saying, what if we just train a deep learning model on the standard household travel surveys that almost every region already collects? The authors, led by Xishun Liao and Jiaqi Ma, they're using a Transformer architecture—the same kind of engine behind large language models—to generate a person's entire day of activities.
Jane: So instead of writing a thousand rules about when people go to work or shop, the model learns the patterns directly from data. It takes your age, your job, your household income, who else lives with you, and it generates a sequence of activities with start times and end times. It's like a language model, but instead of predicting the next word, it predicts the next activity.
Tom: And the coolest part is the input. They don't just feed in your personal attributes. They also feed in the attributes of everyone else in your household, with these special separator tokens. So the model literally learns how your spouse's work schedule or your kid's school day influences your own decisions.
Jane: That's the household interdependency piece, and it's something traditional models really struggle with. We know that if one parent drops the kids at school, the other parent's schedule shifts. This model is trying to capture that automatically from the data.
Tom: And they trained it on the National Household Travel Survey, which is this massive open-source dataset from the US government with over one hundred twenty-nine thousand households. Then they fine-tuned it for California, for the Puget Sound region, and even for Mexico City. The transferability is the real headline here.
Jane: Because that's the pain point. Building a model for a new region from scratch is a multi-year, multi-million-dollar project. If this approach works, you could take the pre-trained model, fine-tune it on a small local survey, and have a working travel demand model in weeks instead of years.
Tom: And they're not just claiming it works. They loaded the generated demand into a full traffic simulation of Los Angeles County and compared it against the official regional model. We'll get into those numbers in a bit, but let me just say, the results are surprisingly tight.
Jane: I love that they're validating at the system level, not just checking if individual activity chains look plausible. They're asking, does this synthetic population, when simulated on real roads, produce realistic congestion patterns? That's the test that matters for actual planning.
Tom: Exactly. And that's what we're going to dig into next—the actual architecture, the loss functions, and how they got a Transformer to behave so well on this kind of data.
Jane: Stay with us, because this is where the engineering gets clever.
Paper discussion segment 2: Tom: So Jane, we've set the stage. Now let's talk about how the Deep Activity Model actually works under the hood, because there are some genuinely clever choices here that make it work with survey data.
Jane: And that's important, because survey data is not like GPS trajectory data. It's sparse. A person might only have three or four activities in a day. So you can't just throw a massive deep learning model at it and hope for the best.
Tom: Right, and the authors made a really smart move with the loss functions. They're predicting activity type, start time, and end time. The type is a straightforward classification problem. But for the times, they use something called soft label loss. Instead of demanding the model predict the exact fifteen-minute time slot, they give partial credit to neighboring time slots.
Jane: That makes so much sense. If someone actually starts work at eight:seven and the model says eight:fifteen that's basically perfect. But a hard loss function would punish that as a full error. The soft label lets the model learn the general temporal pattern without overfitting to noise in the survey responses.
Tom: And they add two more loss terms that enforce temporal logic. One penalizes the model if an activity's end time comes before its start time. The other penalizes it if the end time of one activity overlaps with the start of the next one. So the model learns that time has to flow forward.
Jane: Those are the kind of constraints that make a generative model actually usable. Without them, you'd get nonsense chains where someone goes to work at three PM and comes home before they left.
Tom: They also tackled a classic problem with travel survey data: it's imbalanced. Most people do the same few things—home, work, home. So if you train naively, the model gets really good at generating boring chains and never learns to generate the interesting stuff like errands or recreation.
Jane: Their solution is a data balancing algorithm that iteratively reweights samples so that rare activities get more representation without completely flattening the natural distribution. They show that on the Mexico City dataset, this improved the activity type divergence by over seventy-six percent.
Tom: And that matters because the model isn't just trying to predict what a specific person did on a specific day. It's trying to learn the underlying distribution of human behavior so it can generate new, plausible chains for synthetic people.
Jane: Which is a crucial distinction. If you evaluate this model by asking, "did it predict this exact person's exact day correctly?" you're missing the point. The goal is to generate a population of agents whose aggregate behavior matches reality.
Tom: And on that front, the numbers are strong. On the national test set, the model achieves a Jensen-Shannon Divergence of zero point zero zero two for chain length and zero point zero zero three for start and end times. Those are tiny numbers, meaning the generated distributions are nearly indistinguishable from the real ones.
Jane: And they compared against a bunch of baselines—LSTM, GRU, even GPT-four and LLaMA. The LLMs were actually pretty bad at this task. They could generate plausible-looking text, but the temporal distributions were off, and they missed a lot of activity transitions.
Tom: That's a fascinating finding. The LLMs have all this world knowledge, but they don't have the statistical grounding in actual survey data. The specialized Transformer, trained on the right data with the right loss functions, just crushes them on fidelity.
Jane: And that's the core insight of this paper. It's not about having the biggest model. It's about designing the right inductive biases for the problem. Now, we should talk about what happens when they scale this up to a real city.
Paper discussion segment 3: Tom: So the activity generation works. But a chain of activities isn't a trajectory until you know where each activity happens. And that's where the location assignment module comes in. Jane, this is the part that makes the whole thing practical.
Jane: Right, because travel surveys don't usually include precise GPS coordinates. They tell you someone went to work, but not exactly which office building. So the authors developed what they call the Activity Location Assignment, or ALA, algorithm.
Tom: And the approach is clever. They use the fact that commute distances follow predictable distributions within a region. They split LA County into eight sub-regions, and for each one, they fit distributions for home-to-work distances, non-mandatory trip distances, and angular deviations.
Jane: So for a synthetic person, they first assign their work location based on the commute distance distribution for their home sub-region. Then, for each non-work activity, they assign a distance and a direction relative to the next anchor point. It's a statistical approach, not a rule-based one.
Tom: And they validate it against the SCAG Activity-Based Model, which is the official regional model for Southern California. They generated demand for a million synthetic agents and loaded it into the LASim traffic simulation. The cosine similarity of the origin-destination matrices was zero point nine nine seven.
Jane: That's essentially a perfect match in terms of spatial flow patterns. And the network-level vehicle-miles-traveled had a Mean Absolute Percentage Error of just under five percent compared to SCAG.
Tom: But here's the part that really got me. They compared the simulated traffic against real-world observations from Caltrans PeMS sensors on the I-four hundred five corridor. The traffic volume MAPE was five point eight five percent in the northbound direction, and speed error was around four point four percent.
Jane: And they captured the directional asymmetry in congestion. The northbound corridor gets congested in the morning, the southbound in the afternoon. The model reproduced that asymmetry, which is a very hard thing to get right because it depends on subtle differences in where people live and work.
Tom: Now, I want to bring in Lu and Meng, because I think they'll have strong opinions on this. Lu, you've been working on generative models for years. What do you make of the transfer learning results?
Lu: The transfer learning is the most exciting part for me. They pre-trained on the national survey, then fine-tuned on regional data. For California, they had about fifty-four thousand training chains. For Puget Sound, only about seven thousand eight hundred. And the model still produced distributions that closely matched the regional ground truth.
Meng: But I want to push back on the practical side. The fine-tuning for Mexico City required changing the input feature set and the activity taxonomy. That's not a trivial adaptation. How much manual work went into that?
Jane: That's a fair question. The paper acknowledges that Mexico City has a different activity classification—only ten categories instead of fifteen and some activities like exercise are grouped under recreation. So the transfer isn't fully automatic.
Lu: But the fact that it works at all is remarkable. The model learned the general structure of human daily behavior from the US data, and then it could adapt to a very different cultural context with different survey conventions. That suggests the model is learning something fundamental about how people organize their days.
Meng: And the computational cost? They mention training on a single RTX A5000 GPU. That's not a huge cluster. For a transportation agency, that's actually accessible.
Tom: Exactly. And that's the point. This isn't a research toy. It's a practical tool that could replace or augment the expensive, hand-crafted models that agencies have been using for decades.
Jane: And the implications go beyond transportation. If you can generate realistic synthetic populations, you can use them for public health modeling, for disaster evacuation planning, for energy demand forecasting. The same activity chains that drive traffic also drive electricity consumption and disease spread.
Lu: That's the vision. This paper is a foundation, not the final product.
Conclusion: Tom: Alright, we've covered a lot of ground on the Deep Activity Model. Let's pull it all together, because this paper is genuinely important.
Jane: It is. The core contribution is showing that a Transformer-based generative model, trained on standard household travel survey data, can produce realistic daily activity chains for synthetic populations. And it can do it in a way that transfers across regions.
Tom: And the validation is what sets it apart. They didn't just show that the generated chains look plausible. They loaded them into a real traffic simulation of Los Angeles County and showed that the resulting traffic patterns match both the official regional model and real-world sensor data.
Jane: The numbers bear repeating. A cosine similarity of zero point nine nine seven on origin-destination flows. A vehicle-miles-traveled error under five percent. Traffic volume and speed errors in the low single digits on a major corridor. Those are the kinds of numbers that make transportation planners sit up and take notice.
Tom: And the implications are huge. Building an activity-based model for a new region currently costs millions of dollars and takes years. This approach offers a path to do it in weeks, with just a small local survey and a pre-trained model.
Lu: The household interdependency modeling is also a real step forward. The attention mechanism lets the model learn how family members influence each other's schedules, which is something traditional models handle with rigid rules.
Meng: And from an engineering standpoint, the fact that it runs on a single GPU and uses open-source data means it's actually deployable. Agencies don't need a supercomputer or proprietary datasets.
Jane: There are limitations, of course. The location assignment is zone-level, not building-level. And the model inherits any biases in the survey data. But as a foundation, this is solid.
Tom: And the authors are clear about the future directions. They want to integrate point-of-interest level data, explore other generative architectures like diffusion models, and improve the temporal resolution.
Lu: I think the most exciting possibility is combining this with the emerging work on LLM-based agents. You could have a model like this generate the skeleton of a person's day, and then an LLM fills in the narrative details, the reasons behind the choices.
Jane: That's a compelling vision. For now, though, the Deep Activity Model stands on its own as a significant advance in synthetic human mobility generation.
Tom: We've enjoyed breaking this one down. Thanks to Lu and Meng for joining the conversation. And to our listeners, if you're working on transportation modeling, urban planning, or any field that needs realistic synthetic populations, this paper is worth your time.
Jane: Until next time, keep thinking about how people move, and why.
Tom: See you on the next episode.
Xishun Liao, Qinhua Jiang, Brian Yueshuai He, Yifan Liu, Chenchen Kuai, Jiaqi Ma
University of California, Los Angeles · University of Louisville · Texas A&M University
cs.LG, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 69/100
The gist: The paper introduces the Deep Activity Model, a generative deep learning approach for synthesizing human mobility patterns.
Key concepts
- Activity-Based Models (ABMs)
- Traditional, complex models used by transportation planners to predict city movement. They are extremely expensive and time-consuming to build, often requiring years of survey data and calibration.
- Transformer Architecture
- The deep learning engine behind large language models. Here, it is adapted to predict a person's sequence of activities (like predicting the next word), taking inputs like age, job, and household members.
- Soft Label Loss
- A technique used in training that gives partial credit for time predictions rather than requiring exact matches. This helps the model learn general temporal patterns without overfitting to minor survey noise.
- Synthetic Population
- A generated group of people whose activities and movements are modeled by the AI. These populations are used to simulate realistic traffic flow and test planning scenarios.
Terminology
Summary
The paper introduces the Deep Activity Model, a generative deep learning approach for synthesizing human mobility patterns. The authors state: We propose a generative Transformer model that uses socio-demographic and household attributes to synthesize daily activity chains, with a location module assigning spatial zones to produce complete daily trajectories.
Motivation and Limitations of Existing Approaches: The authors highlight that existing deep learning models tend to overlook the semantic interdependencies among activities and households and rely on restricted GPS data,
while activity-based models (ABMs) depend on rigid assumptions and extensive data, making them costly and difficult to adapt to new regions.
They note that ABMs, developed in the late 1990s, "have significant limitations: data collection, model development, and calibration are costly and time-consuming; their complexity leads to high computational demands; and they rely on numerous assumptions, limiting their adaptability to other regions. Additionally, current DL models
often rely on specialized mobility data sources like GPS, social media, and communication records, which
differ in format and availability, making cross-dataset and regional adaptation challenging."
Core Approach: The model interprets human mobility patterns as chains of activities (i.e., travel demand) and locations that individuals visit, influenced by socio-demographic factors and environmental conditions.
Unlike previous approaches that focus on prediction, our model emphasizes generation, producing synthetic and realistic human mobility patterns based solely on input socio-demographic profiles.
The model is trained on the 2017 National Household Travel Survey (NHTS), which includes over 129,600 US households, which includes demographics, activity patterns, and travel behaviors for each household member.
The activity types in NHTS are aggregated to 15 categories (e.g., Home, Work, School, Buy goods, Exercise, Religious, etc.).
Model Architecture: The architecture is a Transformer-based encoder-decoder architecture.
The encoder receives combined embeddings of personal, household, and prior activity information,
with padding and causal masks applied. The decoder takes the combined activity sequence and the output (memory) from encoder as additional context.
The model uses a unique data concatenation strategy that combines embedded social-demographic data, data of other household members, and embedded activity data within the time domain,
using learnable delimiters `` for separation. The model is designed to accommodate households of up to five members, a decision informed by statistical analysis of the NHTS dataset, which reveals that 95% of households do not exceed this size.
A total of 26 socio-demographic attributes are used (13 personal and 13 household shared attributes), including gender, age, employment status, household income, household size, vehicles owned, and population density.
Loss Functions: The final loss combines five terms: cross-entropy loss for activity type prediction, soft label loss for start and end time prediction (allowing the prediction results to deviate within a small window
), temporal order loss (ensuring the predicted end time of an activity does not precede its start time
), and sequential timing loss (ensuring the end time of a preceding activity does not exceed the start time of the subsequent activity
).
Data Balancing: The authors developed a multivariate, multi-objective data balancing technique
to address imbalances in HTS data, since most people follow similar activity patterns, which can overshadow less common activities and lead to biased models favoring the majority class.
The algorithm iteratively calculates weights for each data sample using the raking algorithm, setting target distributions between the observed and uniform distributions, so that very common patterns are down weighted and rare activities are up weighted without flattening the empirical distribution.
Model Transfer: The model supports transfer learning: The proposed Deep Activity model, initially trained on the NHTS dataset (160,000 training samples), serves as a generic pre-trained model, which can then be fine-tuned using the limited local HTS data.
For California and Puget Sound, the model reuses the NHTS encoder and decoder with fine-tuning. For Mexico City, which has only 40 of the 60 input attributes used in NHTS and a different activity type with 10 locally defined activity categories,
the model applies all three transfer steps, including adding new layers and replacing the final activity classification layer.
Activity Location Assignment (ALA): To address the limitation that precise location information is often missing
in travel survey data, the authors propose an ALA method that assigns zone-level locations Z for each predicted activity by considering the distribution of distances and angular deviations between preceding and subsequent activities.
The method assigns locations for mandatory activities (work/school) based on commute distance distributions, then assigns non-mandatory activity locations based on distance and angular deviation distributions, with iterative refinement to match ground truth spatial distributions across sub-regions.
Key Results:
-
Activity Generation Performance: The proposed model outperforms all baselines (GRU, RNN, LSTM, decoder-only transformer, GPT-4, LLaMA3, DeepSeek) across most metrics, achieving the lowest JSD values for activity chain length (0.002), duration (0.002), start time (0.003), and end time (0.003), and the highest edge completeness (92.2%).
-
Ablation Study: Removing the soft label loss has
the most significant impact across all metrics,
underscoring its critical role in capturing temporal patterns. Removing temporal order loss and sequential timing loss also degrades performance. -
Contextual Variation: The model
effectively captures
weekday-weekend differences and age-related differences (young 0-18, middle-aged 19-65, elderly 65+) in activity patterns. -
Household Interdependency: Attention heatmaps reveal
the interdependencies among the person's household and their activities.
A case study shows that removing a child from the household input featuresleads to changes in both the attention heatmap and the generated activity chain of the target individual.
-
Transferability: Fine-tuned models for California, Puget Sound, and Mexico City achieve accuracy
comparable to those obtained using the full NHTS dataset,
demonstrating the model's effectiveness in capturing diverse regional mobility patterns. -
System-Level Validation: Using 100,000 agents from SCAG ABM for transfer learning and applying to 1 million agents in LA County, the ALA achieves
a cosine similarity of 0.997
for OD matrices compared to SCAG ABM, with network-level VMT MAPE of 4.97% and traffic speed MAPE of 1.16%. At the corridor level on I-405, compared to Caltrans PeMS observations, the simulated traffic achievesMAPE of 5.85% for traffic volume and 4.36% for speed.
-
Model Complexity Analysis: For the NHTS dataset size,
models with fewer layers generally performed better,
andincreasing the number of attention heads had only a moderate effect on performance.
The Transformer showed improved performance with increased data size, while LSTM performance plateaued.
Contributions: The authors list four key contributions: (1) introducing a learning-based generative framework that integrates travel demand generation and travel trajectory creation while incorporating regional mobility characteristics,
ensuring privacy by relying on aggregated survey data; (2) developing a deep learning model that generates synthetic human mobility data based solely on socio-demographic and household characteristics
that is transferable across regions; (3) explicitly modeling the interdependencies among activities of household members
; and (4) exploring a standard technique for multivariate, multi-objective data balancing to process the ubiquitous HTS data.
Future Work: The authors note opportunities for adaptive temporal loss weighting and activity-type-specific temporal priors to improve performance in underrepresented time windows,
addressing potential sampling biases in NHTS data, and incorporating point of interest level location data with detailed activity and socio-demographic information.
They also acknowledge the potential of alternative generative approaches such as CVAEs, GANs, and diffusion models, which they plan to explore in a dedicated follow-up study.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
Improvement: Implement a Transformer-based encoder-decoder architecture that auto-regressively generates daily activity chains (activity type, start time, end time) conditioned on 26 socio-demographic attributes of the target person AND all household members (up to 5), using learnable `` delimiters to fuse heterogeneous inputs.
What the improved system can do:
-
Generate realistic, privacy-preserving synthetic daily activity chains for any individual given only their demographic profile and household composition, without requiring any historical trajectory data
-
Capture household-level coordination effects (e.g., a father's
BuyMeal
activity is influenced by the presence of a child, as shown in attention heatmaps) -
Produce unlimited diverse activity chains by injecting randomness during inference, enabling large-scale population synthesis for regions with sparse data
Improvement: Use the NHTS-trained model as a pre-trained foundation, then apply a three-step transfer learning protocol: (1) add new layers for region-specific features, (2) freeze pre-trained layers, (3) fine-tune with local data while gradually unfreezing. This handles both small datasets (Puget Sound: 7,791 chains) and structurally different datasets (Mexico City: 40 attributes vs. 60, 10 activity types vs. 15).
Improvement: Implement the iterative raking-based balancing algorithm that simultaneously balances three target features (activity type, chain length, duration) by computing sample weights, adjusting target distributions toward uniform, and re-weighting until convergence.
Improvement: Replace hard one-hot cross-entropy for start/end time prediction with a soft-label loss that assigns non-zero probability to adjacent 15-minute time bins within a window, plus two penalty losses: (a) temporal order loss ensuring end time ≥ start time, and (b) sequential timing loss ensuring previous end time ≤ next start time.
Improvement: Implement the iterative location assignment algorithm that: (1) assigns mandatory activity zones based on home-to-work/school distance distributions per sub-region, (2) assigns non-mandatory zones using distance and angular deviation distributions relative to anchor points, (3) iteratively adjusts reference distributions until sub-region activity counts match ground truth.
Improvement: Apply the empirical finding that for datasets of 180,000 samples, Transformer models with fewer layers (e.g., 2 encoder + 2 decoder layers) outperform deeper models (4+ layers), and that increasing attention heads has only moderate impact. Use this to guide architecture selection.
Improvement: Leverage the model's ability to condition on day-of-week and age attributes to generate distinct activity patterns, validated by distributional analysis showing later weekend starts, fewer work activities, and age-specific patterns (school peaks for youth, midday peaks for elderly).
Abstract
Human mobility plays a crucial role in transportation, urban planning, and public health, but current approaches face important limitations. Existing deep learning models tend to overlook the semantic interdependencies among activities and households and rely on restricted GPS data, while activity-based models depend on rigid assumptions and extensive data, making them costly and difficult to adapt to new regions, especially those with limited conventional travel data. To address these limitations, we propose a generative Transformer model that uses socio-demographic and household attributes to synthesize daily activity chains, with a location module assigning spatial zones to produce complete daily trajectories. Trained on open-source and widely available household travel survey data and then fine-tuned with local data, the model captures national activity patterns and transfers effectively to California, Washington, and Mexico City. This approach offers potential for advancing synthetic human mobility modeling and provides urban planners and policymakers with improved tools for simulating transportation systems and supporting decisions in urban development and public health. Its practical utility is demonstrated through large-scale traffic simulations in Los Angeles County. Compared to the SCAG Activity-Based Model (ABM), the proposed method produces consistent spatial demand patterns, achieving an activity location cosine similarity of 0.997 and a network-level vehicle-miles-traveled Mean Absolute Percentage Error (MAPE) of 4.97%. Compared to real-world observations from Caltrans PeMS, the simulated traffic achieves MAPE of 5.85% for traffic volume and 4.36% for speed on California's I-405 corridor, presenting the practicality of learning-based synthetic mobility for regional simulation and planning.
Sources
- Reconstructing Human Mobility Pattern: A Semi-Supervised Approach for Cross-Dataset Transfer Learning
- Deciphering Human Mobility: Inferring Semantics of Trajectories with Large Language Models
- Mobility-LLM: Learning Visiting Intentions and Travel Preferences from Human Mobility Data with Large Language Models
- TrajLLM: A Modular LLM-Enhanced Agent-Based Framework for Realistic Human Trajectory Simulation
- How Powerful are Decoder-Only Transformer Neural Models?
- GPT-4o System Card
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks