A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds
summary
The gist
Preferred Speed (W1), Slow-to-Stop (W2), Slow (W3), and Fast (W4).
This episode discusses
- A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds · Paper Radio
The paper
A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds · Read on arXiv
Robyn Larracy, Angkoon Phinyomark, Ala Salehi, Eve MacDonald, Saeed Kazemi, Shikder Shafiul Bashar, Aaron Tabor, Erik Scheme
University of New Brunswick
Gait refers to the patterns of limb movement generated during walking, which are unique to each individual due to both physical and behavioral traits. Walking patterns have been widely studied in biometrics, biomechanics, sports, and rehabilitation. While traditional methods rely on video and motion capture, advances in plantar pressure sensing technology now offer deeper insights into gait. However, underfoot pressures during walking remain underexplored due to the lack of large, publicly accessible datasets. To address this, we introduce the UNB StepUP-P150 dataset: a footStep database for gait analysis and recognition using Underfoot Pressure, including data from 150 individuals. This dataset comprises high-resolution plantar pressure data (4 sensors per cm-squared) collected using a 1.2m by 3.6m pressure-sensing walkway. It contains over 200,000 footsteps from participants walking with various speeds (preferred, slow-to-stop, fast, and slow) and footwear conditions (barefoot, standard shoes, and two personal shoes), supporting advancements in biometric gait recognition and presenting new research opportunities in biomechanics and deep learning. UNB StepUP-P150 establishes a new benchmark for plantar pressure-based gait analysis and recognition.
DOI: 10.1038/s41597-025-05792-1
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds".
Jane: The paper was written by Robyn Larracy, Angkoon Phinyomark, Ala Salehi, Eve MacDonald, Saeed Kazemi et al. from University of New Brunswick.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and First Impressions: Tom: Welcome back to the show, everybody. Today we’re cracking open a brand new paper from arXiv, and it’s called “A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds.” Jane, I have to say, just reading that title got me excited, because we finally have a big, public dataset for foot pressure.
Jane: Tom, I’m right there with you. Plantar pressure just means the pressure under your foot, and this paper is all about capturing that in incredible detail. They used a walkway that’s covered in sensors, and it’s not just a tiny mat either. This thing is over three meters long, so people can take multiple natural steps in a row.
Tom: Right, and that’s the part that got me. Most of the older datasets, they had people stepping on a single plate or a very short mat, so you’d get maybe one or two footsteps. Here, they’ve got a one point two by three point six meter grid, which means they’re catching four to six steps per pass. That’s a huge leap forward for studying how people actually walk.
Jane: And the resolution is wild. We’re talking about four sensors per square centimeter. That’s like having a tiny grid of pressure detectors every five millimeters. For comparison, some of the older public datasets had sensors that were maybe five times bigger, so you’d lose all the fine detail of how your heel strikes or how your toes push off.
Tom: The team behind this is from the University of New Brunswick, and they went all in. They recruited one hundred fifty people, and they had them walk under four different footwear conditions—barefoot, a standard shoe they provided, and then two pairs of the participant’s own shoes. Plus, they varied the walking speed: preferred, slow, fast, and even a slow-to-stop condition.
Jane: That slow-to-stop one is clever, because that mimics real life, like when you’re approaching a door or a security checkpoint. You don’t just walk at a constant speed all day. So having that in the dataset makes it much more realistic for building systems that work in the real world.
Tom: And the sheer volume of data is staggering. Over two hundred thousand footsteps. The previous biggest dataset of this kind had about twenty thousand so this is a tenfold increase. That’s the kind of scale you need if you want to train modern deep learning models, which are hungry for data.
Jane: Exactly. And they didn’t just dump raw sensor readings on us. They also provide preprocessed data, where each footstep is cut out, aligned, and normalized. That saves researchers weeks of work just cleaning the data, so they can jump straight into the science.
Tom: I love that they thought about the user. They give you both raw trial files and these segmented footsteps, plus a spreadsheet with all the participant demographics. So you can study how age, sex, body size, or shoe type changes the pressure patterns.
Jane: And that’s the hook for our next segment, because it’s not just about having a big pile of data. It’s about what you can actually do with it. We’re going to dig into the methods and how they made sure all those footsteps were labeled correctly.
Tom: Stay with us, because this dataset could change how we think about gait recognition and biomechanics.
Summary and Methodology: Tom: Welcome back. We’re still on “A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds,” and Jane, I want to get into the nitty-gritty of how they actually built this thing, because the summary in the paper is dense.
Jane: It is, but the core idea is simple. They set up this long walkway with pressure tiles, and they had people walk back and forth for ninety seconds at a time. But the clever part is how they turned that continuous stream of pressure frames into individual, labeled footsteps.
Tom: Right, because you can’t just look at one frame and say, "that’s a foot." They used a tracking algorithm called SORT, which is usually used for tracking objects in video. Here, they applied it to track blobs of pressure across time, so they could follow a foot from heel strike to toe-off.
Jane: And that gives you a three dee bounding box for each step, with time, height, and width. That’s the raw extraction. But then they had to figure out which foot it was, left or right, and which direction the person was walking. They used the center of pressure trajectory for that, which is basically the weighted average of where the pressure is at each moment.
Tom: I read that they got the left/right classification right ninety-nine point seven percent of the time with an automated algorithm. But they didn’t stop there. They had two human reviewers visually inspect every single footstep. That’s a massive quality control effort, and it’s exactly what you need for a benchmark dataset.
Jane: They even used the video recordings they took from seven cameras around the room to double-check uncertain cases. So if the algorithm flagged something weird, they could look at the video and see what actually happened. That’s how you catch the edge cases, like when someone shuffles their feet during the slow-to-stop trials.
Tom: And they thought about what happens when people step partially off the mat. Those are called incomplete footsteps, and they’re flagged in the metadata. They also flagged standing footsteps, which happen during that slow-to-stop condition, and they used an outlier detection method to flag any steps that just look abnormal.
Jane: That outlier detection is interesting. They used something called an R-score, which basically compares each footstep to the median footstep in that trial. If a step is too different in terms of size, duration, or force profile, it gets flagged as an outlier. That way, researchers can easily filter out the noisy steps if they want to train a clean model.
Tom: And here’s the thing I really appreciate. They provide two different preprocessing pipelines. Pipeline one keeps the foot size and rotation information, so you can still use that for recognition. Pipeline two resizes everything to a common foot size and normalizes the amplitude, which is better for comparing pressure patterns across people.
Jane: Because there’s no single right way to preprocess pressure data. If you resize everyone’s foot to the same size, you lose information about foot size, which might be useful for identifying someone. But if you keep the original size, you can’t directly compare pressure at the same anatomical location across people. So they give you both options and a script to generate your own.
Tom: That flexibility is huge. It means researchers can test different preprocessing choices and see which one works best for their specific task, rather than being locked into one format. And that brings us to the next segment, because I want to talk about what improvements this enables and what new research doors it opens.
Jane: Absolutely, because a dataset like this isn’t just a bigger version of what we had. It’s a different beast entirely.
Improvements and Implications: Tom: So we’ve established that this dataset is big and well-annotated. But Jane, the real question is, what can we do with it that we couldn’t do before? What improvements does “A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds” actually enable?
Jane: The biggest one is that we can finally train deep learning models properly. Before, with only twenty thousand footsteps, you’d overfit quickly. Now, with two hundred thousand you can train convolutional neural networks or transformers to recognize individuals just from their foot pressure patterns, and actually expect them to generalize.
Tom: And the diversity of the data matters just as much as the volume. They have people from nineteen to ninety-one years old, a balanced mix of men and women, and a wide range of body sizes and foot shapes. That’s critical for building biometric systems that don’t fail on people who aren’t in the training set.
Jane: But it’s not just about recognition. This dataset lets us study how footwear changes your gait. They have the same person walking barefoot, in a standard shoe, and in two of their own shoes. So you can isolate the effect of the shoe itself versus the person’s natural gait.
Tom: That’s a point I want to push on, because the paper shows some really interesting visualizations. They have peak pressure images, and you can see how a steel-toe work boot creates a completely different pressure pattern than a stiletto heel. But the person underneath is the same, so you can start to separate the identity component from the footwear component.
Jane: And that’s the hard problem in gait recognition. When someone changes shoes, their pressure pattern changes, but there’s still something consistent about how they walk. With this dataset, you can actually train models to be invariant to footwear, which is what you need for a real security system.
Tom: There’s also the walking speed variation. They have slow, preferred, fast, and that slow-to-stop condition. The paper shows that walking speed changes the ground reaction force profile significantly, especially the peak force at heel strike. So if you want a system that works in the real world, it has to handle people walking at different speeds.
Jane: And the implications go beyond security. This could be used in rehabilitation. If you have a patient with a foot injury, you can track their pressure patterns over time and see if they’re improving. Or you could use it to design better shoes, because you can see exactly where the pressure is concentrated.
Lu: I’d like to jump in here, Tom. The scale of this dataset also opens up the possibility of generative models. You could train a model to synthesize realistic foot pressure patterns for people who aren’t in the dataset. That could be used to augment training data even further, or to simulate how a new shoe design would affect gait before you even build a prototype.
Meng: And from an engineering standpoint, the fact that they provide both raw and preprocessed data is a godsend. We can test our own preprocessing pipeline against theirs and see if it makes a difference. Plus, the metadata includes spatiotemporal parameters like step length and step width, which are useful for feature engineering in traditional machine learning.
Tom: That’s a great point, Meng. The paper even compares step length and step width across datasets, and they show that the CASIA-D dataset had participants that were almost perfectly separable using just those two features. But in this new dataset, there’s much more overlap, which means the recognition problem is harder and more realistic.
Jane: So it’s not just a bigger dataset. It’s a harder dataset, which is exactly what we need to push the field forward. And that leads us to our final thoughts.
Conclusion: Tom: We’ve spent the whole show on “A dataset of high-resolution plantar pressures for gait analysis across varying footwear and walking speeds,” and I think we’ve only scratched the surface. Jane, what’s the big takeaway for our listeners?
Jane: The big takeaway is that this dataset is a game-changer for anyone studying how we walk. It’s the largest, most detailed public dataset of underfoot pressure ever released, with one hundred fifty participants and over two hundred thousand footsteps. It covers multiple footwear types and walking speeds, and it’s meticulously annotated.
Tom: And it’s not just for biometrics. It’s for biomechanics, sports science, rehabilitation, shoe design, and even generative AI. The fact that they provide raw data, preprocessed data, and a flexible preprocessing script means that researchers can adapt it to almost any task.
Jane: I also want to highlight the effort they put into quality control. Two human reviewers checked every footstep, and they used video to confirm uncertain labels. That level of care is rare, and it makes the dataset trustworthy as a benchmark.
Tom: And the authors made it easy to use. It’s available on the Federated Research Data Repository, with both Python and MATLAB formats. So whether you’re a deep learning researcher or a biomechanist, you can start using it right away.
Lu: I’ll add that this dataset will likely become the standard benchmark for footstep recognition. The previous datasets were too small to train modern models, so this fills a critical gap. I expect to see a wave of new papers using this data within the next year.
Meng: And from a practical standpoint, the inclusion of the slow-to-stop condition is brilliant. That’s the exact scenario you’d encounter at a secure access point, so it makes the dataset directly relevant for deployment.
Tom: Alright, we’ve said our goodbyes to this paper. It’s been a pleasure, and we’re already looking forward to the next one. Thanks for listening, and we’ll see you on the next episode.
Jane: Goodbye, everyone. Keep walking, and keep measuring.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization