Design and Embedded Validation of Compact ML Models for Affective Touch Classification in a Soft Interactive Companion
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Design and Embedded Validation of Compact ML Models for Affective Touch Classification in a Soft Interactive Companion".
Jane: The paper was written by Aleksandrs Vališevskis, Aleksandrs Okss, Inese Tīģere, Aleksejs Kataševs, Dina Bethere et al. from Institute of Architecture and Design, Riga Technical University and Institute of Mechanical and Biomedical Engineering, Riga Technical University and Center for Pedagogy and Social Work, Riga Technical University Liepaja Academy and Institute of Digital Humanities, Riga Technical University and Masaryk University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Jane, I think I need a glass of water just to finish reading this title.
Jane: You mean "Design and Embedded Validation of Compact ML Models for Affective Touch Classification in a Soft Interactive Companion"?
Tom: Exactly, it is quite the mouthful.
Jane: It sounds intimidating, but the core idea is actually really heartwarming.
Tom: Are you talking about making stuffed animals that can feel emotions?
Jane: Well, they can feel how we touch them, like a hug or a scratch, and respond accordingly.
Tom: This research comes out of Riga Technical University, right?
Jane: Yes, Aleksandrs Vališevskis and his team are looking at how to put this intelligence inside a soft, plush toy.
Tom: I love that they are focusing on soft companions instead of those scary, rigid metal robots.
Jane: It makes sense because a soft toy is much safer for a child to interact with.
Lu: Imagine a plushie that doesn't just sit there, but actually understands your mood through your touch.
Tom: That sounds like something straight out of a sci-fi movie, Lu.
Lu: It could be so much more than a toy; it could be a reactive, living presence in a room.
Meng: I wonder how they actually fit all that math into something as small as a stuffed animal.
Jane: That is the "compact" part of the title, Meng.
Meng: They aren't just running it on a giant computer in the cloud, are they?
Tom: No, they want this to run directly on the toy's own tiny internal brain.
Jane: It is all about making the artificial intelligence small enough to live inside the toy itself.
Lalam: This could change how we support people with autism by providing a non-verbal way to connect.
Tom: You think the emotional connection would be that strong, Lalam?
Lalam: If the toy can recognize a gentle stroke versus a rough hit, it creates a sense of mutual understanding.
Jane: It moves us toward a world where technology feels more like a companion and less like a gadget.
Tom: Let's see if the actual data supports this big vision.
Summary: Tom: We are moving from the big picture into the actual numbers of this study.
Jane: They collected a massive amount of data to make sure the toy actually learns correctly.
Tom: They used one thousand three hundred twenty-six different gesture sequences, right?
Jane: That is correct, and they got those from twenty-five different people.
Tom: And it wasn't just adults; they included teenagers and even kindergarten-aged children.
Jane: That diversity is huge because a child's touch is very different from an adult's.
Lu: The way they organized that dataset is a work of art for researchers.
Tom: They used something called a 1D CNN to process all those touches, didn't they?
Jane: Yes, a one-dimensional convolutional neural network is perfect for reading signals that change over time.
Tom: And they managed to make the model incredibly tiny.
Jane: It only has about thirteen thousand two hundred parameters.
Tom: That is tiny compared to the massive models we usually hear about.
Jane: It achieved seventy-five percent accuracy on their tests, which is quite impressive for something so small.
Meng: I am looking at the hardware requirements they mentioned for the ESP32 microcontroller.
Tom: What caught your eye there, Meng?
Meng: They estimated it needs about three point two million multiply-accumulate operations per window.
Jane: That sounds like a lot of math for a little chip.
Meng: It is, but they say it can still run in real-time at twenty Hz.
Tom: So the toy can react almost instantly when you touch it.
Lalam: Running everything locally on the chip is the smartest move for privacy.
Jane: You mean because the touch data never leaves the toy?
Lalam: Exactly, the personal interactions stay between the human and the companion.
Tom: That makes the whole thing feel much more secure and intimate.
Jane: Let's look at how they actually improved the system compared to what existed before.
Improvements: Tom: The researchers weren't just starting from scratch; they were fixing a broken system.
Jane: The previous version of this toy used a simple "heuristic" method, which is just a set of basic rules.
Tom: Like, if the sensor hits a certain number, then do this?
Jane: Exactly, but it couldn't tell the difference between a complex hug and a simple hold.
Tom: So it was basically guessing when things got subtle.
Jane: It was, and the researchers found that the new CNN model is much better at those nuanced gestures.
Tom: But they didn't just throw the old rules away, did they?
Jane: No, they actually proposed a "hybrid" strategy.
Tom: A hybrid strategy sounds like they are using the best of both worlds.
Jane: They use the fast, simple rules to catch big, energetic things like a hit or a pull.
Tom: And then the smart CNN takes over for the gentle stuff like scratching or stroking?
Jane: Precisely, the CNN handles the subtle social touches that require more thought.
Lu: That is such a clever way to balance speed and intelligence.
Meng: It's a great engineering decision because it saves power.
Tom: How does it save power, Meng?
Meng: The chip doesn't have to run the heavy math for every single tiny movement.
Jane: It only kicks in the complex processing when it needs to interpret something meaningful.
Tom: That is a massive improvement over just running a heavy model constantly.
Lalam: It creates a much more natural social rhythm for the interaction.
Jane: It prevents the toy from being overwhelmed by random noise.
Lalam: A toy that reacts too slowly or too incorrectly loses the emotional bond immediately.
Tom: This hybrid approach seems like the real secret sauce here.
Jane: It really is, and it sets a high bar for future smart companions.
Conclusion: Tom: We have covered a lot of ground with this paper.
Jane: From the tiny thirteen thousand two hundred parameter models to the way they can help children with autism.
Tom: It is amazing how much intelligence you can pack into a soft plush toy.
Jane: And they've even shared their dataset and software so others can build on this.
Lu: I can see these companions becoming part of every household to help with emotional regulation.
Meng: From my side, seeing this work on an ESP32 proves that edge AI is ready for the real world.
Lalam: It shows that technology can be soft, private, and deeply human all at once.
Tom: We are moving on to the next paper now, but this one was a standout.
Jane: We'll be watching for that follow-up study on the actual clinical trials.
Tom: Thanks for joining us to discuss "Design and Embedded Validation of Compact ML Models for Affective Touch Classification in a Soft Interactive Companion."
Jane: See you next time!
Institute of Architecture and Design, Riga Technical University · Institute of Mechanical and Biomedical Engineering, Riga Technical University · Center for Pedagogy and Social Work, Riga Technical University Liepaja Academy · Institute of Digital Humanities, Riga Technical University · Masaryk University
cs.AI
Submitted: 2026-04-16
Updated: 2026-09-12
Comments: 25 pages, 11 figures
Code: https://github.com/makeabilitylab/arduino
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: This paper presents a "complete open-source MATLAB-based framework" for the development and validation of compact deep learning models designed to recognize affective touch in soft, sensorized
Key concepts
- 1D CNN
- A one-dimensional convolutional neural network is a type of model perfect for reading signals that change over time. In this study, it is used to process touch data, allowing a small, compact model to identify various human gestures within a soft, interactive companion.
- Hybrid Strategy
- This approach combines simple, rule-based heuristic methods with advanced machine learning. The toy uses fast, basic rules to catch large, energetic movements like hits or pulls, while the smart CNN takes over to interpret subtle, social touches like scratching or gentle stroking.
- Edge AI
- This involves running artificial intelligence directly on a device's own internal hardware, such as an ESP32 microcontroller, rather than using a giant computer in the cloud. This enables real-time interaction, saves power, and ensures privacy because personal touch data never leaves the toy.
Terminology
Summary
This paper presents a complete open-source MATLAB-based framework
for the development and validation of compact deep learning models designed to recognize affective touch in soft, sensorized interactive companions. It addresses the technical challenges posed by the deformability and multichannel tactile sensing
of plush devices, providing a pathway for emotionally meaningful, privacy-preserving touch interpretation
to be implemented directly within therapeutic hardware.
Data Collection and Preprocessing
The study utilizes a diverse FAIR-compliant dataset of 1326 labelled gesture sequences
collected from 25 participants, including children, teenagers, and adults. The sensor suite integrated into the plush companions includes multiple capacitive sensors located in the extremities, ears, head, neck, nose, tail, back, and belly, alongside a 3-axis accelerometer. The recorded gestures are categorized into three emotional types:
-
Positive interactions (e.g., stroking or scratching).
-
Neutral interactions (e.g., holding or resting).
-
Negative interactions (e.g., hitting or pulling the tail).
Data is processed through a modular workflow involving anonymization and organization,
quality control
to flag outliers, and truncation and alignment
to remove initial latency caused by human reaction time. The signals are then segmented using a sliding window of 250 samples (2.5 s at 100 Hz)
to expand the dataset for improved model generalization.
Model Development and Optimization
To move beyond heuristic threshold-based algorithms
that often misclassified complex gestures, the researchers implemented a One-Dimension Convolutional Neural Network (1D CNN) optimized for real-time embedded execution.
Through systematic exploration of 468 models across 13 different architectures, the study identified that compact dilated one-dimensional convolutional neural networks (1D CNNs)
were the most effective solution.
The optimization process involved testing various hyperparameters to balance predictive performance with computational load:
-
Filter size sets ranging from compact to large models.
-
Epoch counts (100, 250, and 1000) to evaluate overfitting.
-
Training mini-batch sizes.
-
Architectural variations, specifically utilizing
dilated convolutions
to expand the receptive field without increasing the number of parameters.
Results and Validation
The research identifies a significant trade-off between accuracy and complexity, noting that increasing the number of parameters even by orders of magnitude, does not increase the performance of the model substantially.
A highly compact model with only 13.2k-parameter[s]
achieved 75% test accuracy and 85% mean leave-one-subject-out cross-validation accuracy,
demonstrating robust inter-subject generalization
to unseen users.
In real-time PC-based simulations, the CNN demonstrated a clear advantage in resolving subtle social touches
that previous heuristic systems failed to detect. However, the study found that high-force negative interactions are captured more reliably by trivial threshold-based logic.
Consequently, the authors propose a hybrid inference pipeline
consisting of:
-
Instantaneous heuristic filtering for high-energy events.
-
CNN-based nuanced gesture classification for complex touches.
Embedded Feasibility
The study concludes with a theoretical analysis of the model's suitability for deployment on an ESP32-class microcontroller.
The selected compact models are highly efficient, occupying only approximately 14-53KB of memory
before quantization.
To ensure real-time performance, the authors calculated the computational requirements for a single inference:
-
The model requires approximately 3.2 MMAC per window.
-
Quantized deployment is estimated to be compatible with
20 Hz real-time operation on the target microcontroller.
This demonstrates that emotionally responsive interactive systems
can operate autonomously without external computing resources or network connectivity,
which is critical for reliability and privacy in therapeutic settings.
Improvements for AI systems
-
The Improvement: Implement a dual-stage processing architecture that uses a low-latency, rule-based heuristic layer as a
first responder
followed by a deep learning classifier. -
What the improved system can do: It can provide near-instantaneous reaction times (< 200ms) for high-energy, abrupt, or negative interactions (e.g., hits or violent pulls) via simple thresholding, while simultaneously utilizing a 1D CNN to resolve subtle, low-amplitude affective gestures (e.g., stroking or scratching) that require temporal context. This optimizes both responsiveness and computational efficiency.
-
The Improvement: Replace standard convolutional layers with dilated 1D convolutions in lightweight neural architectures designed for microcontrollers (MCUs).
-
What the improved system can do: It can expand the temporal receptive field to capture long-duration gestures (e.g., a 2.5s stroke) without increasing the parameter count or memory footprint. This allows an extremely compact model (e.g., about 13k parameters) to achieve high-accuracy gesture recognition on resource-constrained hardware like the ESP32, enabling real-time, on-device affective computing without cloud dependency.
-
The Improvement: Shift from random window-level splitting to a Leave-One-Subject-Out Cross-Validation (LOSO-CV) protocol during the model validation and optimization phase.
-
What the improved system can do: It can provide a robust guarantee of performance when encountering entirely new users (e.g., children with ASD in a clinical setting) who were not present in the training dataset, preventing the AI from overfitting to specific individual tactile signatures or hand sizes.
-
The Improvement: Utilize an ACR-based selection metric—balancing classification accuracy against the total number of learnable parameters—rather than optimizing for raw accuracy alone.
-
What the improved system can do: It can identify
sweet spot
architectures that provide maximum predictive power within a strict memory and power budget, ensuring the model fits within the limited Flash/RAM of an embedded system while maintaining sufficient accuracy (e.g., >70%) for reliable social interaction. -
The Improvement: Integrate multichannel capacitive touch data, proximity sensing, and 3-axis inertial (accelerometer) magnitude into a unified 1D CNN input stream.
-
What the improved system can do: It can distinguish between contextually similar gestures by correlating spatial touch patterns with physical movement (e.g., differentiating a
back stroke
performed while the device is on a table versus while it is being held in hands), significantly reducing class confusion in complex, real-world interactions.
Abstract
Soft plush companions provide a safe and intuitive platform for affective human-robot interaction, but their deformable structure and distributed tactile signals make reliable gesture recognition difficult. This study presents a complete workflow for developing and validating compact affective-touch classifiers for an interactive plush companion. A newly collected dataset comprised 1,326 labelled recordings before curation, including interactions from 25 children, teenagers, and adults. Each classifier received 2.5-s windows containing ten capacitive channels and one accelerometer-magnitude channel. MATLAB supported acquisition, quality control, window generation, and a 468-run exploratory study of dilated one-dimensional convolutional neural networks (1D CNNs). A closely matched Python workflow then preserved participant provenance, fitted preprocessing inside each fold, and evaluated shortlisted models by 25-fold leave-one-subject-out cross-validation (LOSO-CV). On the operational 10-class task, the compact dilated CNN achieved 81.96% mean macro-F1 and 85.42% mean accuracy. The depthwise-separable CNN achieved the strongest neural result (84.50% macro-F1), whereas a linear support-vector machine using 66 predefined time-domain features achieved the best overall result (87.98% macro-F1 and 91.05% mean accuracy). Paired fold analysis showed that the linear support-vector machine and constrained random forest outperformed the compact dilated CNN, whereas differences among the tested neural models were not statistically significant after correction. Direct measurements on a 240-MHz ESP32-S3 confirmed valid deployment of the dilated CNN, depthwise-separable CNN, temporal convolutional network, linear support-vector machine, and random forest; the linear model required a 3.2 kB serialized payload and 1.405 ms mean end-to-end classifier time.
Sources
- Multi-Scale Context Aggregation by Dilated Convolutions
- CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection