Frequency-aware decomposition learning for sensorless wrench estimation in vibration-rich robotic contact

arXiv:2604.12905 · cs.RO, cs.LG · Submitted 2026-04-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Frequency-aware decomposition learning for sensorless wrench estimation in vibration-rich robotic contact".

Dev: Force and torque (F/T) sensing is critical for robot-environment interaction, but physical F/T sensors impose constraints in size, cost, and fragility.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Let's talk about the title and authors of this work now, "Frequency-aware decomposition learning for sensorless wrench estimation in vibration-rich robotic contact." It’s clear right away that they are tackling a very specific problem.

Dev: They're focusing on sensorless wrench estimation in situations where there’s a lot of vibration involved during contact, which is exactly the kind of scenario we struggle with when relying on physical sensors.

Taro: I see the focus is on overcoming those physical sensor limitations by using internal robot states to estimate forces and torques, which is a big step for making robots more versatile in unstructured settings.

Rosa: The authors are proposing this Frequency-aware Decomposition Network as their solution, suggesting it’s not just about estimation but about how the estimation process itself should be structured spectrally.

Dev: Their approach seems to be rooted in time-series forecasting, using historical proprioceptive data to predict future wrench components based on a decomposition strategy.

Taro: The authors are essentially saying that simply feeding raw data into a standard model isn't enough; you need a model that understands the underlying frequency content of the interaction.

Rosa: They’re integrating concepts like spectral decomposition and frequency-awareness directly into the network architecture to better capture those transient events.

Dev: It seems they are aiming for short-term forecasting, which is crucial because we need to know what forces are coming in the immediate future for stable control loops.

Taro: I’m curious if this means their method is useful only for very quick interactions, or if it can handle slower but still complex dynamic events.

Rosa: They specifically validate the framework on real-world grinding excavation with a six-DoF hydraulic manipulator, where the target wrench exhibits substantial high-frequency vibrations arising from rapid contact transients.

Dev: So it’s not just theoretical; they tested it against a concrete example of high-vibration scenarios, which is important for validating the method's practical applicability.

Taro: That real-world validation on grinding excavation gives us a good baseline for understanding its performance in messy, dynamic contact situations.

Rosa: And as we look at the authors, they seem to be coming from a place where deep learning methods are being applied to solve these kinds of complex physical estimation problems.

Dev: I’m interested in seeing what their background is, because the success of any model often depends on how well it understands the underlying physics it’s trying to emulate.

Taro: If their expertise is strong in autonomy, we might see this framework applied to more complex social navigation or manipulation tasks down the line.

Rosa: Ultimately, they are presenting a sophisticated way to use machine learning tools—the FDN—to estimate forces and torques when traditional physical sensors are impractical.

Dev: It’s a very interesting direction for sensorless control, moving beyond just simple force estimation to modeling the dynamics of the interaction.

Taro: So, we’re looking at a system that uses historical state data and learned frequency priors to predict wrench in complex contact situations.

Rosa: That’s right; it's about making the estimation process aware of the frequency content of what's happening during robot interaction.

The paper's summary: Dev: Now that we’ve looked at the title and authors, let’s get into the actual summary of this paper, "Frequency-aware decomposition learning for sensorless wrench estimation in vibration-rich robotic contact." Essentially, they outline their core method.

Rosa: Their central idea is to propose the Frequency-aware Decomposition Network, or FDN. They break down the wrench into a trend component and a residual component over each time step horizon.

Taro: So it’s not just one prediction; it's two separate forecasts that they model differently, which allows them to target different parts of the signal spectrum with different models.

Dev: The trend head uses a low-pass filter, FPFlow, while the residual head models that high-frequency component as a learned conditional distribution rather than just a regression output.

Rosa: That’s right; they use FPFlow to impose a frequency band prior on the trend component and then model the residual using an asymmetric deterministic and probabilistic head.

Taro: It sounds like they are trying to model the low-frequency dynamics with high accuracy while letting the residual handle those tricky, rapidly changing details that we usually miss.

Dev: The input processing involves modality-specific encoders—four for position and velocity, plus one MLP for initial conditions—to process the different state variables effectively.

Rosa: These encoders first enhance time-varying modalities in the frequency domain before applying patch embedding to create a latent representation z in R 6xNxd.

Taro: I’m interested in that frequency enhancement step because it suggests they are trying to make sure the input features are optimally represented across different frequencies before they even enter the main forecasting network.

Dev: And then from that representation, they use two separate linear heads: one for the trend and one for predicting both the mean and log-variance of a Gaussian distribution for the residual.

Rosa: So, by summing those predictions together, they get a total forecast which combines a smooth trend estimate with a probabilistic high-frequency prediction.

Taro: It seems like they’re building a very structured way to handle the inherent uncertainty in these dynamic systems without just throwing everything into one black box.

Dev: That structuring through decomposition and frequency-aware modeling is what makes this paper stand out compared to simpler time-series models for wrench estimation.

Rosa: It really is about providing a detailed mechanism for how a model can learn to separate the predictable motion from the complex, transient forces during robot interaction.

The paper's improvements: Taro: Moving on to what they actually improved in this paper, the authors suggest several key architectural enhancements to make the FDN even more effective.

Rosa: One major suggestion is incorporating frequency-aware layers directly into the network to enhance inputs and refine outputs in the frequency domain, using FPFlow for trends and an FPFhigh filter for residuals.

Dev: That means they are not just applying filters after training; they are designing layers that actively work within the frequency domain to shape what information is being processed.

Taro: I also see a suggestion for a learnable frequency enhancement filter, FEF, which uses a mixture of experts over M learnable filters fm(·) to dynamically select experts based on the input context.

Rosa: That MoE layer before the patch embedding is designed to provide dynamic allocation of computational resources, allowing different "experts" to specialize in amplifying specific frequency bands relevant to the current interaction.

Dev: That sounds like a way to make the input processing much more adaptive, which should help with robustness against varying noise levels during operation.

Taro: These architectural choices seem aimed at improving the separation of signal from noise based on spectral characteristics before feeding it into the main forecasting network.

Rosa: The overall improvement is this asymmetric modeling for low-frequency and high-frequency bands of the wrench signal, which is modeled with deterministic and probabilistic heads.

Dev: So they are addressing that typical trade-off where you either get a good smooth prediction or a good transient prediction, but not both effectively in one model.

Taro: This approach allows them to model the low-frequency dynamics with high accuracy while using the learned conditional distribution for the residual to capture those difficult high-frequency vibrations.

Rosa: The pretraining strategy also provides an improvement by ensuring that their representations transfer well across different robot platforms, reducing the need for extensive task-specific fine-tuning.

Dev: That transfer analysis is valuable because it shows they’re not just overfitting to one specific robot setup but learning general principles of wrench dynamics.

Conclusion: Rosa: So to wrap up on this paper, the main implication is that by using the Frequency-aware Decomposition Network, we get a method for sensorless estimation that explicitly models spectral characteristics of vibration-rich interactions.

Dev: This means we can get a forecast for short-term wrench with better accuracy than what we might achieve with traditional methods when dealing with rapid contact transients.

Taro: The big picture is that this could enable much more precise robotic manipulation in dynamic environments where the ability to sense these transient forces accurately is critical for safe operation.

Rosa: I think it gives us a structured way to handle the inherent uncertainty in these dynamic systems by separating the trend and residual forecasts using deterministic and probabilistic modeling.

Dev: It’s definitely a significant step forward because it provides a detailed mechanism for forecasting that explicitly handles the low-frequency smooth dynamics versus the high-frequency noise.

Taro: I’m really excited about how this framework could be used to create systems that are proactive rather than reactive in response to dynamic forces.

Rosa: Indeed, this Frequency-aware Decomposition Network offers a detailed framework for sensorless estimation that focuses on the spectral characteristics of the interaction itself, and I think it will be a useful foundation for future research.

Dev: We should keep an eye on how they implement those frequency-aware layers in deployment to see if they can maintain stable loop rates in real-world scenarios.

Taro: I just hope that as these models move out of the lab, this level of robustness holds up when faced with the unexpected variability we see in the field.

Kyung Hee University

cs.RO, cs.LG

Submitted: 2026-04-14

Updated: 2026-09-30

Comments: revised. 27 pages, 10 figures, 10 tables. Code: https://github.com/leehyeonbeen/FDN Data: https://doi.org/10.5281/zenodo.23026020

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 67/100

The gist: Force and torque (F/T) sensing is critical for robot-environment interaction, but physical F/T sensors impose constraints in size, cost, and fragility.

Key concepts

Frequency-aware Decomposition Network (FDN)
A neural network designed to predict wrench by splitting it into a smooth trend part and a residual part. It uses frequency filtering techniques to improve input data and constrain the output prediction, allowing it to accurately estimate complex, high-frequency forces.
Wtrend_t+1:tf
The predicted trend component of the wrench at a future time step. This is calculated using a low-pass filter (FPFlow) on historical wrench data. It represents the slow, predictable movement or overall force trend over a short time horizon.
Wres_t+1:tf
The predicted residual component of the wrench at a future time step. This is the difference between the actual measured wrench and its predicted trend component. The model learns to forecast this residual, which captures sudden, high-frequency interactions like grinding.
FPFlow function
A differentiable low-pass filter used in the model to impose a frequency band prior on the trend prediction. It smooths out the input data based on a cutoff frequency (fc), ensuring that the predicted trend component remains within expected frequency limits.

Terminology

Summary

Force and torque (F/T) sensing is critical for robot-environment interaction, but physical F/T sensors impose constraints in size, cost, and fragility. This work proposes a Frequency-aware Decomposition Network (FDN) for short-term forecasting of vibration-rich wrench from proprioceptive history to address the underexplored estimation of high-frequency wrench components induced by rapid interactions like grinding.

The gist: FDN predicts spectrally decomposed wrench with asymmetric deterministic and probabilistic heads, modeling the high-frequency residual as a learned conditional distribution, and incorporates frequency-awareness to adaptively enhance input spectra with learned filtering and impose a frequency-band prior on the outputs.

How it works

The proposed Frequency-aware Decomposition Network (FDN) models the wrench by decomposing it into trend and residual components over each T-step horizon. For a prediction time t, the ground-truth trend and residual sequences are defined using spectral decomposition:

  1. Wtrend t+1:tf = FPFlow(Wt+1:tf)

  2. Wres t+1:tf = Wt+1:tf − Wtrend t+1:tf

The FPFlow function is a differentiable non-recursive low-pass filter defined as F−1(Hl(f; fc) ⊙ F(·)), where Hl(f; fc) = 1/(1 + (f / fc) 2r), imposing a frequency band prior. The model Mθ then learns to forecast these decompositions given the history xth:t, yielding:

Wˆ trend t+1:tf, Wˆ res t+1:tf ∼ Mθ(· xth:t)

Modality-specific encoders

To reduce sensitivity to episode-dependent initial positions, the input vector xt is redefined as xt = [∆qt, q˙t, q¨t, ut, qe0] ∈ R 5n. The model utilizes four modality-specific PatchTST encoders (Enc∆q, Encq˙, Encq¨) and one MLP encoder for the episode-dependent initial positions (Encqe0). These encoders process time-varying modalities by first enhancing them in the frequency domain as described in Section III-D, then applying patch embedding FEF(xδth:t) to a latent dimension D. The final representation z is obtained by concatenating and projecting these embeddings into the 6 wrench channels, resulting in z ∈ R 6×N×D.

Asymmetric forecasting heads

From the final representation z, two separate linear heads are used for forecasting:

  1. Trend Head: W˜ trend t+1:tf = ¯zWtrend + btrend.

  2. Residual Head: µ˜res t+1:tf = ¯zWµ + bµ and vˆ res t+1:tf = ¯zWv + bv, predicting the mean and log-variances of a step- and channelwise Gaussian distribution for the residual. The final forecast is obtained by summing these predictions: Wˆ t+1:tf = Wˆ trend t+1:tf + Wˆ res t+1:tf.

Frequency-awareness

The model incorporates frequency-aware layers to enhance inputs and refine outputs in the frequency domain. The predicted trend and residual are filtered using FPFlow for the trend and FPFhigh for the residual, which acts as a denoising high-pass filter defined as F−1(Hh(f; fc, fdn c) ⊙ F(·)), where Hh is a band-pass response between fc and fdn c. Additionally, a learnable frequency enhancement filter (FEF) is applied to the input sequence xδth:t before patch embedding, which uses a mixture of experts (MoE) over M learnable filters fm(·).

Proprioception-to-Wrench pretraining

The model is pretrained on the RH20T dataset using proprioception and wrench data. Since the dataset covers both 6- and 7-DoF robots, n=7 is used for input construction. To avoid negative transfer from the modality gap between motor-based actuation and downstream hydraulic actuation, the actuation signal u is masked with zero in both xδth:t and FEF(xδth:t), and the corresponding encoder representation zu is zeroed out. The learned proprioception-to-wrench representations are then transferred by initializing Enc∆q, Encq˙, Encq¨, and their embedding layers with pretrained parameters. Fine-tuning is performed end-to-end after a linear probing step.

Evaluation methodology

The model is compared against baselines from force estimation (e.g.

Improvements for AI systems

Here are specific improvements to AI systems, derived from the proposed Frequency-aware Decomposition Network (FDN) for sensorless wrench forecasting, and what these improved systems can achieve:


The core improvement lies in developing a class of predictive models that move beyond simple instantaneous estimation or standard sequence-to-sequence forecasting by explicitly modeling the spectral characteristics (frequency content) of dynamic interactions.

Here are the specific improvements and their capabilities:

  1. Decomposition-Based Probabilistic Modeling:

A system can be improved to decompose wrench forecasts into two distinct components: a smooth, predictable Trend component and a volatile, high-frequency Residual component.

  • Instead of predicting a single trajectory, the model predicts parameters (mean and variance) for these two components separately.

  • The system can now explicitly model the low-frequency dynamics (via the trend head) with high accuracy while using a learned conditional distribution (the residual head) to capture transient, high-frequency vibrations that are otherwise difficult to estimate.

Frequency-Aware Filtering and Prior Imposition:

The AI system can be enhanced with differentiable filtering layers (like the Low-Pass Filter and Denoising High-Pass Filter) that operate directly on the predicted outputs in the frequency domain.

  • The system can adaptively enhance noisy input spectra based on learned frequency priors, effectively suppressing noise outside of known relevant bands while sharpening predictions within task-critical frequencies.

  • This allows the AI to impose a frequency-band prior, ensuring that high-frequency predictions are modeled probabilistically rather than relying solely on potentially overfitting regression heads.

Modality-Specific Encoder Architectures with Frequency Enhancement:

The input processing pipeline can be upgraded by integrating a Mixture of Experts (MoE) layer before the modality encoders to perform learnable frequency enhancement across different input modalities (joint positions, velocities, accelerations, and actuation signals).

  • The system can dynamically allocate computational resources to experts that specialize in amplifying specific frequency bands relevant to the current interaction context.

  • This allows the AI to better separate signal components from noise based on their spectral characteristics before they are fed into the main forecasting network.

Large-Scale Proprioception-to-Wrench Pretraining:

The system can benefit from a pretraining phase using massive, open-source datasets (like RH20T) to learn robust, general representations of robot states that map directly to wrench dynamics.

  • This allows the AI to achieve strong generalization across different robot platforms (e.g., 6-DoF vs. 7-DoF arms) and task domains (e.g., manipulation vs. grinding) without needing extensive task-specific fine-tuning on the target environment data alone.

The resulting improved AI system can perform the following tasks:

  1. Predict short-term wrench trajectories for complex, vibration-rich robotic interactions (like grinding or milling) with high fidelity, especially during rapid contact transients where traditional methods fail.

  2. Provide reliable force and torque estimations in sensorless settings under delayed estimation conditions by effectively utilizing historical proprioceptive data and learned frequency priors.

  3. Achieve superior performance in a balanced manner across both low-frequency (smooth interaction) and high-frequency (vibration) wrench components simultaneously, overcoming the typical trade-off observed in existing time-series models.

  4. Function as a robust component in predictive control loops for robotic manipulation, enabling proactive adjustments based on forecasted transient forces rather than reactive error correction.

  5. Transfer learned knowledge of fundamental robot dynamics and contact structures from one robot platform to another, significantly reducing the data requirements for deploying wrench estimation systems on new hardware.

Abstract

Force and torque (F/T) sensors enable contact-aware control by providing reactive feedback, but they are often fragile and expensive. To overcome these limitations, sensorless methods estimate F/T or wrench solely from robot proprioception, and have shown success in slow interaction tasks such as grasping. However, their low-pass characteristics limit the estimation of high-frequency signals, which are critical in rapid-contact tasks such as grinding. Communication delays can also make their estimates outdated during deployment, but few methods address this directly. To bridge these gaps, we propose a Frequency-aware Decomposition Network (FDN) to estimate vibration-rich wrench in a sensorless, multi-step-ahead manner. Considering higher-frequency stochasticity, FDN spectrally decomposes the wrench horizon into a low-frequency trend and a high-frequency residual, and estimates each by pointwise regression and a learned conditional distribution, respectively. The frequency-aware layers impose band decomposition priors on the outputs and adaptively enhance frequency amplitudes of the inputs. FDN requires neither an identified robot model nor an F/T sensor during estimation. On real-world grinding data from our 6-DoF hydraulic manipulator, FDN reduces high-frequency amplitude error by up to 47% over the baselines under assumed time delays and maintains competitive low-frequency pointwise accuracy, while the baselines fail to balance these two. We also find multi-step-ahead estimation feasible, with FDN estimating a 1,000 ms horizon within 11 ms on a single CPU thread. Ablation studies further support our design choices. In an exploratory study, transferring wrench dynamics learned from an open-source everyday manipulation dataset reduces low-frequency error by 8%, while high-frequency dynamics appear domain-specific.

Sources

Related papers