Research papers — 2026-10-05

Today's work centers on improving how vision language action models generalize instructions, which is vital because current models often fail when faced with new tasks they haven't seen before. We are trying to fix this by using Action Expert Pretraining, or APT, a two-stage method that separates the visual-action knowledge from the language understanding.

This approach works by first training an action expert as a Vision-Action prior using only balanced vision and action data. This builds a solid foundation of visuomotor skills without language interference. Then, we fine-tune this expert with language tokens to steer its actions toward specific instructions, ensuring it follows complex commands reliably.

The core mechanism involves decoupling the policy into two specialized experts, EMove and EOperate, mediated by a phase selection router that mimics human motor strategies. This structural disentanglement is key because it prevents conflicting updates between coarse relocation and fine manipulation phases from destabilizing the learning process.

We are also using an automated pipeline where a multimodal large language model segments video data to create high-fidelity move and operate phase labels. These labels are then used for supervised routing learning to enforce this specialization.

The results show that APT significantly outperforms monolithic baselines on benchmarks like RoboTwin2, achieving an average success rate of sixty-eight point nine percent. This is a twenty-four percent improvement over the standard pi0 baseline. This demonstrates that explicitly separating these behavioral phases leads to substantial gains in performance and efficiency on complex manipulation tasks.

The most critical work here is the security threat modeling framework because it addresses the immediate danger of deploying emerging AI agent protocols before their underlying structural weaknesses are understood. This systematic analysis examines Model Context Protocol, Agent2Agent, Agora, and Agent Network Protocol to create a catalog of design-induced threats grounded in trust boundaries.

We saw that while ANP offered strong initial security features like W3C DID and E2E encryption during the creation phase, MCP exhibited significant risks due to weak or absent identity verification mechanisms. This ambiguity translated into a concrete failure when we measured wrong-provider tool execution under multi-server composition, showing a non-zero risk when identity was not cryptographically bound to the provider.

The analysis also highlighted that in the operation stage, both MCP and Agora carried a high overall risk because they lacked runtime code-integrity enforcement. A2A presented only moderate risk due to its lack of guarantees on semantic validation or strict token lifetime management. This comparison shows that no single protocol offers complete protection across the entire lifecycle.

Furthermore, we established a lifecycle-based framework following NIST SP 800-30, which maps protocol activities to creation, operation, and update stages to calculate risk using the formula R equals L times I. This methodology helps us see how protocols manage identity validation and namespace governance across these different phases.

The most significant finding from this work is that low-frequency motion control policies are sufficient to achieve robust and dynamic quadrupedal locomotion without needing dynamics randomization or explicit actuation modeling for sim-to-real transfer. This matters because it simplifies the deployment of complex robotic systems by showing that simpler, slower controllers can handle real-world dynamics effectively.

The core architecture involves a floating base model where the robot state includes its global position and orientation, and the control system uses an impedance control model for joint torques. The motion controller runs at a frequency fm to generate desired joint states. An actuation tracker operates at a higher frequency fa to generate the actual torques based on those desired states.

The training methodology frames this as a sequential Markov decision process solved using Proximal Policy Optimization, and notably, no dynamics randomization was performed during the training of the blind policies. Empirical evaluations showed that low-frequency motion control policies are less sensitive to actuation dynamics when the settling time is shorter than the control step time. Lower-frequency policies exhibited reduced vibrations compared to high-frequency ones. This suggests that these slower controllers function more like motion planners than predictive controllers at high frequencies.

The most significant development here is the SAVE framework, which addresses the critical issue of uncertainty quantification in vision-language-action models when they are deployed in unpredictable, non-stationary environments. Without confidence metrics, these models cannot reliably signal when their actions might be wrong in the real world. The SAVE method uses velocity field disagreement to estimate epistemic uncertainty and then guides active fine-tuning to collect expert demonstrations more efficiently.

The core idea is deriving an efficient way to measure epistemic uncertainty in flow-matching models by looking at velocity field disagreement across a small ensemble of models. This score, denoted as ue(y; V), is calculated from the pairwise KL divergence between two flow-matching models by sampling intermediate ODE states and measuring the difference in their learned velocity fields. This estimate is then used to prioritize tasks and initial states for expert demonstration collection within the SAVE framework.

This prioritization involves computing VFD uncertainties for all candidate tasks and initial observations, then prioritizing them based on their mean VFD uncertainty using a categorical sampling distribution with a temperature parameter to balance exploration and exploitation. Based on this prioritization, the framework queries an expert demonstration starting from the most uncertain initial observation within each sampled task to collect new data. This new data is then used to iteratively fine-tune the vision-language-action ensemble on a mixture of pre-training and newly collected data, carefully managing the replay ratio to prevent catastrophic forgetting.

This uncertainty estimation method, VFD, was compared against several other techniques like Action-L2 and ACE. It proved to be better calibrated than those baselines when measured by Spearman rank correlation with task success rates. Furthermore, high epistemic uncertainty during deployment can signal imminent failure with a 67 percent accuracy rate in detecting failures. This process of guided fine-tuning is what allows the SAVE framework to require at least twenty-two percent fewer samples than previous methods to achieve similar performance levels.

The work that matters most here is SCALARFEDLQR because it tackles the massive communication bottleneck in deploying sophisticated control algorithms across many different agents. This is a huge hurdle for real-world applications of policy optimization in LQR control. This algorithm proposes reducing per-agent uplink communication from O(d) to O(1) by having each agent only send a scalar projection of its local gradient estimate. The server then aggregates these to find a global descent direction.

This is significant because it solves the problem where communication costs scale poorly with fleet size and policy dimension. The core mechanism involves agents computing a local zeroth-order gradient estimate, gtilde t n, and instead of sending the full vector, they sample a random Rademacher direction vt n and send only the scalar projection rt n along with a seed xi t n. The server then reconstructs these directions deterministically from the seeds to form a global descent direction gbar t. This is used to update the shared policy gain K t.

This process ensures that every agent transmits just one real-valued scalar and an integer-valued seed per round, achieving the desired O(1) uplink cost regardless of how large the policy dimension d is. The stability analysis shows that under standard conditions including a Polyak–Łojasiewicz condition, the algorithm guarantees global stability on a set Sc. Furthermore, by combining this with assumptions about gradient heterogeneity and using a PL condition, Theorem 2 proves that SCALARFEDLQR converges linearly to the optimal cost Jstar avg.

Numerical results confirm this efficiency in practice; specifically, when comparing SCALARFEDLQR against FedLQR under fixed bit budgets, it consistently achieves a higher recovery percentage than FedLQR. This shows superior efficiency when measured against communication cost.

The work on heterogeneous air-ground robot teams matters because it provides a real-world foundation for how different sensing modalities can be fused in complex outdoor settings. This dataset, GA3T, was built using a Clearpath Husky UGV equipped with 3D LiDAR and stereo cameras alongside an Autel EVO II UAV carrying RGB and thermal imagery across five distinct environments. This setup allows researchers to study cross-view perception where the ground robot offers precise 3D LiDAR data while the aerial platform provides rich RGB and infrared observations. This enables occlusion-aware collaboration through sparse tree canopies in early spring.

The annotation pipeline for this data is key because it uses a foundation model assisted human-in-the-loop process. It starts with SAM three to generate initial segmentation masks which are then refined by annotators who correct local errors. This iterative refinement process substantially reduces the effort needed compared to drawing pixel-accurate segmentations manually from scratch. This allows researchers to test cross-view semantic prediction and collaborative traversability estimation using the dataset.

The benchmark evaluation showed that adapting SAM three on GA3T improved performance on both views, with the largest gains seen on the UAV data. This suggests that this specific dataset captures domain characteristics not covered by generic priors. This success is supported by the fact that synchronized sensor streams, including joystick commands and perception data, are recorded to support future studies in learning from demonstration for off-road navigation.

In contrast to this perception work, research into deception against data-driven linear-quadratic control explores how an adversary can be misled into learning a suboptimal attack when injected with deceptive feedback. The optimal deception gain is found by minimizing the distance between the learned policy's gain and a target benign gain while keeping the deceptive feedback as small as possible to maintain stability.

This optimization problem is solved numerically using a block successive over-relaxation algorithm, which iteratively solves coupled algebraic Riccati and Lyapunov equations to find the solution. The simulation validated this approach by showing that deception can force an adversary to learn a more benign attack even when the nominal optimal attack is destabilizing.

Another area of study involves material science where measurements on monolayer TaIrTe4 using microARPES confirm its insulating ground state. This is consistent with density functional theory calculations using the Heyd-Scuseria-Ernzerhof hybrid functional. The research also uncovered a pronounced electron–hole asymmetry in doping response, showing that adding electrons fundamentally alters the electronic structure by driving band renormalization rather than just shifting the Fermi level rigidly.

This material understanding is further refined by investigating spin-orbit coupling, which was found to be the determining factor driving the system from a semimetal to a quantum spin Hall insulator. The study also demonstrated how alkali-metal deposition can create new electron pockets on the surface of TaIrTe4, leading to superstructure formation.

On a more fundamental physics level, work characterizing Haag duality for quantum spin systems uses an entropic criterion based on conditional mutual information to prove the property model-independently. This method shows that for two or more disjoint cones in a two-dimensional system, Haag duality is equivalent to the vanishing of the topological entanglement entropy.

The study on step-edge anomalies in topological metals predicts a robust step-edge conductance that assumes a non-integer value dependent on bulk topology and step height. This anomalous conductance arises from a combination of quantized response at the edge and a non-quantized response carried by bulk modes, which is confirmed by lattice simulations.

The work on kinetically trapped nanocrystals is particularly important because understanding how shape control happens at different stages of crystal growth gives us a blueprint for precisely synthesizing desired morphologies, like cubes or octahedra. The study found that the primary factors driving the formation of cubic nanocrystal shapes are the adatom nucleation energies and the geometry of growth islands. This means we can now guide synthesis by controlling these specific kinetic steps.

This is supported by observations that transient sites dominate growth, leading to metastable shapes such as surface roughening alongside symmetry preservation in various crystal forms. In terms of trapped ions for quantum computing, the WISER framework provides comparative lower-bound estimates of logical clock speed and logical error rate. This is crucial because it identifies viable operating regions without needing absolute hardware predictions.

This framework systematically explores design choices, showing that an order of sixteen multiplexing balances logical clock speed at fifty-three hertz with a power consumption of zero point five watts per logical qubit. Furthermore, the study recommends bivariate-bicycle codes as the only scheme offering low logical error rates with low power on WISE architectures.

The development of a novel compiler for trapped ions is significant because it performs optimal qubit mapping and routing using a SAT approach that jointly models global odd-even routing and multiplexed control. This compiler minimizes total reconfiguration time by decomposing circuits into native WISE operations, effectively finding the best way to place interacting ion pairs into the same trap within minimum passes. This contrasts with existing solvers which cannot express this combination of routing and control constraints.

The research on world models for gradient-based planning addresses a fundamental mismatch where world models trained on expert trajectories fail during planning due to compounding model errors in out-of-distribution states. Online World Modeling is proposed to fix this by iteratively correcting trajectories produced by gradient descent and finetuning the world model on these corrected rollouts. This method, when combined with Adversarial World Modeling, enables gradient-based planning to match or exceed the performance of search-based planners with a ten times reduction in computation time.

The study on adaptive quantum-safe cryptography is important because it proposes a framework that dynamically selects the best post-quantum cryptographic algorithm based on predicted mobility and channel variations in 6G vehicular networks. This Context-Aware Adaptive PQC framework uses an adaptive predictive multi-objective evolutionary algorithm to balance latency, computational cost, and security requirements in real time. The framework demonstrates significant performance gains, reducing end-to-end latency by up to twenty seven percent while lowering communication overhead by sixty five percent compared to static baselines.

The work on on-chip calibrated radio frequency measurement at cryogenic temperatures is paramount because it directly addresses the critical need to accurately characterize strontium titanate based varactors for quantum information processing systems operating at four kelvin. This system overcomes errors caused by long radio frequency circuit lines, which is a major hurdle when dealing with commercial components that often fail under cryogenic conditions.

The calibration procedure involves measuring Smith charts under open, short, and load conditions at the PCB calibration port while calibrated near the vector network analyzer. This process allows for the observation of ideal results across all conditions for a reference capacitor with a known capacitance of fifteen pico farads measured at four kelvin. This calibration enables evaluation of capacitance across the frequency range typically used in radio frequency reflectometry, extending up to twenty-six point five gigahertz.

The investigation into strontium titanate varactor properties examined how annealing conditions, crystal orientation, and calcium doping affected their characteristics. Annealing at one thousand two hundred fifty degrees celsius for thirty hours resulted in a thicker device of three hundred thirty micrometers compared to the two hundred sixty micrometer device without annealing. Furthermore, comparing devices with one thousand one zero orientation versus one thousand one one revealed that the non-doped material exhibited higher relative permittivity values than its (one thousand one) counterpart.

The introduction of calcium doping into strontium titanate resulted in reduced dielectric constant values when compared to the non-doped version. This might be due to oxygen vacancies and interfacial dielectric characteristics. Slight hysteresis observed during voltage sweeping suggests a ferroelectric transition at this specific doping concentration and temperature. This calibration technique developed can now be applied to characterize various other cryogenic microwave components, such as superconducting inductors.

Today's papers

The papers

Important terms

Action Expert Pretraining (APT)
A two-stage method that first trains an action expert using only visual and action data to build motor skills, then fine-tunes it with language tokens to follow complex instructions reliably.
Phase Selection Router
A mechanism that mimics human motor strategies by decoupling the policy into EMove and EOperate experts, preventing conflicting updates during coarse relocation and fine manipulation phases.
SAVE Framework
A method to quantify uncertainty in vision-language-action models using velocity field disagreement to estimate epistemic uncertainty, guiding active fine-tuning for better performance.
SCALARFEDLQR
An algorithm that reduces communication costs in multi-agent control from O(d) to O(1) by having agents send only a scalar projection of their gradient estimate to the server.