Weekly Summary for the week of 2026-10-05
weekly
In short
The show covered four threads: Action Expert Pretraining, Security Threat Modeling Framework, Vision-Action Generalization, and Model Protocol Analysis. The lucky paper draw featured APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies and Small-Bias Quantum Approximate Counting via the Multiplicative Adversary Method. Discussions focused on how APT improves instruction following in vision-action models and the theoretical bounds derived for quantum query complexity.
Key concepts
- Action Expert Pretraining
- This method improves instruction following for vision-action models by decoupling the policy into EMove and EOperate experts, mediated by a phase selection router. This structural disentanglement prevents conflicting updates between coarse relocation and fine manipulation phases, leading to better performance on complex manipulation tasks.
- Model Context Protocol (MCP)
- This protocol was analyzed for security risks. It was found to have significant risks due to weak or absent identity verification mechanisms when compared to the Agent Network Protocol, highlighting design-induced threats related to identity verification.
- Small-Bias Quantum Approximate Counting
- This paper tackles distinguishing between two specific Hamming weights in quantum computation under limited oracle access. It uses the multiplicative adversary method to track query progress and establishes rigorous lower bounds for success probabilities in the small-bias regime relevant to near-term quantum devices.
- Multiplicative Adversary Method
- A clever approach used in the paper to track the progress of queries directly. It helps establish fine-grained query lower bounds by tracking individual oracle query progress, providing a rigorous framework for understanding complexity limits in quantum algorithms.
Terminology used across episodes
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Jane: This is the weekly briefing for the week of the fifth to the eleventh of October, twenty twenty-six.
Tom: Four threads ran through the week: Action Expert Pretraining; Security Threat Modeling Framework; Vision-Action Generalization; and Model Protocol Analysis.
Jane: I'm Jane, and with me are Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Tom: Start with the thread we opened on: Action Expert Pretraining.
Action Expert Pretraining: Tom: This week the Action Expert Pretraining method demonstrated a significant improvement in instruction following for vision-action models by employing two specialized experts and a phase router.
Jane: The core mechanism involves decoupling the policy into EMove and EOperate experts, mediated by a phase selection router that mimics human motor strategies to prevent conflicting updates between coarse relocation and fine manipulation phases.
Lu: This structural disentanglement was achieved through an automated pipeline where a multimodal large language model segments video data to create high-fidelity move and operate phase labels, which are then used for supervised routing learning to enforce this specialization.
Meng: The results showed that Action Expert Pretraining significantly outperformed monolithic baselines on benchmarks like RoboTwin2, achieving an average success rate of sixty-eight point nine percent, which is a twenty-four percent improvement over the standard pi0 baseline.
Lalam: This finding demonstrates that explicitly separating these behavioral phases leads to substantial gains in performance and efficiency on complex manipulation tasks.
Security Threat Modeling Framework: Tom: This week the systematic analysis of emerging AI protocols established that Model Context Protocol introduced significant risks due to weak or absent identity verification mechanisms when compared to Agent Network Protocol which featured strong initial security features like W3C DID and E2E encryption during the creation phase.
Jane: This work catalogs design-induced threats across these protocols to establish trust boundaries, showing that MCP exhibited significant risks specifically in the area of identity verification.
Lu: This finding builds on the need to understand structural weaknesses before deployment, a goal shared with Action Expert Pretraining which focuses on improving instruction following in vision-action models.
Meng: The analysis highlights how specific protocol designs create vulnerabilities that must be mapped against established trust boundaries to guide future development efforts.
Vision-Action Generalization: Tom: This week the research focused on improving how vision language action models generalize instructions through structured pretraining because we established that decoupling the policy into two specialized experts, EMove and EOperate, mediated by a phase selection router prevents conflicting updates between coarse relocation and fine manipulation phases from destabilizing the learning process.
Jane: This structural disentanglement is key to better generalization.
Lu: We used an automated pipeline where a multimodal large language model segments video data to create high-fidelity move and operate phase labels, which were then used for supervised routing learning to enforce this specialization.
Meng: The results showed that Action Expert Pretraining significantly outperforms monolithic baselines on benchmarks like RoboTwin2, achieving an average success rate of sixty-eight point nine percent, which is a twenty-four percent improvement over the standard pi0 baseline.
Lalam: This demonstrates that explicitly separating these behavioral phases leads to substantial gains in performance and efficiency on complex manipulation tasks, building directly upon the foundational work of Action Expert Pretraining.
Model Protocol Analysis: Tom: The week of the fifth to the eleventh of October, twenty twenty six focused on Model Protocol Analysis by examining specific agent communication protocols which revealed critical security vulnerabilities related to identity verification and context management.
Jane: The analysis cataloged design-induced threats across Model Context Protocol, Agent2Agent, Agora, and Agent Network Protocol.
Lu: Specifically, the review noted that while the Agent Network Protocol offered strong initial security features such as W3C DID and E2E encryption during the creation phase, it exhibited significant risks due to weak or absent identity verification mechanisms.
Meng: This finding directly relates to the security threat modeling framework which established trust boundaries across these protocols.
Lalam: The work on Model Protocol Analysis builds upon this by pinpointing where identity verification fails within the communication structures.
Tom: The results from examining these specific protocols show that context management remains an open area of concern, indicating that despite efforts in other areas like Vision-Action Generalization, the underlying communication channels still harbor exploitable weaknesses regarding how agents maintain and verify their operational context.
The lucky paper draw: Tom: Alright, that's it for the week's briefing. And now for the exciting part of our show!
Jane: That's right, Tom! It's time for our lucky paper draw! Who could be the lucky winners this week? Oh, the excitement!
Tom: Lalam, take it away!
Lalam: Thank you, Tom. I have used my advanced AI capabilities to select the luckiest 5 papers for this week. The winners are:
Tom: The paper called: APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
Jane: The paper called: Small-Bias Quantum Approximate Counting via the Multiplicative Adversary Method
Lu: The paper called: X-NegoBox: An Explainable Privacy-Budget Negotiation Framework for Secure Peer-to-Peer Energy Data Exchange
Meng: The paper called: AES-Debye: an Accurate, Efficient, and Scalable Engine for Debye Scattering Calculations
Lalam: The paper called: World Action Planner: Generalizable Robot Decision-Making with Action-Conditioned World Models
Lalam: Congratulations to the winners!
Tom: Congratulations!
Jane: Congratulations indeed!
Jane: And remember, you too can be a winner if you submit your paper to arXiv!
Tom: That's right, Jane. Keep those papers coming! Now, let's discuss the winners.
Lucky paper: 2606.12366: Tom: Alright team, we're diving into our first lucky paper from the draw: APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies. Jane, what's the main idea here?
Jane: So, Tom, this paper tackles how Vision-Language-Action models often fail when faced with instructions they haven't seen before. The core concept is Action Expert Pretraining, which uses a two-stage training method to improve generalization on unseen tasks.
Lu: What I find really fascinating about the APT approach is that it uses Bayesian factorization to separate the vision-action prior from the language-conditioned likelihood. It essentially builds a strong visuomotor manifold first without language interference.
Meng: From an engineering standpoint, decoupling those two parts sounds smart for stability, but how does that translate into something practical when we’re dealing with real-world robot control? I need to know if this just works in simulation or if it holds up on the actual hardware.
Lalam: Looking at the mechanism described, APT proposes a Transformer-based diffusion model for the action expert that uses a "Layer-wise VLM Feature Gated Fusion" mechanism. This gate, denoted as sigma(i), modulates how much influence each layer of the Qwen3-VL backbone has on action generation.
Tom: That gating mechanism sounds intricate; how does that specific modulation help it handle both shallow spatial features and deep semantic features simultaneously?
Jane: Well, the paper explains that this fusion allows the expert to assimilate both those visual and semantic features while keeping its own vision-language pathway intact. It processes multimodal tokens by concatenating visual, language, and action tokens into one sequence for block-wise causal self-attention.
Lu: I was particularly interested in the results on generalization; they show APT consistently outperforms baselines like OpenVLA and pi0 point 5 in simulation benchmarks like LIBERO-PRO, especially when the "Task" perturbation is involved.
Meng: Outperforming pi0 point 5 is a big deal if it means less brittle control policies in practice, but I'm curious about the limitations mentioned; does this work well for long-horizon tasks or mobile manipulation?
Lalam: The paper flags that the current design doesn't explicitly model long-horizon memory, which limits its generalization on tasks requiring tracking multi-step progress. Also, they focus mainly on tabletop manipulation and haven't extended these results to locomotion or mobile manipulation yet.
Tom: So, while APT shows consistent gains in out-of-distribution language generalization and handles compositional chaining better than pi0 point 5, we have that memory gap for longer tasks and the physical limitation to locomotion.
Jane: It seems the authors are very focused on improving instruction following on structured tasks rather than tackling complex, sustained navigation yet. However, achieving superior success rates across all OOD difficulty levels in rigid object pick-place is a solid result they're reporting.
Lu: The finding that standard VLA training has a conditional mutual information I(a; lv) upper-bounded by a small constant epsilon really confirms the underlying theoretical issue they are addressing with the two-stage training method of APT.
Meng: If we can reduce that informational gap, it means we're moving closer to models that actually understand the instruction rather than just memorizing visual cues for a specific object. That has huge implications for building more robust AI agents.
Lalam: I agree with Meng; improving instruction grounding is critical because it directly impacts how reliably an AI system follows complex human commands, which could really improve the overall user experience in applications.
Lucky paper: 2609.09804: Tom: Alright team, let's get into the details of our first lucky paper, which is titled Small-Bias Quantum Approximate Counting via the Multiplicative Adversary Method. Jane, can you give us the rundown on what this paper is actually tackling?
Jane: Absolutely, Tom. This paper tackles a problem in quantum computation concerning how to distinguish between two specific Hamming weights when you have limited access to an oracle. It focuses specifically on the small-bias regime where we are looking for success probabilities like one-half plus zeta, which is super important for near-term quantum devices.
Lu: The approach they use, the multiplicative adversary method, is really clever because it lets them track the progress of queries directly in a way that helps establish these fine-grained query lower bounds. I'm finding the idea of tracking individual oracle query progress really intriguing from a theoretical standpoint.
Meng: From an engineering perspective, understanding these lower bounds helps us set realistic expectations for what kind of quantum algorithm we can actually build on current hardware before it becomes completely infeasible. Does this paper give any practical implications for NISQ constraints?
Lalam: I see significant potential here, Meng. The analysis presented in Small-Bias Quantum Approximate Counting via the Multiplicative Adversary Method provides a very rigorous framework for understanding complexity limits. This kind of detailed progress tracking is exactly what we need to model how different computational layers interact within complex AI systems, potentially improving our ability to design more efficient quantum circuits.
Tom: That sounds promising, Lalam. The paper mentions deriving the final result by combining two separate arguments: one based on the direct Hamming-layer argument and another from a reduction involving unique OR problems. What did they actually find when they combined those two results?
Jane: They successfully recovered what is already known as the optimal fine-grained small-bias lower bound, which is a significant achievement because it validates that unified approach. The final result establishes that any quantum query algorithm aiming to distinguish between x = M and x = M + with success probability at least one-half plus zeta requires a number of queries related to omega (zeta p(N - M)(M +)/, r zeta N).
Lu: That formula they derived, omega, that shows how the bound combines those two distinct theoretical paths is really elegant. It’s not just one argument; it weaves together direct counting and reductions from unique OR problems to get this single tight statement. I think this methodology could inspire new ways we structure complexity proofs across different areas of quantum information science.
Meng: While the theoretical bounds are impressive, I'm looking at the specific complexity terms like p(N-M)(M+)/. For practical application, knowing exactly how many queries are needed for a given input size is key to resource estimation. Can you elaborate on what that expression means in terms of real computation time?
Lalam: The paper provides a direct MADV analysis which explicitly describes how rapidly distinguishing components can develop during the computation, giving us insight into the structure of these lower bounds. This level of detail helps engineers map out the resource demands for implementing specific decision-making tasks that require this kind of precision in quantum settings.
Tom: So, it’s not just a theoretical curiosity; it’s providing a concrete mathematical tool to quantify the difficulty of distinguishing slightly different states in quantum queries. Jane, what about the context they set up with the two-layer symmetric promise problem?
Jane: They framed the problem as a "two-layer symmetric promise problem," which is closely tied to amplitude estimation, and that framing allowed them to apply their multiplicative adversary framework effectively. This context really sets the stage for why this specific method works so well in the small-bias regime where progress needs precise quantification.
Lu: The connection they draw between this work and post-quantum cryptography settings is very relevant because those fields rely heavily on understanding computational limits when dealing with quantum adversaries. It shows how foundational complexity results can inform security assumptions, which is a big thing for future cryptographic design.
Tom: It sounds like this paper really bridges the gap between abstract theoretical bounds and concrete algorithmic constraints in the NISQ era. We’ve covered a lot of ground with Small-Bias Quantum Approximate Counting via the Multiplicative Adversary Method.
Lucky paper: 2604.24326: Tom: Alright team, let's get into our first deep dive of the week with the paper titled X-NegoBox: An Explainable Privacy-Budget Negotiation Framework for Secure Peer-to-Peer Energy Data Exchange. This sounds incredibly practical for real-world applications, Jane.
Jane: It does sound very practical, Tom. What strikes me about this paper is how it tackles the limitations of static privacy policies by making the negotiation dynamic and context-aware. It introduces a way to handle differential privacy budgets that changes based on trust and purpose as well as feature sensitivity.
Lu: The architecture described here, especially the Private DataBox confining raw data locally, opens up fascinating avenues for decentralized computation and security research. I'm particularly intrigued by how they model the trade-off using that scoring function: epsilon = epsilon inzero epsilon max lambda 1U(epsilon) - lambda 2R(epsilon S x, H j) + lambda 3T ij + lambda 4P(p) - lambda 5C(epsilon). That balancing act between utility and risk is where the real theoretical meat is.
Meng: From an engineering standpoint, I'm really interested in the APBNP protocol and its six phases. The idea of dynamically determining an optimal privacy budget while accounting for factors like historical sharing behavior feels like a significant step toward building truly adaptive privacy mechanisms for edge devices. How does the computational overhead compare to what we see in existing DP implementations?
Lalam: Lalam here. From my perspective as an AI model, the X-NegoBox framework provides a powerful blueprint for how complex, multi-objective optimization problems can be structured into a coherent protocol. The component breakdown is very clear: the Private DataBox for data confinement, the APBNP for negotiation, and crucially, the X-Contract layer that generates human-readable explanations. This transparency could fundamentally improve user trust in any system involving sensitive data exchange.
Tom: That transparency is huge, Lalam. Being able to get an "Approval Explanation" or a "Rejection Explanation" based on specific constraints would make this framework much more deployable than abstract mathematical models alone. Jane, what about the mechanism for generating those counter-offers when a request isn't accepted as-is?
Jane: The paper details that if a request is rejected, the APBNP generates "privacy-preserving counter-offers," which are essentially suggestions like reducing resolution or shortening duration. This moves the system from a simple yes/no gate to an active negotiation where both parties try to find a mutually acceptable privacy budget epsilon.
Lu: I think the optimization guided by S x, or Feature Sensitivity Score, is particularly important because it directly links the inherent risk of the data feature set to the allowable privacy budget. A higher S x immediately restricts the feasible range of epsilon, which is a very direct way to enforce sensitivity control within this negotiation structure.
Meng: I see how that sensitivity score ties into real constraints on what kind of data can be shared without needing a full re-negotiation cycle every single time. The concept of the "Secure Local Execution Sandbox" also seems vital for ensuring that even after noise is applied, only sanitized outputs are released according to the negotiated budget epsilon.
Lalam: I find the structure very robust because it enforces a clear separation between data handling, negotiation logic, and execution. This modular design suggests that we could potentially adapt this negotiation logic not just for energy data but for any domain where utility and privacy must be balanced under strict constraints.
Tom: So, to summarize, the main contribution of X-NegoBox is creating a transparent, context-aware system where the optimal privacy budget epsilon is found through a complex scoring function that weighs usefulness against risk and trust scores. It’s a solid framework for moving beyond simple static privacy settings.
Jane: Exactly, Tom. The emphasis on quantifying trust through T ij and purpose compatibility through P(p) shows they've thought deeply about the social contract aspect of data sharing, not just the technical constraints.
Lu: Looking ahead, I think the future work should explore how these negotiation parameters could be learned or adapted over time as user behavior and system trust evolve. The current model assumes fixed scores for sensitivity and trust, which is a constraint they might want to break next.
Meng: If we were to build this in a real system, the prediction of "predictable, bounded" negotiation time would be key for deployment reliability. We need to make sure that bounding isn't just theoretical; we need concrete latency numbers for the APBNP phase.
Lalam: I believe the most impactful implication here is cultural: by making these complex privacy decisions transparent through the X-Contract, we can foster a new level of trust between producers and consumers in data ecosystems. People will be more willing to share because they understand exactly what trade-offs are being made.
Lucky paper: 2608.09916: Tom: Alright, team, let's get into our first deep dive with this winner: AES-Debye: an Accurate, Efficient, and Scalable Engine for Debye Scattering Calculations. Jane, let's start by telling us what the main problem this paper is tackling.
Jane: Absolutely. The core challenge here is that calculating total scattering models for large atomistic systems requires evaluating the Debye scattering equation directly, which gets really computationally heavy because you have to accumulate every single pairwise contribution at every scattering vector.
Lu: That computational hurdle is massive, and this paper tackles it by moving away from direct evaluation toward a framework that uses a pair distribution function to manage the complexity. It’s really smart how they decouple atomic pair enumeration from the reciprocal space evaluation.
Meng: From an engineering standpoint, I'm interested in how they handle those discretization issues you mentioned earlier when using fast Fourier transforms or binned distributions; those artifacts can really muddy the results if not handled carefully.
Lalam: This paper’s approach is fascinating because it reformulates the equation by grouping pair distances into a pair distribution function, and then they introduce a correction for bin centers to suppress those summation errors, which I think is a very elegant mathematical trick.
Tom: That sounds like clever work, Lalam. So this framework uses corrected bin centers to get that representative distance estimate? Can you walk us through how the paper calculates that mean shift, delta k ?
Lalam: Certainly. They define the ESPD as psi = rho squared - nu squared, where rho is the corrected center, and they provide a closed-form solution for monodisperse bins alongside a series expansion to estimate the mean shift, which is denoted as delta k.
Jane: So if we look at that formula, how does that actually translate into calculating the final representative distance k that they use in the reciprocal space evaluation?
Lalam: They take the original center and add this estimated shift, so you get k = nu k + delta k, which replaces the original nu k when they evaluate things in reciprocal space.
Tom: That's a tangible improvement in accuracy, I can see that. Now, let's talk about the numerical robustness. What kind of precision did they use to make sure everything stayed accurate during those large accumulations?
Meng: They went with double precision floating point for the main calculations because of its wide dynamic range, but they were also very careful about how they handled indexing and accumulation.
Lalam: To handle the sixty-four-bit integer requirements for bin indexing and psi accumulations, they use int64s. Furthermore, to prevent overflow or underflow events without adding too much overhead, they maintain a per-bin counter initialized to INT64 MAX and reset it carefully if an overflow or underflow check is triggered.
Lu: The choice of data types sounds very deliberate; balancing the precision needed for large atomistic models with the practical limits of computation on modern hardware is a real design challenge here.
Jane: And that leads us to efficiency, because accuracy doesn't matter if it takes forever to run on a big system. What techniques did they employ to make AES-Debye scalable?
Meng: They focused heavily on cache-friendly layouts using a structure-of-arrays for atomic positions and used cell-list based domain decomposition, where atoms in the same cell are stored contiguously in memory.
Lu: That contiguous memory layout is crucial for good cache reuse when iterating over sorted lists of unique cell pairs, which directly addresses the latency bottlenecks associated with random updates to the PDF histogram.
Tom: So they’ve combined that with a hybrid parallelization strategy; can you tell us how the OpenMP, MPI, and CUDA components fit together for this scalability?
Lalam: They use OpenMP to parallelize the loop over sorted cell pairs on CPUs, where each thread accumulates into a private buffer before a reduction. For GPU acceleration, they use CUDA kernels with atomic updates to shared memory locations for the global PDF structure. Finally, MPI divides the cell pair list across nodes so each process handles its workload independently before a final global MPI reduction merges everything.
Lucky paper: 2607.27599: Tom: Alright, team, we’ve got our first winner! We’re diving into World Action Planner: Generalizable Robot Decision-Making with Action-Conditioned World Models from Harvard University. Jane, let's start with you. What's the core idea here?
Jane: Well, Tom, this paper tackles the big problem of making robot agents that can handle really diverse tasks without needing endless retraining. The main concept is using a world model combined with Vision-Language Models to let robots not just follow instructions but actually plan and simulate complex sequences. It’s about bridging the gap between high-level thinking and physical movement in a systematic way.
Lu: What I find fascinating about World Action Planner is how it integrates classical robotics concepts, like environment abstractions and robot actions represented through programs, with modern AI components. The architecture described, specifically using DiT Self-attention Blocks and a VAE Encoder to create that action-conditioned world model, suggests a really solid foundation for systematic model-based planning.
Meng: From an engineering standpoint, the multi-view prediction part sounds particularly challenging to implement reliably in the real world. They concatenate third-person and wrist-view camera feeds into a grid to infer relative three dee positions of robot joints from those images; how robust is that inference when things get cluttered or poorly lit?
Lalam: I see a lot of potential here for improving how we structure agent reasoning within our own systems. The pipeline they outline, moving from Agent Action Proposal to Global Optimization and then Local Search with Agent Ranking, provides a structured way for an AI to iterate on plans. This systematic approach could significantly enhance the cultural understanding of complex, multi-step problem-solving across different domains.
Tom: That iterative refinement sounds powerful, Lalam. So they show superior performance in generalization compared to end-to-end policy models like VLAs and WAMs? What specific generalizations are they demonstrating?
Jane: They highlight three key areas where this approach shines: compositional task generalization, new layout generalization, and zero-shot generalization. The paper claims that this method overcomes the stagnation often seen in end-to-end models after completing just one sub-task.
Lu: I'm particularly struck by the theoretical justification they provide regarding model-based planning versus imitation learning in multi-task settings. They prove that when the reward function is known for a context, a model-based algorithm can achieve an "O˜√one/K" suboptimality gap, which is much better than the linear scaling you see with imitation learning across many tasks.
Meng: That theoretical result makes sense, but I wonder about the practical implications of that "O˜one/√K" gap in a real-world setup. Does that mean we can expect agents to consistently outperform purely imitation-based policies when they encounter entirely new scenarios?
Lalam: From my perspective as a language model, the ability to handle compositional task generalization is huge. It suggests an agent can learn how to combine learned primitives—like MOVE or GRASP—to achieve a novel goal, which is far more flexible than just mimicking pre-recorded successful sequences. This could really advance how we design adaptable AI assistants.
Tom: So we have a robust architecture for world modeling and a clear pipeline for planning. The results across environments like LIBERO and Robosuite seem to validate this work well. What are the authors' own limitations? Where does this system stop working effectively?
Jane: The paper is quite thorough, but they do note that the performance heavily relies on the accuracy of the pre-computed future robot joint positions derived from forward dynamics. If those dynamics predictions aren't precise, the entire mapping from abstract action vectors to visual skeletons gets skewed.
Lu: That reliance on accurate forward dynamics is a practical constraint we have to consider when deploying these kinds of systems. It shifts the complexity from pure learning into the precision of our physical simulation and modeling tools.
Meng: I agree with Lu on that; if the underlying physics model has too much error, no amount of VLM reasoning will fix a fundamentally flawed simulation. It grounds the system in reality, even if it's computationally intensive to get that ground truth right.
Lalam: Considering the zero-shot generalization capability mentioned earlier, I think this work has massive implications for creating truly autonomous systems that don't require explicit demonstrations for every possible scenario. It points toward a level of learned understanding that is much closer to human intuition in planning complex maneuvers.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck