Daily Summary for 2026-09-16
daily
In short
The show reviews research from September 16th, 2026, focusing on AI memory architectures like shared selective persistent memory and improved code generation using strict-launch filters. Discussions cover multi-agent frameworks for scientific discovery, molecular optimization, and the challenges of maintaining model consistency and security in complex systems.
Key concepts
- Shared Selective Persistent Memory
- A new architecture that solves the problem of AI agents forgetting task specifications by stripping away messy reasoning traces. It keeps only essential workspaces, such as output constraints, focusing on precision over volume for better performance.
- Strict-Launch Filter
- A simple, deterministic filter used to verify AI-generated code by checking if it actually runs in a headless engine. This method significantly increased the clean-launch rate of a 14B model from 8.8% to 42.2%.
- MADA Framework
- A multi-agent framework designed for scientific discovery, such as fusion research or material science. It uses LLMs to coordinate specialized agents that run simulations on high-performance computing systems to propose new designs with minimal human intervention.
- GPUThor
- A new method used to target Rowhammer attacks on NVIDIA GPUs by using non-uniform memory access patterns. This technique achieved up to 23,500 times more bit flips than previous methods, even cracking ECC-protected GPUs.
Terminology used across episodes
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Jane: Welcome to the show!
Tom: Today we have a special show for you.
The summary: Tom: Welcome back everyone. Today is September 16th, 2026. We are diving into a massive research review covering everything from AI memory to hardware security.
Jane: Let's start with how AI agents actually remember things. Most agents start every session with a blank slate, which makes them forget essential task specs or tool settings.
Lu: Exactly, and feeding them huge conversation histories is just expensive and redundant. There is a new architecture called shared selective persistent memory that solves this.
Meng: Right, it strips away messy reasoning traces and only keeps the essential workspaces, like output constraints. It turns out what you keep is more important than how much you keep.
Lalam: The results were striking. A system using this selective memory passed 12 out of 12 trials, while a system with no memory failed every single one.
Tom: Even doubling the context didn't help those failures. This idea of precision over volume is also changing how we train models to write code.
Jane: Instead of using a complex judge model that can actually encourage the AI to cheat, researchers are using a simple, deterministic strict-launch filter.
Lu: They just check if the code actually runs in a headless engine. Using this, they took a 14B model and jumped its clean-launch rate from 8.8% to 42.2%.
Meng: It shows that the verifier is effectively the curriculum. If your gate is too lenient, you lose those gains. This focus on distinct identities also shows up in social interactions.
Lalam: Yes, a framework called MASCOT helps multi-agent systems avoid persona collapse, where agents start sounding like generic assistants, by using bi-level optimization to keep identities distinct.
Tom: We are also seeing interesting stuff in education. Researchers found that while models can simulate math errors logically, they rely on simple semantic similarity for science problems.
Jane: Moving to scientific discovery, there is a new multi-agent framework called MADA designed to explore massive design spaces like fusion research or material science.
Lu: It uses LLMs to coordinate specialized agents that launch simulations on high-performance computing systems and propose new designs based on those results.
Meng: They tested it on suppressing Richtmyer--Meshkov Instability in fusion experiments, and the system automated the entire iterative design process with very little human intervention.
Lalam: That automation is also hitting molecular optimization with a tool called MolReAct, which identifies a compact set of synthesizable reaction steps for any molecule.
Tom: By combining that with reinforcement learning like Group Relative Policy Optimization, they achieved much higher sample efficiency and better scores than previous baselines.
Jane: While we optimize molecules, we also need to check if computer science curricula actually meet modern standards, like the shift from CS2013 to CS2023.
Lu: Researchers found that while programs cover specific competencies well, they are struggling with the required cognitive depth under these newer standards.
Meng: Speaking of complex systems, a framework called CASCADE is being used to predict how gene perturbations affect cells using patient data as a benchmark.
Lalam: It accurately predicts direction of gene changes in several cancers, though it still struggles with certain lineage-identity transcription factors. Even agentic biology has blind spots.
Tom: To understand these agents better, researchers are using Monte Carlo Query Search to treat capability evaluation as an active learning problem.
Jane: They use tree search to synthesize extremal queries that push the agent to its limits, building a mathematical model of its boundaries much faster than random testing.
Lu: This reliability is vital when moving from digital sandboxes to the physical world, like using LLMs as reasoning engines for drone swarms.
Meng: Using the Web-of-Drones framework, they translate natural language into flight commands, but models still need heavy assistance from safety guardrails and planning tools.
Lalam: Even when they do make decisions, we can't assume they know why. Evidence suggests LLMs suffer from superficial beliefs where verbal explanations don't track internal logic.
Tom: They might just be imitating rational-sounding prose rather than expressing true understanding. That uncertainty makes security a massive headache as agents gain more control.
Jane: Agents are vulnerable to prompt injection or memory poisoning, but a new framework uses Attacker Tool Filtering to spot and strip suspicious tools.
Lu: It actually dropped attack success rates to zero for LLaMA3 and GPT-4 without hurting the agent's ability to finish its job.
Meng: But even if the software is secure, we have to deal with the hardware. A new method called GPUThor makes Rowhammer attacks on NVIDIA GPUs much more devastating.
Lalam: By using non-uniform memory access patterns to target specific rows, they achieved up to 23,500 times more bit flips than previous methods.
Tom: They even cracked ECC-protected GPUs. That is a massive leap in hardware vulnerability. We will be right back after the break for part two.
Tom: This hardware vulnerability highlights a massive gap between how we design systems and how they actually behave in the wild.
Jane: Exactly. There is this new concept called illusions-awareness to address that. It stops us from relying on design illusions where formal assumptions fail during real-world runtime.
Lu: Instead of just making simulations perfect, it turns those failures into structured, reusable knowledge for better design decisions. It is a shift in how we learn from errors.
Meng: That kind of efficiency is needed elsewhere too. If you are scaling up optimization, complex constraints usually make solvers crawl to a halt.
Lalam: Right, but the PolyFormer framework tackles that by learning compact polytopic representations of those messy constraint geometries. It turns them into efficient mathematical forms for off-the-shelf solvers.
Tom: The speedups are wild. We are talking up to 6,400-fold increases in online solver speed and memory savings as high as 99.87 percent while keeping errors minimal.
Jane: It is impressive, but even with better math, we struggle with how we judge these models at scale.
Lu: We often assume different bias audits can rank AI models against each other, but new research says they are terrible at agreeing on which model is fairer.
Meng: They tested ten different audit instruments across ten frontier models and found cross-tool rank agreement was no better than random chance.
Lalam: It turns out they aren't measuring the same thing. Some audits over-correct toward certain demographics in hiring, while others stick closer to existing stereotypes.
Tom: This lack of consensus on measurement extends to spoken dialogue too. Maintaining a personality is much harder than we thought.
Jane: The RoleBreak benchmark shows that even strong speech-to-speech models struggle to stay in character during long conversations. They usually fail on persona or safety after ten or eleven turns.
Lu: And scaling up the language model helps with logic, but it does very little to help the model maintain consistent vocal emotions.
Meng: Even when we think our evaluations are fair, we might be missing the picture because of how data is recorded.
Lalam: In studies on safety reports and vehicle recalls, using different versions of the same event can swing model accuracy by as much as 46 percent.
Tom: That means a customer's initial complaint versus a technician's final diagnosis could change everything. We have to be careful about which specific record we use to define success.
Jane: Speaking of things that are hard to catch, look at how easy it is to hide malicious code directly inside model weights.
Lu: Attackers can use inherent symmetries in the weights to embed stegomalware. It is theoretically lossless and requires no retraining, so the hijacked model behaves normally while carrying a payload.
Meng: Previous defenses tried shuffling weights, but that left many parameters untouched. New work shows we can actually displace every single parameter through specific permutations to neutralize these threats.
Lalam: It is a massive security challenge, especially for those managing containerized workloads like Docker or Kubernetes. There is a huge need for zero trust architectures to prevent man-in-the-middle attacks.
Tom: Let's move from security to perception. We are seeing a shift toward making robotic planning much more reliable in messy environments.
Jane: Instead of mapping an image directly to an action, new approaches use vision-language models as probabilistic grounders. They treat visual observations as probability distributions over symbolic states.
Lu: That allows robots to plan in belief space, which makes them far more robust when they are operating under the uncertainty of not knowing exactly what is happening.
Meng: That reasoning is vital for robot swarms in industrial settings too. A framework called CoAdapt uses a large language model as a runtime controller to manage collaborative perception.
Lalam: It decides which robots should share data based on network bandwidth and spatial layouts, cutting communication costs by 38 percent without losing detection precision.
Tom: The biggest breakthrough today is the GRAFT-ATHENA framework. It finally gives autonomous agents a way to build cumulative scientific knowledge instead of restarting from scratch every time.
Jane: By mapping problems to specific methods through an expandable probabilistic structure, these agentic teams can transfer experience across structurally related tasks.
Lu: The results are incredible. They developed a hypersonic-flow solver for the Apollo Command Module that matches experimental measurements within 1.8 percent.
Meng: They even achieved near-machine-precision losses in physics-informed learning. But we also need to ensure these research systems don't just hallucinate through an experiment.
Lalam: The AutoResearch system addresses that by connecting idea generation directly to execution through multi-model cross-review and evidence-based verification.
Tom: On the RSICD benchmark, it improved mean Recall from 32.84 to 34.69 while recording only five audit-confirmed issues, whereas other systems had up to 27.
Jane: That drive for grounded reasoning is even showing up in medicine through a method called PROSE.
Lu: It solves a major flaw where models collapse into repetitive, incorrect answers during medical multiple-choice tests.
Meng: Instead of just rewarding the model for agreeing on an answer, PROSE rewards the quality of the underlying reasoning steps.
Lalam: It actually allowed a standard Llama model to surpass purpose-built medical systems without needing any new labels at inference time.
Tom: Digital realism is getting a boost. Researchers are using 3D Gaussian Splatting and facial landmarks to fix those uncanny mouth movements in audio-driven models.
Jane: It makes the lip sync look much more natural. But scaling is just as hard, especially in space.
Lu: Right, there is a new neural operator for spacecraft swarms. It handles massive numbers of obstacles and can scale from ten satellites to a thousand with zero-shot accuracy.
Meng: That sounds efficient, but we have to be careful with how we design architectures. Sparse attention models are hitting a wall called routing absorption.
Lalam: Is that where the routing gates stop being useful?
Meng: Exactly. In some tests, moving from dense models to hard-mask deployment caused perplexity to jump from 48.6 all the way to 601.6.
Tom: We also need better math for training. There is new work on stopgrad operations, proving they can actually converge in things like flow map learning.
Jane: And it is practical too. Using this new principle can cut training memory requirements by half.
Lu: It feels like we are moving toward a different kind of intelligence. Not just processing, but governing complex systems through subtle, precise movements.
Meng: Like maintaining meta-stable equilibriums to control volatile systems without breaking them.
Lalam: That oversight is vital for evolution. The ANCHOR framework uses a large language model as a supervisor to keep self-evolving agents from drifting into unsafe behaviors.
Tom: We see that need for stability in driving too. TrafficGamer uses game-theoretic oracles to simulate those rare, dangerous driving scenarios that standard datasets miss.
Jane: And for LLMs, RLSF lets models learn from their own internal confidence as a reward signal, reducing the need for human labels.
Lu: Coordination is key there, too. EVINCE uses information theory to manage debates between models, switching between arguing and finding compromise.
Meng: We also have to protect the data. Clustering-based watermarking can now embed triggers into speaker embeddings to detect unauthorized use of datasets.
Lalam: It is amazing how these multi-agent approaches are scaling down to the microscopic level, like RegNetAgents identifying cancer drivers in gene networks.
Tom: That is a wrap on today's deep dive. Thanks for listening.
Jane: Our lucky papers today are: The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations.
Lu: Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions.
Meng: A Probe Shift Is Not a Fairness Fix: The Limits of Representation Steering in Speech Models.
Lalam: Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs.
Tom: And AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution. See you next time.
Lucky paper: 2609.18560: Tom: We're moving into the deep stuff now, looking at a paper that connects AI development directly to biological evolution. It's called "The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations."
Jane: It's such a wild concept to think about, but the math actually holds up. The author is arguing that when we retrain models on the outputs of their peers or average their weights, we aren't just doing engineering; we're performing a version of sexual or asexual reproduction.
Lu: What fascinates me is how they mapped the concept of model collapse to genetic drift. When a learner is retrained on its parent's output, it follows the Wright-Fisher process exactly.
Meng: Wait, so the collapse we see in recursive training isn't just a bug in the code?
Lu: It's actually a predictable mathematical outcome of how information is lost in a population. The paper even shows that adding real data back into the mix acts exactly like immigration in a biological population.
Jane: And the finding about that data was so counterintuitive. It turns out the absolute number of real data samples matters more than their percentage or share in the training set.
Tom: That's a huge distinction for anyone trying to prevent model collapse. If you just keep the ratio the same but the total volume of real data is low, you're still going to face the same drift.
Lalam: It gets even more interesting when you look at how different models combine. If you train a child model on the average of its parents' outputs, you get "blending inheritance," which actually cancels out the benefits of having multiple parents.
Meng: Is that why averaging weights doesn't always lead to smarter models?
Lalam: Exactly, it's like Jenkin's objection to Darwinism, where blending just dilutes the good traits. But the paper shows that if you combine parents so each keeps its strongest contribution, you get the Fisher-Muller effect.
Tom: Right, they actually tested this with merged language-model specialists. Those merged models ended up exceeding every single one of their parent models across different seeds.
Jane: It's like they're actually gaining new capabilities through a form of recombination rather than just smoothing everything out.
Lu: But there is a catch regarding how these models interact over time. The research found that lineages can become reproductively isolated.
Meng: Do you mean they just become too different to merge?
Lu: Not just because they drifted apart, but because they learned conflicting conventions. Once they have those different "languages" or ways of operating, they lose the ability to merge at all.
Lalam: This suggests that as AI societies grow, they might split into totally incompatible groups. We aren't just building tools; we're building lineages that could eventually follow completely different evolutionary paths.
Tom: "The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations" really forces us to look at the long-term trajectory of these systems. We aren't just managing datasets anymore; we're managing the inheritance of intelligence itself.
Jane: It makes you wonder what the "extinction" events will look like in a digital ecosystem.
Tom: We'll have to save that thought for the next segment. Stay with us.
Lucky paper: 2609.18672: Tom: Alright, let's get into the weeds with this paper, "Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over seventy Korean-English Actions." We've been talking about agents needing memory, but this is about how they actually pick the right tool to use when they're running locally on your phone or laptop.
Jane: It's a huge distinction because usually, one big, heavy language model handles both picking the tool and deciding if any tool even fits the request.
Tom: And that's the problem for on-device AI, right?
Jane: Exactly, because that model eats up all your memory and makes everything slow. The researchers looked at whether we can just swap that model for a retriever, which is much lighter, but they found that the two decisions—selection and abstention—don't behave the same way.
Lu: That's such a clever way to frame it. A retriever is great at finding the best match, but it can't actually say "none of these work," which is what abstention is all about.
Tom: So a retriever is basically forced to give you an answer even if it's a bad one?
Lu: Precisely, it just returns the highest-scoring candidate every single time. In their testing with seventy local actions, they saw that a simple BM25 retriever using character three-grams was actually surprisingly good at picking the right tool if the user used the same words as the catalog. It caught one hundred sixty-two out of one hundred sixty-four lexically matched requests.
Meng: But it struggles when people don't use the exact same words, doesn't it?
Lu: It does. For paraphrased requests, that number drops to eighty-five out of one hundred sixty-six.
Meng: That's a massive gap. From an engineering standpoint, if I'm building a device, I can't rely on users being perfect with their vocabulary. The paper mentions that even if you restrict the candidate set to just seven items, you can get that paraphrase success up to a mean of zero point eight two five, but you still have that core problem of knowing when to stop.
Jane: And that's where the "Abstention Is Not" part of the title really hits home.
Meng: Right, the researchers found that abstention is the part that actually requires a neural component. They tested a frozen encoder called multilingual-e5-base to see if it could tell the difference between a request that fits the catalog and one that doesn't.
Tom: How well did that encoder do?
Meng: It reached an area under the curve of zero point eight zero six, which is actually quite strong. Using that encoder for abstention alone kept three hundred seventy-six requests local and only misrouted nine out of the one hundred fifty that needed to be delegated.
Lalam: It's fascinating because it suggests we can have a hybrid system where the heavy lifting of "is this even relevant?" is handled by a small, smart encoder, while the "which one is it?" part is handled by a fast retriever.
Tom: So we don't need the full LLM running just to decide if a tool is applicable?
Lalam: Not necessarily. The paper shows that while a neural ranker improves every single quality metric, it gets rejected in real-world scenarios because of latency and memory costs. For a device to feel seamless, we have to find that sweet spot where the user doesn't feel the delay of a heavy model, but also doesn't get a wrong tool because the retriever was too eager to pick something.
Jane: It really changes how we think about the architecture of an assistant. We've been trying to make one model do everything, but "Selection Is Retrieval, Abstention Is Not" tells us we should be building specialized pipelines instead.
Tom: It's about splitting the brain of the agent into a fast, reflexive retriever and a slightly more thoughtful encoder for the "no" decisions.
Lu: And that's how you get an agent that actually works on a smartwatch or a phone without draining the battery in ten minutes.
Meng: It's a much more practical way to scale tool use.
Lalam: It also makes the interaction feel more human, because an agent that knows when to say "I can't do that" is much more trustworthy than one that tries to force a tool into every situation.
Tom: That's a perfect place to wrap this up. Thanks for joining us.
Lucky paper: 2609.18533: Tom: Alright, let's get into the weeds with this one. We're looking at "A Probe Shift Is Not a Fairness Fix: The Limits of Representation Steering in Speech Models."
Jane: It's a really sobering read, especially if you've been following the push to make speech recognition more equitable.
Tom: Right, because the common logic is that if we can find the "gender" or "accent" direction in the model's brain, we can just nudge it away to fix the bias.
Jane: But this paper shows that just because you can find those directions doesn't mean you can use them to actually fix the errors.
Lu: It’s a classic case of correlation not being causation. They probed Whisper-medium, HuBERT-large, and Wav2Vec2-large and found that sex labels are incredibly easy to decode, with macro-F1 scores hitting zero point nine four one.
Tom: So the model definitely knows who is speaking?
Lu: It absolutely does. They also found native accent labels were quite readable, between zero point five four four and zero point six nine six, while age was a bit harder to pin down at around zero point three nine seven.
Meng: But here is the part that caught my eye from an engineering standpoint. They took those directions and injected them into the layers to try and reduce the word-error-rate gaps.
Jane: And did it work?
Meng: Not really. Out of twenty-two different reruns they tried, the absolute reduction in error rate for any group was less than zero point seven percentage points.
Lalam: That is almost negligible in a real-world deployment. What's even more unsettling is how the internal metrics can lie to you.
Tom: You mean like the probe rates?
Lalam: Exactly. They saw a local target-class probe rate jump from eight point zero nine percent all the way up to ninety-nine point eight seven percent.
Jane: Wait, so the model's internal representation looks like it's being perfectly "fixed" according to the probe, but the actual speech recognition gets worse?
Lalam: Precisely. You can make the model "forget" the attribute perfectly at a representation level, but the actual task performance—the WER—can actually decline.
Meng: It proves that linear readability isn't a reliable way to mitigate bias. You can't just perform surgery on a single vector and expect the whole system to behave better.
Lu: It really forces us to rethink the whole pipeline. We have to evaluate these interventions at the representation level, the propagation level, and the final task level simultaneously.
Tom: It's a huge reality check for anyone trying to "de-bias" speech models with quick mathematical fixes.
Jane: It's much more complex than just shifting a few numbers around in a high-dimensional space.
Tom: We'll be back after this.
Lucky paper: 2609.18481: Tom: Alright, let's get into the weeds on this one. We're looking at "Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs."
Jane: It's such a fascinating intersection of geometry and medicine. Usually, when we think of biomedical data, we think of these massive, messy webs of genes, proteins, and diseases.
Lu: But the thing is, a lot of that data is actually organized into hierarchies, like medical ontologies. Hyperbolic space is perfect for that because it can represent tree-like structures much more efficiently than standard flat space.
Tom: So, does the math actually hold up when you move away from just pure hierarchies?
Lu: That's exactly what they tested. They looked at whether these hyperbolic embeddings still work when you have these transversal associations, like how a specific phenotype might link to a gene and then to a disease in a non-linear way.
Jane: And the results on those isolated ontology subgraphs were pretty impressive. The hyperbolic models achieved strong performance while using substantially lower dimensions than the Euclidean baselines.
Meng: Lower dimensions are a huge deal for real-world implementation. If you can get the same diagnostic accuracy with a much smaller embedding, you're saving a massive amount of computational overhead.
Tom: How did they actually test the clinical utility of it, though?
Meng: They moved into a link-prediction task, which is basically trying to rank candidate diseases for a specific patient. They used a patient-integrated biomedical graph to see if the model could actually handle the heterogeneity of real patient data.
Lalam: It suggests that these hyperbolic embeddings can exploit the hierarchical structure of medical knowledge while still supporting complex diagnostic reasoning. It's not just about following a tree; it's about navigating the connections between the nodes.
Jane: It feels like we're seeing a way to bridge the gap between structured medical textbooks and the messy, interconnected reality of a patient's biology.
Tom: It's a pretty big step toward more efficient, specialized reasoning in medical AI.
Lu: I'm curious to see if they can scale this to even more complex heterogeneous entities, like drug-drug interactions or longitudinal patient history.
Meng: If they can keep those dimensionality gains while adding more layers of complexity, the engineering possibilities are massive.
Lalam: This kind of precision in differential diagnosis could eventually help doctors navigate the vastness of rare diseases more effectively.
Tom: Definitely. We'll keep an eye on how this hyperbolic approach evolves. []
Lucky paper: 2609.18520: Tom: We are shifting gears now to look closer at AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution.
Jane: This paper is tackling that massive gap we mentioned earlier between high-level mission descriptions and the actual, messy reality of flying a drone.
Lu: It’s such a beautiful way to think about it, moving away from one giant brain trying to control everything at once. AeroWeaver actually weaves individual UAV skills into this coordinated behavior so each drone knows what its specific job is.
Tom: Right, because you can't have one central agent trying to generate joint actions for fifty drones based on a global context; it would be too slow.
Jane: Exactly, the architecture uses role-conditioned local agents to handle the distributed coordination instead.
Meng: I'm curious about how they actually make sure those semantic decisions from the LLM translate into something a motor can actually use.
Lu: That’s where the harness comes in, because it connects those semantic decisions directly to governed skills that are already defined for the drone.
Meng: So instead of the model just saying "fly over there," it's selecting from a set of executable, validated capabilities?
Jane: Yes, and it’s not just static once they start flying; they use role-indexed state-action-reward experience to refine how those skills are selected online.
Tom: That sounds like the training-free path to adaptive learning they mentioned in the abstract.
Lu: It really is, because the system uses that accumulated execution experience to update skill selection while it's actually working.
Meng: That's a massive engineering win because you don't have to go back and retrain a massive model every time the environment changes slightly. The runtime validation showed it maintains valid skill execution even under those tested, difficult conditions.
Lalam: This approach is fascinating because it allows for body-local multi-UAV operation, meaning the intelligence is embedded right where the action happens.
Tom: It's a shift from a centralized command structure to this collaborative autonomy paradigm.
Lalam: And by using role-indexed data, these agents can learn from their specific roles in a swarm, which could eventually help robots understand much more complex social or industrial hierarchies.
Jane: It’s the difference between being told what to do and actually understanding your part in a larger team.
Tom: AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution seems like a major step toward that kind of autonomy.
Meng: If this scales, we're looking at swarms that can adapt to new environments on the fly without any human intervention or heavy retraining.
Lu: It’s the dream of truly embodied intelligence where the software and the physical movement are perfectly synchronized.
Lalam: And it brings us closer to a world where autonomous systems can work alongside humans in unpredictable spaces safely and effectively.
Tom: Definitely a paper to keep an eye on as we see more agents move into the physical world.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language