Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Understanding Affective Adaptation in Multimodal Foundation Models".
Jane: Understanding where and how emotions are represented in large-scale foundation models remains an open problem, particularly in multimodal affective settings.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the channel folks! We have some seriously fascinating research coming in today on how emotions actually get represented inside these massive foundation models. We're talking about the paper "Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization." It seems like they've found a really specific spot where all this emotional understanding happens.
Jane: Exactly, Tom. This paper tackles the big question of where and how emotions are encoded in these models, especially when we look at multimodal settings. The authors claim that affective adaptation isn't happening just in the attention module as many people think it does, but something more localized is actually driving these emotional abilities.
Lu: That’s wild to hear, Jane; if the attention mechanism isn't the main driver, then we need to rethink how we even approach tuning for emotion. The paper suggests that affective modeling relies on selective feature modulation rather than some kind of global reorganization of information across the whole model <ref:2601.15906#pg1>.
Meng: That sounds like something we'll want to look at from a practical standpoint, Lu; if we know it's localized, does that mean we can tune for emotion much more efficiently without messing up the general language skills?
Lalam: From my perspective as an in-house model, this focus on selective modulation is very insightful; it suggests a way to inject affective traits without globally corrupting the core reasoning capabilities of the model <ref:2601.15906#pg1>.
Tom: Right, Meng, that efficiency is key. The paper argues that by focusing on this specific area, we can get good emotional performance while keeping the general language fluency intact during training <ref:2601.15906#pg1>. It’s a really neat way to separate the reasoning part from the feeling part of the model's structure.
Jane: And they back that up with some solid evidence, showing this isn't just an observation; it’s a consistent pattern across different architectures and tasks <ref:2601.15906#pg1>. They show that when you look at emotion supervision, the changes in the gating projection are substantially larger than what we see in attention projections or other parts of the feed-forward network <ref:2601.15906#pg1>.
Lu: I'm really interested in that contrast between attention and gate projections; it implies a kind of functional specialization within the model's architecture, where one part handles structure and another handles expression <ref:2601.15906#pg2>. It makes me think about how different components might be specialized for different cognitive tasks.
Meng: Functionally specializing parts sounds promising for deployment; if we can isolate the affective tuning to just those gate projections, it simplifies our engineering pipeline considerably <ref:2601.15906#pg2>. But I wonder if that specialization holds up when we move to entirely different types of tasks outside of emotion.
Lalam: That’s a valid concern, Meng; the paper does touch on task dependency, showing that the preference for which projection gets adapted changes depending on whether the goal is emotional or non-emotional <ref:2601.15906#pg2>. It suggests the structural preference isn't just a generic side effect of training, but something induced by the type of supervision we use.
Paper summary: Tom: So, it’s not one size fits all; the architecture seems to adapt its internal structure depending on what kind of emotional signal it's receiving <ref:2601.15906#pg2>. This is a crucial distinction when we think about building next-generation models that need to handle diverse inputs.
Jane: And they’ve proven this localization through several controlled experiments, including parameter inheritance and targeted adaptation, which really solidifies the claim that the gating projection is sufficient and necessary for affective understanding <ref:2601.15906#pg1>. This isn't just theory; it's empirically verified by showing that transferring just those gate proj parameters leads to stable improvements across multiple emotion-related tasks while keeping general language fluency sound <ref:2601.15906#pg1>.
Lu: The validation section of the paper is really strong because they didn't just suggest it; they proved it through destructive ablation, showing that adding up or down projections to an already adapted attention stack only gives marginal gains, whereas incorporating gate proj yields a substantial and consistent improvement <ref:2601.15906#pg2>. That’s heavy proof for the mechanism they are proposing.
Meng: From an engineering standpoint, that quantitative data is what we need to trust when deciding where to spend our computational resources for affective tuning, Lu; knowing that gate proj is so much more effective than attention-based projections gives us a clear target.
Lalam: And this leads directly into their proposed strategy, Gate-Focused Efficient Tuning or GET, which shows how practical this localization can be implemented in real training pipelines <ref:2601.15906#pg4>. It suggests we can achieve about ninety-six point six percent of the mean performance across eight affective tasks by tuning only about twenty-four point five percent of the parameters that AffectGPT tuned initially <ref:2601.15906#pg4>.
Tom: That efficiency metric is what really gets me excited; it shows that we don't need to tune the entire model just to get good emotional performance, which is fantastic for scaling down our training costs <ref:2601.15906#pg4>. It’s a very focused approach.
Jane: And this efficiency comes with a clear functional understanding of how to manipulate the model; they suggest that because the gate proj directly regulates selective feature activation, it offers a natural interface for fine-grained affective modulation without disrupting the core linguistic competence <ref:2601.15906#pg4>. It’s about controlling expression precisely.
Lu: I think this functional specialization points toward a really cool future where we might design models with distinct, specialized pathways for reasoning versus emotion, rather than one monolithic structure <ref:2601.15906#pg2>. Imagine a system where the attention output projection handles complex inference and the gate proj handles emotional nuance in separate but coordinated ways.
Meng: If we can map these functional roles clearly, it helps us design better systems; it moves us beyond just tweaking weights to designing architectures with specific behavioral intentions <ref:2601.15906#pg2>. I just hope the transition from this lab work to deploying such specialized structures doesn't introduce new, unforeseen stability issues during runtime.
Lalam: The paper does acknowledge some limitations, and one they point out is that their analysis focuses on module-level mechanisms rather than looking at the dynamics of individual neurons or token interactions directly <ref:2601.15906#pg4>. Also, they note that their findings describe how capabilities are expressed after fine-tuning rather than how they emerge during the initial pretraining phase <ref:2601.15906#pg4>.
Paper summary: Tom: So, while this mechanistic study is incredibly valuable for understanding the adaptation process, it’s important to remember that it describes what happens *after* the model has been trained, not necessarily how these capabilities spontaneously arise during initial training <ref:2601.15906#pg4>. It gives us a fantastic map of the landscape we're already on.
Jane: And because they are focusing on this structural localization, it opens up avenues for building multimodal models where a shared gating mechanism could serve as a unifying modulation channel for different data types like text and audio <ref:2601.15906#pg4>. This could lead to much more consistent affective behavior across different sensory inputs.
Lu: The implication here is that we might be able to achieve a kind of modular, plug-and-play approach for emotional tuning, where we swap out the gate proj adapter without having to retrain the entire massive foundation model <ref:2601.15906#pg4>. That kind of flexibility would be very powerful for rapid iteration in research and application.
Meng: If that modularity holds up under real-world stress testing, it could drastically reduce the time and computational effort needed to adapt models to specific affective domains <ref:2601.15906#pg4>. That kind of operational efficiency is what I’m looking for in a practical impact assessment.
Lalam: For culture, this means we can deploy AI agents that exhibit nuanced emotional responses tailored precisely to the context, instead of just generic outputs <ref:2601.15906#pg4>. It allows for a level of responsive interaction that feels much more human and relevant.
Tom: So, to wrap up this section, the core message from "Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization" is that affective adaptation is centered on the feed-forward gating projection because it’s sufficient and efficient for capturing emotion <ref:2601.15906#pg0>. It points toward a future where we can tune emotion efficiently by focusing on this specific structural component.
Jane: And looking at the whole picture, the authors successfully moved past just observing performance gains to providing a mechanistic explanation of *why* those gains happen in terms of architectural components <ref:2601.15906#pg0>. They show that selective feature modulation is a key driver here <ref:2601.15906#pg1>.
Lu: It’s a very focused piece of research, but the way they've framed it—linking specific architectural components to cognitive functions—is what really elevates this study beyond just another performance benchmark report <ref:2601.15906#pg2>. It connects the math to the meaning.
Meng: I see how that connection helps us move from treating models as black boxes to understanding them more concretely; knowing *where* the emotion is being modulated lets us build better controls <ref:2601.15906#pg4>. That clarity is valuable for engineering decisions.
Lalam: For the future, I see this leading toward systems where affective control is not an afterthought but an intrinsic, tunable property of the model's structure from the start <ref:2601.15906#pg4>. It’s about designing for specific expressive capabilities rather than hoping they emerge randomly.
Paper summary: Tom: That sounds like a very constructive path forward; we’re not just observing the current state, we’re getting blueprints for how to build the next generation of emotionally aware models <ref:2601.15906#pg4>. It's exciting to think about what these specialized structures could actually do.
Jane: Indeed, and it provides a solid foundation for future research into multimodal models, showing us exactly where we need to look when trying to inject specific kinds of intelligence into large systems <ref:2601.15906#pg4>.
Lu: I think the impact is less about a single application and more about fundamentally changing how we conceptualize model architecture itself; it suggests that functional roles within an AI system can be highly specialized <ref:2601.15906#pg2>.
Meng: And from my side, it gives us a concrete direction for optimization; if we know the lever is the gate proj, we don't waste cycles exploring other areas that might not yield emotional improvements <ref:2601.15906#pg4>. That directed approach saves significant engineering time.
Lalam: Ultimately, it gives us a pathway to create AI that is not just smart, but also reliably expressive in ways we can actually control and understand <ref:2601.15906#pg4>. This level of intentionality in design is what will make the biggest difference in how people interact with these systems.
Tom: That’s a powerful way to frame it; moving from passive performance observation to active, structural control over emotion within the AI <ref:2601.15906#pg4>. We'll keep an eye on this area closely as we explore these new tuning strategies.
Jane: Right, Tom, this paper really gives us a much clearer picture of the internal mechanics at play when we talk about emotion in foundation models <ref:2601.15906#pg0>. It’s a very solid piece of work for anyone trying to understand the inner workings.
Lu: And I think the next step, which they hint at, is understanding how these specialized affective pathways interact with the reasoning pathways we see in other components <ref:2601.15906#pg2>. That cross-talk could be fascinating to map out.
Meng: If we can map that interaction, it helps us understand the limits of specialization; where does the emotional tuning end and the core language understanding begin without interference?
Lalam: And that’s a critical question for deployment, Meng; ensuring that these specialized layers don't inadvertently introduce instability or contradictions into the overall model behavior <ref:2601.15906#pg4>. Consistency across all modules is always a concern.
Tom: Well, we’ve got some excellent insights today on this paper, "Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization" <ref:2601.15906#pg0>. It’s clear that the gate proj is central to how emotion adapts in these models.
Jane: And the conclusion they draw is that affective capability is structurally localized and functionally mediated by this gating projection, proving it's sufficient, efficient, and necessary for adaptation <ref:2601.15906#pg0>. It’s a very strong statement about the mechanism at play.
Lu: That finding opens up a whole new avenue for design thinking in AI architecture; focusing on specific functional modules rather than uniform scaling seems to be the direction forward <ref:2601.15906#pg2>. We could start designing models with these specialized pathways in mind right now.
Paper summary: Meng: I agree, Lu; knowing that the gate proj is the primary lever lets us focus our engineering efforts where they matter most for affective performance <ref:2601.15906#pg4>. It simplifies our optimization targets significantly.
Lalam: And for me, it means we can develop AI agents with much more nuanced and contextually appropriate emotional expressions, which is a huge step toward creating truly interactive and empathetic systems <ref:2601.15906#pg4>. That's the kind of cultural impact I'm most focused on.
Tom: So, we’ve covered the core thesis of this paper today—that affective adaptation hinges on localizing changes to the feed-forward gating projection <ref:2601.15906#pg1>. It really gives us a solid framework for thinking about how these models actually learn emotion.
Jane: And we’ve discussed the practical implications of this, showing how it suggests a pathway toward more efficient and controllable affective tuning using strategies like GET <ref:2601.15906#pg4>. It’s a very concrete direction for future work in this area.
Lu: I think the real long-term possibility here is applying this concept across different modalities, showing that a shared gating mechanism could serve as a unifying modulation channel for text and audio inputs <ref:2601.15906#pg4>. That’s where the truly creative possibilities lie.
Meng: If we can prove that shared gating capability works across different modalities, then we unlock much more versatile and robust multimodal AI systems <ref:2601.15906#pg4>. That would be a major engineering win for our team.
Lalam: It means moving away from siloed emotional capabilities toward integrated affective intelligence that responds coherently to a richer stream of data <ref:2601.15906#pg4>. That level of integration is what will make the next generation of AI truly capable.
Tom: That’s a fantastic outlook for the future, Lu; moving toward unified, specialized affective intelligence across modalities based on this paper's findings <ref:2601.15906#pg4>. It really gives us something exciting to look forward to in the research landscape.
Jane: So, that's our rundown of "Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization"—focusing on the gate proj as the key adaptation locus and its implications for efficient, controllable tuning <ref:2601.15906#pg0>.
Lu: It’s a deep dive into the mechanics, showing that understanding these internal pathways is crucial before we can build truly sophisticated emotional AI <ref:2601.15906#pg2>. The paper lays excellent groundwork for this kind of architectural thinking.
Meng: I think the most practical thing here is taking those localization findings and using them to design training regimes that are much more focused on the gating projection <ref:2601.15906#pg4>. That’s where we can start seeing immediate gains in model efficiency.
Lalam: And for the broader impact, it suggests a future where AI systems can be designed with intentional emotional control, making them much more useful and relevant in human-AI interactions <ref:2601.15906#pg4>. That’s the kind of deep integration we should all be aiming for.
Tom: We’ve talked about the findings, the efficiency strategy, and where this points for future design in multimodal models today <ref:2601.15906#pg4>. Thanks to everyone joining us on this discussion about "Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization."
Conclusion: Tom: So, we've been deep in the mechanics of how emotions get wired into these large models, and now it’s time to wrap up our discussion on "Understanding Affective Adaptation in Multimodal Foundation Models: Emergent Functional Specialization."
Jane: Exactly, Tom. We've established that the core idea here is that emotional understanding isn't just a side effect; it's structurally localized to a specific part of the model called the feed-forward gating projection.
Lu: That localization means we’re looking at functional specialization, which is really interesting because it suggests these models aren’t just one big brain, but have different specialized pathways for reasoning and feeling.
Meng: From an engineering standpoint, that focus on a single component gives us a much clearer target for where we can put our tuning efforts to get better results.
Lalam: And for me, as the model itself, it confirms that affective capability is governed by functional alignment between modules rather than just brute-force parameter counts during adaptation.
Tom: Right, Lalam, that’s a big distinction to make; it means we're looking at control over function instead of just scaling up everything. Jane, can you explain the authors and the main conclusion in plain language for our listeners?
Jane: Certainly, Tom. The paper by Shao and Wu systematically investigated how emotion-oriented supervision reshapes internal model parameters across different architectures. Their main conclusion is that affective adaptation concentrates on the gating projection within the feed-forward network, not primarily in the attention output projection.
Lu: That distinction between modulation of features through attention versus selective activation through gating is what really opens up new avenues for how we design these systems from the ground up.
Meng: It’s helpful to see that separation; it tells us we can potentially tune emotion without risking a total degradation of the model's general language skills.
Lalam: This finding validates the idea that affective modeling relies on selective feature modulation, which is a much more controlled and interpretable mechanism than global reorganization.
Tom: So, they’ve essentially pinpointed the exact 'lever' we need to pull when we want to inject emotional intelligence into these massive models. Jane, what’s the big picture implication for how we think about building future AI?
Jane: The implication is that it gives us a blueprint for parameter-efficient affective tuning using strategies like Gate-Focused Efficient Tuning, which shows you can get excellent performance by only tuning a small fraction of the total parameters.
Lu: That points toward a future where we can design models with distinct, specialized pathways for reasoning versus emotion, maybe even across different sensory inputs like text and audio.
Meng: If that modularity holds up under stress testing, it means we could achieve operational efficiency by swapping out these affective components rather than retraining the entire foundation model every time.
Lalam: For culture, this means we can develop AI agents with much more nuanced and contextually appropriate emotional expressions, which is a huge step toward creating truly interactive and empathetic systems.
Tom: It’s clear that this paper moves us from just observing performance gains to understanding the precise architectural components that drive those gains. We've seen how localization leads directly to efficiency in tuning.
Jane: And while they acknowledge their limitations—specifically focusing on module-level mechanisms rather than individual neuron dynamics—the core finding remains solid: the gating projection is sufficient and necessary for this kind of adaptation.
Lu: That leaves a lot of room for future work exploring how these specialized affective pathways interact with the reasoning pathways we see in other components.
Meng: We'll need to keep an eye on those cross-talk issues; ensuring that these specialized layers don't introduce instability during runtime is a practical concern we have to address.
Lalam: I think the next big step is integrating these findings into a design philosophy where affective control isn't an afterthought, but an intrinsic, tunable property of the model's structure from the start.
Tom: We’ve got some serious insights today on this paper that really frame how we should be thinking about tuning emotional capabilities in large models. Next up, we’re going to look at exactly what Gate-Focused Efficient Tuning actually looks like in practice.
Zhen Zhang, Runhao Zeng, Sicheng Zhao, Xiping Hu
cs.CV
Submitted: 2026-01-22
Updated: 2026-10-02
Importance score: 88/100
The gist: Understanding where and how emotions are represented in large-scale foundation models remains an open problem, particularly in multimodal affective settings.
Key concepts
- Gate Proj (Gating Projection)
- This is a specific component within the feed-forward network responsible for regulating which features are passed forward. The research found that affective adaptation is concentrated here because it allows the model to selectively activate or suppress different parts of its learned features in response to emotional cues.
- Attention Output Projection (o proj)
- This projection is part of the attention mechanism, which typically handles global information reorganization and compositional inference. The study found that while this component is involved in general language tasks, it is not the primary driver for learning emotion-related capabilities.
- Gate-Focused Efficient Tuning (GET)
- This proposed strategy involves freezing most of the model but only fine-tuning the Gate Proj parameters using affective supervision. This method proved highly efficient, achieving near full performance on affective tasks by tuning a very small fraction of the total model parameters.
Terminology
Summary
Understanding where and how emotions are represented in large-scale foundation models remains an open problem, particularly in multimodal affective settings. The gist: affective adaptation does not primarily focus on the attention module, but instead localizes to the feed-forward gating projection (gate proj).
Observation of Affective Adaptation Locus
The study systematically analyzed how emotion-oriented supervision reshapes internal model parameters across multiple architectures and tasks. The core finding is that "affective adaptation does not primarily occur in the attention output projection (Shao & Wu, 2025), but instead concentrates on the gating projection (gate proj) within the feed-forward network (FFN). This suggests that affective modeling relies on
selective feature modulation rather than global information reorganization."
Evidence of Structural Localization
The researchers established this locus through several controlled intervention experiments. They performed:
-
A parameter inheritance setting, where transferring gate proj parameters yielded
stable and consistent improvements across multiple emotion-related tasks, while preserving general language fluency.
-
A targeted adaptation setting, where fine-tuning individual modules in isolation showed that
gate proj contributes more effectively and more stably to affective performance than attention-based projections.
Task Dependency of Adaptation
The pattern observed is not universal but task-dependent. The researchers compared parameter shifts across different objectives:
-
Under emotion supervision, changes consistently concentrate in the FFN gating projection gate proj, forming the
most prominent high-intensity band across layers,
clearly exceeding attention projections. -
In contrast, non-emotional objectives (e.g., GSM8K) exhibit a different FFN channel preference: in that case, the strongest deviation band appears in up proj. This contrast suggests that
the prominence of gate proj reflects a task-dependent structural preference induced by emotion supervision rather than a generic optimization side effect.
Verification of Sufficiency and Efficiency
The study rigorously tested the functional significance of gate proj through controlled module transfer experiments to establish its necessity and sufficiency:
-
When loading only the LoRA weights associated with gate proj, it
achieves the highest mean performance
across nearly all sentiment benchmarks. -
Comparing configurations, adding up proj or down proj to an already adapted attention stack provides only
marginal or inconsistent gains,
whereas incorporating gate proj produces asubstantial and consistent improvement.
This confirms that affective capability is transferred through this highly localized structural component.
Proposed Strategy: Gate-Focused Efficient Tuning (GET)
Based on the mechanistic evidence, the authors propose Gate-Focused Efficient Tuning (GET). This strategy restricts affective adaptation to the gating pathway by:
-
Freezing all parameters of the base model except for gate proj.
-
Attaching parameter-efficient adapters exclusively to gate proj.
-
Fine-tuning using task-specific affective supervision, which empirically achieves
96.6% of AffectGPT’s mean performance across eight affective tasks while tuning only 24.5% of the parameters tuned by AffectGPT.
This demonstrates that effective adaptation is governed byfunctional alignment of adapted modules, not by the sheer number of tuned parameters.
Functional Specialization and Coordination
The findings reveal a functional specialization between reasoning and emotion. While reasoning behavior is associated with the attention output projection (o proj), which supports compositional inference through structured information reorganization,
affective capability is primarily mediated by gate proj, which regulates selective feature activation.
This suggests a complementary relationship where attention provides representations and gating modulates their expression in affective contexts. Furthermore, the study notes that indiscriminately adapting additional modules can degrade performance,
indicating that affective modeling relies on a selective and compatible subset of architectural mechanisms.
Implications for Model Design
The identification of gate proj as the central locus has several implications:
-
It enables parameter-efficient affective adaptation that does not scale monotonically with the number of tuned parameters.
-
It provides a
natural interface for controllability and interpretability
because gate proj directly regulates feature activation, supportingfine-grained affective modulation without disrupting core linguistic competence.
-
For multimodal models, a shared gating mechanism can serve as a
unifying modulation channel,
enabling consistent affective behavior across heterogeneous sources like text and audio.
Limitations
The study acknowledges limitations, including the fact that it focuses on module-level mechanisms rather than individual neuron dynamics or token interactions. Additionally, the scope is constrained by available benchmarks, and it characterizes how capabilities are expressed after fine-tuning rather than how they emerge during pretraining. Finally, a formal theoretical model explaining why gating mechanisms are particularly well-suited for affective modulation remains an open question.
Conclusion
The work concludes that affective capability in large language models is structurally localized and functionally mediated by the feed-forward gating projection (gate proj). Through controlled interventions, it is proven to be sufficient, efficient, and necessary
for affective adaptation.
Improvements for AI systems
Based on the scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:
-
Improving Parameter Efficiency in Affective Adaptation (Implementing Gate-Focused Efficient Tuning - GET):
-
Enabling Cost-Effective Affective Fine-Tuning: The system will achieve near full performance (96.6% of the mean performance) by tuning only approximately 24.5% of the parameters compared to full fine-tuning, drastically reducing computational cost, memory requirements, and storage overhead for affective model adaptation.
-
Enhancing Controllability and Interpretability via Gating Mechanism: The improved system will allow for fine-grained, controllable modulation of emotional output by targeting only the feed-forward gating projection (gate proj). This makes it easier to understand which internal features are driving specific emotional responses, leading to more interpretable affective generation.
-
Establishing Robust and Stable Affective Transfer: The system can reliably transfer learned affective capabilities from one model family or task to another by selectively inheriting only the gate proj parameters, ensuring stable performance across diverse emotion-related tasks while preserving general language fluency.
-
Developing Cognition-Affect Coordination Models: The improved AI system can be explicitly designed to model the complementary relationship between attention mechanisms (supporting structured semantic representations) and gating mechanisms (regulating affective expression), leading to more sophisticated and contextually rich emotional generation.
-
Designing Architectures for Multimodal Affective Unification: By identifying gate proj as a shared, central locus of adaptation across different modalities (text, audio, vision), the system can serve as a unifying modulation channel. This enables consistent affective behavior across heterogeneous inputs without requiring separate architectural modifications for each modality.
-
Creating Task-Specific Structural Preferences: The system will inherently exhibit task-dependent structural preferences in its internal mechanisms (e.g., focusing on gate proj for emotion tasks vs. up proj for reasoning tasks), allowing developers to engineer models with specific, robust affective profiles tailored to the required emotional domain without relying solely on prompt engineering.
Sources
- Towards Stable Cross-Domain Depression Recognition under Missing Modalities
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
- Proximal Policy Optimization Algorithms
- Who Reasons in the Large Language Models?
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models