RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates".
Jane: The paper was written by Donggen Li from Sichuan University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everyone. Today we are digging into a brand new arXiv paper, and the title alone is a mouthful: "RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates."
Jane: It really is a mouthful, Tom. But I promise the ideas underneath are actually pretty intuitive once you unpack them. And we have our full crew here today to do just that. Lu, you’ve been staring at this thing, what’s the one-sentence pitch?
Lu: The one-sentence pitch is that when a vision-language model looks at a picture and some text at the same time, the way it encodes "where" things are is often mathematically sloppy, and this paper proposes a cleaner rule for when spatial coordinates should even be trusted.
Meng: And I’ll jump in because my first question is always, does this change the model’s architecture or just the math? And the answer here is mostly the math, which is good news for anyone who wants to actually run this thing.
Tom: So it’s not a new model, it’s a new way of positioning tokens inside an existing model architecture. That’s the rotary part, right? The RoPE part of the title.
Jane: Exactly. RoPE, or rotary position encoding, is how modern models tell tokens apart by their position. It’s like giving every word and every image patch a little clock hand that spins based on where it sits in the sequence.
Lu: And the problem this paper tackles is that we’ve been spinning those clock hands for image patches using coordinates that don’t always mean what we think they mean. Two patches from different images might have the same "screen coordinate" but no real geometric relationship.
Meng: Which is a fancy way of saying the model might think two things are close together when they’re actually from completely different photos. That’s a recipe for confusion.
Tom: So the paper’s fix is to be more careful about when we apply that spatial rotation, and that’s the "relation-stratified" part of the title. It’s about grouping pairs of tokens by whether they actually share a meaningful spatial relationship.
Jane: And we’ll get into all the details, but I love that the paper is honest about its own limits. It says outright, this is a theoretical framework, we haven’t run the giant benchmarks yet. That’s rare and refreshing.
Lu: It’s a "here’s the problem, here’s the math, here’s how you’d test it" kind of paper. And for researchers like me, that’s gold, because it gives us a clear roadmap.
Tom: Alright, we’re going to start unpacking the actual mechanics next. But first, let’s just sit with the fact that a paper this technical is asking a very human question: when should a model trust that two things are actually near each other?
Summary: Jane: So we’ve established that "RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates" is about being careful with spatial math in vision-language models. Now let’s talk about what the paper actually proposes, because it’s not just one fix, it’s a whole framework.
Tom: Right, and the core idea is that the model should treat different kinds of token pairs differently. Text-to-text pairs, image-patch-to-image-patch pairs from the same image, and pairs that cross modalities or cross different images, they all get different positional treatment.
Lu: And the key word there is "stratified." The paper splits attention into groups based on the relationship between the query and the key. Same image? That’s one group. Text to text? Another group. Anything else, like text looking at an image or one image looking at a different image, that’s a third group.
Meng: And the reason that matters is that the current standard approach, which is called M-RoPE, applies the same spatial rotation math to all of those pairs. It just assumes every token has a height and width coordinate, even when that coordinate is meaningless.
Jane: Meaningless is a strong word, but the paper proves it. It shows that if you take two different images and subtract their coordinates, the result changes depending on how you set up the coordinate system. It’s not a stable, real property of the images.
Tom: So it’s like if I tell you my house is at ten on a map and your house is at twenty you might think we’re close. But if my map uses miles and yours uses kilometers, that subtraction is garbage.
Lu: That’s exactly the gauge argument in the paper. They call it "gauge freedom," which is a physics term. Each image can independently shift its origin, and that shift changes the apparent distance between patches from different images. So the raw number isn’t intrinsic.
Meng: And here’s the part I really appreciate as an engineer. The paper doesn’t just say "this is broken." It gives you a concrete alternative. For pairs that don’t have a valid spatial relationship, you just don’t apply the spatial rotation. You apply the temporal rotation, which is about sequence order, but you skip the height and width part.
Jane: And then, to make sure those unrotated pairs don’t dominate the attention, they use a separate normalization step. Each group gets its own softmax, and then a gate decides how much total attention mass each group gets.
Tom: So it’s not that the model ignores cross-image pairs entirely. It still attends to them, it just doesn’t pretend they have a spatial relationship. That’s a really clean way to think about it.
Lu: Clean, and mathematically principled. They even prove that if all the scores were the same, this whole stratified structure collapses back to the standard global softmax. So it’s a generalization, not a completely different beast.
Meng: And that’s important for anyone who wants to retrofit an existing model. You’re not throwing away the old behavior, you’re adding structure on top of it.
Jane: We’re going to dig into the second big piece of the paper next, which is this "traversal coordinate" idea. That’s the part that deals with how the model measures distance through a sequence that mixes text, images, and video.
Improvements: Tom: So we’ve covered the spatial side of "RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates." But the paper has a second big idea, and it’s about time, or at least about how the model measures progress through a sequence.
Jane: Right, and this is the "representation-aware traversal coordinates" part. The problem is that when you have a text token, then an image, then more text, the model needs a single number to represent "how far along" each token is. And the standard way of doing that has some quirks.
Lu: The quirk is that an image’s position advance often depends on its resolution. If you have a big image, it pushes the next text token further away than a small image would. That might be fine, but the paper argues it’s happening for the wrong reason.
Meng: The wrong reason being that the advance is computed from the maximum coordinate of the image grid, which is a side effect of the local spatial layout, not a deliberate choice about how much context the image should consume.
Tom: So the paper proposes a cleaner rule. Text tokens each take one step. An image takes a step that scales with its linear size, not its area. And a video takes a step for each temporal slice, plus a smaller spatial correction.
Jane: And the key word there is "ordered." The paper makes a distinction between axes that are ordered, like time in a video, and axes that are parallel, like the height and width of an image. Ordered axes add up. Parallel axes only contribute a sublinear amount.
Lu: That’s the part I find elegant. If you split a video into two halves, the total extent of the two halves should equal the extent of the whole video. That’s additivity. And the paper proves their construction satisfies that property exactly.
Meng: And it also fixes a weird artifact of the naive approach. If you just take the cube root of the total token count, splitting a video into two pieces changes the total extent. That’s a bug, and this paper’s coordinate system doesn’t have it.
Tom: So it’s a more principled ruler for measuring context distance. And the paper is careful to say this is "representation time," not physical time. It’s about how many tokens the model actually processes, not how many seconds the video lasts.
Jane: That’s a crucial distinction, because a video with more frames per second will have more temporal tokens, and that will stretch the representation time. The paper is upfront that this is a design choice, not a claim about physics.
Lu: And they even offer an optional variant that uses actual timestamps if you want physical time. But the default is representation time, which is content-independent and deterministic given a fixed tokenizer.
Meng: Which is great for reproducibility. You don’t need to run the model to compute the coordinates. You just need the tokenizer output and the grid sizes. That’s a very engineer-friendly property.
Tom: Alright, so we have the spatial fix and the temporal fix. Next we need to talk about how this all fits together in practice, and what the paper says about actually testing it.
First Page: Jane: We’re looking at the opening of "RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates," and I want to go back to the abstract, because it sets up the whole paper with a really clear statement of intent.
Tom: The abstract basically says two things are structurally ambiguous in current multimodal position encoding. First, a spatial displacement between two different images might be numerically available but not geometrically meaningful. Second, the way we advance position across visual blocks is often inherited from local coordinate extrema rather than being a deliberate choice.
Lu: And I love that the abstract immediately gives the remedy. It says, "A gauge argument shows why raw cross-instance coordinate subtraction depends on independent chart choices." That’s the physics-flavored proof we talked about earlier.
Meng: And then it says something that I think is the most practical sentence in the whole paper: "A self-aligned analysis and a distributional extension show when replacing a missing spatial relation by zero rotation favors the unregistered branch." That’s a warning that the naive fix, just setting the rotation to zero, is actually biased.
Jane: Right, because if you just say "no spatial relation, so no rotation," the identity rotation happens to be the one that maximizes self-similarity. The paper proves that. So the naive fix is secretly giving unregistered pairs a boost.
Tom: That’s such a subtle point, and it’s the kind of thing that would never show up in a benchmark but could quietly skew a model’s behavior. The paper catches it with math.
Lu: And the response is the relation-stratified attention. You don’t let the unregistered pairs compete with registered pairs in the same softmax. You give them their own normalization, and then you use a separate, spatial-neutral statistic to decide how much mass each group gets.
Meng: The paper calls that the "common evidence" gate, and it’s built from a score that ignores height and width entirely. So the spatial displacement can’t influence how much total attention an unregistered group receives. That’s the calibration fix.
Jane: And the abstract also mentions the traversal coordinates, which we covered, and it emphasizes that the method adds no new learned parameters under a fixed configuration. That’s a big deal for adoption.
Tom: It also says, right at the end, "We establish the theoretical properties and a validation protocol without claiming empirical superiority." That’s the paper being honest about what it is and isn’t.
Lu: And that honesty is why I trust the math. They’re not overselling. They’re saying, here’s a structural risk, here’s a fix, here’s how to test it. That’s how good science should work.
Meng: And from my side, the fact that they include a detailed validation protocol, with specific probes for gauge invariance and traversal consistency, means someone can actually implement this and check it without guessing.
Tom: So the first page alone gives us the problem, the proposed solution, and the caveats. That’s a dense opening. We’ve got one more segment to pull it all together.
Conclusion: Tom: We’ve spent the whole show on "RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates," and I think we can all agree it’s a paper that rewards careful reading.
Jane: It really does. We started with the title, which is intimidating, but the core message is simple: don’t pretend two things are spatially related when they’re not, and measure context distance with a ruler that respects the difference between ordered time and parallel space.
Lu: And the paper backs that up with real theorems. The gauge argument shows why cross-image coordinates are unreliable. The null-relation analysis shows why the naive zero-rotation fix is biased. And the traversal construction has provable additivity and consistency properties.
Meng: From an implementation standpoint, I’m impressed that it adds no learned parameters and keeps the same asymptotic complexity. The overhead is a few extra normalization states per query row, which is manageable in a fused kernel.
Tom: And the paper is refreshingly honest about what it doesn’t do. No large-scale benchmarks, no claims of state-of-the-art accuracy. Just a clear problem statement, a principled solution, and a detailed protocol for validation.
Jane: That’s the kind of paper that moves the field forward even before the big empirical results come in, because it gives researchers a shared language and a set of checks to run.
Lu: And I think the impact could be significant. As models process more interleaved images and videos, the way we encode position becomes more important. This paper offers a way to do that that’s mathematically grounded rather than ad hoc.
Meng: The one thing I’ll be watching for is the follow-up. The paper mentions a "subsequent version" with large-scale validation. If the empirical results hold up, this could become a standard component in the next generation of vision-language models.
Tom: Well, we’ll be watching the arXiv feed for that. For now, we’ve got a solid theoretical contribution that’s worth a read, especially if you work on multimodal systems.
Jane: And that’s a wrap on "RIG-RoPE." Thanks to Lu and Meng for joining us, and to all our listeners for sticking with us through the rotary geometry and the traversal coordinates.
Tom: Next up, we’ve got a paper on efficient video understanding that I think is going to be a fun one. Until then, keep your coordinate systems honest, and we’ll see you on the next episode.
Sichuan University
cs.CL, cs.CV
Submitted: 2026-05-21
Updated: 2026-08-25
Comments: 24 pages, 2 figures. Major theoretical revision: reformulated cross-instance geometry, null-relation analysis, relation-stratified normalization, and representation-aware traversal coordinates; expanded related work and implementation details. Preliminary technical report; empirical validation is left to future work
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 58/100
The gist: The paper introduces RIG-RoPE, a framework for multimodal rotary positional encoding that addresses two structural ambiguities in existing approaches.
Key concepts
- Relation-Stratified Attention
- A method that groups token pairs based on their actual relationship (e.g., same image vs. cross-image). This prevents the model from incorrectly assuming a spatial connection when one doesn't exist.
- Rotary Position Encoding (RoPE)
- A technique used in modern models to encode the position of tokens. It mathematically assigns a 'spin' or clock hand to every word or image patch based on where it appears in the sequence.
- Gauge Freedom
- A physics argument showing that raw coordinate subtraction between patches from different images is unreliable because the resulting distance changes depending on how the coordinate system's origin is set up.
- Representation-Aware Traversal Coordinates
- A method for calculating context distance in mixed media. It proposes a principled way to measure progress through a sequence (text, images, video) that respects the difference between ordered time and parallel space.
Terminology
Summary
The paper introduces RIG-RoPE, a framework for multimodal rotary positional encoding that addresses two structural ambiguities in existing approaches. First, the authors identify that a cross-instance spatial displacement may be numerically available under preprocessing conventions while lacking an intrinsic geometric meaning unless a shared chart or registration is explicitly declared.
Second, they note that the scalar advance across visual blocks is often inherited from local coordinate extrema rather than defined as a representation-level traversal measure.
The method combines instance-local rotary geometry, relation-stratified attention, and representation-aware traversal coordinates.
The paper argues that the central issue is not only how to choose better static coordinates, but when a coordinate difference should be treated as geometrically meaningful.
The authors explain that "a visual patch has a height/width coordinate relative to a chart induced by its preprocessing pipeline. A text token does not inhabit that chart. Patches from different images may share a conventional normalized screen coordinate, yet that coordinate is generally not invariant to independent cropping, padding, resizing, or origin choices. They conclude that
a cross-instance displacement becomes geometrically usable only when the model or task explicitly declares a shared chart, registration, or canonical-coordinate contract."
A second issue concerns the temporal subspace: The scalar coordinate used across interleaved blocks should express how far the model traverses through its represented context, not the raster length of a visual block and not necessarily physical or cognitive time.
The authors note that M-RoPE's approach, where the next modality starts after the maximum position ID of the preceding modality,
produces a resolution-dependent global advance
that is produced indirectly by the maximum of local coordinate axes
and does not state which axes are ordered, which tokens are parallel, or how video duration should behave under temporal partitioning.
The framework follows two simple rules: "Apply height/width rotary encoding only when a shared visual chart is declared. Do not let spatially registered and unregistered relations compete through uncalibrated H/W phases. Measure global context separation on a traversal axis that is additive along ordered slices and sublinear along parallel spatial slices."
The paper proves that raw cross-instance displacement is not intrinsic. Under independent translations of two coordinate charts, the apparent cross-instance displacement becomes (x j(n) + b n) − (x i(m) + b m) = (x j(n) − x i(m)) + (b n − b m). Because b m and b n are independent, b n − b m can be arbitrary.
Therefore, the value of Eq. (14) is determined by chart choices rather than by instance-independent geometry.
The paper proves that setting an unregistered displacement to zero (identity rotation) creates a structural bias. For identical content vectors u in query and key, s A = u T R(ω∆)u = ∥u∥2 cos(ω∆)
for a registered pair, while s B = u T Iu = ∥u∥2
for an unregistered pair with identity fallback, so s B ≥ s A
with strict inequality whenever ω∆ ∉ 2πZ and u ≠ 0. The authors emphasize this is a deterministic existence result for the self-aligned limiting case
and does not state that identity rotation maximizes q T R(θ)k for arbitrary q and k.
The paper proves that static assignment incompatibility
arises when one static H/W coordinate assignment and shared non-degenerate RoPE frequencies are used for all pairs, without relation-conditioned routing.
The requirements of intra-instance spatial faithfulness
and text–vision spatial invariance
cannot both hold for all content vectors, so the operator must depend on relation type, not only on static token coordinates.
The paper introduces a multimodal traversal coordinate τ that measures represented context extent on the RoPE axis
and is not defined as wall-clock time, human reading time, semantic information content, or flattened token count.
The traversal structure satisfies: Ordered-axis additivity
(concatenating consecutive ordered slices adds their extents), Parallel-axis sublinearity
(increasing simultaneous spatial tokens increases extent no faster than their characteristic linear spatial scale), Content independence
(same modality label and post-tokenization grid receive the same coordinates), Text compatibility
(each text token contributes one unit), and Phase-range compatibility
(anchored to the base model's position range).
For each block B b with U b ordered slices, the traversal extent is:
D b = d⋆ m b Σ u=1 U b (l(H b,u, W b,u) / l⋆ m b) β m b
where l(H, W) is a characteristic spatial scale satisfying scale covariance l(cH, cW) = c·l(H, W), with useful choices including l area = √(HW), l diag = √(H2 + W2), and l max = max(H, W). The exponent satisfies 0 ≤ β m ≤ 1.
For text: D b T = 1. For images with β I = 1: D b I = D I⋆ · l(H b, W b)/l(H I⋆, W I⋆), which with l = √(HW) becomes D b I = D I⋆ √(H b W b / (H I⋆ W I⋆)). For video: D b V = d⋆ V Σ f=1 F b (l(H b,f, W b,f)/l(H V⋆, W V⋆)) β V, where F b ≡ the number of temporal tokens in video block B b after visual tokenization.
The paper proves: Pure-text preservation
(text-only sequences recover ordinary one-dimensional relative offsets exactly), Image simultaneity
(any two patches in the same image have τ j − τ i = 0), Temporal-partition additivity
(partitioning a video preserves total extent), Local video-scale consistency
(the offset between temporal slices f and f+k is independent of total video length), and Structural monotonicity.
The paper explicitly states: The formulation does not impose the cross-modal boundary condition D V(F = 1, H, W) = D I(H, W).
A one-temporal-token video remains in the video observation regime: its modality label, tokenizer path, calibration constant, and exponent are those of video.
The paper distinguishes representation time from physical time: If the same physical clip is tokenized into more temporal tokens, its model-space extent increases. This behavior is intentional in the core method.
An optional physical-time-aware extension is provided: D V b,phys = d⋆ V Σ (∆t f/t 0)(l(H b,f, W b,f)/l(H V⋆, W V⋆)) β V.
Each visible key is assigned to exactly one relation class: J i T = j ∈ C i: M i = M j = Text
(native text order), J i S = j ∈ C i: SharedChart G(i, j)
(declared spatial relation), and J i U = C i (J i T ∪ J i S)
(unregistered relation). Under the default instance-local contract, SharedChart G(i, j) reduces to M i = M j = Vision and I i = I j.
For text relations: s T ij = (1/√d)⟨q i, R 1D(∆τ ij)k j⟩. For spatially registered visual relations: s S ij = (1/√d)[⟨q t i, R t(∆τ ij)k t j⟩ + ⟨q h i, R h(∆h G ij)k h j⟩ + ⟨q w i, R w(∆w G ij)k w j⟩]. For unregistered relations: s U ij = (1/√d)[⟨q t i, R t(∆τ ij)k t j⟩ + ⟨q h i, k h j⟩ + ⟨q w i, k w j⟩]. The identity on H/W-selected blocks is an explicitly typed absence of H/W transformation, not a claim that ∆h = ∆w = 0.
Each group is normalized internally: a r ij = exp(s r ij) / Σ k∈J i r exp(s r ik), and o r i = Σ j∈J i r a r ij v j. The spatial phase ranks keys only against keys governed by the same spatial semantics.
The common score is c ij = (1/√d)⟨q i, G(∆τ ij)k j⟩, where G(∆τ) applies traversal rotation to blocks in Ω t and identity to blocks in Ω h ∪ Ω w.
This statistic is relation independent, contains no H/W coordinate difference, remains sensitive to traversal order, and is well defined for contiguous, interleaved, and head-wise frequency assignments.
The default evidence is the common-score LogSumExp: e r i = log Σ j∈J i r exp(c ij). With fixed positive relation priors π r, the gate is γ r i = π r exp(e r i) / Σ u∈A i π u exp(e u i). The fixed default uses uniform π r = 1 and introduces no learned gate parameters.
The final head output is o i = Σ r∈A i γ r i o r i.
An optional balanced variant is defined: e r i,bal (τ g) = τ g log[(1/J i r) Σ j∈J i r exp(c ij/τ g)], which encodes the deliberate prior that groups with equal average evidence deserve equal total mass regardless of key count.
The paper proves: Full connectivity and normalization
(every visible key receives positive weight summing to one), Common-score consistency
(if all relation-specific scores reduce to the common score, the final weight equals standard global softmax on c), Explicit relation-mass control
(Σ j∈J i r α ij = γ r i), and Gauge invariance of unregistered interactions.
Traversal coordinates are preprocessing metadata computed after visual tokenization and before the transformer stack.
The construction is a prefix scan over block metadata and is O(L) in the packed sequence length.
The implementation requires only M ∈ 0,1 L, I ∈ N 0 L, τ ∈ R L
plus ordinary H/W coordinates. In a FlashAttention-style tiled kernel, the default relation class is computed in registers from (M i, M j, I i, I j).
The complexity remains O(L2d)
asymptotically with O(L)
persistent metadata overhead under the default instance-local contract.
The paper specifies lightweight analytic and model-level checks including: a Null-Privilege and Mass-Control Probe
with synthetic vectors, a Gauge Perturbation Test
applying independent coordinate-origin shifts, Traversal-Coordinate Consistency Tests
verifying text offsets, image simultaneity, temporal-partition additivity, and local video-scale consistency, and a Small-Scale Model Patch
evaluating text-only reasoning, single-image spatial VQA, interleaved multi-image QA, and video pairs with equal total token counts but different (F, H, W) factorizations.
The paper acknowledges: This work establishes a geometric argument and a concrete algorithmic rule, but it does not yet include large-scale empirical results.
Additional limitations include Empirical uncertainty
(theorems identify structural bias, not net effect on downstream accuracy), Gate sufficiency and cardinality prior,
Changed attention semantics
(relation-stratified normalization intentionally changes attention semantics, so pretrained weights are not guaranteed to remain calibrated after a training-free patch
), Unregistered H/W content,
Traversal calibration
(the power-law form is a structured inductive bias rather than a uniquely derived law
), Representation-time semantics
(not invariant to sampling density), Cross-modal boundary,
Chart-contract dependence,
Video and registration boundaries,
and Kernel engineering.
The paper concludes: RIG-RoPE is based on a simple principle: rotary phases and softmax denominators should reflect meaningful relation structure.
The method preserves full one-dimensional RoPE for text–text interactions and M-RoPE H/W rotations wherever G declares a shared visual chart, while replacing the temporal-axis coordinate by τ.
Under a fixed configuration, the method adds no learned parameters and has explicit invariance, mass-control, temporal-partition, and local-scale consistency properties; empirical superiority remains an open question.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
-
Replace flat softmax with relation-homogeneous softmaxes: Partition attention keys into three groups—textual (T), spatially registered (S), and unregistered (U)—based on modality and instance/chart identity.
-
Apply separate within-group normalization: Compute attention weights independently within each group, preventing unregistered keys from competing with registered spatial keys through a shared denominator.
-
Use a common H/W-neutral gate: Allocate total attention mass across groups using a LogSumExp over a relation-independent score that applies traversal rotation only to temporal blocks, with identity on H/W blocks.
-
Declare shared-chart contracts explicitly: Use a function
SharedChartG(i,j)that returns true only when both tokens are visual and a common coordinate chart is declared (e.g., same instance ID, or explicit registration). -
Apply H/W RoPE only for spatially comparable pairs: For same-instance visual pairs, retain full M-RoPE H/W rotations. For cross-instance or text–vision pairs, use identity on H/W blocks—not as a zero displacement, but as an explicitly typed absence of spatial transform.
-
Ensure gauge invariance: Independent coordinate-origin shifts of different visual instances must not change any logits, within-group weights, or gate outputs.
-
Replace axis-maximum advance with an explicit scalar metric: For each block, compute traversal extent as:
-
Text:
D = 1per token. -
Image:
D = D I* · (l(H,W)/l I*)where l is a characteristic scale (e.g., √(HW)). -
Video:
D = d V* · Σ f (l(H f,W f)/l V*) βVover temporal tokens, with 0 ≤ βV ≤ 1. -
Assign coordinates: Text tokens and image patches share their block center; video temporal slices are placed additively with sublinear spatial correction.
-
Preserve properties: Pure-text offsets must exactly match 1D RoPE; image patches must have zero traversal offset; video partition must be additive; local slice offsets must be independent of total video length.
-
Compute gate evidence as
e i r = log Σ j∈J i r exp(c ij)wherec ijis the common score using traversal rotation on temporal blocks and identity on H/W blocks. -
Set group mass as
γ i r = π r·exp(e i r) / Σ u π u·exp(e i u)with uniform priors by default. -
Guarantee exact recovery: If all relation-specific scores equal the common score, the final attention weights must exactly match standard global softmax.
-
Implement cardinality-normalized evidence:
e i r,bal = τ g · log((1/J i r) · Σ j exp(c ij/τ g))to give equal total mass to groups with equal average evidence, independent of key count.
-
Eliminate cross-instance spatial leakage: When processing multiple images in one context, the model no longer treats normalized screen coordinates across different images as geometrically meaningful. It only applies H/W rotations within the same image instance, preventing artificial spatial proximity between unrelated patches.
-
Prevent null-relation privilege: Unregistered text–image or cross-image pairs no longer receive inflated attention mass merely because their H/W phase is identity. The relation-stratified normalization ensures that a missing spatial relation cannot outrank a valid spatial relation through self-alignment bias.
-
Maintain text-only performance exactly: On pure-text inputs, traversal coordinates reduce to standard 1D RoPE offsets, so the model's language understanding is unchanged.
-
Preserve single-image spatial reasoning: Same-instance visual tokens retain full H/W rotary encoding, so tasks like
what is above the red car?
continue to work correctly. -
Handle video with proper temporal structure: The model treats video as ordered temporal slices with simultaneous spatial tokens. Splitting a video into two segments yields additive traversal extents, and local temporal offsets are independent of total clip length. A one-temporal-token video is not forced to behave like an image.
-
Provide explicit control over cross-group mass: The common H/W-neutral gate lets the model decide how much attention to allocate to text, registered visual, and unregistered groups based on traversal order alone—not on potentially uncalibrated spatial phases.
-
Support both global-mass and group-balanced semantics: The default gate preserves key-count effects (larger groups get more mass if scores are equal), while the optional balanced variant gives equal mass to groups with equal average evidence.
-
Remain computationally efficient: The method adds no learned parameters, requires only O(L) metadata (modality, instance ID, traversal coordinate), and can be implemented in a fused FlashAttention-style kernel with at most two active relation groups per query row.
-
Interleaved multi-image QA: The model can reason about
compare the leftmost object in image 1 with the rightmost in image 2
without confusing their independent coordinate systems. -
Text–image–text reasoning: The model correctly orders text tokens around an image block using traversal extent, avoiding the ambiguity of whether an image advances the position counter by one or by its grid size.
-
Video temporal reasoning: The model can answer
what happened after the first 5 seconds?
with consistent relative ordering regardless of whether the video is tokenized into 10 or 30 temporal tokens. -
Long-context multimodal tasks: The explicit traversal metric prevents resolution-dependent global advances from distorting later text positions, improving coherence in long interleaved documents.
The improved system can be verified to satisfy:
-
Gauge invariance under independent chart translations.
-
Exact text-only offset recovery.
-
Zero traversal offset within an image.
-
Temporal-partition additivity for videos.
-
Local video-scale consistency (offset between slices f and f+k is independent of total length).
-
Exact global-softmax recovery when all branch scores equal the common score.
-
Full connectivity with normalized final weights summing to 1.
Sources
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- Qwen3-VL Technical Report
- MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
- Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding
- VRoPE: Rotary Position Embedding for Video Large Language Models
- VideoRoPE: What Makes for Good Video Rotary Position Embedding?
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
- Qwen3-Omni Technical Report
- Qwen3.5-Omni Technical Report
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering