State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
GuangDong Police College · Department of Computer Science and Technology, School of Informatics, Xiamen University
cs.AI
Submitted: 2026-08-04
Updated: 2026-09-25
Code: https://github.com/hilhert/CSP
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 37/100
The gist: The paper "State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking" proposes the Complex State Propagator (CSP), a minimalistic recurrent architecture
Terminology
Summary
The paper State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
proposes the Complex State Propagator (CSP), a minimalistic recurrent architecture designed for deterministic state tracking tasks such as parity checking, modular counting, and parenthesis matching. The central thesis of the work is that state propagation alone is sufficient
for these tasks, arguing that for a class of deterministic state tracking tasks... attention may be overkill.
Architecture and Design Principles
The CSP is designed around four core principles: Temporal integration: information from the past must be aggregated over time
; Input gating: each incoming symbol should influence the state in a controlled manner
; Phase accumulation: if the state is complex-valued, the natural way to represent cyclic or periodic patterns is through phase rotation
; and State-Only Propagation: Hidden states are passed directly between layers, avoiding intermediate output projection overhead.
The architecture consists of stacked CSP blocks, where each block performs four specific operations:
-
Rotate: The input z t is rotated by a learned angle theta t via
an independent rotation to each complex dimension across layers,
expressed as t = e i theta t z t. -
Recur: The rotated input is processed through a complex-valued recurrence: h t = alpha t h t-1 + gamma t t, where alpha t controls the decay of past information and gamma t scales the current input.
-
Skip: A skip connection is used to
preserve the input signal and facilitate gradient flow,
defined as t = SiLU(h t) + sigma(g) z t. -
Normalize: The residual output is
normalized element-wise by its complex modulus: h t = t / t,
whichprojects each complex unit onto the unit circle, ensuring that information is encoded primarily in the phase.
Experimental Results
The model was evaluated on three deterministic tasks: Parity Check, Mod-3 Counting, and Parenthesis Matching. The authors report that Applied with Focal Loss, CSP achieves 100% accuracy with perfect F1 scores across canonical tasks.
Ablation studies revealed the necessity of the design components:
-
Rotation:
Without rotation, the model fails entirely on all three tasks, performing at chance level,
confirming thatthe phase accumulation mechanism is the core inductive bias that enables state tracking.
-
Structural Components: Replacing complex normalization with standard LayerNorm
destroys performance on parenthesis matching,
and removing block skip connectionscauses a noticeable drop on Mod-3 and Parenthesis.
Furthermore,step-wise nonlinearities would otherwise distort the phase information accumulated across time,
justifying the choice to restrict nonlinearities toblock-level boundaries.
-
Focal Loss: While standard cross-entropy was sufficient for Parity and Mod-3, Focal Loss was
essential for learning the minority class
in the Parenthesis Matching task.
Grokking Observation
The paper observes the phenomenon of grokking
across all tasks, described as long near-random performance followed by abrupt perfect generalization.
The authors hypothesize that grokking is amplified in CSP due to its structured parameterization,
where the model must learn precise angles and decay rates
without shortcuts. They suggest that the steep gradient near the plus or minus pi boundary of the atan2 function plays a catalytic role,
acting as a strong directional signal that pushes the model out of the saddle region once it approaches the decision boundary.
Improvements for AI systems
1. Integration of Complex-Valued Phase-Accumulation Layers
- What it can do: Enables AI models to track cyclical, periodic, or modular patterns (such as clock signals, modular arithmetic in cryptography, or rhythmic musical structures) with extreme precision and minimal parameter overhead, replacing heavy attention mechanisms with efficient phase rotations.
2. Implementation of State-Only Propagation
with Complex Unit-Circle Normalization
- What it can do: Creates ultra-lightweight, low-latency recurrent models for edge devices that can perform complex formal language tasks (like syntax parsing or parenthesis matching) by eliminating the computational cost of intermediate output projections and preserving information density within the phase.
3. Structured Phase-Based Parameterization (utilizing atan2 gradient signals)
- What it can do: Accelerates the
grokking
process in small-scale models, allowing them to transition rapidly from rote memorization to perfect algorithmic generalization on tasks like parity checking and counting by providing strong directional gradients during training.
4. Phase-Preserving Nonlinearity Constraints (Restricting nonlinearities to block boundaries)
- What it can do: Prevents the distortion of temporal state information in high-fidelity signal processing, allowing AI systems to maintain stable, long-term tracking of complex waveforms or electromagnetic signals without the
signal washing
effect caused by standard step-wise nonlinearities.
5. Hybrid CSP-Focal Loss Architectures for Rare Event Detection
- What it can do: Improves the detection of minority-class state transitions in imbalanced time-series data (e.g., identifying a single specific error pattern in a massive stream of valid network traffic) by forcing the model to prioritize the phase shifts associated with rare, critical events.
Sources
- Efficiently Modeling Long Sequences with Structured State Spaces
- Diagonal State Spaces are as Effective as Structured State Spaces
- Simplified State Space Layers for Sequence Modeling
- Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- The doubly librating Plutinos
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Provable Benefits of Complex Parameterizations for Structured State Space Models
- Hungry Hungry Hippos: Towards Language Modeling with State Space Models
- Retentive Network: A Successor to Transformer for Large Language Models
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Toward Horizon-scale Accretion Onto Supermassive Black Holes in Elliptical Galaxies
- Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
- Linear Transformers Are Secretly Fast Weight Programmers
- Gated Delta Networks: Improving Mamba2 with Delta Rule
- Test-time regression: a unifying framework for designing sequence models with associative memory
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection