State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking

arXiv:2608.03425 · cs.AI · Submitted 2026-08-04 · Read on arXiv

GuangDong Police College · Department of Computer Science and Technology, School of Informatics, Xiamen University

cs.AI

Submitted: 2026-08-04

Updated: 2026-09-25

Code: https://github.com/hilhert/CSP

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 37/100

The gist: The paper "State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking" proposes the Complex State Propagator (CSP), a minimalistic recurrent architecture

Terminology

Summary

The paper State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking proposes the Complex State Propagator (CSP), a minimalistic recurrent architecture designed for deterministic state tracking tasks such as parity checking, modular counting, and parenthesis matching. The central thesis of the work is that state propagation alone is sufficient for these tasks, arguing that for a class of deterministic state tracking tasks... attention may be overkill.

Architecture and Design Principles

The CSP is designed around four core principles: Temporal integration: information from the past must be aggregated over time; Input gating: each incoming symbol should influence the state in a controlled manner; Phase accumulation: if the state is complex-valued, the natural way to represent cyclic or periodic patterns is through phase rotation; and State-Only Propagation: Hidden states are passed directly between layers, avoiding intermediate output projection overhead.

The architecture consists of stacked CSP blocks, where each block performs four specific operations:

  1. Rotate: The input z t is rotated by a learned angle theta t via an independent rotation to each complex dimension across layers, expressed as t = e i theta t z t.

  2. Recur: The rotated input is processed through a complex-valued recurrence: h t = alpha t h t-1 + gamma t t, where alpha t controls the decay of past information and gamma t scales the current input.

  3. Skip: A skip connection is used to preserve the input signal and facilitate gradient flow, defined as t = SiLU(h t) + sigma(g) z t.

  4. Normalize: The residual output is normalized element-wise by its complex modulus: h t = t / t, which projects each complex unit onto the unit circle, ensuring that information is encoded primarily in the phase.

Experimental Results

The model was evaluated on three deterministic tasks: Parity Check, Mod-3 Counting, and Parenthesis Matching. The authors report that Applied with Focal Loss, CSP achieves 100% accuracy with perfect F1 scores across canonical tasks.

Ablation studies revealed the necessity of the design components:

  • Rotation: Without rotation, the model fails entirely on all three tasks, performing at chance level, confirming that the phase accumulation mechanism is the core inductive bias that enables state tracking.

  • Structural Components: Replacing complex normalization with standard LayerNorm destroys performance on parenthesis matching, and removing block skip connections causes a noticeable drop on Mod-3 and Parenthesis. Furthermore, step-wise nonlinearities would otherwise distort the phase information accumulated across time, justifying the choice to restrict nonlinearities to block-level boundaries.

  • Focal Loss: While standard cross-entropy was sufficient for Parity and Mod-3, Focal Loss was essential for learning the minority class in the Parenthesis Matching task.

Grokking Observation

The paper observes the phenomenon of grokking across all tasks, described as long near-random performance followed by abrupt perfect generalization. The authors hypothesize that grokking is amplified in CSP due to its structured parameterization, where the model must learn precise angles and decay rates without shortcuts. They suggest that the steep gradient near the plus or minus pi boundary of the atan2 function plays a catalytic role, acting as a strong directional signal that pushes the model out of the saddle region once it approaches the decision boundary.

Improvements for AI systems

1. Integration of Complex-Valued Phase-Accumulation Layers

  • What it can do: Enables AI models to track cyclical, periodic, or modular patterns (such as clock signals, modular arithmetic in cryptography, or rhythmic musical structures) with extreme precision and minimal parameter overhead, replacing heavy attention mechanisms with efficient phase rotations.

2. Implementation of State-Only Propagation with Complex Unit-Circle Normalization

  • What it can do: Creates ultra-lightweight, low-latency recurrent models for edge devices that can perform complex formal language tasks (like syntax parsing or parenthesis matching) by eliminating the computational cost of intermediate output projections and preserving information density within the phase.

3. Structured Phase-Based Parameterization (utilizing atan2 gradient signals)

  • What it can do: Accelerates the grokking process in small-scale models, allowing them to transition rapidly from rote memorization to perfect algorithmic generalization on tasks like parity checking and counting by providing strong directional gradients during training.

4. Phase-Preserving Nonlinearity Constraints (Restricting nonlinearities to block boundaries)

  • What it can do: Prevents the distortion of temporal state information in high-fidelity signal processing, allowing AI systems to maintain stable, long-term tracking of complex waveforms or electromagnetic signals without the signal washing effect caused by standard step-wise nonlinearities.

5. Hybrid CSP-Focal Loss Architectures for Rare Event Detection

  • What it can do: Improves the detection of minority-class state transitions in imbalanced time-series data (e.g., identifying a single specific error pattern in a massive stream of valid network traffic) by forcing the model to prioritize the phase shifts associated with rare, critical events.

Sources

Related papers