QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting".
Jane: The paper was written by Junkai Lin, Siqi Hou and Raymond Lee from Faculty of Science and Technology, Beijing Normal-Hong Kong Baptist University and Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science (BNBU).
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We just introduced this paper, and I'll admit the title alone made me do a double take. That name packs a lot of loaded words into one line. But the more I look at it, the more each piece maps to something concrete in the work.
Jane: It does. Quantformer is a transformer built specifically for numerical data rather than language, so it uses linear projections instead of word embeddings. The chaotic gating part means the network can switch between a smooth, calm response and a wild, oscillation-driven one depending on the input. And the volatile time series part is the domain: electricity loads, transformer temperatures, stock indices — things that move in sudden bursts.
Tom: The authors are Junkai Lin, Siqi Hou, and Raymond Lee, from Beijing Normal-Hong Kong Baptist University in Zhuhai, with a provincial key lab for data science behind them. Raymond Lee's name jumps out at me, because the oscillator at the heart of the method comes from his own earlier work. This is decades of his research feeding into a modern transformer framework.
Lu: Exactly. The Lee Oscillator is a discrete-time system where an excitatory unit and an inhibitory unit push against each other, and with the right parameters that tension produces genuinely chaotic behavior. Lee has been building neural architectures around that idea since at least the early 2000s. So the paper sits on a very deliberate research tree, not a random mashup.
Meng: And linked to that lineage is how upfront they are about the quantum and fractal language. They say explicitly that "quantum-fractal-inspired" is a computational analogy, not a formal quantum-mechanical derivation. The superposition is just a soft weighted average of eight oscillator families, and the fractal side is shorthand for multi-scale responses.
Jane: That honesty matters, because the words "quantum" and "fractal" usually set off alarm bells in forecasting circles. By stating exactly what is and isn't mathematics here, they invite scrutiny rather than avoiding it. There's also an IEEE copyright notice on the paper, so this is clearly aimed at a peer-reviewed venue, not just a preprint posting.
Lalam: Stepping back, that careful naming is how chaotic dynamics gets taken seriously in applied fields. The history of chaotic neural networks has its share of hype and overselling. So when a paper states its analogies plainly, it gives reviewers and readers a clean target — they can argue with the results instead of the vocabulary.
Tom: So the title is bold but the claims are measured. The open question is what all this chaos actually buys you, and that's where the abstract's numbers start to matter.
Summary: Tom: After unpacking the title, the natural next step is the abstract and what it actually promises. The central claim is that transformer forecasters handle long-range dependencies beautifully, but their feed-forward blocks rely on smooth, static activations that can't react to abrupt regime changes.
Jane: And the proposed fix has three main pieces. First, a Quantformer-style encoder that processes the numerical input directly with a linear embedding, skipping the whole NLP pipeline. Second, the Lee Oscillator activation, which runs each incoming value through a short chaotic time evolution and compresses it with Max-over-Time pooling. Third, a learnable gate that blends that chaotic response with the standard smooth GELU activation.
Lu: The Max-over-Time step is worth pausing on, because it keeps things tractable. If you let the oscillator trajectory run for dozens of internal steps, you'd get dozens of values per input and the computation would explode. Instead the model takes the maximum over those steps, so the dimension stays the same and what you're extracting is essentially the strongest response the oscillator produced.
Tom: And the headline results are genuinely strong. On the ETTh2 dataset, which is the harder electricity transformer temperature benchmark, the 24-step horizon shows a 43 point 9 percent MSE improvement over the HAT baseline and 41 point 3 percent over COTN. Those are large jumps for a forecasting paper.
Meng: Context matters here. ETTh2 comes from the State Grid Corporation of China, and it's known for being noisier and more volatile than its sibling dataset ETTh1. The fact that the biggest gains land on the most volatile data is exactly what the authors would predict if their mechanism works as intended.
Jane: There's also a subtle pattern in the error metrics across the tables. The MSE improvements are consistently larger than the MAE ones. Since MSE squares the error, it punishes big mistakes disproportionately, so a large MSE drop means the model is specifically eliminating the worst forecasts rather than shaving a little off everything.
Lalam: That error profile matters for real deployment. A grid operator or a trader can live with small errors; the expensive events are the occasional catastrophic misses. So even beyond the average improvement, the shape of the improvement is arguably the better story — the model is more reliable precisely when things go haywire.
Tom: Which leaves an obvious mechanical question on the table. What does a chaotic oscillator actually do to a single number as it passes through the network?
Improvements: Tom: The abstract gave us those headline numbers, but to understand what's genuinely new here, we need to place the work relative to what already existed. Chaotic activations in forecasters aren't brand new — the paper builds directly on COTN, which already used a Lee Oscillator, Max-over-Time pooling, and gated fusion.
Jane: Right. So the first real innovation is the mixture. Instead of one fixed oscillator, the framework uses eight different Lee Oscillator families with different parameter settings, and the weights on those families are learned. A single oscillator gives you one response shape; eight of them let the network switch between response styles depending on the local data regime.
Lu: And those eight families genuinely behave differently. The parameter table shows some types running with small coefficients and gentle dynamics, while others sit right at the edge of chaotic behavior. The bifurcation plots in the paper display distinct trajectories for each type, so the mixture is doing real work — it can choose a wild response for a volatile stretch and a calm one for a quiet stretch.
Meng: The second innovation is architectural. COTN hung its oscillator unit on a standard transformer, but this paper adopts the Quantformer-style backbone, with direct linear embedding and no masking or autoregressive decoding. And the authors make a strong claim about this: simply introducing a chaotic unit is not sufficient, because how you integrate it into the numerical backbone matters just as much.
Tom: The ablation study backs that up. When they strip out the oscillator modules, the special attention, and replace the forecasting head with a plain linear one, performance drops. On ETTh1 the drop is almost negligible — MSE goes from 0 point 484 to 0 point 487. On ETTh2 it's much steeper, from 0 point 403 to 0 point 478, and that's the volatile dataset where the complex components earn their keep.
Jane: And even the stripped-down version stays competitive with HAT and COTN at the short horizon. So the Quantformer backbone alone is a solid foundation, and the chaotic gating is the extra layer that pushes past the strong baselines.
Lalam: For anyone considering adopting the approach, that ablation might be the most useful part of the paper. It shows where the value concentrates, and it suggests you could build a simplified version of this model for calmer data and only pay for the full chaotic machinery when your problem is genuinely wild.
Tom: So we've got a smarter oscillator mixture, a better numerical backbone, and a gate that keeps training stable. Now let's see how the first page frames the original forecasting problem that all of this is responding to.
First Page: Tom: The first page opens with the standard pain points of forecasting — long-range dependencies, volatility clustering, structural shifts — and then walks through why existing tools fall short. The argument builds very cleanly from there.
Jane: The critique of recurrent models is that LSTM and GRU capture sequential dependence but pay for it with sequential computation, which limits parallelism and makes long-horizon optimization awkward. Transformers solved the parallelism problem with self-attention. But the paper's central observation is that most transformer forecasters kept redesigning attention while leaving the feed-forward nonlinearity untouched.
Lu: So you get an imbalance where the attention mechanism receives all the research attention, and the pointwise activation that processes each value stays frozen. Under volatile conditions, smooth static activations simply can't respond fast enough to abrupt transitions. That single observation drives the entire paper.
Meng: And page one also includes a positioning statement that deserves credit. The authors say their contribution is a practical integration and extension of existing ideas, not a new formal theory of quantum or fractal dynamics. They borrow Quantformer's linear embedding and COTN's oscillator dynamics, then combine them in a new way. No overreach.
Jane: The dataset descriptions on that page also signal the intended use case. The ETT data comes from the State Grid Corporation of China, with two hourly subsets of 8,640 records each. Then there's a high-frequency A-share stock dataset with more than 17,000 one-minute records, including open, high, low, and close prices plus trading volume. That's genuinely chaotic financial data.
Tom: And the problem formulation is refreshingly simple on paper. Given an input window with T time steps and C variables, produce a forecast with H steps and D target variables. The entire framework is a learned mapping between those two tensors, with all the complexity hidden inside that mapping.
Lalam: What strikes me is how the first page sets expectations for the whole paper. You know what's borrowed, what's new, what problem it targets, and where the evidence will come from. By the time you hit the results tables, there's no ambiguity left about what was actually tested.
Tom: So the motivation is solid, the design is coherent, and the evidence is on the table. The last question is whether this approach has staying power, which is exactly what the conclusion considers.
Conclusion: Tom: So where does this leave us? I think the paper's strongest contribution is the argument that activation design deserves far more attention in transformer forecasting. The feed-forward block has been treated as a fixed part to reuse unchanged, and this work shows it can be a genuine source of gains. That's a subtle reframing of where research effort should go.
Jane: And the evidence hangs together. A learned mixture of eight Lee Oscillator families, compressed with Max-over-Time pooling and balanced against GELU by a learnable gate, outperforms the strong baselines across both electricity and financial data. The biggest wins land on the most volatile benchmark, exactly where the theory says they should. For a paper whose central trick is swapping an activation function, that's a strong result.
Lu: What I'll remember is how each component has a clear job. The oscillator provides a short internal time evolution that can catch sharp local changes. The pooling keeps the computation tractable. The gate decides how much chaos to allow at any given moment. There's no dead weight in the design.
Meng: And every number tells the same story. The dramatic MSE improvement over HAT on ETTh2, the larger ablation drop on the volatile dataset, the smaller but consistent gains on the stock data — all of it points one direction: chaos-aware activation helps when conditions get rough and doesn't hurt when things are calm. That consistency across datasets is what makes the claim believable.
Lalam: The wider lesson is that the forecasting community has poured enormous effort into attention mechanisms, and this paper is a reminder that the feed-forward path deserves equal scrutiny. The authors also mention extending the approach to broader sequence modeling tasks in future work, which sounds genuinely promising. If the method generalizes that far, it could outgrow the time-series niche entirely.
Tom: I know I'll be looking at plain old GELU differently from now on. And I'm curious whether these oscillator activations will show up beyond forecasting — in anomaly detection or control systems, for instance.
Jane: Either way, this one goes on the watch list. That wraps up our discussion, thanks for staying with us, and we'll be back with the next paper shortly.
Junkai Lin, Siqi Hou, Raymond Lee
Faculty of Science and Technology, Beijing Normal-Hong Kong Baptist University · Guangdong Provincial Key Laboratory of Interdisciplinary Research and Application for Data Science (BNBU)
cs.AI
Submitted: 2026-08-07
Updated: 2026-08-10
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 41/100
The gist: The paper is authored by Junkai Lin, Siqi Hou, and Raymond Lee from the Faculty of Science and Technology, Beijing Normal-Hong Kong Baptist University, Zhuhai, China (arXiv:2608.07363v1 [cs.AI], 7
Key concepts
- Lee Oscillator
- A discrete-time system with excitatory and inhibitory units that can produce chaotic behavior. In QFCQT, it's used as an activation function, running each input through a short chaotic time evolution and compressing it with Max-over-Time pooling to capture sharp local changes.
- Quantformer
- A transformer variant designed for numerical data, using linear projections instead of word embeddings. QFCQT adopts this backbone for direct processing of time-series inputs, avoiding the NLP pipeline and enabling efficient handling of numerical sequences.
- Chaotic gating
- A learnable mechanism that blends the chaotic oscillator response with the smooth GELU activation. The gate decides how much chaos to allow based on the input, enabling the model to switch between calm and wild responses depending on data volatility.
- Max-over-Time pooling
- A technique that takes the maximum value over the internal steps of the oscillator trajectory, keeping the output dimension the same as the input. This prevents computational explosion while extracting the strongest response from the chaotic dynamics.
Terminology
Summary
The paper is authored by Junkai Lin, Siqi Hou, and Raymond Lee from the Faculty of Science and Technology, Beijing Normal-Hong Kong Baptist University, Zhuhai, China (arXiv:2608.07363v1 [cs.AI], 7 Aug 2026).
Abstract: The paper states that "Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility bursts, structural shifts, and nonlinear oscillatory behaviors. Although Transformer-based forecasters are effective for modeling long-term temporal dependencies, their feed-forward blocks typically rely on smooth static activations that are insufficiently sensitive to abrupt regime changes."
The authors propose QFCQT, short for Quantum-Fractal-inspired Chaotically Gated Quantformer, for robust forecasting under complex volatile dynamics.
The authors clarify an important scope condition: 'quantum-fractal-inspired' denotes a computational analogy based on soft oscillator superposition and multi-scale nonlinear responses, rather than a formal quantum-mechanical or fractal-theoretic derivation.
The abstract summarizes the three main components:
-
a Quantformer-style numerical encoder that directly processes multivariate inputs via linear embedding
-
a learnable Lee-oscillator activation module that maps scalar pre-activations to dynamic oscillatory responses and summarizes them through Max-over-Time pooling
-
a smooth-chaotic gated fusion mechanism that adaptively balances conventional smooth activations and chaos-sensitive responses
Additionally, instead of using a single fixed oscillator, QFCQT employs a soft superposition of eight parameterized Lee oscillator families to adaptively capture different nonlinear response patterns across regimes.
Key results: "Experiments on ETTh1, ETTh2, and A-share Stock Index benchmarks show that QFCQT consistently outperforms strong baselines, including Informer, LogTrans, LSTMa, HAT, and COTN. On highly volatile settings, the proposed method achieves up to 43.9% MSE improvement over HAT and 41.3% over COTN on ETTh2 with prediction horizon 24."
The paper frames its contribution as a practical integration and extension of existing ideas rather than a new formal theory of quantum or fractal dynamics.
The authors identify two relevant prior directions: Quantformer (Quantitative Transformer) [2] shows that numerical time series can be modeled more naturally by direct linear embedding instead of following the standard NLP-oriented Transformer pipeline,
and COTN [3] shows that Lee Oscillator [4] dynamics, Max-over-Time pooling, and gated fusion can improve responsiveness to chaotic and highly volatile signals.
The paper notes a critical gap in prior Transformer forecaster designs: "most existing Transformer forecasters mainly improve attention design or computational efficiency, while the nonlinear transformation in the feed-forward block remains largely unchanged. Under highly volatile conditions, such smooth static activations may be insufficient to capture abrupt local transitions."
"Given a multivariate historical sequence X ∈ R(T×C), where T is the look-back window and C is the number of input variables, the objective is to predict a future sequence Y ∈ R(H×D), where H is the forecasting horizon and D is the number of target variables. The model learns a mapping fΘ: R(T×C) → R(H×D)."
The framework follows a four-stage pipeline: numerical sequence initialization, quantum-fractal-inspired feature encoding, chaotically gated Quantformer encoding, and horizon projection.
The chaotically gated encoder combines temporal self-attention with oscillator-based gated activations to capture both long-range dependencies and local volatile transitions.
"Following Quantformer, we adopt a standard linear projection to embed the input sequence, instead of using token embedding and masking-based decoding. The embedding is defined as H(0) = XWe + be, where We ∈ R(C×d) and be ∈ R d."
Multi-head self-attention captures long-range dependencies: Qh = HWh Q, Kh = HWh K, Vh = HWh V, Attnh(H) = Softmax(QhK⊤h/√dh)Vh.
The outputs are concatenated: MHSA(H) = Concat(Attn1,..., AttnM)WO.
The oscillator evolves via a discrete-time excitatory–inhibitory process
:
-
E(t+1) = f(a1L(t) + a2E(t) − a3I(t) + a4S(t) − ξE)
-
I(t+1) = f(b1L(t) − b2E(t) − b3I(t) + b4S(t) − ξI)
-
omega(t+1) = f(S(t))
-
L(t) = [E(t) − I(t)] exp(−kS(t)2) + omega(t)
where f(·) is implemented using a bounded nonlinear response, and S(t) is the external stimulus induced by the pre-activation value.
Since a direct use of the full oscillator trajectory would dramatically increase dimensionality and computational cost,
the authors compress the internal temporal evolution: fkMoT(x) = max Lk(t; x), 1≤t≤N, where k indexes the oscillator family and N is the number of internal evolution steps.
Instead of a single fixed oscillator, the model uses a learnable soft mixture: πk = exp(ck) / Σ exp(cj), fchaos(x) = Σ πk fkMoT(x).
The eight oscillator types exhibit distinct bifurcation patterns, which enrich the response space of QFCQT for non-stationary time-series modeling.
To preserve training stability: "fsmooth(x) = GELU(x), g = σ(λ), fQFCQT(x) = g fsmooth(x) + (1 − g) fchaos(x), where λ is a learnable scalar gate. This fusion preserves the training stability of GELU while enabling the model to respond to highly volatile local perturbations."
The block uses a channel expansion/projection structure with 1 × 1 convolution
: U = Conv1D↑(H̃⊤), V = fQFCQT(U), Z = Conv1D↓(V), H(l) = LayerNorm(H̃(l) + Dropout(Z⊤)).
Mean squared error is optimized: LMSE = (1/HD) Σ Σ (Ŷt,d − Yt,d)2,
with MAE as an additional evaluation metric.
Datasets: The ETT (Electricity Transformer Temperature) dataset from the State Grid Corporation of China, using two hourly subsets ETTh1 and ETTh2, each containing 8,640 records. Both subsets exhibit strong volatility and nonlinear temporal patterns.
The financial dataset is a high-frequency A-share stock dataset from a private data vendor, containing more than 17,000 one-minute records over several months. Each record includes OHLC prices and trading volume.
The dataset captures key intraday market characteristics, such as volatility clustering, liquidity variation, and rapid short-term fluctuations.
Preprocessing: "Timestamp continuity is maintained by forward-filling short missing intervals and removing longer discontinuities. Duplicate records are discarded, and abnormal points are filtered using domain-specific rules together with Z-score thresholds. Based on the OHLCV data, we further construct standard temporal features, including log-returns, moving averages, and volatility-related indicators."
Baselines: Informer, LogTrans, LSTMa, HAT, COTN, and TimesNet. These baselines cover efficient Transformer forecasting, recurrent sequence modeling, and oscillator-enhanced Transformer modeling.
"QFCQT achieves the best or competitive performance in most settings across ETTh1, ETTh2, and A-share Stock Index datasets. The improvement is most significant on ETTh2, especially at horizons 24 and 168, indicating that the proposed chaotically gated activation is more beneficial under highly volatile and nonlinear dynamics. On the A-share dataset, the gains are more moderate but consistent, suggesting that the framework can transfer from electricity benchmarks to financial forecasting scenarios."
Key improvements in MSE over baselines (from Table IV):
-
ETTh2, horizon 24: 43.9% over HAT, 41.3% over COTN, 23.7% over TimesNet
-
ETTh2, horizon 48: 19.5% over HAT, 14.8% over COTN
-
ETTh1, horizon 336: 21.9% over HAT, 15.2% over COTN
-
A-share, horizon 24: 41.79% over HAT, 37.74% over COTN
The authors note a critical insight: Compared with COTN, QFCQT demonstrates that simply introducing a chaotic unit is not sufficient; the way it is integrated into the numerical backbone also matters.
Furthermore, QFCQT generally achieves larger gains in MSE than in MAE, indicating that it is especially effective in reducing large prediction deviations and improving robustness under difficult segments.
The ablation removes three components: "1) the quantum-gated chaotic activation module, including QuantumSuperposition LORS and VectorizedLeeOscillator; 2) the fractal-modulated attention module, i.e., FractalModulatedAttention; and 3) the original forecasting head, which is replaced with a simpler linear regression head."
Results: removing these components leads to consistent performance degradation on both ETTh1 and ETTh2, which confirms that each of them contributes positively to the forecasting capability of QFCQT.
Notably, "the performance drop is small on ETTh1 but more evident on ETTh2, suggesting that the oscillator-based gated design is particularly useful under stronger volatility. Even after ablation, the simplified variant remains competitive with HAT and COTN at horizon 24."
The authors attribute the gains to three factors:
-
The Quantformer-style numerical embedding avoids unnecessary NLP-style assumptions and is better aligned with ordered quantitative observations.
-
"The Lee Oscillator activation introduces a richer nonlinear response landscape than conventional pointwise activations. Through internal temporal evolution and Max-over-Time pooling, the model can transform each scalar pre-activation into a dynamic meta-activation that is more sensitive to sharp local variations."
-
The smooth-chaotic gated fusion stabilizes training. Instead of letting chaotic responses dominate all hidden transformations, the learnable gate adaptively balances the oscillator response and the smooth GELU component.
The authors emphasize again the non-formal scope: "QFCQT does not claim a formal quantum or fractal theory; rather, the term 'quantum-fractal-inspired' refers to a computational analogy based on soft oscillator-family superposition and multi-scale nonlinear response organization. By constraining chaotic responses through Max-over-Time pooling and smooth-chaotic gated fusion, the model improves sensitivity to abrupt local changes while maintaining stable optimization."
The conclusion states: "Unlike conventional Transformer-based forecasting models relying on fixed smooth activations, QFCQT integrates a Lee-oscillator-based dynamic activation, Max-over-Time pooling, and a smooth-chaotic gated fusion strategy into a Quantformer-style encoder, enabling effective modeling of long-range dependencies and nonlinear temporal dynamics. The experiments
showed that QFCQT consistently outperformed representative baselines... across multiple forecasting horizons, with particularly notable improvements under highly volatile settings. The paper concludes that
the feed-forward nonlinear transformation plays a crucial role in Transformer forecasting, and future work
will explore its extension to broader sequence modeling tasks."
Improvements for AI systems
Improvements to AI systems:
-
Replace static feed-forward activations (e.g., GELU) in Transformer encoders with Lee-oscillator-based dynamic activations that evolve over internal time steps and are compressed via Max-over-Time pooling — enabling sensitivity to abrupt, local volatility bursts instead of smooth pointwise responses.
-
Use a soft superposition of eight parameterized Lee-oscillator families with learnable mixture weights, allowing the model to adaptively select the most suitable nonlinear response patterns for different regimes within a single sequence.
-
Add a smooth-chaotic gated fusion (
g = σ(λ)) between the conventional GELU activation and the chaotic oscillator response, preserving training stability while still letting chaotic dynamics influence hidden states when needed. -
Adopt Quantformer-style direct linear embedding for numerical inputs instead of NLP-style tokenization, avoiding unnecessary assumptions about discrete token structure and better aligning with ordered quantitative observations.
-
Integrate a fractal-modulated attention mechanism (as validated by ablation) to capture multi-scale temporal dependencies alongside the oscillator-enhanced feed-forward block.
What the improved AI system can do:
-
Forecast highly volatile, non-stationary multivariate time series (e.g., electricity transformer temperature, high-frequency stock indices) with long-range dependencies and structural shifts.
-
Significantly reduce large prediction errors — e.g., up to 43.9% MSE improvement over HAT and 41.3% over COTN on ETTh2 with a 24-step horizon, and consistent gains on A-share financial data.
-
Maintain stable optimization despite chaotic activation components, avoiding training divergence common with purely chaotic or oscillatory units.
-
Dynamically balance smooth and chaos-sensitive responses per layer/regime, adapting to both stable segments and abrupt transitions in a single time series.
-
Generalize across different forecasting horizons (24, 48, 168, 336) and transfer from energy benchmarks to financial forecasting scenarios.
Sources
- Quantformer: from attention to profit with a quantitative transformer trading strategy
- COTN: A Chaotic Oscillatory Transformer Network for Complex Volatile Systems under Extreme Conditions
- TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis
- Transformers in Time Series: A Survey
- Deep Transformer Models for Time Series Forecasting: The Influenza Prevalence Case
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection