Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration
Sabin Roman, Ljupco Todorovski, Saso Dzeroski
Jožef Stefan Institute · University of Ljubljana
cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 15 pages, 4 figures. Accepted for oral presentation at the 29th International Conference on Discovery Science (DS 2026), Mainz, Germany, October 5-9, 2026. To appear in Springer Lecture Notes in Computer Science (LNCS)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper develops the Sparse Orthogonal Regression Technique (SORT), "a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data." SORT "estimates
Terminology
Summary
The paper develops the Sparse Orthogonal Regression Technique (SORT), a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data.
SORT estimates expansion coefficients directly from observations using L1-regularized regression, avoiding explicit quadrature or analytic inner-product evaluation.
The central application is data-driven discovery of ordinary differential equations: vector fields are represented in chosen orthogonal bases and learned as sparse coefficient expansions.
This "provides a complementary route to symbolic regression, grammar-based discovery, and SINDy-style sparse identification by first recovering a compact spectral representation, which can later guide searches for simpler analytic forms."
The authors state that "SORT combines two standard ingredients: orthonormal basis representations and sparsity-promoting regression. We therefore do not claim algorithmic novelty in sparse regression on orthogonal features. The contribution is application-level and representational: SORT treats the learned sparse orthonormal expansion as a reusable object for data-driven dynamical-system reconstruction, numerical integration by coefficient readout, and nonlinear function approximation with order-consistent model growth."
For equation discovery, "SORT provides a complementary route: each component of the vector field is represented in a fixed orthonormal basis and estimated through sparse regression. The output is not necessarily a compact symbolic expression, but a sparse spectral representation of the vector field. The experiments show that
Across the dynamical-system experiments, SORT matches or improves upon library-based sparse-regression baselines when the basis is well adapted to the problem, and shows more stable degradation under sparse sampling, noisy derivative estimates, and representation mismatch. Specifically,
as sampling becomes coarse, Lotka–Volterra and Van der Pol show abrupt SINDy rollout failure, while the orthogonal sparse model remains bounded and degrades more gradually. In the cyclic-system experiments with the Thomas attractor and a Bessel-driven variant,
SORT remains robust across training fractions and noise levels while SINDy degrades under representation mismatch for the Bessel system, because
the true nonlinearity is not sparse in the selected library, making coefficient selection brittle under noise."
For numerical integration, "In an orthonormal expansion, coefficients are inner products between the function and basis elements. If a basis direction is chosen to represent an integration functional, then the corresponding coefficient gives the integral, up to a known normalization factor. SORT
estimates them from scattered or noisy point samples by sparse regression and then evaluates the desired integral by coefficient readout. The experiments show that
SORT accurately tracks the closed-form values across a wide range of α, including strongly oscillatory Fresnel cases where accurate quadrature would require resolving rapid oscillations, and
the estimates remain accurate in moderate dimensions, but the error increases with dimension."
For approximation, increasing the model order refines the representation without changing the meaning of previously learned low-order coefficients.
The experiments show that SORT gives the lowest test error and, more importantly, preserves a stable low-order coefficient hierarchy as model size increases,
while "OLS in a monomial basis changes substantially with order; OLS in a Legendre basis benefits from orthogonality but remains noise-sensitive; and RBF/RFF ridge models produce dense weights without an interpretable hierarchy."
The authors conclude that "SORT provides a compact intermediate representation between raw samples and later symbolic or analytic description. It can first recover a stable expansion in a suitable orthogonal basis; a subsequent symbolic stage may then search for simpler closed-form expressions consistent with that expansion. They also note that
basis design is part of the scientific modeling problem, not an implementation detail, and that
SORT is not immune to mismatch, but it shifts the problem away from brittle selection among generic terms to basis design adapted to the problem domain. Future work
should develop adaptive and anisotropic basis selection, empirical orthogonalization for irregular samples, and larger-scale applications."
Improvements for AI systems
Improvements to AI systems:
-
Robust dynamical-system discovery from sparse/noisy data – Implement SORT as a spectral surrogate module in neural ODE or equation-discovery pipelines. The AI system can learn vector fields as sparse orthonormal expansions directly from irregularly sampled, noisy observations, avoiding brittle library selection. It maintains bounded, gradual degradation under coarse sampling or derivative noise, unlike SINDy-style methods that fail abruptly.
-
Stable, order-consistent function approximation – Use SORT’s coefficient hierarchy to build AI models that refine predictions by increasing basis order without altering previously learned low-order coefficients. This enables incremental model growth, interpretable feature importance, and noise-robust fitting in high-dimensional regression tasks, outperforming OLS, RBF, and random-feature ridge models.
-
Quadrature-free numerical integration from scattered samples – Integrate SORT into AI systems that must compute integrals (e.g., expected values, physical quantities) from point clouds. The system estimates expansion coefficients via L1-regularized regression and reads off the integral from a designated basis coefficient, handling oscillatory or high-frequency integrands without explicit quadrature or dense sampling.
-
Basis-adaptive symbolic regression preprocessor – Enhance symbolic regression or grammar-based discovery systems by first using SORT to recover a compact spectral representation. The AI can then search for analytic forms consistent with that expansion, reducing search space and improving interpretability, especially when true nonlinearities are not sparse in generic libraries.
-
Anisotropic and adaptive basis selection – Build an AI system that automatically learns problem-specific orthogonal bases (e.g., via empirical orthogonalization or dictionary learning) and applies SORT for sparse coefficient estimation. This shifts the burden from brittle term selection to adaptive basis design, improving generalization under representation mismatch.
What the improved AI system can do:
-
Discover governing equations from sparse, noisy, irregularly sampled time-series data with graceful performance decay, even when the underlying dynamics are not sparse in standard polynomial or trigonometric libraries.
-
Approximate complex functions with a stable, interpretable hierarchy of coefficients, allowing users to increase model capacity without retraining from scratch or losing prior knowledge.
-
Compute integrals of functions from scattered, noisy samples without requiring dense grids or analytic inner-product evaluation, including strongly oscillatory cases.
-
Provide a reusable spectral representation that can be fed into symbolic regression tools, enabling faster and more reliable discovery of closed-form laws.
-
Automatically adapt its basis to the data’s structure (e.g., anisotropic features, irregular domains) and maintain robustness when the assumed basis is imperfect.
Abstract
We develop the Sparse Orthogonal Regression Technique (SORT), a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data. SORT estimates expansion coefficients directly from observations using L1-regularized regression, avoiding explicit quadrature or analytic inner-product evaluation. The central application is data-driven discovery of ordinary differential equations: vector fields are represented in chosen orthogonal bases and learned as sparse coefficient expansions. This provides a complementary route to symbolic regression, grammar-based discovery, and SINDy-style sparse identification by first recovering a compact spectral representation, which can later guide searches for simpler analytic forms. Across the dynamical-system experiments, SORT matches or improves upon library-based sparse-regression baselines when the basis is well adapted to the problem, and shows more stable degradation under sparse sampling, noisy derivative estimates, and representation mismatch. Specific examples illustrate why this representation is useful: if a finite library misses the problem-specific nonlinearity, the resulting model can fail. SORT is not immune to mismatch, but it shifts the problem away from brittle selection among generic terms to basis design adapted to the problem domain. The experiments also show that dominant low-order coefficients persist as model order increases, supporting order-consistent model growth. Beyond equation discovery, the same learned expansion supports nonlinear approximation and estimation of complex, high-dimensional integrals by coefficient readout. Overall, SORT provides a reusable intermediate representation for system identification, approximation, and integration, while making basis design an explicit part of the scientific modeling problem.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks