When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision
The University of Hong Kong
stat.ME, cs.CL
Submitted: 2026-08-11
Updated: 2026-09-12
Code: https://github.com/Jinsong-Chen/vbpm
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 95/100
The gist: Author: Jinsong Chen, Faculty of Education, The University of Hong Kong arXiv:2608.10731v1 [stat.ME] 11 Aug 2026 --- The paper addresses whether an additional general dimension is necessary beyond
Terminology
Summary
Author: Jinsong Chen, Faculty of Education, The University of Hong Kong
arXiv:2608.10731v1 [stat.ME] 11 Aug 2026
The paper addresses whether an additional general dimension is necessary beyond correlated first-order factors in factor analysis models. The central observation is that "whether a general factor is empirically distinguishable from correlated factors is a property of the population loading pattern. Specifically, it depends on the proportionality between general and group loadings within clusters. This distinguishability is
a property of the population covariance matrix, not of any estimator or design."
The paper establishes three key results: (1) when general and group loadings are proportional within every cluster, the bifactor structure is covariance-equivalent to correlated factors (Proposition 1); (2) when proportionality fails in every cluster, three items per cluster with mild regularities leave no K-factor model with diagonal uniquenesses able to reproduce the covariance matrix (Theorem 1); (3) between these lies a mixed boundary, located numerically, that turns on cluster resistance.
"If wk = 0 in every cluster, so that b gk = ck bkˢ with ck = αk/‖bkˢ‖, then C = AΦA′ exactly, with second-order loadings γk = ck/√(1 + ck2). Equivalently, the ratio defined above is b s,j/b g,j = 1/ck within cluster k. The same uniquenesses serve both representations, so Σ admits an exact K-factor representation. A reducible bifactor is indistinguishable from correlated factors at any sample size."
"Let Σ = C + Ψ arise from an orthogonal bifactor structure with K ≥ 2 clusters, a nonzero general loading on every item, at least three items per cluster, and at least two nonzero group loadings per cluster. If wk ≠ 0 in every cluster, then no K-factor model reproduces Σ under any diagonal uniquenesses: rank(C + Δ) ≥ K + 1 for every diagonal Δ."
The proof uses a decomposition within each cluster: b gk = αkuk + wk, where uk = bkˢ/‖bkˢ‖ and wk ⟂ bkˢ. A cluster with wk = 0 has exactly proportional loadings and is called silent
—it contributes nothing to the general-factor question at any sample size.
When some clusters are proportional, the theorem's hypothesis fails. A cluster is called resistant
if it has at least four items and its within-cluster ratio vector ρi = b s,i/b g,i contains two disjoint unequal pairs. Numerical analysis shows: "a lone non-proportional cluster with only three items, or with a constant-except-one ratio pattern, is traded away into the uniquenesses exactly when the remaining clusters are proportional, while a lone resistant cluster is not."
Distinguishability is graded at the Σ level, measured by the population distance to the K-factor class:
D K(Σ) = min over Λ ∈ R J×K, Ψ ≥ 0 diagonal, ΛΛ′ + Ψ ≻ 0 of F ML(Σ, ΛΛ′ + Ψ)
where F ML(Σ, Ω) = tr(Ω−1Σ) − logΩ−1Σ − J.
The paper defines three tiers: Tier I (theorem-covered distinguishability: wk ≠ 0 in every cluster), Tier II (mixed boundary: at least one silent and at least one non-proportional cluster), and Tier III (reducible: wk = 0 everywhere; exactly higher-order by Proposition 1).
Two geometric facts drive intuition: "Cross-cluster covariance blocks are rank-one under both models, so between-cluster information can never distinguish a general factor from factor correlations. And an additional general dimension is distinguishable from correlated first-order factors only through within-cluster structure."
The paper separates four questions: identifiability (can a specified model's parameters be recovered), count selection (which candidate counts deserve examination), structural stability (whether approximately the same configuration persists across adjacent counts), and distinguishability (whether D K(Σ) > 0). The middle two are logically independent, since a unanimously selected count can carry a structure that fails to persist and adjacent counts can share a persisting core.
The paper develops a two-step procedure within partially exploratory factor analysis (PEFA):
Step 1 runs a PEFA sweep in oblique mode only, launching from a backbone design with two anchor items per suspected cluster, over a window of candidate counts. The target is the stable group-structure core, the configuration of first-order clusters that persists across adjacent counts.
The stability analysis compares anchored solutions at adjacent counts using minimum Tucker congruence over matched columns.
Step 2 poses the operational model comparison conditionally on the delivered structure: The delivered oblique K-factor model is compared with the anchored bifactor (K* factors) built from the same design Q, with two markers per cluster taken from the step-1 sweep.
The delivery decision uses two axes—count evidence and structural evidence—defining three layers: L1 (resolved count, stable structure: deliver the structure at the selected count), L2 (uncertain count, stable common core: deliver the core and carry count uncertainty into step 2), and L3 (unstable: report that no stable structure was delivered).
Six fixed populations with K = 5 clusters and J = 20 items (four per cluster) were used, crossing the tier taxonomy with group-loading magnitude. Key findings:
-
Finding 1: The anchoring design matters only where estimation is already hard. "AO designs trail AZ designs slightly and both trail full-primary specification, but the gaps appear only under weak loadings at small N and on the reducible population, and the errors are omissions rather than false discoveries."
-
Finding 2: At the generating count, the step-2 comparison has favorable operating characteristics in both directions.
-
Finding 3:
The oblique sweep counts well and the bifactor sweep does not, under either specification, because surviving columns' free entries can absorb an omitted group factor's within-cluster block.
-
Finding 4: The two specifications (AO and AZ) count almost identically.
-
Finding 5:
On the moderate and reducible populations, when the oblique sweep errs it errs upward, and a K+1 solution still contains the K structure beside a thin surplus column.
-
Finding 6:
A bifactor-mode under-count takes the structure down with it.
Seven populations crossed three structural bases (bifactor, higher-order, oblique) with interference in two forms (minor factor, doublet residual covariance). Key findings:
Finding 7: The specification the structural check runs on decides whether a correct structure is delivered, and the anchor-zero design rejects correct structures on bifactor populations.
Under anchor-zero, the structural check falsely rejects 42.5% of correctly estimated structures on bifactor populations versus 2.4% under anchor-only, rising to 68.4% and 70.2% on doublet populations at N = 2000. Anchored zeros close the outlet through which an irreducible general dimension would otherwise be absorbed at K̂+1, so the larger solution reorganizes the retained columns instead.
Finding 8: A borderline doublet converts added information into non-delivery, and a delivery rate that falls with N is the signature of no stable structure.
Finding 9: Absorbed local dependence imitates a general factor; the error grows with sample size, and no step-1 indicator detects it.
In the adversarial population (higher-order base with absorbed doublet), the bifactor, wrong on this base, is preferred in a rising share of replications as N grows
—from 2.0% at N = 500 to 85.0% at N = 2000 under BIC. Theorem 1 presumes diagonal uniquenesses, the residual doublet violates exactly that, and the absorbed within-cluster covariance is signal only the general column can carry.
Four datasets were analyzed: ICAR ability (16 × 4), Holzinger–Swineford 24 tests (24 × 5), PID-5 domain-primary facets (15 × 5), and bfi (25 × 5).
-
Ability delivers at three factors (L1): the 3–4 step is the first whose ELBO gain falls below the 20% cut (−2.3), with minimum matched congruence.880 clearing the.85 threshold.
-
Holzinger 24 delivers at four factors (L1): the 4–5 step is first below the cut (−68.2), with congruence.964.
-
PID-5 delivers at five factors (L2): gains stay above the cut through 4–5 (47.8), and the 5–6 step is first below it (2.5), with congruence.896. However,
it is the four-factor solution that persists against every larger candidate in the window,
so both counts are carried into step 2. -
bfi delivers nothing (L3):
the only step that clears the congruence threshold, 4–5 at.949, carries the largest gain in the window, while every step whose gain has fallen below the cut fails the threshold.
In step 2, Every delivered structure prefers the added dimension, and it does so on both criteria and under both specifications.
The PID-5 gives the clearest result at both counts (BIC margins of −77.0 at five and −326.9 at four under anchor-only).
The paper states six practical choices:
-
Sweep in oblique mode rather than bifactor mode.
-
Use the anchor-only (AO) specification at both steps.
-
Deliver on the joint criterion: the first adjacent step whose operative gain falls below the 20% cut and whose stability index clears its threshold (φ ≥.85, aggregate RMSD ≤.10, or worst-column RMSD ≤.20).
-
When delivered count and persisting count differ, report the pair rather than choosing between them (L2 core).
-
The step-2 anchored comparison is the operative way to ask whether a general dimension is necessary; a reversal across specifications or counts is reason to withhold.
-
When nothing is delivered (L3), report non-delivery and treat local dependence as the first thing to examine.
The paper acknowledges: "The simulations judge recovery against generating truth, and their populations use fixed loading patterns with four-item clusters and correct anchors, chosen to isolate the reducibility tiers at the cost of the heterogeneity real batteries show. The persistence threshold is
calibrated within the same scenario families and checked on an untouched half, not externally validated. Step 2 is
a conditional, anchored, variational-BIC comparison at a selected structure rather than an estimator or formal test of the unrestricted distance D K."
Selection is not delivery: the gain rule identifies a neighborhood worth examining, not by itself a structure worth interpreting.
The paper's through-line is that the bifactor's practical troubles—fit advantages that do not replicate, anomalous solutions, and spurious general factors
—are "not estimator pathologies to be tuned away. They are surface expressions of a population-level property with a sharp sufficient condition and a partly characterized boundary. Distinguishability lives in the within-cluster loading-ratio patterns. It dies at the reducible boundary that the higher-order model occupies (Proposition 1), it survives whenever every cluster departs from proportionality (Theorem 1), and mixed configurations require the finer pattern analysis developed in the Appendix."
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems and their resulting capabilities:
Improvement: Implement a two-step selection procedure (oblique sweep first, then anchored bifactor comparison) instead of fitting a single pre-specified model class. Use the 20% gain cut
and minimum Tucker congruence ≥.85
as joint delivery criteria, rather than relying on a single information criterion.
Capability: The AI can now distinguish between count selection
and structural delivery
—it will report when a model count is uncertain (L2) or when no stable structure exists (L3), rather than forcing a choice. It will avoid false discoveries of general factors in reducible populations and will detect when local dependence (e.g., doublet residuals) is masquerading as a general factor.
Improvement: Replace anchor-zero (AZ) specifications with anchor-only (AO) specifications in all exploratory factor analysis routines. The paper shows AZ falsely rejects correct structures at rates up to 70.2% on bifactor populations with doublets, while AO performs near-nominal.
Improvement: Add a pre-processing check for within-cluster loading-ratio proportionality (wk ≠ 0) before deciding between bifactor and higher-order representations. Use the paper's Theorem 1 condition: if every cluster has at least three items, two nonzero group loadings, and non-proportional general-to-group ratios, then no K-factor model can reproduce the covariance—so a general factor is necessary.
Improvement: When a model fails to deliver (L3) or when fit improves with sample size but structure does not stabilize, automatically flag residual doublets or absorbed local dependence as the primary suspect, per Finding 9.
Improvement: Replace single-count selection with a persistence analysis across adjacent counts, using the paper's stability indices (minimum matched congruence, aggregate RMSD ≤.10, worst-column RMSD ≤.20). Report the persisting core
when counts differ from the selected count.
Improvement: Implement the population distance D K(Σ) as a diagnostic tool, computed via the minimum of F ML over the K-factor class, to measure graded distinguishability rather than binary yes/no decisions.
Improvement: Use the paper's finding that anchoring design matters only under weak loadings and small N—and that errors are omissions, not false discoveries—to set default anchor strategies (two markers per suspected cluster) that minimize computational cost without sacrificing accuracy in well-powered settings.
Improvement: Add a reversal check
as a mandatory step: if the step-2 comparison reverses across specifications (AO vs. AZ) or across adjacent counts, withhold delivery and report ambiguity, per the paper's recommendation.
Abstract
Whether an additional general dimension is necessary beyond correlated first-order factors is a property of the population covariance matrix, not of any estimator or design. This research establishes when that property can be decided. Where the general and group loadings are proportional within every cluster the bifactor structure is covariance-equivalent to correlated factors, so no sample size separates them (Proposition 1); where that proportionality fails in every cluster, three items per cluster and some mild regularities leave no K-factor model with diagonal uniquenesses able to reproduce the covariance matrix (Theorem 1); and between them lies a mixed boundary, located numerically here and turning on cluster resistance. Distinguishability is therefore graded, measured by the population distance to the K-factor class. Because that question is conditional on a first-order structure which is itself uncertain, a two-step procedure is developed within partially exploratory factor analysis, delivering a structure only when it reproduces across adjacent counts and treating non-delivery as legitimate. Simulation shows that a unanimous count can accompany a structure that fails to reproduce, and that absorbed local dependence can imitate a general factor, the error growing with sample size while stability indicators stay clean. Four empirical datasets illustrate the possible outcomes.
Sources
- Recovering Latent Structures after Variational Bayesian Variable Selection: Fit Assessment and Factor-Number Selection in Partially Exploratory Factor Analysis
- Exploratory Hierarchical Factor Analysis with an Application to Psychological Measurement
Related papers
- Doubly robust inference via calibration
- Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries
- Flexible Nonparametric Inference for Causal Effects under the Front-Door Model
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- A Survey on Archetypal Analysis
- Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States