Bio papers — 2026-09-17

Today's focus is understanding how structure in foundation models dictates function across different domains because figuring out what makes a parameter important is key to building better systems everywhere. We looked closely at high-gain rows within gated feed-forward networks in both text and genomic foundation models to see if structural prominence translates into causal importance.

We found that candidates derived from activation were enriched compared to random or top-norm controls. However, neither spectral concentration nor operator magnitude predicted the actual causal effect size, and these associations disappeared when we looked only at the text decoder subset.

A within-layer sweep across thirty-six rows in one genomic decoder and one text decoder revealed two distinct regimes. Below a certain threshold, a row offered no positive information about causal damage, while above it ordered rows strongly but did not show how severe the damage was.

This sweep also uncovered a second row that was individually catastrophic and invisible to our initial census method. Furthermore, it showed instances where critical rows displayed non-additive damage when they were located near each other. Case studies demonstrated different ways these critical structures operate, such as a robust super-additive pair interaction in DNABERT-2 versus a sharply localized dependence in GENERator where keeping the beginning of the sequence contribution rescued most of the native loss.

This suggests that high-gain gated feed-forward network rows are an architectural feature that acts as an enrichment signal rather than a direct measure of functional criticality or causal organization. This mechanism varies depending on the specific model architecture, meaning this structural prominence is general, but how it manifests depends entirely on the specific model architecture.

Today's papers

The papers

Important terms

High-gain rows
These are specific rows within gated feed-forward networks that show strong activation patterns, suggesting they might be important structural features in foundation models.
Causal importance
This refers to determining which model components actually cause a specific outcome or effect in the system, which is the main goal of this research.
Activation-derived candidates
The researchers found that rows selected based on their activation patterns were more likely to be important than random or top-ranked control rows.
Architectural feature vs. functional criticality
The study suggests that high-gain rows are a structural property of the model, not necessarily a direct measure of how critical or functionally important they are for the task.