Probing and steering biology across Boltz-1s trunk-diffusion boundary

arXiv:2608.11475 · q-bio.QM, cs.LG · Submitted 2026-08-11 · Read on arXiv

Piotr Jedryszek, Tongmeng Xie, Adam Winnifrith, Alexander Hasson, Weronika Ślesak, George Wicks, Toby Winnifrith, Oliver M. Crook

University of Oxford · Evolvere Biosciences · Independent Researcher

q-bio.QM, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: This paper investigates how biological information is represented and transformed across the architectural boundary between the Pairformer trunk and the diffusion module in Boltz-1, an

Terminology

Summary

This paper investigates how biological information is represented and transformed across the architectural boundary between the Pairformer trunk and the diffusion module in Boltz-1, an AlphaFold3-class protein structure predictor. The authors analyze per-residue activations from both modules using linear probes, sparse autoencoders (SAEs), and causal interventions.

Key findings:

  1. Representational split between trunk and diffusion module: The Pairformer trunk represents both geometry and sequence chemistry. Probe-raw (logistic probe on raw activations) decodes geometry strongly at every trunk depth: helix, strand, and coil reach F1 scores of 0.79–0.90, and disorder 0.86. Sequence chemistry is also accessible, with signal peptides at 0.76 and disulfide-bond residues at 0.43. Amino-acid identity reaches 0.996 at the trunk output. In contrast, the diffusion module retains geometry but not sequence chemistry: comparing final layers (trunk L47 vs. diffusion L22), "secondary structure is essentially preserved, with helix, strand, and coil shifting by at most 0.02, whereas signal peptide falls from 0.76 to 0.35, disulfide bond from 0.43 to 0.10, and amino-acid identity from a perfect 1.00 to 0.66."

  2. Causal steerability is distinct from decodability: Steering experiments add a unit direction (supervised probe or SAE decoder direction) to the final trunk single representation that conditions the diffusion module. Both the helix and the coil directions raise their own state, dose-dependently and against a null of 19 matched-norm random directions, while mean pLDDT stays stable at ∼79. Helix and coil lie on a single antiparallel axis (cos = −0.69) and trade off against each other. However, adding a highly predictive strand probe direction (F1 0.82, precision ≥ 0.93) to the single representation produces no measurable increase in strand content; instead, it converts helix into coil. The authors hypothesize that β-strand pairing is a property of residue pairs, which are carried in z, but in Boltz-1 information flows z → s but not the reverse, so steering s cannot reach it. Ablation results confirm this: removing the direction from the trunk, from the diffusion module, or from both fails to reduce the concept's content.

  3. Sparse curated annotations underestimate model representations: Across all readouts, recall consistently exceeds precision. Comparing identical trunk activations scored against dense DSSP vs. sparse SwissProt annotations: helix at 0.85 versus 0.42 and strand at 0.83 versus 0.33. A self-consistency check computing DSSP on Boltz-1's own predicted structures changes helix probe-raw by only ∼0.06–0.08, confirming the large SwissProt gap reflects under-annotation rather than structural disagreement.

  4. SAEs provide no direct advantage over supervised probes for known concepts: "On the trunk's secondary-structure concepts, probe-raw beats SAE-1feat, the best single SAE latent, by 0.13–0.35 F1 (helix 0.85 vs. 0.72; strand 0.82 vs. 0.47), and probes steer more reliably: the coil probe moves coil content while the best coil latent does not. However, SAE latents may hold value for hypothesis generation, as an unsupervised latent can surface candidate features that existing annotation vocabularies lack, including a putative zinc-coordination cluster and a kinase catalytic-histidine motif with no matching SwissProt label."

Methodological recommendations: The authors suggest three recommendations: (1) claims about represented information should specify the readout class and avoid treating linear decodability as direct evidence of causal use; (2) causal conclusions should be restricted to the intervention site and intervention family that were tested; (3) evaluation against sparse biological annotations should include controls for missing labels whenever possible.

Limitations: The study analyzes one AF3-class model, one main SAE recipe, and a limited set of residue-level labels. Steering is applied only to secondary-structure axes. Probe folds are grouped by protein but not clustered by sequence identity, though train–test sequence-identity leakage is negligible (only 3.3% of held-out proteins had any training-fold neighbor above 30% sequence identity).

Improvements for AI systems

Improvements to AI systems:

  1. Architecture-aware steering interfaces: Build protein structure prediction systems with separate, explicit steering channels for residue-level (single representation) and pair-level (pair representation) features. This prevents failed interventions like the strand-steering case, where information encoded in pair representations is inaccessible via single-representation steering. The improved system can steer β-strand content by directly modifying pair-representation activations, enabling controlled generation of sheet-rich or sheet-poor structures.

  2. Readout-class-conditional interpretability modules: Integrate a dual-readout diagnostic layer into protein models that distinguishes between linearly decodable features and causally actionable features. The system can automatically flag when a feature is decodable but not steerable (e.g., strand in the trunk), preventing users from relying on misleading interpretability signals. This enables more reliable model editing and debugging.

  3. Dense-annotation-aware training and evaluation: Replace or supplement sparse biological annotations (e.g., SwissProt) with dense structural annotations (e.g., DSSP) during fine-tuning or evaluation. The improved system can detect under-annotation artifacts, automatically correcting for missing labels when computing concept-level metrics. This yields more accurate assessments of what the model truly represents, avoiding false negatives in feature discovery.

  4. Hybrid probe-SAE feature discovery pipeline: Combine supervised probes (for known concepts with high precision) and sparse autoencoders (for novel hypothesis generation) into a single interpretability framework. The system can automatically propose candidate features from SAE latents that lack existing annotations, then validate them via causal steering and dense structural checks. This enables discovery of new biological motifs (e.g., zinc-coordination clusters, kinase catalytic histidines) without prior label knowledge.

  5. Intervention-site-restricted causal inference: Implement a causal analysis system that explicitly tracks and reports the intervention site (trunk vs. diffusion) and intervention family (probe direction vs. SAE latent) for every steering experiment. The improved system can automatically constrain conclusions to the tested intervention site, preventing overgeneralization of causal claims across architectural boundaries. This yields more trustworthy model-editing protocols.

  6. Sequence-identity-aware evaluation controls: Add automatic clustering of training and test proteins by sequence identity (e.g., 30% threshold) before computing interpretability metrics. The improved system can flag potential leakage and adjust confidence intervals accordingly, ensuring that reported feature decodability is not inflated by memorization. This produces more robust benchmarks for protein representation analysis.

  7. Geometry-chemistry disentangled representation learning: Modify the trunk architecture to explicitly separate geometric features (secondary structure, disorder) from chemical features (amino-acid identity, disulfide bonds) into distinct latent subspaces. The improved system can then selectively transfer geometric knowledge to the diffusion module while retaining chemical information in the trunk, enabling more flexible downstream tasks like sequence-conditioned structure generation with controlled chemical property preservation.

Abstract

AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates. How biological information changes as it crosses this architectural boundary remains poorly understood. We analyze per-residue activations from the Pairformer trunk and diffusion module of Boltz-1 using linear probes, sparse autoencoders (SAEs), and causal interventions. From the trunk, both geometry (secondary structure, disorder) and sequence chemistry (amino-acid identity, signal peptides, disulfide-bond annotations) are linearly decodable. In the diffusion module, the two diverge. Secondary structure transfers essentially unchanged, whereas sequence chemistry is strongly attenuated. We then test whether decodable directions can steer the model, intervening on the final trunk single representation that conditions the diffusion module. Helix and coil directions change predicted structure dose-dependently against matched-norm random controls, but a beta-strand direction that is highly predictive (F1 =0.82) produces no measurable increase in strand content: linear decodability does not imply causal influence at the site we tested. The same probes also score markedly lower against sparse SwissProt annotations than against dense DSSP labels, because unannotated residues that the model gets right are charged as false positives; such scores are therefore lower bounds. Finally, supervised probes outscore single SAE features wherever a label already exists. We release the trained trunk and diffusion SAEs, Boltz-1 per-residue activations, and the analysis code.

Sources

Related papers