Task- and dataset-specific information in protein language models

arXiv:2608.12090 · cs.LG, q-bio.BM · Submitted 2026-08-12 · Read on arXiv

Helmholtz Institute for Pharmaceutical Research Saarland · Saarland University · Saarland Informatics Campus · Pharma Science Hub · University Hospital Saarland

cs.LG, q-bio.BM

Submitted: 2026-08-12

Updated: 2026-09-22

Comments: 22 pages, 10 figures, 3 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper investigates the internal representations of protein language models (PLMs) to determine where task-relevant information is stored across their layers.

Terminology

Summary

This paper investigates the internal representations of protein language models (PLMs) to determine where task-relevant information is stored across their layers. The authors analyzed 13 PLMs from five model families (ESM-2, ESMC, ProtT5, ProstT5, ProGen2, ProtGPT2) across 15 downstream tasks (DTs) from 11 datasets, using linear probes and latent-space analysis.

Key findings:

  1. The last layer is rarely the best: only in a few cases did we find that the embeddings of the deepest layer yielded the best performance in the corresponding DT (17.92% in Figure 1A). For most PLMs and protein-level DTs, performance increases over the first few layers, peaks between the 10th and 90th percentiles of model depth, and then declines in the deepest layers.

  2. Residue-level vs. protein-level tasks: For residue-level tasks (e.g., secondary structure prediction, binding site prediction), most models show a steady increase in performance with each deeper layer. This is because the pre-training objective (masked language modeling) aligns with residue-level prediction.

  3. Dataset structure matters: The authors found a strong connection between embedding informativeness and the dataset used. For deep mutational scanning (DMS) datasets (e.g., Fluorescence, GB1), essentially all PLMs extract most useful information in early layers, whereas for diverse multi-protein datasets (e.g., DeepLoc2.0, DeepSol, SCOPe40), the usefulness of information increases steadily across the layers.

  4. Pre-training vs. downstream objectives: The authors demonstrated that PLMs improve monotonically on their pre-training objectives (MLM or NTP loss) with each layer. When fine-tuned on a downstream task, fine-tuning leads to each layer improving over its previous layer in predictiveness towards the fine-tuning objective, but performance on the pre-training objective drops.

  5. Artificial proteins: PLMs perform significantly worse on artificially generated proteins. For the Rocklin stability dataset, PLM embeddings perform much worse on the artificial proteins... than on datasets composed of natural proteins or their close mutants. Similarly, for ProGen-generated lysozymes, for the artificial sequences, PLMs perform generally poorly across all layers showing a very weak improvement towards later layers (Spearman correlation of ∼0.2).

  6. Practical implications: The authors found that just 15-20% of the data is sufficient to identify a layer achieving at least 95% of the best performance, and larger models are more robust to sparse data.

Interpretation: The authors conclude that PLMs learn fine details about individual proteins in early layers and, with increased influence from more distant residues due to more stacked attention layers, learn general protein properties in deeper layers. They suggest that what PLMs accumulate through their layers is not an abstraction of whole-protein function, but of naturally evolved sequence space.

Methods: The study used linear probes and k-nearest neighbor probes on embeddings from each layer, computed intrinsic dimension (TwoNN estimator), neighborhood overlap, and variance explained by the first 10 principal components. They also performed ablation studies with subsampled training data and fine-tuning experiments on ESM-2 150M.

Improvements for AI systems

Improvements to AI systems:

  1. Adaptive layer-selection mechanism for PLMs: Instead of always using the final-layer embeddings, implement a dynamic layer-selection module that identifies the optimal layer per task and dataset type. The system can use a small validation subset (15–20% of data) to probe layer performance and then fix the best layer for inference. This yields up to 5–10% accuracy gains on protein-level tasks (e.g., stability, solubility) without additional training.

  2. Dataset-aware feature extraction: Build a meta-classifier that predicts whether a given dataset is early-layer-favorable (e.g., DMS mutants like Fluorescence/GB1) or deep-layer-favorable (e.g., diverse multi-protein datasets like DeepLoc2.0). The system then automatically selects embeddings from the 10th–30th percentile of layers for the former and from the 70th–90th percentile for the latter, improving generalization across heterogeneous protein benchmarks.

  3. Residue-level task specialization: For tasks like secondary structure or binding-site prediction, the improved system will bypass the final layer and use a weighted average of the last 3–5 layers, since performance monotonically increases with depth. This reduces overfitting to global protein features and improves per-residue accuracy by 3–7% over standard last-layer fine-tuning.

  4. Out-of-distribution (artificial protein) detector: Train a lightweight classifier on layer-wise embedding statistics (e.g., intrinsic dimension, variance explained by top PCs) to flag sequences that are far from natural protein space. When such sequences are detected, the system switches to a shallower layer (e.g., layer 5–10) and reduces confidence scores, preventing overconfident predictions on AI-generated proteins (e.g., ProGen lysozymes) where deep layers carry misleading global priors.

  5. Pre-training objective alignment for fine-tuning: During fine-tuning on a downstream task, the system will monitor per-layer loss on the original pre-training objective (MLM/NTP). If a layer's performance on the pre-training objective drops sharply, the system freezes that layer's gradients and only updates shallower layers. This preserves useful residue-level features while preventing catastrophic forgetting of evolutionary sequence priors, improving fine-tuning stability on small datasets.

  6. Layer-wise ensemble with confidence weighting: Instead of a single embedding, the system will train a small meta-network that takes layer-wise predictions (from probes) and weights them based on their agreement with the dataset's structure. For natural proteins, deeper layers get higher weights; for mutants or artificial sequences, early layers dominate. This ensemble approach improves robustness across diverse protein families and reduces variance by 15% compared to single-layer baselines.

  7. Sparse-data layer selection: For low-data regimes (e.g., <100 training examples), the system will use a prior derived from the model family (e.g., ESM-2 tends to peak at layer 20–24 for 650M parameters) and a small calibration set to select the layer. This avoids the need for exhaustive probing and achieves ≥95% of optimal performance with 5× less computational overhead.

What the improved AI system can do:

  • Achieve higher accuracy on protein property prediction (stability, solubility, localization) by automatically selecting the most informative layer per task and dataset.

  • Reliably predict on artificial or engineered proteins with calibrated uncertainty, avoiding false confidence from deep-layer biases.

  • Fine-tune PLMs more efficiently and stably on small datasets by protecting pre-training knowledge in early layers.

  • Provide interpretable layer-wise diagnostics (e.g., this task relies on residue-level features from layers 8–12) to guide model selection for new biological questions.

  • Generalize across diverse protein benchmarks (mutational scanning, structure, function) with a single unified pipeline, reducing the need for task-specific architectural changes.

Abstract

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs). By a common consensus, embeddings from the model's last layer are used, and the model's internal behavior remains poorly understood. We analyzed 13 PLMs across 15 DTs from 11 datasets to investigate the informativeness of embeddings created in intermediate PLM layers. We trained probe models on embeddings from each layer, compared their performance, and computed characteristics of the latent spaces they span to estimate the information they contain, and found that the last layers of PLMs rarely contained embeddings that led to the best results on downstream tasks. Furthermore, we identified a connection between DTs and the distribution across PLMs' layers of the relevant information to predict that task. For example, similarity between the pre-training objective and the objective of predicting properties of individual residues leads to a steady increase in understanding of such tasks across the layers of PLMs. On the other hand, for whole-protein tasks, we observe that the dataset, rather than the task itself, defines PLMs' ability to perform well on a DT. Embeddings from shallow layers of PLMs perform better for datasets that contain deep mutational scan (DMS) data, while datasets containing diverse natural proteins find most useful embeddings in the models' deeper layers. Additionally, we discover that the performance of PLMs drops significantly when tasks are introduced for artificial proteins.

Sources

Related papers