Bio papers — 2026-09-16

Today's focus centers on how we are evaluating different methods for attributing genomic sequence models because the way we currently measure success isn't truly valid across various transcription factors. We looked at how much of a known motif these methods recover or how their predictions degrade when evidence is removed.

Neither metric makes sense without knowing what would happen purely by random chance. The core issue is that the uniform chance level for contiguous motif overlap varies significantly across the twenty-six eight transcription factors in UniBind, spanning from 0.0118 to 0.0427. This spread is a three point six fold difference determined only by motif length and window size.

This spread means that raw recovery rates are not comparable quantities for some factors because their bootstrap intervals around the chance levels do not overlap at all. Correcting for this chance level changes how we classify the factors. For instance, one factor previously reported as a resolution failure moves up to be second highest, while two others reported as complete failures fall below the chance level.

Furthermore, we found that perturbation-based evaluation can fail its own test of faithfulness. This is because for one specific factor, even when all input is masked, the score remains above the decision boundary and the curve isn't consistently increasing with more masked positions. We are providing closed form chance levels and a corrected score to address these comparability issues.

Today's papers

The papers

Important terms

Genomic Sequence Models
These are computational models used to predict how transcription factors bind to DNA sequences. The research focuses on fairly evaluating these models across different transcription factors.
Contiguous Motif Overlap
This metric measures how much of a known DNA motif is successfully recovered by the prediction methods. Its performance varies significantly depending on the factor being studied.
Chance Level
This is the baseline level of recovery expected purely by random chance. The research shows this baseline changes greatly between different transcription factors.
Perturbation-based Evaluation
A method used to test model faithfulness by masking parts of the input sequence. This technique sometimes fails to accurately assess performance for certain factors.