Bio papers — 2026-09-16
Today's focus centers on how we are evaluating different methods for attributing genomic sequence models because the way we currently measure success isn't truly valid across various transcription factors. We looked at how much of a known motif these methods recover or how their predictions degrade when evidence is removed.
Neither metric makes sense without knowing what would happen purely by random chance. The core issue is that the uniform chance level for contiguous motif overlap varies significantly across the twenty-six eight transcription factors in UniBind, spanning from 0.0118 to 0.0427. This spread is a three point six fold difference determined only by motif length and window size.
This spread means that raw recovery rates are not comparable quantities for some factors because their bootstrap intervals around the chance levels do not overlap at all. Correcting for this chance level changes how we classify the factors. For instance, one factor previously reported as a resolution failure moves up to be second highest, while two others reported as complete failures fall below the chance level.
Furthermore, we found that perturbation-based evaluation can fail its own test of faithfulness. This is because for one specific factor, even when all input is masked, the score remains above the decision boundary and the curve isn't consistently increasing with more masked positions. We are providing closed form chance levels and a corrected score to address these comparability issues.
Today's papers
- Recovery Rates Are Not Comparable Across Transcription Factors: Chance Correction for Attribution Evaluation This paper shows that recovery rates for transcription factors cannot be directly compared without correcting for random chance. [paper]
The papers
Important terms
- Genomic Sequence Models
- These are computational models used to predict how transcription factors bind to DNA sequences. The research focuses on fairly evaluating these models across different transcription factors.
- Contiguous Motif Overlap
- This metric measures how much of a known DNA motif is successfully recovered by the prediction methods. Its performance varies significantly depending on the factor being studied.
- Chance Level
- This is the baseline level of recovery expected purely by random chance. The research shows this baseline changes greatly between different transcription factors.
- Perturbation-based Evaluation
- A method used to test model faithfulness by masking parts of the input sequence. This technique sometimes fails to accurately assess performance for certain factors.