BlockFormer: Transformer-based inference from genomic contact maps

arXiv:2605.21617 · cs.LG, q-bio.QM · Submitted 2026-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "BlockFormer: Transformer-based inference from genomic contact maps".

Jane: BlockFormer introduces a novel, data-driven transformer architecture designed to infer per-entity parameters from interaction maps, such as those derived from Hi-C techniques, which is crucial for tasks like centromere localization.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we've got the paper "BlockFormer: Transformer-based inference from genomic contact maps," and it sounds like they are tackling a really tricky problem where you have interaction maps, like Hi-C data, and you need to figure out hidden parameters.

Jane: That's right, Tom; essentially, this paper is looking at how to take those complex interaction maps—which show how different parts of the genome touch each other—and use a transformer structure to infer specific things about those entities.

Lu: What's fascinating here is that they are treating this as a generic inverse problem where you have these maps summarizing pairwise interactions, and the goal is to figure out the parameters given those blocks of variable numbers and sizes.

Meng: It sounds like they’re trying to solve something that existing methods struggle with, which is a lot when you're dealing with real biological data that's never exactly the same twice.

Lalam: From my perspective, the core idea of this paper is powerful because it moves away from older techniques that were too slow or too rigid for real-world genomic data.

Tom: Exactly, and what they're doing is using a transformer architecture to handle those variable block numbers and sizes, which is a big deal because most models prefer fixed inputs.

Jane: The paper introduces this "three dimensional positional encoding" as their main trick to manage that variability while still capturing the patterns within each block.

Lu: They achieve this by cutting the input map into squared patches and adding position embeddings defined by a three dee-position vector, which lets the model see both the block index and where it sits inside that block.

Meng: So, instead of treating every interaction in a uniform way, this architecture allows for per-block pattern learning while still aggregating information across those blocks.

Lalam: And I think that ability to model structure explicitly rather than just relying on global message passing is what makes this approach interesting for understanding the genome.

Tom: Moving on, the paper talks about how they train this BlockFormer architecture using synthetic data to make it robust across different map resolutions and entity sizes.

Title and authors: Jane: They are training it on map/parameter pairs indexed by their entities, and a key part of their strategy is training on small interaction maps while expecting it to generalize to larger ones.

Lu: The training loss they use is a regression loss, L(phi) = (one/M) sum m=one M X one m M BF phi(C m i) - m i two two which means they are directly optimizing for the parameter estimation k.

Meng: From an engineering standpoint, that training strategy makes sense because it forces the model to learn how to handle different amounts of parameter information and map resolutions, which is a huge practical consideration.

Lalam: It’s smart that they are using synthetic maps generated by a custom simulator instead of relying solely on real data for training, because that lets them control the exact variability in block numbers and sizes.

Tom: Speaking of real data, they tested this method on centromere localization from Hi-C maps across several species including yeast and *A. thaliana*, showing results like a near-resolution accuracy of zero point five eight for *S.M.*.

Jane: That near-resolution accuracy is impressive, especially when they demonstrate that the estimations are often under the resolution threshold across various chromosomes.

Lu: It’s interesting that they specifically compared their method against Centurion, which infers centromere positions by fitting Gaussian profiles to interaction peaks, showing BlockFormer is faster in both initialization and fitting.

Meng: Speed matters immensely when you're dealing with large-scale biological screening; if this inference engine is significantly quicker than the traditional non-amortized optimization problem Centurion, it opens up a lot of practical avenues.

Lalam: And that speed advantage means we can run these kinds of inferences much more frequently and at a larger scale without hitting computational bottlenecks.

Tom: Beyond just finding the position, the paper also discusses how BlockFormer can be used within a Bayesian framework to provide an estimate of uncertainty for those parameter choices.

Jane: They decompose the inference by running one inference per entity in parallel and using BlockFormer as a summary statistic S phi that approximates the conditional expectation of theta given the map C.

Lu: This summary statistic approach is clever because it's designed to preserve first-order information when summarizing the map C as the threshold epsilon approaches zero.

Title and authors: Meng: That sounds like a solid way to handle uncertainty quantification without having to run a massive number of full simulations for every possible parameter set.

Lalam: And they use Sequential Neural Posterior Estimation, which allows them to iteratively refine those posterior estimates, giving us a principled way to quantify the uncertainty in our biological guesses.

Tom: So, looking at the overall picture of this paper titled "BlockFormer: Transformer-based inference from genomic contact maps," we see a model that is specifically engineered to handle the messy variability in genomic interaction data.

Jane: It seems like they've managed to build something that is both flexible enough for diverse biological maps and accurate enough for practical applications like centromere prediction.

Lu: The architecture itself, with its three dee positional encoding, represents a significant step away from the localized graph approaches that usually dominate this field.

Meng: From an engineering standpoint, the ability to handle heterogeneous inputs without requiring manual re-processing for every new dataset is what makes this system scalable.

Lalam: I think the real cultural impact here is that if we can reliably infer critical genomic features with better speed and accuracy, it speeds up discovery across many biological systems.

Tom: So, to wrap this up on the implications of BlockFormer: Transformer-based inference from genomic contact maps, it gives us a much faster and more flexible way to extract key genomic parameters from complex contact maps.

Jane: It shows that transformer architectures are not just for image processing anymore; they can be tailored effectively to infer hidden structures in biological data.

Lu: The flexibility in handling variable block sizes and numbers across different species suggests a broad applicability to other structural biology problems, not just centromeres.

Meng: For us at the startup, the implication is that we can build inference engines that work reliably on messy data structures without needing a custom pre-processing pipeline for every new type of interaction map.

Lalam: Ultimately, this paper gives us a powerful tool for making sense of genomic organization with quantified uncertainty, which is crucial when we are trying to trust the results in biological research.

The paper's summary: Tom: So we've got this paper, "BlockFormer: Transformer-based inference from genomic contact maps," and basically, they've developed a transformer that can take messy interaction data from genomics and reliably pull out the hidden parameters we need.

Jane: Exactly, Tom; think of it as giving a sophisticated AI a map of how biological entities are touching and asking it to figure out exactly where things like centromeres are located, even when that map is full of weird block sizes.

Lu: What's really interesting is the mechanism they use to handle those variable block structures with that three-dimensional positional encoding; it lets the model look at both the block number and where a patch sits within that block, which is a significant departure from standard vision transformers.

Meng: From an engineering standpoint, that ability to handle such heterogeneous inputs without needing custom pre-processing for every new dataset is what makes this system scalable; I'm curious how robust the simulator they used really was for generating that synthetic data.

Lalam: The fact that they are training it to generalize across different map resolutions and entity sizes shows a real leap in flexibility, which I think will profoundly impact how we build AI models for biological data.

Tom: And the results are pretty solid; they showed near-resolution accuracy even under noisy conditions across several species, which means this isn't just theoretical stuff for a lab; it’s actually performing well on real biological data.

Jane: That level of accuracy is what makes me excited because it suggests we can move past rough estimates and start getting precise positional information for key genomic structures.

Lu: Furthermore, they included a Bayesian framework to quantify uncertainty, which is important because knowing the confidence level in an inference is as critical as the estimate itself.

Meng: I'm looking at that uncertainty quantification part; if we can get a probability distribution of possible positions instead of just a single guess, it changes how we trust the output for actual experimental design.

Lalam: It’s about building systems that don't just give an answer but tell you *how sure* they are about that answer, and I think this level of principled uncertainty modeling is a cultural step forward in how we build reliable scientific AI.

Tom: So, to sum up, BlockFormer gives us a powerful inference engine that’s fast, handles variable genomic map structures gracefully through its unique encoding, and provides quantifiable uncertainty for critical tasks like centromere localization across diverse organisms.

Jane: It truly shows how specialized transformer architectures can be designed to solve specific inverse problems in biology, moving beyond general image tasks into the realm of complex structural inference.

Lu: The potential here is huge; imagine applying this same architecture to other complex biological organization problems, like protein folding or chromatin looping structures, where block variability is a constant challenge <ref:two thousand six hundred five point two one six one seven#pg0.

Meng: I'm seeing the practical implication as a system that can run efficiently on existing high-throughput sequencing data without requiring us to build entirely new, slow pipelines for every single species we want to analyze.

Lalam: This kind of robust and accurate inference capability could fundamentally improve the speed and reliability of biological discovery across the entire life sciences sector, making complex genomic maps much more accessible to researchers globally <ref:two thousand six hundred five point two one six one seven#pg0.

The paper's improvements: Tom: So, to wrap up on the practical side of BlockFormer, the authors didn't just stop at inference; they proposed ways to make it even better by focusing on how we handle uncertainty and how we can adapt it for different biological scenarios.

Jane: They suggested decomposing the inference task into running one estimation per entity in parallel, which helps us get a clearer picture of what each parameter is doing, and using BlockFormer as a summary statistic to capture that information efficiently.

Lu: That summary statistic approach is quite clever because it’s designed to keep the first-order information intact even when we condense the entire interaction map into just one number or vector, which is really smart for managing complexity.

Meng: I like that idea of summarizing the data as a threshold epsilon approaches zero; that suggests a way to maintain precision while simplifying the model's workload during high-throughput analysis.

Lalam: From my perspective, this moves us closer to building AI agents that don't just give an answer but can provide a principled probability distribution around that answer, which is essential for deploying AI in high-stakes biological research.

Tom: And they also used Sequential Neural Posterior Estimation to iteratively refine those initial estimates, meaning the model gets better and more accurate as it runs through the process.

Jane: That iterative refinement sounds like a solid way to handle noise in the data because instead of trying to get one perfect guess at once, it allows the AI to correct its course based on what it finds.

Lu: It really highlights that this isn't just about a single forward pass; it’s about building an entire inference pipeline that is adaptive and self-correcting, which opens up avenues for much more sophisticated data interpretation.

Meng: That adaptability is what I need for deployment; if the system can iteratively improve its estimate based on new information, it becomes much more reliable when applied to messy, real-world genomic contact maps.

Lalam: This kind of principled quantification of uncertainty means we are building trust into the AI's output, which is a huge step toward making scientific AI tools truly trustworthy and influential in shaping our understanding of life.

Conclusion: Tom: So, to wrap up on "BlockFormer: Transformer-based inference from genomic contact maps," we’ve seen how this architecture uses its unique positional encoding to handle variable block structures in interaction maps very effectively.

Jane: It really is a neat way for an AI to navigate the messy reality of biological data when the inputs aren't perfectly uniform, Tom; it teaches us a lot about building flexible models.

Lu: I think what stands out most about this work is how it treats the problem as a sequence processing task, which allows us to apply these powerful transformer concepts directly to structural biology problems in a way that was previously much harder.

Meng: I'm impressed by how fast this inference engine runs compared to traditional methods like Centurion; having something that's both accurate and quick is what makes it viable for actual high-throughput genomic analysis.

Lalam: This paper gives us a powerful tool for making sense of genomic organization with quantified uncertainty, which is crucial when we are trying to trust the results in biological research.

Tom: Exactly; the ability to provide that probability distribution around our inferences means we’re not just getting a guess, but a reliable estimate of confidence.

Jane: And that confidence level is what separates good data processing from truly meaningful scientific discovery, Tom; it lets us know when to trust the AI and when to dig deeper into the raw data.

Lu: The potential here is huge; imagine applying this same framework to other complex biological organization problems, like protein folding or chromatin looping structures, where block variability is a constant challenge.

Meng: I’m seeing the practical implication as a system that can run efficiently on existing high-throughput sequencing data without requiring us to build entirely new, slow pipelines for every single species we want to analyze.

Lalam: This kind of robust and accurate inference capability could fundamentally improve the speed and reliability of biological discovery across the entire life sciences sector, making complex genomic maps much more accessible to researchers globally.

Univ. Grenoble Alpes · inria

cs.LG, q-bio.QM

Submitted: 2026-05-20

Updated: 2026-09-30

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 78/100

The gist: BlockFormer introduces a novel, data-driven transformer architecture designed to infer per-entity parameters from interaction maps, such as those derived from Hi-C techniques, which is crucial for

Key concepts

BlockFormer Architecture
This is the core transformer model tailored for interaction maps with irregular block structures. It uses a 'three-dimensional positional encoding' to process input blocks of varying sizes and numbers effectively, allowing it to capture local patterns and aggregate information across multiple blocks simultaneously.
3D Positional Encoding
This mechanism allows the model to understand where every part of the interaction map is located. It uses a 3D vector (block index, relative patch position) to encode the location, which helps the transformer process variable block sizes and capture non-local relationships across different parts of the map.
Synthetic Data Training Strategy
The model is trained on synthetic 'map/parameter pairs' rather than real data alone. This strategy involves training on small maps while forcing generalization to larger ones, ensuring the model learns to handle varying amounts of parameter information and map resolutions robustly.

Terminology

Summary

BlockFormer introduces a novel, data-driven transformer architecture designed to infer per-entity parameters from interaction maps, such as those derived from Hi-C techniques, which is crucial for tasks like centromere localization. This method addresses the challenge of handling interaction maps exhibiting variable block numbers and sizes by leveraging a transformer's ability to process sequences of blocks while incorporating block-aware positional encoding. By utilizing a custom simulator to generate abundant synthetic data, BlockFormer enables fast and accurate inference across a wide range of species and genome sizes, outperforming traditional methods like Centurion.

BlockFormer Architecture

The core innovation is the BlockFormer architecture, which is a transformer-based model tailored for interaction maps with variable block structures. The input to the model is any sequence of blocks of interactions between entity k and others, denoted as Ck. The architecture employs a three dimensional positional encoding that allows handling variable block sizes and numbers while capturing per-block patterns and aggregating non-local information across multiple blocks. This mechanism is achieved through:

  1. Cutting the input map into squared patches of size (P, P), respecting the block-wise structure via asymmetric zero-padding.

  2. Projecting these patches to a constant latent vector size D, resulting in patch embeddings of size (N, D).

  3. Adding position embeddings of size (N + 1, D) to retain the position of each patch in the original map. These are defined as a 3D-position vector (i, j, k) where i is the block index, and (j, k) the relative position of the patch in the block i.

  4. The resulting sequence serves as input to a transformer adopting the classical series of B blocks alternating Multi-head self-attention, Layernorm, and MLP presented in vision transformer (ViT) [8].

Training Strategy for Flexibility

To ensure generalization across varying map structures, BlockFormer employs a specific training strategy designed for flexibility. The model is trained on synthetic data consisting of map/parameter pairs (Ci, θi), indexed by their corresponding entity i. Key aspects of this strategy include:

  1. Training on small interaction maps while allowing it to generalize to larger ones, ensuring the model handles varying amounts of parameter information and map resolutions.

  2. Ensuring training maps are as representative as possible by choosing, per batch, a number of blocks, the size of each block, and the entity to test.

  3. The natural training loss is a regression loss: L(ϕ) = (1/M) Σ X [1≤m≤M] BFϕ(C m i) − ˜θ m i22, where ˜θ i is any normalized parameter.

Genomic Application and Performance

BlockFormer is applied to the problem of centromere localization from Hi-C maps, where entities are chromosomes and parameters are centromere positions. The method's performance is validated across diverse genome sizes and species, including seven yeast species, the parasite P. falciparum, and the plant A. thaliana. Key findings include:

  1. The model achieves near-resolution accuracy (e.g., 0.58 for S.M.) in centromere prediction across various chromosomes, demonstrating robustness to varying block sizes and numbers of blocks (most of the estimations are under the resolution).

  2. BlockFormer is consistently faster than Centurion in both initialization and fitting, showing a significant speedup (e.g., 2.64 s versus 12.37 s for S.C.).

  3. The method is robust to various challenges, including varying numbers and sizes of blocks, different sequencing depths, and diverse spot pattern structures (Gaussian, square, elliptical, ring spots), often achieving sub-resolution precision even under noisy map conditions (BlockFormer is more accurate and faster than Centurion).

Inference with Uncertainty Quantification

Beyond point estimation, BlockFormer can serve as an informative summary statistic within a Bayesian framework to quantify uncertainty. The inference strategy targets the posterior density p(θCref) by decomposing the problem into I subproblems, inferring θi from Ci. This is achieved through:

  1. Decomposing the inference: we perform in parallel one inference per entity, estimating θi from Ci.

  2. Using BlockFormer as a summary statistic Sϕ that approximates the conditional expectation E [θC], which preserves first-order information when summarizing C as the threshold ϵ approaches 0.

  3. Employing Sequential Neural Posterior Estimation (SNPE) to iteratively refine the posterior estimate, allowing for a principled quantification of uncertainty in parameter estimates.

Ablations and Robustness

The paper details extensive ablations to confirm the necessity of its components.

Improvements for AI systems

Here are specific improvements to AI systems derived from the BlockFormer methodology, along with what these improved systems can achieve:


)Based on BlockFormer's capabilities, here are specific enhancements for various AI systems:

  1. The core improvement is a shift from computationally expensive, non-amortized optimization methods (like Centurion) to a fast, transformer-based inference model.

  2. The use of a custom simulator allows for massive data generation without needing slow biophysical simulations, enabling the training of robust models on diverse and varied data structures.

Here are the specific improvements and resulting system capabilities:

  1. The architecture is specifically designed to handle variable block numbers and sizes using a 3D positional encoding, moving beyond fixed-size patch inputs common in Vision Transformers (ViT).

  2. The training strategy explicitly enforces generalization across varying map resolutions, block counts, and genomic entity sizes by sampling diverse configurations during training.

  3. The model is inherently suited for inverse problems where the goal is to infer hidden parameters (like centromere positions) from summary statistics (interaction maps).

Improved AI Systems and Capabilities:

  1. A fast, robust, and scalable system for genome structure inference that can be applied across a wide range of species and genome sizes.

  2. A system capable of identifying critical genomic features (like centromeres) with sub-resolution precision (e.g., 0.08 error or better for yeast) in significantly less time than traditional methods, making it suitable for high-throughput biological screening or large-scale data analysis across multiple organisms.

  3. A flexible inference engine capable of handling heterogeneous interaction maps—those with varying numbers and sizes of entities—without requiring manual retraining or complex pre-processing steps for each new dataset.

  4. A system that can perform loop localization (identifying CTCF-mediated interactions) by adapting the architecture to use localized cis-blocks, offering superior performance over existing methods like Centurion in terms of speed and accuracy on complex, noisy maps with diverse spot patterns (Gaussian, square, elliptical).

  5. A Bayesian inference framework allowing for principled quantification of uncertainty in inferred parameters (posterior estimation), providing not just a point estimate but a full probability distribution that reflects the inherent biological noise and heterogeneity of the interaction map data.

Abstract

Locating genomic features from genomic contact maps, such as centromere identification from genome-wide chromosome conformation capture techniques, notably Hi-C, can be formulated as an inverse problem: infer one parameter per entity given a map summarizing pairwise interactions through blocks of variable numbers and sizes. In this work, we introduce a data-driven approach that leverages shared structure between these contact maps, such as global alignment between localized patterns, while handling the variability in number and size of chromosomes arising in real-world data. Our approach relies on a transformer architecture capable of handling such variability and a custom simulator to generate abundant, yet computationally cheap synthetic data for training. Applied to the problem of centromere localization, the method recovers genomic positions at or below the map resolution across species with various genome sizes using a coarse-to-fine refinement step when chromosome sizes fall outside the training range. We concentrate our evaluation on Hi-C data, where the block structure is well characterized and ground truth is available.

Sources

Related papers