Deep Learning for BioImaging: What Are We Learning?

arXiv:2603.13377 · cs.CV, cs.LG · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Deep Learning for BioImaging: What Are We Learning?".

Jane: The paper was written by Ivan Svatko, Maxime Sanchez, Ihab Bendidi, Gilles Cottrell and Auguste Genovesio from Université Paris Cité and Institut de Biologie de l'École Normale Supérieure and École Normale Supérieure and Université PSL and Institut Curie and Iktos and Inserm and Valence Labs and Recursion.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the bioimaging world, and it’s called “Deep Learning for Bioimaging: What Are We Learning?” And Jane, I have to say, that title just gets me. It’s almost cheeky, right?

Jane: It really is, Tom. It’s like the authors are standing up and saying, “We’ve been building all these fancy models, but do we actually know what they’re learning?” And that’s such an important question, especially in a field like biology where getting it wrong has real consequences.

Tom: Exactly. And the team behind it is pretty impressive. You’ve got researchers from École Normale Supérieure in Paris, plus folks from Institut Curie, and even people from companies like Iktos and Recursion. So it’s a mix of academic rigor and industry practicality.

Jane: That mix is crucial here because they’re not just theorizing. They’re testing these massive foundation models that companies are actually deploying. And the title is almost a challenge to the whole field. It’s asking us to slow down and check our assumptions.

Tom: Right. And the implications are huge. If we can’t trust that these models are learning real biology, then every drug discovery pipeline or diagnostic tool built on them is on shaky ground. This paper is essentially a reality check.

Jane: A reality check we desperately needed. And the way they go about it is by introducing some really clever baselines, which we’ll get into in a bit. But for now, let’s just sit with that title. What are we learning? Apparently, maybe not what we think.

Tom: And that’s the hook that got me. Let’s keep that question in mind as we break down what they actually found.

Summary: Tom: So, Jane, we’ve got the title. Now let’s talk about what this paper actually did. And the short version is, they ran a bunch of state-of-the-art models against some surprisingly simple baselines, and the results are humbling.

Jane: Humbling is the right word. They looked at two main types of microscopy data. You’ve got cell culture images, like the Cell Painting assays, and then you’ve got tissue images, like the histology slides from the HEST-1k dataset. And they tested these against things like untrained neural networks and even just pixel statistics.

Tom: And the wild part is, on the cell culture benchmarks, those untrained models were competitive with the big pretrained foundation models. I mean, we’re talking about a randomly initialized ViT going toe-to-toe with something like OpenPhenom, which was trained on millions of images.

Jane: That’s the headline finding for me. It suggests that a lot of the performance we’ve been celebrating on these benchmarks might be coming from low-level cues, like intensity or texture, rather than any deep biological understanding. The models are basically taking shortcuts.

Tom: And it gets even more interesting when you look at the tissue data. They introduced this baseline where they strip away all the visual appearance and just look at the spatial arrangement of cells, like a graph of cell positions. And that alone was enough to match or beat some out-of-domain image models on certain cancer datasets.

Jane: So the structure of the tissue, where the cells are, carries a ton of information. And that’s a really important insight. But it also means that when a model does well, we can’t just assume it’s because it understands cell morphology or staining patterns. It might just be counting cells.

Tom: Right. And that’s the core message of “Deep Learning for Bioimaging: What Are We Learning?”. The benchmarks we have are often not diagnostic enough to tell us what’s actually being learned. They’re too easy to game with simple tricks.

Jane: And that’s why they’re pushing for better baselines and more careful evaluation. Because if we can’t measure what’s learned, we can’t make progress. We’re just guessing.

Tom: So the summary is, we’ve been overestimating our models, and this paper gives us the tools to see that clearly. Next up, we need to talk about what they suggest we do about it.

Improvements: Tom: Alright, Jane, so we know the problem. These models are taking shortcuts. But what does this paper actually suggest we do about it? Because just saying “your benchmarks are bad” isn’t enough.

Jane: No, it isn’t. And they do offer a path forward. The biggest suggestion is to make baselines a standard part of any benchmark. Not as an afterthought, but as a core part of the evaluation. If your fancy foundation model can’t beat a randomly initialized network, you haven’t really learned anything.

Tom: And they also push for more diagnostic benchmarks. Ones that are designed to test specific capabilities, like whether a model is actually using cell morphology or just counting cells. That way, when a model does well, you know why it’s doing well.

Jane: Exactly. And they show how to do this with their structure-only baseline. By separating the spatial organization from the visual appearance, they can isolate what signal is actually being used. That’s a really powerful tool for understanding model behavior.

Tom: And they’re not just talking about it. They’re showing how to do it. They even have these synthetic experiments where they generate cell graphs with controlled properties to see if models can pick up on specific patterns. It’s like a controlled lab experiment for representation learning.

Jane: That’s the kind of rigor we need. And there’s another improvement they highlight, which is the need to look at intermediate layers, not just the final output. They found that performance can actually decrease in deeper layers for biological tasks, which is the opposite of what you see on ImageNet.

Tom: Right. On natural images, deeper layers are better because they build up these abstract concepts. But on microscopy, that trend doesn’t hold. So we need to think about where in the network the useful features are, and maybe we shouldn’t just be using the last layer.

Jane: And that has practical implications for how we build and use these models. It’s not just about the architecture, it’s about how we extract the features and what we expect them to represent. The paper is really asking us to be more thoughtful.

Tom: Thoughtful and humble. So the improvements are about better baselines, better benchmarks, and a better understanding of how these models actually work. Which brings us to the big picture.

Conclusion: Tom: So, Jane, we’ve covered a lot of ground on “Deep Learning for Bioimaging: What Are We Learning?”. Let’s wrap this up and say goodbye to this paper.

Jane: It’s been a great discussion. The core takeaway for me is that we need to be skeptical of our own successes. The paper shows that simple baselines can rival state-of-the-art models, which means we have to work harder to prove that our models are actually learning biology.

Tom: And that’s not a bad thing. It’s a call to action. We need better benchmarks that can distinguish between learning real biological features and just exploiting dataset quirks. And we need to embrace baselines as a critical part of the evaluation process.

Jane: The implications are huge for the field. If we can build more diagnostic benchmarks, we can actually make progress on understanding what makes a good representation. And that will lead to better models for drug discovery, disease diagnosis, and all the other amazing applications of bioimaging.

Tom: Absolutely. And it’s exciting to think about where this goes next. The authors are already hinting at future work on interpretability and on integrating different levels of biological abstraction. This paper is really just the beginning of a conversation.

Jane: A conversation we’re glad to be part of. So, with that, we’re going to say goodbye to this paper and get ready to dive into the next one. Thanks for listening, everyone.

Tom: And remember, when you see a great result, ask yourself, what’s the baseline? It might just save you from a very embarrassing mistake. See you next time.

Ivan Svatko, Maxime Sanchez, Ihab Bendidi, Gilles Cottrell, Auguste Genovesio

Université Paris Cité · Institut de Biologie de l'École Normale Supérieure · École Normale Supérieure · Université PSL · Institut Curie · Iktos · Inserm · Valence Labs · Recursion

cs.CV, cs.LG

Submitted: 2026-08-17

Updated: 2026-08-18

Code: https://github.com/bioptimus/releases

Project page: https://www.rxrx.ai/rxrx2

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: This paper investigates what microscopy foundation models truly learn by evaluating foundation models on curated cell-culture and tissue benchmarks, using ImageNet-1k as a natural-image reference.

Key concepts

Deep Learning for BioImaging
This refers to the application of deep learning models to analyze biological imaging data, such as cell culture images and tissue histology slides. The paper questions whether these models are truly learning biological concepts or just superficial patterns.
Baselines
In this context, baselines are simple comparison models, like untrained neural networks or pixel statistics. The paper argues that if a complex foundation model cannot outperform these simple baselines, it suggests the model is not learning meaningful biological information.
Spatial Arrangement
This concept involves looking at the structural organization of cells in tissue images rather than just their visual appearance. The paper found that analyzing the spatial arrangement of cells alone was sufficient to match or beat some out-of-domain image models on certain cancer datasets.

Terminology

Summary

This paper investigates what microscopy foundation models truly learn by evaluating foundation models on curated cell-culture and tissue benchmarks, using ImageNet-1k as a natural-image reference. The authors introduce two simple but informative baselines to make performance easier to interpret: (1) untrained models that use post-processed features from randomly initialized networks to test how much signal comes from architectural inductive biases and weak pixel correlations rather than learned biological content, and (2) disentangled tissue structures that represent histology images through the spatial organization of cells, testing whether models capture biological information beyond tissue morphology.

The paper states: In this work, we investigate what microscopy foundation models truly learn by evaluating fondation models on curated cell-culture and tissue benchmarks, using ImageNet-1k as a natural-image reference.

The authors evaluate widely used foundation models on these tasks and compare them to both baselines, finding that many models fall short of expectations.

Cellular Level Results: On RxRx3-core gene-gene retrieval (recall@5% of cosine similarities), the paper reports: "Unexpectedly, untrained ViTs perform comparably to pretrained ViTs and foundation models like OpenPhenom. Moreover, a minimal untrained SingleConv baseline is competitive with the best-performing methods, while pixel-statistics baselines recover a substantial fraction of the signal. In contrast, ResNet representations, whether pretrained or untrained, perform poorly on this task."

The results are stable across three random seeds and three folds, indicating that weight initialization does not impact untrained models.

Layer-wise Analysis: "By averaging performance across architectures and layers, we observe that performance generally decreases or remains stable in deeper layers for biological recall, in contrast to the monotonic improvements typically observed on natural image benchmarks. The paper notes: Here, all models except for the untrained ones demonstrate an increase in accuracy when evaluating hidden representations from deeper layers" for ImageNet classification.

JUMP-CP Results: Per-compound mean average precision (mAP) results show that "Several compounds (e.g., JCP 050797 and JCP 085227) exhibit near-random retrieval performance across all models, suggesting little to no detectable phenotypic effect. Other compounds (e.g., JCP 012818) are retrieved almost equally well by untrained baselines and pretrained models, indicating easily detectable phenotypic changes. In contrast, some compounds (e.g., JCP2022 025848) are only reliably retrieved by pretrained or foundation models, suggesting that these models capture more subtle discriminative features."

Tissue Level Results: On HEST-1k-1NN, in-domain foundation models offer a substantial increase in performance across most datasets. Surprisingly, however, for COAD and PRAD the performance of structure-based models is competitive with OOD vision encoders. The paper notes that On PRAD and COAD, we observe a surprisingly competitive scores between a structure-only encoder and H-Optimus-1 for a large subset of genes.

Representational Similarity: Pretrained and untrained ViTs yield highly correlated rankings, indicating shared relational structure; their rankings also correlate with DINOv3 and the SingleConv baseline.

In-Domain Models: The paper reports a cautiously optimistic perspective evaluating RxRx3-core embeddings from MAE-L/8 and MAE-G/8 models: Benchmarking on a balanced split of RxRx3-core shows a considerable improvement over all pretrained and untrained baselines, reaching 1.5 times higher averaged top 5% recall.

The paper states: Across cell culture (RxRx3-core, JUMP-CP) and tissue (HEST-1k, HEST-1k-1NN), the results show that benchmark performance does not consistently track acquisition of high-level biological abstractions.

The authors conclude: "Therefore, strong simple baselines are necessary to interpret a score. On RxRx3-core, untrained ViTs and SingleConv are competitive with pretrained models, while pixel-statistics features achieve non-trivial recall despite discarding spatial structure."

Regarding tissue structure, the paper states: "On HEST-1k-1NN, structure-only models (cell-centroid graphs) approach and sometimes exceed OOD image baselines on selected tissue categories (notably PRAD and COAD), and gene-wise analyses show subsets of genes that are comparatively predictable from cell coordinates."

The paper's overall message: Together, our results suggest that progress in microscopy image representation learning requires not only stronger models, but also more diagnostic benchmarks that measure what is actually learned.

Improvements for AI systems

Based on the paper's findings, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: I will integrate three mandatory baseline families into any microscopy representation learning evaluation framework:

  • Pixel-level statistics (channel-wise mean, std, skewness)

  • Untrained models (randomly initialized ViTs, CNNs)

  • Structure-only baselines (cell-centroid graphs, cell-count features)

What the improved system can do: Automatically calibrate benchmark scores by reporting the gap between a model's performance and these baselines. This prevents overclaiming—if a foundation model scores 0.65 recall but an untrained ViT scores 0.66, the system flags that the model has learned no biologically meaningful features beyond architectural priors.


In summary: The improved AI system will not just rank models—it will diagnose what they learn, flag shortcut exploitation, and provide interpretable, controlled baselines. This prevents costly mistakes in biological discovery where a model that appears to perform well may actually be exploiting batch effects or low-level intensity cues, leading to false conclusions about drug targets or disease mechanisms.

Abstract

Representation learning has driven major advances in natural image analysis by enabling models to acquire high-level semantic features. In microscopy imaging, however, it remains unclear what current representation learning methods actually learn. In this work, we conduct a systematic study of representation learning for the two most widely used and broadly available microscopy data types, representing critical scales in biology: cell culture and tissue imaging. To this end, we introduce a set of simple yet revealing baselines on curated benchmarks, including untrained models and simple structural representations of cellular tissue. Our results show that, surprisingly, state-of-the-art methods perform comparably to these baselines. We further show that, in contrast to natural images, existing models fail to consistently acquire high-level, biologically meaningful features. Moreover, we demonstrate that commonly used benchmark metrics are insufficient to assess representation quality and often mask this limitation. In addition, we investigate how detailed comparisons with these benchmarks provide ways to interpret the strengths and weaknesses of models for further improvements. Together, our results suggest that progress in microscopy image representation learning requires not only stronger models, but also more diagnostic benchmarks that measure what is actually learned.

Sources

Related papers