Deep Learning for BioImaging: What Are We Learning?
summary
The gist
This paper investigates what microscopy foundation models truly learn by evaluating foundation models on curated cell-culture and tissue benchmarks, using ImageNet-1k as a natural-image reference.
In short
The episode discusses Ivan Svatko et al.'s paper, "Deep Learning for BioImaging: What Are We Learning?" which critically examines what deep learning models actually learn in bioimaging. The hosts find that state-of-the-art models can perform well on simple baselines, suggesting they often rely on low-level cues rather than deep biological understanding. The paper suggests improving evaluation by using stronger baselines and designing more diagnostic benchmarks.
Key concepts
- Deep Learning for BioImaging
- This refers to the application of deep learning models to analyze biological imaging data, such as cell culture images and tissue histology slides. The paper questions whether these models are truly learning biological concepts or just superficial patterns.
- Baselines
- In this context, baselines are simple comparison models, like untrained neural networks or pixel statistics. The paper argues that if a complex foundation model cannot outperform these simple baselines, it suggests the model is not learning meaningful biological information.
- Spatial Arrangement
- This concept involves looking at the structural organization of cells in tissue images rather than just their visual appearance. The paper found that analyzing the spatial arrangement of cells alone was sufficient to match or beat some out-of-domain image models on certain cancer datasets.
Terminology used across episodes
This episode discusses
- Deep Learning for BioImaging: What Are We Learning? · Paper Radio
- Benchmarking Transcriptomics Foundation Models for Perturbation Analysis: one PCA still rules them all
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- A simple yet effective baseline for non-attributed graph classification
- Vision Transformers Need Registers
- Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
- A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics
- Revisiting the Platonic Representation Hypothesis: An Aristotelian View
- RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy
- Do Vision Transformers See Like Convolutional Neural Networks?
- What's Hidden in a Randomly Weighted Neural Network?
- Weakly supervised cross-modal learning in high-content screening
- TxPert: Leveraging Biochemical Relationships for Out-of-Distribution Transcriptomic Perturbation Prediction
The paper
Deep Learning for BioImaging: What Are We Learning? · Read on arXiv
Ivan Svatko, Maxime Sanchez, Ihab Bendidi, Gilles Cottrell, Auguste Genovesio
Université Paris Cité · Institut de Biologie de l'École Normale Supérieure · École Normale Supérieure · Université PSL · Institut Curie · Iktos · Inserm · Valence Labs · Recursion
Representation learning has driven major advances in natural image analysis by enabling models to acquire high-level semantic features. In microscopy imaging, however, it remains unclear what current representation learning methods actually learn. In this work, we conduct a systematic study of representation learning for the two most widely used and broadly available microscopy data types, representing critical scales in biology: cell culture and tissue imaging. To this end, we introduce a set of simple yet revealing baselines on curated benchmarks, including untrained models and simple structural representations of cellular tissue. Our results show that, surprisingly, state-of-the-art methods perform comparably to these baselines. We further show that, in contrast to natural images, existing models fail to consistently acquire high-level, biologically meaningful features. Moreover, we demonstrate that commonly used benchmark metrics are insufficient to assess representation quality and often mask this limitation. In addition, we investigate how detailed comparisons with these benchmarks provide ways to interpret the strengths and weaknesses of models for further improvements. Together, our results suggest that progress in microscopy image representation learning requires not only stronger models, but also more diagnostic benchmarks that measure what is actually learned.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Deep Learning for BioImaging: What Are We Learning?".
Jane: The paper was written by Ivan Svatko, Maxime Sanchez, Ihab Bendidi, Gilles Cottrell and Auguste Genovesio from Université Paris Cité and Institut de Biologie de l'École Normale Supérieure and École Normale Supérieure and Université PSL and Institut Curie and Iktos and Inserm and Valence Labs and Recursion.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the bioimaging world, and it’s called “Deep Learning for Bioimaging: What Are We Learning?” And Jane, I have to say, that title just gets me. It’s almost cheeky, right?
Jane: It really is, Tom. It’s like the authors are standing up and saying, “We’ve been building all these fancy models, but do we actually know what they’re learning?” And that’s such an important question, especially in a field like biology where getting it wrong has real consequences.
Tom: Exactly. And the team behind it is pretty impressive. You’ve got researchers from École Normale Supérieure in Paris, plus folks from Institut Curie, and even people from companies like Iktos and Recursion. So it’s a mix of academic rigor and industry practicality.
Jane: That mix is crucial here because they’re not just theorizing. They’re testing these massive foundation models that companies are actually deploying. And the title is almost a challenge to the whole field. It’s asking us to slow down and check our assumptions.
Tom: Right. And the implications are huge. If we can’t trust that these models are learning real biology, then every drug discovery pipeline or diagnostic tool built on them is on shaky ground. This paper is essentially a reality check.
Jane: A reality check we desperately needed. And the way they go about it is by introducing some really clever baselines, which we’ll get into in a bit. But for now, let’s just sit with that title. What are we learning? Apparently, maybe not what we think.
Tom: And that’s the hook that got me. Let’s keep that question in mind as we break down what they actually found.
Summary: Tom: So, Jane, we’ve got the title. Now let’s talk about what this paper actually did. And the short version is, they ran a bunch of state-of-the-art models against some surprisingly simple baselines, and the results are humbling.
Jane: Humbling is the right word. They looked at two main types of microscopy data. You’ve got cell culture images, like the Cell Painting assays, and then you’ve got tissue images, like the histology slides from the HEST-1k dataset. And they tested these against things like untrained neural networks and even just pixel statistics.
Tom: And the wild part is, on the cell culture benchmarks, those untrained models were competitive with the big pretrained foundation models. I mean, we’re talking about a randomly initialized ViT going toe-to-toe with something like OpenPhenom, which was trained on millions of images.
Jane: That’s the headline finding for me. It suggests that a lot of the performance we’ve been celebrating on these benchmarks might be coming from low-level cues, like intensity or texture, rather than any deep biological understanding. The models are basically taking shortcuts.
Tom: And it gets even more interesting when you look at the tissue data. They introduced this baseline where they strip away all the visual appearance and just look at the spatial arrangement of cells, like a graph of cell positions. And that alone was enough to match or beat some out-of-domain image models on certain cancer datasets.
Jane: So the structure of the tissue, where the cells are, carries a ton of information. And that’s a really important insight. But it also means that when a model does well, we can’t just assume it’s because it understands cell morphology or staining patterns. It might just be counting cells.
Tom: Right. And that’s the core message of “Deep Learning for Bioimaging: What Are We Learning?”. The benchmarks we have are often not diagnostic enough to tell us what’s actually being learned. They’re too easy to game with simple tricks.
Jane: And that’s why they’re pushing for better baselines and more careful evaluation. Because if we can’t measure what’s learned, we can’t make progress. We’re just guessing.
Tom: So the summary is, we’ve been overestimating our models, and this paper gives us the tools to see that clearly. Next up, we need to talk about what they suggest we do about it.
Improvements: Tom: Alright, Jane, so we know the problem. These models are taking shortcuts. But what does this paper actually suggest we do about it? Because just saying “your benchmarks are bad” isn’t enough.
Jane: No, it isn’t. And they do offer a path forward. The biggest suggestion is to make baselines a standard part of any benchmark. Not as an afterthought, but as a core part of the evaluation. If your fancy foundation model can’t beat a randomly initialized network, you haven’t really learned anything.
Tom: And they also push for more diagnostic benchmarks. Ones that are designed to test specific capabilities, like whether a model is actually using cell morphology or just counting cells. That way, when a model does well, you know why it’s doing well.
Jane: Exactly. And they show how to do this with their structure-only baseline. By separating the spatial organization from the visual appearance, they can isolate what signal is actually being used. That’s a really powerful tool for understanding model behavior.
Tom: And they’re not just talking about it. They’re showing how to do it. They even have these synthetic experiments where they generate cell graphs with controlled properties to see if models can pick up on specific patterns. It’s like a controlled lab experiment for representation learning.
Jane: That’s the kind of rigor we need. And there’s another improvement they highlight, which is the need to look at intermediate layers, not just the final output. They found that performance can actually decrease in deeper layers for biological tasks, which is the opposite of what you see on ImageNet.
Tom: Right. On natural images, deeper layers are better because they build up these abstract concepts. But on microscopy, that trend doesn’t hold. So we need to think about where in the network the useful features are, and maybe we shouldn’t just be using the last layer.
Jane: And that has practical implications for how we build and use these models. It’s not just about the architecture, it’s about how we extract the features and what we expect them to represent. The paper is really asking us to be more thoughtful.
Tom: Thoughtful and humble. So the improvements are about better baselines, better benchmarks, and a better understanding of how these models actually work. Which brings us to the big picture.
Conclusion: Tom: So, Jane, we’ve covered a lot of ground on “Deep Learning for Bioimaging: What Are We Learning?”. Let’s wrap this up and say goodbye to this paper.
Jane: It’s been a great discussion. The core takeaway for me is that we need to be skeptical of our own successes. The paper shows that simple baselines can rival state-of-the-art models, which means we have to work harder to prove that our models are actually learning biology.
Tom: And that’s not a bad thing. It’s a call to action. We need better benchmarks that can distinguish between learning real biological features and just exploiting dataset quirks. And we need to embrace baselines as a critical part of the evaluation process.
Jane: The implications are huge for the field. If we can build more diagnostic benchmarks, we can actually make progress on understanding what makes a good representation. And that will lead to better models for drug discovery, disease diagnosis, and all the other amazing applications of bioimaging.
Tom: Absolutely. And it’s exciting to think about where this goes next. The authors are already hinting at future work on interpretability and on integrating different levels of biological abstraction. This paper is really just the beginning of a conversation.
Jane: A conversation we’re glad to be part of. So, with that, we’re going to say goodbye to this paper and get ready to dive into the next one. Thanks for listening, everyone.
Tom: And remember, when you see a great result, ask yourself, what’s the baseline? It might just save you from a very embarrassing mistake. See you next time.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization