Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation

arXiv:2608.07562 · cs.CV, cs.LG · Submitted 2026-08-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation".

Jane: The paper was written by Nafis Fuad, Xiaodong Qian and Dongxiao Zhu from Wayne State University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper with a real mouthful of a title — "Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision–Language Models for Centimeter-Level Flood Depth Estimation." Jane, I need you to translate that for our listeners.

Jane: Happy to, Tom. So the core idea is teaching a computer to look at a photo of a flooded street and tell you exactly how deep the water is — down to the centimeter. Not "oh, that's about knee-high," but "that's twenty-three centimeters." And the fancy title is basically saying they used a clever technique to figure out which parts of the model actually do the work, so they only had to train those parts.

Tom: And that matters because these vision-language models are enormous. We're talking about an eleven-billion-parameter model here. Training all of it takes serious computing power.

Jane: Right, and the team at Wayne State University — Nafis Fuad, Xiaodong Qian, and Dongxiao Zhu — they trained a baseline version that works really well. It gets the depth wrong by less than half a centimeter on average. But then they asked a really interesting question: do we need to train everything, or just the parts that matter?

Tom: And that's where the "mechanistic interpretability" part comes in. They basically opened up the hood of the model and looked at which layers were actually learning the flood-depth information.

Jane: Exactly. They found that one specific layer — layer twenty-three — is where the model first figures out depth. It's like the moment in a detective story where all the clues click together. And once they knew that, they could train just a handful of layers instead of all forty.

Tom: So they went from fifty-four million trainable parameters down to about six or seven million. That's an eighty-eight percent reduction. And the model still performs almost as well.

Jane: Almost as well on synthetic data, and actually better than the existing published baseline on real-world flood photos. That's the part that got me excited — this isn't just a lab trick. It's a practical way to make these models cheaper to train and still deploy them in the real world.

Tom: And we should say why this matters for regular people. When a city floods, knowing whether a road has ten centimeters or thirty centimeters of water is the difference between driving through safely and getting stranded — or worse, for electric vehicles, risking battery damage.

Jane: Right, and that's what we'll dig into next — what the paper actually did to build this system and why the synthetic data they created was such a big part of the story.

Summary: Tom: So Jane, we've got the title unpacked. Now let's talk about what the paper actually did. And the first thing that struck me was the dataset. They didn't go out and photograph real floods — they built a synthetic world.

Jane: Right, they used Unreal Engine five which is a video game engine, to create photorealistic flooded street scenes. And they generated two point eight one million images. That's a staggering number.

Tom: It is. And they varied everything systematically — seven different vehicle types, four weather conditions, and water depths from zero to forty centimeters. The single-vehicle scenes used five-centimeter steps, and the mixed-vehicle scenes used one-centimeter steps.

Jane: And that finer resolution in the mixed scenes turned out to be crucial. When they trained the model first on single-vehicle images, the error was about five point six five centimeters. But after training on the mixed-vehicle scenes, the error dropped to zero point four zero centimeters. That's a ninety-three percent improvement.

Tom: Why do you think the mixed scenes helped so much?

Jane: Think about it this way — if you see one car in a flooded street, you have one reference point for where the waterline sits. But if you see three cars, you can compare them. A sedan sits lower than an SUV, so the water might be at the sedan's headlights but only at the SUV's tires. That gives the model multiple clues to triangulate the actual depth.

Tom: That makes sense. And the model they fine-tuned is LLaMA three point two-11B Vision, which is an open-source model from Meta. They used a technique called QLoRA to keep the training manageable on a single GPU.

Jane: Right, QLoRA is a way of fine-tuning that quantizes the model to four-bit precision and only trains small adapter matrices. It's like renovating a house by changing the furniture instead of rebuilding the walls.

Tom: And the results on the synthetic test set were impressive — ninety-seven point six percent of predictions were within five centimeters of the true depth. That's well within the tolerance needed to decide whether a road is passable.

Jane: But here's the thing that really sets this paper apart — they didn't just build a model and stop. They asked why it works. And that's what led them to the mechanistic interpretability analysis, which is the heart of the paper.

Tom: So next we should talk about that analysis and what they found when they looked inside the model.

Jane: Exactly — and it's genuinely surprising what they discovered about which layers matter and which ones don't.

Improvements: Tom: So Jane, we've established the model works. Now let's talk about what the paper improves over previous approaches. And I think the key word here is "continuous."

Jane: Right, and that's a big deal. Previous systems for estimating flood depth from images mostly put water into categories — tire-level, bumper-level, door-level. That's coarse. You might get a range like ten to twenty centimeters, but not an exact number.

Tom: And the paper points out that those categories aren't good enough for real decisions. A difference of a few centimeters can determine whether an electric vehicle's battery is at risk, or whether a road is passable at all.

Jane: Exactly. And there were also systems using large multimodal models like GPT-4V, but those had errors above eight centimeters and they're proprietary black boxes. You can't inspect how they make decisions, which matters for safety-critical applications.

Tom: So the improvement here is threefold — centimeter-level precision, open-source transparency, and the interpretability analysis that tells you where the model stores its knowledge.

Jane: And that last part led to a genuinely clever improvement. They used four different analysis techniques — linear probing, logit lens, CKA similarity, and attention entropy — to rank which cross-attention layers actually encode depth information.

Tom: And the finding was that layer twenty-three is the pivot point. Before that layer, depth isn't decodable from the model's representations. At layer twenty-three it suddenly is. And the layers that changed the most during training weren't necessarily the ones doing the encoding.

Jane: That's the counterintuitive part. Layers thirteen through twenty-two changed a lot — their representations drifted significantly. But they didn't contribute to depth encoding at all. Meanwhile, layers twenty-three through thirty-eight stayed closer to the original model but were the ones where depth became readable.

Tom: So they were restructuring the visual features early on and then encoding the depth later. The paper calls this a "restructure-then-encode" pattern.

Jane: Right. And that insight let them build two efficient variants — MI5 and MI6 — that only fine-tune five or six of the eight cross-attention layers. That cuts trainable parameters by eighty-six to eighty-eight percent while keeping accuracy almost as high.

Tom: And on real-world flood photos, those efficient models actually beat the published STURM-FloodDepth baseline — ninety-eight point six percent accuracy versus eighty-six point six percent. That's a twelve-point improvement.

Jane: So the improvement isn't just efficiency for its own sake. It's efficiency that also generalizes better to real-world conditions. That's the story we should dig into next when we look at the first page of the paper.

First Page: Tom: Alright, let's actually look at the opening of the paper. The abstract lays out the whole arc — the dataset, the three models, the interpretability findings, and the real-world validation.

Jane: And one thing that jumped out at me on that first page is the motivation. They cite research showing that just two point two percent of flood-induced road failures caused seventeen point seven percent of the network collapse during Hurricane Harvey. Small disruptions can cascade into massive failures.

Tom: That's a sobering stat. And they also mention the percolation transition — where gradual increases in water depth suddenly cause the road network to fragment. It's not gradual damage; it's a cliff.

Jane: Right, and that's why centimeter-level precision matters. The difference between fifteen and twenty-five centimeters can be the difference between a road that's passable and one that's completely cut off. And for electric vehicles, depths above about fifteen to twenty-five centimeters put the battery pack at risk of thermal runaway.

Tom: And there's another angle on that first page — autonomous vehicles. Self-driving cars currently just exclude flooded roads from their operational domain because they can't estimate water depth. This work could change that.

Jane: Exactly. If an AV can reliably tell that the water is only eight centimeters deep, it might be able to drive through safely instead of rerouting. That's a huge operational improvement.

Tom: The first page also sets up the three generations of flood depth estimation — reference objects, vehicle detection, and now vision-language models. And the paper argues that VLMs are the right path because they generalize without needing a catalog of reference objects.

Jane: And they're the first to fine-tune an open-source VLM for continuous depth regression. That's the claim, and the results back it up — zero point four zero centimeter MAE on synthetic data.

Tom: So when we look at the implications, what excites you most?

Jane: Honestly, it's the interpretability angle. They didn't just build a model that works — they built a model they understand. And that understanding let them train it more efficiently. That's a template for how we should approach fine-tuning large models for specialized tasks.

Tom: So before we wrap up, let's bring in the rest of the team to get their take on the bigger picture.

Conclusion: Tom: So let's pull it all together. We've been discussing "Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision–Language Models for Centimeter-Level Flood Depth Estimation," and I want to get the team's final thoughts.

Jane: I'll start with what I think is the lasting contribution — they showed that you can use interpretability tools to decide which layers to fine-tune, rather than guessing or tuning everything. That's a methodology that could apply far beyond flood detection.

Lu: I completely agree, Jane. And I'd push it further. This restructure-then-encode pattern — where early layers reorganize features and a specific middle layer becomes the encoding leverage point — that's probably not unique to flood depth. I'd bet we see similar patterns in other regression tasks, like estimating vehicle speed from video or predicting road friction.

Meng: From an engineering standpoint, the eighty-eight percent parameter reduction is the headline. Training a fifty-four-million-parameter adapter on a single GPU takes real time and money. Cutting that to six point five million parameters while keeping accuracy within tolerance — that makes deployment practical for smaller cities and agencies that don't have big compute budgets.

Tom: And the real-world validation is what seals it for me. Beating the published baseline by twelve points on actual flood photos, not just synthetic scenes, that's the proof that the approach transfers.

Lalam: If I can add a cultural perspective — this technology could change how cities communicate flood risk to residents. Instead of a map that says "flooded" or "not flooded," imagine a system that tells you "this street has eighteen centimeters of water, your sedan can pass but your EV should avoid it." That's actionable information that respects people's time and safety.

Jane: That's a beautiful way to put it, Lalam. And it ties back to the paper's motivation — the difference between passable and impassable is often just a few centimeters.

Tom: So we've got a paper that builds a massive synthetic dataset, trains a strong baseline, dissects it with interpretability tools, and uses those insights to build efficient variants that actually generalize better to the real world. That's a complete arc.

Lu: And it's reproducible. The model is open-source, the method is described in detail, and the dataset generation pipeline is documented. That's how the field advances.

Meng: My only caution is the test set size — three hundred real-world images is small. I'd want to see validation on thousands of images across more diverse conditions before deploying this in a production system.

Jane: Fair point, and the paper acknowledges that limitation. Night scenes, extreme depths, and scenes without vehicles are all under-represented.

Tom: Well, we've covered a lot today. To summarize — this paper shows that understanding why a model works can make it cheaper to train and better at generalizing. That's a win on every axis.

Jane: And it brings us one step closer to a world where we don't have to guess whether a flooded road is safe. We can know.

Tom: Thanks for joining us, everyone. We'll be back next time with another paper from the arXiv. Until then, stay curious.

Jane: And stay dry.

Nafis Fuad, Xiaodong Qian, Dongxiao Zhu

Wayne State University · Wayne State University · Wayne State University

cs.CV, cs.LG

Submitted: 2026-08-02

Comments: This Paper is accepted in International Conference on Machine Learning and Application (ICMLA) 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 46/100

Key concepts

Mechanistic Interpretability
This is a technique used to open up the internal workings of a large AI model. The researchers used it to find which specific layers within an eleven-billion-parameter model were actually responsible for learning the information needed for flood depth estimation.
Selective Fine-Tuning
Instead of training the entire massive vision-language model, the team identified and trained only a small subset of layers—specifically five or six cross-attention layers. This process reduces trainable parameters by about eighty-eight percent while keeping accuracy high.
Restructure-then-Encode Pattern
This is an observation about how the model learns features. The researchers found that early layers restructure visual features, and a specific middle layer acts as the leverage point where depth information is finally encoded and made readable by the model.

Terminology

Summary

Summary

This paper presents three vision–language models (VLMs) fine-tuned for continuous flood depth estimation from street-level imagery: FloodLlama-Dense, a fully fine-tuned QLoRA baseline; and FloodLlama-MI5 and FloodLlama-MI6, interpretability-guided sparse variants that fine-tune only the top 5 and top 6 causally relevant cross-attention layers identified through mechanistic interpretability analysis, respectively.

Dataset Generation: The authors construct a 2.81M-image synthetic flood corpus using Unreal Engine 5, combining a 1.73M single-vehicle subset (5 cm depth steps) and a 1.08M mixed-vehicle subset (1 cm steps). The corpus spans 7 vehicle types (BUS, Trailer, VAN as Big; SUV, Truck, Mini-van as Medium; Sedan as Small), 4 weather conditions (Sunny, Afternoon, Night, Rainy), and 0–40 cm depths. Approximately 610K images are used for fine-tuning. Sim-to-real fidelity was assessed via pixel-level KS statistics; the mixed-vehicle subset achieves substantially better alignment with real flood imagery than the single-vehicle subset (KS brightness: 0.308 vs. 0.445; KS saturation: 0.167 vs. 0.443).

FloodLlama-Dense Baseline: The model fine-tunes LLaMA 3.2-11B Vision Instruct via QLoRA (4-bit NF4 quantization, LoRA r=16, α=16) inserted into Q/K/V/O projections of all 40 attention modules (32 self-attention + 8 cross-attention), yielding 54.4M trainable parameters (0.49% of the 10.72B base). Training proceeds in two phases: Phase 1 fine-tunes on a balanced 340,416-image subsample of the single-vehicle subset for 3 epochs (peak LR 2×10−4, cosine schedule); Phase 2 adapts to mixed-vehicle scenes via nine sequential 30,000-image chunks at reduced LR 5×10−5. After Phase 2, FloodLlama-Dense achieves MAE = 0.40 cm, RMSE = 1.97 cm, R2 = 0.9512 on a 16,000-image mixed-vehicle test set, with exact-match accuracy of 94.92% and Acc@5cm = 97.59%. Depth-stratified analysis shows accuracy exceeding 96.8% in Low (1–10 cm), Medium (11–20 cm), and High (21–30 cm) categories, with a modest decline to 93.7% in the Very High (31–40 cm) regime.

Mechanistic Interpretability Analysis: Four methods are applied to a held-out subset of 480 synthetic images: linear probing, logit lens, centered kernel alignment (CKA), and cross-attention entropy, anchored by LoRA adapter weight analysis. Key findings:

  • LoRA weight analysis: Five of eight cross-attention layers (L18, 23, 28, 33, 38) absorb 2.5–3 times more adaptation than the all-layer average; the largest weight updates concentrate on query and output projections.

  • Linear probing: The base model exhibits R2 < 0 at every layer. After fine-tuning, probe R2 jumps from −0.355 at L18 to +0.513 at L23 (a 0.772 absolute increase concentrated in a single layer), with probe MAE dropping from 9 cm to 5 cm. Downstream layers (L28, L33, L38) maintain R2 ≈ 0.55, indicating L23 is the depth-encoding leverage.

  • Logit lens: Layers L3–L23 produce incoherent subword fragments; from L28 onward, the fine-tuned model produces numeric tokens at L33 and output-schema tokens at L38, confirming depth-relevant language emerges only after the decodability transition.

  • CKA: The largest representational drift occurs at decoder layers L13–L22 (minimum CKA = 0.395 at L17), not at the depth-encoding layers L23–L38 (CKA 0.68–0.78). This reveals a two-stage adaptation: early-to-mid layers reorganize visual features, late layers convert the reorganized representation into a numerical depth estimate.

  • Cross-attention entropy: The largest entropy reductions occur at L13 (ΔH = −0.30), L18 (−0.25), and L28 (−0.20), with L13 showing the sharpest reduction but the lowest probe R2 — spatial selectivity is necessary but not sufficient for task encoding.

A composite ranking (Table 5) synthesizes the four methods: L23 ranks first (0.768), L18 second (0.679), L28 third (0.615), L33 fourth (0.512), L38 fifth (0.505), L13 sixth (0.457), L8 seventh (0.331), L3 eighth (0.180).

FloodLlama-MI5 and MI6: Based on the composite ranking and empirical search, MI5 tunes cross-attention layers L8, 18, 23, 33, 38 (6.55M trainable parameters, 88% reduction vs. Dense), and MI6 adds L28 (7.86M parameters, 86% reduction). Both use the same QLoRA configuration and two-phase progressive curriculum as Dense. On the synthetic holdout, MI5 achieves MAE = 0.80 cm, RMSE = 2.64 cm, Acc@5cm = 95.1%; MI6 achieves MAE = 0.78 cm, RMSE = 2.51 cm, Acc@5cm = 95.8% — both well within the 5 cm operational tolerance.

Real-World Benchmark: On a 300-image real-world test set from the STURM-FloodDepth benchmark, FloodLlama variants are compared against the STURM baseline using three prompt styles (Simple, Detailed, Chain-of-Thought). FloodLlama-MI6 with CoT achieves 98.62% overall accuracy vs. 86.61% for STURM — a 12.01 percentage point improvement — with 100.00% Level 2 accuracy (vs. 88.27% for STURM). All FloodLlama variants outperform STURM on overall accuracy and nearly every category, with the single exception of MI5-Simple on Level 2. The paper recommends CoT for safety-critical deployment, noting that explicit reasoning prompts act as a calibration mechanism.

Key Contributions: (1) A 2.81M-image synthetic flood corpus with quantitative sim-to-real fidelity assessment; (2) FloodLlama-Dense, the first open-source VLM fine-tuned for continuous flood depth regression; (3) the first mechanistic interpretability analysis of cross-attention layers for a flood depth model, identifying L23 as the depth-encoding leverage; and (4) FloodLlama-MI5 and MI6, achieving 86–88% parameter reduction while outperforming the published STURM-FloodDepth baseline on real-world data.

Limitations: The analysis is conducted on a single VLM architecture (LLaMA 3.2-11B Vision); the 300-image real-world test set under-represents night-time scenes, extreme depths above 30 cm, and scenes lacking vehicle reference objects; depth-stratified errors above 30 cm confirm reliance on wheel-arch cues becomes brittle when those features submerge.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: Replace heuristic layer selection (e.g., tune all layers or tune last k layers) with a mechanistic interpretability-driven selection process. The paper demonstrates that a composite score from linear probing, logit lens, CKA, and cross-attention entropy identifies the minimal set of layers that encode task-relevant information.

What the improved system can do:

  • Achieve 86–88% reduction in trainable parameters (54.4M → 6.55–7.86M) while maintaining 95–96% accuracy within ±5 cm tolerance on flood depth estimation.

  • Generalize to other regression tasks (e.g., vehicle speed, road friction) by identifying the encoding leverage layer without exhaustive combinatorial search.

  • Reduce training cost by 8× on a single GPU (RTX 4090), enabling fine-tuning of 11B-parameter VLMs in resource-constrained settings.

Improvement: Implement the paper's Phase 1 (single-vehicle, coarse 5 cm depth steps) → Phase 2 (mixed-vehicle, fine 1 cm steps) curriculum. This exploits the finding that mixed-vehicle scenes provide cross-referenceable geometric cues (multiple tire heights, waterlines) that enable triangulation of water level.

Improvement: Use the paper's finding that prompt style (Simple vs. Detailed vs. Chain-of-Thought) shifts the decision boundary in depth classification. CoT prompts yield the highest overall accuracy (98.62%) and best Level 2 (deep flood) accuracy (100%), while Simple prompts bias toward moderate-depth predictions.

Improvement: Leverage the paper's 2.81M-image synthetic corpus (Unreal Engine 5) with systematic variation of vehicle type (7), weather (4), and depth (0–40 cm at 1–5 cm steps). The mixed-vehicle subset achieves better sim-to-real alignment (KS brightness: 0.308 vs. 0.445) than single-vehicle.

Improvement: Apply the paper's finding that layers L13–L22 restructure visual representations (low CKA) without contributing to depth encoding (zero probe ΔR2), while L23 is the sharp transition layer where depth first becomes linearly decodable (ΔR2 = +0.772 in one layer).

Improvement: Use the MI6 variant (adds L28 to MI5's L8, L18, L23, L33, L38) to provide a redundant encoding pathway around L23. This improves robustness when L23's representation is degraded by visual occlusion.


The improved system can:

  • Estimate continuous flood depth at centimeter resolution (MAE = 0.40–0.80 cm) from street-level imagery in real time, with 95–98% accuracy within ±5 cm.

  • Operate on a single consumer GPU (24 GB) during fine-tuning, making it deployable for municipal agencies, insurance companies, and autonomous vehicle fleets.

  • Generalize from synthetic to real-world data with 12 percentage point improvement over state-of-the-art baselines.

  • Provide interpretable layer-level diagnostics, enabling engineers to identify exactly where and why the model fails.

  • Adapt to new regression tasks (e.g., road friction, vehicle speed) by re-running the interpretability protocol (cost: 40 forward passes + closed-form ridge probes) rather than exhaustive hyperparameter search.

Abstract

Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estimates at centimeter resolution. This paper presents three vision-language models fine-tuned for continuous flood-depth estimation from street-level imagery: FloodLlama-Dense, a fully fine-tuned QLoRA baseline, and FloodLlama-MI5 and FloodLlama-MI6, interpretability-guided sparse variants that fine-tune only the top five and six causally relevant cross-attention layers identified through mechanistic interpretability analysis, respectively. Training uses an approximately 610,000-image subset of a 2.81-million-image synthetic corpus generated in Unreal Engine 5. The dataset combines single-vehicle subsets with 5 cm depth increments and mixed-vehicle subsets with 1 cm depth increments, spanning seven vehicle types, four weather conditions, and flood depths from 0 to 40 cm. FloodLlama-Dense achieves an MAE of 0.40 cm, an RMSE of 1.97 cm, and an Acc@5cm of 97.59%. Mechanistic interpretability analysis combining linear probing, logit lens, centered kernel alignment (CKA), and cross-attention entropy reveals a two-stage adaptation pattern: layers L13-L22 restructure visual representations, while depth first becomes linearly decodable at layer L23. FloodLlama-MI5 and FloodLlama-MI6 leverage this insight by fine-tuning only five or six of the eight cross-attention layers, achieving an 86-88% reduction in trainable parameters (6.55-7.86 million versus 54.4 million) with minimal accuracy loss. On a real-world benchmark, FloodLlama-MI6 achieves 98.62% accuracy, compared with 86.61% for the published STURM-FloodDepth baseline.

Sources

Related papers