When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models".
Jane: The paper was written by Renjie Liang and Zijian Xu from University of Florida.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the arXiv radio hour, folks. I'm Tom, and I've got Jane here with me. Today we're looking at a paper that's got a title that just rolls off the tongue: "Cheap Probes Predict Expensive Training in three dee-CT Vision–Language Models." Jane, I gotta say, that title is doing a lot of work. It's basically the whole thesis right there.
Jane: It really is, Tom. And I love that because in most papers you have to dig through pages of math to figure out what the point is. Here, they're telling you upfront: we've got a cheap way to test something, and it predicts what an expensive way would tell you. The authors are from the University of Florida, and they're tackling a very practical problem in medical imaging.
Tom: Right, and let's unpack what "three dee-CT vision–language models" even means for our listeners who aren't living in radiology land. These are systems that look at a three dee CT scan, like a chest scan, and then generate a text report or answer questions about it. Think of it like a radiologist's assistant that can read the scan and describe what it sees.
Jane: Exactly. And the expensive part is training these things. You've got a vision encoder that processes the scan, and then a large language model that generates the text. Fine-tuning that language model on each candidate encoder and compression scheme takes about a GPU-day per configuration. That's a lot of compute.
Tom: So the authors are saying, what if we could predict which encoder and compression scheme will work best without doing all that expensive training? And their answer is to use a "probe," which is a lightweight read-out head that just looks at the frozen embeddings, the features the encoder already produces, and tries to predict a clinical attribute.
Jane: And the beautiful part is that these probes train in seconds to minutes. So instead of spending a GPU-day per candidate, you spend a few minutes. And their central claim is that the probe results correlate with the expensive fine-tuning results at a Pearson r of zero point nine five. That's a huge number.
Tom: That's a massive correlation. I mean, in my world, if you see a zero point nine five correlation between a cheap proxy and the real thing, you start to get very excited. But Jane, let's be careful here. The paper itself is very honest about this being preliminary. They've only measured six cells, six encoder-compression combinations, on one dataset.
Jane: Right, and that's the responsible way to do it. They're staking a claim, saying here's a methodology that looks very promising, here's the evidence so far, and here's what we need to do to confirm it. They're not overselling it. And that's refreshing.
Tom: Absolutely. So we've got the title, we've got the authors, we've got the big idea. Next we need to dig into how they actually built this benchmark and what the probing setup looks like. That's where the real meat is.
Jane: And I'm curious to hear what our engineer friend Meng thinks about the practical side of this. But that's for the next segment. Tom, let's keep the momentum going.
Summary and Core Idea: Tom: Welcome back. We're still on "Cheap Probes Predict Expensive Training in three dee-CT Vision–Language Models." Jane, last segment we set the stage. Now let's get into the actual summary of what they did. The paper has three parts: building the benchmark, comparing read-out heads, and then the big correlation result.
Jane: Right. So first, they built a grid of cells. Each cell is one encoder combined with one compression scheme. They used three encoders: CoLiPri, CT-CLIP, and BTBthree dee. And on top of the grid encoders, they applied three compressions: none, which is just native tokens, uniform pooling to a fixed budget, and a learned bottleneck they call "adp."
Tom: And on each of these cells, they have a set of clinical attributes. Fifteen regression attributes across four families: size, density, radiomics, and location. Plus eighteen disease findings from the CT-RATE dataset. The key thing here is that all these labels are image-grounded. They're measured straight from the CT scan using segmentation masks and Hounsfield units, not from any report.
Jane: That's a critical design choice. Because if you derive labels from reports, you're inheriting all the noise and potential shortcuts in that text. By measuring directly from the image, the labels are clean and reproducible. And they have this great phrase for it: the probe target and the VQA target share one identical label.
Tom: And then they have these two validation gates, which they call the methodology contribution. Gate one is scale sanity. They check that a proposed clinical threshold actually produces a balanced split in the data. If a threshold gives you ninety-nine percent one class, that's a degenerate question, so they reject it. Only three clinical thresholds survive: aortic dilation, emphysema, and aortic calcification.
Jane: Gate two is probe-separability. They take a strong linear probe and check if it can separate the attribute's buckets above chance. If a strong probe can't separate them, the attribute is ill-posed. Thirteen of fifteen attributes pass this gate. The two location targets are marginal, so they keep them but flag them as "spatial-position canaries," which is a great name.
Tom: Canaries. Because a merge-based compressor is expected to lose spatial position information first. So those attributes are the early warning system. Then they compare ten different read-out heads on the same frozen tokens. And the finding there is that most heads perform about the same, in a narrow band from zero point seven two to zero point seven five AUROC.
Jane: But one head, TransMIL, collapses to zero point six five six. It's over-parameterized, so it's measuring the information ceiling of the representation rather than what the LLM can actually use. And later, it's the only head that inverts the probe-to-downstream correlation. That's a beautiful diagnostic story.
Tom: So the summary so far: a careful benchmark, a read-out comparison that shows most heads are fine but one is pathological, and then the money figure. The disease-probe AUROC predicts report-generation clinical micro-F1 at r equals zero point nine five, Spearman rho equals zero point eight nine, across six matched cells. And it's robust across read-out choice, except for that one pathological head.
Jane: And they're very explicit that this is an ordinal claim. The probe predicts the ranking of cells, not the exact F1 scale. So if you train only the probe-preferred cells, you're very likely training the cells that full fine-tuning would also prefer.
Tom: I love that honesty. They even disclose where the ranking mismatches. Within an encoder's compression variants, the probe can't reliably break near-ties. For CoLiPri, the probe ranks pool above adp, but the full training ranks adp above pool. But the difference is statistically indistinguishable, so it's a tie, not a real inversion.
Jane: Right. The probe reliably picks the right encoder, but you shouldn't trust it to choose between near-duplicate compression variants. That's a very actionable piece of guidance for anyone who wants to use this methodology.
Tom: So that's the summary. Now, next segment, I want to talk about the improvements this suggests. Not just for this paper, but for the field. What does this unlock?
Improvements and Implications: Tom: We're back on "Cheap Probes Predict Expensive Training in three dee-CT Vision–Language Models." Jane, we've covered the setup and the results. Now let's talk about what this paper actually improves. What does it change for people working in this space?
Jane: The biggest improvement is that it turns a days-long search into a minutes-long screening. Right now, if you're a lab that wants to pick an encoder for a three dee CT vision-language model, you have to fine-tune the LLM on every candidate. That's a GPU-day per cell. If you have ten candidates, that's ten GPU-days. Most groups can't afford that.
Tom: And that's a real bottleneck. It means only well-funded labs can do this kind of exploration. This paper democratizes that process. You cache the embeddings, run a probe for a few minutes, and you get a ranking that correlates at zero point nine five with the expensive result. So you can screen dozens of candidates and only spend the GPU-days on the top two or three.
Jane: And there's a methodological improvement too. The two validation gates, scale-sanity and probe-separability, give the community a reusable protocol. Anyone building a probing benchmark for medical imaging can use these gates to make sure their attributes are well-posed before they ship them. That's a contribution that outlives this specific paper.
Tom: And the probe-audit rubric, those six dimensions, that's another reusable piece. It forces you to be explicit about label validity, input validity, read-out strength, metric and null, construct validity, and robustness. That's a checklist that makes claims falsifiable. And we need more of that in this field.
Jane: I want to bring in Lu here, because I think there's a bigger picture. Lu, what does this mean for the field beyond just saving compute?
Lu: Thanks, Jane. The compute savings are real, but I think the deeper implication is that it changes the search space. When fine-tuning is cheap to evaluate, you can explore much more aggressively. You can try more encoders, more compression schemes, more token budgets. The combinatorial space explodes, and this paper says you can navigate it with a cheap proxy.
Tom: That's a great point. The paper mentions that the combinations grow fast. Encoders times compression schemes times token budgets. With probes, you can sweep that whole space and find the sweet spots that you would never have found if you had to fine-tune every cell.
Meng: And from an engineering standpoint, I want to know about the practical pipeline. The paper says probes train in seconds to minutes. But what does that actually look like in practice? You cache the embeddings once, and then you can run all your probes on that cache?
Jane: That's exactly right, Meng. The embeddings are cached, so the probe training is just a lightweight read-out on those cached features. You don't need to re-run the encoder. And the paper shows that normalization matters, because encoders differ by an order of magnitude in embedding scale. They use ZCA whitening for low-dimensional variants and LayerNorm for high-dimensional ones.
Meng: So the practical workflow is: pick your candidate encoders, run them once to cache embeddings, run your probes, rank the cells, and then fine-tune only the top ones. That's a pipeline I could actually implement.
Tom: And that's the improvement in a nutshell. It's not just a paper result; it's a workflow change. And I think that's what makes this paper exciting. It's not just saying "here's a correlation." It's saying "here's how you should do model selection from now on."
Jane: And it's honest about the limitations. It's preliminary, it's on one dataset, it's on six cells. But the signal is strong enough that it's worth pursuing. And that's exactly what a stake-a-claim marker should do.
Conclusion: Tom: And that brings us to the close of our discussion on "Cheap Probes Predict Expensive Training in three dee-CT Vision–Language Models." Jane, let's wrap this up for our listeners who might have just tuned in.
Jane: Sure, Tom. The paper's central claim is that a cheap probe on frozen encoder embeddings can predict the results of expensive LLM fine-tuning for three dee CT vision-language models. They built a careful benchmark with validation gates, compared ten read-out heads, and found a strong correlation of zero point nine five between probe AUROC and downstream clinical micro-F1.
Tom: And the key caveat is that this is an ordinal claim, not an exact estimate. The probe predicts the ranking, not the magnitude. And it's robust across read-out choice, except for one over-parameterized head that fails. The paper is very honest about being preliminary, with only six cells on one dataset.
Jane: But the implications are significant. If this holds up, model selection for medical imaging becomes a minutes-long screening process instead of a days-long fine-tuning sweep. That could open up this kind of research to groups that can't afford the compute.
Tom: And I think the methodological contributions, the validation gates and the audit rubric, are just as valuable as the correlation result. They give the community a way to build probing benchmarks that are rigorous and falsifiable.
Jane: So we're saying goodbye to this paper, but we're excited to see where it goes. The authors are planning to re-run on the gated benchmark, enlarge the cell grid, add seed error bars, and replicate on a second LLM backbone. That's the next step.
Tom: And we'll be watching for it. Thanks to Lu and Meng for joining us today, and thanks to all our listeners. Next up on the arXiv radio hour, we've got a paper on reconstruction-based tokenizers for three dee medical imaging. That should be a fun one.
Jane: See you then, folks. Keep probing.
University of Florida
cs.CV, cs.AI
Submitted: 2026-07-24
Updated: 2026-09-25
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 76/100
Key concepts
- Three Dee-CT Vision–Language Models
- These are systems that process three-dee CT scans, like chest scans, and generate text reports or answers to questions about the scan. They involve a vision encoder for the image and a language model for generating the text.
- Cheap Probes
- A lightweight read-out head used to test frozen embeddings produced by an encoder. These probes train in seconds to minutes, offering a fast way to predict expensive training outcomes without needing full fine-tuning.
- Validation Gates
- Two checks used in the benchmark: scale sanity, which ensures clinical thresholds create balanced data splits, and probe-separability, which checks if a strong probe can separate attribute buckets above chance.
Terminology
Summary
Summary
This paper proposes and provides preliminary evidence for a methodology to screen 3D CT vision–language model (VLM) encoder and token-compression choices using cheap frozen-token probes, rather than expensive full LLM fine-tuning. The authors state: "Picking the frozen image encoder for a 3D CT vision–language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates... Comparing them the usual way means fine-tuning a large language model (LLM) on each combination, and running the whole sweep this way needs far more compute than most groups can spend. We ask whether a cheap probe on the encoder's cached embeddings can stand in for that comparison."
The paper is organized into three parts. First, the benchmark construction: "a grid of (encoder × compression) cells, 15 image-grounded regression attributes plus 18 disease findings, two question types, two validation gates (scale-sanity and probe-separability), explicit input normalization, and a six-dimension probe-audit rubric." Cells are built on CT-RATE non-contrast chest CT, with encoders including CoLiPri, CT-CLIP, and BTB3D, and compressions including none, uniform pool (to 216 tokens), and adp (a learned pack-then-linear bottleneck). All attributes are image-grounded, measured straight from the CT image (segmentation masks or Hounsfield units), never from a report, so labels carry no report or LLM noise and are exactly reproducible.
The 15 regression attributes span four families: size (5), density (4), radiomics (3), and location (3). Two question types are used: catalog (letter-MCQ, either 3-class tertile bucketing or clinical-threshold split) and value (open numeric generation). The two validation gates are the main methodological contribution. Gate 1 (scale sanity) reject[s] any threshold that produces a near-degenerate split (say, 99%-one-class). Only 3 of the candidate clinical catalogs survive: aortic dilation, emphysema, and aortic calcification.
Gate 2 (probe-separability) requires that "a strong CoLiPri linear probe must separate each attribute's tertile buckets above chance (0.333). If a strong probe cannot separate an attribute at that granularity, the attribute is ill-posed, and we drop or coarsen it. Thirteen of 15 attributes are well-posed (separability 0.58–0.87); two location targets are marginal (≈0.45) and kept as
spatial-position canaries. Input normalization is treated as an explicit variable:
ZCA-whitening (fit once on training features) for the low-dimensional variants, and per-token LayerNorm for the high-dimensional learned-projector variant, with raw features as an ablation. A six-dimension probe-audit rubric is also defined:
(1) label validity; (2) input validity; (3) read-out strength; (4) metric & null; (5) construct validity; (6) robustness."
Second, the paper compares ten read-out heads on identical frozen embeddings: attention pooling (ABMIL, ABMIL-multi), simple pools (mean-linear, top-k-linear at two k), flatten variants (flatten-linear, top-k-flatten), a global-local head (GLocal), and TransMIL.
The authors find that read-out choice matters little among sensible heads. Nine of the ten heads sit in a narrow 0.72–0.75 band, with attention pooling (ABMIL-multi 0.749) and a plain mean-linear pool (0.738) roughly tied at the top.
However, "one head is pathological. TransMIL collapses to 0.656, and (Section 5) it is the only head whose probe→downstream correlation inverts. That is itself diagnostic. An over-parameterized read-out measures the representation's information ceiling rather than the usable information the LLM can exploit."
Third, the central predictive-validity result: We pair the disease probe (ABMIL AUROC) with report-generation clinical micro-F1 over six matched cells (CoLiPri / CT-CLIP / BTB3D under adp, pool, or none).
The relationship is strong and monotone, at Pearson r = 0.95, Spearman ρ = 0.89. A minute-long probe on cached embeddings orders the cells almost exactly as a GPU-day-per-cell report-generation training run would.
The authors emphasize this is an ordinal claim: Predictive validity here is ordinal. The probe's job is to predict the downstream ranking of cells, not to reproduce the F1 scale. Spearman ρ = 0.89 is the number that matters.
The correlation is robust to read-out choice: Every strong head agrees (ρ = 0.66–0.89, r = 0.91–0.95), and only the pathological TransMIL inverts (ρ = −0.30).
The authors disclose a limitation: "the within-encoder ranking of compression variants matches downstream for CT-CLIP and BTB3D (adp > pool/none in both probe and F1) but flips for CoLiPri, where the probe ranks pool > adp while report-gen ranks adp > pool. This is a tie, not a real inversion. The two CoLiPri variants are statistically indistinguishable in both probe AUROC (0.853 vs. 0.846) and F1 (0.489 vs. 0.501), so the cross-encoder signal, the one you would act on, is unaffected. A negative control is also reported:
Probing for shuffled disease labels collapses to chance (AUROC 0.51), which confirms the real signal is not a read-out artifact."
The paper concludes: "We staked out a probing methodology for 3D-CT encoder and compression selection and gave preliminary evidence for its central promise... (i) read-out choice matters little among sensible heads, though one over-parameterized head is pathological, and (ii) the cheap disease-probe AUROC predicts report-generation clinical micro-F1 at Pearson r = 0.95, Spearman ρ = 0.89, robustly across read-outs. Read ordinally, with the probe ranking predicting the downstream ranking, this says encoder and compression choices can be screened with minute-long frozen-token probes instead of a full fine-tuning sweep, and full training paid only for the survivors."
Limitations are explicitly stated: "This is a preliminary marker. The correlation rests on six matched cells from a single dataset (CT-RATE, single-institution, non-contrast chest CT, model-extracted labels), and it was computed on an earlier label revision of this pipeline, not on the exact 15-attribute / 18-finding gated benchmark of Section 3, so we state it at the methodology level. Within-encoder near-ties are not reliably ordered, and the per-family regression-attribute pairing against the image-grounded VQA is not yet reported here. We are re-running the correlation on the gated benchmark, enlarging the cell grid, adding n=3 seed error bars, replicating across a second LLM backbone, and validating the cut-points with clinician review. External validity to other sites, contrast protocols, and body regions is still to be shown."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:
-
Implementation: Insert a lightweight read-out head (ABMIL or mean-linear pool) on the frozen encoder’s cached embeddings. Train it in minutes on image-grounded clinical attributes (disease findings, size/density/radiomics/location regression targets). Use the probe’s AUROC as a gate: only cells scoring above a threshold proceed to full LLM fine-tuning.
-
What it does: Cuts compute by 100–1000× for model selection. Instead of spending 1 GPU-day per candidate cell, the system screens dozens of (encoder × compression) combinations in minutes and fine-tunes only the top 1–2 survivors.
-
Implementation: (a) Scale-sanity gate: reject any clinical threshold that produces a near-degenerate class split (e.g., 99% one class) on the ground-truth distribution; fall back to data-driven tertiles. (b) Probe-separability gate: require a strong linear probe to separate the attribute’s buckets above chance (0.333 for 3-class); if not, drop or coarsen the attribute.
-
What it does: Prevents the system from training on ill-posed, degenerate, or undecodable labels. This eliminates silent failures where the model appears to learn but is actually fitting noise or a constant.
-
Implementation: Automatically report: (1) label validity (leakage-free, image-grounded), (2) input validity (identical preprocessing across cells), (3) read-out strength (documented head choice, avoid over-parameterized heads), (4) metric & null (explicit chance level, negative control with shuffled labels), (5) construct validity (probe capability matches paired downstream metric), (6) robustness (sensitivity to read-out, compression, seeds).
-
What it does: Makes every probing claim falsifiable and reproducible. The system flags its own confidence and discloses where rankings are near-ties (e.g., within-encoder compression variants) versus robust cross-encoder signals.
-
Implementation: Default to ABMIL-multi or mean-linear pool. Explicitly detect and reject heads like TransMIL that collapse toward chance (0.656 vs. 0.72–0.75 band) and invert the probe→downstream correlation (ρ = −0.30).
-
What it does: Avoids the pathological case where a too-powerful probe measures the information ceiling rather than the usable information the LLM can exploit. The system’s selection decisions become robust across read-out choice (ρ = 0.66–0.89 for all strong heads).
-
Implementation: Apply ZCA-whitening (fit on training features) for low-dimensional tokenizers (e.g., BTB3D 18-d) and per-token LayerNorm for high-dimensional learned-projector variants (e.g., CT-CLIP 512-d). Treat raw features as an ablation.
-
What it does: Prevents the probe from silently penalizing low-variance encoders. Example from the paper: CT-CLIP mean read-out decodes 0.60 AUROC raw vs. 0.72 under ZCA, cleanly separating it from BTB3D. Without this, the ranking is distorted by scale artifacts.
-
Implementation: Report Spearman ρ (0.89) as the primary selection signal, not Pearson r (0.95). Treat the probe as a ranking predictor. Add a tie-detection rule: if two cells’ probe AUROCs differ by less than 0.01, declare them a near-tie and do not commit to a ranking.
-
What it does: The system correctly avoids over-claiming precision. It reliably picks the right encoder (cross-encoder signal) but honestly flags when it cannot break near-ties between an encoder’s own compression variants (e.g., CoLiPri pool vs. adp: 0.853 vs. 0.846 AUROC, F1 0.489 vs. 0.501 — statistically indistinguishable).
-
Implementation: Shuffle disease labels, re-run the probe, and require AUROC ≈ 0.51 (chance). If the shuffled probe deviates significantly, flag the read-out as broken or the features as leaking.
-
What it does: Confirms the real signal is not a read-out artifact. This is a cheap, automatic sanity check that catches data leakage, label bugs, or head pathologies before they propagate to expensive fine-tuning decisions.
-
Screen 3D-CT encoder and compression choices in minutes — from a grid of 10+ candidates, it identifies the top 1–2 cells with frozen-token probes, then spends full LLM fine-tuning compute only on those survivors. This makes large-scale model selection feasible for groups with limited GPU budgets.
-
Make trustworthy model-selection decisions with disclosed confidence — it tells you not only which cell to pick but how sure it is, flagging near-ties and refusing to over-claim ordinal precision where the signal is genuinely ambiguous.
-
Avoid silent training failures — by enforcing scale-sanity and probe-separability gates, it never trains on degenerate labels or undecodable attribute buckets, eliminating a class of subtle, expensive-to-discover bugs.
-
Produce reproducible, auditable probing results — every experiment ships with the six-dimension rubric, negative controls, and normalization details, so claims can be independently verified and failures can be traced to a specific dimension (label, input, read-out, metric, construct, or robustness).
-
Transfer the methodology to new domains — the same gates, rubric, and probe-screening pipeline apply to other 3D modalities (MRI, PET), other body regions, and other generative downstream tasks (VQA, segmentation-guided generation), not just chest CT report generation.
Bottom line: The system becomes a compute-efficient, self-auditing model selector that replaces a brute-force fine-tuning sweep with a minute-long probe screen, while explicitly guarding against the three failure modes the paper identifies: degenerate labels, over-parameterized read-outs, and cross-encoder scale artifacts.
Abstract
Picking the frozen image encoder for a 3D CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their tokens, and several token budgets, and the combinations grow fast. Comparing them the usual way means fine-tuning a large language model (LLM) on each combination, and running the whole sweep this way needs far more compute than most groups can spend. We ask whether a cheap probe on the encoder's cached embeddings can stand in for that comparison. We build an image-grounded probing benchmark over (encoder times compression) cells, with clinical attribute families and two validation gates, scale-sanity and probe-separability, that keep each attribute well-scaled and decodable. These gates are the main methodological contribution. On this benchmark we compare a range of read-out heads, and in a preliminary study we pair each probe with its matched downstream task. The early signal is encouraging: the cheap probe orders the candidates in close agreement with expensive fine-tuning, at about r about0.95 on the cells measured so far. We read this as an ordinal claim, a ranking predictor rather than an exact estimate, and we are explicit about where it stays preliminary. If it holds up, encoder and compression choices can be screened in minutes with frozen-token probes, with full training spent only on the finalists.
Sources
- Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
- CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation
- M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
- RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis
- Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
- Comprehensive language-image pre-training for 3D medical image understanding
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models