SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking".
Jane: The paper was written by Muhammad Taha Mukhtar, Syed Musa Ali Kazmi, Khola Naseem, Muhammad Ali Chattha, Andreas Dengel et al. from National University of Sciences and Technology and German Research Center for Artificial Intelligence.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the arXiv Review Hour, everyone! I'm Tom, and as always, I'm here with my brilliant co-host, Jane. Jane, we have a paper today that really caught my eye — it's called "SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking."
Jane: Tom, this one is genuinely exciting. And the title alone tells you a lot. We're talking about using satellite imagery and machine learning to map informal settlements — slums — in cities across the developing world. The team is from NUST in Pakistan and DFKI in Germany, and they've built a whole benchmark dataset for this.
Tom: And the name — SLUM-i — I love that. It's a play on "slum" and "semi-supervised learning," but it also sounds like "salami," which is oddly fitting because they're slicing up cities into tiles to analyze them.
Jane: Ha! I hadn't thought of that. But you're right, they do slice things up. They're working with five hundred twelve by five hundred twelve pixel tiles of satellite imagery, and for each tile, they're trying to label every single pixel as either "informal settlement" or "everything else." That's the segmentation task at the heart of this paper.
Tom: And why does that matter? Because in cities like Lahore, Karachi, and Mumbai, millions of people live in informal settlements, but governments often don't have accurate maps of where those settlements actually are. Without maps, you can't deliver services, you can't plan infrastructure, you can't track changes over time.
Jane: Exactly. And the traditional way to get this data is to send surveyors out into the field, which is slow, expensive, and often politically complicated. Satellite imagery is cheap and always available, but the challenge is teaching a computer to recognize a slum from space.
Tom: And that's harder than it sounds, right? Because a slum doesn't look like a forest or a lake. It looks like a dense urban area — which is exactly what a formal neighborhood also looks like.
Jane: That's the crux of it. The paper actually shows side-by-side images of a verified informal settlement in Lahore and a planned formal development, and honestly, to my eye, they look remarkably similar. Dense buildings, narrow streets, similar roof colors. The visual distinction is subtle, which is why this is a genuinely hard computer vision problem.
Tom: So the authors had to do something clever. They built a dataset from scratch for Lahore, using official government records from the Katchi Abadis Directorate — that's the informal settlements directorate — and then had human annotators carefully trace the boundaries of two hundred sixty-six settlements using high-resolution Google Earth imagery.
Jane: And they didn't stop there. They also pulled together companion datasets for Karachi and Mumbai, plus four more cities from earlier research — two in Sudan, one in Nairobi, one in Medellín. So the final benchmark spans seven cities across three continents, roughly nine hundred square kilometers of urban area.
Tom: Seven cities, three continents — that's a serious benchmark. And they didn't just throw the data together. They did a systematic analysis of each city's complexity: how intricate the settlement boundaries are, how much the imagery from different cities diverges, how noisy the annotations are, and how imbalanced the classes are.
Jane: And that complexity analysis is going to be really important for the rest of the conversation, because it turns out those factors predict how well different machine learning methods perform. But before we get into the methods — Tom, what struck me most is the scale of the problem they're addressing.
Tom: Yeah, the UN estimates that over a billion people live in informal settlements, and that number is only going to grow as cities expand. This paper is trying to give researchers and policymakers a reliable tool to see where those settlements are, how big they are, and how they're changing.
Jane: And that's the hook for our next segment — because having the data is one thing, but actually getting a computer to learn from it is another challenge entirely. Stick around, we're just getting started.
Summary: Tom: Welcome back to the arXiv Review Hour. We're talking about "SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking," and Jane, we just covered the dataset. Now let's talk about the actual problem they're trying to solve.
Jane: Right. So here's the situation: they have all these satellite images, but only a tiny fraction of them have human-verified labels. In the real world, getting experts to annotate thousands of images is expensive and slow. So the question is — can you train a good model using only a small amount of labeled data plus a large pile of unlabeled images?
Tom: That's the classic semi-supervised learning setup. And the standard approach is something called pseudo-labeling. You train a model on your labeled data, then you use that model to make predictions on the unlabeled images. Where the model is very confident, you treat those predictions as if they were ground truth and add them to your training set. Then you repeat.
Jane: And that works pretty well for general computer vision tasks. But this paper found that it breaks down for slum detection, and they identified two specific reasons why.
Tom: Let's hear them.
Jane: First, the class imbalance problem. In most of these cities, the slum pixels are a tiny minority — maybe ten or twenty percent of the total. And the standard pseudo-labeling approach uses a fixed confidence threshold, usually ninety-five percent. The model only accepts a pseudo-label if it's at least ninety-five percent sure. But when a class is rare, the model tends to be less confident about it, so those slum pixels get rejected. The model ends up learning to predict "background" almost everywhere, and the slum class gets suppressed entirely.
Tom: So the model gets good at saying "not a slum" and terrible at saying "slum." That's a real problem.
Jane: Exactly. And the second issue is what they call covariate shift. The unlabeled pool — the pile of images the model is supposed to learn from — inevitably contains tiles that are just different from the labeled data. Maybe they're mostly empty land, or dense forest, or a highway interchange. The model tries to generate pseudo-labels for those tiles, but they're out of distribution, so the labels are garbage, and the garbage contaminates the training.
Tom: So what did the authors do about it? This is where it gets interesting.
Jane: They proposed two fixes, and they're both elegant. The first is a filtering step that happens before training even begins. They use a frozen DINOv2 model — that's a self-supervised vision transformer that's been trained on massive amounts of images — to embed every tile into a vector space. Then they compute the average embedding of the labeled tiles, and for each unlabeled tile, they measure how similar it is to that average. They keep the top eighty percent most similar tiles and discard the rest.
Tom: So they're using the labeled data as a reference point to curate the unlabeled pool. If a tile looks nothing like the labeled tiles, it's probably not relevant, so they throw it out.
Jane: Precisely. And the second fix is called Class-Aware Adaptive Thresholding — CAAT for short. Instead of a fixed ninety-five percent threshold, the threshold adapts per class. The model tracks a moving average of its confidence for each class, and if the slum class is consistently getting lower confidence, the threshold for that class gets lowered automatically.
Tom: So the slum pixels don't get rejected just because the model is generally less confident about them.
Jane: Right. The threshold is still capped at ninety-five percent, but it can drop as low as needed for the minority class. It's a dynamic curriculum — early in training, the model accepts lower-confidence slum pseudo-labels, and as it gets better, the threshold rises.
Tom: And the results? We've got numbers from the paper — they tested this across all seven cities, with three different label budgets — ten twenty and thirty percent labeled data — and they repeated everything over five random seeds.
Jane: The gains are substantial. In Medellín, at the ten percent label budget, they got a five point nine percentage point improvement in mIoU over the state-of-the-art UniMatch baseline. In El Daein, three point six points. And in several cities, at the thirty percent budget, their semi-supervised model actually beat the fully supervised model that had access to one hundred percent of the labels.
Tom: That's remarkable. Using less data to beat a model that had all the data. That's the kind of result that makes people sit up and take notice.
Jane: And it tells you something important — that the unlabeled data, when curated properly, carries real signal. But we haven't even talked about how they integrated this into modern transformer-based architectures. That's coming up in the next segment.
Improvements: Tom: We're back with "SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking," and Jane, we've covered the dataset and the two core ideas. Now I want to dig into how these improvements actually play out in practice — and I've got Lu and Meng joining us for this one.
Jane: Welcome, both of you! Lu, you're our AI researcher — what struck you about the way they integrated these components?
Lu: Thanks, Jane. What impressed me is that they didn't just bolt these ideas onto one architecture. They tested their components on two completely different backbones. The first is a classic ResNet-one hundred one with a DeepLabV3+ decoder — that's the workhorse of semantic segmentation. The second is a modern DINOv2 transformer with a DPT decoder — that's the kind of foundation model architecture that's all the rage right now.
Meng: And that matters for practical reasons, right? Because if your fix only works on one specific model, it's not really a fix — it's a hack.
Lu: Exactly, Meng. And the results show the components transfer cleanly. The DINO filter and CAAT both improve performance on the transformer pipeline too. They got the best results in five out of seven cities at the twenty and thirty percent label budgets with the DINOv2 backbone. That tells me these are addressing fundamental problems in the pseudo-labeling process, not just quirks of a particular network.
Meng: I want to push on the engineering side, though. The DINO filter — you're running a frozen DINOv2 model over every unlabeled tile to compute embeddings. That's a real computational cost. Did they talk about that overhead?
Jane: They did address it, actually. They don't do exhaustive pairwise comparisons between every unlabeled tile and every labeled tile. Instead, they compute a single prototype embedding — the average of all labeled tile embeddings — and then measure cosine similarity to that one vector. So it's a single pass over the unlabeled pool, which is cheap relative to training the segmentation model itself.
Meng: That's smart. And the CAAT mechanism — that's just a few moving averages and a comparison operation per pixel. That's essentially free. So the inference-time cost is zero — the model runs exactly the same at deployment.
Tom: And that's the kind of thing our listeners care about — can you actually use this in the real world?
Lu: Yes, and I think the bigger story is what happens when you look at which cities benefit most. The paper does this beautiful complexity analysis — they measure boundary complexity, annotation noise, domain shift between cities. And the cities with the worst annotation quality and the highest domain shift — El Daein, El Geneina, Medellín — those are exactly where the semi-supervised gains are largest.
Meng: So the method is most valuable where the data is messiest. That's actually the real-world scenario. If you had clean, perfectly annotated data, you wouldn't need semi-supervised learning in the first place.
Jane: And that's the key insight. But there's a flip side too — in Lahore, where the annotations came from official government records and are very high quality, the semi-supervised methods barely beat the supervised baseline. The unlabeled data just doesn't add much when the labels are already great.
Tom: So the value proposition depends on the data quality. That's a nuanced finding.
Lu: It is. And it's actually a really useful diagnostic for practitioners. If you're starting a mapping project in a new city, you can run their complexity analysis first — measure boundary displacement, domain shift — and get a sense of whether semi-supervised learning will help you or whether you should just invest in more annotations.
Meng: And that's the kind of practical guidance that papers often skip. They just report numbers on benchmark datasets and leave you to figure out whether it applies to your problem.
Jane: Right. And speaking of practical guidance — the authors released all their code, all their dataset splits, and all their analysis scripts publicly. Anyone can reproduce their experiments or apply the framework to a new city.
Tom: That's huge for the community. But before we wrap up, I want to bring in Lalam — our in-house language model — to talk about the broader implications. Lalam, what does this mean for the world beyond computer vision?
Conclusion: Tom: We're in the final stretch of our discussion on "SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking," and I want to bring in Lalam to help us zoom out.
Lalam: Thank you, Tom. I've been listening to the whole conversation, and what strikes me is that this paper is really about making the invisible visible. Informal settlements are home to over a billion people, yet they're often absent from official maps. That absence has real consequences — it means no formal address, no reliable access to services, no recognition in urban planning.
Jane: And this paper gives us a tool to change that. Not just for one city, but across continents. The benchmark spans Pakistan, India, Sudan, Kenya, and Colombia — and the method works across all of them.
Lalam: Exactly, Jane. And I think the cultural dimension is worth emphasizing. When we map informal settlements, we're not just drawing polygons on satellite images. We're creating a shared reference point for policymakers, researchers, and communities themselves. A map can be a tool for advocacy — it can show that a settlement exists, that it has boundaries, that it deserves services and legal recognition.
Meng: That's a powerful framing. But I want to keep one foot on the ground. The paper has limitations — they acknowledge that two of their test sets are very small. N. Nairobi has only one hundred eleven tiles, Medellín has just thirty-five. The results for those cities should be read with caution.
Lu: And they also note that the fixed eighty percent retention threshold for the DINO filter might not be optimal for every city. An adaptive threshold could squeeze out more performance. But that's future work, not a flaw in the current results.
Tom: So what's the bottom line for our listeners? If you're working on urban mapping, or remote sensing, or semi-supervised learning — what should you take away from this paper?
Jane: I'd say three things. First, the dataset is a gift to the community — seven cities, carefully characterized, with public code and splits. Second, the two proposed components — the DINO filter and CAAT — are simple, modular, and architecture-agnostic. They solve real problems that were holding back pseudo-labeling in imbalanced, noisy settings. And third, the complexity analysis gives you a way to predict when semi-supervised learning will actually help you.
Lalam: And I'd add a fourth: this is a reminder that the most impactful applications of AI are often in domains we don't think about every day. Mapping informal settlements isn't flashy, but it touches the lives of a billion people. That's the kind of work that matters.
Tom: Beautifully said, Lalam. And with that, we're going to say goodbye to "SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking." It's been a fantastic paper — rigorous, practical, and socially meaningful. We'll be back next episode with a new paper, but until then, keep your eyes on the sky and your models well-calibrated.
Jane: Thanks for listening, everyone. See you next time.
Muhammad Taha Mukhtar, Syed Musa Ali Kazmi, Khola Naseem, Muhammad Ali Chattha, Andreas Dengel, Sheraz Ahmed, Muhammad Naseer Bajwa, Muhammad Imran Malik
National University of Sciences and Technology · German Research Center for Artificial Intelligence
cs.CV, cs.AI
Submitted: 2026-08-16
Updated: 2026-08-18
Comments: 10 pages, 8 figures, 5 tables
Code: https://github.com/tahamukhtar20/Slum-i
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 62/100
The gist: the scarcity of annotated data and inherent data quality issues, specifically high spectral ambiguity between formal and informal structures and significant annotation noise.
Key concepts
- Semi-supervised Learning
- This approach trains a machine learning model using a small amount of labeled data and a large amount of unlabeled data. The goal is to use the unlabeled images to improve the model's performance on the task, which is useful when obtaining expert labels is slow or expensive.
- Pseudo-labeling
- This technique involves training a model on labeled data and then using that trained model to predict labels for unlabeled images. Predictions with high confidence are treated as ground truth and added back into the training set, allowing the model to learn from more data.
- Covariate Shift
- This occurs when the unlabeled images used for training are different from the labeled data. For example, if a model is trained on images of one type of environment but then tested on completely different environments, the resulting pseudo-labels will be inaccurate and contaminate the training process.
- Class-Aware Adaptive Thresholding (CAAT)
- This is a dynamic thresholding method that replaces a fixed confidence level for accepting pseudo-labels. Instead of using one fixed number, it adjusts the acceptance threshold differently for each class, lowering it automatically for minority classes like slums when the model's confidence in those specific classes is low.
Terminology
Summary
Summary
The paper introduces SLUM-i, a semi-supervised learning framework and multi-city benchmark dataset for satellite-based semantic segmentation of informal settlements (slums). The authors address two critical challenges in this domain: the scarcity of annotated data and inherent data quality issues, specifically high spectral ambiguity between formal and informal structures and significant annotation noise.
The paper states: "We address this by introducing a benchmark dataset for Lahore, constructed from scratch, along with companion datasets for Karachi and Mumbai, which were derived from verified administrative boundaries, totaling approximately 900 km2 of urban area. This collection is supplemented by four cities from prior literature across Sub-Saharan Africa and Latin America, with comprehensive data quality assessments provided for each city."
The dataset comprises seven cities spanning three continents: Lahore, Karachi, and Mumbai in South Asia; El Daein, El Geneina, and N. Nairobi in Africa; and Medellín in South America. The tile composition varies significantly across cities, with total tiles ranging from 35 (Medellín) to 4,363 (Karachi). The paper notes: "The dataset exhibits a highly heterogeneous tile profile and a stark class imbalance that mirrors real-world urban topographies. A significant majority of the tiles across all domains consist purely of background features."
The authors conduct a systematic complexity analysis across four dimensions: boundary morphology, domain shift, annotation quality, and class imbalance. Key findings include: "Boundary complexity correlates inversely with annotation quality... Cities with more intricate settlement outlines tend to carry smaller label displacement... N. Nairobi achieves the tightest alignment (median displacement 1.0 px; 70.5% of boundary pixels within 2 px), while El Daein exhibits the largest systematic offset (median 7.0 px; only 23.6% within 2 px). The pairwise Jensen-Shannon divergence analysis reveals that
N. Nairobi is the single largest source of distributional shift, lying at least JS=0.28 from every other city, while the Medellín–N. Nairobi pair reaches JS=0.65."
The proposed method extends the UniMatch pipeline with two modular components: (1) a DINOv2-based unlabeled pool filter that removes out-of-distribution tiles prior to training, and (2) a Class-Aware Adaptive Threshold (CAAT) mechanism that dynamically adjusts confidence thresholds per class via an Exponential Moving Average to prevent minority class suppression. The paper explains: Standard UniMatch relies on a fixed global threshold (τ = 0.95). This approach disproportionately suppresses the minority class when the class imbalance is significant.
The DINO filter computes a prototype embedding from labeled data and retains only the top 80% of unlabeled tiles by cosine similarity to this centroid. The CAAT mechanism maintains two EMA quantities—a global mean-confidence EMA and a per-class mean-softmax EMA—to compute adaptive per-pixel thresholds: This scales the global threshold down for underrepresented classes (low φc) and up for dominant ones, with an upper bound of 0.95.
Experiments are conducted under three label scarcity protocols (10%, 20%, 30% labeled data) with nested sampling, using both ResNet-101/DeepLabV3+ and DINOv2-Small/DPT backbones. All experiments are repeated over five random seeds. The paper reports: Our method shows the most substantial gains in El Daein, El Geneina, and Medellín... Under the strict 10% label budget, our approach improves over the UniMatch baseline by +3.6 pp in El Daein and +5.9 pp in Medellín.
At the 30% label budget, the method achieves mIoU of 0.896 in Medellín and 0.747 in El Daein, exceeding their corresponding fully supervised results (0.881 and 0.726). The paper notes: This suggests that the curated semi-supervised pipeline can, in some settings, compensate for annotation noise more effectively than standard fully supervised training.
For South Asian cities, results are more nuanced. In Karachi, gains of +1.7 pp and +3.7 pp are observed at 20% and 30% budgets respectively. Mumbai shows variable performance, while Lahore is the only city where supervised training consistently matches or outperforms all semi-supervised methods, attributed to the high fidelity of Lahore's official Katchi Abadis registry annotations.
Ablation studies isolate the contribution of each component. The DINO filter alone produces gains in Medellín (+5.1 pp at 10%) and Karachi (+3.1 pp at 30%), while CAAT alone delivers gains in El Geneina (+3.6 pp at 10%) and Karachi (+3.2 pp at 30%). The combined method shows complementarity: This interaction is particularly visible in Karachi at the 20% budget, where neither component alone improves over UniMatch (F= −1.7 pp; C= −0.4 pp), yet their combination yields a positive gain (F+C= +1.7 pp).
The paper concludes: "Cities with high boundary displacement and substantial domain shift appear to benefit most from unlabeled data curation, whereas cities with high-fidelity official annotations may gain less from semi-supervised augmentation under the current label-budget regime." Limitations include the fixed k=80% retention threshold, small test sets for N. Nairobi (111 tiles) and Medellín (35 tiles), restriction to binary segmentation, and independent per-city training rather than joint multi-city training.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:
1. Add a DINOv2-based unlabeled pool filter (pre-training data curation)
-
Implementation: Before training a semi-supervised segmentation model, extract CLS-token embeddings from a frozen DINOv2-Small backbone for all labeled and unlabeled tiles. Compute the mean embedding of the labeled set as a centroid. Score each unlabeled tile by cosine similarity to that centroid. Retain only the top 80% of unlabeled tiles; discard the bottom 20% as out-of-distribution noise.
-
Resulting capability: The model trains only on unlabeled data that is visually consistent with the labeled domain. This reduces covariate shift, prevents wasted capacity on irrelevant tiles (e.g., dense vegetation, highways), and improves mIoU by up to +5.1 pp in high-shift cities (e.g., Medellín) without any inference overhead.
2. Add a Class-Aware Adaptive Threshold (CAAT) for pseudo-label gating
-
Implementation: Replace the fixed confidence threshold (τ = 0.95) in the unsupervised loss with a per-pixel adaptive threshold. Maintain two EMA quantities: a global mean-confidence (p̃) and a per-class mean-softmax (μ). Compute a class modulator φ c = μ c / max(μ). For each pixel with predicted class ŷ, set the threshold τ = min(p̃ · φ ŷ, 0.95). Gate the unsupervised loss using a binary mask M = 1[max(p) ≥ τ].
-
Resulting capability: The model dynamically lowers the acceptance threshold for the minority slum class during early training, preventing systematic suppression of low-confidence but valid slum pixels. This yields gains of up to +3.6 pp mIoU in class-imbalanced cities (e.g., El Geneina) and reduces false-negative regions in qualitative outputs.
3. Combine both components into a unified semi-supervised pipeline
-
Implementation: Integrate the DINO filter (step 1) and CAAT (step 2) into an existing SSL framework (e.g., UniMatch or UniMatch-v2). Use the filtered unlabeled pool for weak-to-strong consistency training, with CAAT gating both strong-augmentation and feature-perturbation losses.
-
Resulting capability: The combined system achieves the best overall performance in 5 of 7 cities across label budgets (10%, 20%, 30%) and both backbone families (ResNet-101/DeepLabV3+ and DINOv2-Small/DPT). At the 30% label budget, it matches or exceeds the fully supervised ceiling in 4 of 7 cities (El Daein, Medellín, Mumbai, Karachi) while using only 30% of the annotations.
4. Add a dataset complexity diagnostic for city-level SSL deployment
-
Implementation: Before applying SSL to a new city, compute four metrics: (a) boundary complexity (fractal dimension ratio of settlement outlines), (b) feature contrast (separability between slum and background pixels), (c) annotation quality (median label displacement in pixels), and (d) class imbalance (proportion of slum vs. background tiles). Use these to predict whether SSL will provide gains.
-
Resulting capability: The system can automatically flag cities where SSL is likely to be beneficial (high boundary displacement, high domain shift) versus cities where supervised training may suffice (high-fidelity annotations, well-matched distributions). This prevents wasted compute and avoids performance regressions (e.g., Lahore, where SSL underperforms supervised training).
-
Segment informal settlements from satellite imagery with higher accuracy than current state-of-the-art SSL baselines, particularly under severe class imbalance and label scarcity (10% labels).
-
Automatically curate unlabeled satellite data to remove out-of-distribution tiles, ensuring that semi-supervised training focuses only on relevant urban morphology.
-
Adapt its confidence thresholds per class in real time, preventing minority-class suppression and producing spatially coherent predictions with fewer false negatives.
-
Transfer across geographic regions (South Asia, Sub-Saharan Africa, Latin America) without retuning, because the DINO filter and CAAT are architecture-agnostic and require no inference-time changes.
-
Provide a diagnostic tool for urban planners and remote sensing teams to decide whether semi-supervised learning will outperform supervised training for a given city, based on quantitative complexity metrics rather than trial-and-error.
These improvements are directly implementable from the paper’s methodology (Section 2.2) and validated by the ablation study (Figure 5) and quantitative results (Table 3). No additional data collection or architectural changes are required beyond the two modular components described.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models