Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging

arXiv:2605.01742 · cs.CV · Submitted 2026-05-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging".

Jane: The gist The proposed multi-axis framework achieves more than 10× improvement in throughput along with over 10× reductions in parameter count, FLOPs,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're looking at this paper, "Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging." The main idea here is that they tackle the big problems with Vision Transformers being too slow and too expensive for real industrial use.

Jane: Right. It’s about this one holistic framework that tries to optimize three different things at once: the architecture, how the tokens are handled, and how much precision we use for calculations.

Lu: What's interesting is that they aren't just picking one technique; they are jointly optimizing these three axes to see if you get better overall deployment results while keeping accuracy high on a tough industrial task.

Meng: So, the core claim seems to be that this multi-axis approach gives them over ten times the improvement in throughput and a tenfold reduction in parameter count, GFLOPs, and energy consumption compared to their baseline.

Tom: That's what they claim—a massive jump in efficiency metrics while still maintaining the accuracy needed for industrial work. It matters because it shows how you can squeeze much more performance out of existing models before you have to design a brand new architecture.

Jane: The paper is looking at Vision Transformers specifically because, as the authors say, their high computational cost and memory needs are really limiting their use in resource-constrained industrial settings.

Lu: They start by using something called Neural Architecture Search, AutoFormer, to find these compact backbones. It’s a systematic way to discover smaller transformer designs that fit the task better than just picking one architecture randomly.

Meng: Finding a compact backbone is one step, but then they move on to token compression using something called Token Merging, or ToMe. This suggests they think reducing the sequence length of the tokens themselves is another key lever for cutting computation.

Tom: And it's not just about those two things working separately; they combine them with fp16 mixed-precision inference to accelerate the actual math on deployment hardware. It’s a three-pronged attack on efficiency, architecture, and bit-width simultaneously.

Jane: They use ImageNet 1K as a benchmark initially to figure out which specific combinations of backbone and token compression ratio were best before moving to their real application <ref:2605.01742#pg1>.

Lu: The results they show from that analysis point toward models like AutoF-Tr=fifteen being particularly strong, showing throughput increases up to one hundred sixty-one point one frames per second compared to the baseline's thirteen point one FPS.

Meng: That speedup on the hardware side is what engineers care about most, especially when you look at the energy consumption reduction; they saw a drop from about fifteen thousand three hundred thirty-nine Joules down to one thousand one hundred fifty-one Joules for that top performer.

Tom: So, for someone just listening who isn't deep in the math, it means if you’re deploying an AI model in a factory setting, this approach could translate directly into running a process much faster and using significantly less power on the machines.

Jane: It shifts the focus from just building bigger models to intelligently optimizing how small models are structured and run on real hardware.

Paper summary: Lu: They've done something early in combining architecture-level compression with token-level compression, which is interesting because they suggest that token merging still helps even after you’ve already made the backbone smaller.

Meng: I wonder about the practicality of that token merging step; can you actually implement progressive merging on top of a fixed backbone without adding too much complexity during deployment?

Tom: That’s a good point, Meng. The paper implies it's designed to be applied to existing transformer architectures, which is what makes it attractive for industrial applications where you aren't starting from scratch.

Jane: They are showing that this joint optimization isn't just theoretical; they fine-tuned these selected variants on a real-world three dee X-ray semiconductor defect classification dataset to prove the efficiency gains hold up in practice <ref:2605.01742#pg1>.

Lu: That transfer and fine-tuning step is crucial because it confirms that the efficiency improvements aren't just theoretical numbers on ImageNet 1K, but they actually work for the specific industrial task they care about <ref:2605.01742#pg1>.

Meng: So, what this means practically is that a company designing an inspection system for semiconductor chips could use this framework to make their current ViT models run much more efficiently without losing the accuracy needed to spot those tiny defects.

Tom: Exactly. And looking at the authors, Phat Nguyen and his team are among the first to do this joint optimization across all three dimensions—architecture, token, and bit-width—which is a significant contribution compared to previous work that only looked at one aspect.

Jane: The title of "Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging" really sums up the ambition of this paper. It’s about attacking the ViT problem from every angle at once.

Lu: I think what makes this work particularly creative is how they structure it to address the different bottlenecks: architecture compression handles the backbone, token merging handles sequence cost, and bit-width reduction handles per-operation speed.

Meng: From an engineering standpoint, the result on AutoF-Tr=fifteen showing an eight times throughput increase and a nine times energy reduction is quite compelling data for deployment readiness.

Tom: It really is concrete numbers that sell the idea here; they aren't just talking about efficiency gains in general, they are showing exactly how much faster and more efficient this specific pipeline runs on their test setup.

Jane: So, to wrap up this part, the paper confirms that combining compact backbone design with token compression creates a deployment-ready Vision Transformer for image classification tasks.

Lu: It’s a very practical approach because it addresses different sources of inference cost in a coordinated way that previous single-axis optimizations didn't manage.

Meng: The limitation they mention is that while they optimize these three axes, they are focused on vision transformer deployment, so this method might not be directly applicable to other types of deep learning models.

Tom: That’s the caveat—it’s specialized for ViTs right now—but for anyone working on industrial image inspection with transformers, this paper lays out a clear and effective path forward.

Conclusion: Tom: So we've seen how they tackle three different knobs on Vision Transformers to make them run way faster and use less power for industrial inspection tasks today.

Jane: Yeah, this paper is titled "Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging," and the authors are focused on making these models practical for real manufacturing.

Lu: What they’re doing is combining finding a smaller model design with cutting down on how much data the model processes at each step and also lowering the precision they use to do the math.

Meng: It really makes sense because we're dealing with very specific, resource-heavy problems in semiconductor packaging right now.

Lalam: From a cultural standpoint, this shows that we can make complex AI tools accessible to more different kinds of factories without needing massive supercomputers for every single task.

Tom: Right. So the main point is that they managed to get over ten times better speed and reduced energy use while still hitting the accuracy target for defect detection on those X-ray images.

Jane: They used ImageNet 1K to test different combinations, and then fine-tuned those optimized models on a real dataset of semiconductor defects <ref:2605.01742#pg1>.

Lu: The paper suggests that architecture compression, token merging, and bit-width reduction work together in a way that single optimizations don't manage alone.

Meng: That joint optimization is what makes this different; it’s not just one trick to make things snappy, it's a coordinated strategy for efficiency.

Tom: And the results show that the best configuration they found, AutoF-Tr=fifteen gave them an eight times speed boost and nearly a ten times energy cut on their test hardware.

Jane: So what this means for you is that if you’re working on industrial inspection with AI, this framework gives you a very clear roadmap for how to shrink and speed up your models effectively.

Lu: It opens up the idea that we can move beyond just tweaking one part of the model and start thinking about these three dimensions together.

Meng: The limitation they mention is that this specific method is tailored to Vision Transformers, so applying it directly to, say, a different kind of deep learning model might need some adjustments.

Tom: Exactly. But for ViTs in manufacturing right now? This is a very strong starting point for making them actually useful on the factory floor.

Singapore University of Technology and Design (SUTD) · Agency for Science, Technology, and Research (A*STAR)

cs.CV

Submitted: 2026-05-03

Updated: 2026-10-08

Importance score: 83/100

The gist: The gist The proposed multi-axis framework achieves more than 10× improvement in throughput along with over 10× reductions in parameter count, FLOPs, and energy consumption while maintaining the

Key concepts

AutoFormer
This is a search method used to find small, efficient versions of Vision Transformer backbones. It explores different design choices like depth and attention structure. A key feature is weight-sharing training, which helps create usable models with good initial weights for later fine-tuning.
Token Merging (ToMe)
This technique reduces the amount of information the transformer processes during inference by merging similar tokens. By compressing sequences in later layers, it effectively lowers the computational cost associated with processing long input sequences without drastically harming performance.
Multi-Axis Optimization
The framework optimizes three distinct dimensions simultaneously: model architecture (AutoFormer), token representation (ToMe), and numerical precision (fp16). This holistic approach targets different bottlenecks in Vision Transformer inference, leading to superior overall efficiency compared to optimizing only one aspect.

Terminology

Summary

The gist The proposed multi-axis framework achieves more than 10× improvement in throughput along with over 10× reductions in parameter count, FLOPs, and energy consumption while maintaining the required accuracy on the downstream industrial task

How it works

The proposed holistic framework jointly optimizes three complementary axes for Vision Transformer deployment: architecture, token, and bit-width Specifically, the framework identifies compact backbones via Neural Architecture Search (AutoFormer), reduces information processing via token merging (ToMe), and accelerates per-operation execution via fp16 mixed-precision inference Starting from a DeiT-B/16 baseline, the process begins by replacing the backbone with compact architectures discovered by AutoFormer, specifically adopting AutoFormer-T and AutoFormer-S to reduce parameter count and computational cost relative to the baseline

The framework further reduces inference cost through token compression using Token Merging (ToMe) ToMe progressively merges similar tokens during inference, reducing the effective sequence length in later transformer layers and thereby lowering computation This combination of AutoFormer and ToMe is central to the method, as AutoFormer reduces model complexity at the backbone level while ToMe further reduces the sequence-processing cost within that backbone Subsequently, the model is compressed under fp16 mixed-precision inference setting for practical deployment This design is motivated by the observation that architecture-level, token-level, and bit-width-level compression address different bottlenecks of Vision Transformer inference

Optimization Axes

The study analyzes accuracy-efficiency trade-offs on ImageNet-1K under aggressive compression to select optimal configurations The evaluation considers multiple metrics, including classification accuracy, parameter count, GFLOPs, throughput, and energy consumption Accuracy reflects downstream utility while parameter count and GFLOPs characterize model compactness and nominal computational cost Throughput and energy consumption capture practical deployment behavior on hardware

The analysis of ImageNet-1K revealed a clear trade-off between efficiency and top-1 accuracy, as more aggressive token compression led to a larger accuracy drop on ImageNet1K The selected models are then transferred and fine-tuned on a realworld in-house 3D X-ray semiconductor defect classification dataset for IC chip packaging inspection

Experimental Setup and Results

The experiments use two datasets with complementary roles: ImageNet-1K serves as the benchmark for selecting compact backbones and token compression settings, and the selected variants are then transferred to the in-house 3D X-ray semiconductor defect dataset for deployment evaluation For deployment, the retrained AutoFormer models are exported to ONNX with opset 17 and dynamic batch support, and then converted to TensorRT Benchmarking is performed on an NVIDIA RTX 6000 Ada Generation GPU using TensorRT10

The deployment results on the 3D X-ray semiconductor defect dataset show that all selected compressed variants substantially outperform the baseline in throughput, energy consumption, parameter count, and GFLOPs Specifically, AutoF-Tr=15 provides the strongest deployment result by increasing throughput from 13.1 to 161.1 FPS and reducing energy consumption from 15339.6 J to 1151.1 J This corresponds to a speedup of about 8× in throughput and a reduction in energy consumption by approximately 9×

Conclusion

The study confirms the main hypothesis that combining compact backbone selection with token compression produces a practical deploymentready ViT for image classification The multi-axis optimization of ViT, combining compact backbone design with token compression and bit-width reduction, is an effective and practical approach for resource-efficient, deployment-oriented ViT optimization for 3D X-ray semiconductor defect classification This work is among the earliest works to jointly optimize architecture, token, and bit-width dimensions in Vision Transformers and the first such resource-efficient, deployment-focused study for semiconductor manufacturing

How it works

The framework systematically combines three optimization axes: architecture (via AutoFormer), token merging (ToMe), and bit-width reduction (fp16 mixed-precision inference) This holistic approach addresses different bottlenecks of Vision Transformer inference The results demonstrate that the selected pipeline not only satisfies but exceeds the target of 10× deployment improvement on the downstream task The combination of AutoFormer and ToMe is expected to be complementary because they act on different sources of inference cost This confirms that token compression remains beneficial even after architecture-level compression

How it works

The evaluation protocol involves analyzing the trade-off between accuracy and efficiency on ImageNet-1K to select favorable operating points The selected configurations are then transferred and fine-tuned on the in-house 3D X-ray semiconductor defect dataset to assess whether the efficiency gains can be retained while recovering downstream task performance This process allows for a practical deploymentready ViT by prioritizing variants that provide strong efficiency gains while remaining within a reasonable accuracy range under aggressive compression The overall study shows that multi-axis optimization of ViT combining compact backbone design with token compression and bit-width reduction can be an effective and practical approach for resource-efficient, deployment-oriented ViT optimization for 3D X-ray semiconductor defect classification The results show that the selected pipeline not only satisfies but exceeds the target of 10× deployment improvement on the downstream task This work is among the earliest works to jointly optimize architecture, token, and bit-width dimensions in Vision Transformers and the first such resource-efficient, deployment-focused study for semiconductor manufacturing The remainder of the paper is organized as follows: Section II reviews prior work on compact Vision Transformers, token compression, and model quantization Section III presents the proposed framework Section IV describes the experimental setup and evaluation protocol, and reports the main experiments and 3D X-ray semiconductor defect classification experiments Section V concludes the work The remainder of the paper is organized as follows

How it works

The framework identifies compact backbones via Neural Architecture Search (AutoFormer), which searches transformer design dimensions such as depth, embedding width, and attention structure AutoFormer is particularly attractive because its weight-sharing supernet training allows sampled subnetworks to inherit well-optimized weights, improving the practical usability of the searched backbones and providing them strong initialization weights for downstream transfer learning The selection process prioritizes variants that provide strong efficiency gains while remaining within a reasonable accuracy range under aggressive compression This systematic approach leads to the identification of configurations that are suitable for real-world industrial applications The study confirms the main hypothesis of this work: combining compact backbone selection with token compression produces a practical deploymentready ViT for image classification

How it works

The experimental setup involves evaluating models on both ImageNet-1K and the in-house 3D X-ray semiconductor defect dataset The evaluation protocol involves fine-tuning all selected compressed variants to recover the highest possible downstream classification performance This allows for a comparison of efficiency improvement after adaptation rather than comparing raw

Improvements for AI systems

  1. textbf Holistic Multi-Axis Optimization Framework for ViTs in Semiconductor Inspection (SUTD/A∗STAR): This framework jointly optimizes architecture, token, and bit-width by using Neural Architecture Search (AutoFormer) for compact backbones, token merging (ToMe) to reduce sequence length by combining similar or redundant tokens, and fp16 mixed-precision inference to accelerate execution. This enables a system that achieves a more than 10× improvement in throughput along with over 10× reductions in parameter count, FLOPs, and energy consumption while maintaining accuracy for industrial tasks like IC chip packaging inspection.

  2. textbf Real-World Deployment on Semiconductor Defect Classification: The optimized model can perform IC chip packaging inspection using a real-world in-house 3D X-ray semiconductor defect dataset, demonstrating that the efficiency gains transfer effectively to the target application. This allows for high-speed, low-power automated quality control in manufacturing environments.

  3. textbf Extreme Resource Efficiency Gains: The system can achieve significant hardware savings on deployment, specifically a 24.1× reduction in GFLOPs and a 15.1× smaller model compared to the baseline while delivering over 10x throughput improvement, directly addressing constraints in resource-constrained industrial settings.

Sources

Related papers