Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute".
Jane: The paper was written by Nikita Kozodoi, Zainab Afolabi and Jack Butler from Amazon Web Services.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: So we are starting with a paper that has quite a mouthful of a title — Test-Time Augmentation for LLMs, with that subtitle about input diversity beating output diversity at matched compute. Jane, I have to say, just reading that title got me curious about what "matched compute" even means here.
Jane: It's a great place to start, Tom. Essentially, the authors from Amazon Web Services are asking a very practical question: if you have a fixed budget for extra compute at inference time, where should you spend it? Do you spend it on generating multiple different answers from the same question, or do you spend it on asking the question in multiple different ways?
Tom: And that's the input versus output distinction, right? Output diversity is what a lot of us know as self-consistency — you give the model the same prompt, sample multiple reasoning paths, and vote. Input diversity, which is the test-time augmentation approach, means you actually change the question itself — rephrase it, perturb it, and then aggregate predictions across those variants.
Jane: Exactly. And the phrase "matched compute" is really the backbone of this paper because it means they're not just comparing raw accuracy. They're comparing accuracy per unit of cost, per unit of compute. That's the metric that actually matters when you're deploying models at scale.
Tom: The authors are Nikita Kozodoi, Zainab Afolabi, and Jack Butler, all at AWS. And they've published this at a COLM workshop on efficient reasoning. That's a great venue for this kind of work because efficiency is the whole game there.
Jane: Right, and this isn't a purely theoretical exercise either. They have a public implementation on GitHub, so practitioners can actually use this. But the deeper point is that they're questioning the default assumption that more reasoning paths is the best way to spend extra inference compute.
Tom: The implications are pretty significant. If input diversity is indeed more efficient, it would mean that many practitioners who are cranking up sampling budgets for self-consistency might be better served by generating paraphrases of their questions instead.
Jane: And the title claims they've found evidence for that — that input diversity actually beats output diversity. I want to know exactly how they tested that, because that's a bold claim in a field where self-consistency is so dominant.
Tom: Well, we're about to get into precisely that. The abstract promises this systematic, matched-compute comparison across six datasets, and the results apparently show semantic rephrasing delivering about one point eight times more accuracy per dollar. That's a number that should make anyone who's been blindly using self-consistency sit up and pay attention.
Jane: It really should. And I'm curious to see whether those gains hold up across different types of tasks, from math reasoning to multilingual knowledge to multimodal question answering. That breadth is what separates a real insight from a one-off trick.
Tom: So let's dig into how they designed this study, because the methodology is where the credibility of that headline number lives or dies.
Summary: Jane: So we've set the stage with the core question: does varying the input convert inference compute into accuracy more efficiently than varying the reasoning path? Tom, how did the authors actually go about testing that?
Tom: They designed a systematic comparison across six benchmarks — MMLU for general knowledge, MMMLU for multilingual knowledge, MMMU for multimodal reasoning, HLE for expert-level questions, Math500 for mathematics, and IMDB for sentiment classification. On each of those, they compared three augmentation strategies against the standard baselines of chain-of-thought prompting and self-consistency.
Jane: And the three strategies are semantic rephrasing, lexical perturbation, and visual transformation. Semantic rephrasing is exactly what it sounds like — you take a question and generate paraphrased versions of it using an LLM. Lexical perturbation means adding typos and character-level noise. And visual transformation applies small rotations, brightness, and contrast changes to images.
Tom: Right. And the key methodological choice is that everything is matched at the same number of augmentations, k. So if self-consistency gets k samples, then semantic TTA also gets k answers, just from k different phrasings. That way, the comparison isolates the contribution of input diversity versus output diversity, rather than just throwing more compute at one side.
Jane: Lu, you've been quiet there — is there something about that setup that strikes you?
Lu: Actually, yes. What really impressed me is that they didn't just look at accuracy. They looked at cost-accuracy frontiers and measured accuracy gain per extra dollar and per extra LLM call. So even though semantic TTA incurs an extra cost for the rephrasing step — you need one additional LLM call to generate the paraphrases — they still found that semantic TTA delivers roughly one point eight times more accuracy per dollar than self-consistency.
Meng: And that's not a marginal result. It's a statistically significant improvement on five of the six tasks. The average gain for semantic TTA over single-call CoT was about 1 point 8 percentage points, while self-consistency only managed about 0 point 9. So the input-side diversity is contributing something that output-side diversity cannot capture.
Tom: That's the key claim of this paper right? That paraphrasing captures a different kind of variance. When you sample multiple reasoning paths from the same question, the model is still constrained by the surface form of that question. But when you rephrase the question itself, you're forcing the model to approach the problem from genuinely different linguistic perspectives.
Jane: And the really interesting nuance, Meng just mentioned it, is that the gap was largest on Math500 and MMMLU. That makes sense — math problems and multilingual questions are exactly the domains where phrasing sensitivity tends to be high.
Lu: I want to add something about the methodology though. They also ran paired statistical tests, both parametric t-tests and a paired bootstrap over 2,400 pooled questions. The 95 percent confidence interval for semantic TTA was plus 0 point 88 to plus 2 point 71 percentage points, so it's comfortably above zero. That gives me confidence that this isn't just noise.
Meng: But the full picture matters too. For a model that's already near its ceiling of performance, the absolute gains are small. The paper is honest about that — the gains of one to two percentage points might not justify a two to six times cost increase in every deployment setting.
Tom: Which brings us nicely to the cost-effectiveness analysis. They actually constructed cost-accuracy frontiers on Math500 and MMLU, and semantic TTA achieved the highest accuracy at every cost level. But I want to dig into those trade-offs more, because that's where the practical guidance really emerges.
Improvements: Tom: So we've covered the headline results, Jane, but this paper goes deeper than just saying "semantic TTA works." It actually gives us practical guidance on when and how to use it. What improvements over existing approaches are they really proposing?
Jane: Right, and one of the most interesting findings is about the number of augmentations. They did ablations with k up to 10, and they found that semantic TTA peaks at around k equals 4 or 5, with diminishing returns after that. Self-consistency, on the other hand, keeps improving all the way up to k equals 10.
Lu: That's a really practical insight. It means that semantic TTA gives you diminishing returns earlier, so you don't need as many calls to reach its peak accuracy. The average optimal k for semantic TTA across datasets was 4 point 33, while self-consistency needed 4 point 67 and lexical TTA needed 5 point 33. So semantic TTA reaches its best accuracy with less compute, which is exactly what you want.
Meng: And then there's the multimodal dimension, which I find particularly interesting. On the MMMU benchmark, they compared text-based augmentation against image-based augmentation. The result was that text-based semantic TTA achieved 68 point 09 percent accuracy, while visual TTA only reached 67 point 59 percent even with more augmentations.
Tom: That's a notable finding. The model is apparently more sensitive to how the question is phrased than to mild image transformations like small rotations or brightness adjustments. But then they found something even more counterintuitive — combining both text and image augmentation actually hurt performance. It dropped to 65 point 08 percent, which is even below the single-call baseline.
Jane: I find that fascinating. You'd think more diversity would be better, but in this case, simultaneous perturbation of both modalities introduced enough inconsistent variation that majority voting got confused. They actually recommend text-based augmentation alone for multimodal tasks.
Lu: And I should add that the visual augmentation they used was deliberately mild — rotations of plus or minus 3 degrees, brightness and contrast shifts of 5 percent. So this conclusion is scoped to gentle transformations. Stronger visual augmentations might behave differently, and the paper acknowledges that.
Meng: But the most practically significant finding, I think, is the model scaling analysis. They ran the same experiments on three different model sizes — Claude Haiku, Sonnet, and Opus. The gains from semantic TTA were 2 point 75 percentage points on Haiku, 2 points on Sonnet, and only 0 point 25 points on Opus.
Tom: That pattern tells a clear story: TTA is most valuable when the baseline accuracy leaves room for improvement. For a strong model that's already near ceiling, there's not much variance left to average out. And for lexical TTA, it actually degraded performance on Opus — the typos hurt a model that was already performing well.
Jane: And this is where the authors make a really important distinction. They're not claiming TTA is a substitute for upgrading to a stronger model. Even with TTA, Haiku doesn't match Sonnet's baseline accuracy. So TTA is positioned as a compute-efficiency tool for the mid-tier regime — when you can't afford the larger model, you can extract more from the one you have.
Lu: This is such a careful and honest analysis. They're not overclaiming. They're saying: if you're stuck with a mid-tier model, here's a way to get more efficiency from your compute budget. But if you can afford the stronger model, that's still the better investment.
Tom: So we have this detailed picture of where TTA helps and where it doesn't. I want to step back now and look at the presentation of the paper itself — the framing and the experimental setup they chose. There are some design decisions there that make this work particularly convincing.
First Page: Jane: Let's look at the first page more closely, because there's a lot packed into the framing. Tom, what stands out to you about how they've set up the problem right from the abstract?
Tom: What strikes me is that they immediately establish the economic framing. They say test-time scaling "improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment." That's a very deployment-oriented way to frame a research question.
Lu: And there's a nice acknowledgment of the lineage. They explicitly position TTA as an extension of self-consistency, adding input-side diversity on top of output-side diversity. That framing helps isolate exactly what they're contributing — they're not proposing a radically new method, they're asking where the compute budget is best spent within an existing framework.
Meng: The abstract also makes a specific quantitative promise: roughly 1 point 8 times more accuracy per dollar compared to self-consistency, and outperforming self-consistency on five of six tasks. That's a sharp, falsifiable claim. And the first page sets up the intuition for why that might be true — LLMs are sensitive to surface form, so aggregating across phrasings reduces variance from any single formulation.
Jane: Right. And they make a point of saying the augmentation techniques they study are "deliberately simple and established" — paraphrasing, character noise, image transforms. That's a smart choice because it isolates the efficiency question itself. They're not claiming a new augmentation method; they're claiming that the input-side regime deserves attention.
Tom: The figure on that first page also shows the TTA pipeline visually — input question and image get transformed into k variants, each processed independently by the LLM, then answers are aggregated via majority voting. It's a clean visualization that makes the method immediately accessible.
Lu: I should mention that the framework includes a formal definition too. The final prediction is the arg max over the sum of indicator functions for each candidate answer across the k predictions. Simple majority voting, with ties broken randomly. Nothing fancy, which is exactly the point.
Meng: And they're honest about the limitations even in the abstract. They note that TTA applies only to tasks where answer equivalence is well-defined — tasks with discrete answers where majority voting makes sense. Open-ended generation like summarization would need different aggregation mechanisms.
Jane: That honesty carries through the whole paper. The conclusion has a whole section of limitations covering the single model family, the multilingual aggregation across fourteen languages, and the potential for majority voting to amplify confidently wrong answers.
Tom: And I think that's actually a strength. This is a paper that's making a practical claim about efficiency, and it's being very explicit about where that claim holds and where it doesn't. That's the kind of work practitioners can actually use.
Lu: One more thing about the first page — the fact that they're publishing this at a workshop on efficient reasoning signals that the community is starting to take input-side methods seriously. This isn't a fringe idea; it's a direct challenge to the dominance of output-side sampling.
Meng: And it's backed by a public implementation on GitHub, which means the results are reproducible. That's the gold standard for this kind of empirical work.
Tom: So we've covered the methods, the results, the ablations, and the framing. I think we should wrap this up by reflecting on what this means for the broader landscape of inference-time scaling.
Conclusion: Jane: So here we are at the end. Let's pull it all together, Tom. What do we actually know now that we didn't know before reading this paper?
Tom: We know that if you're using a mid-tier model and you have a fixed budget for extra inference compute, spending at least some of that budget on rephrasing the input appears to convert compute into accuracy more efficiently than spending it all on additional reasoning samples. The evidence is a consistent gain of about 1 point 8 percentage points on average, statistically significant, and it Pareto-dominates self-consistency on cost-effectiveness.
Lu: And we know the practical parameters. Semantic TTA reaches near-optimal accuracy at around four augmentations, while self-consistency keeps improving up to ten. For multimodal tasks, text-based augmentation is the way to go, and combining text with image augmentation can actually hurt. And the benefit diminishes as the model gets stronger.
Meng: I'd add that this is a particularly useful result because it's so simple to implement. No retraining, no parameter updates, no self-verification loops. You just generate a few paraphrases, run the model on each, and majority vote. That's something any practitioner with an API budget can try tomorrow.
Lalam: I want to zoom out a bit, if I may. This paper is part of a broader shift in how we think about scaling. For a long time, the assumption was that getting better answers meant getting a bigger model. But the test-time compute literature is showing that you can get meaningful gains by being smarter about how you spend inference compute on a fixed model. This paper contributes to that direction by showing that input diversity is a legitimate and cost-effective axis of that scaling, not just an afterthought.
Jane: That's a good way to frame it. And the authors themselves are careful to say that TTA isn't a substitute for a stronger model — it's a tool for the regime where a stronger model is unavailable or too expensive. That's a realistic and honest scope.
Tom: There are still open questions, of course. The paper notes that decoupling the rephrasing model from the answering model might yield further gains. And balancing input-side with output-side diversity rather than using one exclusively is a natural next step.
Lu: There's also the open question of whether these findings transfer to open-weight models. This study used the Claude family, and the authors are explicit that claims about LLMs in general should be read as claims about current mid-tier models.
Meng: And the multilingual dimension deserves more attention. The gains on MMMLU were substantial, but paraphrase quality varies by language, especially for lower-resource languages. That's a real deployment consideration.
Lalam: But as a research direction, I think the message is clear: input-side scaling deserves a seat at the table. The compute you spend on paraphrasing may well be the highest-value compute you spend at inference time.
Jane: Well said. I think we've covered the important ground here. And as always, the implementation is public, so listeners can test these findings on their own workloads.
Tom: That's where we'll leave it for this paper. Thanks for joining us, everyone. We'll be back soon with the next one.
Nikita Kozodoi, Zainab Afolabi, Jack Butler
Amazon Web Services
cs.LG, cs.AI, stat.ML
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: Published at the COLM 2026 Workshop on Efficient Reasoning
Code: https://github.com/aws-samples/sample-genai-reflection-for-bedrock
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Test-Time Augmentation (TTA) for LLMs is a training-free inference-time technique that extends self-consistency by perturbing the input before sampling, aggregating predictions across transformed
Key concepts
- Test-time augmentation (TTA)
- A method where you create multiple variants of the input (e.g., rephrased questions, slightly altered images) and run the model on each, then aggregate the answers via majority voting. It aims to improve accuracy without changing the model, by reducing sensitivity to any single input formulation.
- Self-consistency
- A technique where you give the model the same prompt multiple times, sample different reasoning paths, and take a majority vote over the final answers. It's a common way to spend extra inference compute, but the paper compares it against input diversity to see which is more efficient.
- Matched compute
- Comparing methods under the same compute budget, meaning each method gets the same number of model calls or augmentations (k). This ensures that any accuracy differences are due to the method itself, not just throwing more compute at one side.
- Cost-accuracy frontier
- A curve showing the best accuracy achievable for a given cost (e.g., dollars or LLM calls). The paper uses this to show that semantic TTA achieves higher accuracy at every cost level compared to self-consistency, making it a more efficient use of inference budget.
Terminology
Summary
Test-Time Augmentation (TTA) for LLMs is a training-free inference-time technique that extends self-consistency by perturbing the input before sampling, aggregating predictions across transformed versions of the input via majority voting. The paper systematically compares input-side diversity (TTA) against output-side diversity (self-consistency) at matched compute across six benchmarks: MMLU, MMMLU, MMMU, HLE, Math500, and IMDB Reviews, using Claude 4.5 Haiku as the primary model with k ∈ 2, 4, 6 augmentations.
Three TTA strategies are evaluated: semantic TTA (LLM-generated paraphrases preserving meaning while varying surface form), lexical TTA (character-level perturbations with 5% probability per word, capped at 10 perturbations, requiring no extra LLM call), and visual TTA (mild image transformations: rotation ±3°, contrast and brightness ±5%). The self-consistency baseline samples k reasoning paths from the same input at temperature 0.75 with CoT prompting.
Main results show semantic TTA delivers the largest average accuracy gain of 1.8 percentage points (pp) over single-call CoT prompting, outperforming self-consistency on five of six benchmarks. The gap is largest on Math500 (+1.52 pp) and MMMLU (+1.50 pp). Lexical TTA achieves smaller gains (1.2 pp average), while self-consistency improves over baseline but is consistently outperformed by semantic TTA. Paired t-tests show semantic TTA achieves statistically significant improvements over both CoT (p < 0.01) and self-consistency (p < 0.05), with a paired bootstrap 95% confidence interval of [+0.88, +2.71] pp.
On cost-effectiveness, semantic TTA Pareto-dominates self-consistency, delivering roughly 1.8× more accuracy per dollar despite its higher per-call cost due to the rephrasing call. Lexical TTA traces a frontier close to self-consistency at lower cost, making it a reasonable fallback when augmentation overhead is undesirable. The paper notes gains of 1–2 pp may not justify a 2–6× cost increase in all settings; TTA is most valuable when baseline accuracy is moderate (40–80%).
Ablation studies reveal semantic TTA reaches peak accuracy with fewer augmentations on average (optimal k = 4.33 vs. 4.67 for self-consistency and 5.33 for lexical TTA), with accuracy non-monotonic in k due to stochastic paraphrase generation and tie-breaking dynamics. On Math500, semantic TTA peaks at k = 5 with diminishing returns, while self-consistency continues improving up to k = 10, nearly matching semantic TTA.
For multi-modal tasks (MMMU), text-based semantic TTA achieves the highest accuracy (68.09% at k = 2), outperforming visual TTA (67.59% at k = 6) and self-consistency (67.09%). Combining text and image TTA decreases performance (65.08%, below the 66.08% baseline), as simultaneous perturbation of both modalities accumulates noise. The paper recommends text-based semantic augmentation alone for multi-modal tasks.
Base model scaling experiments on MMMLU across Claude Haiku, Sonnet, and Opus show semantic TTA gains diminish with model size: +2.75 pp for Haiku, +2.00 pp for Sonnet, and only +0.25 pp for Opus. Lexical TTA degrades Opus (87.00% vs. 88.25% baseline), indicating character-level noise can hurt strong models near ceiling. TTA is not a substitute for upgrading to a stronger model—Haiku with TTA (82.50%) does not match Sonnet's single-call accuracy (85.50%)—but is most cost-effective for mid-tier models where a stronger model is unavailable or too expensive.
The paper concludes that for current mid-tier LLMs, varying the input converts inference compute into accuracy more efficiently than varying the reasoning path alone. Limitations include: TTA applies only to tasks with well-defined answer equivalence, evidence is limited to a single proprietary model family, multilingual gains may not transfer uniformly across the fourteen MMMLU languages, the same model is used for rephrasing and answering, majority voting can amplify confidently wrong answers, and semantic TTA adds latency and cost.
Improvements for AI systems
Improvements to AI Systems:
-
Adaptive Inference Router: Implement a dynamic decision module that selects between semantic TTA, lexical TTA, self-consistency, or single-pass CoT based on (a) estimated baseline accuracy (activate TTA only when predicted accuracy is 40–80%), (b) model size (disable lexical TTA for frontier models like Opus to avoid degradation), and (c) task modality (use text-only semantic TTA for multi-modal inputs, never combine with visual perturbations).
-
Cost-Aware Augmentation Scheduler: Build a budget-constrained controller that allocates augmentations per query. For low-stakes or high-volume tasks, use lexical TTA (cheapest, no extra LLM call). For high-stakes tasks with moderate model accuracy, escalate to semantic TTA with k=4 (optimal average) and stop early if confidence threshold is met, reducing cost by up to 33% vs. k=6.
-
Paraphrase Quality Gate: Add a lightweight verification step after LLM-generated paraphrases—check semantic equivalence via embedding cosine similarity (e.g., threshold >0.85) and discard paraphrases that drift too far. This prevents noisy augmentations from amplifying wrong answers, directly addressing the majority-voting failure mode.
-
Tie-Breaking with Confidence Weighting: Replace naive majority voting with confidence-weighted aggregation. Each sampled answer’s vote is weighted by the model’s softmax probability or logit margin. This reduces the impact of confidently wrong paraphrases and improves robustness when k is small (e.g., k=2), where ties are more frequent.
-
Model-Size-Aware TTA Policy: For mid-tier models (e.g., Haiku, Sonnet), automatically enable semantic TTA with k=5 for math and multilingual tasks (largest gains: +1.5 pp). For frontier models (e.g., Opus), disable TTA entirely and rely on single-pass CoT, since gains are negligible (+0.25 pp) and lexical TTA is harmful.
-
Multi-Modal Perturbation Isolation: For vision-language tasks, apply TTA only to the text modality (paraphrase the question) and never to images. The system should explicitly block visual augmentation when text TTA is active, as combined perturbations reduce accuracy by 3 pp below baseline.
-
Early-Exit with Augmentation Count Prediction: Train a small meta-model to predict the optimal k per query (based on input complexity and model confidence). For simple queries, use k=2; for complex ones, escalate to k=5–6. This matches the observed non-monotonic accuracy curve, avoiding wasted compute on easy questions.
-
Fallback to Self-Consistency for High k: When the system detects that semantic TTA’s accuracy plateaus or declines (e.g., after k=5 on Math500), automatically switch to self-consistency with higher k (up to 10) to capture additional gains, combining both strategies sequentially rather than exclusively.
What the Improved AI System Can Do:
-
Achieve up to +1.8 pp accuracy gains over standard CoT at the same compute budget, with statistical significance (p<0.01), by intelligently choosing the best augmentation strategy per query.
-
Reduce inference cost by 30–50% compared to naive TTA by using lexical TTA for easy tasks, early-exit at k=2–4, and skipping TTA for frontier models—while retaining 90% of the accuracy benefit.
-
Avoid performance regressions on strong models (e.g., Opus) and multi-modal tasks by automatically disabling harmful perturbations (lexical noise, image transforms).
-
Handle ambiguous answers better via confidence-weighted voting, reducing the risk of majority-vote amplification of confidently wrong paraphrases.
-
Scale cost-effectively: For a mid-tier model, the system delivers accuracy comparable to a stronger model’s single pass (e.g., Haiku+TTA ≈ Sonnet baseline) at lower total cost, making high-accuracy inference accessible without model upgrades.
Sources
- Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
- Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection
- Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities
- Exploring LLM Reasoning Through Controlled Prompt Variations
- Teaching Large Language Models to Self-Debug
- RoParQ: Paraphrase-Aware Alignment of Large Language Models Towards Robustness to Paraphrased Questions
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
- Frustratingly Easy Test-Time Adaptation of Vision-Language Models
- M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models
- Measuring Massive Multitask Language Understanding
- Mirror-Consistency: Harnessing Inconsistency in Majority Voting
- Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
- Test-time Augmentation for Factual Probing
- Improved Text Classification via Test-Time Augmentation
- ReFT: Reasoning with Reinforced Fine-Tuning
- Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation
- Humanity's Last Exam
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks