ClusterFewshot: Improving Few-shot Optimization for LLMs workflow
cs.CL
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/omrirh/clusterfewshot
License: http://creativecommons.org/licenses/by/4.0/
The gist: The performance of large language model (LLM) workflows often depends on selecting a small set of in-context demonstrations to guide model behavior on new tasks.
Terminology
Abstract
The performance of large language model (LLM) workflows often depends on selecting a small set of in-context demonstrations to guide model behavior on new tasks. Recent methods improve this process by augmenting prompts with successful reasoning paths. However, their demonstration selection relies on random sampling or metric-based rankings, overlooking the semantic structure of the task. We propose ClusterFewshot, a strategy that combines semantic structuring with utility-aware scoring to construct representative and effective few-shot demonstration sets. Evaluated within DSPy-based pipelines, ClusterFewshot substantially reduces optimization cost across multiple benchmarks, while consistently improving accuracy relative to prior bootstrap-based methods in both standalone prompt tuning and hybrid prompt-weight optimization.
Sources
- How Many Demonstrations Do You Need for In-context Learning?
- Training Verifiers to Solve Math Word Problems
- Prompt Optimization via Adversarial In-Context Learning
- ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
- Self-Generated In-Context Learning: Leveraging Auto-regressive Language Models as a Demonstration Generator
- Diverse Demonstrations Improve In-context Compositional Generalization
- ReAct: Synergizing Reasoning and Acting in Language Models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- SGLang: Efficient Execution of Structured Language Model Programs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering