MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning
cs.LG, cs.AI
Submitted: 2025-10-06
Updated: 2026-08-27
Terminology
Sources
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- Pushing the boundaries of Structure-Based Drug Design through Collaboration with Large Language Models
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design
- Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
- Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language Model
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
- ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- ExLLM: Experience-Enhanced LLM Optimization for Molecular Design and Beyond
- CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
- Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning
- Efficient Evolutionary Search Over Chemical Space with Large Language Models
- Fusing Models with Complementary Expertise
- TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks