Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
cs.AI
Submitted: 2026-08-31
Updated: 2026-09-01
Comments: 72pages
Code: https://github.com/EleutherAI/lm-evaluation-harness
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Nemotron-4 340B Technical Report
- Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning
- CS4: Measuring the Creativity of Large Language Models Automatically by Controlling the Number of Story-Writing Constraints
- Program Synthesis with Large Language Models
- Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
- Constitutional AI: Harmlessness from AI Feedback
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
- Humans or LLMs as the Judge? A Study on Judgement Biases
- Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
- Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
- Self-Questioning Language Models
- Evaluating Large Language Models Trained on Code
- SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Learning to Reason for Factuality
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection