Large-scale factor analysis shows machine intelligence is only partially interpretable
cs.CL, cs.AI, q-bio.NC
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/open-compass/opencompass
Project page: http://skylion007.github.io/OpenWebTextCorpus
Terminology
Sources
- Program Synthesis with Large Language Models
- Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
- On the Opportunities and Risks of Foundation Models
- Revealing the structure of language model capabilities
- Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
- Evaluating Large Language Models Trained on Code
- MedDialog: Two Large-scale Medical Dialogue Datasets
- The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
- CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
- On the Measure of Intelligence
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- No Language Left Behind: Scaling Human-Centered Machine Translation
- WebApp1K: A Practical Code-Generation Benchmark for Web App Development
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
- Improving LLM Leaderboards with Psychometrical Methodology
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering