Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
cs.IR, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/karpathy/nanochat
Terminology
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- Large Language Models Reflect the Ideology of their Creators
- Thomson: Continual Learning of Frontier Models for SovereignAI
- OPEN-1B: A Fully Auditable Training Run
- Gemma 4 Technical Report
- Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation
- Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- K2-V2: A 360-Open, Reasoning-Enhanced LLM
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
- DataComp-LM: In search of the next generation of training sets for language models
- R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
- Textbooks Are All You Need II: phi-1.5 technical report
- Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
- A Proposed Conceptual Framework for a Representational Approach to Information Retrieval
- Building a Culture of Reproducibility in Academic Research
- LLM360: Towards Fully Transparent Open-Source LLMs
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG