Foundations of Large Language Models
cs.CL, cs.AI, cs.LG
Submitted: 2025-01-16
Updated: 2026-10-08
Code: https://github.com/NiuTrans/NLPBook
Project page: https://niutrans.github.io/NLPBook
Terminology
Sources
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
- The Falcon Series of Open Language Models
- A General Language Assistant as a Laboratory for Alignment
- Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- Efficient Prompting Methods for Large Language Models: A Survey
- Unleashing the potential of prompt engineering for large language models
- AlpaGasus: Training A Better Alpaca with Fewer Data
- Extending Context Window of Large Language Models via Positional Interpolation
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- PaLM: Scaling Language Modeling with Pathways
- Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
- Scaling Instruction-Finetuned Language Models
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Universal Transformers
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- A Survey on In-context Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering