BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
Ru Peng, Haokai Xu, Xijun Gu, Tianyu Zhao, Zhiting Fan, Yawen Zeng, Yihong Zhuang, Jinyang Zhang, Kexin Yang, Jian Wu, Hao Chen, Junyang Lin, Dayiheng Liu, Junbo Zhao
cs.CL
Submitted: 2026-08-14
Updated: 2026-08-18
Code: https://github.com/huggingface/datatrove
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program Synthesis with Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
- Large Language Models as Tool Makers
- ToMBench: Benchmarking Theory of Mind in Large Language Models
- The Llama 3 Herd of Models
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Hierarchical Neural Story Generation
- Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Measuring Massive Multitask Language Understanding
- FastText.zip: Compressing text classification models
- LongForm: Effective Instruction Tuning with Reverse Instructions
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models
- Dual-Difficulty Curriculum Learning for Direct Preference Optimization
- Towards General Text Embeddings with Multi-stage Contrastive Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering