Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- Instruction Retrieval at Inference Time for Small Language Models
- Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
- Evaluating and Benchmarking the System One Model Jev
- Evaluating Decision Models for Text Annotation in Computational Social Science
- JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- Jev in Medicine: A Benchmark Evaluation
- RouteLLM: Learning to Route LLMs with Preference Data
- Qwen3-VL Technical Report
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It
- JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges
- Qwen3 Technical Report
- Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering