Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation
cs.SE, cs.AI, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: When a hint turns a failing generated program into a passing one, does it provide missing information or merely steer the model toward a solution it could already produce? We test these hypotheses on
Terminology
Abstract
When a hint turns a failing generated program into a passing one, does it provide missing information or merely steer the model toward a solution it could already produce? We test these hypotheses on HumanEval+ and MBPP+ using executable evaluation. For Qwen2.5-3B-Instruct, adaptive relevant hints rescue 36 of 79 selected failures; an unrelated hint rescues 19, while eight unhinted samples solve 46 and recover 31 of the 36 relevant-hint rescues. Phi-3.5-mini shows the same pattern: relevant hints rescue 42 of 101 failures, an unrelated hint rescues 17, and unhinted sampling solves 57, including 36 of the 42 relevant-hint rescues. Because the hint conditions use different attempt budgets, these comparisons do not isolate a purely semantic effect. Mechanistic tests on Qwen identify a stable activation direction shared by relevant and unrelated hints. Persistently adding this direction yields 14 rescues and 18 regressions, with no detectable net accuracy gain; learned low-rank interventions have a positive but imprecise estimated effect. Full textual specifications solve 22 of 24 context-defined problems, versus 5-11 for tested virtual-KV prefixes. Post-generation hidden-state probes transfer across benchmarks, with pooled AUROC 0.806 and 0.780, but their top-one selection advantage over token confidence is statistically unresolved. Overall, relevant hints can rescue failures, but most rescued solutions are already reachable through ordinary sampling, and the internal interventions tested here do not establish task-general capability transfer.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
- Program Synthesis with Large Language Models
- The Internal State of an LLM Knows When It's Lying
- Skill Neologisms: Towards Skill-based Continual Learning
- Discovering Latent Knowledge in Language Models Without Supervision
- CodeT: Code Generation with Generated Tests
- Evaluating Large Language Models Trained on Code
- Code Correctness Is Linearly Decodable from LLM Hidden States Before Generation
- KV-Skill: Forging Expertise in the Model's Native Language
- In-Context Learning Creates Task Vectors
- Qwen2.5-Coder Technical Report
- Language Models (Mostly) Know What They Know
- From Noise to Diversity: Random Embedding Injection in LLM Reasoning
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- What Am I Missing? Question-Answering as Hidden State Probing
- Thought Branches: Interpreting LLM Reasoning Requires Resampling
- Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
- Learning to Compress Prompts with Gist Tokens
- Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties