LLM Conditioning Study Finds Steering Methods Trade Fluency for Effectiveness
On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study
AISummary
Apple researchers systematically tested LLM conditioning methods and found efficient activation steering often degrades fluency. Steering is far less effective on instruction-tuned models than base models, while prompting and full supervised fine-tuning work for concept injection but are weaker at concept removal. Cheap textual metrics correlate highly with costly LLM-as-judge scores.
Source: Apple Machine Learning Research · machinelearning.apple.com