Skip to content
All cold checks
Post-trainingExpertEvaluation design8 min

Situation

During a preference-optimization run, training preference accuracy keeps rising. Blind human evaluation and factuality peak early, then fall. Responses also become longer.

Design the smallest evaluation plan that separates reward overoptimization from length bias, label or template overfit, and judge drift. Name the checkpoint-level curves, untouched splits, and ablations you need before changing the objective.

Keep going

A different failure surface.