During a preference-optimization run, training preference accuracy keeps rising. Blind human evaluation and factuality peak early, then fall. Responses also become longer.
Design the smallest evaluation plan that separates reward overoptimization from length bias, label or template overfit, and judge drift. Name the checkpoint-level curves, untouched splits, and ablations you need before changing the objective.