Teacher-to-reward guidance
Knowing when to listen to the teacher.
Explore ATOD’s annealing and turn-level gates using editable synthetic signals.
Move through the schedule
Training progress linearly interpolates the teacher and reward coefficients. The two weights need not sum to one.
Inspect each turn
Normalize supplied disagreement and uncertainty within the trajectory, then combine them with Soft-OR. Constant proxies normalize to 0.5, matching the paper’s small-denominator convention.
Read this as a mechanism study
This implements equations 10–15 and the combined signal on synthetic, editable turn inputs. It does not train an agent or reproduce reported ATOD results. Original submission: June 2026; the linked revision is August 2026.
The equation behind the motion
κ = κstart + (κend − κstart)p ρ = ρstart + (ρend − ρstart)p wₖ = 1 − (1 − d̃ₖ)(1 − h̃ₖ) Aₖ = κ wₖ Δlog pₖ + ρ Aᴿᴸₖ
A mechanism illustration of ATOD equations 10–15 and its combined signal. Original June 2026 submission; August 2026 revision.
Read the primary source ↗