Skip to content
← All mechanisms

Teacher-to-reward guidance

Knowing when to listen to the teacher.

Explore ATOD’s annealing and turn-level gates using editable synthetic signals.

Edit this visualization ↗

Move through the schedule

Training progress linearly interpolates the teacher and reward coefficients. The two weights need not sum to one.

Inspect each turn

Normalize supplied disagreement and uncertainty within the trajectory, then combine them with Soft-OR. Constant proxies normalize to 0.5, matching the paper’s small-denominator convention.

Read this as a mechanism study

This implements equations 10–15 and the combined signal on synthetic, editable turn inputs. It does not train an agent or reproduce reported ATOD results. Original submission: June 2026; the linked revision is August 2026.

The equation behind the motion

κ = κstart + (κend − κstart)p
ρ = ρstart + (ρend − ρstart)p
wₖ = 1 − (1 − d̃ₖ)(1 − h̃ₖ)
Aₖ = κ wₖ Δlog pₖ + ρ Aᴿᴸₖ

A mechanism illustration of ATOD equations 10–15 and its combined signal. Original June 2026 submission; August 2026 revision.

Read the primary source ↗