Loading your workspace…
Loading your workspace…
Situation
A run moves from 8 to 64 GPUs. Per-device batch size and gradient accumulation stay unchanged, the loss uses mean reduction, and the learning-rate schedule is still expressed in epochs. Throughput improves, but the loss curve and final quality move.
Keep going