text_both_refwarmup_0724Last checkpoint of training job 388261, on 5 held-out val_stage2 clips.
Every clip has generated audio as well as video — unmute to judge the full result.
a video with {object};
the model must put the object and its sound back.a video without {object}; the model must take it out.--prepend-ref, stripped again before decoding);
no first frame forces it off (--no-prepend-ref).add direction overshoots badly at
--cfg-scale 5.0: the latents leave the VAE's range and decode to a flat pink wash.
It is monotonic in cfg and reproduces across seeds, while remove survives cfg 5.0
fine. The main matrix below is therefore generated at cfg 1.0; the last section shows
the same clip at cfg 1.0 / 2.5 / 5.0 so the failure mode stays visible.