JavisDiT audio-visual object edit — text_both_refwarmup_0724

Last checkpoint of training job 388261, on 5 held-out val_stage2 clips. Every clip has generated audio as well as video — unmute to judge the full result.

How to read this
Guidance scale matters a lot here. The add direction overshoots badly at --cfg-scale 5.0: the latents leave the VAE's range and decode to a flat pink wash. It is monotonic in cfg and reproduces across seeds, while remove survives cfg 5.0 fine. The main matrix below is therefore generated at cfg 1.0; the last section shows the same clip at cfg 1.0 / 2.5 / 5.0 so the failure mode stays visible.