Safety

ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

arXiv:2608.00769v1 Announce Type: new Abstract: One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes

DGX agentpaper
safetyarxiv-cs-cv

arXiv:2608.00769v1 Announce Type: new Abstract: One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applied independently to video frames, however, it produces temporal flicker and edit-strength drift. We introduce extbf{ChordVideo}, which extends the same low-energy principle to video time through shared noise, motion-aligned causal aggregation of per-frame Chord fields, and an optional temporally smoothed proximal correction. We derive a warping-error bound that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows. On TGVE/DAVIS with two one-step backbones, ChordVideo reduces warping error by extbf{78%} and flicker by extbf{49%}, improves CLIP frame consistency by extbf{9--10 points}, and increases background PSNR by about extbf{1.5,dB}, while retaining extbf{2 NFE/frame}. Compared with seven multi-step editors, it achieves competitive temporal consistency and source preservation using extbf{10--60imes fewer model steps per clip

Source: arXiv cs.CV | 2026-08-04

Loading related sources…