Skip to content

ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

Zhiqiang Lao

cs.CVarXiv:2608.00769

Abstract

One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applied independently to video frames, however, it produces temporal flicker and edit-strength drift. We introduce ChordVideo, which extends the same low-energy principle to video time through shared noise, motion-aligned causal aggregation of per-frame Chord fields, and an optional temporally smoothed proximal correction. We derive a warping-error bound that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows. On TGVE/DAVIS with two one-step backbones, ChordVideo reduces warping error by 78\% and flicker by 49\%, improves CLIP frame consistency by 9--10 points, and increases background PSNR by about 1.5,dB, while retaining 2 NFE/frame. Compared with seven multi-step editors, it achieves competitive temporal consistency and source preservation using 10--60× fewer model steps per clip

Create a lesson