ECCV 2026 | Accepted Paper

FlexComposer

Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

Songchun Zhang1, Sitong Guo2, Xianghao Kong1, Pengwei Liu2, Yuwei Guo3, Lvmin Zhang4, and Anyi Rao1

1 Hong Kong University of Science and Technology 2 Zhejiang University 3 The Chinese University of Hong Kong 4 Stanford University

01Static and dynamic foregrounds

02Flexible 3D trajectory control

03End-to-end scene harmonization

Paper overview

One framework.
Any foreground.

Generative video compositing should preserve the identity and motion of an inserted asset while giving creators precise control over where it moves. Existing systems typically trade one for the other.

FlexComposer standardizes static images and dynamic footage in a canonical foreground space, transports their latent features along a user trajectory, and synthesizes the final video with coherent occlusion, lighting, and shadows.

How FlexComposer creates a controlled composite video A static image or dynamic clip is converted into a canonical foreground, moved through latent space along a trajectory with occlusion awareness, and composed into a coherent output video. 01 / REPRESENT Canonical foreground IMAGE VIDEO Decouple local motion 02 / TRANSPORT Spatial-aware latent injection VISIBILITY GATE Follow the user trajectory 03 / GENERATE Coherent video composition Relight, occlude, harmonize Static or dynamic asset Controllable global trajectory Background-aware generation
FlexComposer preserves a foreground's identity and intrinsic motion while giving the user direct control over its global path.
Static foreground

Trajectory with depth awareness

Dynamic foreground

Intrinsic motion preserved

Harmonization

Lighting and shadow adaptation

Core design

Control without sacrificing fidelity.

01

Unified canonical foreground

Stabilizes dynamic clips and expands static images into one centered representation, separating intrinsic motion from global displacement.

02

Spatial-aware latent injection

Transports canonical VAE features directly onto user-defined trajectories with a parameter-free mapping and visibility-aware occlusion.

03

Synthetic-to-real curriculum

Combines procedural geometry, cinematic real footage, and generative data to learn control, realism, and open-domain composition in stages.

Method

Canonicalize, transport, compose.

A single conditional generation pipeline replaces brittle reconstruction, lighting estimation, and rendering stages.

The FlexComposer pipeline from input assets through latent injection to final video
Foreground motion is decoupled, mapped into the background latent, and fused as dense spatiotemporal conditioning.

01 / RepresentCenter and encode heterogeneous foreground assets.

02 / TransportProject the trajectory and move features in latent space.

03 / GenerateDiffuse a coherent composite with background context.

Video results

Composition in motion.

Occlusion

Airplane across a steel bridge

Trajectory

Car through a dynamic landscape

Micro-motion

Animated subject in a natural scene

Static composition

One scene, multiple inserted assets.

Background

Balloon composition

Chair composition

Trajectory control

The same background supports different objects and paths.

Background

Teapot / Path A

Teapot / Path B

Motorcycle / Path A

Motorcycle / Path B

More cases

Across subjects, scenes, and trajectories.

Opening highlight A
Opening highlight B
Case 01
Case 02
Case 03
Case 04
Case 05
Case 06
Case 07
Case 08
Case 09
Case 10
Case 11
Case 12

Quantitative results

Better control, consistency, and realism.

2.15EPE on DAVIS

Best trajectory error in the V2V setting.

468.30FVD on dynamic composition

Lower than AnyV2V, VACE, and GenCompositor.

91.20%Subject consistency

Preserves asset identity through motion and placement.

0.974Harmonization CLIP score

Balances photometric and geometric scene consistency.

01

Static foreground compositing

Ours vs. Kling 1.5, Tora, and Wan-Move

Case ATrajectory following and occlusion
FlexComposer
Kling 1.5
Tora
Wan-Move
Case BIdentity and scene consistency
FlexComposer
Kling 1.5
Tora
Wan-Move
Case CSmall-object control in an open scene
FlexComposer
Kling 1.5
Tora
Wan-Move
02

Dynamic foreground compositing

Ours vs. AnyV2V, VACE, and GenCompositor

Case APreserving intrinsic foreground motion
Background input
Foreground input
FlexComposer
AnyV2V
VACE
GenCompositor
Case BMotion fidelity under global displacement
Background input
Foreground input
FlexComposer
AnyV2V
VACE
GenCompositor

Ablation

Every component has a visible job.

Background context, trajectory conditioning, visibility reasoning, and canonical representation each address a distinct failure mode.

01

Component ablation

One controlled case, five model variants

W/o BG context

Background structure becomes less stable.

Text only

Spatial control is weakened.

W/o visibility

Occlusion ordering is less reliable.

W/o canonical

Local and global motion interfere.

Full model

Control and appearance remain coherent.

02

Canonical representation

A second dynamic foreground case

W/o canonical representation

Global motion conflicts with the subject's local dynamics.

Full model

Local motion remains stable along the target trajectory.

Failure analysis

A challenging composition case.

Complex foreground motion and close scene interaction remain difficult. The same input is shown across commercial systems and FlexComposer for direct inspection.

Input video
Background
Pixelverse
Runway
FlexComposer

Paper

FlexComposer

Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

Songchun Zhang1, Sitong Guo2, Xianghao Kong1, Pengwei Liu2, Yuwei Guo3, Lvmin Zhang4, and Anyi Rao1

1 HKUST2 Zhejiang University3 CUHK4 Stanford University

ECCV 2026

BibTeX
@inproceedings{zhang2026flexcomposer,
  title={FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control},
  author={Zhang, Songchun and Guo, Sitong and Kong, Xianghao and Liu, Pengwei and Guo, Yuwei and Zhang, Lvmin and Rao, Anyi},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}