OmniShowcase: Omnidirectional Subject-Consistent
Video Generation from Multi-View Images

Junyeong Ahn1*, Geonwoo Kim2*, Kinam Kim1*, Kihwi Kim3, Hoonjin Jung3, Hyojin Jang1, Jaegul Choo1
1KAIST AI   2Inha University   3FLIPTION
Orbit 360°

The video begins with a snowman wearing a wizard hat and a black cloak, hugging a small bird, standing on a floating ice block in a serene, icy landscape with a backdrop of glaciers and snow-capped mountains. The video concludes with a frontal view of the snowman and the bird, highlighting the peaceful and picturesque setting. + [Orbit 360°]

Abstract

Generating videos of a specific subject from a few reference images remains challenging when the camera reveals surfaces that are not visible in the input. Early subject-to-video (S2V) methods condition on a single reference image and must therefore infer the subject's unseen appearance and geometry as the viewpoint changes. Recent methods extend S2V to multiple reference views, providing more information about the subject from different viewpoints, but they do not explicitly specify the 3D structure that connects these observations. In addition, existing training data do not guarantee that viewpoint changes preserve the same underlying 3D subject. We present OmniShowcase, built on two components. First, a fully automated data synthesis pipeline produces training videos along diverse camera trajectories, with viewpoint variation driven by underlying subject geometry, even spanning full 360-degree orbits. Second, a geometry-conditioning method provides the video model with a turntable of surface-normal maps reconstructed from the reference views, without requiring temporal alignment with the generated frames. We construct Showcase-4K, a corpus of 3D-asset-based videos covering twelve camera motions across 88 trajectory configurations with reconstruction-derived surface-normal conditioning, and ShowcaseBench, a held-out evaluation set that ships the source meshes and camera parameters used to render each video. Experiments show state-of-the-art subject identity and 3D consistency while maintaining competitive video quality and motion.

Method

Four reference views are reconstructed in 3D and rendered as a surface-normal turntable. References and normals are concatenated with the target span and read in context by a LoRA-tuned image-to-video diffusion transformer.

Automated 3D-consistent data synthesis pipeline
Data synthesis pipeline. Blender renders the target orbit and four reference views from one curated 3D asset. The target is composited into a prompt-driven scene, and the four views are reconstructed and rendered as the normal turntable used as the condition.
Architecture overview
Architecture. The reference images and the normal condition enter the reference pathway of the diffusion transformer, with no frame-wise alignment between the normal turntable and the generated video.

Beyond the turntable

The normal condition is a fixed yaw sweep, but the generated camera is not bound to it. The target follows whichever trajectory the prompt names.

Arc-Up R→C

The video begins with the wooden teddy bear statue standing on a hand-painted blue-and-yellow majolica tile ledge on an Amalfi Coast terrace at bright midday, surrounded by lemon trees heavy with yellow lemons, with pastel yellow and coral houses stacked on the cliffside and a sparkling turquoise sea below, every leaf and tile in crisp sharp focus. + [Arc-Up R→C]

Comparison with baselines

Same references, same prompt and same trajectory for every method.

ShowcaseBench

Arc-Down L→C
Phantom
MAGREF
MV-S2V
3DreamBooth
Ours

The video begins with the wooden teddy bear statue standing small on the wide wooden porch of a snow-covered log cabin at dusk, seen from a distance so the whole porch is visible, with glowing string lights along the railing, a stack of firewood, a lantern by the door, tall snow-laden pine trees, and warm light spilling from a frosted window behind. + [Arc-Down L→C]

NAVI

Orbit 360°
Phantom
MAGREF
MV-S2V
3DreamBooth
Ours

The video starts with a small decorative ice cream cart model with a green base and a red and white striped canopy, set on the painted ledge of a fairground arcade stall. + [Orbit 360°]

Qualitative results

Qualitative comparison with baselines
Each method carries three marks for whether the camera follows the trajectory the prompt names (Traj.), whether the subject keeps its identity (ID), and whether it stays 3D-consistent (3D).

Quantitative results

Bold is best, underline is second best.

Method Identity 3D consistency Text Video quality
DINO↑CLIP↑ MEt3R↓GT-CD↓RotCov↓ ViCLIP↑ Imaging↑Aesthetic↑Dynamic↑
Held-outPhantom-MV0.6210.8710.4440.28430.9060.23672.520.5950.460
MAGREF-MV0.6160.8760.4090.17891.1770.24273.880.6140.760
MV-S2V0.6880.8960.3510.23060.7870.24473.210.5900.960
3DreamBooth0.7070.9070.3250.17170.7650.23870.670.5981.000
OmniShowcase0.7990.9250.1900.08240.6650.25974.290.5741.000
NAVIPhantom-MV0.6900.8820.4620.20141.4580.25571.170.5720.029
MAGREF-MV0.6650.8920.4570.21451.2840.25169.810.5740.343
MV-S2V0.7700.9120.3500.18161.1630.25071.730.5651.000
3DreamBooth0.7410.9120.3860.22721.1790.24470.340.5591.000
OmniShowcase0.8280.9180.2490.09751.0740.26074.580.5511.000

Ablation

Ablation on the geometry condition
Removing the normals, replacing them with a textured render, or thinning the turntable all degrade identity and 3D consistency.
Geometry condition Identity 3D consistency Text Video quality
DINO↑CLIP↑ MEt3R↓GT-CD↓RotCov↓ ViCLIP↑ Imaging↑Aesthetic↑
Held-outNormals (ours)0.7990.9250.1900.08240.6650.25974.290.574
None0.7520.9030.2650.20920.7270.24869.850.547
Textured0.7550.9100.2700.16980.7470.23673.410.564
Sparse0.7520.9120.2630.17770.6980.25071.840.552
NAVINormals (ours)0.8280.9180.2490.09751.0740.26074.580.551
None0.7930.9040.3150.19871.1070.23867.670.532
Textured0.7800.9070.3230.18421.0890.24071.470.544
Sparse0.7720.9090.2880.17691.0900.24067.720.531

BibTeX