ECCV 2026

LVSPM

Long Sequence View Synthesis
and Pose Estimation Model

Xi Chen · Yachi Zhang · Linghao Chen · Minghua Liu
Hao Su · Zexiang Xu · Xiaoshuai Zhang
UC San Diego · Sudo AI
LVSPM output · cubic-spline camera path · 64 unposed inputs
RESULT FIRST64 unposed imagesposes + continuous novel views

Start with what the model produces.

Indoor sequence
Large interior
Long camera motion

All clips are rendered by LVSPM along cubic-spline interpolated cameras; examples were selected after rendering the 135-scene canonical split.

Long image collections should not need
a separate calibration pipeline.

one real DL3DV input image
16 → 256 imagesunknown camera poses
LVSPMshared sequence model
LVSPM novel-view prediction
camera poses + RGBat queried target rays
No pose inputNo dense 3D labelsNo per-scene optimization

Appearance and cameras share one long-context backbone.

LVSPM architecture from the paper
01Patch + camera tokens
02LaCT sequence backbone
03Plücker-ray queries
04Pose + RGB heads

Local attention resolves views.
A learned state carries the scene.

123256
View attentionexchange evidence in local view windows
TTT hidden stateupdated over large chunks to carry scene context24 LaCT blocks

The model is trained with sequences up to 128 views and evaluated directly at 256 views.

LVSPM stays ahead as the observed scene grows.

26242220181664128256 viewsPSNR ↑
+3.59 dBover the pose-free baseline
at 256 views
PSNR ↑1664128256
LVSPM25.8424.3623.3222.15
AnySplat23.1920.5419.4518.56
DepthSplat†24.8118.40OOMOOM

† receives ground-truth camera poses. AnySplat and LVSPM are pose-free.

Canonical-only reduction: exactly 135 tracked scenes per reported setting. Values are recomputed from saved per-scene outputs.

The advantage is not PSNR-only.

PSNR

1664128256
LVSPM25.8424.3623.3222.15
AnySplat23.1920.5419.4518.56
DepthSplat†24.8118.40OOMOOM

SSIM

1664128256
LVSPM.815.757.714.661
AnySplat.759.647.597.548
DepthSplat†.849.618OOMOOM

LPIPS

1664128256
LVSPM.150.180.206.244
AnySplat.157.247.293.338
DepthSplat†.133.362OOMOOM
At 256 views+3.59 dB PSNR+.113 SSIM−.094 LPIPSLVSPM vs. AnySplat

Scene means over the same 135-scene canonical split. † DepthSplat receives ground-truth poses; missing cells are measured OOM, not omitted runs.

The average is backed by nearly every scene.

Per-scene PSNR difference between LVSPM and AnySplat on 135 DL3DV scenes
134 / 135higher PSNR
129 / 135higher SSIM
124 / 135lower LPIPS
+3.73 dBmedian PSNR difference

Direct scene-ID join between the complete LVSPM spline run and AnySplat’s preserved 64-view outputs; no unmatched or extra scenes enter the counts.

LVSPM first. Baselines on the same held-out target.

LVSPM prediction on strong DL3DV case
LVSPMpose-free · model output
ground truth
Ground truth
input image
One of 64 inputs
AnySplat prediction
AnySplatpose-free
DepthSplat prediction
DepthSplatground-truth poses

Scene 389a… · LVSPM scene PSNR 27.51 · held-out target verified by pixel-matching saved ground truth.

The comparison is not only a best-case gallery.

LVSPM prediction on median DL3DV case
LVSPMpose-free · model output
ground truth
Ground truth
input image
One of 64 inputs
AnySplat prediction
AnySplatpose-free
DepthSplat prediction
DepthSplatground-truth poses

Scene 444d… · exact median of the 135 scenes by stored LVSPM PSNR · scene PSNR 24.19.

Failure modes remain visible.

LVSPM prediction on hard DL3DV case
LVSPMpose-free · model output
ground truth
Ground truth
input image
One of 64 inputs
AnySplat prediction
AnySplatpose-free
DepthSplat prediction
DepthSplatground-truth poses

Scene 073f… · lowest LVSPM scene PSNR, 19.32. Fine structure and appearance are visibly degraded.

One path, three reconstructions.

64 shared input frames18 sorted held-out camerassame physical spline path

All three render the exact spline through the same sorted held-out cameras. AnySplat's path is mapped into its reconstructed Gaussian frame by one Sim(3) fitted from context cameras; DepthSplat receives ground-truth context poses.

The same test in a clean, structured interior.

64 shared input frames12 sorted held-out camerasfull forward-and-back path

The camera motion is synchronized by normalized spline position, not by cherry-picked frames. The complete round trip is shown, with method and pose-input labels embedded in the video.

Every baseline, every canonical view count.

Relative-pose AUC@3 ↑163264128256
LVSPM88.492.392.090.090.1
Depth Anything 386.286.274.684.588.8
VGGT74.389.478.487.785.0
FLARE43.159.680.073.374.6
TTT3R58.162.854.837.830.4
Cut3R35.955.447.927.217.3
Fast3R17.535.324.517.39.7
53 / 53 scenes for every cell+1.3 points over the next method at 256 views90+ from 32 through 256 views

Re-aggregated from preserved full output trees; the deck does not substitute manuscript values where stored outputs differ.

Every baseline trajectory, revealed in sequence.

Similarity-aligned ground-truth and predicted camera centers for LVSPM and six baselines

Scene 1881… was selected by LVSPM’s margin to the strongest baseline. Camera centers are similarity-aligned only for visualization; AUC values are the saved scene metrics.

LVSPM is the fastest measured model in this set.

LVSPM1.78 s
Cut3R12.41 s
Depth Anything 316.38 s
TTT3R17.25 s
VGGT21.26 s
Fast3R21.81 s
FLARE23.13 s
13.0×faster than the slowest measured baseline58.8 FPSnovel-view rendering after encoding

Encoder-time fields from the saved 256-view output summaries; LVSPM uses the canonical paper/release measurement.

Joint RGB supervision and long-context capacity both matter.

camera predictionnovel-view prediction

Rendering supplies dense geometric evidence without depth or point-map labels.

SettingAUC@3 ↑PSNR ↑SSIM ↑LPIPS ↓
DL3DV · 32 views
without synthetic pretraining79.920.08.589.267
LVSPM80.221.12.638.222
ASE · 32 views
without RGB loss50.9
half TTT chunk67.924.98.702.263
12 TTT layers59.324.18.673.288
LVSPM full71.925.01.701.257
Remaining limitsfine structurereflectionsweak texturemulti-stage training

Long-sequence, pose-free
view synthesis in one model.

01

Unified output
camera poses and target RGB

02

Long context
directly evaluated to 256 views

03

Measured impact
stronger NVS, pose, and speed

All comparisons in this deck come from canonical saved outputs or documented release measurements.