LVSPM
Long Sequence View Synthesis
and Pose Estimation Model
Hao Su · Zexiang Xu · Xiaoshuai Zhang UC San Diego · Sudo AI
Start with what the model produces.
All clips are rendered by LVSPM along cubic-spline interpolated cameras; examples were selected after rendering the 135-scene canonical split.
Long image collections should not need
a separate calibration pipeline.
Appearance and cameras share one long-context backbone.
Local attention resolves views.
A learned state carries the scene.
The model is trained with sequences up to 128 views and evaluated directly at 256 views.
LVSPM stays ahead as the observed scene grows.
at 256 views
| PSNR ↑ | 16 | 64 | 128 | 256 |
|---|---|---|---|---|
| LVSPM | 25.84 | 24.36 | 23.32 | 22.15 |
| AnySplat | 23.19 | 20.54 | 19.45 | 18.56 |
| DepthSplat† | 24.81 | 18.40 | OOM | OOM |
† receives ground-truth camera poses. AnySplat and LVSPM are pose-free.
Canonical-only reduction: exactly 135 tracked scenes per reported setting. Values are recomputed from saved per-scene outputs.
The advantage is not PSNR-only.
SSIM ↑
LPIPS ↓
Scene means over the same 135-scene canonical split. † DepthSplat receives ground-truth poses; missing cells are measured OOM, not omitted runs.
The average is backed by nearly every scene.
Direct scene-ID join between the complete LVSPM spline run and AnySplat’s preserved 64-view outputs; no unmatched or extra scenes enter the counts.
LVSPM first. Baselines on the same held-out target.





Scene 389a… · LVSPM scene PSNR 27.51 · held-out target verified by pixel-matching saved ground truth.
The comparison is not only a best-case gallery.





Scene 444d… · exact median of the 135 scenes by stored LVSPM PSNR · scene PSNR 24.19.
Failure modes remain visible.





Scene 073f… · lowest LVSPM scene PSNR, 19.32. Fine structure and appearance are visibly degraded.
One path, three reconstructions.
All three render the exact spline through the same sorted held-out cameras. AnySplat's path is mapped into its reconstructed Gaussian frame by one Sim(3) fitted from context cameras; DepthSplat receives ground-truth context poses.
The same test in a clean, structured interior.
The camera motion is synchronized by normalized spline position, not by cherry-picked frames. The complete round trip is shown, with method and pose-input labels embedded in the video.
Every baseline, every canonical view count.
| Relative-pose AUC@3 ↑ | 16 | 32 | 64 | 128 | 256 |
|---|---|---|---|---|---|
| LVSPM | 88.4 | 92.3 | 92.0 | 90.0 | 90.1 |
| Depth Anything 3 | 86.2 | 86.2 | 74.6 | 84.5 | 88.8 |
| VGGT | 74.3 | 89.4 | 78.4 | 87.7 | 85.0 |
| FLARE | 43.1 | 59.6 | 80.0 | 73.3 | 74.6 |
| TTT3R | 58.1 | 62.8 | 54.8 | 37.8 | 30.4 |
| Cut3R | 35.9 | 55.4 | 47.9 | 27.2 | 17.3 |
| Fast3R | 17.5 | 35.3 | 24.5 | 17.3 | 9.7 |
Re-aggregated from preserved full output trees; the deck does not substitute manuscript values where stored outputs differ.
Every baseline trajectory, revealed in sequence.
Scene 1881… was selected by LVSPM’s margin to the strongest baseline. Camera centers are similarity-aligned only for visualization; AUC values are the saved scene metrics.
LVSPM is the fastest measured model in this set.
Encoder-time fields from the saved 256-view output summaries; LVSPM uses the canonical paper/release measurement.
Joint RGB supervision and long-context capacity both matter.
Rendering supplies dense geometric evidence without depth or point-map labels.
| Setting | AUC@3 ↑ | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|
| DL3DV · 32 views | ||||
| without synthetic pretraining | 79.9 | 20.08 | .589 | .267 |
| LVSPM | 80.2 | 21.12 | .638 | .222 |
| ASE · 32 views | ||||
| without RGB loss | 50.9 | — | — | — |
| half TTT chunk | 67.9 | 24.98 | .702 | .263 |
| 12 TTT layers | 59.3 | 24.18 | .673 | .288 |
| LVSPM full | 71.9 | 25.01 | .701 | .257 |
Long-sequence, pose-free
view synthesis in one model.
Unified output
camera poses and target RGB
Long context
directly evaluated to 256 views
Measured impact
stronger NVS, pose, and speed
All comparisons in this deck come from canonical saved outputs or documented release measurements.