GenPD project page STATUS: manuscript

Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models

A zero-shot evaluation of modern generative video models and simple short-horizon predictors for latency-sensitive teleoperation.

Aws Khalil · Jaerock Kwon
Bio-Inspired Machine Intelligence (BIMI) Lab, University of Michigan – Dearborn

Research Overview

Problem. Vision-based teleoperation suffers from stale visual feedback when communication latency delays the operator's view of the remote scene.

Question. Can current off-the-shelf generative video models provide sufficiently accurate and computationally efficient short-horizon predictions for predictive display?

The benchmark uses a two-stage design: an initial N=5 five-model generator comparison followed by an expanded N=30 validation comparing LTX-2B with Persistence and corrected 9-frame History Farneback. The lowest-error generator in the initial benchmark, LTX-2B, is still outperformed on average by Persistence and History Farneback in the expanded N=30 validation.

Benchmark Design

Original benchmark
N=5 logical clips
Scene
Town01
Source rate
15 FPS
Original protocol
9 observed + 8 future
Original scope
5 generators, 2 resolutions
Expanded validation
N=30 clips
Expanded sampling
3 sequences, 10 clips each
Expanded methods
Persistence, History Farneback, LTX-2B

The expanded validation uses 512x320 clips. Evaluation reports MAD, N=30 SSIM, paired clip-level differences, t-based 95% confidence intervals, and runtime characterization.

Initial Five-Model Benchmark

Model Resolution MAD Sec/Frame Rollout (s) VRAM (GB)
LTX 13B256x16022.343.7329.8646.68
LTX 13B512x32019.573.8530.8146.68
LTX 2B256x16014.233.1224.9525.97
LTX 2B512x32011.553.1725.3825.97
SVD 1.1256x16078.681.3110.453.15
SVD 1.1512x32083.241.5412.363.44
Wan I2V 1.3B256x16026.902.6521.1610.80
Wan I2V 1.3B512x32023.522.9223.3210.80
Wan VACE 1.3B256x16089.422.5420.3011.70
Wan VACE 1.3B512x32032.083.8230.5313.58

Values are rounded from the frozen aggregate CSV. Rollout time is computed as 8 predicted frames times the reported seconds per rollout frame.

Original benchmark aggregate MAD and inference-time tradeoff
Aggregate MAD and inference-time tradeoff across the initial benchmark. LTX-2B achieves the lowest error among the evaluated generators, while all generators exceed the source-frame latency target.

Temporal Behavior

Original benchmark per-step MAD
Per-step MAD across the original benchmark. Different generators exhibit different temporal failure modes; LTX shows comparatively modest drift, while SVD can lose scene structure despite its aggregate MAD behavior.

Expanded Temporal Error

Expanded N=30 per-step MAD with confidence intervals
Per-step MAD with 95% confidence intervals in the expanded validation. History Farneback is lowest at every horizon, Persistence is second, and LTX-2B is highest among the three throughout the rollout.

Expanded N=30 Baseline Validation

Method N MAD ↓ SSIM ↑
Persistence308.600.834
History Farneback (9-frame)306.830.853
LTX-2B3010.190.779

History Farneback achieves the lowest average MAD and highest SSIM, Persistence also outperforms LTX-2B on average, and the initial LTX-2B advantage among generators does not generalize to the expanded baseline comparison.

Original Generator Comparison

Representative original five-model qualitative rollout
One representative rollout from the initial five-model benchmark, illustrating visual differences across model families. The example is qualitative context and does not by itself establish the aggregate ranking.

Expanded Baseline Comparison

Expanded N=30 qualitative baseline comparison
Rows compare GT, Persistence, History Farneback, and LTX-2B. The clip was selected using a deterministic q25 paired-difference rule, so the example is tied to numerical criteria rather than visual cherry-picking.

Paired N=30 Comparison

Differences are LTX-2B minus baseline. For MAD, positive differences favor the baseline. For SSIM, negative differences favor the baseline. Wins report clips where LTX-2B is better under the metric direction.

Comparison Metric Mean Difference 95% CI LTX Wins
LTX-2B vs PersistenceMAD+1.59[+0.40, +2.79]12/30
LTX-2B vs PersistenceSSIM-0.055[-0.081, -0.028]11/30
LTX-2B vs History FarnebackMAD+3.36[+2.58, +4.14]2/30
LTX-2B vs History FarnebackSSIM-0.074[-0.096, -0.052]4/30

Runtime and Deployment Context

Category Metric Value
Source videoSource rate15 FPS
Source videoFrame intervalapproximately 66.7 ms
Prediction horizonEight-frame temporal horizonapproximately 533 ms
History FarnebackCPU OpenCV rolloutapproximately 0.79 s / 8-frame rollout
LTX-2BRTX 6000 Ada rolloutapproximately 23.5 s / 8-frame rollout

These measurements characterize the tested implementations and are not a hardware-normalized efficiency comparison. The eight-frame horizon describes the prediction window, not a universal frame-throughput budget.

Key Takeaways

  • LTX-2B is the lowest-error generator in the initial benchmark.
  • Increasing generative model scale does not improve short-horizon prediction in this comparison.
  • In N=30, simple persistence and history-aware optical-flow baselines outperform LTX-2B on average.
  • None of the evaluated implementations satisfies both predictive-fidelity and latency requirements in this Town01 zero-shot, open-loop setting.

Citation

If you use this work, please cite:

@article{khalil2026towards,
  title={Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models},
  author={Khalil, Aws and Kwon, Jaerock},
  journal={arXiv preprint arXiv:2605.09670},
  year={2026}
}