A zero-shot evaluation of modern generative video models and simple short-horizon predictors for latency-sensitive teleoperation.
Problem. Vision-based teleoperation suffers from stale visual feedback when communication latency delays the operator's view of the remote scene.
Question. Can current off-the-shelf generative video models provide sufficiently accurate and computationally efficient short-horizon predictions for predictive display?
The benchmark uses a two-stage design: an initial N=5 five-model generator comparison followed by an expanded N=30 validation comparing LTX-2B with Persistence and corrected 9-frame History Farneback. The lowest-error generator in the initial benchmark, LTX-2B, is still outperformed on average by Persistence and History Farneback in the expanded N=30 validation.
The expanded validation uses 512x320 clips. Evaluation reports MAD, N=30 SSIM, paired clip-level differences, t-based 95% confidence intervals, and runtime characterization.
| Model | Resolution | MAD | Sec/Frame | Rollout (s) | VRAM (GB) |
|---|---|---|---|---|---|
| LTX 13B | 256x160 | 22.34 | 3.73 | 29.86 | 46.68 |
| LTX 13B | 512x320 | 19.57 | 3.85 | 30.81 | 46.68 |
| LTX 2B | 256x160 | 14.23 | 3.12 | 24.95 | 25.97 |
| LTX 2B | 512x320 | 11.55 | 3.17 | 25.38 | 25.97 |
| SVD 1.1 | 256x160 | 78.68 | 1.31 | 10.45 | 3.15 |
| SVD 1.1 | 512x320 | 83.24 | 1.54 | 12.36 | 3.44 |
| Wan I2V 1.3B | 256x160 | 26.90 | 2.65 | 21.16 | 10.80 |
| Wan I2V 1.3B | 512x320 | 23.52 | 2.92 | 23.32 | 10.80 |
| Wan VACE 1.3B | 256x160 | 89.42 | 2.54 | 20.30 | 11.70 |
| Wan VACE 1.3B | 512x320 | 32.08 | 3.82 | 30.53 | 13.58 |
Values are rounded from the frozen aggregate CSV. Rollout time is computed as 8 predicted frames times the reported seconds per rollout frame.
| Method | N | MAD ↓ | SSIM ↑ |
|---|---|---|---|
| Persistence | 30 | 8.60 | 0.834 |
| History Farneback (9-frame) | 30 | 6.83 | 0.853 |
| LTX-2B | 30 | 10.19 | 0.779 |
History Farneback achieves the lowest average MAD and highest SSIM, Persistence also outperforms LTX-2B on average, and the initial LTX-2B advantage among generators does not generalize to the expanded baseline comparison.
Differences are LTX-2B minus baseline. For MAD, positive differences favor the baseline. For SSIM, negative differences favor the baseline. Wins report clips where LTX-2B is better under the metric direction.
| Comparison | Metric | Mean Difference | 95% CI | LTX Wins |
|---|---|---|---|---|
| LTX-2B vs Persistence | MAD | +1.59 | [+0.40, +2.79] | 12/30 |
| LTX-2B vs Persistence | SSIM | -0.055 | [-0.081, -0.028] | 11/30 |
| LTX-2B vs History Farneback | MAD | +3.36 | [+2.58, +4.14] | 2/30 |
| LTX-2B vs History Farneback | SSIM | -0.074 | [-0.096, -0.052] | 4/30 |
| Category | Metric | Value |
|---|---|---|
| Source video | Source rate | 15 FPS |
| Source video | Frame interval | approximately 66.7 ms |
| Prediction horizon | Eight-frame temporal horizon | approximately 533 ms |
| History Farneback | CPU OpenCV rollout | approximately 0.79 s / 8-frame rollout |
| LTX-2B | RTX 6000 Ada rollout | approximately 23.5 s / 8-frame rollout |
These measurements characterize the tested implementations and are not a hardware-normalized efficiency comparison. The eight-frame horizon describes the prediction window, not a universal frame-throughput budget.
The repository includes benchmark code, model configs, exact clip and matrix manifests, frozen metric artifacts, baseline implementations, tests, and reproduction documentation.
If you use this work, please cite:
@article{khalil2026towards,
title={Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models},
author={Khalil, Aws and Kwon, Jaerock},
journal={arXiv preprint arXiv:2605.09670},
year={2026}
}