# Evaluation

This document summarizes the benchmark protocol and metrics implemented in the release code.

## Rollout Protocol

All paper evaluations use:

- conditioning length `L = 9`
- prediction horizon `K = 8`
- total clip length `17`
- source frame rate `15 FPS`

At 15 FPS, each frame interval is approximately 66.7 ms and the 8-frame horizon spans approximately 533 ms. This horizon describes the prediction window; it is not a universal real-time throughput budget.

## Decoded-MP4 Evaluation Domain

Metrics are computed on decoded uint8 RGB MP4 frames. This matches the historical paper metric domain and avoids mixing raw source PNG and decoded-video domains.

Generated model videos are aligned as:

- frames 1-9: conditioning prefix
- frames 10-17: predicted rollout

Ground-truth future videos contain the 8 future frames. Evaluation compares generated frames 10-17 against GT frames 1-8.

## MAD

Mean absolute difference is computed over RGB channels and pixels.

- Per-step MAD: MAD for each prediction step `t+1` through `t+8`.
- Rollout MAD: average absolute RGB error across all 8 future frames.

Original N=5 and expanded N=30 results use decoded-MP4 MAD.

## RGB SSIM

RGB SSIM is reported for the expanded N=30 validation. The implementation uses `skimage.metrics.structural_similarity` with RGB channel handling and `data_range=255`.

## Persistence Baseline

Persistence repeats the final conditioning frame for all 8 future prediction steps. It uses no future ground truth.

## Corrected 9-Frame History Farneback

The corrected History Farneback baseline uses exactly the 9 conditioning frames:

1. Convert the 9 conditioning frames to grayscale.
2. Estimate 8 dense forward Farneback flow fields between consecutive conditioning frames.
3. For each historical flow, advect source-frame pixel coordinates forward through the observed trajectory toward the final conditioning frame.
4. Bilinearly sample subsequent flow fields during advection.
5. Bilinearly splat the transported historical velocity into final-frame coordinates.
6. Aggregate transported candidate velocities with a per-pixel median over valid candidates.
7. Assign zero velocity to pixels with no valid transported candidates.
8. Synthesize an 8-step constant-motion rollout from the final conditioning frame using `cv2.remap`.
9. Use `BORDER_REPLICATE` at boundaries.

No future ground-truth frames are used to estimate the prediction.

The implementation is in `scripts/compute_extended_metrics.py`, and the behavior is covered by CPU tests in `tests/test_extended_metrics.py`.

## Paired Statistics

N=30 paired comparisons use exact clip identity. For each clip, model and baseline metrics are paired by `clip_id`.

Reported paired outputs include:

- mean clip-level difference
- sample standard deviation
- t-based 95% confidence interval
- wins out of N=30

The release does not report p-values.

## Runtime Interpretation

Generative model timing was collected on RTX 6000 Ada GPU hardware. Classical baselines were measured with CPU OpenCV. This is a deployment characterization, not a hardware-normalized efficiency comparison.

For CPU baselines, decode time is tracked separately; the primary algorithm timing excludes decode where applicable.

## Historical MAD Consistency Gate

When recomputing extended metrics from saved generated MP4s, `scripts/compute_extended_metrics.py` recomputes model MAD and compares it against historical `aggregate_rollout_mad`. Nontrivial discrepancies stop aggregate generation.

## Qualitative Candidate Selection

N=30 qualitative candidates are selected from deterministic numerical rules in `scripts/make_vtc_n30_artifacts.py`, including the q25 pair-difference rule. The selected q25 qualitative figure is therefore tied to a reproducible metric-based rule rather than visual cherry-picking.
