# Models

This repository includes final-paper model configs and thin wrappers only. It does not include model weights, downloaded checkpoints, or external model repositories. Users must comply with upstream model licenses and terms.

## LTX-Video 2B Distilled

- Config: `configs/models/ltx/ltx-2b.yaml`
- Family: LTX-Video
- Conditioning regime: multi-frame video conditioning over the 9 observed frames
- Checkpoint filename in config: `ltxv-2b-0.9.8-distilled.safetensors`
- Spatial upscaler filename in config: `ltxv-spatial-upscaler-0.9.8.safetensors`
- External dependency: installed `ltx_video` package/API
- Weights included: no

Visible auxiliary model references in the config:

- `PixArt-alpha/PixArt-XL-2-1024-MS`
- `MiaoshouAI/Florence-2-large-PromptGen-v2.0`
- `unsloth/Llama-3.2-3B-Instruct`

## LTX-Video 13B

- Config: `configs/models/ltx/ltx-13b.yaml`
- Family: LTX-Video
- Conditioning regime: multi-frame video conditioning over the 9 observed frames
- Checkpoint filename in config: `ltxv-13b-0.9.8-distilled.safetensors`
- Spatial upscaler filename in config: `ltxv-spatial-upscaler-0.9.8.safetensors`
- External dependency: installed `ltx_video` package/API
- Weights included: no

Visible auxiliary model references match the LTX-2B config.

## Stable Video Diffusion 1.1

- Config: `configs/models/svd/svd-1.1.yaml`
- Family: Stable Video Diffusion through Diffusers
- Model ID: `stabilityai/stable-video-diffusion-img2vid-xt-1-1`
- Conditioning regime: final conditioning frame only
- External dependency: Diffusers `StableVideoDiffusionPipeline`
- Weights included: no

The wrapper receives the same 9-frame conditioning clip as other models, selects the final observed frame, and emits 8 predicted future frames. Hugging Face access or model-term acceptance may be required depending on the upstream model.

## Wan I2V 1.3B

- Config: `configs/models/wan/wan-i2v.yaml`
- Family: Wan image-to-video through Diffusers
- Model ID: `engineerA314/Wan2.1-Fun-V1.1-1.3B-InP-Diffusers`
- Conditioning regime: final conditioning frame only
- External dependency: Diffusers `WanImageToVideoPipeline`
- Weights included: no

The wrapper drops Wan's native first output frame because it is anchored by the conditioning image, then uses the next 8 frames as the rollout.

## Wan VACE 1.3B

- Config: `configs/models/wan/wan-vace-1.3b.yaml`
- Family: Wan VACE through Diffusers
- Model ID: `Wan-AI/Wan2.1-VACE-1.3B-diffusers`
- Conditioning regime: masked multi-frame video conditioning
- External dependency: Diffusers `WanVACEPipeline`
- Weights included: no

The wrapper builds a 17-frame canvas with the 9 observed frames and 8 masked future placeholders, then extracts the future rollout for evaluation.

## Inference Requirements and Notes

- Generative-model inference requires a CUDA-capable GPU and the model-specific dependencies described in [Setup](SETUP.md).
- Runtime measurements reported in the paper characterize the tested hardware and software configurations and should not be interpreted as hardware-normalized efficiency comparisons.
- Dependency requirements differ across model families and may change across upstream releases.
- Model checkpoints are obtained from their respective upstream sources and are not distributed with this repository.
- Some Hugging Face models may require authentication or acceptance of upstream access terms. Use standard Hugging Face authentication mechanisms; credentials should never be stored in repository files.
