Video catalogue · Verified local path

Self Gradient Forcing local guide

Apache-licensed autoregressive video diffusion release for minute-scale local text-to-video extrapolation from 5-second training windows.

Choose an app

Start with Recommended. No terminal commands are shown.

Compare video models

Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.

What it does

Text To VideoLong Video GenerationAnimationStreaming Video

The official GitHub release documents Python 3.10, PyTorch, FlashAttention, CUDA setup, a Hugging Face weight downloader, and framewise or chunkwise inference scripts. The public Hugging Face repository exposes Apache-2.0 model.pt checkpoints for both framewise and chunkwise SGF plus Causal-Forcing initialization weights, with direct unauthenticated downloads around 5.4 GB each. The launcher uses eight GPUs when available but falls back to single-GPU serial inference; LocalClaw records 64 GB RAM and 24 GB NVIDIA VRAM as a conservative entry floor for reduced single-GPU experiments because the default 240-second examples can be much heavier.

Hardware figures are practical entry floors, not performance guarantees. Resolution, duration, precision, offloading and runtime versions can materially change memory use.

Strengths

  • Official public code, training scripts and checkpoints
  • Framewise and chunkwise long-video extrapolation modes
  • Single-GPU fallback path despite multi-GPU default launcher

Limits to know

  • Default 963-latent-frame examples are slow and memory-heavy
  • Research stack depends on Wan and Causal-Forcing components
  • No official consumer VRAM table is published yet