Video catalogue · Verified local path

4DAnyone local guide

Video-to-video model that turns a casual monocular human video into multiview-consistent videos for free-viewpoint 4D reconstruction.

Choose an app

Start with Recommended. No terminal commands are shown.

Compare video models

Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.

What it does

Video To VideoMultiview Video GenerationNovel View Synthesis4d Human ReconstructionAnimation

Ant Research publishes official code, a local CLI, a GUI, and a public Hugging Face model repository with 4danyone/model.safetensors plus VAE, prompt context, SMPL-X regressor and support assets. The README reports peak CUDA memory below 24 GB, 27 seconds per 121-frame target-view video on one RTX 4090, and automatic model/example downloads; LocalClaw records 64 GB RAM and 24 GB NVIDIA VRAM as the conservative practical floor for the Turbo inference path.

Hardware figures are practical entry floors, not performance guarantees. Resolution, duration, precision, offloading and runtime versions can materially change memory use.

Strengths

  • Official Ant Research weights, CLI and interactive GUI
  • Runs under the 24 GB CUDA memory threshold reported by the authors
  • Produces synchronized multiview outputs for downstream 4D Gaussian Splatting

Limits to know

  • Human-centric reconstruction model, not a prompt-only generator
  • Repository combines Apache-2.0 first-party weights with third-party non-commercial, research, AGPL and attribution-licensed assets
  • Input quality and pose recovery strongly affect final multiview consistency