Video catalogue · Verified local path
4DAnyone local guide
Video-to-video model that turns a casual monocular human video into multiview-consistent videos for free-viewpoint 4D reconstruction.
Start with Recommended. No terminal commands are shown.
Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.
What it does
Ant Research publishes official code, a local CLI, a GUI, and a public Hugging Face model repository with 4danyone/model.safetensors plus VAE, prompt context, SMPL-X regressor and support assets. The README reports peak CUDA memory below 24 GB, 27 seconds per 121-frame target-view video on one RTX 4090, and automatic model/example downloads; LocalClaw records 64 GB RAM and 24 GB NVIDIA VRAM as the conservative practical floor for the Turbo inference path.
Strengths
- Official Ant Research weights, CLI and interactive GUI
- Runs under the 24 GB CUDA memory threshold reported by the authors
- Produces synchronized multiview outputs for downstream 4D Gaussian Splatting
Limits to know
- Human-centric reconstruction model, not a prompt-only generator
- Repository combines Apache-2.0 first-party weights with third-party non-commercial, research, AGPL and attribution-licensed assets
- Input quality and pose recovery strongly affect final multiview consistency