Video catalogue · Verified local path

SoulX-LiveAct local guide

Apache-licensed audio-driven human-animation model for hour-scale talking, music, podcast and FaceTime-style avatar video.

Choose an app

Start with Recommended. No terminal commands are shown.

Compare video models

Desktop app links require the app to be installed. If nothing opens, LocalClaw will show app-download and model-file fallbacks.

What it does

Speech To VideoAudio To VideoImage To VideoHuman AnimationEmotion Editing

Soul-AILab publishes the LiveAct safetensors checkpoint, official inference code, GUI demo commands and a chinese-wav2vec2-base companion download path. The README documents Python 3.10, SageAttention, vLLM, LightVAE and local generate.py/demo.py entrypoints; two H100/H200 GPUs are the realtime 20 FPS reference, while the authors added RTX 4090/RTX 5090 support through FP8 KV cache, block offload and T5 CPU offload and report 6 FPS on one RTX 5090. LocalClaw records 64 GB RAM and 24 GB NVIDIA VRAM as the conservative consumer-GPU offload floor.

Hardware figures are practical entry floors, not performance guarantees. Resolution, duration, precision, offloading and runtime versions can materially change memory use.

Strengths

  • Official Soul-AILab weights and inference repository
  • Streaming and GUI demo paths for audio-driven human video
  • Consumer NVIDIA path with FP8 KV cache and block offload

Limits to know

  • Realtime reference performance still uses two H100/H200 GPUs
  • Specialized human animation model, not general scene generation
  • Depends on Wan2.1 I2V, wav2vec and compiled CUDA attention/runtime components