Unable to Generate Full-Body Walking(e.g., TED Talk Style)
Hi team ,
First of all, thank you for open-sourcing InfiniteTalk — it’s an incredible project with impressive lip-sync and expression alignment results!
While testing the model, I noticed that most generated outputs (including the official examples) are limited to upper-body or half-body shots — usually up to the waist or chest level.
Issue
I’ve been trying to generate full-body, dynamic videos where a person is:
Standing and walking on stage (like giving a TED Talk)
Moving around while talking and addressing different sides of an audience
However, even after several attempts, the generated video still remains half-body, and no movement like walking or turning is observed — even when I specify it clearly in the prompt or use different modes (streaming/clip) and model resolutions (480P, 720P).
Questions
Is InfiniteTalk limited to upper-body motion because of its training data or architecture?
Are there any configuration options or input formats (e.g., specific JSON fields or motion control parameters) that could help enable full-body motion?
Would it be possible to extend the model or retrain it to support larger body movements and locomotion (e.g., walking, gesturing, turning)?
If InfiniteTalk is designed mainly for face and torso-level dubbing, could you please confirm that — so I can explore combining it with motion synthesis models for the lower body?
Tried prompts describing:
“A person giving a TED talk, walking on stage, turning to the audience, using hand gestures.”
But the video still remains static beyond upper body movement.
Request
Could you please guide me on how to generate such full-body, walking-talking style videos — or confirm if this limitation is due to the model design/training data?
Thank you again for your great work and open release! Looking forward to your guidance.
Source: MeiGen-AI/InfiniteTalk