How I Tested Five Talking-Head Tools and One Actually Worked
What I was trying to do
What I was trying to do
One still portrait, one voice track, render a short talking-head avatar video. I wanted to run it locally where I could. I expected the hosted options to win — they have more compute behind them, someone else is managing the infrastructure, and one of them wraps a face-swap step on top of a polished generation pipeline. I was wrong about almost all of that.
The failure that cost me the most time
Wan-Animate went first. It took roughly 73 minutes of compute to produce 3 seconds of video. I could have lived with that if the output were usable, but the face that came out did not look like the person in the source image, and the background drifted and warped visibly across the clip. That is not a stylistic complaint — the warping made the clip unwatchable. I cut it after one render.
LivePortrait, using a static-image masking approach, had a different problem: ghosting. On any head movement, a faint afterimage of the face trailed by about 2.1 seconds. It looked like a double-exposure gone wrong. I tried adjusting the masking parameters and got nowhere useful fast.
A distilled "fast" variant of InfiniteTalk collapsed on lip-sync. The mouth simply stopped tracking the audio partway through. I am not sure whether the distillation broke something specific or whether my input was outside the model's comfort zone, but the result was a face mouthing nothing while the voice continued.
The hosted option that almost worked
The pipeline that got furthest before failing was a hosted model — OmniHuman 1.5 via fal — followed by an inswapper face-swap step to restore the original likeness. Technically it completed without crashing. The output was blurry, low-resolution, and the expression had drifted enough that it no longer read as the same person. The face-swap step was supposed to fix the identity problem. It didn't. What I got was a slightly different blurry person.
I had also dropped LoRA-fine-tuned source images earlier in the process because they looked obviously AI-generated in the final composite. A plain GPT-4o still frame held up better as the base input than any fine-tuned variant I tried.
What actually worked
InfiniteTalk, the full open-source model, Apache-2.0, run locally. Output around 720p. Face fidelity was better than anything the hosted pipeline produced, including the version with the face-swap correction layer on top. I did not expect a local model to beat the hybrid setup I had been running.
It has one problem I have not solved: the lower half of the face moves a little stiffly at times. The mouth opens and closes correctly for the phonemes, but the motion looks mechanical rather than natural — something is off in how the jaw and lips transition between positions. It is not unwatchable the way the ghosting or the warping was, but it is noticeable if you are looking for it.
The pattern across all five tools was the same failure mode dressed differently: the options that looked more capable all broke on the one thing a talking-head clip cannot survive without — keeping the face the same real person from the first frame to the last.
One rule from this
Before you spend time on a hosted pipeline or a distilled fast variant, render 3 seconds with the boring local model first and check whether the identity holds across the clip — every other quality problem is fixable downstream, and that one usually is not.
Editor's note: AI tools were used to help draft parts of this story. The techniques it describes come from work originally written by human practitioners, and the story was reviewed and edited by a human before publishing.
About the Creator
Yohei Maruyama
I turn proven AI workflows from Japan's tech community (Zenn, Qiita, and peers) into practical English guides—not translations.
Posts are AI-assisted for clarity; the workflows and judgment are mine.
Based in Tokyo.
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed.
Comments
There are no comments for this story
Be the first to respond and start the conversation.