AI
“LongCat-Video is an open-source model from Meituan that can turn a single photo and an audio track into a talking video with synchronized facial and mouth movements. The project supports more than realistic human faces. The same system can animate anime characters, animals, and other stylized subjects, while handling both single-speaker…”
Plain restatementMeituan has publicly released a video generation system, referred to as LongCat-Video, that accepts a reference image plus an audio track and outputs a lip-synced talking video; it is said to work on stylized subjects such as anime characters and animals as well as realistic humans, to accept single and multiple audio streams, and to use Whisper-Large-v3 as its audio encoder, with lip sync maintained over long clips. Code and weights are publicly downloadable.
Distortion codes this site does not recognise yet: scale_conflation, capability_extrapolation. Not collectible until the field guide has an entry.
This post is mostly right but it names the wrong model. Meituan has indeed open-sourced a system that turns one photo plus an audio track into a lip-synced talking video, and it does support anime characters, animals, and conversations with two audio streams. However, that system is called LongCat-Video-Avatar 1.5, released in May 2026. LongCat-Video, the name used in the post, is the underlying 13.6 billion parameter base model, and on its own it only does text-to-video, image-to-video, and video continuation with no audio input at all. The Whisper-Large-v3 detail is also version-specific: it applies to version 1.5, while the earlier December 2025 avatar release used a different audio encoder called wav2vec2. The GitHub link in the post is correct, since the avatar code lives in that same repository, and the licence really is MIT, so anyone can run and build on it. One caveat worth keeping in mind is that all the quality claims, including the anime and animal support and the stability over long clips, come from Meituan's own model card and its own unreviewed technical report, with no independent testing found.
[drifted from the evidence:] LongCat-Video is an open-source model from Meituan that [drifted from the evidence:] can turn a [drifted from the evidence:] single photo and an audio track [drifted from the evidence:] into a talking video [drifted from the evidence:] with synchronized facial and mouth movements. The project supports more than realistic human faces. The same system can animate anime characters, [drifted from the evidence:] animals, and [drifted from the evidence:] other stylized subjects, while handling both single-speaker and [drifted from the evidence:] multi-speaker audio. [drifted from the evidence:] For speech alignment, the pipeline uses Whisper-Large-v3 to [drifted from the evidence:] help extract audio [drifted from the evidence:] features and keep lip movements matched with [drifted from the evidence:] the spoken words, including across longer generated clips. [drifted from the evidence:] Because the code and [drifted from the evidence:] model are publicly [drifted from the evidence:] available, developers can run, test, and build on the technology themselves instead of relying entirely on closed commercial avatar platforms." (Source given: https://github.com/meituan-longcat/LongCat-Video)
Meituan [added by the neutral restatement:] has publicly released a video generation system, referred to as LongCat-Video, that [added by the neutral restatement:] accepts a [added by the neutral restatement:] reference image plus an audio track [added by the neutral restatement:] and outputs a [added by the neutral restatement:] lip-synced talking video; [added by the neutral restatement:] it is said to work on stylized subjects such as anime characters and [added by the neutral restatement:] animals as well as realistic humans, to accept single and [added by the neutral restatement:] multiple audio [added by the neutral restatement:] streams, and to [added by the neutral restatement:] use Whisper-Large-v3 as its audio [added by the neutral restatement:] encoder, with [added by the neutral restatement:] lip sync maintained over long clips. Code and [added by the neutral restatement:] weights are publicly [added by the neutral restatement:] downloadable.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- Meituan's LongCat team has publicly released an open-source, MIT-licensed audio-driven avatar system that turns a reference image plus audio into a lip-synced talking video (Audio-Text-Image-to-Video).
- It supports stylized subjects: the card lists robust generalization to anime, animals and complex real-world conditions such as multi-person interactions.
- It supports both single-stream and multi-stream audio inputs, with documented merge and concatenation modes for two speakers.
- The v1.5 checkpoint does use a Whisper-large-v3 audio encoder, replacing wav2vec2, for better lip sync.
- Long-clip handling is a stated design goal: the report describes accurate lip-synchronization, full-body temporal stability and robust long-video generation with strict identity consistency.
- Code and weights are publicly downloadable from the GitHub repo the post cites, so the "developers can run and build on it" framing is correct.
- Scale conflation: the post attributes all of this to "LongCat-Video." That name belongs to the 13.6B foundation model whose documented tasks are Text-to-Video, Image-to-Video and Video-Continuation, with no audio conditioning. The audio-driven behaviour belongs to a separate checkpoint line, LongCat-Video-Avatar and LongCat-Video-Avatar-1.5. A reader who downloads the weights named in the post gets a model that cannot do what the post describes; they need both weight directories. The mitigating fact is that the avatar code lives in the linked repo, so the URL is right even though the model name is not.
- Omitted qualifier: "the pipeline uses Whisper-Large-v3" is true only of v1.5. v1.0, the default model_type, uses wav2vec2. Stated without the version, it describes the family incorrectly for the checkpoint released in December 2025.
- Marketing as evidence: every capability statement in the post traces to Meituan's own model card and its own unrefereed technical report. The anime and animal generalization, the long-clip stability and the lip-sync accuracy are vendor self-descriptions. That is fine as evidence that the features are offered; it is not independent evidence that they work well.
- Capability extrapolation (mild): "keep lip movements matched with the spoken words, including across longer generated clips" is stated flatly. The underlying source frames long-video stability as an engineering goal validated on the authors' own 508-case benchmark, not as an unconditional property.
- How well the stylized-subject and long-clip claims hold under ordinary user conditions. No independent quantitative evaluation of LongCat-Video-Avatar-1.5's lip sync or identity consistency was located, only vendor benchmarks and informal hands-on reports.
- Whether the post intended v1.0 or v1.5. It is version-ambiguous, and the Whisper detail only fits v1.5, so I resolved it to v1.5.
- The reported parity or superiority against HeyGen, OmniHuman 1.5 and Kling Avatar 2.0 is entirely vendor-run on a vendor-designed benchmark and is not independently verified. The post does not make this claim, but it is the implicit backdrop of the "instead of closed commercial platforms" line.
- Exact release date of v1.0 varies between sources (the repo news log says Dec 16, 2025; at least one third-party page says November 2025). I used the repo.
There are two distinct things in this family, and the claim merges them. The base model. LongCat-Video is a foundational video generation model with 13.6B parameters covering Text-to-Video, Image-to-Video and Video-Continuation, and it unifies those three tasks in a single framework. It was released Oct 25, 2025. Audio conditioning is not among its listed tasks; its Hugging Face card tags it text-to-video, image-to-video and video-continuation. The avatar model. The audio-driven system is a separate checkpoint line. LongCat-Video-Avatar was announced Dec 16, 2025 as a unified model for audio-driven character animation supporting Audio-Text-to-Video, Audio-Text-Image-to-Video and Video Continuation, with compatibility for single-stream and multi-stream audio inputs. On May 21, 2026 Meituan released LongCat-Video-Avatar-1.5, which replaces Wav2Vec2 with Whisper-Large for more accurate lip synchronization, generalizes to stylized domains (anime, animals, complex real-world conditions), supports single-stream and multi-stream audio, and accelerates inference to 8 steps via step distillation. Its card states it is built on the LongCat-Video foundation model and supports AT2V, ATI2V and Video Continuation, and lists stylized domain generalization to anime, animals and multi-person interactions. The encoder detail is version-gated. The card states that --model_type avatar-v1.0 uses the wav2vec2 audio encoder by default, while --model_type avatar-v1.5 uses the Whisper-large-v3 audio encoder for better lip sync quality. The avatar code ships inside the repo the post links to: LongCat-Video-Avatar extends the foundational LongCat-Video pipeline by adding audio conditioning, reusing the core model components while introducing audio processing modules, and supports single-audio input for one character and multi-audio input with two streams for multi-character dialogues. A third party notes the dependency explicitly: it depends on two weight directories, LongCat-Video as the base video generation model and LongCat-Video-Avatar-1.5 as the avatar model. Licensing: the LongCat-Video model weights are released under the MIT License, and a practitioner review describes 1.5 as an open-weight audio-driven human video model built on LongCat-Video and released under MIT. Quality claims are vendor-run. The tech report states v1.5 achieves competitive or superior performance versus leading closed-source systems such as HeyGen, OmniHuman 1.5 and Kling Avatar 2.0 on human-likeness and expert quality assessments on the authors' own benchmark, a benchmark of 6 application scenarios, 2 languages and 2 visual styles totalling 508 image-audio pairs, scored by 770 crowdsourced evaluators producing 13,240 judgments plus 10 domain experts.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/f83779d06df2/qF0l4WBz1t4XiHO3UpE5-uo8nYh
Ask this case
Answers come only from the case file above; nothing is added.
Does LongCat-Video actually turn a photo and audio into a talking video?
Not by itself. LongCat-Video is Meituan's 13.6 billion parameter base model, and its documented tasks are text-to-video, image-to-video, and video continuation, with no audio input. The photo-plus-audio talking video feature belongs to a separate checkpoint called LongCat-Video-Avatar, which is built on top of the base model.
So is the claim wrong about the anime, animal, and multi-speaker support?
No, those features are real, but they belong to LongCat-Video-Avatar-1.5, not to LongCat-Video itself. The avatar model's own card lists generalization to anime, animals, and complex real-world conditions, plus support for single-stream and multi-stream audio.
Does the model really use Whisper-Large-v3 for lip sync?
Only in the newer version. LongCat-Video-Avatar-1.5, released in May 2026, uses Whisper-Large-v3 for better lip sync, but the earlier December 2025 version used a different audio encoder called wav2vec2.
Is the GitHub link and open-source claim accurate?
Yes. The avatar code lives in the same repository the post links to, and the weights are released under the MIT License, so developers can download, run, and build on it themselves.
How trustworthy are the quality claims about lip-sync accuracy and stability on long clips?
The investigation did not find independent testing of these claims. All the performance and quality statements come from Meituan's own model card and its own unreviewed technical report, evaluated on the company's own benchmark.