Meta Muse Video

Meta Super Intelligence Laboratory’s first self-developed Vincent video model, native audio and images are generated at once

In-depth Report

  • On July 7, 2026, Meta Superintelligence Labs (MSL) released an early preview version of the video model Muse Video while releasing the self-developed image model Muse Image. This is Meta's first completely self-developed Vincent video model. The biggest highlight is the "native audio" - sound and images are generated at once in the same generation process, instead of producing a silent video first and then soundtracking it later. On Arena Vincent Video's artificial preference list, it ranked third at its debut. The model is still in preview status and has not yet been officially released. Meta itself admits that there are still obvious shortcomings in the physical accuracy of audio and video synchronization and rapid movement.

  • Muse Video is operated by the Meta Super Intelligence Laboratory led by Alexandr Wang. It is the third important model release of MSL after the language model Muse Spark in April 2026 and the image model Muse Image in June. It shares the same pre-trained base with Muse Image, but extends its capabilities from static images to moving images. The strategic significance of this release is clear: Meta wants to get rid of its dependence on third-party graphics tools such as Midjourney and Black Forest Labs, integrate image and video generation into its own "Muse" brand matrix, and tie it with the Muse Spark inference model and Meta AI products into a complete multi-modal creation stack. MSL’s official account on X, @AIatMeta, confirmed the release in a post on July 8.

  • The core capability of Muse Video is to generate videos from text, and it also supports image-based videos. Unlike most Vincent video models that first generate silent clips and then use audio models to fill in the sounds, it calculates dialogue, environmental sounds, soundtracks and images together in the same generation process. In theory, the sound and images are aligned from the beginning. Meta officially emphasizes that it performs competitively in three dimensions: prompt adherence, visual fidelity, and temporal consistency across frames (temporal consistency, that is, character identity, scene, and camera movement remain stable in multiple shots). Judging from the online trial generator on the product page musevideo.dev, users can directly describe the scene (subject, camera movement, lighting, dialogue, sound effects), choose Vincent video or Tusheng video, the resolution is 480p and 720p, the duration is from 4 seconds to 15 seconds, the aspect ratio covers 16:9, 9:16, 1:1, 21:9, etc., and there are two modes: Fast (cheap) and Higher quality (720p). The bottom layer uses MSL's agentic media framework. It will perform reasoning, call tools, and self-correct during reasoning, which is in the same vein as Muse Image. However, Meta itself made the shortcomings very straightforward: First, the audio and video synchronization in fast scenes, especially lip-sync, is not stable; second, the physical rationality of fast movements (sports, action scenes) is insufficient, and motion blur and temporal artifacts will occur. These are exactly the directions the team will focus on before the preview version is released.

  • Muse Video is currently in preview, and official pricing has not yet been announced. You can see its billing idea from the trial generator on the product page: press credit, 480p costs about 40 credits/second, and generates about 160 credits at a time. Registration will give you a free credit. As a reference, basic use of Muse Image in Meta AI is free, and heavy use is blocked behind monthly subscription plans such as Meta One; while MSL has reduced the price of Muse Spark API by up to 80% during the same period, and it is generally expected that the price of Muse Video will be quite aggressive when it is officially released. For developers, Meta’s plan is to package Muse Image, Muse Video, and Muse Spark into the same API, so that teams working on multi-modal applications no longer need to connect multiple suppliers.

  • After the preview was released, the feedback from those who got started early was obviously divided. On the positive side, I feel that its picture quality is close to Byte's Seedance 2.0, and sometimes even better, obviously better than Google Veo 3.1; the native audio is the real killer, some people have cut their Reels workflow from three tools to one prompt box, and there is no need to record separate narration for spoken-word videos. On the other hand, many testers pointed out that it is currently stuck at the upper limit of about 10 seconds, while Seedance 2.0 can already produce 15-second, up to 4K movie-like clips; some also complained that body movements and sounds in fast-paced shots did not match up, which was traced by AI generation at first glance.

  • In the Vincent video landscape in July 2026, the first tier is OpenAI's Sora 2 Pro (cinematic ceiling, only available in ChatGPT Pro file), Google's Veo (prompt word compliance and physics are the most stable, fully available in Google Cloud), Byte's Seedance 2.5 (native 30 seconds, local editing), plus Runway, Kling, and Pika. Horizontally, Muse Video's unique advantage is native audio plus Meta's distribution coverage - once it is officially released on Instagram, WhatsApp, and Meta AI, it may become the most used video model overnight with 3 billion daily users. Its weakness is that it is still in the preview period and there is no clear GA date. The cinematic feel is temporarily different from Sora 2 Pro, and audio-visual synchronization and fast motion are weaknesses that Meta itself claims. Most industry analysts believe that the scenarios it can really compete in are social short videos and native audio content, and professional film and television will have to wait for it to make up for its shortcomings.

  • Muse Video itself didn't make too big a deal, but it shared the architecture with Muse Image and followed the same consumer-grade product path, so Muse Image's privacy controversy couldn't be avoided. After Muse Image went online, Instagram "checked" users' public accounts into a function by default, allowing others to use these public photos to generate new AI images without the owner of the photos knowing. Groups such as the actors' union SAG-AFTRA criticized it as opening the door to harassment and impersonation. Meta removed the entire cross-account remix function on Instagram on July 10, 2026, admitting that it "missed the mark"; however, the Muse Image itself has always been available in the Meta AI app and WhatsApp. This lesson from the past means that when Muse Video officially enters Instagram, the issues of data authorization and portrait rights are likely to be put on the table again. In addition, as a closed-source model, it does not have a public API and cannot be audited like open source weights, so it occasionally suffers from professional accuracy.

  • If you are making "short and fast" content such as social short videos, advertising slices, product warm-ups, and educational materials, and are already active in the Instagram/WhatsApp ecosystem, Muse Video is worth waiting for its official version - native audio can really save a process. If you are pursuing professional film and television with cinematic quality, long shots or precise movements, Sora 2 Pro or Veo are more reliable at this stage; if you need interactive editing workflow, choose Runway; if you have a limited budget and pursue rapid iteration, you can look to Kling or Pika. It should be reminded that if all generated content is used commercially, it must ensure that the prompt words do not infringe the rights of third parties and comply with the platform content policy; before the official version is released, it is not realistic to use it for production-level delivery.

  • Muse Video is a key step for Meta to use the same set of bases to combine images and videos into a creative stack. Its native audio and three billion distribution channels give it a natural advantage that others do not have; but it is still a preview version. Until the shortcomings of audio and video synchronization and fast motion are repaired, it is more suitable to "wait for it to grow up" in social short videos rather than use it for professional work now.

User Reviews

  • 头像
    刘然桂
    The native audio is really good, eliminating the need for post-production soundtracking.

  • 头像
    BStephens_2022
    Arena comes in third, and Meta does have something this time around.

  • 头像
    WOgom
    Being stuck at 10 seconds is too uncomfortable to do long shots.

  • 头像
    ya4prrkzml
    I compared it carefully with Seedance 2.0. The picture quality is indeed very close, and sometimes slightly better, but the duration is quite different - Seedance 2.0 can already produce 15-second, up to 4K movie-like clips in one breath, while the Muse Video preview version is still stuck at about 10 seconds, and the resolution is only 720p. A short product preview video is enough, but a serious advertisement is not enough.

  • 头像
    MJenkinsQ
    Our team reduced the Reels workflow from three tools to one prompt box. For spoken-word videos, there is no need to record separate narrations. This step saves real time. However, the mouth shape and voice still don't match up in fast-paced shots, and the transitions when walking in the wind are occasionally blurry, which can be seen as traces of AI. My suggestion is that before the official release, it is best not to touch scenes such as interviews, oral broadcasts, and fast movements. It is more stable to use it for static or medium-low speed social media slicing.

  • 头像
    AmyThomas_2021
    The Instagram farce of using other people’s photos to generate AI pictures has just subsided. When Muse Video enters Instagram, the issue of portrait rights will probably explode again.

  • 头像
    Dennis.RichardsonIII
    Preview version, wait until the party wins.

  • 头像
    Pamela_Roberts_7481
    It’s good to have the same base as Muse Image. I make the image first and then throw it into Muse Video to move it. The process is very smooth.

  • 头像
    Scott_Parker_2023855
    Closed source does not have an API, so you have to wait to connect it to your own products.

  • 头像
    AAnderson_66
    Better than Veo 3.1, at least that's what early testing says.

  • 头像
    8_3rtf
    From a horizontal comparison, its unique brand is native audio plus 3 billion daily active distribution. Sora 2 Pro and Veo do have better image quality, but that requires payment and a different platform. Once Meta officially plugged Muse Video into Instagram and WhatsApp, it became the most used video model overnight by default. This is Meta's real approach - not to simply fight for picture quality, but to lock creation into its own ecosystem.

  • 头像
    MarkHendersonSr511
    Credit billing is 40 credits/second, and registration will give you free credit. If you use it heavily, you may still need to subscribe. The price has not been announced yet. We will wait and see.

  • 头像
    PPerry_99
    Audio-visual synchronization and motion physics are shortcomings that Meta itself acknowledged in its release, which is quite honest. Action scenes, sports, dancing, etc. have a high probability of overturning, and the lip-sync alignment is not yet stable. Now, if you use it to make something, you must manually review it, don't send it directly. But it iterates very quickly. For reference, Muse Image only took three days from release to fixing privacy issues. Looking at it after three months, these pitfalls will most likely be fixed.

  • 头像
    Olivia511
    Sora 2 Pro is still a ceiling.

  • 头像
    Katherine_MorganX5
    The prompt follows well. I wrote "handheld follow-up shooting with shallow depth of field during backlit golden hours", and it really gave me the language of the corresponding lens.

  • 头像
    DragoslavDaničić
    The proof of concept is sufficient, but production delivery is still early.

  • 头像
    赵兰浩
    The most cost-effective thing for social media creators and small teams is to bring their own sound. Product warm-up and teaching clips can be directly produced with audio and video from scratch, without the need for editing and matching BGM. However, there are several pitfalls that you should keep in mind when using it commercially: Do not include other people’s portraits or copyrighted materials in the prompts, abide by the platform’s content policy, and if the generated content involves real people, it is best to obtain authorization first. It’s okay to test the waters during the preview period. If you really want to commercialize it in batches, you should wait for the official version and clear licensing terms.

  • 头像
    KBell369
    During the Muse Image privacy scandal, Meta removed the cross-account remix from the shelves on July 10, admitting that it missed the mark. This is a valuable reference for judging the risks of Muse Video on Instagram.

  • 头像
    卢萱_1
    Waiting for GA, now I can only play the preview.