HyperAIHyperAI

Command Palette

Search for a command to run...

Make-A-Video: Text-to-Video Generation without Text-Video Data

Abstract

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in aesthetic, fantastical depictions, etc.) of today's image generation models. We design a simple yet effective way to build on T2I models with novel and effective spatial-temporal modules. First, we decompose the full temporal U-Net and attention tensors and approximate them in space and time. Second, we design a spatial temporal pipeline to generate high resolution and frame rate videos with a video decoder, interpolation model and two super resolution models that can enable various applications besides T2V. In all aspects, spatial and temporal resolution, faithfulness to text, and quality, Make-A-Video sets the new state-of-the-art in text-to-video generation, as determined by both qualitative and quantitative measures.

One-sentence Summary

Researchers at Meta AI propose Make-A-Video, which decouples the learning of visual appearance from text-image data and motion dynamics from unsupervised video by decomposing the full temporal U-Net and attention tensors spatially and temporally, then integrates a spatial-temporal pipeline with a video decoder, an interpolation model, and two super-resolution models, achieving state-of-the-art text-to-video generation without paired text-video data, as demonstrated by both qualitative and quantitative measures.

Key Contributions

  • Make-A-Video adapts a text-to-image diffusion model for text-to-video generation by fine-tuning with pseudo-3D convolutions and temporal attention layers, eliminating the need for paired text-video data.
  • A spatial-temporal pipeline combines a video decoder, an interpolation model, and two super-resolution models to generate high-resolution, high-frame-rate videos.
  • Make-A-Video sets a new state-of-the-art in text-to-video generation, achieving superior spatial and temporal resolution, text faithfulness, and overall quality as measured by qualitative and quantitative evaluations.

Introduction

The recent explosion in text-to-image (T2I) generation relies on billions of alt-text and image pairs harvested from the web, but a similarly scaled text-to-video (T2V) dataset does not exist. Training T2V models from scratch is both wasteful and impractical given the scarcity of high-quality paired text-video data and the high dimensionality of video. Prior T2V approaches are either confined to narrow domains, require large private paired datasets, or freeze image-model weights, limiting adaptability. The authors propose Make-A-Video, which circumvents paired data entirely by bootstrapping a diffusion-based T2I model with unsupervised video learning. They extend the spatial network with temporal attention modules via function-preserving initialization, transferring image knowledge to video instantly, and add spatial and temporal super-resolution strategies to generate high-definition, high-frame-rate videos from text alone.

Dataset

The authors rely on a mix of public image-text and video-only datasets, along with curated evaluation sets, to train and assess their text-to-video generation system.

  • Image training data: A 2.3B‑pair English subset of the large-scale dataset from Schuhmann et al. (likely LAION). Pairs are filtered to remove NSFW images, toxic words in the captions, and images with a watermark probability above 0.5. This filtered collection is used to train the image models.

  • Video training data (no text):

  • WebVid-10M: the full 10M‑video dataset, used to train the decoder Dt and the interpolation model.

  • HD-VILA-10M: a 10M‑video subset randomly sampled from HD-VILA-100M. The super‑resolution model SRtl is trained on both WebVid-10M and HD-VILA-10M. All video datasets are used as raw frames only; no aligned captions or text annotations are employed.

  • Zero‑shot evaluation benchmarks:

  • UCF-101: 101 action classes; a single fixed template sentence per class is used for evaluation. 10,000 samples are generated following the training‑set class distribution, and FVD and IS are reported.

  • MSR-VTT: all 59,794 captions from the official test set are used; FID and CLIPSIM (average CLIP similarity) are computed.

  • Human evaluation prompts:

  • AMT‑300: 300 prompts collected from Amazon Mechanical Turk. Annotators were asked to propose interesting T2V prompts. Incomplete, overly abstract, and offensive prompts were removed, and the remaining prompts were grouped into five categories: animals, fantasy, people, nature and scenes, and food and beverage. The prompts were fixed without any video generation.

  • DrawBench: the set of prompts originally introduced in Imagen is also used for human evaluation of video quality and text‑video faithfulness.

Method

Make-A-Video consists of three main components: a base text-to-image (T2I) model trained on text-image pairs, spatiotemporal convolution and attention layers that extend the network building blocks to the temporal dimension, and spatiotemporal networks that include a frame interpolation network for high frame rate generation.

The final text-to-video inference scheme can be formulated as: y^t=SRh∘SRlt∘↑F∘Dt∘P∘(x^,Cx(x))\hat{y}_t = \mathrm{SR}_h \circ \mathrm{SR}_l^t \circ \uparrow_F \circ \mathrm{D}^t \circ \mathrm{P} \circ (\hat{x}, \mathrm{C}_x(x))y^​t​=SRh​∘SRlt​∘↑F​∘Dt∘P∘(x^,Cx​(x)) where y^t\hat{y}_ty^​t​ is the generated video, SRh\mathrm{SR}_hSRh​ and SRlt\mathrm{SR}_l^tSRlt​ are the spatial and spatiotemporal super-resolution networks, ↑F\uparrow_F↑F​ is the frame interpolation network, Dt\mathrm{D}^tDt is the spatiotemporal decoder, P\mathrm{P}P is the prior, x^\hat{x}x^ is the BPE-encoded text, Cx\mathrm{C}_xCx​ is the CLIP text encoder, and xxx is the input text.

As shown in the figure below:

Prior to adding temporal components, the authors train the backbone T2I model. This model uses a prior network P\mathrm{P}P that generates image embeddings yey_eye​ given text embeddings xex_exe​ and BPE encoded text tokens x^\hat{x}x^. A decoder network D\mathrm{D}D generates a low-resolution 64×6464 \times 6464×64 RGB image y^l\hat{y}_ly^​l​ conditioned on the image embeddings yey_eye​. Two super-resolution networks SR1\mathrm{SR}_1SR1​ and SRh\mathrm{SR}_hSRh​ then increase the generated image resolution to 256×256256 \times 256256×256 and 768×768768 \times 768768×768 pixels respectively.

To expand the two-dimensional conditional network into the temporal dimension, the authors modify the convolutional and attention layers. Refer to the framework diagram for the architecture and initialization scheme of these layers:

Motivated by separable convolutions, a 1D convolution is stacked following each 2D convolutional layer. This facilitates information sharing between the spatial and temporal axes without the heavy computational load of 3D convolutions. Given an input tensor h∈RB×C×F×H×Wh \in \mathbb{R}^{B \times C \times F \times H \times W}h∈RB×C×F×H×W, where B,C,F,H,WB, C, F, H, WB,C,F,H,W are the batch, channels, frames, height, and width dimensions respectively, the Pseudo-3D convolutional layer is defined as: ConvP3D(h):=Conv1D(Conv2D(h)∘T)∘T\mathrm{Conv}_{P3D}(h) := \mathrm{Conv}_{1D}(\mathrm{Conv}_{2D}(h) \circ T) \circ TConvP3D​(h):=Conv1D​(Conv2D​(h)∘T)∘T where the transpose operator ∘T\circ T∘T swaps between the spatial and temporal dimensions. For smooth initialization, the Conv2D\mathrm{Conv}_{2D}Conv2D​ layer is initialized from the pre-trained T2I model, while the Conv1D\mathrm{Conv}_{1D}Conv1D​ layer is initialized as the identity function.

Similarly, temporal attention layers are stacked following spatial attention layers to approximate full spatiotemporal attention. Given an input tensor hhh, the Pseudo-3D attention layer is defined as: ATTNP3D(h)=unflatten(ATTN1D(ATTN2D(flatten(h))∘T)∘T)\mathrm{ATTN}_{P3D}(h) = \text{unflatten}(\mathrm{ATTN}_{1D}(\mathrm{ATTN}_{2D}(\text{flatten}(h)) \circ T) \circ T)ATTNP3D​(h)=unflatten(ATTN1D​(ATTN2D​(flatten(h))∘T)∘T) where flatten and unflatten are matrix operators that manipulate the spatial dimensions. The ATTN2D\mathrm{ATTN}_{2D}ATTN2D​ layer is initialized from the pre-trained T2I model and the ATTN1D\mathrm{ATTN}_{1D}ATTN1D​ layer is initialized as the identity function. Additionally, the authors add a frame rate conditioning parameter fpsfpsfps to enable augmentation and provide control over the generated video at inference time.

To increase the frame rate within memory and compute constraints, the authors train a masked frame interpolation and extrapolation network ↑F\uparrow_F↑F​. The spatiotemporal decoder Dt\mathrm{D}^tDt is fine-tuned on masked frame interpolation by zero-padding the masked input frames. When fine-tuning, an additional 4 channels are added to the U-Net input: 3 channels for the RGB masked video input and an additional binary channel indicating which frames are masked. This enables video upsampling with variable frame-skips and fpsfpsfps conditioning.

The different components of Make-A-Video are trained independently. The prior P\mathrm{P}P is trained on paired text-image data. The decoder, prior, and two super-resolution components are first trained on images alone. After training on images, the new temporal layers are added and initialized, then fine-tuned over unlabeled video data. During this phase, 16 frames are sampled from the original video with random fpsfpsfps ranging from 1 to 30. The masked-frame interpolation component is subsequently fine-tuned from the temporal decoder.

Experiment

The evaluation setup involves training on public image-text and unlabeled video datasets, with zero-shot automatic assessment on MSR-VTT and UCF-101 and human evaluations using collected prompts and DrawBench to judge video quality, text-video faithfulness, and interpolation realism. Results show that Make-A-Video generalizes significantly better than prior text-to-video models, surpassing even those fine-tuned on the same benchmarks, and produces videos with stronger motion consistency and text alignment. Human raters consistently prefer Make-A-Video over CogVideo and VDM, and its interpolation model offers more semantically meaningful motion than FILM. Qualitative examples confirm the model's versatility across tasks like image animation and video variation, with richer content and better real-world motion understanding.

Make-A-Video establishes a new state-of-the-art on MSR-VTT in the zero-shot setting, achieving the lowest FID (13.17) and the highest CLIPSIM (0.3049) while generating only a single sample per prompt. It substantially outperforms prior zero-shot model CogVideo in both Chinese and English input modes, and also surpasses fully supervised methods GODIVA and NÜWA, demonstrating far stronger generalization. Make-A-Video’s FID of 13.17 is roughly half that of CogVideo’s best score (23.59 for English) and over three times lower than NÜWA’s 47.68, all with just one generated sample per input. Make-A-Video obtains a CLIPSIM of 0.3049, exceeding both zero-shot CogVideo (0.2631) and the non-zero-shot GODIVA (0.2402) and NÜWA (0.2439), indicating better text-video alignment.

On UCF-101, Make-A-Video achieves the highest Inception Score and lowest FVD among zero-shot class-conditional models, notably outperforming CogVideo in both Chinese and English settings even when using a lower resolution. Fine-tuning further reduces FVD dramatically, setting a new state-of-the-art and generating more coherent videos than prior work. Zero-shot Make-A-Video obtains a much higher Inception Score and much lower FVD than CogVideo, despite generating at 256×256 rather than 480×480. Make-A-Video's fine-tuned model reduces FVD well below the previous best of 577, establishing new state-of-the-art video coherence.

In human evaluations, raters consistently preferred Make-A-Video over prior models for both video quality and text-video faithfulness. Against VDM on its own curated set of prompts, Make-A-Video received strong majority preference. When compared to CogVideo, preference rates remained well above chance across two benchmarks and two prompt languages, confirming robust zero-shot generation. Make-A-Video was preferred over VDM by 84.4% of raters for quality and 78.1% for faithfulness on the 28 prompts from VDM's website, without cherry-picked outputs. Across DrawBench and the authors' evaluation set with Chinese and English prompts, Make-A-Video surpassed CogVideo with preference rates between 68.8% and 77.2% for both quality and faithfulness.

Make-A-Video achieves a new state-of-the-art in zero-shot video generation, significantly outperforming prior models like CogVideo and even fully supervised methods on MSR-VTT and UCF-101 with substantially better FID, CLIPSIM, and FVD scores, while using only a single sample per prompt. Fine-tuning further improves coherence, setting a new state-of-the-art on UCF-101. Human evaluations confirm a strong preference for Make-A-Video over VDM and CogVideo across multiple benchmarks and languages for both visual quality and text-video faithfulness.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp