Lumiere
Lumiere is a text-to-video diffusion model from Google Research built to synthesize videos with realistic, diverse, and coherent motion. The project page showcases sample outputs for text-to-video, image-to-video, stylized generation, video stylization, cinemagraphs, and video inpainting.
Its Space-Time U-Net architecture generates an entire video clip in one model pass, rather than creating distant keyframes and filling gaps with temporal super-resolution. The model uses spatial and temporal down- and up-sampling and builds on a pre-trained text-to-image diffusion backbone to produce full-frame-rate video across multiple space-time scales.
The site is a research demo and paper companion, not a consumer app. It is aimed at researchers, engineers, and creators who want to study state-of-the-art video diffusion and see what the model can do across generation and editing tasks.
Turns text prompts into full motion clips with dozens of sample outputs on the demo page
Animates a still image into video when you pair it with a short text description
Stylized generation copies a reference image style using fine-tuned text-to-image weights
Video stylization applies text-based image editing methods frame by frame for consistent looks
Cinemagraph mode animates only a masked region while the rest of the image stays still
Video inpainting fills masked areas or edits clothing and accessories from text prompts
Generates full video duration in one pass with a Space-Time U-Net instead of keyframe stitching
Covers multiple tasks on one model: text-to-video, image-to-video, inpainting, and stylization
Demo gallery shows a wide range of prompts, styles, and editing examples with visible inputs
Published research paper and author credits make the technical approach easy to verify
No public interface to generate custom videos; the site is a research showcase only
Model outputs shown at low resolution per the paper's full-frame-rate, low-resolution design
No pricing, support channel, or product roadmap for teams looking to deploy it in production
Who developed Lumiere?
Lumiere was developed by researchers at Google Research, with collaborators from the Weizmann Institute, Tel-Aviv University, and the Technion. The project page lists the full author team and credits Google Research as the primary institution.
Is Lumiere available as a public app or API?
No. Lumiere is presented as a Google Research project with a demo gallery and paper link. The site does not offer sign-up, downloads, or an API for generating your own videos.
What architecture does Lumiere use?
Lumiere uses a Space-Time U-Net that generates the full temporal duration of a video in a single pass. It applies spatial and temporal down- and up-sampling and relies on a pre-trained text-to-image diffusion model as its backbone.
Can Lumiere turn an image into a video?
Yes. Lumiere supports image-to-video generation. The demo page shows still images paired with text prompts that are animated into short video clips.
Does Lumiere support video inpainting and editing?
Yes. Lumiere demonstrates video inpainting that fills masked regions in footage and text-driven edits such as changing clothing, accessories, or poses within a video sequence.
Where can I read the Lumiere research paper?
The Lumiere paper is published on arXiv at https://arxiv.org/abs/2401.12945. The project homepage links directly to that paper under Read Paper.

