
Last updated 08-14-2026
Category:
Overall Rating:
5.0 🏆
Reviews:
Join thousands of AI enthusiasts in the World of AI!
Text-To-4D
Text-To-4D is a Meta AI research project (MAV3D) that generates three-dimensional dynamic scenes from text descriptions. The method uses a 4D dynamic Neural Radiance Field optimized for appearance, density, and motion consistency by querying a text-to-video diffusion model.
Unlike static 3D generators that output a single mesh, Text-To-4D produces scenes you can view from any camera angle and composite into other 3D environments. The approach needs no 3D or 4D training data; the underlying text-to-video model trains only on text-image pairs and unlabeled videos.
The project page hosts demo samples for text-to-4D prompts like "a corgi playing with a ball" and image-to-4D conversions from still photos. It is a research showcase, not a commercial product with a public API or signup. Researchers and 3D artists interested in NeRF-based dynamic scene generation can explore the paper and sample outputs on the site.
Generates 3D dynamic scenes from text prompts via MAV3D (Make-A-Video3D)
4D dynamic NeRF optimized for appearance, density, and motion consistency
View generated scenes from any camera location and angle
Image-to-4D mode converts still photos into dynamic video scenes
Trained only on text-image pairs and unlabeled videos, no 3D/4D data required
Published research paper on arXiv (2301.11280) with interactive demo samples
Pioneering approach to generating full 3D dynamic scenes from text alone
No 3D or 4D training data required for the underlying model
Output viewable from any camera angle and compositable into 3D environments
Free research demo with interactive sample outputs
Research demo only with no public generation API or tool
Paper published in 2023 with no indication of ongoing commercial development
Thin site with only demo samples, no documentation or user support
Requires technical knowledge to reproduce results from the paper
What is Text-To-4D?
Text-To-4D is a Meta AI research method called MAV3D that generates three-dimensional dynamic scenes from text descriptions. Text-To-4D uses a 4D dynamic Neural Radiance Field optimized by a text-to-video diffusion model, and the demo page shows sample outputs you can explore.
Is Text-To-4D free to use?
Text-To-4D is a free research demo hosted on GitHub Pages with no signup or payment required. There is no public API or commercial product; you can view sample generations and read the paper, but you cannot generate your own scenes through the site.
How does Text-To-4D work?
Text-To-4D builds a 4D dynamic NeRF that queries a text-to-video diffusion model for scene appearance, density, and motion. The output is a dynamic video viewable from any camera angle and compositable into 3D environments. Text-To-4D also supports image-to-4D from a single input photo.
Who created Text-To-4D?
Text-To-4D was developed by researchers at Meta AI including Uriel Singer, Shelly Sheynin, Adam Polyak, and others. The work was published on arXiv in 2023 as "Text-To-4D Dynamic Scene Generation" (arXiv:2301.11280).
Can I generate my own scenes with Text-To-4D?
The Text-To-4D demo page shows pre-generated samples like "a panda dancing" and "a space shuttle launching" but does not offer a public generation interface. Text-To-4D is a research showcase; implementing the method requires following the published paper.
What is the difference between Text-To-4D and Make-A-Video?
Make-A-Video generates 2D videos from text, while Text-To-4D (MAV3D) extends that to full 3D dynamic scenes viewable from any angle. Text-To-4D adds a 4D NeRF layer so the output can be composited into 3D environments, not just watched as flat video.
