CM3leon
CM3leon is a multimodal foundation model from Meta AI, introduced in July 2023 on ai.meta.com. Pronounced like "chameleon," it combines text-to-image and image-to-text generation in one architecture instead of separate specialized models.
Meta trained CM3leon with a recipe adapted from text-only language models: large-scale retrieval-augmented pre-training followed by multitask supervised fine-tuning. The causal masked mixed-modal (CM3) design lets the model generate sequences of text and images conditioned on mixed image and text inputs, expanding beyond earlier models that handled only one direction.
On the MS-COCO text-to-image benchmark, CM3leon reported a zero-shot FID score of 4.88 while using five times less compute than prior transformer-based text-to-image methods. Meta positions it for researchers and teams working on unified vision-language generation, editing, captioning, and visual question answering.
One foundation model covers text-to-image and image-to-text in both directions
Hit a zero-shot MS-COCO FID of 4.88, beating Google's Parti on Meta's published benchmark
Trained with five times less compute than earlier transformer text-to-image models
Text-guided image editing runs on the same model, not a separate fine-tuned editor
Multitask instruction tuning covers captioning, VQA, editing, and conditional generation
Retrieval-augmented pre-training plus SFT on a three-billion-token text dataset
Unifies image generation and image understanding in one foundation model.
Published strong MS-COCO benchmark numbers with lower training compute than prior transformer approaches.
Supports editing, captioning, and VQA from the same weights through multitask instruction tuning.
Documented only as a Meta research announcement with no public CM3leon product or API listed.
Initial blog post dates to July 2023 with limited same-domain documentation beyond ai.meta.com.
No pricing or consumer access path is published on the official announcement page.
What is CM3leon?
CM3leon is a multimodal foundation model from Meta AI that handles both text-to-image and image-to-text generation. Meta introduced it in a July 2023 research blog post on ai.meta.com.
How do you pronounce CM3leon?
Meta says CM3leon is pronounced like "chameleon." The name reflects its CM3 (causal masked mixed-modal) architecture.
What tasks can CM3leon perform?
CM3leon supports text-to-image generation, text-guided image editing, image captioning, image-to-text generation, visual question answering, and conditional image generation. Meta documents all of these running from a single model.
Does CM3leon have a public API or app?
Meta's CM3leon announcement is a research showcase on ai.meta.com. The blog post describes model capabilities and benchmark results but does not list a standalone public product, pricing page, or developer API for CM3leon.
How does CM3leon compare to other text-to-image models?
On zero-shot MS-COCO, Meta reports CM3leon achieved an FID score of 4.88, which it describes as state of the art at announcement time and ahead of Google's Parti model on that benchmark.
Who developed CM3leon?
CM3leon was developed by Meta AI researchers and published on ai.meta.com as part of Meta's generative AI research program.
What dataset was CM3leon trained on?
Meta says CM3leon was trained on a licensed dataset rather than the web-scraped mixes many image models use. Meta highlights that choice as evidence strong performance is possible with a different data distribution from prior models.

