Mixtral of experts - Mistral AI

Mixtral of experts - Mistral AI

Mixtral 8x7B is an open-weight large language model from Mistral AI, released as a sparse mixture-of-experts decoder with Apache 2.0 licensing. You can download the weights for self-hosted inference, fine-tune them on your data, or call the model through Mistral's hosted API. It targets general text generation, code writing, and instruction following.

Most open models at this performance tier run dense architectures where every parameter activates on every token. Mixtral routes each token through only two of eight expert feedforward layers at every decoder block, combining their outputs additively. That sparse routing delivers 46.7 billion total parameters while activating just 12.9 billion per token, which Mistral reports yields 6x faster inference than Llama 2 70B on most benchmarks.

The model handles a 32k token context and performs well on multilingual tasks across English, French, Italian, German, and Spanish. Code generation is a documented strength, and the separate Mixtral 8x7B Instruct variant reaches 8.30 on MT-Bench after supervised fine-tuning and direct preference optimization.

ML engineers, researchers, and product teams use Mixtral when they want open weights with a permissive license and competitive performance against closed models like GPT-3.5. Self-hosters can deploy through vLLM with Megablocks CUDA kernels or Skypilot on cloud instances, while API users pay $0.70 per million input and output tokens on the open-mixtral-8x7b endpoint.

Top Features:
  1. Router activates 2 of 8 expert layers per token, using 12.9B of 46.7B total parameters

  2. 32k token context window for long documents and multi-turn conversations

  3. Apache 2.0 open-weight license for research and commercial self-hosting

  4. Supports English, French, Italian, German, and Spanish

  5. API endpoint open-mixtral-8x7b at $0.70 per million input and output tokens

  6. Mixtral 8x7B Instruct scores 8.30 on MT-Bench after DPO fine-tuning

  7. Deployable via vLLM with Megablocks CUDA kernels or Skypilot cloud endpoints

Pros:
  1. Apache 2.0 open weights let you self-host without vendor lock-in.

  2. Sparse routing activates 12.9B of 46.7B parameters per token, cutting inference cost versus dense 70B models.

  3. 32k token context window handles long documents and extended chat sessions.

  4. Five-language support across English, French, Italian, German, and Spanish.

  5. Documented vLLM and Skypilot deployment paths for production self-hosting.

Cons:
  1. Mistral's current model lineup has moved past Mixtral to newer Small and Medium releases.

  2. Hosted API access requires a Mistral Studio account and pay-per-token billing.

  3. Self-hosting needs GPU infrastructure with CUDA support for efficient inference.

FAQs:

What is Mixtral 8x7B?

Mixtral 8x7B is Mistral AI's open-weight sparse mixture-of-experts language model with 46.7 billion total parameters and Apache 2.0 licensing. At each layer, a router selects two of eight expert feedforward groups to process each token, activating 12.9 billion parameters per token for faster inference than dense models of similar quality.

How much does Mixtral 8x7B cost on the Mistral API?

On Mistral Studio, Mixtral 8x7B is available at the open-mixtral-8x7b endpoint for $0.70 per million input tokens and $0.70 per million output tokens. Batch processing receives a 50% discount, and cached input tokens receive a 90% discount on input costs.

Can you self-host Mixtral 8x7B?

Yes. Mixtral 8x7B ships with open weights under the Apache 2.0 license, so you can download and run it on your own GPU infrastructure. Mistral documents deployment through vLLM with Megablocks CUDA kernels, and Skypilot supports launching vLLM endpoints on cloud instances.

What languages does Mixtral 8x7B support?

Mixtral 8x7B handles English, French, Italian, German, and Spanish. Mistral's release notes describe strong multilingual benchmark performance across all five languages, making it useful for European-language text generation and translation workflows.

How does Mixtral 8x7B compare to Llama 2 70B?

Mistral reports that Mixtral 8x7B matches or outperforms Llama 2 70B on most standard benchmarks while running 6x faster at inference time. The sparse architecture activates only 12.9 billion of its 46.7 billion parameters per token, which cuts compute cost relative to dense 70B models.

What license does Mixtral 8x7B use?

Mixtral 8x7B is released under the Apache 2.0 license with open weights available for download. Mistral describes it as the strongest open-weight model with a permissive license at launch, suitable for both research and commercial self-hosted deployments.

Pricing:

Freemium

Tags:

Mixture of Experts
Open Weights
Apache 2.0
Decoder Model
Multilingual
Sparse Model
vLLM
Sparse Mixture-of-Experts

Tech used:

Netlify
Google Tag Manager
Font Awesome
Discord
Tailwind CSS

Reviews:

Give your opinion on Mixtral of experts - Mistral AI :-

Overall rating

Join thousands of AI enthusiasts in the World of AI!

Best Free Mixtral of experts - Mistral AI Alternatives (and Paid)

By Rishit