mistralai-Mixtral-8x7B-v01

mistralai-Mixtral-8x7B-v01

Mistral AI
Version: 16

The Mixtral-8x7B-v0.1 Large Language Model (LLM) is a pretrained generative text model with 7 billion parameters.
Mixtral-8x7B-v0.1 outperforms Llama 2 70B on most benchmarks with 6x faster inference.

For full details of this model please read release blog post .

Model Architecture

Mixtral-8x7B-v0.1 is a decoder-only model with 8 distinct groups or the "experts". At every layer, for every token, a router network chooses two of these experts to process the token and combine their output additively. Mixtral has 46.7B total parameters but only uses 12.9B parameters per token using this technique. This enables the model to perform with same speed and cost as 12.9B model.

TaskUse caseDatasetPython sample (Notebook)CLI with YAML
Text GenerationSummarizationSamsum summarization_with_text_gen.ipynb text-generation.sh

Quick facts

Model providerMistral AI
TypeText generation
LifecycleGenerally available (GA)