Mercury-2

Mercury-2

State-of-the-art diffusion based reasoning model.
Inception
Version: 1

Mercury 2 is the fastest reasoning LLM and the first reasoning dLLM — delivering 5–10x faster inference than speed-optimized autoregressive models like Claude 4.5 Haiku and GPT-5 Mini. Built on the diffusion LLM architecture Inception pioneered, Mercury 2 generates output through parallel refinement rather than sequential decoding, achieving ~1,000 tokens per second on NVIDIA GPUs. It's production-grade, API-accessible, and priced at a fraction of comparable models.

Mercury 2 is OpenAI client API compatible and supports a 128K context length, tool calling, structured outputs - plus a rich set of tuning options to tailor deployments to your needs.

Quick facts

Model providerInception
TypeChat completion
LifecycleGenerally available (GA)
Input typetext
Output typetext
Context window128k
Token limits50000 output
PricingUnit price varies depending on your deployment type