FW-Qwen3.5-397B-A17B

FW-Qwen3.5-397B-A17B

Qwen3.5 397B A17B is a 397B-parameter Mixture-of-Experts language model from Alibaba with 17B activated parameters, featuring Gated Delta Networks, 201 language support, and 262K native context. This is the text-only variant of the model.
Fireworks
Version: 2

Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.

About this model

Qwen3.5 397B A17B is Alibaba's flagship Mixture-of-Experts language model with 397 billion total parameters, activating 17 billion per token for efficient inference. It features a hybrid architecture combining Gated Delta Networks with sparse MoE for high-throughput inference with minimal latency. This is the text-only variant of the Qwen3.5 397B A17B model; for image understanding capabilities, see the Qwen3.5 397B A17B Vision model card. It supports 201 languages and dialects, with a native context length of 262,144 tokens extensible up to 1,010,000 tokens via YaRN.

Key model capabilities

  • Text-only language model (for image input support, see the Vision variant)
  • Mixture of Experts architecture: 397B total parameters with 17B activated per token (512 experts, 10 routed + 1 shared)
  • Gated Delta Networks combined with sparse MoE for efficient inference
  • Native 262K context, extensible to 1M+ tokens via YaRN
  • Thinking mode with configurable enable/disable
  • Multi-Token Prediction (MTP) support
  • Multilingual support (201 languages and dialects)
  • Function calling, tool use, and agentic workflows
  • Streaming support

Quick facts

Model providerFireworks
TypeChat completion
LifecycleGenerally available (GA)
Input typetext
Output typetext
Context window262.144k