FW-Qwen3.5-4B

FW-Qwen3.5-4B

Qwen3.5 4B is a 4B-parameter causal language model with a vision encoder, hybrid Gated DeltaNet and Gated Attention architecture, thinking mode by default, 201 language support, and a native 262K token context window.
Fireworks
Version: 1

Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.

About this model

Qwen3.5-4B is a post-trained causal language model with a vision encoder from Alibaba's Qwen3.5 family. Qwen3.5 integrates multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility. It uses a 32-layer hybrid stack with Gated DeltaNet and Gated Attention blocks, supports text and image inputs in the official model card, and supports 201 languages and dialects.

Key model capabilities

  • Unified vision-language foundation with early-fusion multimodal training
  • Hybrid Gated DeltaNet and Gated Attention architecture
  • Thinking mode by default, with configurable non-thinking mode
  • Tool use through Qwen-Agent, SGLang, and vLLM examples
  • Support for 201 languages and dialects
  • Native 262,144-token context, extensible to 1,010,000 tokens with YaRN
  • Multi-Token Prediction (MTP) support

Quick facts

Model providerFireworks
TypeChat completion
LifecycleGenerally available (GA)
Input typetext, image
Output typetext
Context window262.144k