FW-DeepSeek-V4.1-Flash
Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.
About this model
DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text and generates text. Fireworks supports a 1,048,576-token context window and function calling. Its Causal Encoder-Decoder (CED) architecture activates 8B parameters per token during prefill and 16B during decode. Fireworks describes a KV cache footprint of roughly one quarter of DeepSeek-V4-Flash, supporting cost-efficient agentic workloads.
Key model capabilities
- Native image and text input with text output
- 552B backbone parameters with 8B active during prefill and 16B during decode
- Causal Encoder-Decoder architecture for input-heavy agentic workloads
- 1,048,576-token context window on Fireworks
- Compressed Sparse Attention 2 (CSA2) for efficient long-context processing
- DSpark speculative decoding
- Reasoning capabilities
- Function calling and streaming responses