Skip to main content
Microsoft Foundry
FW-GLM-5.3-Flash

FW-GLM-5.3-Flash

GLM 5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and 18B active parameters, a hybrid sparse and linear attention architecture, Manifold-Constrained Hyper-Connections, text and image input, function calling, and a 1,
Fireworks
Version: 1

Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.

About this model

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. It starts from a newly trained base model and uses a hybrid architecture combining sparse and linear attention to reduce long-context serving costs while preserving precise long-context capabilities. The model also uses Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency. On Fireworks, GLM 5.3 Flash supports text and image input, text output, function calling, serverless and on-demand deployment, and a 1,048,576-token context window.

Key model capabilities

  • Native multimodal support for text and image input
  • 320B total parameters with 18B active parameters
  • Hybrid sparse and linear attention architecture
  • Manifold-Constrained Hyper-Connections for scaling efficiency
  • 1,048,576-token context window on Fireworks
  • Function calling and streaming responses
  • Serverless and on-demand availability on Fireworks

Quick facts

PublisherFireworks
AuthorzAI
TypeChat completion
LifecycleGenerally available (GA)
Hosted onFireworks infrastructure
Input typetext, image
Output typetext
Context window1048.576k