Skip to main content
Microsoft Foundry
FW-GLM-5.2-Fast

FW-GLM-5.2-Fast

GLM 5.2 Fast is Fireworks' high-throughput serverless path for GLM 5.2, using the same weights and quality as Standard with the full 1M-token context and about 2x Standard throughput.
Fireworks
Version: 1

Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.

About this model

GLM 5.2 Fast delivers the same GLM 5.2 model quality and capabilities as the standard deployment while providing approximately 2× higher token-generation throughput on Fireworks' serverless infrastructure. It retains key features of the base model, including the 1M-token context window, long-context support, structured outputs, tool calling, and prompt caching, with no model-weight changes or quality tradeoffs reported by Fireworks. Compared with the standard deployment, GLM 5.2 Fast is optimized for lower response completion times and improved agent workflow efficiency, making it particularly well suited for coding and long-context agentic applications.

Key model capabilities

  • Same GLM 5.2 weights and quality as the Standard path
  • About 2x Standard generated-token throughput on the same workload
  • Full 1M-token context window
  • Same structured-output, function-calling, and parsing contract as Standard
  • Coding and long-horizon agentic workloads
  • Reasoning with flexible effort levels
  • Streaming and tool use

Quick facts

PublisherFireworks
AuthorzAI
TypeChat completion
LifecycleGenerally available (GA)
Hosted onFireworks infrastructure
Input typetext
Output typetext
Context window1048.576k