Skip to main content
Microsoft Foundry
FW-Inkling

FW-Inkling

Inkling is a 975B-parameter sparse Mixture-of-Experts model from Thinking Machines Lab with 41B active parameters, controllable thinking effort, and a 1M-token context window.
Fireworks
Version: 1

Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.

About this model

Inkling is Thinking Machines Lab's first open-weights model, a 975B-parameter sparse Mixture-of-Experts model with 41B active parameters and a 1M-token context window. The underlying model was natively trained across text, image, and audio, while the current Fireworks endpoint supports text input only.

Key model capabilities

  • Text input
  • 1M-token context window
  • Sparse Mixture-of-Experts architecture with 975B total parameters and 41B active parameters
  • Hybrid local and global attention
  • Function calling and serverless inference support
  • Agentic and reasoning capabilities with controllable thinking effort

Quick facts

Model providerFireworks
TypeChat completion
LifecycleGenerally available (GA)
Hosted onFireworks infrastructure
Input typetext
Output typetext
Context window1048.576k