Skip to main content
Microsoft Foundry
Flux.1-Kontext-pro

Flux.1-Kontext-pro

Generate and edit images through both text and image prompts. FLUX.1 Kontext is a multimodal flow matching model that enables both text-to-image generation and in-context image editing. Modify images while maintaining character consistency and performing l
Black Forest Labs
Direct from Azure
Version: 1

Direct from Azure models are a select portfolio curated for their market-differentiated capabilities:

  • Secure and managed by Microsoft: Purchase and manage models directly through Azure with a single license, consistent support, and no third-party dependencies, backed by Azure's enterprise-grade infrastructure.
  • Streamlined operations: Benefit from unified billing, governance, and seamless PTU portability across models hosted on Azure - all part of Microsoft Foundry.
  • Future-ready flexibility: Access the latest models as they become available, and easily test, deploy, or switch between them within Microsoft Foundry; reducing integration effort.
  • Cost control and optimization: Scale on demand with pay-as-you-go flexibility or reserve PTUs for predictable performance and savings.

Learn more about Direct from Azure models .

About this model

Generate and edit images through both text and image prompts. Flux.1 Kontext is a multimodal flow matching model that enables both text-to-image generation and in-context image editing. Modify images while maintaining character consistency and performing local edits up to 8x faster than other leading models.

Key model capabilities

The model provides powerful text-to-image and image-editing capabalities:

  1. Change existing images based on an edit instruction.
  2. Have character, style and object reference without any finetuning.
  3. Robust consistency allows users to refine an image through multiple successive edits with minimal visual drift.

FLUX.1 Kontext [pro] consistently ranks among the top performers across all tasks, achieving the highest scores in text editing and character preservation while consistently outperforming competing state-of-the-art models in inference speed.

In-Context Performance

We show evaluation results across six in-context image generation tasks. FLUX.1 Kontext [pro] consistently ranks among the top performers across all tasks, achieving the highest scores in Text Editing and Character Preservation.

Speed Performance

FLUX.1 Kontext models consistently achieve lower latencies than competing state-of-the-art models for both text-to-image generation (left) and image-editing (right)

FLUX.1 Kontext models demonstrate competitive performance across aesthetics, prompt following, typography, and realism benchmarks.

Genre Performance

T2I Evaluation

left: input image; middle: edit from input: "tilt her head towards the camera", right: "make her laugh"

I2I Evaluation

left: input image; middle: edit from input: "change the 'YOU HAD ME AT BEER' to 'YOU HAD ME AT CONTEXT'", right: "change the setting to a night club"

FLUX.1 Kontext exhibits some limitations in its current implementation. Excessive multi-turn editing sessions can introduce visual artifacts that degrade image quality. The model occasionally fails to follow instructions accurately, ignoring specific prompt requirements in rare cases. World knowledge remains limited, affecting the model's ability to generate contextually accurate content. Additionally, the distillation process can introduce visual artifacts that impact output fidelity.

Failure

Illustration of a FLUX.1 Kontext failure case: After six iterative edits, the generation is visually degraded and contains visible artifacts.

Quick facts

PublisherBlack Forest Labs
AuthorBlack Forest Labs
TypeText to image, Image to image
LifecycleGenerally available (GA)
Hosted onAzure
Input typetext, image
Output typeimage
Context window131.072k
Token limits4096 output