Skip to main content
Microsoft Foundry
Phi-4-multimodal-instruct

Phi-4-multimodal-instruct

First small multimodal model to have 3 modality inputs (text, audio, image), excelling in quality and efficiency
Microsoft
Version: 2

Models from Microsoft, Partners, and Community models are a select portfolio of curated models both general-purpose and niche models across diverse scenarios by developed by Microsoft teams, partners, and community contributors

  • Managed by Microsoft: Purchase and manage models directly through Azure with a single license, world class support and enterprise grade Azure infrastructure
  • Validated by providers: Each model is validated and maintained by its respective provider, with Azure offering integration and deployment guidance.
  • Innovation and agility: Combines Microsoft research models with rapid, community-driven advancements.
  • Seamless Azure integration: Standard Microsoft Foundry experience, with support managed by the model provider.
  • Flexible deployment: Deployable as Managed Compute or Serverless API, based on provider preference.

Learn more about models from Microsoft, Partners, and Community

About this model

The model is intended for broad multilingual and multimodal commercial and research use. The model provides uses for general purpose AI systems and applications which require:

  • Memory/compute constrained environments
  • Latency bound scenarios
  • Strong reasoning (especially math and logic)
  • Function and tool calling
  • General image understanding
  • Optical character recognition
  • Chart and table understanding
  • Multiple image comparison
  • Multi-image or video clip summarization
  • Speech recognition
  • Speech translation
  • Speech QA
  • Speech summarization
  • Audio understanding

Key model capabilities

Phi-4-multimodal-instruct demonstrated strong performance in speech tasks:

  • Surpassed expert ASR model WhisperV3 and ST models SeamlessM4T-v2-Large in automatic speech recognition (ASR) and speech translation (ST).
  • Ranked number 1 on the Huggingface OpenASR leaderboard with a word error rate of 6.14% compared to the current best model at 6.5% as of February 18, 2025.
  • First open-sourced model capable of performing speech summarization, with performance close to GPT4o.
  • Exhibited a gap with closed models like Gemini-2.0-Flash and GPT-4o-realtime-preview on the speech QA task. Efforts are ongoing to improve this capability in future iterations.

Phi-4-multimodal-instruct can process both image and audio together. The table below shows the model quality when the input query for vision content is synthetic speech on chart/table understanding and document reasoning tasks. Compared to other state-of-the-art omni models, Phi-4-multimodal-instruct achieves stronger performance on multiple benchmarks.

BenchmarksPhi-4-multimodal-instructInternOmni-7BGemini-2.0-Flash-Lite-prv-02-05Gemini-2.0-FlashGemini-1.5-Pro
s_AI2D68.953.962.069.467.7
s_ChartQA69.056.135.551.346.9
s_DocVQA87.379.976.080.378.2
s_InfoVQA63.760.359.463.666.1
Average72.262.658.266.264.7

To understand the vision capabilities, Phi-4-multimodal-instruct was compared with a set of models over a variety of zero-shot benchmarks using an internal benchmark platform. Below is a high-level overview of the model quality on representative benchmarks:

DatasetPhi-4-multimodal-insPhi-3.5-vision-insQwen 2.5-VL-3B-insIntern VL 2.5-4BQwen 2.5-VL-7B-insIntern VL 2.5-8BGemini 2.0-Flash Lite-prv-0205Gemini2.0-FlashClaude-3.5-Sonnet-2024-10-22Gpt-4o-2024-11-20
Popular aggregated benchmark55.143.047.048.351.850.654.164.755.861.7
MMBench (dev-en)86.781.984.386.887.888.285.090.086.789.0
MMMU-Pro (std / vision)38.521.829.932.438.734.445.154.454.353.0
ScienceQA Visual (img-test)97.591.379.496.287.797.385.088.381.288.2
MathVista (testmini)62.443.960.851.267.856.757.647.256.956.1
InterGPS48.636.348.353.752.754.157.965.447.149.1
AI2D82.378.178.480.082.683.077.682.170.683.8
ChartQA81.481.880.079.185.081.073.079.078.475.1
DocVQA93.269.393.991.695.793.091.292.195.290.9
InfoVQA72.736.677.172.182.677.673.077.874.371.9
TextVQA (val)75.672.076.870.977.774.872.974.458.673.1
OCR Bench84.463.882.271.687.774.875.781.077.077.7
POPE85.686.187.989.487.589.187.588.082.686.5
BLINK61.357.048.151.255.352.559.364.056.962.4
Video MME (16 frames)55.050.856.557.358.258.758.865.560.268.2
Average72.060.968.768.873.371.170.274.369.172.4

Below are the comparison results on existing multi-image tasks. On average, Phi-4-multimodal-instruct outperforms competitor models of the same size and is competitive with much bigger models on multi-frame capabilities. BLINK is an aggregated benchmark with 14 visual tasks that humans can solve very quickly but are still hard for current multimodal LLMs.

DatasetPhi-4-multimodal-instructQwen2.5-VL-3B-InstructInternVL 2.5-4BQwen2.5-VL-7B-InstructInternVL 2.5-8BGemini-2.0-Flash-Lite-prv-02-05Gemini-2.0-FlashClaude-3.5-Sonnet-2024-10-22Gpt-4o-2024-11-20
Art Style86.358.159.865.065.076.976.968.473.5
Counting60.067.560.066.771.745.869.260.865.0
Forensic Detection90.234.822.043.937.931.874.263.671.2
Functional Correspondence30.020.026.922.327.748.553.134.642.3
IQ Test22.725.328.728.728.728.030.720.725.3
Jigsaw68.752.071.369.353.362.769.361.368.7
Multi-View Reasoning76.744.444.454.145.155.641.454.954.1
Object Localization52.555.753.355.758.263.967.258.265.6
Relative Depth69.468.568.580.676.681.572.666.173.4
Relative Reflectance26.938.838.832.838.833.634.338.138.1
Semantic Correspondence52.532.433.828.824.556.155.443.947.5
Spatial Relation72.780.486.088.886.774.179.074.883.2
Visual Correspondence67.428.539.550.044.284.991.372.782.6
Visual Similarity86.767.488.187.485.287.480.779.383.0
Overall61.348.151.255.352.559.364.056.962.4

Quick facts

PublisherMicrosoft
AuthorMicrosoft
TypeChat completion
LifecycleGenerally available (GA)
Input typeaudio, image, text
Output typetext
Context window128k
Token limits4096 output