embed-v-4-0

embed-v-4-0

Embed 4 transforms texts and images into numerical vectors
Cohere
Direct from Azure
Version: 6

Direct from Azure models are a select portfolio curated for their market-differentiated capabilities:

  • Secure and managed by Microsoft: Purchase and manage models directly through Azure with a single license, consistent support, and no third-party dependencies, backed by Azure's enterprise-grade infrastructure.
  • Streamlined operations: Benefit from unified billing, governance, and seamless PTU portability across models hosted on Azure - all part of Microsoft Foundry.
  • Future-ready flexibility: Access the latest models as they become available, and easily test, deploy, or switch between them within Microsoft Foundry; reducing integration effort.
  • Cost control and optimization: Scale on demand with pay-as-you-go flexibility or reserve PTUs for predictable performance and savings.

Learn more about Direct from Azure models .

About this model

Cohere’s Embed 4 is a multilingual multimodal embedding model. It is capable of transforming different modalities such as images, texts, and interleaved images and texts into a single vector representation. Embed 4 offers state-of-the-art performance across all modalities (texts, images, interleaved texts and image) and in both English and multilingual settings.

Embed 4 supports a 128k context length and an images can have a maximum of 2MM pixels. Embed 4 is capable of vectorizing interleaved texts and images and capturing key visual features from screenshots of PDFs, slides, tables, figures, and more, thereby eliminating the need for complex document parsing. Embed 4 offers a variety of ways for compression both on the number of dimensions and the number-format precision. The model offers byte and binary quantization and matryoshka embeddings for further compression.

Key model capabilities

  • Multilingual multimodal embedding capabilities
  • Transform different modalities such as images, texts, and interleaved images and texts into a single vector representation
  • State-of-the-art performance across all modalities (texts, images, interleaved texts and image) in both English and multilingual settings
  • Support for 128k context length
  • Process images with a maximum of 2MM pixels
  • Vectorize interleaved texts and images
  • Capture key visual features from screenshots of PDFs, slides, tables, figures, and more
  • Eliminate the need for complex document parsing
  • Variety of compression options including byte and binary quantization
  • Matryoshka embeddings for further compression

Quick facts

Model providerCohere
TypeEmbeddings, Summarization
LifecycleGenerally available (GA)
Input typeimage, text
Output typeimage, text
Context window131.072k
Token limits4096 output