Skip to main content
Microsoft Foundry
NVIDIA-Nemotron-Nano-12B-v2-VL-NIM-microservice

NVIDIA-Nemotron-Nano-12B-v2-VL-NIM-microservice

Nvidia
Version: 1

Description

The NVIDIA Nemotron Nano 12B v2 VL NIM microservice enables multi-image reasoning and video understanding, along with strong document intelligence, visual Q&A and summarization capabilities. This model is ready for commercial use.

Nemotron Nano 12B V2 VL is a model for multi-modal document intelligence. It would be used by individuals or businesses that need to process documents such as invoices, receipts, and manuals. The model is capable of handling multiple images of documents, up to four images at a resolution of 1k x 2k each, along with a long text prompt. The expected use is for tasks like summarization and Visual Question Answering (VQA). The model is also expected to have a significant advantage in throughput.

Input

Type(s): Image, Video, Text
Format: Image (PNG, JPG), Video (MP4, MKV, FLV, 3GP), Text (String)
Parameters: Image (2D), Video (3D), Text (1D)

Other Properties Related to Input:

  • Input Images Supported: 4
  • Language Supported: English only
  • Input + Output Token: 128K
  • Minimum Resolution: 32 * 32 pixels
  • Maximum Resolution: Determined by a 12-tile layout constraint, with each tile being 512 X 512 pixels. This supports aspect ratios such as:
    • 4 X 3 layout: up to 2048 X 1536 pixels
    • 3 X 4 layout: up to 1536 X 2048 pixels
    • 2 X 6 layout: up to 1024 X 3072 pixels
    • 6 X 2 layout: up to 3072 X 1024 pixels
    • Other configurations allowed, provided total tiles ≤ 12
  • Channel Count: 3 channels (RGB)
  • Alpha Channel: Not supported (no transparency)
  • Frames: 2 FPS with min of 8 frames and max of 128 frames

Output

Type(s): Text
Format: String
Parameters: 1D

Other Properties Related to Output:

Input + Output Token: 128K

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

NVIDIA AI Enterprise
NVIDIA AI Enterprise is an end-to-end, cloud-native software platform that accelerates data science pipelines and streamlines development and deployment of production-grade co-pilots and other generative AI applications. Easy-to-use microservices provide optimized model performance with enterprise-grade security, support, and stability to ensure a smooth transition from prototype to production for enterprises that run their businesses on AI.

Quick facts

Model providerNvidia
TypeChat completion
LifecycleGenerally available (GA)
Input typetext, image, video
Output typetext