Skip to main content
Microsoft Foundry
bosonai-higgs-audio-3-instruct

bosonai-higgs-audio-3-instruct

14B audio-in instruct LLM with voice-agent reflexes — audio-native tool calling, multi-turn state tracking, and interruption-aware instruction following — with text-model-grade IF.
BosonAI
Version: 1

Higgs Audio v3 Instruct is Boson AI’s production-quality, audio-native instruct model: a 14B instruction-tuned LLM that can understand audio, text, or both, and generate high-quality, instruction-following text responses. It functions as a strong text LLM when given text alone, while bringing native audio understanding to speech and multimodal audio-text inputs.

Unlike today’s omni-chat audio LLMs — such as GPT-4o audio mode, Gemini 2.5 audio, and Qwen2.5-Omni — Higgs Audio v3 Instruct is fine-tuned specifically for voice-agent reflexes: audio-native tool calling, multi-turn state tracking, and interruption-aware instruction following. These capabilities are trained directly into the model weights, rather than prompt-engineered on top.

Key Capabilities

  1. 14B Audio-in LLM with Text-Model-Grade Instruction Following:
    On the release checkpoint: IFEval 85.5, IFBench 30.4, MultiChallenge 27.6, MultiChallenge-Audio 31.3. A 14B audio-in model is now competitive with strong text models on instruction following — the unlock for voice agents that previously had to choose between audio understanding and IF.

  2. Audio-Native Tool Calling:
    Function-calling is part of the model's behavior, not a prompt hack. The API supports standard OpenAI-style function calling. The release checkpoint reaches 22.1 success / 55.2 call_acc on the audio-converted ComplexFuncBench — the strongest tool-use surface we have shipped on a 14B audio LLM.

  3. Voice-Agent Reflexes — Multi-Turn, Interruption-Aware:
    The model holds the instruction frame across turns and interruptions, and stays in flow rather than resetting after every turn. AudioMultiChallenge 27.9, Interruption-v2 (true-follow / re-query / resume) 46.1 / 69.9 / 88.3.

  4. Streaming-Friendly Audio Path:
    Chunk-prefill audio input — VAD-segmented at up to 4-second chunks, 16 kHz, robust to noise and accent. ASR / AST is supported as an inherited capability, not the product.


Model Details

  • Model Type: Multimodal Audio and Text (Audio/Text-in, Text-out Instruct LLM)
  • Tensor Type: BF16
  • Parameters: ~14B
  • Tasks: Audio-Native Instruction Following, Audio-Native Tool / Function Calling, Multi-Turn Audio Chat, Interruption Handling, ASR / AST (inherited)
  • Language Support: English, Chinese (primary); additional languages supported by the underlying audio encoder, with instruction-following primarily validated on English and Chinese.
  • Developed by: Boson AI

Quick facts

PublisherBosonAI
TypeChat completion
LifecycleGenerally available (GA)
Input typetext, audio
Output typetext
Context window32768
Token limits4096 output
PricingUnit price varies depending on your deployment type