Skip to main content
Microsoft Foundry
Higgs-Audio-v2.5

Higgs-Audio-v2.5

1B parameter high-fidelity audio generation model.
BosonAI
Version: 2

Higgs Audio v2.5 is the latest iteration of the Higgs Audio series, succeeding Higgs Audio v2 . This release represents a paradigm shift toward efficient high-fidelity synthesis.

We have distilled the architecture into a 1B parameter model that outperforms its 3B predecessor. This efficiency is achieved through a new alignment strategy utilizing Group Relative Policy Optimization (GRPO) on our curated "Voice Bank" dataset, alongside enhanced voice cloning and granular style control.

Key Improvements (v2.5 vs v2.0)

Higgs Audio v2.5 introduces four major architectural and data-centric improvements:

  1. Lightweight Architecture (3B $\rightarrow$ 1B):
    We reduced the model size from 3 billion to 1 billion parameters. This results in significantly faster inference speeds and lower hardware requirements while maintaining (and often exceeding) the audio quality of the v2 model.

  2. Voice Bank & GRPO Alignment:
    We curated Voice Bank, a high quality self-cleaned voice dataset from public source. Then we aligned the model through utilizing GRPO (Group Relative Policy Optimization). This allowed us to optimize the model for English (En), Chinese (Zh), Korean (Ko), and Japanese (Ja), while also significantly boosting zero-shot generalization to other major global languages.

  3. Enhanced Voice Cloning:
    The pretraining objective was refined to improve speaker consistency. Higgs Audio v2.5 captures subtle timbre and prosody details from shorter reference audio clips compared to v2.

  4. Expressiveness Control Tags:
    We introduced explicit control tags to dictate speech style. Users can now guide the generation with tags (<|higher_expressiveness|>, <|lower_expressiveness|>) to achieve context-aware prosody.


Model Details

  • Model Type: Autoregressive Audio Transformer
  • Tensor Type: F16
  • Parameters: 1B
  • Language Support:
    • Primary (GRPO-Aligned): English, Chinese, Korean, Japanese.
    • Secondary: Generalizes to major global languages.
  • Developed by: Boson AI

Quick facts

Model providerBosonAI
TypeAudio generation
LifecycleGenerally available (GA)
Input typetext, audio
Output typeaudio
Context window8192
Token limits4096 output
PricingUnit price varies depending on your deployment type