All posts
2 min readby Romiel Inolino

Nvidia's free diarization model labels up to eight speakers live. Better call notes for your CRM

Voice AINvidiaCRMCall TranscriptionOpen Models

A transcript that says who said what is worth far more to a CRM than a wall of text.

What happened

Nvidia released Nemotron 3 Diarization, an open weight model of about 100 million parameters that identifies which speaker is talking at each moment of a conversation. According to Nvidia's Hugging Face post, it handles up to eight speakers, works on recordings and live audio, and detects overlapping speech.

Key details from the announcement:

  • It takes 16 kHz single channel audio in formats including wav, flac, opus and mp3.
  • It offers four input buffer settings, from 30.4 seconds for offline use down to 0.32 seconds for ultra low latency. Nvidia notes these figures exclude compute, network and transcription time.
  • It ranks first on VoiceArena's Diarization-Bench with a 14.72% diarization error rate, versus 19.3% for the next system.
  • Paired with a speech recognition model such as Parakeet, it produces transcripts with speaker labels.

The limits matter. Labels are generic, like speaker_2, not names. Nvidia says noise, reverb, far field recording and long conversations can increase errors, and the maximum is eight speakers. It was trained on conversations in 21 languages, though the main benchmarks focus on English. The license is the OpenMDW License Agreement, version 1.1.

My take

The weak point in most call to CRM automations is not the transcript. It is attribution. If the summary cannot tell the prospect's objection from the rep's pitch, the "next steps" field fills up with nonsense.

A pipeline I would test with this:

  1. Recording lands from Zoom, a dialer or GoHighLevel call tracking.
  2. Diarization plus transcription produce speaker labelled text.
  3. A simple rule maps speakers to people, for example whichever label matches the rep's known channel or opening.
  4. An LLM step extracts objections, budget signals and commitments by speaker and writes them to the deal record.

Because the model is small and open, self hosting is a realistic option for teams that do not want call audio leaving their own infrastructure. Test it on your own recordings first, especially noisy ones.

More posts