Model Gallery

14 models from 1 repositories

Filter by type:

Filter by tags:

parakeet-cpp-nemotron-3-diarization
Nemotron-3-Diarization (Sortformer), Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Speaker diarization only: served through /v1/audio/diarization, returns per-segment start, end and speaker label ("0", "1", ...). It does not transcribe; pair it with an ASR model and set asr_model to get speaker-attributed text from the same call. num_speakers, min_speakers, max_speakers and clustering_threshold are not supported by Sortformer and are ignored.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-nemotron-3-diarization-asr
Nemotron-3-Diarization (Sortformer) paired with the Parakeet TDT+CTC 110M ASR model through the asr_model option, both Q8_0/F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Served through /v1/audio/diarization with include_text: each speaker segment comes back with its transcribed text in one call. Diarization model is OpenMDW-1.1, ASR model is CC-BY-4.0.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-nemotron-3-diarization-speakers
Nemotron-3-Diarization (Sortformer) with WeSpeaker ResNet34 speaker identification, for the parakeet-cpp backend. Speakers you register with /v1/voice/register (using the voice-detect-wespeaker-resnet34 model) come back by name in /v1/audio/diarization, next to the SPEAKER_NN label. Speakers that are not registered keep only their SPEAKER_NN label. The diarization model is OpenMDW-1.1, the speaker model is CC-BY-4.0. Naming was measured on one two-voice fixture only; check the threshold on your own audio.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-nemotron-3-diarization-asr-speakers
Nemotron-3-Diarization (Sortformer) paired with the Parakeet TDT+CTC 110M ASR model through the asr_model option, both Q8_0/F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Served through /v1/audio/diarization with include_text: each speaker segment comes back with its transcribed text in one call. Diarization model is OpenMDW-1.1, ASR model is CC-BY-4.0. Also loads WeSpeaker ResNet34 (CC-BY-4.0) through the speaker_model option: speakers registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) come back by name, next to the SPEAKER_NN label.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-realtime-scene-speakers
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0, WeSpeaker ResNet34 CC-BY-4.0. Also loads WeSpeaker ResNet34 through the speaker_model option, so live speaker segments carry the name of a voice registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) once the speaker is identified.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene-tdt
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options: one parakeet-cpp backend transcribes, labels speakers and tags sound events. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). TDT is not a streaming model, so in a realtime pipeline use it with server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments (conversation.item.input_audio_transcription.segment, with text) and sound tags (conversation.item.sound_detection). Also labels speakers on /v1/audio/transcriptions. Speaker labels are per turn. License per model: transcription model CC-BY-4.0, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime-scene-base
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Base (86M, the largest CED; more confident sound tags than CED-Tiny at a small extra cost: on CPU the live diarization + sound stream runs at 0.125 of real time against 0.103 with CED-Tiny) through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Base Apache-2.0.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-realtime-scene-tdt-base
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with Nemotron-3-Diarization and CED-Base (86M, the largest CED; about 0.03 s of CPU per second of audio per committed turn, against 0.005 for CED-Tiny) through the diarization_model and sound_model options: one parakeet-cpp backend transcribes, labels speakers and tags sound events. GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). TDT is not a streaming model, so in a realtime pipeline use it with server_vad: set it as both transcription and sound_detection and turn on pipeline.diarization, and each committed turn gets speaker segments (conversation.item.input_audio_transcription.segment, with text) and sound tags (conversation.item.sound_detection). Also labels speakers on /v1/audio/transcriptions. Speaker labels are per turn. License per model: transcription model CC-BY-4.0, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: cc-by-4.0

audio-cpp-sortformer-diarization
Sortformer Diarization 4-speaker v1 (audio.cpp, Q8_0) - speaker diarization for up to four speakers, served by the audio-cpp backend through /v1/audio/diarization. Returns per-segment start, end and speaker label; it does not transcribe, so pair it with an ASR model for text. Q8_0 because upstream records it as a clean Pass, the same as 16-bit, at two thirds of the size. Licensing: the base checkpoint nvidia/diar_sortformer_4spk-v1 is CC BY-NC 4.0, so commercial use is not permitted.

Repository: localaiLicense: cc-by-nc-4.0

nemo-speech-cpp-sortformer-diarization-v2
Streaming Sortformer Diarization 4-speaker v2 (nemo-speech-cpp, Q8_0) - speaker diarization for up to four speakers, served by the nemo-speech-cpp backend through /v1/audio/diarization. Returns per-segment start, end and speaker label; it does not transcribe, so pair it with an ASR model for text. This is the streaming variant: use it when you cannot wait for the full recording. For offline batch work the non-streaming v1 is faster and more accurate.

Repository: localaiLicense: cc-by-4.0

nemo-speech-cpp-nemotron-3.5-asr-streaming
Nemotron 3.5 ASR Streaming 0.6B (nemo-speech-cpp, Q8_0) - multilingual streaming speech to text for 40 language-locales, with native punctuation and capitalization. Served by the nemo-speech-cpp backend through /v1/audio/transcriptions and the streaming transcription path. A standalone ASR entry. For per-word speaker tags, attach the sortformer diarization model with the diar_model option.

Repository: localaiLicense: openmdw-1.1

nemo-speech-cpp-nemotron-3.5-asr-streaming-diarized
Nemotron 3.5 ASR Streaming 0.6B with streaming Sortformer Diarization 4-speaker v2 (nemo-speech-cpp, Q8_0). Multilingual streaming ASR with per-word speaker tags: the ASR model transcribes and the attached sortformer model labels each word with its speaker. Served through /v1/audio/transcriptions; segments are cut at each speaker change. Q8_0 for both models. The diarization model is the streaming v2 variant, for use when you cannot wait for the full recording.

Repository: localaiLicense: openmdw-1.1

nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized
Parakeet TDT 0.6B v3 with streaming Sortformer Diarization 4-speaker v2 (nemo-speech-cpp, Q8_0). Multilingual streaming ASR with per-word speaker tags: the ASR model transcribes and the attached sortformer model labels each word with its speaker. Served through /v1/audio/transcriptions; segments are cut at each speaker change. Q8_0 for both models. The diarization model is the streaming v2 variant, for use when you cannot wait for the full recording.

Repository: localaiLicense: cc-by-4.0