Model Gallery

34 models from 1 repositories

Filter by type:

Filter by tags:

nemo-parakeet-tdt-0.6b
NVIDIA NeMo Parakeet TDT 0.6B v3 is an automatic speech recognition (ASR) model from NVIDIA's NeMo toolkit. Parakeet models are state-of-the-art ASR models trained on large-scale English audio data.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt_ctc-110m
Hybrid TDT+CTC FastConformer, 110M. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-realtime_eou_120m-v1
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M. Use with streaming transcription. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-ctc-0.6b
CTC FastConformer, 0.6B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-rnnt-0.6b
RNNT FastConformer, 0.6B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt-0.6b-v2
TDT FastConformer, 0.6B (v2). F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt-0.6b-v3
TDT FastConformer, 0.6B (v3, multilingual). F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

orukeet
Orukeet is a 0.6B, 25-language fine-tune of Parakeet TDT v3 by Oruk. Q8 GGUF for the nemo-speech-cpp backend. Runs locally on CPU, with optional GPU acceleration through the backend's gpu option.

Repository: localaiLicense: cc-by-sa-4.0

parakeet-cpp-ctc-1.1b
CTC FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-rnnt-1.1b
RNNT FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt-1.1b
TDT FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-tdt_ctc-1.1b
Hybrid TDT+CTC FastConformer, 1.1B. F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo Parakeet), byte-identical to NeMo at WER 0. Faster than NeMo on CPU and GPU.

Repository: localaiLicense: cc-by-4.0

parakeet-cpp-nemotron-3.5-asr-streaming-0.6b
Multilingual (40+ locales), prompt-conditioned, cache-aware streaming FastConformer RNN-T, 0.6B. Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Byte-identical to NeMo at WER 0 offline and streaming, about 2.5x faster than NeMo on CPU with no GPU. Select a language with the request "language" field (for example en, de, es, ja-JP), or leave it empty for automatic detection. License OpenMDW-1.1.

Repository: localaiLicense: other

parakeet-cpp-nemotron-3-diarization
Nemotron-3-Diarization (Sortformer), Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Speaker diarization only: served through /v1/audio/diarization, returns per-segment start, end and speaker label ("0", "1", ...). It does not transcribe; pair it with an ASR model and set asr_model to get speaker-attributed text from the same call. num_speakers, min_speakers, max_speakers and clustering_threshold are not supported by Sortformer and are ignored.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-nemotron-3-diarization-asr
Nemotron-3-Diarization (Sortformer) paired with the Parakeet TDT+CTC 110M ASR model through the asr_model option, both Q8_0/F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Served through /v1/audio/diarization with include_text: each speaker segment comes back with its transcribed text in one call. Diarization model is OpenMDW-1.1, ASR model is CC-BY-4.0.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-nemotron-3-diarization-speakers
Nemotron-3-Diarization (Sortformer) with WeSpeaker ResNet34 speaker identification, for the parakeet-cpp backend. Speakers you register with /v1/voice/register (using the voice-detect-wespeaker-resnet34 model) come back by name in /v1/audio/diarization, next to the SPEAKER_NN label. Speakers that are not registered keep only their SPEAKER_NN label. The diarization model is OpenMDW-1.1, the speaker model is CC-BY-4.0. Naming was measured on one two-voice fixture only; check the threshold on your own audio.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-nemotron-3-diarization-asr-speakers
Nemotron-3-Diarization (Sortformer) paired with the Parakeet TDT+CTC 110M ASR model through the asr_model option, both Q8_0/F16 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Served through /v1/audio/diarization with include_text: each speaker segment comes back with its transcribed text in one call. Diarization model is OpenMDW-1.1, ASR model is CC-BY-4.0. Also loads WeSpeaker ResNet34 (CC-BY-4.0) through the speaker_model option: speakers registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) come back by name, next to the SPEAKER_NN label.

Repository: localaiLicense: openmdw-1.1

parakeet-cpp-realtime-scene-speakers
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0, WeSpeaker ResNet34 CC-BY-4.0. Also loads WeSpeaker ResNet34 through the speaker_model option, so live speaker segments carry the name of a voice registered with /v1/voice/register (voice-detect-wespeaker-resnet34 model) once the speaker is identified.

Repository: localaiLicense: nvidia-open-model-license

parakeet-cpp-ced-tiny
CED-Tiny sound event tagger, Q8_0 GGUF for the parakeet-cpp backend (C++/ggml, loaded through third_party/ced.cpp). Served through /v1/audio/classification: 10 s windows are scored and averaged over the clip, then sorted by score with threshold and top_k applied. Smallest and fastest of the CED sizes; use ced-base for higher accuracy.

Repository: localaiLicense: apache-2.0

parakeet-cpp-ced-base
CED-Base sound event tagger, Q8_0 GGUF for the parakeet-cpp backend (C++/ggml, loaded through third_party/ced.cpp). Served through /v1/audio/classification: 10 s windows are scored and averaged over the clip, then sorted by score with threshold and top_k applied. Larger and more accurate than ced-tiny, still CPU-friendly.

Repository: localaiLicense: apache-2.0

parakeet-cpp-realtime-scene
Cache-aware streaming RNNT FastConformer with end-of-utterance (EOU) detection, 120M, paired with Nemotron-3-Diarization and CED-Tiny through the diarization_model and sound_model options. F16/Q8_0 GGUF for the parakeet-cpp backend (C++/ggml port of NVIDIA NeMo). Use with streaming transcription: while a turn is live, closed speaker segments and sound events are surfaced alongside the ASR text (realtime conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events). Live speaker/sound events only fire during speech turns under semantic_vad; sounds between turns are not seen by this path. License per model: transcription model NVIDIA Open Model License, diarization model OpenMDW-1.1, CED-Tiny Apache-2.0.

Repository: localaiLicense: nvidia-open-model-license

Page 1