parakeet-cpp-realtime-scene-tdt-base
Parakeet TDT 0.6B v3 (multilingual, 25 European languages) paired with
Nemotron-3-Diarization and CED-Base (86M, the largest CED; about 0.03 s of
CPU per second of audio per committed turn, against 0.005 for CED-Tiny)
through the diarization_model and
sound_model options: one parakeet-cpp backend transcribes, labels speakers
and tags sound events. GGUF for the parakeet-cpp backend (C++/ggml port of
NVIDIA NeMo). TDT is not a streaming model, so in a realtime pipeline use
it with server_vad: set it as both transcription and sound_detection and
turn on pipeline.diarization, and each committed turn gets speaker segments
(conversation.item.input_audio_transcription.segment, with text) and
sound tags (conversation.item.sound_detection). Also labels speakers on
/v1/audio/transcriptions. Speaker labels are per turn. License per model:
transcription model CC-BY-4.0, diarization model OpenMDW-1.1, CED-Tiny
Apache-2.0.