Automatic Speech Recognition
NeMo
PyTorch
speaker-diarization
speaker-recognition
speech
audio
Transformer
FastConformer
Conformer
NEST
NeMo
Eval Results (legacy)
Instructions to use nvidia/diar_streaming_sortformer_4spk-v2.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use nvidia/diar_streaming_sortformer_4spk-v2.1 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("nvidia/diar_streaming_sortformer_4spk-v2.1") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -291,6 +291,7 @@ This model is a streaming version of Sortformer diarizer. [Sortformer](https://a
|
|
| 291 |
|
| 292 |
Sortformer resolves permutation problem in diarization following the arrival-time order of the speech segments from each speaker.
|
| 293 |
|
|
|
|
| 294 |
## Discover more from NVIDIA:
|
| 295 |
For documentation, deployment guides, enterprise-ready APIs, and the latest open models—including Nemotron and other cutting-edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at [developer.nvidia.com](https://developer.nvidia.com/).
|
| 296 |
Join the community to access tools, support, and resources to accelerate your development with NVIDIA’s NeMo, Riva, NIM, and foundation models.<br>
|
|
|
|
| 291 |
|
| 292 |
Sortformer resolves permutation problem in diarization following the arrival-time order of the speech segments from each speaker.
|
| 293 |
|
| 294 |
+
This speaker diarization model can be used to enable the [NeMo Voice Agent](https://github.com/NVIDIA-NeMo/NeMo/tree/main/examples/voice_agent) to recognize speakers in conversations. See the [NeMo Voice Agent](https://github.com/NVIDIA-NeMo/NeMo/tree/main/examples/voice_agent) and the [YAML configuration](https://github.com/NVIDIA-NeMo/NeMo/blob/316aea20fc2aac4e41bc08123cc77118a9f0b82a/examples/voice_agent/server/server_configs/default.yaml#L25) for more details.
|
| 295 |
## Discover more from NVIDIA:
|
| 296 |
For documentation, deployment guides, enterprise-ready APIs, and the latest open models—including Nemotron and other cutting-edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at [developer.nvidia.com](https://developer.nvidia.com/).
|
| 297 |
Join the community to access tools, support, and resources to accelerate your development with NVIDIA’s NeMo, Riva, NIM, and foundation models.<br>
|