Sound Event Detection MIT/ast-finetuned-audioset-10-10-0.4593 Audio Classification • 86.6M • Updated Sep 6, 2023 • 532k • 360 laion/larger_clap_music_and_speech Feature Extraction • Updated Oct 31, 2023 • 31.7k • • 41
MIT/ast-finetuned-audioset-10-10-0.4593 Audio Classification • 86.6M • Updated Sep 6, 2023 • 532k • 360
Speaker Diarization and Transcripts openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 6.9M • • 3.13k pyannote/segmentation Voice Activity Detection • Updated May 10, 2024 • 2.73M • 681 pyannote/segmentation-3.0 Voice Activity Detection • Updated May 10, 2024 • 6.47M • 1.26k pyannote/speaker-diarization-3.1 Automatic Speech Recognition • Updated May 10, 2024 • 8.25M • 2.59k
openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 6.9M • • 3.13k
Sound Event Detection MIT/ast-finetuned-audioset-10-10-0.4593 Audio Classification • 86.6M • Updated Sep 6, 2023 • 532k • 360 laion/larger_clap_music_and_speech Feature Extraction • Updated Oct 31, 2023 • 31.7k • • 41
MIT/ast-finetuned-audioset-10-10-0.4593 Audio Classification • 86.6M • Updated Sep 6, 2023 • 532k • 360
Speaker Diarization and Transcripts openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 6.9M • • 3.13k pyannote/segmentation Voice Activity Detection • Updated May 10, 2024 • 2.73M • 681 pyannote/segmentation-3.0 Voice Activity Detection • Updated May 10, 2024 • 6.47M • 1.26k pyannote/speaker-diarization-3.1 Automatic Speech Recognition • Updated May 10, 2024 • 8.25M • 2.59k
openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 6.9M • • 3.13k