Site map
All 1014 pages on this site, grouped. Search engines can use the machine-readable version at /sitemap.xml.
Datasets 469
- /datasets/
- /datasets/hindi-call-center-speech/
- /datasets/arabic-msa-call-center-speech/
- /datasets/indonesian-call-center-speech/
- /datasets/thai-call-center-speech/
- /datasets/turkish-call-center-speech/
- /datasets/vietnamese-call-center-speech/
- /datasets/filipino-call-center-speech/
- /datasets/persian-call-center-speech/
- /datasets/tamil-call-center-speech/
- /datasets/telugu-call-center-speech/
- /datasets/malayalam-call-center-speech/
- /datasets/nepali-call-center-speech/
- /datasets/sinhala-call-center-speech/
- /datasets/bangla-call-center-speech/
- /datasets/khmer-call-center-speech/
- /datasets/hindi-conversational-speech/
- /datasets/arabic-msa-conversational-speech/
- /datasets/indonesian-conversational-speech/
- /datasets/thai-conversational-speech/
- /datasets/turkish-conversational-speech/
- /datasets/vietnamese-conversational-speech/
- /datasets/filipino-conversational-speech/
- /datasets/persian-conversational-speech/
- /datasets/tamil-conversational-speech/
- /datasets/telugu-conversational-speech/
- /datasets/malayalam-conversational-speech/
- /datasets/nepali-conversational-speech/
- /datasets/sinhala-conversational-speech/
- /datasets/bangla-conversational-speech/
- /datasets/khmer-conversational-speech/
- /datasets/hindi-singing-voice/
- /datasets/arabic-msa-singing-voice/
- /datasets/indonesian-singing-voice/
- /datasets/thai-singing-voice/
- /datasets/turkish-singing-voice/
- /datasets/vietnamese-singing-voice/
- /datasets/filipino-singing-voice/
- /datasets/persian-singing-voice/
- /datasets/tamil-singing-voice/
- /datasets/telugu-singing-voice/
- /datasets/malayalam-singing-voice/
- /datasets/nepali-singing-voice/
- /datasets/sinhala-singing-voice/
- /datasets/bangla-singing-voice/
- /datasets/khmer-singing-voice/
- /datasets/hindi-speech-commands/
- /datasets/arabic-msa-speech-commands/
- /datasets/indonesian-speech-commands/
- /datasets/thai-speech-commands/
- /datasets/turkish-speech-commands/
- /datasets/vietnamese-speech-commands/
- /datasets/filipino-speech-commands/
- /datasets/persian-speech-commands/
- /datasets/tamil-speech-commands/
- /datasets/telugu-speech-commands/
- /datasets/malayalam-speech-commands/
- /datasets/nepali-speech-commands/
- /datasets/sinhala-speech-commands/
- /datasets/bangla-speech-commands/
- /datasets/khmer-speech-commands/
- /datasets/hindi-read-speech/
- /datasets/arabic-msa-read-speech/
- /datasets/indonesian-read-speech/
- /datasets/thai-read-speech/
- /datasets/turkish-read-speech/
- /datasets/vietnamese-read-speech/
- /datasets/filipino-read-speech/
- /datasets/persian-read-speech/
- /datasets/tamil-read-speech/
- /datasets/telugu-read-speech/
- /datasets/malayalam-read-speech/
- /datasets/nepali-read-speech/
- /datasets/sinhala-read-speech/
- /datasets/bangla-read-speech/
- /datasets/khmer-read-speech/
- /datasets/hindi-podcast-speech/
- /datasets/arabic-msa-podcast-speech/
- /datasets/indonesian-podcast-speech/
- /datasets/thai-podcast-speech/
- /datasets/turkish-podcast-speech/
- /datasets/vietnamese-podcast-speech/
- /datasets/filipino-podcast-speech/
- /datasets/persian-podcast-speech/
- /datasets/tamil-podcast-speech/
- /datasets/telugu-podcast-speech/
- /datasets/malayalam-podcast-speech/
- /datasets/nepali-podcast-speech/
- /datasets/sinhala-podcast-speech/
- /datasets/bangla-podcast-speech/
- /datasets/khmer-podcast-speech/
- /datasets/hindi-multilingual-speech/
- /datasets/arabic-msa-multilingual-speech/
- /datasets/indonesian-multilingual-speech/
- /datasets/thai-multilingual-speech/
- /datasets/turkish-multilingual-speech/
- /datasets/vietnamese-multilingual-speech/
- /datasets/filipino-multilingual-speech/
- /datasets/persian-multilingual-speech/
- /datasets/tamil-multilingual-speech/
- /datasets/telugu-multilingual-speech/
- /datasets/malayalam-multilingual-speech/
- /datasets/nepali-multilingual-speech/
- /datasets/sinhala-multilingual-speech/
- /datasets/bangla-multilingual-speech/
- /datasets/khmer-multilingual-speech/
- /datasets/hindi-noisy-speech/
- /datasets/arabic-msa-noisy-speech/
- /datasets/indonesian-noisy-speech/
- /datasets/thai-noisy-speech/
- /datasets/turkish-noisy-speech/
- /datasets/vietnamese-noisy-speech/
- /datasets/filipino-noisy-speech/
- /datasets/persian-noisy-speech/
- /datasets/tamil-noisy-speech/
- /datasets/telugu-noisy-speech/
- /datasets/malayalam-noisy-speech/
- /datasets/nepali-noisy-speech/
- /datasets/sinhala-noisy-speech/
- /datasets/bangla-noisy-speech/
- /datasets/khmer-noisy-speech/
- /datasets/hindi-code-switching-speech/
- /datasets/arabic-msa-code-switching-speech/
- /datasets/indonesian-code-switching-speech/
- /datasets/thai-code-switching-speech/
- /datasets/turkish-code-switching-speech/
- /datasets/vietnamese-code-switching-speech/
- /datasets/filipino-code-switching-speech/
- /datasets/persian-code-switching-speech/
- /datasets/tamil-code-switching-speech/
- /datasets/telugu-code-switching-speech/
- /datasets/malayalam-code-switching-speech/
- /datasets/nepali-code-switching-speech/
- /datasets/sinhala-code-switching-speech/
- /datasets/bangla-code-switching-speech/
- /datasets/khmer-code-switching-speech/
- /datasets/hindi-children-speech/
- /datasets/arabic-msa-children-speech/
- /datasets/indonesian-children-speech/
- /datasets/thai-children-speech/
- /datasets/turkish-children-speech/
- /datasets/vietnamese-children-speech/
- /datasets/filipino-children-speech/
- /datasets/persian-children-speech/
- /datasets/tamil-children-speech/
- /datasets/telugu-children-speech/
- /datasets/malayalam-children-speech/
- /datasets/nepali-children-speech/
- /datasets/sinhala-children-speech/
- /datasets/bangla-children-speech/
- /datasets/khmer-children-speech/
- /datasets/hindi-audiobook/
- /datasets/arabic-msa-audiobook/
- /datasets/indonesian-audiobook/
- /datasets/thai-audiobook/
- /datasets/turkish-audiobook/
- /datasets/vietnamese-audiobook/
- /datasets/filipino-audiobook/
- /datasets/persian-audiobook/
- /datasets/tamil-audiobook/
- /datasets/telugu-audiobook/
- /datasets/malayalam-audiobook/
- /datasets/nepali-audiobook/
- /datasets/sinhala-audiobook/
- /datasets/bangla-audiobook/
- /datasets/khmer-audiobook/
- /datasets/hindi-short-video-speech/
- /datasets/arabic-msa-short-video-speech/
- /datasets/indonesian-short-video-speech/
- /datasets/thai-short-video-speech/
- /datasets/turkish-short-video-speech/
- /datasets/vietnamese-short-video-speech/
- /datasets/filipino-short-video-speech/
- /datasets/persian-short-video-speech/
- /datasets/tamil-short-video-speech/
- /datasets/telugu-short-video-speech/
- /datasets/malayalam-short-video-speech/
- /datasets/nepali-short-video-speech/
- /datasets/sinhala-short-video-speech/
- /datasets/bangla-short-video-speech/
- /datasets/khmer-short-video-speech/
- /datasets/hindi-live-stream-speech/
- /datasets/arabic-msa-live-stream-speech/
- /datasets/indonesian-live-stream-speech/
- /datasets/thai-live-stream-speech/
- /datasets/turkish-live-stream-speech/
- /datasets/vietnamese-live-stream-speech/
- /datasets/filipino-live-stream-speech/
- /datasets/persian-live-stream-speech/
- /datasets/tamil-live-stream-speech/
- /datasets/telugu-live-stream-speech/
- /datasets/malayalam-live-stream-speech/
- /datasets/nepali-live-stream-speech/
- /datasets/sinhala-live-stream-speech/
- /datasets/bangla-live-stream-speech/
- /datasets/khmer-live-stream-speech/
- /datasets/hindi-voice-assistant/
- /datasets/arabic-msa-voice-assistant/
- /datasets/indonesian-voice-assistant/
- /datasets/thai-voice-assistant/
- /datasets/turkish-voice-assistant/
- /datasets/vietnamese-voice-assistant/
- /datasets/filipino-voice-assistant/
- /datasets/persian-voice-assistant/
- /datasets/tamil-voice-assistant/
- /datasets/telugu-voice-assistant/
- /datasets/malayalam-voice-assistant/
- /datasets/nepali-voice-assistant/
- /datasets/sinhala-voice-assistant/
- /datasets/bangla-voice-assistant/
- /datasets/khmer-voice-assistant/
- /datasets/hindi-elderly-voice/
- /datasets/arabic-msa-elderly-voice/
- /datasets/indonesian-elderly-voice/
- /datasets/thai-elderly-voice/
- /datasets/turkish-elderly-voice/
- /datasets/vietnamese-elderly-voice/
- /datasets/filipino-elderly-voice/
- /datasets/persian-elderly-voice/
- /datasets/tamil-elderly-voice/
- /datasets/telugu-elderly-voice/
- /datasets/malayalam-elderly-voice/
- /datasets/nepali-elderly-voice/
- /datasets/sinhala-elderly-voice/
- /datasets/bangla-elderly-voice/
- /datasets/khmer-elderly-voice/
- /datasets/hindi-accented-english/
- /datasets/arabic-msa-accented-english/
- /datasets/indonesian-accented-english/
- /datasets/thai-accented-english/
- /datasets/turkish-accented-english/
- /datasets/vietnamese-accented-english/
- /datasets/filipino-accented-english/
- /datasets/persian-accented-english/
- /datasets/tamil-accented-english/
- /datasets/telugu-accented-english/
- /datasets/malayalam-accented-english/
- /datasets/nepali-accented-english/
- /datasets/sinhala-accented-english/
- /datasets/bangla-accented-english/
- /datasets/khmer-accented-english/
- /datasets/hindi-speech-translation/
- /datasets/arabic-msa-speech-translation/
- /datasets/indonesian-speech-translation/
- /datasets/thai-speech-translation/
- /datasets/turkish-speech-translation/
- /datasets/vietnamese-speech-translation/
- /datasets/filipino-speech-translation/
- /datasets/persian-speech-translation/
- /datasets/tamil-speech-translation/
- /datasets/telugu-speech-translation/
- /datasets/malayalam-speech-translation/
- /datasets/nepali-speech-translation/
- /datasets/sinhala-speech-translation/
- /datasets/bangla-speech-translation/
- /datasets/khmer-speech-translation/
- /datasets/hindi-in-the-wild-speech/
- /datasets/arabic-msa-in-the-wild-speech/
- /datasets/indonesian-in-the-wild-speech/
- /datasets/thai-in-the-wild-speech/
- /datasets/turkish-in-the-wild-speech/
- /datasets/vietnamese-in-the-wild-speech/
- /datasets/filipino-in-the-wild-speech/
- /datasets/persian-in-the-wild-speech/
- /datasets/tamil-in-the-wild-speech/
- /datasets/telugu-in-the-wild-speech/
- /datasets/malayalam-in-the-wild-speech/
- /datasets/nepali-in-the-wild-speech/
- /datasets/sinhala-in-the-wild-speech/
- /datasets/bangla-in-the-wild-speech/
- /datasets/khmer-in-the-wild-speech/
- /datasets/arabic-egyptian-call-center-speech/
- /datasets/arabic-egyptian-conversational-speech/
- /datasets/arabic-egyptian-singing-voice/
- /datasets/arabic-egyptian-speech-commands/
- /datasets/arabic-egyptian-read-speech/
- /datasets/arabic-egyptian-podcast-speech/
- /datasets/arabic-egyptian-multilingual-speech/
- /datasets/arabic-egyptian-noisy-speech/
- /datasets/arabic-egyptian-code-switching-speech/
- /datasets/arabic-egyptian-children-speech/
- /datasets/arabic-egyptian-audiobook/
- /datasets/arabic-egyptian-short-video-speech/
- /datasets/arabic-egyptian-live-stream-speech/
- /datasets/arabic-egyptian-voice-assistant/
- /datasets/arabic-egyptian-elderly-voice/
- /datasets/arabic-egyptian-accented-english/
- /datasets/arabic-egyptian-speech-translation/
- /datasets/arabic-egyptian-in-the-wild-speech/
- /datasets/arabic-gulf-saudi-call-center-speech/
- /datasets/arabic-gulf-saudi-conversational-speech/
- /datasets/arabic-gulf-saudi-singing-voice/
- /datasets/arabic-gulf-saudi-speech-commands/
- /datasets/arabic-gulf-saudi-read-speech/
- /datasets/arabic-gulf-saudi-podcast-speech/
- /datasets/arabic-gulf-saudi-multilingual-speech/
- /datasets/arabic-gulf-saudi-noisy-speech/
- /datasets/arabic-gulf-saudi-code-switching-speech/
- /datasets/arabic-gulf-saudi-children-speech/
- /datasets/arabic-gulf-saudi-audiobook/
- /datasets/arabic-gulf-saudi-short-video-speech/
- /datasets/arabic-gulf-saudi-live-stream-speech/
- /datasets/arabic-gulf-saudi-voice-assistant/
- /datasets/arabic-gulf-saudi-elderly-voice/
- /datasets/arabic-gulf-saudi-accented-english/
- /datasets/arabic-gulf-saudi-speech-translation/
- /datasets/arabic-gulf-saudi-in-the-wild-speech/
- /datasets/kannada-call-center-speech/
- /datasets/kannada-conversational-speech/
- /datasets/kannada-singing-voice/
- /datasets/kannada-speech-commands/
- /datasets/kannada-read-speech/
- /datasets/kannada-podcast-speech/
- /datasets/kannada-multilingual-speech/
- /datasets/kannada-noisy-speech/
- /datasets/kannada-code-switching-speech/
- /datasets/kannada-children-speech/
- /datasets/kannada-audiobook/
- /datasets/kannada-short-video-speech/
- /datasets/kannada-live-stream-speech/
- /datasets/kannada-voice-assistant/
- /datasets/kannada-elderly-voice/
- /datasets/kannada-accented-english/
- /datasets/kannada-speech-translation/
- /datasets/kannada-in-the-wild-speech/
- /datasets/punjabi-call-center-speech/
- /datasets/punjabi-conversational-speech/
- /datasets/punjabi-singing-voice/
- /datasets/punjabi-speech-commands/
- /datasets/punjabi-read-speech/
- /datasets/punjabi-podcast-speech/
- /datasets/punjabi-multilingual-speech/
- /datasets/punjabi-noisy-speech/
- /datasets/punjabi-code-switching-speech/
- /datasets/punjabi-children-speech/
- /datasets/punjabi-audiobook/
- /datasets/punjabi-short-video-speech/
- /datasets/punjabi-live-stream-speech/
- /datasets/punjabi-voice-assistant/
- /datasets/punjabi-elderly-voice/
- /datasets/punjabi-accented-english/
- /datasets/punjabi-speech-translation/
- /datasets/punjabi-in-the-wild-speech/
- /datasets/burmese-call-center-speech/
- /datasets/burmese-conversational-speech/
- /datasets/burmese-singing-voice/
- /datasets/burmese-speech-commands/
- /datasets/burmese-read-speech/
- /datasets/burmese-podcast-speech/
- /datasets/burmese-multilingual-speech/
- /datasets/burmese-noisy-speech/
- /datasets/burmese-code-switching-speech/
- /datasets/burmese-children-speech/
- /datasets/burmese-audiobook/
- /datasets/burmese-short-video-speech/
- /datasets/burmese-live-stream-speech/
- /datasets/burmese-voice-assistant/
- /datasets/burmese-elderly-voice/
- /datasets/burmese-accented-english/
- /datasets/burmese-speech-translation/
- /datasets/burmese-in-the-wild-speech/
- /datasets/amharic-call-center-speech/
- /datasets/amharic-conversational-speech/
- /datasets/amharic-singing-voice/
- /datasets/amharic-speech-commands/
- /datasets/amharic-read-speech/
- /datasets/amharic-podcast-speech/
- /datasets/amharic-multilingual-speech/
- /datasets/amharic-noisy-speech/
- /datasets/amharic-code-switching-speech/
- /datasets/amharic-children-speech/
- /datasets/amharic-audiobook/
- /datasets/amharic-short-video-speech/
- /datasets/amharic-live-stream-speech/
- /datasets/amharic-voice-assistant/
- /datasets/amharic-elderly-voice/
- /datasets/amharic-accented-english/
- /datasets/amharic-speech-translation/
- /datasets/amharic-in-the-wild-speech/
- /datasets/ukrainian-call-center-speech/
- /datasets/ukrainian-conversational-speech/
- /datasets/ukrainian-singing-voice/
- /datasets/ukrainian-speech-commands/
- /datasets/ukrainian-read-speech/
- /datasets/ukrainian-podcast-speech/
- /datasets/ukrainian-multilingual-speech/
- /datasets/ukrainian-noisy-speech/
- /datasets/ukrainian-code-switching-speech/
- /datasets/ukrainian-children-speech/
- /datasets/ukrainian-audiobook/
- /datasets/ukrainian-short-video-speech/
- /datasets/ukrainian-live-stream-speech/
- /datasets/ukrainian-voice-assistant/
- /datasets/ukrainian-elderly-voice/
- /datasets/ukrainian-accented-english/
- /datasets/ukrainian-speech-translation/
- /datasets/ukrainian-in-the-wild-speech/
- /datasets/swahili-call-center-speech/
- /datasets/swahili-conversational-speech/
- /datasets/swahili-singing-voice/
- /datasets/swahili-speech-commands/
- /datasets/swahili-read-speech/
- /datasets/swahili-podcast-speech/
- /datasets/swahili-multilingual-speech/
- /datasets/swahili-noisy-speech/
- /datasets/swahili-code-switching-speech/
- /datasets/swahili-children-speech/
- /datasets/swahili-audiobook/
- /datasets/swahili-short-video-speech/
- /datasets/swahili-live-stream-speech/
- /datasets/swahili-voice-assistant/
- /datasets/swahili-elderly-voice/
- /datasets/swahili-accented-english/
- /datasets/swahili-speech-translation/
- /datasets/swahili-in-the-wild-speech/
- /datasets/kurdish-call-center-speech/
- /datasets/kurdish-conversational-speech/
- /datasets/kurdish-singing-voice/
- /datasets/kurdish-speech-commands/
- /datasets/kurdish-read-speech/
- /datasets/kurdish-podcast-speech/
- /datasets/kurdish-multilingual-speech/
- /datasets/kurdish-noisy-speech/
- /datasets/kurdish-code-switching-speech/
- /datasets/kurdish-children-speech/
- /datasets/kurdish-audiobook/
- /datasets/kurdish-short-video-speech/
- /datasets/kurdish-live-stream-speech/
- /datasets/kurdish-voice-assistant/
- /datasets/kurdish-elderly-voice/
- /datasets/kurdish-accented-english/
- /datasets/kurdish-speech-translation/
- /datasets/kurdish-in-the-wild-speech/
- /datasets/hausa-call-center-speech/
- /datasets/hausa-conversational-speech/
- /datasets/hausa-singing-voice/
- /datasets/hausa-speech-commands/
- /datasets/hausa-read-speech/
- /datasets/hausa-podcast-speech/
- /datasets/hausa-multilingual-speech/
- /datasets/hausa-noisy-speech/
- /datasets/hausa-code-switching-speech/
- /datasets/hausa-children-speech/
- /datasets/hausa-audiobook/
- /datasets/hausa-short-video-speech/
- /datasets/hausa-live-stream-speech/
- /datasets/hausa-voice-assistant/
- /datasets/hausa-elderly-voice/
- /datasets/hausa-accented-english/
- /datasets/hausa-speech-translation/
- /datasets/hausa-in-the-wild-speech/
- /datasets/uzbek-call-center-speech/
- /datasets/uzbek-conversational-speech/
- /datasets/uzbek-singing-voice/
- /datasets/uzbek-speech-commands/
- /datasets/uzbek-read-speech/
- /datasets/uzbek-podcast-speech/
- /datasets/uzbek-multilingual-speech/
- /datasets/uzbek-noisy-speech/
- /datasets/uzbek-code-switching-speech/
- /datasets/uzbek-children-speech/
- /datasets/uzbek-audiobook/
- /datasets/uzbek-short-video-speech/
- /datasets/uzbek-live-stream-speech/
- /datasets/uzbek-voice-assistant/
- /datasets/uzbek-elderly-voice/
- /datasets/uzbek-accented-english/
- /datasets/uzbek-speech-translation/
- /datasets/uzbek-in-the-wild-speech/
Reference 35
- /technical/
- /technical/word-error-rate/
- /technical/character-error-rate/
- /technical/speaker-diarization/
- /technical/forced-alignment/
- /technical/voice-activity-detection/
- /technical/diarization-error-rate/
- /technical/sample-rate-for-speech-recognition/
- /technical/mfcc-vs-mel-spectrogram/
- /technical/audio-augmentation/
- /technical/mean-opinion-score/
- /technical/far-field-speech-recognition/
- /technical/pii-redaction/
- /technical/multilingual-speech-recognition/
- /technical/signal-to-noise-ratio/
- /technical/mel-spectrogram/
- /technical/transcription-accuracy/
- /technical/speaker-verification/
- /technical/speaker-identification-vs-speaker-verification/
- /technical/speaker-embedding/
- /technical/phoneme-recognition/
- /technical/code-switching-speech-recognition/
- /technical/accented-speech-recognition/
- /technical/noise-robust-speech-recognition/
- /technical/specaugment/
- /technical/room-impulse-response/
- /technical/dereverberation/
- /technical/asr-evaluation-metrics/
- /technical/zero-crossing-rate/
- /technical/speech-enhancement/
- /technical/audio-sample-rate/
- /technical/wav-vs-mp3/
- /technical/opus-audio-codec/
- /technical/beamforming-microphone-array/
- /technical/neural-speech-codec/
Compliance 27
- /compliance/ai-training-data-licensing/
- /compliance/ai-training-data-and-copyright/
- /compliance/gdpr-voice-data/
- /compliance/gdpr-voice-recording-consent/
- /compliance/is-voice-personal-data/
- /compliance/is-voice-biometric-data/
- /compliance/biometric-privacy-laws/
- /compliance/two-party-consent-states/
- /compliance/voice-recording-consent-by-state/
- /compliance/eu-ai-act-transparency-requirements/
- /compliance/eu-ai-act-article-50/
- /compliance/data-provenance/
- /compliance/data-licensing-agreement/
- /compliance/ai-training-data-laws/
- /compliance/ai-training-data-governance/
- /compliance/california-ai-training-data-transparency-act/
- /compliance/gdpr-biometric-data/
- /compliance/bipa/
- /compliance/dpia/
- /compliance/data-processing-agreement/
- /compliance/vendor-due-diligence-checklist/
- /compliance/voice-anonymization/
- /compliance/anonymisation-vs-pseudonymisation/
- /compliance/ethical-data-collection/
- /compliance/ai-training-consent/
- /compliance/children-voice-dataset/
- /compliance/synthetic-data/
Main 7
Buying guides 18
- AI Data Brokerage
- AI Training Data Providers
- AI Training Data Marketplace
- Buy AI Training Data
- AI Data Licensing
- Embodied AI Data Collection
- Robot Data Collection Services
- Physical AI Data
- Sell AI Training Data
- Sell Data to AI Companies
- Sell Data to AI Labs
- Sell Data to OpenAI
- Sell My Own Data
- Is Selling Data Profitable
- Sell Speech Data
- Sell Training Data
- Voice Data Supplier
- Data Supplier Program
Topics 1
Languages 179
- Afrikaans
- All languages
- Amharic
- Amharic — TTS datasets
- Arabic (Egyptian)
- Arabic (Egyptian) — ASR datasets
- Arabic (Gulf)
- Arabic (Gulf) — ASR datasets
- Arabic (Iraqi)
- Arabic (Moroccan)
- Arabic (Moroccan) — Voice datasets
- Arabic (MSA)
- Arabic (MSA) — ASR datasets
- Arabic (MSA) — TTS datasets
- Arabic (MSA) — Voice datasets
- Arabic (Sudanese)
- Arabic (Tunisian)
- Arabic (Yemeni)
- Armenian
- Assamese
- Assyrian Neo-Aramaic
- Aymara
- Azerbaijani
- Bambara
- Bangla
- Bangla — ASR datasets
- Bangla — TTS datasets
- Bangla — Voice datasets
- Bhojpuri
- Bicolano
- Burmese
- Burmese — ASR datasets
- Cantonese
- Catalan
- Cebuano
- Chechen
- Chichewa
- Czech
- Danish
- Dari
- Dhivehi
- Dutch
- Fijian
- Filipino
- Filipino — ASR datasets
- Filipino — TTS datasets
- Filipino — Voice datasets
- Finnish
- French
- Fula (Fulfulde)
- Georgian
- German
- Greek
- Guarani
- Gujarati
- Haitian Creole
- Hausa
- Hebrew
- Hindi
- Hindi — ASR datasets
- Hindi — TTS datasets
- Hindi — Voice datasets
- Hmong
- Hungarian
- Igbo
- Ilocano
- Indonesian
- Indonesian — ASR datasets
- Indonesian — TTS datasets
- Indonesian — Voice datasets
- Italian
- Japanese
- Javanese
- Kabyle
- Kannada
- Kannada — ASR datasets
- Kashmiri
- Kazakh
- Kazakh — TTS datasets
- Khmer
- Khmer — ASR datasets
- Khmer — TTS datasets
- Khmer — Voice datasets
- Kinyarwanda
- Korean
- Kurdish
- Kurdish — TTS datasets
- Kyrgyz
- Lao
- Levantine Arabic
- Luganda
- Maithili
- Malagasy
- Malay
- Malayalam
- Malayalam — ASR datasets
- Malayalam — TTS datasets
- Malayalam — Voice datasets
- Mandarin Chinese
- Maori
- Marathi
- Mongolian
- Nahuatl
- Navajo
- Nepali
- Nepali — ASR datasets
- Nepali — TTS datasets
- Nepali — Voice datasets
- Norwegian
- Odia
- Oromo
- Pashto
- Persian
- Persian — ASR datasets
- Persian — TTS datasets
- Persian — Voice datasets
- Polish
- Portuguese (Brazil)
- Portuguese (Portugal)
- Punjabi
- Punjabi — ASR datasets
- Quechua
- Romanian
- Russian
- Samoan
- Sesotho
- Setswana
- Shona
- Sindhi
- Sinhala
- Sinhala — ASR datasets
- Sinhala — TTS datasets
- Sinhala — Voice datasets
- Somali
- Spanish (Latin America)
- Spanish (Spain)
- Sundanese
- Swahili
- Swahili — ASR datasets
- Swedish
- Tajik
- Tamazight (Berber)
- Tamil
- Tamil — ASR datasets
- Tamil — TTS datasets
- Tamil — Voice datasets
- Tatar
- Telugu
- Telugu — ASR datasets
- Telugu — TTS datasets
- Telugu — Voice datasets
- Tetum
- Thai
- Thai — ASR datasets
- Thai — TTS datasets
- Thai — Voice datasets
- Tigrinya
- Turkish
- Turkish — ASR datasets
- Turkish — TTS datasets
- Turkish — Voice datasets
- Turkmen
- Twi (Akan)
- Ukrainian
- Ukrainian — TTS datasets
- Urdu
- Uyghur
- Uzbek
- Uzbek — TTS datasets
- Vietnamese
- Vietnamese — ASR datasets
- Vietnamese — TTS datasets
- Vietnamese — Voice datasets
- Wolof
- Xhosa
- Yoruba
- Yoruba — TTS datasets
- Yucatec Maya
- Zulu
Categories 73
- Academic Data
- Academic Data — datasets
- Accented English
- Accented English — datasets
- All categories
- Audiobook
- Audiobook — datasets
- Autonomous Driving Data
- Autonomous Driving Data — datasets
- Call Center Speech
- Call Center Speech — datasets
- Children Speech
- Children Speech — datasets
- Code-Switching Speech
- Code-Switching Speech — datasets
- Computer Vision Data
- Computer Vision Data — datasets
- Conversational Speech
- Conversational Speech — datasets
- Dictionary Data
- Dictionary Data — datasets
- Elderly Voice
- Elderly Voice — datasets
- Embodied AI Data
- Embodied AI Data — datasets
- First-Person View Data
- First-Person View Data — datasets
- In-the-Wild Speech
- In-the-Wild Speech — datasets
- Legal Data
- Legal Data — datasets
- Live Stream Speech
- Live Stream Speech — datasets
- LLM Training Data
- LLM Training Data — datasets
- Multilingual Language Data
- Multilingual Language Data — datasets
- Multilingual Speech
- Multilingual Speech — datasets
- News Data
- News Data — datasets
- NLP Data
- NLP Data — datasets
- Noisy Speech
- Noisy Speech — datasets
- Patent Data
- Patent Data — datasets
- Podcast Speech
- Podcast Speech — datasets
- Read Speech
- Read Speech — datasets
- Satellite Imagery Data
- Satellite Imagery Data — datasets
- Short Video Speech
- Short Video Speech — datasets
- Singing Voice
- Singing Voice — datasets
- Speech Commands
- Speech Commands — datasets
- Speech Translation
- Speech Translation — datasets
- Synthetic Data
- Synthetic Data — datasets
- Text Data
- Text Data — datasets
- Translation Data
- Translation Data — datasets
- Video Data
- Video Data — datasets
- Voice Assistant
- Voice Assistant — datasets
- World Model Data
- World Model Data — datasets
Glossary 48
- AgiBot World
- Annotation
- Anonymization
- ASR
- Biometric Data
- Code-Switching
- Consent
- Data Compliance
- Data Licensing
- Data Provenance
- Dexterous Hand Data
- Diarization
- DROID Dataset
- Egocentric Video Data
- Embodied AI Data
- GDPR
- Glossary index
- Grasping Data
- Ground Truth
- Humanoid Robot Data
- Imitation Learning
- Inter-Annotator Agreement
- LeRobot
- LIBERO
- Low-Resource Language
- Mobile Robot Navigation Data
- Motion Capture Data
- Open X-Embodiment
- Phoneme
- PII
- Prosody
- Robot Data Annotation
- Robot Data Augmentation
- Robot Gripper Data
- Robot Manipulation Data
- Sample Rate
- Script Split
- Sim-to-Real Transfer
- SNR
- Speaker Identification
- Tactile Sensor Data
- Teleoperation
- Transcription
- TTS
- Vision-Language-Action Model
- Warehouse Robotics Data
- WER
- World Model
Insights 157
- A vendor risk questionnaire that actually discriminates
- Accent labels that annotators can apply the same way twice
- Aligning hour-long recordings: cut points, drift, and cross-chunk boundaries
- An EU AI Act timeline for data teams, ordered by what cannot be recovered
- Annotating robot episodes at scale: throughput, sampling, and what not to label
- Article 53 of the EU AI Act: what general-purpose model providers owe on training data
- ASR confidence scores: what they mean and when to send a human
- Audio augmentation recipes that help, and the ones that waste compute
- Audit rights over a data supplier: what to audit, and what you cannot
- Auditing a dataset before you buy it: a sample-based checklist that finds the real defects
- Avoiding scope creep in a data project
- Build or buy: the real cost of collecting training data in-house
- Building a data governance file for EU AI Act Article 10
- Building a gold set for annotation QC: composition, leakage, and when to retire items
- Building a grapheme-to-phoneme lexicon that holds up
- Building a noise robustness test set that supports a decision
- Building a pronunciation lexicon that survives proper nouns
- Building a speaker-labeled transcript you can ship
- Building an ASR evaluation harness you can re-run next quarter
- Buying data for low-resource languages: what changes
- Calibration records are part of the dataset, not internal housekeeping
- Casing and inverse text normalization: two decisions that leak into your labels
- CER or WER: which one to report, and when the choice flips a ranking
- Children's voice data and AI training: why this is the highest-risk corpus
- Choosing a delivery format for a dataset: audio, labels, and structure
- Choosing a diarization toolkit: pyannote, NeMo, or a commercial API
- Choosing a format for robot data: RLDS, HDF5, or LeRobot-style datasets
- Choosing a masking policy for speech augmentation
- Choosing a robot for data collection: buy for uptime, not for precision
- Choosing a sample rate for a new collection
- Choosing cameras for robot data capture: geometry first, sensor second
- Closing the sim-to-real gap in practice: what to randomize and how much real data to add
- Code-switching: the data problem nobody scopes for
- Computing WER in Python: what jiwer rewrites before it counts
- Continuity and escrow when a data supplier stops supplying
- Converting audio formats for training: a checklist that catches the quiet failures
- Copyright diligence before you license a dataset
- Cross-border transfers of training data: what counts as a transfer, and how to structure one
- Curriculum learning for speech models: what difficulty means and when it pays
- Data labeling vendor selection criteria you can score and defend
- Data marketplaces versus commissioned collection: choosing the right channel
- Data-centric AI, explained as a purchasing decision
- Dereverberation before recognition: when it helps and when it does damage
- Designing a pilot batch before you scale a robot data collection
- Designing a task taxonomy for a robot dataset
- Detecting drift in a long annotation project before the batch is ruined
- Documenting a dataset for internal handover
- Drafting the DPA schedule for a speech data project
- EU AI Act penalties: what actually gets fined, and how the exposure reaches data suppliers
- Evaluating a robot policy before deployment: eval sets, metrics, and regression tests
- Exclusivity clauses in dataset licences: what you are actually buying
- Failure data: the part of a robot dataset most pipelines throw away
- Fair use and AI training data: the argument, the open questions, and what a buyer can control
- Fine-tuning a speech model on domain jargon: legal, financial, industrial
- Fine-tuning wav2vec2 for a low-resource language without wrecking it
- Fine-tuning Whisper on your own audio: what the data has to look like
- Five ways a WER ends up better than the system really is
- Forced alignment with MFA: dictionary preparation, OOV words, and acceptance checks
- Forced alignment with WhisperX: handling words the aligner cannot place
- Getting voice actor consent for AI training, in practice
- Governing law and disputes in cross-border data deals
- Handling a data batch that failed acceptance
- How big an ASR test set needs to be, and how to build it
- How many robot trajectories do you actually need
- How to choose an AI data vendor when there is no track record to check
- How to compare quotes for a data collection project fairly
- How to define acceptance criteria for a dataset
- How to estimate a data collection timeline: two clocks and one critical path
- How to evaluate an annotation provider with a test you control
- How to read the US Copyright Office AI reports without over-reading them
- How to run a vendor onboarding before the first batch
- How to sample for annotation QC: full inspection, stratification, and how much is enough
- How to specify a speech data project so it does not get rejected
- How to write a data requirement brief that cannot be misread
- Human demonstration versus autonomous data: what each one can teach a policy
- Human in the loop: where the person actually belongs in a data pipeline
- Insights index
- IP assignment in a data collection contract: who ends up owning what
- Keeping a provenance log that survives scrutiny
- Labeling robot trajectories: instructions, success labels, and how fine to go
- Language identification data: the segment length decides the job
- Liability and indemnity in a data supply contract
- librosa and torchaudio: which one for which job in a dataset pipeline
- Licensing an existing dataset versus commissioning a new one
- Licensing terms for robot data: the three parties and what each one needs
- Managing a multi-language data program
- Measuring ASR latency and throughput without fooling yourself
- Measuring dataset diversity beyond the total hour count
- Measuring SNR on real recordings with no clean reference
- MFCC or mel spectrogram in practice: picking a front end and pinning its numbers
- Model rights and data rights are two different assets
- Normalize before you score, and record what you normalized away
- Outsourcing data annotation: what the handover really involves
- Overlapping speech: why diarization fails there and how to record it
- Payment terms and milestones in a data project
- Planning a dexterous manipulation dataset: hardware, tasks, and why the yield is low
- Planning a multilingual data project: sequence, volume, and consistency
- Planning audio recording sessions so they produce usable hours
- Publishing a training data transparency disclosure
- Punctuation restoration for ASR output: what the training data has to be
- Quality checks for a robot episode: continuity, timestamps, dropped frames, drift
- Quality metrics for TTS training data: what to measure per file, per session, and per voice
- Questions to ask before buying a speech dataset, and the red flags in the answers
- Reading a diarization error rate: what the three components tell you
- Recording speaker demographics: fields, granularity, and where the privacy line sits
- Redacting PII from transcripts: what to remove, what to put in its place, and what over-redaction costs
- Refresh and maintenance terms for a dataset you will keep using
- Resampling audio without breaking it
- Resolving disagreements in annotation review: arbitration, voting, and what disagreement is telling you
- Retention rules for training data: three layers, one deletion request
- Right of publicity and AI voices: a separate right from copyright
- Running a teleoperation shift: rig setup, operator training, and the data you throw away
- Running consent for field recording, from approach to archive
- Running pyannote diarization: the parameters that matter and the merge with ASR
- Scaling robot data collection across sites without splitting the dataset in two
- Scoping a second-language batch: what transfers from the first
- Score a multilingual model per language, or the headline will mislead you
- Scoring code-switched audio without penalizing the model twice
- Script splits: when one spoken language needs two datasets
- Setting up egocentric video collection: rigs, angles, sync, and privacy
- Silence and non-speech in training audio: what to cut and what to keep
- SLAs for data delivery, and the clauses that make them unusable
- Sourcing speakers of a rare language
- Speaker verification with embeddings: setting a threshold you can defend
- Speech translation data: decide what artifact you are buying
- Streaming and batch ASR need training segments cut different ways
- Sublicensing and resale rights in a data licence
- The costs that quietly accumulate in a data collection project
- The dataset checklist to run before speech recognition training starts
- The Generative AI Copyright Disclosure Act: what it would require, and why it matters before it passes
- The preprocessing pipeline for ASR: what each step does and why the order matters
- Training an ASR model with NeMo: manifests, tokenizers, and the config that ties them
- Transcription guidelines: writing down the rules for numbers, fillers and laughter
- TTS style and prosody data: the recording matrix is the dataset
- Tuning VAD before diarization: settings that change the speaker output
- Union voice agreements and digital replicas: what a data buyer inherits
- Voice cloning: how much audio, what quality, and where consent ends
- Voice conversion versus voice cloning: two problems that look alike
- Wake word data: what keyword spotting needs that ASR data does not
- What a paid pilot batch should prove before you commit to the full project
- What a single WER number hides about the set behind it
- What actually drives the cost of a robot data collection
- What an AI data supplier should tell you before you ask
- What GDPR actually requires for voice data
- What is a foundation model, and why its data requirements invert the usual ones
- What is a voice agent, and the data a demo never has
- What is domain randomization, and what a wider range costs in episodes
- What is instruction tuning, and why its data does not scale like pretraining data
- What is red teaming, and the kind of data it actually needs
- What is RLHF, and why preference data is a different purchase entirely
- When a data contract ends, what happens to the models already trained
- When to stop collecting data
- Why speaker count matters more than hours
- Working with a data broker versus contracting direct
- Writing a training data summary that is actually useful
- Writing an annotation guideline that holds up under production
- Writing an RFP for a data project without collecting unusable bids