Sell Speech Data
Speech is our home category, so this page can be specific: what buyers ask for, what a supplier needs to have, and which kinds of audio are unsellable no matter how good they sound.
What buyers ask a speech supplier for
Speech data demand is precise, and the specification almost always contains the same fields: language and dialect, total hours, minimum distinct speaker count, recording conditions, annotation depth, and delivery format. Speaker count is the field suppliers underestimate. Hours are easy to add by recording the same voices longer; a buyer who asks for 100 hours across 80 speakers is asking for something harder, and that difficulty is where your margin lives.
The second most common requirement is conditions. Studio recording with a low noise floor, home recording, mobile capture, or a defined noisy environment — each is a different product, and a supplier who can only deliver one of them is limited to the projects that match it.
One structural fact shapes everything below: there is no public portal at any major AI company where speech data can be uploaded and paid for. OpenAI, Google, and Anthropic buy through data teams and vendors. Demand reaches a supplier as a specification, which is why the rest of this page is about specifications and rights rather than about where to upload. And the specification is not a formality: audio that matches no live requirement cannot be paid for, no matter how clean the recording is.
What you need before you can sell speech
Four capabilities, and buyers check all of them.
- Native-speaker access — you can recruit people who genuinely speak the target language or dialect, in the numbers the specification needs.
- A recording setup that meets the condition being sold — microphone, room treatment, and a fixed signal chain, applied consistently across all sessions.
- Annotation capability, or a partner who has it — transcription at minimum, and ideally the deeper layers: intent, diarization, phoneme alignment.
- A consent process — a signed release from every speaker, specific enough about AI training use that a buyer's legal review will accept it.
The consent question, in voice
Voice data is personal data, and a recording is the person in a way that a text file is not. Every speaker in your dataset needs a document that says what the recording will be used for, who it may be shared with, how long it is retained, and how they can withdraw. This is the most common reason speech suppliers are turned away — not audio quality, not price, but a rights chain that cannot be shown.
It also defines what you can sell. You can sell recordings you produced, of speakers who consented. You cannot sell recordings of people who did not, and there is no intermediary who can fix that after the fact.
What is not sellable, and how the money works
Three categories are declined by us and by serious buyers everywhere. Recorded telephone calls, because the other party on the call never consented. Medical and clinical recordings, because consent and regulation make them a different business entirely. And audio scraped from podcasts, videos, or social media that you do not own — public availability is not a license, and rights holders can and do pursue it.
On payment: money follows acceptance. You deliver, the buyer's QA reviews against the criteria written at the start, and the batch is paid when it passes. A supplier who prices to survive one round of re-work is safe; one who treats handover as the finish line is not.
If your material is in the sellable categories and you can document your rights, the practical step is a specific first message: which languages and dialects you can produce, your recording conditions, your annotation depth, and your monthly capacity in hours. That is what can be matched against a live requirement — and if there is no match today, it tells us what to bring you next. Send it through the request form, selecting custom data collection, and it goes on file against the requirements we are working.
Questions we get asked
Can I sell speech data recorded on a phone?
Yes, if the specification asks for mobile capture. Many projects do. What you cannot do is sell phone recordings as studio data — the condition is part of the product, and misrepresenting it fails QA.
Do I need to annotate, or can I sell raw audio?
Raw audio sells, at a lower rate and to fewer buyers. Most demand includes at least transcription, and the annotation layers are where the higher rates are.
What if my speakers signed a general release?
A general release that does not mention machine learning use is usually not enough, because consent has to be specific about purpose. It can often be fixed by re-consenting the speakers before sale — better before than after.
Keep reading
-
AI Data Brokerage
A data broker sits between the people who need training data and the people who can produce it. We source to order — no inventory, no license resale, no recycled datasets.
-
AI Training Data Providers
What to check before you sign with a training data provider — and how we compare on each point.
-
AI Training Data Marketplace
Marketplaces sell what already exists. We source what you actually need. Here is when each model makes sense.
-
Full data catalog
36 data categories across 120 languages, and how to specify each one.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.