In-the-Wild Speech deliveries
We publish no fixed listings, because we hold no inventory. What follows is the structure every in-the-wild speech delivery takes — the fields your specification needs, and what actually arrives in the package.
Why there is nothing to add to a cart
A catalog listing answers the question "what do you have?" Our model answers a different question: "what do you need?" Every in-the-wild speech project is quoted from a specification, which means the price, the timeline and the data itself are all derived from your requirement rather than from something already built.
The tradeoff is real and worth stating plainly: the timeline starts when the specification is agreed, not when you click a button. What you get in exchange is a dataset that matches the requirement, and — for anything unusual — one that is not also available to your competitors.
What arrives in a delivery
| Component | What it is |
|---|---|
| Audio files | Format, sampling rate and segmentation length set to your pipeline. Named to a convention you specify. |
| Transcription | Orthographic or phonetic, produced under a guideline you review before production starts. |
| Speaker metadata | Age band, gender, region and dialect background per file, plus a speaker identifier so the dataset can be sliced. |
| Recording conditions | Environment, device and, where relevant, measured signal-to-noise ratio per file. |
| Annotation guideline | The document the annotators worked from, so you can reproduce the conventions on your own data. |
| Consent records | Signed speaker consent covering your intended use, with the transfer mechanism addressed where required. |
| Collection methodology | How speakers were recruited, screened and scheduled — the part that tells you how biased the pool is. |
| Quality report | Pilot outcome, re-work log, and the annotator agreement figures where the task supports measuring them. |
Specification fields for in-the-wild speech
- Volume — total hours, and minimum distinct speaker count. The second number constrains the project far more than the first.
- Language and variety — including dialect where the language has meaningful internal variation, and script where the language has more than one writing system.
- Recording conditions — studio, home, mobile, or a named environment with a target signal-to-noise ratio.
- Annotation depth — Recording device (phone, headset, professional mic) and collection environment must be logged per file.
- Delivery format — audio container and sampling rate, segmentation length, metadata schema.
- Compliance scope — where the speakers are, where the data will be used, and whether the model will be distributed.
Next step
Send the specification and you will get a real number and a real timeline. If the requirement is not something we can source well, we will say so rather than take the order.
Request in-the-wild speech
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.