Insights
Practical writing on sourcing training data. No trend pieces — just the parts of a data project that decide whether the result is usable.
-
Sourcing speakers of a rare language
The constraint is not finding people who speak it. It is finding people who meet every other requirement at the same time, and proving that they do.
-
Planning audio recording sessions so they produce usable hours
Session length, fatigue, and in-session checks decide how much of a recording day survives review. A plan for the day, and the numbers to set it with.
-
Handling a data batch that failed acceptance
A failed batch is a decision with three possible outcomes, and the expensive mistake is choosing the wrong one. How to measure the failure, and what to agree beforehand.
-
When to stop collecting data
More data stops helping before it stops costing. The signals that the curve has flattened, and the stopping rule that makes the decision mechanical.
-
Documenting a dataset for internal handover
The document set that lets a colleague use the data correctly six months later, including the two items teams almost always leave out.
-
Working with a data broker versus contracting direct
An intermediary sells search, consolidation, legal cover, and risk transfer. Whether that is worth its margin depends on the shape of the project, not the size of the order.
-
Avoiding scope creep in a data project
In a corpus project every property is additive, so requests arrive continuously and each one looks small. The fix is a classification step before any pricing.
-
Choosing a delivery format for a dataset: audio, labels, and structure
The audio container, the annotation file format, and the directory layout are three separate decisions, and only one of them is expensive to reverse.
-
Managing a multi-language data program
Quality means something different in every language, and one delayed language can hold a whole release. Both problems are managed with the same three artifacts.
-
Quality metrics for TTS training data: what to measure per file, per session, and per voice
TTS data fails in ways transcription corpora do not. The measurable properties — signal, consistency, pronunciation, recording state — and the gates that catch them.
-
Auditing a dataset before you buy it: a sample-based checklist that finds the real defects
A listing describes the best of a dataset; an audit looks at the actual files. What to request, what to compute, what to compare, and how to end with a decision.
-
How to estimate a data collection timeline: two clocks and one critical path
Most timeline estimates add up work that runs in parallel and ignore the phases that wait on other people. Separating the two clocks fixes both errors.
-
Measuring dataset diversity beyond the total hour count
The same thousand hours can be one voice repeated or a thousand people. The distribution metrics that show which one you have, and how to report them honestly.
-
Recording speaker demographics: fields, granularity, and where the privacy line sits
Demographics are what make a dataset sliceable and auditable. They are also personal data. What to record per speaker, how to store it, and what to leave out.
-
Redacting PII from transcripts: what to remove, what to put in its place, and what over-redaction costs
Redaction is a policy decision before it is a task. What belongs on the removal list, how replacements should behave, and the content destroyed by removing too much.
-
Resolving disagreements in annotation review: arbitration, voting, and what disagreement is telling you
Most disagreement is a message about the guideline, not a failure by an annotator. How to route it: what to arbitrate, when voting helps, and when it hides the problem.
-
Building a gold set for annotation QC: composition, leakage, and when to retire items
A gold set is how you measure annotators instead of trusting them. How to build one that keeps measuring something, and the ways it quietly stops working.
-
Detecting drift in a long annotation project before the batch is ruined
Quality does not fall off a cliff; it slides. The few measurements that catch the slide, and how to tell annotator fatigue apart from a guideline that has loosened.
-
Annotating robot episodes at scale: throughput, sampling, and what not to label
At tens of thousands of episodes the question changes from how to label to what to label. Automatic outcomes, sampled review, and the process that keeps a large annotator pool aligned.
-
How to sample for annotation QC: full inspection, stratification, and how much is enough
Reviewing every item is not the safe option. It is the expensive one, and it misses exactly the errors that matter. A sampling plan that can be defended.
-
Writing an annotation guideline that holds up under production
The guideline is what makes two annotators interchangeable. What belongs in it, how to write boundary cases, when examples earn their place, and how it changes.
-
Calibration records are part of the dataset, not internal housekeeping
Which calibration values to ship with robot data, the fields that make a record usable six months later, and what a dataset without them silently cannot do.
-
Failure data: the part of a robot dataset most pipelines throw away
Rejected episodes, near-misses, and operator aborts are the only source of negative examples and recovery behavior. What each kind is good for, and what it costs to keep.
-
Choosing cameras for robot data capture: geometry first, sensor second
How many pixels the object occupies, how many frames the contact lasts, and what the policy will actually consume. The three questions that decide the camera set.
-
Scaling robot data collection across sites without splitting the dataset in two
What to standardize, what to let vary on purpose, and how to tell site variation from drift before a month of episodes carries an error nobody can name.
-
Human demonstration versus autonomous data: what each one can teach a policy
Demonstrations sample the expert. Autonomous rollouts sample the policy, including the states where it fails. Neither replaces the other, and the order they are used in matters.
-
Licensing terms for robot data: the three parties and what each one needs
A robot episode carries rights from the platform, the site, and the operator. The clauses that get skipped are the ones that decide what can be done with a trained model later.
-
Evaluating a robot policy before deployment: eval sets, metrics, and regression tests
A success rate over a handful of trials is not an evaluation. How to build an eval set that survives repetition, which metrics predict deployment, and how to catch a regression before the robot does.
-
What actually drives the cost of a robot data collection
Episode cost is not operator time. It is the ratio of setup and reset to demonstration, divided by yield, on top of a fixed engineering line that only volume amortizes.
-
Designing a task taxonomy for a robot dataset
One node, one success criterion, a controlled verb list, and facets kept out of the name. The schema that decides whether a collection stays filterable six months later.
-
Audit rights over a data supplier: what to audit, and what you cannot
An audit clause that is too broad gets refused, and an audit that finds nothing proves nothing. How to scope the right, the sample and the follow-up.
-
Governing law and disputes in cross-border data deals
The law you choose decides how the contract is read. It does not decide whether the collection was lawful where the speakers live. What to pick, and why enforcement is the real question.
-
Continuity and escrow when a data supplier stops supplying
You already hold the data. What you lose is the records, the recipe and the ability to prove where it came from. What to escrow, and what is cheaper than escrow.
-
Liability and indemnity in a data supply contract
Five ways a data deal goes wrong, who carries each one by default, and how caps, carve-outs and insurance are actually negotiated.
-
When a data contract ends, what happens to the models already trained
Termination clauses deal with the data. Almost none of them deal with the model that has already learned from it. The options, and what each one costs.
-
Sublicensing and resale rights in a data licence
Four different things get called sharing, and which one you need determines the clause. The resale permission is the one almost never granted.
-
Exclusivity clauses in dataset licences: what you are actually buying
Exclusivity is a scope question before it is a price question. What the clause has to define, what it can never deliver, and when the premium is not worth paying.
-
Refresh and maintenance terms for a dataset you will keep using
A one-time delivery is a snapshot. If the dataset stays in the pipeline, the agreement has to define versions, refresh triggers, and what happens to the older editions.
-
Model rights and data rights are two different assets
Buying data does not buy the model, and licensing a model does not carry the data rights with it. The boundary questions that have to be written into the contract.
-
Payment terms and milestones in a data project
Payments should follow accepted artifacts, not calendar dates. How to phase them, what to hold back, and when withholding money is the wrong tool entirely.
-
How to run a vendor onboarding before the first batch
Between signature and the first delivery: roles, artifacts, access, controls and cadence. The five parts that decide whether the first batch is routine or revealing.
-
IP assignment in a data collection contract: who ends up owning what
A collection contract moves several different assets at once, and each one needs its own clause. Here is the stack of rights, and the gaps that only show up years later.
-
Writing an RFP for a data project without collecting unusable bids
The commercial fields an RFP needs beyond the specification — response template, evaluation method, required artifacts — and what each blank costs you at evaluation time.
-
SLAs for data delivery, and the clauses that make them unusable
Cadence, quality thresholds, repair windows and remedies. What an SLA needs in order to be applied at all, and the definitions that decide whether it can be.
-
How to define acceptance criteria for a dataset
Metrics, sampling, the acceptance window and the disposition of rejected units. The four parts of an acceptance procedure, and the defaults that quietly accept bad data.
-
Data labeling vendor selection criteria you can score and defend
A weighted scorecard with defined anchors, a disqualification floor, and a scoring order that keeps price from quietly deciding the ranking.
-
What an AI data supplier should tell you before you ask
A supplier with a real process has the artifacts on hand. The disclosure set that should arrive unprompted, and how to request each item when it does not.
-
Data marketplaces versus commissioned collection: choosing the right channel
Both put training data into your pipeline. The differences that decide are how the inventory came to exist, what a listing cannot tell you, and where the provenance risk sits.
-
Children's voice data and AI training: why this is the highest-risk corpus
Children's audio sits where the strictest consent rules, biometric-adjacent data and a synthetic voice capability meet. The documentation problem is solved at collection or not at all.
-
Outsourcing data annotation: what the handover really involves
Outsourcing moves the annotation work, not the responsibility. What you hand over, what the provider must be able to decide alone, and what should stay with your own team.
-
How to evaluate an annotation provider with a test you control
A proposal cannot show whether annotators follow a guideline. A short qualification task can, if you build it with planted defects and score it against your own reference.
-
Right of publicity and AI voices: a separate right from copyright
A licence to a recording is not a licence to a person. Publicity law protects the voice itself, varies sharply between states, and is where most of the risk in synthetic voice work actually sits.
-
Cross-border transfers of training data: what counts as a transfer, and how to structure one
A training corpus crosses borders more than once — collection, annotation, storage, compute — and each of those moves is its own legal event. Here is how to map them and what to write down.
-
Retention rules for training data: three layers, one deletion request
Retention law, retention contracts and what a trained model keeps are three different clocks. Deleting the files answers only the first, and the request that arrives after training is the one to plan for.
-
Fair use and AI training data: the argument, the open questions, and what a buyer can control
Fair use is a defence argued after the fact, factor by factor. Here is how the four factors map onto training, why no appellate court has settled the question, and where a buyer's leverage actually sits.
-
How to read the US Copyright Office AI reports without over-reading them
The Office has published a series of reports on AI and copyright. Some of what they contain is examination practice that binds applicants; much of it is a recommendation to Congress. Telling the two apart is the skill.
-
Union voice agreements and digital replicas: what a data buyer inherits
A recording made under a union agreement carries contractual limits that do not travel with the audio files. Here is where the digital replica terms sit in the chain, and what to verify before licensing.
-
Article 53 of the EU AI Act: what general-purpose model providers owe on training data
Article 53 puts four duties on providers of general-purpose models. Two of them reach straight into the training corpus, and both are satisfied or broken by facts a supplier holds.
-
EU AI Act penalties: what actually gets fined, and how the exposure reaches data suppliers
The fine architecture is tiered by conduct rather than by company size. The failures that produce penalties in practice are documentary, and they travel down the supply chain as contract terms.
-
The Generative AI Copyright Disclosure Act: what it would require, and why it matters before it passes
The federal bill would turn the training corpus into a filing with the Copyright Office. It is not law yet, but the facts it names are the same ones buyers already ask suppliers to produce.
-
Voice cloning: how much audio, what quality, and where consent ends
Cloning needs one speaker recorded consistently rather than many speakers recorded widely. The hard parts are coverage, session drift, and consent that answers what happens to the model.
-
Voice conversion versus voice cloning: two problems that look alike
Cloning generates speech in a target voice; conversion re-renders an existing performance. The data barely overlaps, and conversion puts a second person in the rights chain.
-
What is domain randomization, and what a wider range costs in episodes
Randomization defines a distribution rather than a scene, so the range is a data decision. Widen it and the policy hedges; narrow it and the policy cannot transfer.
-
What is red teaming, and the kind of data it actually needs
Red teaming is adversarial testing rather than quality assurance with a longer checklist. It needs trained attackers, a rubric written in advance, and a budget for re-running it.
-
Human in the loop: where the person actually belongs in a data pipeline
The useful question is not whether humans are involved but at which step. Put them where no rule can check the answer, sample the review properly, and count what gets dropped.
-
What is a voice agent, and the data a demo never has
A voice agent is judged on when it speaks, not only on what it heard. That moves turn-taking, interruption, and outcome labels to the center of the specification.
-
What is RLHF, and why preference data is a different purchase entirely
RLHF trains on comparisons rather than correct answers. That changes who can produce the labels, how disagreement is handled, and why the set stays small.
-
What is a foundation model, and why its data requirements invert the usual ones
A narrow model needs more examples of one task. A foundation model needs coverage of everything you have not thought of yet, plus an evaluation slice you cannot buy.
-
Data-centric AI, explained as a purchasing decision
Holding the model fixed and fixing the data is often the cheaper route to accuracy, but only if the work is targeted, iterative, and budgeted in the right shape.
-
TTS style and prosody data: the recording matrix is the dataset
Style labels, repeated takes and session control decide whether a model learns styles or learns the recording order. The matrix belongs in the specification.
-
Speech translation data: decide what artifact you are buying
Speech-to-text, cascaded and speech-to-speech pipelines need different artifacts from the same recordings. Fix the pipeline first, then write the spec.
-
What is instruction tuning, and why its data does not scale like pretraining data
Instruction tuning needs a small number of verified instruction-response pairs rather than a bigger corpus. Here is what makes a pair worth paying for.
-
Building a pronunciation lexicon that survives proper nouns
A lexicon is a maintained artifact, not a one-time export. Pick a phone set, settle heteronyms and numbers, and keep the out-of-vocabulary loop running.
-
Transcription guidelines: writing down the rules for numbers, fillers and laughter
The guideline is the dataset. These are the decisions to settle before annotators start, and the calibration round that finds the ones you missed.
-
Silence and non-speech in training audio: what to cut and what to keep
Trim too aggressively and the model never learns to output nothing. Keep everything and the compute goes to silence. The policy belongs in the data spec.
-
Accent labels that annotators can apply the same way twice
Accent is a continuum, so the label scheme is a lossy choice. Provenance beats perception for the category, and listeners are reliable only on strength.
-
Streaming and batch ASR need training segments cut different ways
A batch model trains on whole utterances. A streaming model has to be trained on the cuts it will meet in production, mid-word boundaries included.
-
ASR confidence scores: what they mean and when to send a human
Confidence is a decoder by-product, not a calibrated probability. Use it to rank a review queue, and validate the ranking against real errors.
-
Fine-tuning a speech model on domain jargon: legal, financial, industrial
Rare words, spelled-out codes, and why overall word error rate hides the improvement you paid for. How to build and score a jargon dataset.
-
Wake word data: what keyword spotting needs that ASR data does not
A wake word model is a detector, not a transcriber. Positives, confusable negatives and false-accept testing are the three parts of the dataset.
-
Language identification data: the segment length decides the job
LID accuracy is a function of how much audio the model hears. Data for a one-second router and a thirty-second triage pass are different products.
-
Casing and inverse text normalization: two decisions that leak into your labels
Whether the transcript says 1200 or twelve hundred, and whether it says White House or white house, is a labeling decision that changes what a model learns.
-
Choosing a masking policy for speech augmentation
Time masks, frequency masks, and the combinations that hurt. How to set the budget, and why one setting behaves differently on tonal languages and on fixed-window models.
-
Curriculum learning for speech models: what difficulty means and when it pays
Ordering training data from easy to hard sounds free and rarely is. How to define difficulty, why the gains are usually small, and the one use that always pays.
-
The dataset checklist to run before speech recognition training starts
Transcript conventions, segment lengths, speaker-disjoint splits, and duplicate checks. The work that decides whether training converges on anything.
-
Building a grapheme-to-phoneme lexicon that holds up
Where the pronunciations come from, how to handle words that have more than one, and how to measure whether the lexicon is any good.
-
Punctuation restoration for ASR output: what the training data has to be
Punctuation models train on text, not audio, but the conventions and the evaluation are where projects go wrong.
-
Fine-tuning Whisper on your own audio: what the data has to look like
Whisper fine-tuning is mostly a data-format problem. What the labels, the segment lengths, and the held-out sets have to be before training starts.
-
Fine-tuning wav2vec2 for a low-resource language without wrecking it
Character sets, frozen layers, and the pretraining mismatch that decides the result. What to fix before running CTC training on a small corpus.
-
Training an ASR model with NeMo: manifests, tokenizers, and the config that ties them
The manifest fields that have to be right, how the tokenizer choice constrains the model, and the config overrides that actually change the result.
-
Converting audio formats for training: a checklist that catches the quiet failures
The interchange format, the order of operations, and the spot checks that catch the four conversion failures that actually happen.
-
Running consent for field recording, from approach to archive
Field consent is done in person, often in seconds. The approach, the form, the refusals, and the bystander problem.
-
Choosing a sample rate for a new collection
Record high, convert down, never the reverse. How to pick a rate for a new collection and verify what the recorder actually produces.
-
Keeping a provenance log that survives scrutiny
One row per file, recorded at capture, with fields that answer who, when, where and under what consent — years later.
-
Dereverberation before recognition: when it helps and when it does damage
Dereverberation helps far-field recognition only under specific conditions. The two strategies that work, and the failures to watch for.
-
A vendor risk questionnaire that actually discriminates
Most questionnaires collect reassurance. The questions that separate a serious data vendor from a reseller, and how to read the answers.
-
Measuring SNR on real recordings with no clean reference
When there is no clean reference, SNR is an estimate with assumptions. Three methods, their biases, and the checks that catch a bad number.
-
Drafting the DPA schedule for a speech data project
A speech data DPA schedule needs different details than a text one: listening access, audio retention, and subprocessors who only listen.
-
Building a noise robustness test set that supports a decision
A noise test set is a table, not a number. How to choose cells, sources, and sample counts that support an actual decision.
-
Getting voice actor consent for AI training, in practice
The briefing, the scope decisions, the withdrawal mechanics — and why the consent form is the least important part of the process.
-
librosa and torchaudio: which one for which job in a dataset pipeline
soundfile to read, torchaudio to train, librosa to verify. Where each library earns its place, and the defaults that cause drift.
-
Copyright diligence before you license a dataset
What to check, what documents to demand, and the red flags that mean walking away — before you sign a license for training data.
-
Resampling audio without breaking it
Anti-aliasing, integer ratios, and the batch bug that relabels every file without checking. What to verify after a conversion run.
-
Publishing a training data transparency disclosure
A disclosure buried in a slide deck is not a disclosure. Where to publish it, what the page carries, and how to keep it current.
-
Audio augmentation recipes that help, and the ones that waste compute
Which augmentation recipes earn their compute, which ones quietly waste it, and how to tell the difference on your own data.
-
An EU AI Act timeline for data teams, ordered by what cannot be recovered
Data work under the AI Act should be sequenced by what becomes impossible to recover later, not by which deadline is nearest.
-
MFCC or mel spectrogram in practice: picking a front end and pinning its numbers
Log-mel is the default front end now. The real work is pinning its parameters and proving two libraries produce the same features.
-
Writing a training data summary that is actually useful
A training data summary has to answer a stranger's questions without handing a competitor your roadmap. Drafting rules for both.
-
Computing WER in Python: what jiwer rewrites before it counts
jiwer does not score the strings you hand it. It scores transformed versions of them, and the transform can move the number further than a model change does.
-
Forced alignment with MFA: dictionary preparation, OOV words, and acceptance checks
Montreal Forced Aligner is only as good as the lexicon you feed it. How to prepare the dictionary, resolve out-of-vocabulary words, and check what came back.
-
The preprocessing pipeline for ASR: what each step does and why the order matters
A working order of operations for turning raw recordings into trainable audio, with the check that catches a problem at every step.
-
Running a teleoperation shift: rig setup, operator training, and the data you throw away
Leader-follower rigs, operator training, session length, and the failure patterns that show up in raw teleoperation data before any filtering.
-
Building a data governance file for EU AI Act Article 10
Article 10 asks for data governance, not paperwork. What goes in the file, who fills each section, and when it has to be updated.
-
How to choose an AI data vendor when there is no track record to check
A proposal tells you what a vendor can promise, not what they can deliver. Here are the dimensions that predict the difference, and how to test them early.
-
What a single WER number hides about the set behind it
One pooled score can sit on top of a test set where four speakers supply half the words. Here are the four breakdowns that turn the headline into a decision.
-
Forced alignment with WhisperX: handling words the aligner cannot place
WhisperX aligns the transcript that Whisper produced, so transcription errors become alignment errors. The per-word scores are how you find them.
-
Choosing a robot for data collection: buy for uptime, not for precision
Repeatability and payload matter less than duty cycle, repair time, and readable joint state. How to pick arms, grippers, and cameras for a collection rig.
-
Questions to ask before buying a speech dataset, and the red flags in the answers
A dataset listing describes what it contains. These questions get at what it leaves out, and each group of answers has a pattern worth stopping for.
-
CER or WER: which one to report, and when the choice flips a ranking
The two metrics are the same calculation with different units. Pick by writing system and by what the product is judged on, and keep the ratio as a diagnostic.
-
Aligning hour-long recordings: cut points, drift, and cross-chunk boundaries
Whole-file alignment fails slowly. Chunking bounds the damage, but only if the cuts come from the audio and you check the residual at every chunk.
-
Designing a pilot batch before you scale a robot data collection
How large a pilot should be, which task variants to include, and the thresholds that tell you whether to scale up or fix the pipeline first.
-
What a paid pilot batch should prove before you commit to the full project
A pilot is not a free sample. It is a test with pass and fail conditions written in advance, and it should prove four things a proposal cannot.
-
How to specify a speech data project so it does not get rejected
Most rejected deliveries trace back to the specification, not the production. Here is the shape of a specification that leaves nothing to interpretation.
-
Building an ASR evaluation harness you can re-run next quarter
A harness is a manifest, a frozen split, cached hypotheses and a report. Here is the order to build it in, and the four things worth freezing.
-
Running pyannote diarization: the parameters that matter and the merge with ASR
The pipeline defaults are tuned for meeting audio. Speaker count, min_duration_off, and the merge step are where your own data decides the result.
-
Quality checks for a robot episode: continuity, timestamps, dropped frames, drift
Four checks to run on every episode before annotation, with the signal each one catches and how to set its threshold.
-
How to write a data requirement brief that cannot be misread
Most specification failures happen at the reading stage, not the writing stage. A brief that survives a misreading test saves a re-annotation.
-
Why speaker count matters more than hours
Hours are the unit everyone quotes and the unit that predicts the least. Here is why the number of distinct speakers determines whether your model generalizes.
-
Normalize before you score, and record what you normalized away
Normalization can move a WER further than a model upgrade. Run the ablation, freeze one function, and stamp its version onto every score you publish.
-
Tuning VAD before diarization: settings that change the speaker output
A VAD change looks local and is not. Thresholds, padding, and minimum silence propagate into speaker turns, timestamps, and every score computed afterwards.
-
How many robot trajectories do you actually need
Published datasets range from tens of demonstrations per task to a million trajectories. What separates the two ends, and how to measure your own number.
-
Build or buy: the real cost of collecting training data in-house
Collecting data in-house looks cheaper per unit until coordination costs land on your engineers. How to compare the two honestly, and where the crossover sits.
-
Scoring code-switched audio without penalizing the model twice
A script mismatch and a genuine misrecognition score identically by default. Separate them, and find out whether the failure is coverage, boundary or script.
-
Reading a diarization error rate: what the three components tell you
The same DER can come from a weak VAD or a weak clustering model. Splitting the number into misses, false alarms, and confusion tells you which to fix.
-
Labeling robot trajectories: instructions, success labels, and how fine to go
Writing instructions a policy can use, deciding what counts as a failure, and what it costs to label at episode, phase, or frame granularity.
-
Licensing an existing dataset versus commissioning a new one
Both give you training data. They differ in what the agreement transfers, how long each takes, whether competitors hold the same audio, and what it can measure.
-
Code-switching: the data problem nobody scopes for
Most of the world mixes languages mid-sentence. If your dataset treats that as noise, it will not match how people actually speak.
-
Five ways a WER ends up better than the system really is
None of these require anyone to cheat. They are ordinary pipeline behavior, and each one has a check you can run in an afternoon with the files you already have.
-
Building a speaker-labeled transcript you can ship
Merging diarization with ASR is a data-model decision before it is a code decision. The format, the boundary cases, and a review loop that stays small.
-
Setting up egocentric video collection: rigs, angles, sync, and privacy
Head, chest, or wrist mounting, what field of view costs you, how to keep cameras aligned, and how to handle the people who walk into frame.
-
The costs that quietly accumulate in a data collection project
Re-collection, rework, format conversion, and delay rarely appear in a project quote. They appear in the final accounting, and they are predictable.
-
How big an ASR test set needs to be, and how to build it
Word count sets the noise floor, but clustered errors mean the effective sample size is closer to the speaker count. Here is how to compute both numbers.
-
Overlapping speech: why diarization fails there and how to record it
Two people talking at once break the one-label-per-frame assumption. What current tools do about it, and the annotation convention that keeps the data usable.
-
Closing the sim-to-real gap in practice: what to randomize and how much real data to add
Domain randomization is a knob, not a solution. What to vary, what to measure first, and how real episodes fit into a simulation-heavy training plan.
-
How to compare quotes for a data collection project fairly
Two quotes for the same project can differ by a factor and still describe identical work. The gap is usually the denominator, and normalizing it is the job.
-
Score a multilingual model per language, or the headline will mislead you
A pooled number follows the word counts, and the word counts follow whichever language has the most data. Here is the scorecard to publish instead.
-
Speaker verification with embeddings: setting a threshold you can defend
EER is a reporting metric, not an operating point. How to build the trial list, normalize the scores, and pick a threshold from a cost ratio.
-
Choosing a format for robot data: RLDS, HDF5, or LeRobot-style datasets
Three formats dominate robot learning, and each one is optimized for a different bottleneck: streaming, inspection, or sharing. What to store per step.
-
Planning a multilingual data project: sequence, volume, and consistency
Multilingual projects do not fail inside a language. They fail at the seams: sequencing, volumes set by round numbers, and guidelines that fork silently.
-
What GDPR actually requires for voice data
Voice recordings are personal data, and consent is not the only requirement. Here is what a defensible chain looks like in practice.
-
Measuring ASR latency and throughput without fooling yourself
Real-time factor is the easy number and the least useful on its own. Measure the warm-up, the batch curve, the beam trade-off and the time to first word.
-
Choosing a diarization toolkit: pyannote, NeMo, or a commercial API
The decision is made by your audio, not by a leaderboard. A validation protocol that takes about a day and answers the question on your own files.
-
Planning a dexterous manipulation dataset: hardware, tasks, and why the yield is low
Multi-fingered hands multiply every cost in a collection project. What to decide before buying hardware, and which tasks justify the effort.
-
Scoping a second-language batch: what transfers from the first
A second-language batch looks like a repeat of the first. It is a new project with recycled parts, and knowing which parts recycle is the scoping work.
-
Script splits: when one spoken language needs two datasets
Hindi and Urdu are mutually intelligible when spoken and mutually unreadable when written. That single fact reshapes how the data has to be built.
-
Buying data for low-resource languages: what changes
For a language with no existing corpus, everything about the project is different — starting with the fact that the timeline is set by speaker recruitment, not recording.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.