Data sourcing and brokerage
Any AI training data, sourced to order.
Tell us what you need — the data type, the languages if they matter, the volume, and the quality bar. We find the producer who can meet that specification, set the terms with both sides, and stay accountable through delivery.
Submit a sourcing request- Data categories
- 36
- Languages
- 120
- Inventory held
- Zero
What we do — and what we do not
The training data market has a matching problem. AI teams know what they need but not who can produce it; studios can produce it but cannot find that buyer. That gap is what we close.
We source
- Training data in any category — speech, text, image, video or embodied AI
- Rare and low-resource languages, including ones with no existing dataset
- Contributor recruitment, screening, collection and annotation
- Data sourced to the specification you write, at any volume
- Consent and licensing documents with every delivery
We do not
- Resell datasets we licensed from someone else
- Touch medical or clinical data
- Touch real phone calls or intercepted recordings
- Quote a price before we have your specification
How a project runs
-
Write the specification
Data type and volume, distinct contributors, languages and dialects where they apply, collection conditions, annotation depth, delivery format. In writing, before anything is collected.
-
Match to a producer
We go to producers we vet against the spec, confirm they can hit it, and get a real cost and timeline. If nobody can, we say so instead of improvising.
-
Pilot batch first
A small batch gets produced and annotated, and you review it before the full run begins. Problems that would cost a redo get caught after a few hours.
-
Deliver with documentation
Consent and licensing records, collection methodology, and the annotation guideline the data was produced under — so the dataset is usable, not just delivered.
What arrives in a delivery
A dataset is not just the files. Most of what decides whether a delivery is usable sits beside the data itself — and most suppliers treat it as an afterthought.
All eight, every delivery. The pilot batch comes first, so none of it is a surprise at the end.
-
The data files
Format, structure and segmentation set to your pipeline. Named to a convention you specify.
-
Labels or transcription
Orthographic, phonetic or categorical, produced under a guideline you review before production starts.
-
Contributor metadata
The attributes that let you slice the set — age band, gender, region, dialect background, or whatever your task turns on.
-
Collection conditions
Environment, device and, where it can be measured, signal quality — recorded per file.
-
Annotation guideline
The document the annotators worked from, so you can reproduce the conventions on your own data.
-
Consent records
Signed contributor consent covering your intended use, plus cross-border transfer terms where they apply.
-
Collection methodology
How contributors were recruited, screened and scheduled — the part that tells you how biased the pool is.
-
Quality report
Pilot outcome, re-work log, and the annotator agreement figures where the task supports measuring them.
Browse by data category
Speech, text, image, video, embodied AI, and the emerging paradigms — 36 categories in all. Each page explains what the data is, what buyers get wrong about it, and which specification detail most often decides whether a delivery is usable.
-
Multilingual Speech
Speech data covering multiple languages within one project, used for multilingual ASR, cross-lingual transfer, and language identification.
-
Conversational Speech
Speech from natural conversation between two or more people, on open or semi-structured topics, used for conversational AI, voice assistants, and small talk models.
-
Call Center Speech
Customer-service calls, simulated rather than intercepted, capturing both the agent and the caller side, used for call center QC, intent recognition, and dialogue systems.
-
Embodied AI Data
Multimodal data generated as robots and embodied agents operate in real environments — vision, action trajectories, and language instructions.
-
Synthetic Data
Data generated by models or simulation engines, used to fill gaps where real data is scarce, cover long-tail scenarios, and control cost.
-
World Model Data
Environment interaction data for training world models, emphasizing temporal consistency, physical plausibility, and predictability.
-
LLM Training Data
Text corpora for LLM pretraining and fine-tuning — instruction data, preference data, and multi-turn dialogue data.
-
NLP Data
Annotated text for NLP tasks — named entities, dependency syntax, sentiment polarity, and relation extraction.
-
Translation Data
Parallel corpora pairing source and target language, used for machine translation training and evaluation.
Browse by language
One axis of one category. Every language page documents what makes that language genuinely hard to collect — tone systems, dialect boundaries, script splits, or a speaker pool too small to fill an order.
-
Hindi
Hindi and Urdu are mutually intelligible at the spoken level, but one is written in Devanagari and the other in Perso-Arabic script. Which script a recording is transcribed into directly determines whether the data can feed a model targeting Urdu.
-
Arabic (MSA)
Almost nobody speaks Modern Standard Arabic as a native, everyday language. The practical result of recording standard Arabic is that every speaker carries their own dialect accent — unless the speaker's dialect background and the tolerated deviation are defined up front, annotation consistency collapses.
-
Indonesian
In everyday speech, Indonesians mix in local languages (Javanese, Sundanese) and English loanwords heavily. Three language components inside a single sentence is normal, so word segmentation and annotation rules have to be fixed before collection begins.
-
Thai
Thai has five tones, and Thai script carries no word spacing — a sentence is one long run of continuous characters. There is no single right answer for where word boundaries go, which lets different annotators produce mutually incompatible corpora.
-
Turkish
Turkish is an agglutinative language: one root can carry a long chain of suffixes, so the word list explodes. Istanbul and eastern accents also differ noticeably — if speaker origin is not recorded clearly, model generalization suffers.
-
Vietnamese
Vietnamese has six tones, marked in writing by a heavy load of diacritics. The three major dialect regions — north, central, and south — realize tones quite differently, and mixing them in one batch gives the acoustic model contradictory tone cues.
-
Filipino
Filipinos speak Taglish in daily life — switching between Tagalog and English sentence by sentence, sometimes word by word. This is how native speakers naturally talk, not a sign of substandard speech. Anyone who wants authentic data has to treat this mixing as the target, not as noise.
-
Persian
Written Persian and spoken Tehrani Persian are far apart — verb endings, pronouns, and connected speech all shift in the colloquial form. Transcribing spoken recordings against the written form produces text that does not match the audio.
-
Tamil
Written and spoken Tamil are two separate systems, and colloquial speech drops or merges a large share of sounds. Accents also differ across Tamil Nadu in India, Sri Lanka, and Malaysia, so speaker region has to be annotated on every record.
-
Telugu
Telugu has four major dialect regions (Coastal Andhra, Rayalaseema, Telangana, and the south) that differ noticeably, and the Telangana dialect carries a large share of Urdu loanwords. Without region labels, you cannot tell dialect variation from annotation error.
-
Malayalam
Malayalam is spoken fast, with dense connected speech, and public speech corpora are far smaller than for Indian languages of comparable size. The pool of qualified annotators is small, so a single project easily reuses the same people — speaker deduplication is mandatory.
-
Nepali
Usable Nepali speech data is surprisingly scarce, and the accent layers of speakers across borders (Nepal, Sikkim in India, Bhutan) are poorly mapped. Speaker metadata — nationality, home region, language of education — has to be recorded item by item.
-
Sinhala
Everyday Sinhala conversation is heavily embedded with English words, especially in the Colombo urban belt. Treating this mixed speech as pure Sinhala strips out high-frequency loanwords as if they were errors.
-
Bangla
The Bangla of Bangladesh and of West Bengal in India has diverged in vocabulary and accent — the everyday word for the same meaning is often different. In search terms and annotation alike, bangla tracks real usage better than bengali.
-
Khmer
Khmer script has no spaces between words, and it contains many letters that are written but not pronounced. The transcription convention has to be set first: write as spelled (orthographic) or write as spoken (phonemic). Data produced under the two conventions cannot be mixed.
Quality and compliance, collected while the data is made
Consent and provenance cannot be reconstructed after a collection run. They are produced alongside the data itself, which is why they hold up when a model is audited.
-
Signed contributor consent
States the intended use concretely, rather than a general release.
-
Collection methodology
How contributors were recruited, screened and scheduled, so you can judge how biased the pool is.
-
Annotation guideline
The document annotators worked from, so the conventions are reproducible.
-
Collection conditions
Per file: environment, device, and signal quality where it can be measured.
-
Cross-border transfer terms
If the data crosses a border, the agreement states how it may be moved — not left unsaid.
How quality is enforced during the run
- You approve a pilot batch before production starts. A misalignment in the specification surfaces after a few hours of work, not after the full run is delivered.
- Checking happens as data comes in. Not as a final pass over the finished set, where the only remaining fix is to collect it all over again.
- Annotator agreement is measured wherever the task supports it, and the figures go into the delivery report.
Where buyers start
-
AI Data Brokerage
A data broker sits between the people who need training data and the people who can produce it. We source to order — no inventory, no license resale, no recycled datasets.
-
AI Training Data Providers
What to check before you sign with a training data provider — and how we compare on each point.
-
AI Training Data Marketplace
Marketplaces sell what already exists. We source what you actually need. Here is when each model makes sense.
-
Buy AI Training Data
A practical guide to buying training data without overpaying for hours you cannot use.
-
AI Data Licensing
What the license actually permits is more important than the dataset specification. Here is what to check.
-
Embodied AI Data Collection
Collection for robots and physical agents, scoped as a task matrix rather than an hour count. What the project involves, what moves the cost, and what lands in your hands at delivery.
-
Robot Data Collection Services
Buying robot data collection as a service means renting someone else's rig, operators, and scenes. How to judge a provider, and what the contract has to settle before work starts.
-
Physical AI Data
Physical AI is the umbrella term for systems that act in the world — robots, vehicles, drones, and the human demonstration data that teaches them. The choice that shapes the budget is machine-collected or human-collected.
-
Sell AI Training Data
If you have recording capability, annotation capacity, or language access that AI teams need, we want to hear from you.
-
Sell Data to AI Companies
How AI companies actually buy data, the routes a supplier can realistically use, and the reasons most first approaches fail before anyone looks at the data.
-
Sell Data to AI Labs
Labs buy differently from product companies: narrower specifications, smaller volumes, and a higher bar on provenance. What a supplier should expect from that process.
-
Sell Data to OpenAI
How large labs procure data, what public routes actually exist, and what a supplier should realistically expect. This page claims no relationship with OpenAI.
-
Sell My Own Data
What an individual actually has that AI companies pay for, what the offers promising money for your personal data really are, and the routes where one person does get paid.
-
Is Selling Data Profitable
For most data, no — it is a commodity with commodity margins. For data that is genuinely hard to collect, yes. The difference between the two is the whole business.
-
Sell Speech Data
Speech is our home category, so this page can be specific: what buyers ask for, what a supplier needs to have, and which kinds of audio are unsellable no matter how good they sound.
-
Sell Training Data
Training data is a category, not a product. Some of it is free, some of it is a service business, and the part that is neither is where suppliers get paid.
-
Voice Data Supplier
For studios, agencies, and teams that can produce voice recordings to a specification — what buyers check, how repeat work actually arrives, and what ends a supplier relationship.
-
Data Supplier Program
Supplier programs exist across the data industry. What they actually are, what to check before joining one, and how our own supplier onboarding works — no portal, no registry.
New to buying training data?
These terms come up in almost every specification conversation. Each one is explained in plain language, without assuming you already work in the field.
- ASR — Automatic Speech Recognition — the task of turning recorded speech into text.
- TTS — Text-to-Speech — the task of generating spoken audio from written text.
- WER — Word Error Rate — the standard accuracy metric for speech recognition. Lower is better.
- Transcription — Writing down what is said in a recording. It is the core annotation task in any speech dataset.
- Annotation — Attaching machine-readable labels to raw data — a transcript, an intent, a speaker identity.
- Speaker Identification — Identifying who is speaking from the voice itself. Also called voice biometrics.
- Diarization — Marking which speaker said which segment in a multi-speaker recording.
- SNR — Signal-to-Noise Ratio — how much louder the speech is than the background noise, measured in decibels (dB).
Questions buyers ask before the first call
These come up in almost every first conversation. If yours is not here, ask it directly — you will get a straight answer rather than a brochure.
More detail
How long does a project take?
Where a language has an established producer pool, a quote usually takes a few days and production is scheduled from the agreed start date. For a low-resource language or an unusual specification, speaker recruitment is the step that sets the timeline and it is genuinely hard to predict in advance. We will give you an honest range rather than a date we cannot hold.
What is a pilot batch?
A small batch recorded and annotated before the full run begins, which you review and can reject. It costs little and it surfaces specification gaps while there is still time to fix them.
What if the delivery does not meet specification?
Re-work is covered in the agreement before production starts. Because the specification is written down and the pilot batch is approved, disputes about what was promised are rare — and when they occur, the specification is the reference.
Do you publish prices?
No. Every project is quoted from its specification. Hours, distinct speaker count, recording conditions and annotation depth all change the cost substantially, so a published number would be fiction.
Can I get a sample before committing?
Yes. Select a free sample in the request form. For most projects the pilot batch serves this purpose and is more informative than a generic sample file.
Do you work with suppliers as well as buyers?
Yes. If you have recording capability, annotation capacity or native-speaker access in a language we cover, see the sell AI training data page. Sell AI training data.
Have a specification? Send it.
Language, hours, and what the data needs to look like. If it is sourceable we will quote it; if it is not, we will say so rather than improvise.
Submit a sourcing request
Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.
- Pilot batch before the full run, so problems surface early.
- Consent documentation delivered with the data.
- No medical or clinical data. No recorded telephone calls.
- You send the spec Language, hours, recording conditions, and the deadline.
- We reply within two business days A real number and a real timeline — not a range.
- You approve a pilot batch A small run first, so problems surface before the full order.
- Full run, then delivery with documents Audio, transcripts, and the consent record for every speaker.