Data sourcing and brokerage

Any AI training data, sourced to order.

Tell us what you need — the data type, the languages if they matter, the volume, and the quality bar. We find the producer who can meet that specification, set the terms with both sides, and stay accountable through delivery.

Submit a sourcing request
Data categories
36
Languages
120
Inventory held
Zero

What we do — and what we do not

The training data market has a matching problem. AI teams know what they need but not who can produce it; studios can produce it but cannot find that buyer. That gap is what we close.

We source

  • Training data in any category — speech, text, image, video or embodied AI
  • Rare and low-resource languages, including ones with no existing dataset
  • Contributor recruitment, screening, collection and annotation
  • Data sourced to the specification you write, at any volume
  • Consent and licensing documents with every delivery

We do not

  • Resell datasets we licensed from someone else
  • Touch medical or clinical data
  • Touch real phone calls or intercepted recordings
  • Quote a price before we have your specification
Two colleagues standing at an audio mixing desk, one pointing at the controls

How a project runs

  1. Write the specification

    Data type and volume, distinct contributors, languages and dialects where they apply, collection conditions, annotation depth, delivery format. In writing, before anything is collected.

  2. Match to a producer

    We go to producers we vet against the spec, confirm they can hit it, and get a real cost and timeline. If nobody can, we say so instead of improvising.

  3. Pilot batch first

    A small batch gets produced and annotated, and you review it before the full run begins. Problems that would cost a redo get caught after a few hours.

  4. Deliver with documentation

    Consent and licensing records, collection methodology, and the annotation guideline the data was produced under — so the dataset is usable, not just delivered.

Rows of channel faders and knobs on a large-format audio mixing console

What arrives in a delivery

A dataset is not just the files. Most of what decides whether a delivery is usable sits beside the data itself — and most suppliers treat it as an afterthought.

All eight, every delivery. The pilot batch comes first, so none of it is a surprise at the end.

  1. The data files

    Format, structure and segmentation set to your pipeline. Named to a convention you specify.

  2. Labels or transcription

    Orthographic, phonetic or categorical, produced under a guideline you review before production starts.

  3. Contributor metadata

    The attributes that let you slice the set — age band, gender, region, dialect background, or whatever your task turns on.

  4. Collection conditions

    Environment, device and, where it can be measured, signal quality — recorded per file.

  5. Annotation guideline

    The document the annotators worked from, so you can reproduce the conventions on your own data.

  6. Consent records

    Signed contributor consent covering your intended use, plus cross-border transfer terms where they apply.

  7. Collection methodology

    How contributors were recruited, screened and scheduled — the part that tells you how biased the pool is.

  8. Quality report

    Pilot outcome, re-work log, and the annotator agreement figures where the task supports measuring them.

Browse by data category

Speech, text, image, video, embodied AI, and the emerging paradigms — 36 categories in all. Each page explains what the data is, what buyers get wrong about it, and which specification detail most often decides whether a delivery is usable.

  • Rows of channel faders and knobs on a large-format audio mixing console

    Multilingual Speech

    Speech

    Speech data covering multiple languages within one project, used for multilingual ASR, cross-lingual transfer, and language identification.

  • Two colleagues standing at an audio mixing desk, one pointing at the controls

    Conversational Speech

    Speech

    Speech from natural conversation between two or more people, on open or semi-structured topics, used for conversational AI, voice assistants, and small talk models.

  • A speaker in profile wearing over-ear headphones at a desk microphone in daylight

    Call Center Speech

    Speech

    Customer-service calls, simulated rather than intercepted, capturing both the agent and the caller side, used for call center QC, intent recognition, and dialogue systems.

  • A row of closed server cabinets along a data centre wall

    Embodied AI Data

    Emerging paradigm

    Multimodal data generated as robots and embodied agents operate in real environments — vision, action trajectories, and language instructions.

  • Dense bundles of network patch cables terminating at a patch panel

    Synthetic Data

    Emerging paradigm

    Data generated by models or simulation engines, used to fill gaps where real data is scarce, cover long-tail scenarios, and control cost.

  • Fibre optic patch cords plugged into a blue termination panel

    World Model Data

    Emerging paradigm

    Environment interaction data for training world models, emphasizing temporal consistency, physical plausibility, and predictability.

  • A long aisle of server racks receding into the distance in a data centre

    LLM Training Data

    Non-speech

    Text corpora for LLM pretraining and fine-tuning — instruction data, preference data, and multi-turn dialogue data.

  • Dense bundles of network patch cables terminating at a patch panel

    NLP Data

    Non-speech

    Annotated text for NLP tasks — named entities, dependency syntax, sentiment polarity, and relation extraction.

  • Rows of channel faders and knobs on a large-format audio mixing console

    Translation Data

    Non-speech

    Parallel corpora pairing source and target language, used for machine translation training and evaluation.

All 36 data categories →

Browse by language

One axis of one category. Every language page documents what makes that language genuinely hard to collect — tone systems, dialect boundaries, script splits, or a speaker pool too small to fill an order.

  • Hindi

    South Asia · Devanagari

    Hindi and Urdu are mutually intelligible at the spoken level, but one is written in Devanagari and the other in Perso-Arabic script. Which script a recording is transcribed into directly determines whether the data can feed a model targeting Urdu.

    • Dual script
  • Arabic (MSA)

    Middle East & North Africa · Arabic

    Almost nobody speaks Modern Standard Arabic as a native, everyday language. The practical result of recording standard Arabic is that every speaker carries their own dialect accent — unless the speaker's dialect background and the tolerated deviation are defined up front, annotation consistency collapses.

    • Dialect split
  • Indonesian

    Southeast Asia · Latin

    In everyday speech, Indonesians mix in local languages (Javanese, Sundanese) and English loanwords heavily. Three language components inside a single sentence is normal, so word segmentation and annotation rules have to be fixed before collection begins.

    • Code-switching
  • Thai

    Southeast Asia · Thai

    Thai has five tones, and Thai script carries no word spacing — a sentence is one long run of continuous characters. There is no single right answer for where word boundaries go, which lets different annotators produce mutually incompatible corpora.

    • Tone system
  • Turkish

    Middle East & Europe · Latin

    Turkish is an agglutinative language: one root can carry a long chain of suffixes, so the word list explodes. Istanbul and eastern accents also differ noticeably — if speaker origin is not recorded clearly, model generalization suffers.

    • Dialect split
  • Vietnamese

    Southeast Asia · Latin (diacritics)

    Vietnamese has six tones, marked in writing by a heavy load of diacritics. The three major dialect regions — north, central, and south — realize tones quite differently, and mixing them in one batch gives the acoustic model contradictory tone cues.

    • Tone system
  • Filipino

    Southeast Asia · Latin

    Filipinos speak Taglish in daily life — switching between Tagalog and English sentence by sentence, sometimes word by word. This is how native speakers naturally talk, not a sign of substandard speech. Anyone who wants authentic data has to treat this mixing as the target, not as noise.

    • Code-switching
  • Persian

    Middle East & North Africa · Arabic (Perso-Arabic)

    Written Persian and spoken Tehrani Persian are far apart — verb endings, pronouns, and connected speech all shift in the colloquial form. Transcribing spoken recordings against the written form produces text that does not match the audio.

    • Dialect split
  • Tamil

    South Asia · Tamil

    Written and spoken Tamil are two separate systems, and colloquial speech drops or merges a large share of sounds. Accents also differ across Tamil Nadu in India, Sri Lanka, and Malaysia, so speaker region has to be annotated on every record.

    • Dialect split
  • Telugu

    South Asia · Telugu

    Telugu has four major dialect regions (Coastal Andhra, Rayalaseema, Telangana, and the south) that differ noticeably, and the Telangana dialect carries a large share of Urdu loanwords. Without region labels, you cannot tell dialect variation from annotation error.

    • Dialect split
  • Malayalam

    South Asia · Malayalam

    Malayalam is spoken fast, with dense connected speech, and public speech corpora are far smaller than for Indian languages of comparable size. The pool of qualified annotators is small, so a single project easily reuses the same people — speaker deduplication is mandatory.

    • Scarce speakers
  • Nepali

    South Asia · Devanagari

    Usable Nepali speech data is surprisingly scarce, and the accent layers of speakers across borders (Nepal, Sikkim in India, Bhutan) are poorly mapped. Speaker metadata — nationality, home region, language of education — has to be recorded item by item.

    • Scarce speakers
  • Sinhala

    South Asia · Sinhala

    Everyday Sinhala conversation is heavily embedded with English words, especially in the Colombo urban belt. Treating this mixed speech as pure Sinhala strips out high-frequency loanwords as if they were errors.

    • Code-switching
  • Bangla

    South Asia · Bengali

    The Bangla of Bangladesh and of West Bengal in India has diverged in vocabulary and accent — the everyday word for the same meaning is often different. In search terms and annotation alike, bangla tracks real usage better than bengali.

    • Dual script
  • Khmer

    Southeast Asia · Khmer

    Khmer script has no spaces between words, and it contains many letters that are written but not pronounced. The transcription convention has to be set first: write as spelled (orthographic) or write as spoken (phonemic). Data produced under the two conventions cannot be mixed.

    • Script conversion

All 120 languages →

Quality and compliance, collected while the data is made

Consent and provenance cannot be reconstructed after a collection run. They are produced alongside the data itself, which is why they hold up when a model is audited.

  • Signed contributor consent

    States the intended use concretely, rather than a general release.

  • Collection methodology

    How contributors were recruited, screened and scheduled, so you can judge how biased the pool is.

  • Annotation guideline

    The document annotators worked from, so the conventions are reproducible.

  • Collection conditions

    Per file: environment, device, and signal quality where it can be measured.

  • Cross-border transfer terms

    If the data crosses a border, the agreement states how it may be moved — not left unsaid.

How quality is enforced during the run

  1. You approve a pilot batch before production starts. A misalignment in the specification surfaces after a few hours of work, not after the full run is delivered.
  2. Checking happens as data comes in. Not as a final pass over the finished set, where the only remaining fix is to collect it all over again.
  3. Annotator agreement is measured wherever the task supports it, and the figures go into the delivery report.

Where buyers start

  • AI Data Brokerage

    A data broker sits between the people who need training data and the people who can produce it. We source to order — no inventory, no license resale, no recycled datasets.

    • Buying
  • AI Training Data Providers

    What to check before you sign with a training data provider — and how we compare on each point.

    • Buying
  • AI Training Data Marketplace

    Marketplaces sell what already exists. We source what you actually need. Here is when each model makes sense.

    • Buying
  • Buy AI Training Data

    A practical guide to buying training data without overpaying for hours you cannot use.

    • Buying
  • AI Data Licensing

    What the license actually permits is more important than the dataset specification. Here is what to check.

    • Buying
  • Embodied AI Data Collection

    Collection for robots and physical agents, scoped as a task matrix rather than an hour count. What the project involves, what moves the cost, and what lands in your hands at delivery.

    • Buying
  • Robot Data Collection Services

    Buying robot data collection as a service means renting someone else's rig, operators, and scenes. How to judge a provider, and what the contract has to settle before work starts.

    • Buying
  • Physical AI Data

    Physical AI is the umbrella term for systems that act in the world — robots, vehicles, drones, and the human demonstration data that teaches them. The choice that shapes the budget is machine-collected or human-collected.

    • Buying
  • Sell AI Training Data

    If you have recording capability, annotation capacity, or language access that AI teams need, we want to hear from you.

    • Selling
  • Sell Data to AI Companies

    How AI companies actually buy data, the routes a supplier can realistically use, and the reasons most first approaches fail before anyone looks at the data.

    • Selling
  • Sell Data to AI Labs

    Labs buy differently from product companies: narrower specifications, smaller volumes, and a higher bar on provenance. What a supplier should expect from that process.

    • Selling
  • Sell Data to OpenAI

    How large labs procure data, what public routes actually exist, and what a supplier should realistically expect. This page claims no relationship with OpenAI.

    • Selling
  • Sell My Own Data

    What an individual actually has that AI companies pay for, what the offers promising money for your personal data really are, and the routes where one person does get paid.

    • Selling
  • Is Selling Data Profitable

    For most data, no — it is a commodity with commodity margins. For data that is genuinely hard to collect, yes. The difference between the two is the whole business.

    • Selling
  • Sell Speech Data

    Speech is our home category, so this page can be specific: what buyers ask for, what a supplier needs to have, and which kinds of audio are unsellable no matter how good they sound.

    • Selling
  • Sell Training Data

    Training data is a category, not a product. Some of it is free, some of it is a service business, and the part that is neither is where suppliers get paid.

    • Selling
  • Voice Data Supplier

    For studios, agencies, and teams that can produce voice recordings to a specification — what buyers check, how repeat work actually arrives, and what ends a supplier relationship.

    • Selling
  • Data Supplier Program

    Supplier programs exist across the data industry. What they actually are, what to check before joining one, and how our own supplier onboarding works — no portal, no registry.

    • Selling

New to buying training data?

These terms come up in almost every specification conversation. Each one is explained in plain language, without assuming you already work in the field.

  • ASR — Automatic Speech Recognition — the task of turning recorded speech into text.
  • TTS — Text-to-Speech — the task of generating spoken audio from written text.
  • WER — Word Error Rate — the standard accuracy metric for speech recognition. Lower is better.
  • Transcription — Writing down what is said in a recording. It is the core annotation task in any speech dataset.
  • Annotation — Attaching machine-readable labels to raw data — a transcript, an intent, a speaker identity.
  • Speaker Identification — Identifying who is speaking from the voice itself. Also called voice biometrics.
  • Diarization — Marking which speaker said which segment in a multi-speaker recording.
  • SNR — Signal-to-Noise Ratio — how much louder the speech is than the background noise, measured in decibels (dB).

Full glossary →

Questions buyers ask before the first call

These come up in almost every first conversation. If yours is not here, ask it directly — you will get a straight answer rather than a brochure.

More detail

How long does a project take?

Where a language has an established producer pool, a quote usually takes a few days and production is scheduled from the agreed start date. For a low-resource language or an unusual specification, speaker recruitment is the step that sets the timeline and it is genuinely hard to predict in advance. We will give you an honest range rather than a date we cannot hold.

What is a pilot batch?

A small batch recorded and annotated before the full run begins, which you review and can reject. It costs little and it surfaces specification gaps while there is still time to fix them.

What if the delivery does not meet specification?

Re-work is covered in the agreement before production starts. Because the specification is written down and the pilot batch is approved, disputes about what was promised are rare — and when they occur, the specification is the reference.

Do you publish prices?

No. Every project is quoted from its specification. Hours, distinct speaker count, recording conditions and annotation depth all change the cost substantially, so a published number would be fiction.

Can I get a sample before committing?

Yes. Select a free sample in the request form. For most projects the pilot batch serves this purpose and is more informative than a generic sample file.

Do you work with suppliers as well as buyers?

Yes. If you have recording capability, annotation capacity or native-speaker access in a language we cover, see the sell AI training data page. Sell AI training data.

Have a specification? Send it.

Language, hours, and what the data needs to look like. If it is sourceable we will quote it; if it is not, we will say so rather than improvise.

Submit a sourcing request

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

  1. You send the spec Language, hours, recording conditions, and the deadline.
  2. We reply within two business days A real number and a real timeline — not a range.
  3. You approve a pilot batch A small run first, so problems surface before the full order.
  4. Full run, then delivery with documents Audio, transcripts, and the consent record for every speaker.
Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email [email protected]