Sarah is a Head of AI at a Series B startup. Her CFO just asked her to put a number on a 10,000-sample RLHF dataset before next Tuesday's board meeting.
She opens four browser tabs. She finds vague ranges, vendor landing pages, and two academic papers from 2019. Nothing she can put in a budget line. Nothing she can defend to a CFO.
That's the problem this post solves. What follows is the most honest, specific breakdown of AI training data pricing available — with actual dollar figures, a cost comparison table, and a practical formula for budgeting your first dataset.
What Actually Drives AI Annotation Cost
Before the numbers: three variables move the cost needle more than anything else.
1. Task Complexity
Simple binary classification ("is this review positive or negative?") is cheap. Complex pairwise ranking ("which of these two LLM responses is more helpful, more accurate, and less harmful?") requires sustained cognitive effort. Long-form evaluation with written rationale — the gold standard for RLHF reward model training — is expensive because it takes 5–15 minutes per item.
The cost range across task types spans 500×. That's not a rounding error. Task design is the biggest budget decision you'll make.
2. Domain Expertise Required
Generalist annotators can handle sentiment tagging, image classification, and basic NLP work. They cannot reliably evaluate a medical diagnosis AI, a legal contract review model, or a financial risk scoring system. Domain-expert annotators — physicians, lawyers, CPAs, licensed engineers — command a significant premium because their expertise is irreplaceable and their supply is genuinely limited.
"Expert" is not a marketing label. It's a pay tier with real credentialing behind it. The data labeling cost per hour for a generalist annotator is $8–$15. For a domain expert in medicine or law, it's $50–$150+.
3. Volume and Turnaround Time
Rush orders cost more — especially for expert tasks. Crowdsourcing platforms can spin up 500 generalist annotators in 48 hours. Finding 50 verified oncologists for a radiology AI labeling task takes weeks even at premium rates. Volume discounts exist above ~5,000 items but are modest for expert work because the bottleneck is annotator supply, not platform capacity.
AI Training Data Pricing Ranges by Annotation Type
Here's what the market actually looks like. These are 2025–2026 ranges across major platforms including Scale AI, Labelbox, Surge AI, Prolific, and expert marketplaces.
| Annotation Type | Cost Range | Who Does It |
|---|---|---|
| Simple classification / tagging | $0.01–$0.05 / label | Crowdsourcing (MTurk, Toloka) — quality varies significantly |
| Standard NLP annotation | $0.10–$0.50 / item | Trained crowdsource, managed services — requires guidelines + QA |
| Expert pairwise ranking (RLHF) | $1–$5 / comparison pair | Skilled annotators, RLHF specialists — core cost for reward model training |
| Domain-expert evaluation (medical, legal, financial) | $5–$25 / item | Verified domain experts — most expensive; lowest volume available |
| Long-form evaluation with rationale | $8–$20 / item | Expert annotators — high signal for RLHF; 5–15 min per item |
Note on the RLHF tier: The $1–$5 range assumes competent generalist annotators evaluating general-purpose LLM outputs. If your model operates in a domain where wrong answers have real-world consequences — medical triage, legal interpretation, financial advice — you need the $5–$25 expert tier. Cutting corners here is how RLHF datasets produce reward models that confidently give incorrect professional advice.
The Hidden Costs Nobody Talks About
The line-item costs above are the visible budget. Here's what doesn't show up until later.
Rework Cost From Low IAA
Inter-annotator agreement (IAA) measures how consistently different annotators label the same item. When IAA is low, 15–30% of crowdsourced annotation budgets are wasted on unusable labels that fail IAA thresholds. On a $10,000 project, that's $1,500–$3,000 in labels you paid for and can't use — plus another 10–15% for the rework itself.
Retraining Cost From Wrong-Tier Annotators
Using generalist annotators for expert-level tasks doesn't produce noisy data. It produces confidently wrong data — labels that look clean but encode incorrect domain knowledge. The model learns false patterns with high confidence. The cost of catching this: a full retraining cycle at $10k–$50k in compute and engineering time. The "savings" from $0.05 labels evaporate.
Coordination Overhead on Anonymous Pools
Large crowdsourcing platforms use anonymous annotator pools. You can't inspect individual annotator quality or give targeted feedback to improve consistency. Quality control becomes a statistics problem — you're hoping the aggregate is good enough. This overhead manifests as extra rounds of guideline revision and calibration tasks. Add 20–40% to your timeline estimate.
Time Cost: Delayed Training Is Real Cost
For expert-level tasks, crowdsourcing platforms take 2–4 weeks to source, screen, and onboard qualified annotators. That's 2–4 weeks before your first labeled item comes back. If your model training is time-sensitive — a product launch, a competitive window, a funding milestone — that delay has a real dollar cost. AI data labeling pricing doesn't include the cost of being slow, but your CFO should.
Expert Annotators vs. Crowdsourcing — The ROI Math
A concrete example: a 1,000-sample RLHF pairwise ranking dataset.
Option A: Crowdsourcing
- Rate: $0.80 / pair
- Gross cost: $800
- Rework rate: 25%
- Extra labels for 1,000 usable: +$200
- Actual total: ~$1,000
- Timeline: 3 weeks
Option B: Expert Marketplace
- Rate: $1.20 / pair
- Gross cost: $1,200
- Rework rate: 5%
- Extra labels for 1,000 usable: +$60
- Actual total: ~$1,260
- Timeline: 1 week
The sticker price difference: $800 vs. $1,200. The real price difference: $1,000 vs. $1,260. The expert route costs 26% more per project. It delivers 3× faster. And it doesn't carry the risk of a retraining cycle if the annotation quality is wrong.
The cheapest per-label price is almost never the cheapest total cost.
The calculation shifts further for domain-expert tasks. A medical AI team that needs 500 annotated radiology reports is not choosing between $0.05 labels and $5 labels. They're choosing between labels that are wrong (cheap, fast, unusable) and labels that are right (expensive, slower, actually useful). That's not an ROI calculation. That's a build-vs-waste decision.
How to Budget Your First Dataset
If you're going back to your CFO with a number, here's the practical formula.
Run a 50-task pilot first (~$50–$250)
Never commit full budget to an annotation vendor without validating quality on your actual task type. A 50-task pilot against a known-answer calibration set tells you whether annotators understand the task and whether IAA is acceptable. If kappa comes back below 0.65, you've spent $100 to avoid spending $10,000 on bad data.
Budget 10% of total for QA and rework
Even expert annotators produce some rework. Labeling guidelines have ambiguities. Edge cases surface that your guidelines didn't anticipate. Budget a 10% QA buffer — on a $5,000 project, that's $500. If you don't need it, it's a pleasant surprise.
Factor in iteration cycles
First-round annotation is not the end of the annotation budget. As your model trains, you'll identify failure modes requiring additional targeted annotation. Expect 2–3 rounds of annotation over the life of a model. Budget accordingly: your initial dataset estimate should be ~40–50% of total annotation spend over a 12-month model development cycle.
Quick Reference: Budget Ranges
| Dataset Size | Task Type | Estimated Range |
|---|---|---|
| 50 items (pilot) | Any type | $50–$250 |
| 1,000 items | RLHF pairwise (generalist) | $800–$1,500 |
| 1,000 items | RLHF pairwise (expert) | $1,200–$5,000 |
| 10,000 items | Standard NLP annotation | $1,000–$5,000 |
| 10,000 items | Domain-expert evaluation | $50,000–$250,000 |