1. Why RFPs Fail for AI Annotation
Most annotation RFPs are adapted from IT procurement templates. They ask about ISO certifications, years in business, client references from Fortune 500 companies, and SLA uptime guarantees. None of these tell you anything useful about annotation quality for RLHF or LLM training data.
Three specific failure modes show up in almost every annotation RFP that produces a bad vendor selection:
Asking about years of experience instead of IAA data
"How long have you been doing AI annotation?" is irrelevant. A vendor can have eight years of image classification experience and zero credible RLHF methodology. The right question is: "Can you share inter-annotator agreement data (Cohen's κ) from a completed preference annotation project with comparable task complexity?" A vendor who answers with marketing copy instead of a κ number has told you everything you need to know.
Pricing before quality guarantees
Standard procurement asks for pricing early in the RFP process — often before any discussion of methodology. This sequence is backwards for annotation. If you anchor on price before establishing quality standards, you end up comparing vendors on cost with no common quality baseline. The result is a dataset with κ = 0.45 that cost 40% less than a dataset with κ = 0.75 — and a reward model trained on noise that costs 10× more to retrain than the annotation savings were worth.
Missing the calibration pilot requirement
Enterprise IT procurement almost never includes a mandatory calibration phase — you buy the software, you deploy it, you live with it. Annotation is different. A calibration pilot (20–100 examples, IAA review before production begins) is not optional overhead — it is the quality gate that prevents 30–40% rework on production batches. Any RFP that doesn't include calibration pilot terms in the contract structure is setting up a project failure.
The template and rubric below are built around these three corrections. They ask the right questions in the right order, gate pricing on quality commitments, and treat the calibration pilot as a contractual requirement, not a vendor courtesy.
2. Before You Send the RFP: What to Have Ready
An RFP is only as good as the inputs you give vendors to respond to. If you send a vague brief, you get vague vendor responses — and you end up scoring vendors on marketing polish rather than actual capability. Before the RFP goes out, you need five things locked down.
A completed annotation brief
The brief is the primary input to the RFP. It specifies task definition, input/output format, quality criteria rubric, edge case handling, IAA targets, annotator requirements, and timeline. If you don't have a complete brief, your RFP responses will be unscoreable — vendors will give you capability statements instead of project-specific answers. See our complete guide to writing an annotation brief before sending this RFP.annotation brief guide
Task taxonomy
What type(s) of annotation does this project require? Preference annotation (RLHF pairwise), safety flagging, SFT completion writing, named entity recognition, image classification, or a combination? Vendors have different capability profiles by task type. A vendor excellent at crowdsourced NER may have no credible RLHF methodology. Your task taxonomy determines which vendors belong on the shortlist.
IAA target by dimension
Your minimum acceptable Cohen's κ (or Fleiss' κ) for each annotation dimension. Example: κ ≥ 0.70 for preference selection, κ ≥ 0.85 for safety flags, κ ≥ 0.65 for helpfulness ratings. These numbers become the acceptance criteria in the contract. Any vendor who won't commit to measurable IAA targets is not a serious vendor for quality-sensitive annotation.
Budget range
Not a single number — a range. "$15,000–$25,000 for 5,000 preference pairs" gives vendors enough context to propose the right methodology without anchoring to the floor. If budget is listed as "competitive" or "to be determined," vendors will quote whatever they think your organization can absorb. For context on what annotation budgets should look like by task type and volume, see our AI annotation pricing guide.
Timeline and volume
Total annotation count, production start date, milestone checkpoints, and final delivery deadline. A vendor who can't respond to specific volume and timeline inputs isn't operationally ready for your project — they're still selling to you, not scoping for you. Require vendors to confirm capacity for your specific volume window, not generic throughput claims.
3. The 6 Evaluation Dimensions
Every annotation vendor should be evaluated across the same six dimensions. These are not generic procurement criteria — they are the specific quality and operational factors that determine whether an annotation vendor delivers data your models can actually train on.
1. Annotator quality
The two questions that matter: What is the domain expertise level of annotators assigned to tasks like mine? And what is their measured κ track record on comparable tasks? "Vetted annotators" is not an answer. κ ≥ 0.72 on medical Q&A preference annotation with clinical reviewer adjudication is an answer. Require specific data, not general claims.
2. Quality methodology
How does the vendor aggregate labels across annotators? For binary classification, majority vote is defensible. For preference annotation and RLHF pairs, majority vote is the wrong methodology — it averages away disagreement rather than resolving it. Consensus with documented rationale from a senior annotator produces more reliable signal. Require the vendor to describe their methodology for your specific task type, not their general process.
3. RLHF / preference annotation capability
RLHF preference annotation is not standard annotation work. It requires annotators who understand model outputs well enough to make reliable comparative judgments, a preference pair format with documented rationale, and an IAA measurement approach that accounts for subjective quality dimensions. Ask for specific examples: How many preference pairs have they delivered? What was measured κ? Can they share anonymized examples of their rationale format?
4. Pricing model
Four models exist: per-label (cheapest headline, most opaque quality incentive), per-hour (transparent but variable), per-project (fixed scope, cleaner budgeting), and platform subscription (best for internal teams with BYO workforce). Each has different implications for your budget and the vendor's quality incentives. A per-label vendor has a financial incentive to maximize throughput; a per-project vendor has an incentive to deliver the agreed scope. Understand the model before comparing prices.
5. Turnaround and scalability
Time to first batch (can you start this week?), peak throughput (can they scale if you need to triple volume mid-project?), and pilot availability (will they do a calibration pilot on your schedule?). Vendors who can't commit to your specific start date or can't confirm capacity for your volume window are operationally unready for your project, regardless of their quality claims.
6. Data security and compliance
SOC 2 Type II certification is the minimum standard for production training data — it means independent auditors have verified the vendor's security controls. For healthcare data, HIPAA BAA. For EU data, GDPR DPA. NDA process and time to execute matters if your prompts and completions contain confidential information. For commodity annotation this may be a lower-stakes requirement; for production RLHF data that contains sensitive user queries, it is not.
For a full comparison of how major vendors score across these dimensions, see our roundups: Scale AI alternatives, Labelbox alternatives, and Appen alternatives. For context on what expert vs. crowdsourced annotators deliver on these dimensions, see our domain experts vs. crowdsourcing guide.
4. Full RFP Template (Ready to Send)
The following 20-question RFP is formatted for direct use. Customize the bracketed fields with your project specifics, remove sections that don't apply to your task type, and send as a document or email to your vendor shortlist. Responses to this RFP should be scoreable using the rubric in Section 5.
AI Annotation Vendor RFP — [Your Company Name]
Project: [Brief project description] · Response deadline: [Date]
A. Vendor Overview
Q1.Provide a one-paragraph description of your organization's annotation services, focusing specifically on the task types most relevant to this project: [preference annotation / safety flagging / SFT completion writing / other].
Q2.How many annotation projects have you completed in the last 24 months for AI/ML training data specifically? What was the average project volume (total labels)? What percentage involved RLHF or preference annotation?
Q3.Provide two client references for completed projects with comparable task type and volume. Include contact name, company, project scope, and approximate completion date.
B. Annotator Qualifications
Q4.What are the qualifications of annotators you would assign to this specific project? Describe domain expertise level, credentials, and language proficiency. Do not describe your general annotator pool — describe the subset you would use for [task type].
Q5.How many annotators would be assigned to this project? How do you ensure consistency across annotators working on the same task?
Q6.Provide inter-annotator agreement data (Cohen's κ or Fleiss' κ) from a completed project with comparable task complexity. Specify the task type, number of annotators, and measured κ by dimension.
C. Quality Assurance Process
Q7.Describe your label aggregation methodology for preference annotation tasks. Do you use majority vote, consensus with rationale, senior annotator adjudication, or another approach? Explain the reasoning behind your choice.
Q8.What is your IAA measurement process? How frequently do you calculate inter-annotator agreement during production? What is your protocol when a batch falls below the agreed IAA threshold?
Q9.How do you handle annotators who consistently fall below IAA threshold during a project? What is your replacement and rework protocol?
Q10.Confirm that your team will complete a calibration pilot of [20–100 examples] before production begins, with IAA review prior to proceeding. Describe your calibration process.
D. RLHF / Preference Annotation Experience
Q11.Describe your specific experience with RLHF preference pair annotation. How many preference pairs have you produced in total? What was the most complex preference annotation task you've completed, and what κ did you achieve?
Q12.What is your rationale format for preference pairs? Do annotators provide written justification for their preference selection? Provide an anonymized example of a completed preference annotation with rationale.
Q13.Describe your experience producing SFT (supervised fine-tuning) instruction-completion examples. What quality review process do you use for SFT completions, and what is your target quality score?
E. Pricing Model and Structure
Q14.Provide an itemized price estimate for this project: [volume] [task type] with the following requirements: [annotator qualifications], [IAA targets], [calibration pilot included]. Specify whether pricing is per-label, per-hour, per-project, or another model.
Q15.What is included in this price? Specifically: Is the calibration pilot included or priced separately? Is IAA reporting included? Are re-annotation costs for batches that fail IAA threshold included or billed separately?
Q16.What are your payment terms? Do you require a deposit? What is your invoicing schedule for a multi-batch project?
F. Project Timeline and Scalability
Q17.Confirm that you can begin the calibration pilot within [X business days] of contract execution. Provide a proposed project schedule with batch checkpoints for the following total volume: [volume and deadline].
Q18.What is your peak throughput for this task type? If we needed to increase volume by 2× mid-project, what is the minimum lead time required and what quality controls would apply to the scaled workforce?
G. Data Security and Compliance
Q19.What security certifications do you hold? (SOC 2 Type II, ISO 27001, HIPAA, GDPR, etc.) Please attach the most recent audit report or summary for SOC 2 if applicable.
Q20.Describe your data handling protocol for annotation tasks containing sensitive content. How is data transmitted to annotators? Where is it stored? What is your data deletion process upon project completion? What is your NDA execution timeline?
5. Vendor Scoring Rubric
Score each vendor 1–5 on each dimension, then multiply by the weight to get the weighted score. Total weighted score out of 5.0. Use this as a decision tool:
| Dimension | Weight | Score (1–5) | Weighted Score | What to assess |
|---|---|---|---|---|
| Annotator quality | 30% | ___ | ___ | Domain expertise match, κ track record with evidence, annotator credential specificity |
| Quality methodology | 25% | ___ | ___ | IAA measurement frequency, aggregation method fit for task type, adjudication protocol clarity |
| RLHF capability | 20% | ___ | ___ | Demonstrated preference pair experience, rationale format, SFT completion quality process |
| Pricing fit | 10% | ___ | ___ | Price vs. quality relative to alternatives, pricing model alignment with project incentives, inclusion of rework |
| Turnaround / scalability | 10% | ___ | ___ | Days to first batch, peak throughput confirmation, pilot availability on your timeline |
| Security / compliance | 5% | ___ | ___ | SOC 2 status, relevant certifications for data type, NDA process timeline |
| Total | 100% | — | ___ / 5.0 | 4.0+ → pilot; 3.0–3.9 → follow-ups; <3.0 → pass |
Scoring guide
The weights above reflect a reasonable prior for RLHF and expert annotation projects. Adjust them based on your specific situation: if data security is critical (healthcare, legal), increase the security weight to 15–20%. If you have an urgent timeline, increase turnaround weight. The key discipline is to set weights before you receive responses — not after, when anchoring bias kicks in.
Skip the RFP for smaller pilots
For annotation pilots under $1,000, a full RFP process is overkill. Submit your annotation brief directly and get a scoped proposal within 24 hours.
View Starter Pack — $49 →6. Red Flags in Vendor Responses
Six specific warning signs that appear in vendor RFP responses and what each one actually tells you:
"Our annotators are highly vetted" with no IAA data
"We use majority vote" for preference annotation
No calibration pilot offered
Flat per-label pricing for expert annotation
No SOC 2 or equivalent for production training data
"Contact us for pricing" with no ballpark range
7. How Human Consensus AI Responds to This RFP
We'll answer this honestly — which means being specific about what we're genuinely good at, and equally clear about where we're not the right vendor.
What we answer well
- ✓Domain expert annotators: We source domain experts for RLHF and expert annotation tasks — not general workforce. For software engineering, medicine, law, and technical domains, annotators have working credentials in the field.
- ✓Consensus-with-rationale methodology: Preference pairs include documented rationale, not majority vote. Disagreements are adjudicated by a senior annotator with written reasoning. IAA is measured and reported per batch.
- ✓Transparent per-project pricing: Fixed project price based on your brief. No surprise re-annotation billing. Calibration pilot included. You know the cost before the project starts.
- ✓Calibration pilot as standard practice: Every project begins with a calibration batch. We don't offer production without pilot — it's part of our methodology, not an optional add-on.
- ✓RLHF and preference annotation focus: This is our primary task type. We've built our methodology around preference pairs, reward model training data, and SFT completion quality — not image bounding boxes or NER at volume.
What we're not right for
- ✗Commodity labeling at 100K+ volume: High-volume image classification, bounding box annotation, and NER at scale are not our use case. If you need 500,000 image labels, you want a crowdsourcing platform, not us.
- ✗Enterprise IT procurement integrations: We don't integrate with Ariba, Coupa, or other enterprise procurement systems. If your vendor onboarding requires procurement system integration and a 90-day contract execution timeline, we're not the right operational fit.
Skip the RFP — send us your brief directly
For RLHF and expert annotation projects, a well-written annotation brief is more useful than an RFP. Send us your brief and we'll scope the right annotator pool, confirm IAA targets, and deliver a fixed-price proposal within 24 hours. No procurement system required.