Procurement Guide14 min read·Human Consensus AI Team

AI Annotation RFP Template: How to Evaluate and Select an AI Training Data Vendor (with Scoring Rubric)

You've written your annotation brief, scoped your budget, and built a vendor shortlist. Now you need a formal evaluation framework — one built for annotation, not IT procurement. This post gives you a complete 20-question RFP, a weighted 6-dimension scoring rubric, and a red-flag guide for reading vendor responses critically.

1. Why RFPs Fail for AI Annotation

Most annotation RFPs are adapted from IT procurement templates. They ask about ISO certifications, years in business, client references from Fortune 500 companies, and SLA uptime guarantees. None of these tell you anything useful about annotation quality for RLHF or LLM training data.

Three specific failure modes show up in almost every annotation RFP that produces a bad vendor selection:

Asking about years of experience instead of IAA data

"How long have you been doing AI annotation?" is irrelevant. A vendor can have eight years of image classification experience and zero credible RLHF methodology. The right question is: "Can you share inter-annotator agreement data (Cohen's κ) from a completed preference annotation project with comparable task complexity?" A vendor who answers with marketing copy instead of a κ number has told you everything you need to know.

Pricing before quality guarantees

Standard procurement asks for pricing early in the RFP process — often before any discussion of methodology. This sequence is backwards for annotation. If you anchor on price before establishing quality standards, you end up comparing vendors on cost with no common quality baseline. The result is a dataset with κ = 0.45 that cost 40% less than a dataset with κ = 0.75 — and a reward model trained on noise that costs 10× more to retrain than the annotation savings were worth.

Missing the calibration pilot requirement

Enterprise IT procurement almost never includes a mandatory calibration phase — you buy the software, you deploy it, you live with it. Annotation is different. A calibration pilot (20–100 examples, IAA review before production begins) is not optional overhead — it is the quality gate that prevents 30–40% rework on production batches. Any RFP that doesn't include calibration pilot terms in the contract structure is setting up a project failure.

The template and rubric below are built around these three corrections. They ask the right questions in the right order, gate pricing on quality commitments, and treat the calibration pilot as a contractual requirement, not a vendor courtesy.

2. Before You Send the RFP: What to Have Ready

An RFP is only as good as the inputs you give vendors to respond to. If you send a vague brief, you get vague vendor responses — and you end up scoring vendors on marketing polish rather than actual capability. Before the RFP goes out, you need five things locked down.

01

A completed annotation brief

The brief is the primary input to the RFP. It specifies task definition, input/output format, quality criteria rubric, edge case handling, IAA targets, annotator requirements, and timeline. If you don't have a complete brief, your RFP responses will be unscoreable — vendors will give you capability statements instead of project-specific answers. See our complete guide to writing an annotation brief before sending this RFP.annotation brief guide

02

Task taxonomy

What type(s) of annotation does this project require? Preference annotation (RLHF pairwise), safety flagging, SFT completion writing, named entity recognition, image classification, or a combination? Vendors have different capability profiles by task type. A vendor excellent at crowdsourced NER may have no credible RLHF methodology. Your task taxonomy determines which vendors belong on the shortlist.

03

IAA target by dimension

Your minimum acceptable Cohen's κ (or Fleiss' κ) for each annotation dimension. Example: κ ≥ 0.70 for preference selection, κ ≥ 0.85 for safety flags, κ ≥ 0.65 for helpfulness ratings. These numbers become the acceptance criteria in the contract. Any vendor who won't commit to measurable IAA targets is not a serious vendor for quality-sensitive annotation.

04

Budget range

Not a single number — a range. "$15,000–$25,000 for 5,000 preference pairs" gives vendors enough context to propose the right methodology without anchoring to the floor. If budget is listed as "competitive" or "to be determined," vendors will quote whatever they think your organization can absorb. For context on what annotation budgets should look like by task type and volume, see our AI annotation pricing guide.

05

Timeline and volume

Total annotation count, production start date, milestone checkpoints, and final delivery deadline. A vendor who can't respond to specific volume and timeline inputs isn't operationally ready for your project — they're still selling to you, not scoping for you. Require vendors to confirm capacity for your specific volume window, not generic throughput claims.

3. The 6 Evaluation Dimensions

Every annotation vendor should be evaluated across the same six dimensions. These are not generic procurement criteria — they are the specific quality and operational factors that determine whether an annotation vendor delivers data your models can actually train on.

👤

1. Annotator quality

The two questions that matter: What is the domain expertise level of annotators assigned to tasks like mine? And what is their measured κ track record on comparable tasks? "Vetted annotators" is not an answer. κ ≥ 0.72 on medical Q&A preference annotation with clinical reviewer adjudication is an answer. Require specific data, not general claims.

📊

2. Quality methodology

How does the vendor aggregate labels across annotators? For binary classification, majority vote is defensible. For preference annotation and RLHF pairs, majority vote is the wrong methodology — it averages away disagreement rather than resolving it. Consensus with documented rationale from a senior annotator produces more reliable signal. Require the vendor to describe their methodology for your specific task type, not their general process.

🤖

3. RLHF / preference annotation capability

RLHF preference annotation is not standard annotation work. It requires annotators who understand model outputs well enough to make reliable comparative judgments, a preference pair format with documented rationale, and an IAA measurement approach that accounts for subjective quality dimensions. Ask for specific examples: How many preference pairs have they delivered? What was measured κ? Can they share anonymized examples of their rationale format?

💰

4. Pricing model

Four models exist: per-label (cheapest headline, most opaque quality incentive), per-hour (transparent but variable), per-project (fixed scope, cleaner budgeting), and platform subscription (best for internal teams with BYO workforce). Each has different implications for your budget and the vendor's quality incentives. A per-label vendor has a financial incentive to maximize throughput; a per-project vendor has an incentive to deliver the agreed scope. Understand the model before comparing prices.

5. Turnaround and scalability

Time to first batch (can you start this week?), peak throughput (can they scale if you need to triple volume mid-project?), and pilot availability (will they do a calibration pilot on your schedule?). Vendors who can't commit to your specific start date or can't confirm capacity for your volume window are operationally unready for your project, regardless of their quality claims.

🔒

6. Data security and compliance

SOC 2 Type II certification is the minimum standard for production training data — it means independent auditors have verified the vendor's security controls. For healthcare data, HIPAA BAA. For EU data, GDPR DPA. NDA process and time to execute matters if your prompts and completions contain confidential information. For commodity annotation this may be a lower-stakes requirement; for production RLHF data that contains sensitive user queries, it is not.

For a full comparison of how major vendors score across these dimensions, see our roundups: Scale AI alternatives, Labelbox alternatives, and Appen alternatives. For context on what expert vs. crowdsourced annotators deliver on these dimensions, see our domain experts vs. crowdsourcing guide.

4. Full RFP Template (Ready to Send)

The following 20-question RFP is formatted for direct use. Customize the bracketed fields with your project specifics, remove sections that don't apply to your task type, and send as a document or email to your vendor shortlist. Responses to this RFP should be scoreable using the rubric in Section 5.

AI Annotation Vendor RFP — [Your Company Name]

Project: [Brief project description] · Response deadline: [Date]

A. Vendor Overview

Q1.Provide a one-paragraph description of your organization's annotation services, focusing specifically on the task types most relevant to this project: [preference annotation / safety flagging / SFT completion writing / other].

Q2.How many annotation projects have you completed in the last 24 months for AI/ML training data specifically? What was the average project volume (total labels)? What percentage involved RLHF or preference annotation?

Q3.Provide two client references for completed projects with comparable task type and volume. Include contact name, company, project scope, and approximate completion date.

B. Annotator Qualifications

Q4.What are the qualifications of annotators you would assign to this specific project? Describe domain expertise level, credentials, and language proficiency. Do not describe your general annotator pool — describe the subset you would use for [task type].

Q5.How many annotators would be assigned to this project? How do you ensure consistency across annotators working on the same task?

Q6.Provide inter-annotator agreement data (Cohen's κ or Fleiss' κ) from a completed project with comparable task complexity. Specify the task type, number of annotators, and measured κ by dimension.

C. Quality Assurance Process

Q7.Describe your label aggregation methodology for preference annotation tasks. Do you use majority vote, consensus with rationale, senior annotator adjudication, or another approach? Explain the reasoning behind your choice.

Q8.What is your IAA measurement process? How frequently do you calculate inter-annotator agreement during production? What is your protocol when a batch falls below the agreed IAA threshold?

Q9.How do you handle annotators who consistently fall below IAA threshold during a project? What is your replacement and rework protocol?

Q10.Confirm that your team will complete a calibration pilot of [20–100 examples] before production begins, with IAA review prior to proceeding. Describe your calibration process.

D. RLHF / Preference Annotation Experience

Q11.Describe your specific experience with RLHF preference pair annotation. How many preference pairs have you produced in total? What was the most complex preference annotation task you've completed, and what κ did you achieve?

Q12.What is your rationale format for preference pairs? Do annotators provide written justification for their preference selection? Provide an anonymized example of a completed preference annotation with rationale.

Q13.Describe your experience producing SFT (supervised fine-tuning) instruction-completion examples. What quality review process do you use for SFT completions, and what is your target quality score?

E. Pricing Model and Structure

Q14.Provide an itemized price estimate for this project: [volume] [task type] with the following requirements: [annotator qualifications], [IAA targets], [calibration pilot included]. Specify whether pricing is per-label, per-hour, per-project, or another model.

Q15.What is included in this price? Specifically: Is the calibration pilot included or priced separately? Is IAA reporting included? Are re-annotation costs for batches that fail IAA threshold included or billed separately?

Q16.What are your payment terms? Do you require a deposit? What is your invoicing schedule for a multi-batch project?

F. Project Timeline and Scalability

Q17.Confirm that you can begin the calibration pilot within [X business days] of contract execution. Provide a proposed project schedule with batch checkpoints for the following total volume: [volume and deadline].

Q18.What is your peak throughput for this task type? If we needed to increase volume by 2× mid-project, what is the minimum lead time required and what quality controls would apply to the scaled workforce?

G. Data Security and Compliance

Q19.What security certifications do you hold? (SOC 2 Type II, ISO 27001, HIPAA, GDPR, etc.) Please attach the most recent audit report or summary for SOC 2 if applicable.

Q20.Describe your data handling protocol for annotation tasks containing sensitive content. How is data transmitted to annotators? Where is it stored? What is your data deletion process upon project completion? What is your NDA execution timeline?

Response format: PDF or Google Doc. Include this question numbering for ease of scoring. Responses due [date]. Questions to [contact].

5. Vendor Scoring Rubric

Score each vendor 1–5 on each dimension, then multiply by the weight to get the weighted score. Total weighted score out of 5.0. Use this as a decision tool:

4.0+Proceed to calibration pilot
3.0–3.9Request follow-up clarification before pilot
Below 3.0Pass — do not proceed
DimensionWeightScore (1–5)Weighted ScoreWhat to assess
Annotator quality30%______Domain expertise match, κ track record with evidence, annotator credential specificity
Quality methodology25%______IAA measurement frequency, aggregation method fit for task type, adjudication protocol clarity
RLHF capability20%______Demonstrated preference pair experience, rationale format, SFT completion quality process
Pricing fit10%______Price vs. quality relative to alternatives, pricing model alignment with project incentives, inclusion of rework
Turnaround / scalability10%______Days to first batch, peak throughput confirmation, pilot availability on your timeline
Security / compliance5%______SOC 2 status, relevant certifications for data type, NDA process timeline
Total100%___ / 5.04.0+ → pilot; 3.0–3.9 → follow-ups; <3.0 → pass

Scoring guide

5Specific, documented, verifiable. Vendor answered the question asked with data.
4Mostly specific with minor gaps. Answers demonstrate operational capability.
3Partial answer. General claims with some specifics. Follow-up required.
2Vague or evasive. Marketing language, no supporting data.
1Non-answer. Failed to respond to the dimension or contradicted itself.

The weights above reflect a reasonable prior for RLHF and expert annotation projects. Adjust them based on your specific situation: if data security is critical (healthcare, legal), increase the security weight to 15–20%. If you have an urgent timeline, increase turnaround weight. The key discipline is to set weights before you receive responses — not after, when anchoring bias kicks in.

Skip the RFP for smaller pilots

For annotation pilots under $1,000, a full RFP process is overkill. Submit your annotation brief directly and get a scoped proposal within 24 hours.

View Starter Pack — $49 →

6. Red Flags in Vendor Responses

Six specific warning signs that appear in vendor RFP responses and what each one actually tells you:

"Our annotators are highly vetted" with no IAA data

This is the annotation equivalent of "we have a rigorous hiring process" with no performance data to back it up. Any vendor with a credible quality methodology can produce historical κ data from comparable tasks. The absence of IAA data is not a documentation gap — it is evidence that the vendor doesn't measure IAA systematically, which means they can't guarantee your threshold either.

"We use majority vote" for preference annotation

Majority vote is statistically defensible for clear-cut classification tasks. For preference annotation — where the signal is subjective and the annotation task is to capture genuine human judgment — majority vote averages out the disagreement rather than resolving it. RLHF preference pairs need documented rationale and senior annotator adjudication for disagreements. A vendor that defaults to majority vote for preference annotation doesn't understand the methodology.

No calibration pilot offered

A vendor who is confident in their methodology will offer a calibration pilot because it protects them too — it confirms the brief is well-specified before they commit 50+ annotators to production work. A vendor who resists calibration pilots is either operationally unready for structured quality measurement, or has learned that calibration batches reveal IAA problems they'd rather not surface until after invoicing. Pass.

Flat per-label pricing for expert annotation

Flat per-label pricing means the vendor is paid the same whether an annotation takes 45 seconds or 8 minutes. For commodity image classification, this is fine. For expert annotation requiring domain knowledge, documented rationale, and careful judgment — including RLHF preference pairs — flat per-label pricing creates an incentive to rush. Ask how annotators are compensated: if the vendor's incentive model pushes throughput over quality, assume throughput over quality.

No SOC 2 or equivalent for production training data

For a small pilot with sanitized, non-confidential prompts, SOC 2 may be acceptable to waive. For production annotation involving real user queries, proprietary model outputs, or any sensitive domain content, the absence of SOC 2 means your training data is passing through systems with no independently verified security controls. This is not a vendor maturity issue — it is a data risk issue for production-grade training datasets.

"Contact us for pricing" with no ballpark range

Legitimate vendors can give you a ballpark range without a full scoping call. "Expert preference annotation at this volume typically runs $X–$Y" is a normal thing to say. Refusing to provide any range before a call is a sales tactic, not a scoping requirement. It forces you into a conversation designed to build relationship before you've evaluated fit. If a vendor won't give you even an order-of-magnitude range in the RFP response, they're optimizing for the deal, not your evaluation process.

7. How Human Consensus AI Responds to This RFP

We'll answer this honestly — which means being specific about what we're genuinely good at, and equally clear about where we're not the right vendor.

What we answer well

  • Domain expert annotators: We source domain experts for RLHF and expert annotation tasks — not general workforce. For software engineering, medicine, law, and technical domains, annotators have working credentials in the field.
  • Consensus-with-rationale methodology: Preference pairs include documented rationale, not majority vote. Disagreements are adjudicated by a senior annotator with written reasoning. IAA is measured and reported per batch.
  • Transparent per-project pricing: Fixed project price based on your brief. No surprise re-annotation billing. Calibration pilot included. You know the cost before the project starts.
  • Calibration pilot as standard practice: Every project begins with a calibration batch. We don't offer production without pilot — it's part of our methodology, not an optional add-on.
  • RLHF and preference annotation focus: This is our primary task type. We've built our methodology around preference pairs, reward model training data, and SFT completion quality — not image bounding boxes or NER at volume.

What we're not right for

  • Commodity labeling at 100K+ volume: High-volume image classification, bounding box annotation, and NER at scale are not our use case. If you need 500,000 image labels, you want a crowdsourcing platform, not us.
  • Enterprise IT procurement integrations: We don't integrate with Ariba, Coupa, or other enterprise procurement systems. If your vendor onboarding requires procurement system integration and a 90-day contract execution timeline, we're not the right operational fit.

Skip the RFP — send us your brief directly

For RLHF and expert annotation projects, a well-written annotation brief is more useful than an RFP. Send us your brief and we'll scope the right annotator pool, confirm IAA targets, and deliver a fixed-price proposal within 24 hours. No procurement system required.

If you're evaluating vendors for RLHF or expert annotation, start with a pilot.

The fastest way to evaluate an annotation vendor isn't an RFP — it's a pilot. Send us your annotation brief and we'll deliver 25–50 expert preference pairs with full IAA documentation so you can evaluate output quality directly.

Start the Pilot — Starter Pack $49 →

Running a larger annotation program? The Enterprise Bundle includes custom rubric design, dedicated domain expert sourcing, calibration management, and IAA monitoring — without the six-figure contract minimums.

View Enterprise Bundle — $299 →