If Scale AI was your only labeling vendor, the fix is not a new sole supplier. It is two vendors behind a gold set you own. For expert preference and evaluation data, make Surge AI your primary. Put Labelbox underneath as the workspace where your rubric, your gold items and your agreement scores live. Use Prolific for the general-population judgments that do not need a credentialed expert. Scale AI is the loser here for any model owner who competes with Meta, which owns 49% of it. Mercor is a strong expert-staffing option, but after its March 2026 breach it goes in the second slot, not the first.
| Vendor | Pick it for | Workforce model | Published price (checked Sep 29, 2026) | Data-handling flag | Verdict |
|---|---|---|---|---|---|
| Surge AI | Expert RLHF, preference and eval data | Managed expert network | Contact sales | Contractor-misclassification class action (May 2025) | Primary expert vendor |
| Labelbox (+ Alignerr) | Your own workspace, plus an on-demand expert bench | Platform + managed or direct-hire experts | 500 free LBUs/month; paid rate via calculator or sales | SOC 2 Type II (vendor claim) | The control plane |
| Prolific | Crowd preference, UX and eval judgments | Self-serve verified participants | $8/hr minimum pay, $12/hr recommended, plus 42.8% corporate platform fee | Participants work on their own devices | Cheap, fast, transparent |
| Mercor | Named domain experts you want to staff directly | Expert marketplace, cost-plus hourly | Contact sales | March 2026 breach; Meta paused its work | Second source only |
| Scale AI | US government and defense programs | Managed, plus the Outlier/Remotasks crowd | Contact sales | 49% Meta-owned; confidential client docs left public in 2025 | Loser for model builders |
What Changed for Scale AI Customers, and Where They Went
The frontier labs left Scale within weeks of Meta's investment, and the vendors that picked up the work were Surge AI and Mercor. Meta bought a 49% stake for $14.3 billion. Scale's June 12, 2025 announcement valued the company at over $29 billion. It moved founder Alexandr Wang to Meta and named Jason Droege interim CEO.
Google had been Scale's largest customer and planned to pay it about $200 million in 2025. It planned to split right after the deal. OpenAI confirmed it was winding down and said it wanted "other providers for more specialized data." Scale's general counsel said the company would not share other customers' confidential information with Meta. The labs left anyway. When a model competitor owns half your data vendor, a contractual promise does not replace structural separation.
The telling detail is what Meta itself did. By August 2025, its superintelligence lab was working with Surge and Mercor alongside Scale, and TechCrunch reported that its researchers saw Scale's data as lower quality. The same report notes that Scale cut 200 roles from its data-labeling business in July 2025.
Scale has not collapsed. It has repositioned. Droege's January 2026 letter reports "well over $1B in new business" in 2025, a now-profitable data business, and new enterprise customers including Mayo Clinic, BP and Allianz. It also says 2026 means moving capabilities "from our data business into applications faster." That is a clear strategy. It is also a signal that labeling for outside model builders is no longer where Scale is investing. Our Scale AI tool page tracks the product line.
Who Should Not Pick Each Vendor
Every option on this list is the wrong answer for someone, and naming who that is matters more than any feature list.
Surge AI. Surge says it sells frontier data, RL environments and an expert workforce. According to Computerworld, it reportedly generated over $1 billion in revenue the year before the Meta deal, and its clients include Meta and OpenAI. That customer base is the strongest evidence of quality you will find in this market. It is also why Surge is not the right pick if you need a few thousand labels this quarter and a price you can see before a sales call. Prolific or a Labelbox workspace will get you there faster. Surge's workforce also carries legal exposure: a May 2025 class action alleges it misclassified its annotators as contractors. If your procurement policy screens suppliers on labor practices, raise that suit in diligence before you sign.
Labelbox. Labelbox says its Alignerr network has "2.6M+ contributors across 200+ domains" and SOC 2 Type II certification. Those are vendor claims; ask for the report. Pricing is metered in Labelbox Units. The billing docs charge one LBU per text or image row annotated and give free accounts 500 LBUs a month. Labelbox's pricing calculator did not load when we checked on September 29, 2026, so get the paid rate from sales. Do not pick Labelbox if nobody on your team will own the rubric. A platform stores your quality process. It does not create one.
Prolific. This is the only vendor here that publishes a full price. The pricing page lists a $8/hour minimum participant rate, a $12/hour recommended rate and a 42.8% platform fee for corporate customers. Prolific claims 300,000+ active participants and 40+ identity checks at onboarding. Do not pick it for credentialed-expert work, such as a claims adjuster grading claims answers or a clinician grading clinical ones. Prescreeners filter on self-reported background. They are not a licence check.
Mercor. Mercor claims "more than five million domain experts". It agreed to buy RL-environment builder Deeptune on July 9, 2026, and TechCrunch reported the same day that it had passed a $2 billion revenue run rate. But on March 31, 2026 it confirmed a breach that started with poisoned LiteLLM packages. Meta then paused its work indefinitely. Reporting on the incident said the exposure may include customers' labeling protocols and data-selection criteria. Do not make Mercor your only vendor for proprietary training data until you have read its post-incident remediation report yourself. It is a good second source for staffing named experts.
Scale AI. Do not pick Scale if you build models that compete with Meta's. The Meta stake is structural, not a rumour. Its data-handling record has a specific 2025 incident: Business Insider found 85 Google Docs of confidential client project material that anyone with the link could open. Scale said it disabled public sharing from its systems. The case for Scale is US government and defense work, where its January 2026 letter reports two nine-figure contracts.
How to Measure Labeling Quality: Agreement Rates and Gold Sets
Measure quality yourself with two numbers: agreement between annotators on the same item, and accuracy against a gold set you wrote. Never accept a vendor's own quality score as the metric. Inter-annotator agreement is how often two independent annotators give the same judgment on the same item. A gold set is a bank of items with adjudicated answers, mixed unannounced into the live queue.
Set a realistic benchmark first. Preference data is inherently noisy. In OpenAI's InstructGPT paper, training labelers agreed with each other 72.6 ± 1.5% of the time, and held-out labelers 77.3 ± 1.3%. Researchers on an earlier summarization task agreed 73 ± 4%. Those labelers were a team of about 40 contractors hired through Upwork and Scale AI, screened with a test built for the task. So if a vendor quotes 95% agreement on subjective preference pairs, the task is too easy, the rubric is leaking the answer, or the number is not measuring what they say it measures.
What predicts regret is not the absolute agreement number. It is these three things:
- Gold accuracy by annotator, not by batch. A batch average hides the two annotators who are coin-flipping. Ask for per-worker gold accuracy weekly, and write the removal threshold into the SOW.
- Agreement drift over time. Agreement that falls in week three usually means rubric drift or new, unscreened staff on the project. That is the moment a vendor has moved your work onto a cheaper pool.
- Disagreement you can adjudicate. Pay for a third annotator on split items and send the disagreements back to your own team. That queue is where your rubric gets better.
The sample-size math we covered for model evals applies here too. A 200-item gold set cannot tell a 78% annotator from an 82% one.
What Domain Experts Cost Compared With the Crowd
The crowd price is public and the expert price is not, so normalise every quote to cost per accepted judgment on the same workload. Our reference workload: 20,000 pairwise preference items on insurance-claims answers, each labeled twice, with a 10% gold set mixed in. At about three minutes per judgment, that is roughly 2,000 annotator-hours.
- Prolific, general population: 2,000 hours at the $12 recommended rate is $24,000 in participant pay. The 42.8% corporate fee adds $10,272, for about $34,300 total. That price is public, so you can put it in a budget today.
- Expert network (Surge, Mercor, Scale, Alignerr): all four quote through sales. Mercor bills customers a cost-plus hourly rate, and its mark-up is not published. Work it out from pay, not from the quote. If a qualified adjuster costs $60 an hour in pay alone (our assumption; use your own), the same 2,000 hours is $120,000 before any vendor margin.
- Labelbox platform: at one LBU per annotated text row, 22,000 rows including gold is about 22,000 LBUs. The platform cost is small compared with the labor cost at any realistic LBU rate.
The decision is not crowd versus expert. Split the task. Send the items that need licensed judgment to experts, and the fluency and tone comparisons to the crowd. Most enterprise preference sets are more than half the second kind. The same logic applies to generating data instead of buying it, which we covered for synthetic data platforms.
Data Handling and Subcontracting: Where Your Prompts Actually Go
Your data goes wherever the vendor's workforce is, and for most managed vendors that is a contractor pool one or two legal entities away. Scale ran its crowd through the Outlier and Remotasks platforms and used HR partners. When the US Department of Labor opened a Fair Labor Standards Act investigation into Scale, it also examined Upwork and HireArt. The investigation was dropped in May 2025. Private lawsuits from former workers continue.
This matters to a buyer for one reason. Your prompts, your rubric and your model's failure cases are shown to people who do not work for your vendor. Mercor showed what happens when the vendor itself is compromised. The attackers claimed about four terabytes of data, and Mercor's customers had to work out which of their own methods had been exposed. We covered how the same LiteLLM compromise pattern reaches credentials in an eval sandbox.
Put these in the contract, not the questionnaire:
- A named-subcontractor list with notice before changes, including HR intermediaries and task platforms.
- Where annotators work: in the vendor's tool with copy/paste and screenshots disabled, or on their own devices with the data in the clear.
- Incident notification to you within a fixed number of hours, separately from any notice to the vendor's other customers.
- Change-of-control termination rights. Scale's customers did not need these in 2024. They needed them in June 2025. Info-Tech's Thomas Randall advises buyers to "insist on contractual firewalls around staff mobility and data reuse".
- No reuse of your rubric or data for other customers or for the vendor's own models and benchmarks.
Staff mobility is not hypothetical either. In September 2025, Scale sued a former employee and Mercor, alleging he took more than 100 customer-strategy documents. Mercor denied using them. Your account plan was one of the documents in someone's Google Drive. Our six-question AI vendor security review covers the questionnaire half.
How to Run Two Labeling Vendors Without Doubling the Ops
Split vendors by task type, not by volume, and keep the rubric, the gold set and the scoring in one place you control. Splitting the same task 50/50 doubles calibration work and gives you two sets of labels that disagree for reasons you cannot separate. Splitting by task type gives each vendor one rubric to learn and gives you a clean comparison when you move work between them.
The operating model that holds up:
- One schema. Every vendor delivers the same JSON record format: item ID, annotator ID (pseudonymous), label, rationale, time spent. Write the converter once.
- One gold bank, rotated. Keep 10% gold in every queue at every vendor, drawn from a bank only your team can edit. A shared gold set is also the only fair test when you send the same item type to both vendors for a trial.
- One scoreboard. Per-vendor, per-annotator gold accuracy and agreement, weekly, in your own workspace. This is the job Labelbox does well. The eval-tooling comparison covers the same instinct for model outputs.
- A standing 10% overlap. Send a small, fixed share of each vendor's task type to the other. It costs little, and it means a failover is a volume change, not a new vendor onboarding. When your primary changes owner, as Scale did, you raise the overlap instead of starting procurement.
This Week: Pull the last 90 days of labels from your current vendor and write down the agreement rate. If you cannot compute it, that is the first finding. This Month: Build a 500-item gold bank with your own domain experts and run a paid 2,000-item trial at two vendors on the same items. Before the Next Renewal: Add change-of-control termination, subcontractor notice and no-reuse clauses. Then move one task type to the second vendor for good.
The Bottom Line
The labeling market used to work like cloud before multi-cloud: one default vendor, because switching looked expensive and nothing had forced the question. The Meta deal forced it for the frontier labs in a single week, and the Mercor breach showed the replacement vendors carry the same kind of risk. The lesson is the one enterprises learned with cloud. The asset worth protecting is not the vendor relationship. It is the rubric, the gold set and the scoreboard that let you change vendors on a Tuesday.
Vendors change owners. Your gold set doesn't have to.
Continue Reading
- AI Vendor Security Review: 6 Questions That Change the Answer
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
- Best Synthetic Data Platforms: Start With the Free SDK
- Braintrust vs Langfuse vs Promptfoo: Don't Pay for the Gate
- Fine-Tuning vs RAG Cost: The Training Bill Isn't the Bill
- Stilla Promised Continuity. Meta Bought It for WhatsApp.
