NewVideo Object Segmentation and Tracking for SAM 2

What is expert data, and when do you actually need it?

Someone on your team has priced expert annotation and the number is four to ten times what generalist labeling costs. The question that follows is harder than it looks, because "do we need experts" is usually asked about the people when it should be asked about the judgement. Some tasks genuinely require years of accumulated calibration that no written instruction can transfer. Others look that way only because nobody has written the instruction carefully enough yet. Those two situations have the same symptom, low agreement, and opposite cures. This piece defines expert data as a category and gives you a test that separates them.

Key takeaways

  • Expert data is defined by whether the judgement can be reproduced from a written specification, not by the credentials printed on an annotator profile.
  • Run the reproducibility test before you buy: give a careful generalist your specification and measure whether their labels track an expert's. Disagreement that survives a better specification is the signal you are paying for.
  • In a HumanSignal census of preference-data literature, only 6 of 24 randomly sampled papers linked judgments to a persistent annotator identity, so most published datasets cannot tell you whose expertise they captured.
  • Expert judgement has a noise floor. Two professional linguists scoring the same material agreed on roughly 72%, so a process demanding near-total agreement from experts is suppressing the judgement it bought.
  • Expertise does not remove the specification problem. You still own the spec, the adjudication capacity, and the audit design.

What separates expert data from labeled data

Expert data is training or evaluation data whose value comes from judgement that is expensive to acquire rather than from volume. The useful definition is operational rather than credential-based: a label is expert data when a competent, motivated non-specialist working from your written instructions cannot reliably reproduce it.

That framing matters because credentials and reproducibility come apart in both directions. Plenty of tasks handed to specialists turn out to be fully specifiable once someone sits down and writes the decision rules, at which point the specialist is an expensive way to apply a procedure. Other tasks resist specification no matter how much effort goes into the document, because the judgement draws on pattern recognition the expert cannot fully articulate. Only the second kind is genuinely expert data, and only the second kind justifies the rate.

It is worth separating this from the activity of labeling. The annotation work itself is a process question about who sits at the interface. Expert data is a property of the resulting dataset: what it encodes, and whether that encoding could have come from anywhere cheaper.

The test that decides it

Before committing a budget, run a small, cheap experiment that answers the question directly.

Can a careful generalist reproduce the label from a written spec

Take 50 to 100 representative items, including the hard ones. Have a specialist label them. Separately, write the best specification you can and have two trained generalists label the same items from it, without seeing the specialist's answers. Then measure how well the generalist labels track the specialist's.

High agreement means your task is specifiable, and the honest conclusion is that you were paying for a document you had not written yet. Low agreement means something in the judgement is not travelling through the instructions, which is the first real evidence that you need expertise rather than clearer prose.

Does disagreement survive a better specification

One round is not enough, because the first specification is almost always underdetermined. Take the items where the generalists diverged from the specialist, work out what distinction was being applied, write it into the document, and run the test again on fresh items.

What you are watching is the trend. Disagreement that falls sharply with each revision is a specification problem that will keep yielding to effort. Disagreement that plateaus while the document keeps growing is the signature of judgement that does not reduce to rules. That plateau is the point at which expert data is the right purchase, and it is a far more defensible basis for the decision than intuition about how hard the domain sounds.

Where expertise is genuinely load-bearing

Three patterns reliably produce that plateau.

Judgement that encodes years of tacit calibration

Radiologists reading subtle presentations, structural engineers assessing fatigue, and senior editors judging whether an argument holds together share a property: the practitioner can state a conclusion confidently and cannot fully reconstruct the reasoning that produced it. The calibration is real and it is not written down anywhere, including in their own head. This is the clearest case for training data that genuinely requires a specialist.

Consequences that are asymmetric or regulated

Sometimes the judgement is specifiable but the cost of a rare error is severe enough that you want a qualified human accountable for it. Clinical, legal, and safety contexts often work this way. The expertise is buying liability coverage and escalation judgement as much as label accuracy, and that is a legitimate reason to pay for it as long as you are clear that is what you are buying.

Domains where no ground truth exists to check against

When the question has no external answer, as with whether a design reads well or an argument is persuasive, there is nothing to grade labels against except other labels. Expertise here is not about accuracy, since accuracy is undefined. It is about whether the judgement is stable and informed enough to be worth modeling. This case is badly served by the usual quality tooling, and it is where dataset documentation is weakest.

What expert data costs you beyond the rate card

The hourly rate is the visible cost and rarely the binding one.

Throughput is the first hidden cost. Specialist pools are small, and a task requiring a particular credential in a particular language can cap your labeling rate regardless of budget. Adjudication is the second: when two qualified experts disagree, resolving it takes a third expert, and that capacity has to be planned rather than improvised.

The third is the one teams consistently underestimate, which is that hiring expertise does not transfer the specification problem. You still have to write the instructions, design an audit you can run without the domain knowledge, and decide what counts as done. Those tasks remain yours no matter who is labeling, and the work of writing a specification for a domain you cannot personally judge is a prerequisite rather than something the vendor absorbs.

Calibrate your expectations too. In a benchmark HumanSignal ran with Custom.MT across 7,817 translation segments, two professional linguists scoring the same segments agreed on roughly 72% of them. Qualified people applying real judgement to hard material diverge, and a review process that treats every divergence as a defect will grind down the judgement you paid a premium to obtain.

When a better specification is the cheaper answer

If the reproducibility test shows agreement climbing as the document improves, the cheaper path is to keep improving the document and run a tiered arrangement of generalists and experts: generalists handle volume against the specification, specialists handle adjudication and periodic audit. This uses the expensive judgement where it is decisive rather than spreading it across every item.

That arrangement also produces something a fully expert pipeline does not, which is a written artifact encoding what the experts know. Each adjudicated case becomes a rule in the document. Over time the specification absorbs judgement that previously lived only in people, and the ratio of expert time to volume falls.

How to buy it without overbuying

Specify the judgement rather than the credential. A vendor can supply almost any qualification on request, so asking for a credential without saying what judgement it must produce gets you an expensive person applying your underdetermined instructions. Ask instead how the provider verifies that a given annotator produces the judgement your task needs, since how a partner verifies credentials is the part that varies most between suppliers.

Then insist on annotator identity in the delivered data. This is the most commonly skipped requirement and the most expensive to retrofit. In a HumanSignal census of the literature on human preference data for creative generative models, only 6 of 24 randomly sampled papers linked judgments to a persistent annotator identity, and just 5 of 24 reported how many items each annotator judged. Without both, you cannot tell whether a pattern in the data reflects one expert's view or a population's, and you cannot reuse the dataset to model individual judgement later. The census scope is creative work rather than annotation generally, and the recording practice it measures is the same practice your delivery will either include or omit.

Find out whether your task needs experts before you buy them

The reproducibility test costs a fraction of a production run and answers the question your budget depends on. HumanSignal Services recruits from a network of more than 3 million experts across 50 or more knowledge domains, and will run a calibration round against your specification first so the decision rests on measurement rather than on how hard the domain sounds. Start a scoping conversation to set one up.

How do I tell whether my task needs experts?

Run the reproducibility test: have a specialist label 50 to 100 representative items, then have trained generalists label the same items from your written specification alone, and compare. If generalist labels track the specialist's closely, the task is specifiable and you were missing a document rather than expertise. If divergence persists after two or three rounds of improving the specification, the judgement is not reducible to instructions and expert data is the right purchase.

What does expert data cost compared with generalist annotation?

Rates commonly run several times higher, but the rate is rarely what governs the budget. Throughput limits matter more, because specialist pools are small and a rare credential in a specific language can cap your labeling rate outright. Add adjudication capacity for expert-versus-expert disagreement, which needs planning rather than improvisation.

Can I mix expert and generalist annotators on one task?

That arrangement usually outperforms either pool alone, with generalists handling volume against the specification and specialists handling adjudication, audit, and the hardest items. It also produces a written record of expert reasoning as adjudicated cases turn into rules. The requirement is that routing be deliberate, with criteria for escalation set before collection rather than decided case by case.

How do I verify a vendor's experts are real experts?

Ask what judgement the credential is supposed to produce, then test it directly with seeded items whose answers you already know. Credential documents are easy to supply and tell you little about whether a given person produces the judgement your task needs. Verification through performance on known items works even when you cannot evaluate the domain yourself.

What happens if I buy expert data and the specification is wrong?

You get expensive labels that are internally consistent and aimed at the wrong target, which is harder to detect than cheap labels that are obviously noisy. This is the main argument for calibration rounds before production volume. Budget two or three small rounds and expect the specification to change after each one.

Is expert data the same as expert annotation?

They describe different things. Expert annotation is the activity of having qualified people label data; expert data is the property of a dataset whose judgement could not have been reproduced from written instructions by anyone cheaper. Expert annotation can produce ordinary labeled data when the task turns out to be fully specifiable, which is exactly the case the reproducibility test is designed to catch.

Related Content