What is red teaming, and the kind of data it actually needs

Red teaming is adversarial testing rather than quality assurance with a longer checklist. It needs trained attackers, a rubric written in advance, and a budget for re-running it.

Red teaming is structured adversarial testing: people deliberately try to make a system fail, in the ways that matter to whoever is deploying it. The output is not a score. It is a set of demonstrated failures with enough detail that someone else can reproduce them and fix them.

It gets confused with two neighboring activities. It is not evaluation on a benchmark, because a benchmark asks whether the system performs a task, while red teaming asks where the system breaks. And it is not quality assurance, because QA checks whether the system does what the specification says, while red teaming goes looking for the cases the specification never imagined.

The artifact is a set of attempts, not a number

A red-team record has a consistent shape: the input that was tried, the system version it was tried against, what the system did, whether the failure reproduces on a second run, and a severity judgment measured against a scale written before the exercise began.

Two details separate a useful record from a pile of screenshots. The first is the negative results. Attempts that failed to break the system belong in the deliverable, because they mark the boundary and they stop the next team from repeating work that was already done. A report listing only successes reads as a longer list of problems than the system actually has.

The second is the taxonomy that emerges. After a few hundred attempts the failures cluster — into framing tricks, multi-turn setups, conflicting instructions, encoded inputs, language pivots, context-length edges. That taxonomy is what gets carried into the next release. The individual attempts expire quickly; the categories do not.

The skill is adversarial, not linguistic

Finding something is not the same skill as writing well, and it is not the same skill as annotating. It requires a working model of how these systems fail: that an instruction can be buried inside a document, that a persona can shift what gets answered, that a harmless request split across turns can assemble into something else, and that a policy enforced in one language is often thinner in another.

A general annotator pool produces shallow attempts, and shallow attempts produce a report that looks thorough and finds nothing new. What works better is a small group trained on past failure write-ups, working from the previous round's taxonomy, and given time to iterate instead of a quota of attempts per hour.

One separation is worth insisting on: the people writing attempts should not be the people grading them. Whoever found a failure has an interest in it being severe, and that interest is easy to satisfy without noticing.

Grading is the part that gets under-specified

"Did this fail?" is a judgment, and severity is a second judgment stacked on top of it. Both need a rubric written before the run, because a rubric written afterwards tends to fit the results.

Three things make grading defensible. A severity scale with concrete anchors — what counts as a minor slip, a harmful output, a systemic failure — rather than a numbered scale with no descriptions. A double-graded sample, so a grader agreement rate can be reported. And an explicit category for over-refusal: a safe request the system declined. A report that counts only successful attacks pushes the next model toward refusing more, which is a real product cost and belongs on the same page as the failures.

The set decays, so budget for the cycle

A red-team set is a snapshot of a moving target. The system changes and some attempts stop working. Published attempts end up in training data and stop working for that reason too. Both mean the value is not in the file.

What holds value is the process: the taxonomy, the trained group, the rubric, and the habit of re-running the set against every release while refreshing the attempt pool each cycle. Plan for that as a recurring line rather than a one-time purchase, and expect the second round to find different things rather than the same things again.

The refresh rate is a decision worth making explicitly. Too frequent, and you are paying for a set that has not had time to become stale. Too rare, and you are testing last year's system against last year's tricks.

What to specify when you commission it

That last point is worth restating, because it is the failure that costs the most and is hardest to notice. A red-team set reused as training data makes the next round look clean, and a clean round is not the same thing as a system that holds.

  • Which risk categories are in scope, in writing, and which are out of scope.
  • The size and composition of the attempt pool, with coverage targets rather than a single total.
  • The rubric and the severity scale, delivered before the exercise starts.
  • Whether the report includes negative results and reproductions, not only findings.
  • Whether the raw attempts may be reused as training data, with a clear default that they should not be.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com