Consent-gated
Only tickets where the customer has consented under your terms and has not opted out.
03 · AI Data Operations
The engineer who fixed the build is the annotator. The failing output and the working one are the pair. Delivered in your schema, redacted to your guidelines, reviewed twice.
Published quality
2 days
Delivery after ticket closure
≥20%
Second-annotator review
90%
First-pass acceptance target
0.8 κ
Inter-annotator agreement
Why this exists
Data vendors label what they are given. They have never seen your model fail on a real customer's build, and the examples they produce are guesses at what failure looks like.
Meanwhile your support team resolves hundreds of real failures a week and records almost none of it in a form your ML team can use. The most valuable dataset in the company is discarded at ticket close.
Data lines
From evaluation capture on every ticket to RLHF pairs, red-teaming and golden sets.
For every ticket, in your schema: classification of the automated diagnosis as confirmed, partially correct or incorrect, with reason codes and the correct diagnosis where wrong; whether the platform output met the customer's objective, on your scale, with evidence; capabilities used and attempts required; customer feedback verbatim, coded to your taxonomy; defects observed with reproduction steps.
For each designated ticket, a record comprising the customer's original prompt and stated objective; the platform output or automated diagnosis; the engineer's diagnosis, remediation steps and final working result; a preference or quality rating against your rubric; where a corrected output exists, a paired rejected/preferred example; and further labels you specify.
Rubric-based scoring of model outputs, pairwise preference sets, release-over-release regression evals on a golden set.
Adversarial prompting, jailbreak discovery, harmful-output reports with reproduction and severity.
Prompt libraries, system-prompt variants with measured outcomes, few-shot sets drawn from real intents.
Intent and error taxonomies; labelling of prompts, code, logs and conversations.
Maintained reference sets; constructed rare-failure scenarios validated by engineers.
Operating model
Selection criteria come from you, or you designate tickets individually.
Engineers record evaluation fields at ticket close.
A data squad prepares RLHF records within 2 business days of closure, applies redaction, and submits to your repository in your schema.
A second annotator reviews at least 20% of records. You reject non-conforming records; we rework at no charge. Monthly calibration; inter-annotator agreement measured against your labelled sample.
Quality
| Metric | Target |
|---|---|
| Delivery after ticket closure | 2 business days |
| Second-annotator review | ≥20% of records |
| First-pass acceptance | 90% (set in SOW) |
| Inter-annotator agreement vs. your sample | 0.8 κ (set in SOW) |
| Rework of rejected records | No charge, within 3 business days |
| Minimum monthly volume | Agreed against forecast |
Controls
Consent-gated, redacted, yours on creation — never used to train our systems.
Only tickets where the customer has consented under your terms and has not opted out.
Personal data, credentials, API keys, secrets and customers' end-user data removed or pseudonymised to your guidelines.
Your work product on creation.
Only for the SOW; never to train or improve any model or system of ours; never for our analytics; not retained beyond the ticket.
Only AI tools you approve in writing touch customer data, code or prompts; usage auditable.
Disclosed; none added without 30 days' notice and approval.
Commercials
Evaluation capture is included in support fees. RLHF records are priced per accepted record. Evaluation and annotation per item; red-teaming per campaign; golden-set curation monthly.
FAQ
Sample deliverable
One RLHF record: prompt, objective, failing output, engineer diagnosis, corrected output, rubric score, rejected/preferred pair, redaction log and reviewer sign-off. Fifty such records with schema, rubric and acceptance notes are available to AI and product leads on request.
Redacted. Client and customer identifiers removed.
First-pass acceptance was 94% against a rubric our own annotators hit 88% on. These are the failures we actually needed.
No commitment either way
Request a sample dataset of fifty RLHF records with schema, rubric and acceptance notes.