Skip to content

03 · AI Data Operations

Evaluation data from real failures, not synthetic ones.

The engineer who fixed the build is the annotator. The failing output and the working one are the pair. Delivered in your schema, redacted to your guidelines, reviewed twice.

  • Delivered in 2 business days
  • Second-annotator review ≥20%
  • Your work product on creation

Published quality

Real failures, labelled twice, delivered on a clock

Request a sample dataset

2 days

Delivery after ticket closure

≥20%

Second-annotator review

90%

First-pass acceptance target

0.8 κ

Inter-annotator agreement

Why this exists

The most valuable dataset is discarded at ticket close

Data vendors label what they are given. They have never seen your model fail on a real customer's build, and the examples they produce are guesses at what failure looks like.

Meanwhile your support team resolves hundreds of real failures a week and records almost none of it in a form your ML team can use. The most valuable dataset in the company is discarded at ticket close.

Data lines

What's included

From evaluation capture on every ticket to RLHF pairs, red-teaming and golden sets.

  1. 01

    Evaluation capture (included with Support)

    For every ticket, in your schema: classification of the automated diagnosis as confirmed, partially correct or incorrect, with reason codes and the correct diagnosis where wrong; whether the platform output met the customer's objective, on your scale, with evidence; capabilities used and attempts required; customer feedback verbatim, coded to your taxonomy; defects observed with reproduction steps.

  2. 02

    RLHF dataset preparation

    For each designated ticket, a record comprising the customer's original prompt and stated objective; the platform output or automated diagnosis; the engineer's diagnosis, remediation steps and final working result; a preference or quality rating against your rubric; where a corrected output exists, a paired rejected/preferred example; and further labels you specify.

  3. 03

    Human evaluation and benchmarking

    Rubric-based scoring of model outputs, pairwise preference sets, release-over-release regression evals on a golden set.

  4. 04

    Red-teaming and safety testing

    Adversarial prompting, jailbreak discovery, harmful-output reports with reproduction and severity.

  5. 05

    Prompt engineering and instruction data

    Prompt libraries, system-prompt variants with measured outcomes, few-shot sets drawn from real intents.

  6. 06

    Annotation and taxonomy

    Intent and error taxonomies; labelling of prompts, code, logs and conversations.

  7. 07

    Golden dataset curation and synthetic edge cases

    Maintained reference sets; constructed rare-failure scenarios validated by engineers.

Operating model

How it runs

  1. 01

    Select

    Selection criteria come from you, or you designate tickets individually.

  2. 02

    Capture

    Engineers record evaluation fields at ticket close.

  3. 03

    Prepare

    A data squad prepares RLHF records within 2 business days of closure, applies redaction, and submits to your repository in your schema.

  4. 04

    Review

    A second annotator reviews at least 20% of records. You reject non-conforming records; we rework at no charge. Monthly calibration; inter-annotator agreement measured against your labelled sample.

Quality

Quality and SLAs

MetricTarget
Delivery after ticket closure2 business days
Second-annotator review≥20% of records
First-pass acceptance90% (set in SOW)
Inter-annotator agreement vs. your sample0.8 κ (set in SOW)
Rework of rejected recordsNo charge, within 3 business days
Minimum monthly volumeAgreed against forecast

Controls

Data governance

Consent-gated, redacted, yours on creation — never used to train our systems.

Consent-gated

Only tickets where the customer has consented under your terms and has not opted out.

Redaction

Personal data, credentials, API keys, secrets and customers' end-user data removed or pseudonymised to your guidelines.

Ownership

Your work product on creation.

Use

Only for the SOW; never to train or improve any model or system of ours; never for our analytics; not retained beyond the ticket.

Tooling

Only AI tools you approve in writing touch customer data, code or prompts; usage auditable.

Jurisdictions

Disclosed; none added without 30 days' notice and approval.

Commercials

Pricing model

Evaluation capture is included in support fees. RLHF records are priced per accepted record. Evaluation and annotation per item; red-teaming per campaign; golden-set curation monthly.

See the full pricing model

FAQ

Questions on data ownership

You, on creation. Every engineer signs an IP assignment before access.

Sample deliverable

Sample deliverable

One RLHF record: prompt, objective, failing output, engineer diagnosis, corrected output, rubric score, rejected/preferred pair, redaction log and reviewer sign-off. Fifty such records with schema, rubric and acceptance notes are available to AI and product leads on request.

evaluation_record.jsonRedacted
ticket_id
PX-2291-████
objective
Ship checkout with saved cards
diagnosis_class
partially_correct
reason_code
dns_propagation_assumed
correct_diagnosis
Edge proxy bound to stale build target
output_met_objective
no
objective_scale
2 / 5
capabilities_used
domain_config, deploy_logs, rebuild
attempts
2
customer_verbatim
"It said it deployed but the domain 502s"
verbatim_taxonomy
deploy.domain.misbinding
defect_repro_steps
4 steps, attached

Redacted. Client and customer identifiers removed.

First-pass acceptance was 94% against a rubric our own annotators hit 88% on. These are the failures we actually needed.
Head of AI, Coding assistant

No commitment either way

Ready to turn tickets into training data?

Request a sample dataset of fifty RLHF records with schema, rubric and acceptance notes.