Skip to content

The Flywheel

The support queue is the most expensive dataset most AI companies throw away.

Every resolved ticket is a labelled example of where the model fell short and what right looked like. This chapter explains what we do with it.

  • Evaluation records in your schema, at ticket close
  • Preference pairs delivered within 2 business days
  • Consent-gated, redacted, yours on creation

The economics

The support fee is paid either way. The data, kits and scenarios are the return on it.

Request a sample dataset

$7.50

Variable fee for a P3 resolution

4.2 → 2.1

Median generation attempts on kit-covered categories

Tracked

Repeat-defect tickets avoided per regression scenario / quarter

Month 1

Share of tickets eligible for RLHF after consent and redaction

Why this exists

The problem

AI companies pay a support team to resolve failures and separately pay a data vendor to label synthetic examples of failures.

The support team sees the real ones every day and writes none of them down in a usable form. The data vendor never sees a real one. Two invoices, no loop.

The artefacts

The four things the loop produces. All redacted; client and customer identifiers removed.

Evaluation record

Recorded by the engineer at ticket close, in your schema. Included with Support.

Preference pair

The failing output and the engineer's corrected output, rated against your rubric and second-reviewed. Delivered within 2 business days of closure.

Kit spec

Fixes that recur become a reusable kit with prompt sequences and test cases. Accepted within 10 business days.

Regression scenario

The failure becomes a minimal reproducible test in your framework, run on every release within 2 business days.

The Flywheel

The loop

  1. 1

    Ticket

    A build fails; the platform gives a diagnosis.

  2. 2

    Evaluation record

    The engineer validates, resolves and records.

  3. 3

    Preference pair

    Failing and corrected outputs, rated and reviewed.

  4. 4

    Starter kit

    Recurring fixes become a reusable accelerator.

  5. 5

    Regression scenario

    The failure becomes a test on every release.

  6. 6

    Fewer tickets

    The queue shrinks; attainment rises.

Step 6 returns to step 1. A smaller queue of harder tickets produces better evaluation data, which improves the model further. The support fee is paid either way; the data, kits and scenarios are the return on it.

The loop

How support becomes training data, starter kits and regression tests.

Ticket

A customer's build fails. The platform produces an automated diagnosis.

Evaluation record

The engineer validates the diagnosis, resolves the issue and records diagnosis accuracy, output quality, capabilities used, verbatim feedback and any defect with reproduction steps — in your schema, at ticket close.

Preference pair

The failing output and the engineer's corrected output become a rejected/preferred pair, rated against your rubric, redacted, second-reviewed, delivered within 2 business days.

Starter kit

Fixes that recur become a reusable kit with prompt sequences and test cases, reducing generation attempts on the next hundred similar builds.

Regression scenario

The failure becomes a minimal reproducible test in your framework, run on every release.

Fewer tickets

The model improves on real failures, kits absorb common builds, and the regression suite catches repeats in staging. The queue shrinks; attainment rises.

The artefacts

The artefacts

The four things the loop produces. All redacted; client and customer identifiers removed.

evaluation_record.jsonRedacted
ticket_id
PX-2291-████
objective
Ship checkout with saved cards
diagnosis_class
partially_correct
reason_code
dns_propagation_assumed
correct_diagnosis
Edge proxy bound to stale build target
output_met_objective
no
objective_scale
2 / 5
capabilities_used
domain_config, deploy_logs, rebuild
attempts
2
customer_verbatim
"It said it deployed but the domain 502s"
verbatim_taxonomy
deploy.domain.misbinding
defect_repro_steps
4 steps, attached

Redacted. Client and customer identifiers removed.

We stopped buying synthetic failure data in month two.
Head of AI, Coding assistant

Reference values

The economics

The support fee is paid either way. The data, kits and scenarios are the return on it.

ItemReference value
Variable fee for a P3 resolution$7.50
Accepted RLHF record from the same ticketPer accepted record, set in SOW
Reduction in median generation attempts on kit-covered categories4.2 → 2.1
Repeat-defect tickets avoided per regression scenario per quarterTracked and reported
Share of tickets eligible for RLHF after consent and redactionMeasured in month 1

Controls

Governance

Consent-gated under your terms; opted-out tickets excluded. Redacted to your guidelines. Second-annotator review on at least 20%. Inter-annotator agreement measured monthly.

Your work product on creation. Never used to train anything of ours.

What you need in place

Most clients have three of the four on day one.

A rating rubric and schema

We can draft both under 06 Enablement.

Consent language

In your customer terms, covering training use.

A repository or evaluation tooling

To receive the records.

A test environment

For regression scenarios.

FAQs on the flywheel

FAQ

Evaluation records from day 10. RLHF records from week 3, once selection criteria and rubric are agreed.

No commitment either way

The support fee is paid either way

The support fee is paid either way. The data, kits and scenarios are the return on it.