Evaluation record
Recorded by the engineer at ticket close, in your schema. Included with Support.
The Flywheel
Every resolved ticket is a labelled example of where the model fell short and what right looked like. This chapter explains what we do with it.
The economics
$7.50
Variable fee for a P3 resolution
4.2 → 2.1
Median generation attempts on kit-covered categories
Tracked
Repeat-defect tickets avoided per regression scenario / quarter
Month 1
Share of tickets eligible for RLHF after consent and redaction
Why this exists
AI companies pay a support team to resolve failures and separately pay a data vendor to label synthetic examples of failures.
The support team sees the real ones every day and writes none of them down in a usable form. The data vendor never sees a real one. Two invoices, no loop.
The artefacts
Recorded by the engineer at ticket close, in your schema. Included with Support.
The failing output and the engineer's corrected output, rated against your rubric and second-reviewed. Delivered within 2 business days of closure.
Fixes that recur become a reusable kit with prompt sequences and test cases. Accepted within 10 business days.
The failure becomes a minimal reproducible test in your framework, run on every release within 2 business days.
The Flywheel
A build fails; the platform gives a diagnosis.
The engineer validates, resolves and records.
Failing and corrected outputs, rated and reviewed.
Recurring fixes become a reusable accelerator.
The failure becomes a test on every release.
The queue shrinks; attainment rises.
Step 6 returns to step 1. A smaller queue of harder tickets produces better evaluation data, which improves the model further. The support fee is paid either way; the data, kits and scenarios are the return on it.
The loop
A customer's build fails. The platform produces an automated diagnosis.
The engineer validates the diagnosis, resolves the issue and records diagnosis accuracy, output quality, capabilities used, verbatim feedback and any defect with reproduction steps — in your schema, at ticket close.
The failing output and the engineer's corrected output become a rejected/preferred pair, rated against your rubric, redacted, second-reviewed, delivered within 2 business days.
Fixes that recur become a reusable kit with prompt sequences and test cases, reducing generation attempts on the next hundred similar builds.
The failure becomes a minimal reproducible test in your framework, run on every release.
The model improves on real failures, kits absorb common builds, and the regression suite catches repeats in staging. The queue shrinks; attainment rises.
The artefacts
The four things the loop produces. All redacted; client and customer identifiers removed.
Redacted. Client and customer identifiers removed.
We stopped buying synthetic failure data in month two.
Reference values
The support fee is paid either way. The data, kits and scenarios are the return on it.
| Item | Reference value |
|---|---|
| Variable fee for a P3 resolution | $7.50 |
| Accepted RLHF record from the same ticket | Per accepted record, set in SOW |
| Reduction in median generation attempts on kit-covered categories | 4.2 → 2.1 |
| Repeat-defect tickets avoided per regression scenario per quarter | Tracked and reported |
| Share of tickets eligible for RLHF after consent and redaction | Measured in month 1 |
Controls
Consent-gated under your terms; opted-out tickets excluded. Redacted to your guidelines. Second-annotator review on at least 20%. Inter-annotator agreement measured monthly.
Your work product on creation. Never used to train anything of ours.
Most clients have three of the four on day one.
We can draft both under 06 Enablement.
In your customer terms, covering training use.
To receive the records.
For regression scenarios.
FAQs on the flywheel
No commitment either way
The support fee is paid either way. The data, kits and scenarios are the return on it.