Score everything. Classify the failures. Prove the fix.
Most AI deployments improve by anecdote: someone hears a bad call, someone edits a prompt, nobody can say whether it helped. The Hub treats improvement as an engineering discipline. Every conversation is scored, low scores become classified defects with an owner, revisions are proposed with the evidence attached, a person approves them, and nothing reaches production until it passes regression against a saved suite.
The rubric is derived from the outcome contract, which means any worker with a contract can be scored without bespoke work. The model produces the evidence. The code produces the number. A score computed by a model is not reproducible, cannot be regression tested and cannot be compared across prompt versions, so the scorer is deliberately dumb arithmetic and nothing about that is negotiable.
| Component | Contribution |
|---|---|
| Primary outcome | Carries the score. Each rung is worth materially less than the one above it, from full marks for the top rung down to nothing for no outcome at all |
| Outcome status | A tentative outcome earns partial credit, rising when a concrete confirmation path was agreed. Honesty is rewarded, and an inflated claim of confirmed is a defect rather than a win |
| Partial upgrade | A recorded, transcript supported path to a better outcome adds credit without allowing the higher rung to be claimed |
| Standing objectives | A fixed bonus for each one carrying evidence, with the total capped. Judged on whether the worker took what the conversation naturally offered, not on whether it ground through a checklist. The ceiling differs per contract, because the objective sets differ in size, and any report that hides that invites the conclusion that one set of callers is underperforming |
| Behavioural failures | Not deductions. Restating, stacked questions, premature endings, missed disclosure and unresolved identity are certification cells, and a blocking cell fails regardless of score. Trading them against a number is exactly what the separate gate exists to prevent |
A score exists to direct attention. It tells you which conversations to look at, which prompt layer to examine, which vertical is underperforming, and which skill someone should practise. That is the whole of its purpose, and the narrowness is deliberate.
Calibration asks whether this worker is good enough to deploy here, and failure looks like a low score, a regressed cell or a missing cue. Certification asks whether this pack may be listed for this channel in this jurisdiction, and failure looks like a compliant-seeming transcript that breaks a rule. They are separate gates because the failures look nothing alike, and a scoring threshold passes both.
A conversation that delivers every required identification line, but delivers one of them in turn three when the rule requires it before the first question, is non-compliant and scores well. A conversation that continues in English on a non-English turn scores well too, and produces something the contact cannot consent within. Both are blocking cells. Neither is reducible to a number.
Every channel crossed with every jurisdiction, each cell passed independently. Disclosure requirements, consent bases and permitted hours differ by jurisdiction, and a channel changes what a ladder rung can even mean, so a pass in one cell says nothing about any other. A jurisdiction with an unanswered dimension cannot be certified at all, because there is nothing complete to certify against.
Alongside the score, the evaluator returns typed findings, each anchored to a specific quote or turn reference. A number tells you a call went badly. A quoted finding tells you which line of which prompt layer caused it.
| Class | Meaning | Where it goes |
|---|---|---|
| Behavioural | The worker knew what to do and did not do it | Prompt revision queue |
| Knowledge | The worker lacked information it needed | Curation task in the knowledge base |
| Asset | The worker needed collateral that did not exist | The request board |
| Capability | The worker needed a tool or channel it does not have | Roadmap |
| Configuration | Wrong engine, thinking level, voice or scenario binding | Control plane |
| Compliance | A guardrail was approached or crossed | Immediate escalation, never queued |
The coach never rewrites a prompt wholesale, and that restraint is deliberate: wholesale rewrites destroy the accumulated calibration of a mature prompt and make regression impossible to attribute. Every version carries its author, its motivating findings, its regression result and its approver, so a prompt in production can always be traced back to the conversations that shaped it.
The behaviours that matter most appear rarely and cost the most when they do: an abrupt ending while a question is still open, a settled topic re-asked, a referral offered and not captured. Catching those needs volume that no human review can supply and no live pipeline can safely generate.
Every scenario against every contact archetype, attitude, difficulty, language and interaction history. Every cell populated rather than sampled, because the point is to find where a worker fails.
Friendly, concerned, busy, sceptical, hostile, gatekeeping, confused, monosyllabic, digressive, the wrong person entirely, and a contact who asks to be removed.
The simulated contact holds facts it will disclose only under stated conditions, so the harness knows the ground truth and can measure what the worker actually extracted.
Cue recall is the proportion of available intelligence the worker actually elicited. Contract fidelity is the proportion of what it elicited that reached the outcome record correctly.
Separating them matters because two very different failures look identical from outside. A worker that never asked the right question has a conversation problem. A worker that heard the answer and did not record it has a contract problem. Only ground truth tells them apart.
Text mode carries scale. Voice carries realism. Neither substitutes for the other.
The internal use of simulation is the one that never gets demonstrated and matters most.
Every candidate prompt version enters a regression run without anybody triggering it. Whether the revision came from the coach, from a steward's own edit or from a pack update, the run starts and the result is attached before a human looks at it. So an approval queue presents a proposed edit with its regression result already in hand, not a proposal somebody has to remember to test. A revision that fails never reaches a steward as a decision. It reaches them as a finding, which is a different and much cheaper conversation.
The cost of promoting a change is set by the cost of undoing one. A platform where reverting is hard develops a culture of not changing anything, and a worker nobody dares improve stops improving.
Every layer is independently versioned and every conversation records the version it ran under, so reverting takes seconds and is itself a recorded event with a reason.
Regression proves a version against the suite, not against the world. Live scores, abrupt endings, compliance findings and outcome distribution are compared against the pre-promotion window, and one that degrades past a declared threshold reverts on its own. Conservative on purpose: an unnecessary rollback costs a day, an unnoticed regression costs a quarter.
Revisions accumulate until the prompt is sediment nobody dares edit. A steward periodically merges them, proves equivalence by regression, prunes what no longer earns its place and re-baselines. Where a prompt gets shorter rather than longer.
The same harness pointed at a campaign rather than a worker, which answers a commercially larger question than prompt quality.
Resolve the campaign's real membership. Generate synthetic counterparties matched to the actual distribution in that list rather than a generic matrix, being the real mix of roles, seniorities, languages, industries and interaction histories it contains. Run at volume in text with a voice sample for realism. Then report.
| Output | The question it answers |
|---|---|
| Projected contact and conversion distribution | Is the coverage target achievable with the capacity allocated? |
| Projected cost, per outcome and in total | Is the cost per outcome viable before we spend anything? |
| Projected duration against the cycle window | Will this finish inside the window or run past it? |
| Failure modes, with the cells that produced them | Where does the ladder break for this population? |
| Messaging findings | Which openings and objection handlers landed against this audience specifically |
| Compliance dry run | How many of these people are actually contactable |
Two messages, two ladders, two list definitions, run against the same synthetic population. That makes it a real experiment because the population is held constant, and it happens before money is spent rather than after half of it is.
Synthetic counterparties are more cooperative, articulate and patient than real ones. So a simulated conversion rate is a relative comparison and an upper bound, never a forecast, and we say so where the number appears rather than in a footnote. Failure modes are treated as real, because one that occurs against a cooperative synthetic contact will certainly occur against a real one. The asymmetry runs one way.
A human sales coach normally works from anecdote, from the calls they happened to sit in on, and from the seller's own account of what happened. A domain coach on the Hub works from every scored conversation, every typed finding, every cue that was available and missed, and the sentiment trajectory of each call.
| Mode | What happens |
|---|---|
| Deal coaching | Works the qualification gaps on a live opportunity: which elements are unfilled, which rest on weak evidence, who the economic buyer likely is given the stakeholder map, and what single question would most reduce uncertainty next |
| Call preparation | The briefing before a human takes a call: what the account intelligence supports, the objections most likely for this role and industry, the insight worth leading with, and what the last conversation left open |
| Post call debrief | The same analysis machinery pointed at a person's call: what earned progress, what was left on the table, which cues were missed, and where the conversation turned |
| Skill development | Longitudinal rather than per call. Talk and listen ratio, question type distribution, discovery depth, objection handling success by category, and whether insights precede asks |
| Objection rehearsal | Live practice against a seeded simulated contact, at a chosen attitude, in a chosen language, as many times as you want, with no prospect at risk |
| Methodology teaching | Explaining a framework in the context of your own live deal rather than in the abstract, which is the only form of methodology training that survives a quarter end |
Rehearsal reuses the simulation harness, pointed at a person instead of a worker. Role play normally needs a manager's time, happens rarely, and is scored on impression. Here it is on demand, scored on the same rubric as production conversations, and repeatable until the pattern changes.
Every conversation your workers hold makes them measurably better, and the improvement is expressed in versioned artefacts you own and can inspect.