Score everything. Classify the failures. Prove the fix.
Most AI deployments improve by anecdote: someone hears a bad call, someone edits a prompt, nobody can say whether it helped. The Hub treats improvement as an engineering discipline. Every conversation is scored, low scores become classified defects with an owner, revisions are proposed with the evidence attached, a person approves them, and nothing reaches production until it passes regression against a saved suite.
The rubric is derived from the outcome contract, which means any worker with a contract can be scored without bespoke work.
| Component | Contribution |
|---|---|
| Primary outcome | Carries the score. Each rung is worth materially less than the one above it, from full marks for the top rung down to nothing for no outcome at all |
| Outcome status | A tentative outcome earns partial credit, rising when a concrete confirmation path was agreed. Honesty is rewarded, and an inflated claim of confirmed is a defect rather than a win |
| Partial upgrade | A recorded, transcript supported path to a better outcome adds credit without allowing the higher rung to be claimed |
| Standing objectives | Bonuses, capped in aggregate. Judged on whether the worker took what the conversation naturally offered, not on whether it ground through a checklist |
| Behavioural compliance | Cascade violations, restating, stacked questions, manufactured urgency, premature endings, missed disclosure and unresolved identity are deductions regardless of outcome |
A score exists to direct attention. It tells you which conversations to look at, which prompt layer to examine, which vertical is underperforming, and which skill someone should practise. That is the whole of its purpose, and the narrowness is deliberate.
Alongside the score, the evaluator returns typed findings, each anchored to a specific quote or turn reference. A number tells you a call went badly. A quoted finding tells you which line of which prompt layer caused it.
| Class | Meaning | Where it goes |
|---|---|---|
| Behavioural | The worker knew what to do and did not do it | Prompt revision queue |
| Knowledge | The worker lacked information it needed | Curation task in the knowledge base |
| Asset | The worker needed collateral that did not exist | The request board |
| Capability | The worker needed a tool or channel it does not have | Roadmap |
| Configuration | Wrong engine, thinking level, voice or scenario binding | Control plane |
| Compliance | A guardrail was approached or crossed | Immediate escalation, never queued |
The coach never rewrites a prompt wholesale, and that restraint is deliberate: wholesale rewrites destroy the accumulated calibration of a mature prompt and make regression impossible to attribute. Every version carries its author, its motivating findings, its regression result and its approver, so a prompt in production can always be traced back to the conversations that shaped it.
The behaviours that matter most appear rarely and cost the most when they do: an abrupt ending while a question is still open, a settled topic re-asked, a referral offered and not captured. Catching those needs volume that no human review can supply and no live pipeline can safely generate.
Every scenario against every contact archetype, attitude, difficulty, language and interaction history. Every cell populated rather than sampled, because the point is to find where a worker fails.
Friendly, concerned, busy, sceptical, hostile, gatekeeping, confused, monosyllabic, digressive, the wrong person entirely, and a contact who asks to be removed.
The simulated contact holds facts it will disclose only under stated conditions, so the harness knows the ground truth and can measure what the worker actually extracted.
Cue recall is the proportion of available intelligence the worker actually elicited. Contract fidelity is the proportion of what it elicited that reached the outcome record correctly.
Separating them matters because two very different failures look identical from outside. A worker that never asked the right question has a conversation problem. A worker that heard the answer and did not record it has a contract problem. Only ground truth tells them apart.
Text mode carries scale. Voice carries realism. Neither substitutes for the other.
A human sales coach normally works from anecdote, from the calls they happened to sit in on, and from the seller's own account of what happened. A domain coach on the Hub works from every scored conversation, every typed finding, every cue that was available and missed, and the sentiment trajectory of each call.
| Mode | What happens |
|---|---|
| Deal coaching | Works the qualification gaps on a live opportunity: which elements are unfilled, which rest on weak evidence, who the economic buyer likely is given the stakeholder map, and what single question would most reduce uncertainty next |
| Call preparation | The briefing before a human takes a call: what the account intelligence supports, the objections most likely for this role and industry, the insight worth leading with, and what the last conversation left open |
| Post call debrief | The same analysis machinery pointed at a person's call: what earned progress, what was left on the table, which cues were missed, and where the conversation turned |
| Skill development | Longitudinal rather than per call. Talk and listen ratio, question type distribution, discovery depth, objection handling success by category, and whether insights precede asks |
| Objection rehearsal | Live practice against a seeded simulated contact, at a chosen attitude, in a chosen language, as many times as you want, with no prospect at risk |
| Methodology teaching | Explaining a framework in the context of your own live deal rather than in the abstract, which is the only form of methodology training that survives a quarter end |
Rehearsal reuses the simulation harness, pointed at a person instead of a worker. Role play normally needs a manager's time, happens rarely, and is scored on impression. Here it is on demand, scored on the same rubric as production conversations, and repeatable until the pattern changes.
Every conversation your workers hold makes them measurably better, and the improvement is expressed in versioned artefacts you own and can inspect.