Bias controls

Score the work. Not the worker.

A live sales simulation is a work-sample assessment, one of the strongest predictors of job performance in industrial-organizational psychology. Here is what Ready measures, what it deliberately excludes, and why the score guides the review while a person always makes the final decision.

The science

Work-sample assessment is among the strongest predictors of job performance.

Schmidt and Hunter's foundational meta-analysis and the Sackett, Zhang, Berry, and Lievens update both rank work-sample tests at the top of the predictive-validity ladder, alongside structured interviews. A recorded sales conversation, scored against a structured rubric, is exactly that combination digitised.

Roth, Bobko, and McFarland showed that work-sample tests also carry lower subgroup differences than cognitive ability tests at equivalent predictive power. That is the scientific reason Ready scores the conversation, not the candidate's traits, demographics, or background.

  • Validation strategy. Content validity, where an expert sales panel rates each rubric dimension as job-relevant, paired with criterion validity, which we build with each customer over time by correlating rubric scores against how those hires actually perform.
  • Reference standard. SIOP Principles for the Validation and Use of Personnel Selection Procedures, the consensus standard in industrial-organizational psychology.
  • Job-relevance test. Every scored dimension must describe behavior the candidate would perform in the actual role. Nothing else makes it into the rubric.
Rubric weighting

Mostly content. Some structure. Almost no delivery.

The rubric weights what the candidate said and how they structured the conversation. Content of answers (specific outcomes, quantified results, customer language) and conversation structure (qualifying questions, objection handling, methodology adherence) account for the great majority of the score. Delivery features are bounded to a small slice, and only on dimensions that are job-relevant and validated.

The same evaluator scores every candidate for a role, against the same KPIs, and you choose those KPIs when you create the role. Everyone is measured by the standard you set, not by who they are or which reviewer they happened to get.

  • Same evaluator, your KPIs. One consistent evaluator, the KPIs you defined for the role, applied identically to every candidate.
  • Discovery discipline. Question sequencing, hypothesis testing, follow-up depth, and methodology pillars covered (MEDDIC, SPIN, GAP, Sandler, Challenger).
  • Stakeholder navigation. Treatment of conflicting incentives across finance, technology, procurement, and end-user personas in a multi-stakeholder buyer committee.
  • Evidence quality. Specificity of claims, requested proof, acknowledgement of uncertainty, and correction after challenge.
  • Objection handling pattern. Acknowledge, clarify, reframe, confirm. Scored against the pattern, not against the candidate's accent or pitch.
What we exclude

No facial analysis. No voiceprint. No accent penalty.

The evaluator sees the conversation and nothing else. No name, no photo, no CV, no background: it has no idea who is behind the simulation. Beyond that, Ready does not score facial expression, emotional inference, or personality. Those signals correlate with disability, gender, age, and cultural background, and they carry near-zero predictive validity for actual job performance. They are excluded by rubric design.

Ready does not store voiceprints or speaker recognition templates beyond the live assessment session. Accent and phonetic markers are not penalty signals. A candidate speaking English as a second language is scored on the qualifying question they asked, not the way they pronounced it.

  • No facial or emotional analysis. Excluded by design. Near-zero predictive validity and high correlation with disability and cultural variation.
  • No voiceprint storage. Voice embeddings and speaker templates are discarded at the end of the session. The only artifacts kept are the transcript, scenario state, and rubric scores.
  • No pitch, warmth, or pace as bias-correlated features. Vocal warmth and pace are gender-correlated and age-correlated. Excluded from the scoring rubric.
  • No demographic inputs. Name, photo, age, gender, nationality, school, and employer never flow into the scoring model.
Fairness monitoring

What we put in place with each customer.

An AI evaluator, like any model, can carry bias. We are honest about that, and we treat it as something to monitor and contain rather than something to claim away. The controls below are commitments we agree with you in writing when we set up a deployment, not silent guarantees already running in the background of the product.

The standard we hold a deployment to is the four-fifths ratio that employment regulators use: whether the lowest-passing group passes at less than eighty percent of the highest-passing group's rate. Today Ready collects no demographic data at all, so this monitoring depends on demographic fields a candidate would opt into, and on the independent third-party bias audit we commission per customer where the law requires it, before the first regulated candidate is screened and refreshed annually. The automated, in-product impact-ratio dashboard and scoring pause are on the roadmap, not live yet; we say so plainly rather than imply a monitor that is not there.

  • The four-fifths threshold. The benchmark we commit to per deployment: the lowest-passing group should not pass at less than eighty percent of the highest-passing group's rate, per protected class.
  • No demographic data collected today. Ready collects no race, gender, age, or disability data and never infers it. Any impact-ratio monitoring relies on demographic fields a candidate opts into, which is why this is a deployment commitment rather than an automatic feature.
  • Independent audit where required. Where a jurisdiction requires it, an independent third-party bias audit is commissioned for the deployment before the first regulated candidate is screened, and refreshed annually.
  • Automated drift dashboard (roadmap). An in-product impact-ratio dashboard with an alert-and-pause on drift is planned. It is not running yet, and we do not present it as if it were.
Human decision

The score guides the review. A person makes the call.

Every simulation is recorded and saved. The full recording, the transcript and the scenario state are kept, so any candidate's call can be listened to and judged by a person, not taken on a number alone. A score is never delivered without the evidence behind it.

Because an AI can carry bias, we do not treat the score as the hiring decision. We recommend using it as navigation: it tells you which candidates to listen to first. Your team makes the call from the actual recording and is free to disagree with the score entirely; the outcome a reviewer sets is recorded against their account.

A strong score may surface a candidate for review sooner, but it never advances or ends an application on its own. A person decides every outcome.

  • Recorded and reviewable. Every call is saved in full, so a person can listen back and assess it directly, not just read a score.
  • The score is navigation. Use it to decide who to review first. It orders the work, it does not make the decision.
  • A human always decides. Ready never auto-rejects and never auto-hires. The hiring decision sits with your team.
  • Decisions are recorded. The outcome a reviewer sets is logged against their account. Granular per-line score override with a mandatory written reason is on the roadmap.

Want to see the rubric and how review works?

Talk to us. We can walk you through the rubric, the KPIs you would set for your role, and how your team reviews the recordings and the scores.