Per answer, then per skill, then overall
Each answer is evaluated on its own against what the question was examining, and the evaluation records why — what the answer demonstrated, what it missed, and which part of the response the judgement rests on.
Per-skill scores fold the answers that examined that skill. The overall score folds the per-skill scores using the weights you set on the position — a skill you marked mandatory counts for more than one marked nice-to-have.
The reasoning is not decoration
Every score carries its reasoning, and the report shows both. This is deliberate: a number without reasoning cannot be argued with, cannot be audited, and cannot be corrected. Where you disagree with a score, the reasoning tells you whether the model misread the answer or whether you and it are weighting the role differently — and those need different responses.
Evaluation leniency
Set per position or per flow. It moves the whole scale rather than individual judgements: a strict setting expects more evidence for the same score. It does not change what is assessed.
What is deliberately excluded
Accent, fluency for its own sake, and speech disfluency are not scored. Nor is anything demographic. A candidate reasoning aloud in their second language is assessed on the reasoning.
Recalculating
If you change a scoring configuration and want an existing interview re-scored, you can recalculate from the report. It re-runs the evaluation over the stored answers; it does not re-interview anyone, and the transcript is unchanged.