Skip to main content

METHODOLOGY

How a conversation becomes an evidence-based score

This is about how scores are derived. For what happens during the interview itself, see /how-it-works. For the 16 dimensions themselves, see the competency framework.

What counts as evidence

Not every answer counts the same. Every claim you make during the interview is weighed against five tiers, from strongest to weakest:

  1. Direct impact you owned

    A decision you personally made, with a tradeoff you can describe and an outcome you can point to. This is the strongest form of evidence.

  2. Influenced impact you contributed to

    You changed the outcome without owning the final call -- real evidence of influence, weighted differently than direct ownership.

  3. Team or company impact you participated in

    You were part of the effort without a clear individual contribution to point to. Useful context, not strong evidence on its own.

  4. Qualitative evidence when metrics were unavailable

    Not every real decision has a clean number attached. Judgment, tradeoffs, and stakeholder outcomes count when they are specific.

  5. Unsupported impact claims

    A claim without a specific example, tradeoff, or verifiable detail behind it. Recorded as a gap, not scored as if it were proven.

Confidence, scored separately from the number itself

Every dimension score carries its own confidence level, so a report never presents a thin score with the same certainty as a well-evidenced one.

Low confidence
The evidence points in a direction, but it is thin, single-example, or partially unsupported.
Medium confidence
A real, specific example exists and holds up under a follow-up question, but broader pattern evidence is limited.
High confidence
Multiple specific examples, consistent under follow-up, ideally spanning more than one situation.

Reading a 1-to-10 score

Every one of the 16 dimensions is scored on the same 1-to-10 scale. Here is roughly what each band reflects:

1-2Little relevant evidence surfaced
3-4Conceptual understanding, limited real application
5-6Demonstrated within a bounded scope
7-8Demonstrated repeatedly across meaningful complexity
9-10Sustained demonstration at broad organizational scale

A dimension the interview never meaningfully probed is marked not evaluated rather than scored -- the report hides the number in that case instead of showing a mid-range score that would read as a real, demonstrated rating it isn’t.

What a report does, and does not, conclude

A report states the highest scope of responsibility the evidence supports today, not a prediction of your ultimate potential or career ceiling. It separates what was clearly demonstrated from what was claimed but not backed up when asked to be specific, and it names the evidence still needed rather than guessing past a gap.

For where things honestly stand on formal validation and bias-audit work, see /trust.

See it applied to a real conversation