The measurement approach

Meaisure is built on a small number of principles from the assessment literature — and on an equally explicit list of what we do not claim.

Observed behavior beats self-report

Questionnaires and interviews measure what people say about themselves, which bends under self-presentation and simple lack of self-knowledge. A live simulation measures the behavior itself: under time pressure and conflicting demands, the default way you operate shows up whether you intend it to or not. It is also much harder to fake a whole session than to pick the right answer.

One anchored rubric, applied the same way every time

Each dimension is scored against behaviorally anchored criteria — descriptions of what stronger and weaker responses concretely look like — rather than a rater's general impression. The same rating model is applied to every session, which is what makes sessions comparable at all.

Process data, not just transcript

Compliance is an act, not a statement: the system records whether you actually consulted the guidelines before invoking them, how long decisions took, and which timed decisions expired unanswered. These signals feed the evaluation alongside what you wrote, and they are the part a polished writing style cannot cover for.

Scenarios from a calibrated bank

Scenarios are drawn from a growing bank of templates, and an ongoing calibration program measures each template's difficulty, the consistency of rankings across scenarios, and the stability of the AI scorer across repeated evaluations of the same session. The goal is that two people playing different scenarios can still be measured on a common footing — the standard problem of any exam with multiple forms, taken seriously.

Levels, not verdicts

Results are reported as development levels with an explicit behavioral diagnosis and growth moves anchored to specific moments. Fine-grained scores exist inside the measurement machinery — the debrief is deliberately built for learning, not for labeling.

What we don't claim

  • · No predictive promises. We do not claim that a Meaisure result predicts job performance. Validating that kind of claim takes longitudinal evidence we do not have.
  • · The AI scorer is under active validation. We monitor its stability across repeated runs, and comparison against expert human raters is part of the validation roadmap — until then, treat results as structured, evidence-anchored feedback.
  • · One scenario is one sample. A single session shows how you behaved in one situation. Confidence grows with repeated sessions across different scenarios, not from any single debrief.