From the Interview to the Profile: Facts, Not Impressions
A competency profile is only valuable if you can identify its constituent elements. This page describes what the assessment module considers evidence of a given competency, what it does not consider evidence, and why a single successful conversation is not enough to influence the profile.
Metaskills assessment methodology, 2026
What counts as an evidence event
Four things we ask about every piece of evidence
Spontaneity
Did the behaviour appear on its own initiative, or only after the avatar applied explicit pressure? Both are recorded. They do not mean the same thing.
Timing
Did it land at the stage of the conversation where it could still change the outcome? The right sentence three minutes too late is a different result from the right sentence on time.
Fit to the other person
Was the style calibrated to who was actually there - what they needed emotionally and how much detail they could take? A well-formed explanation aimed at the wrong person is not good communication.
Proportionality
Was there too much of it? Excessive empathy undermines clarity as reliably as excessive assertiveness escalates tension. More of a good behaviour is not automatically a better score.
Some behaviours cap the result
How confident the profile is, stated out loud
Insufficient data
Too few samples, or everything from one and the same kind of situation. The competency is shown as not yet assessable. We would rather show a gap than a number nobody should rely on.
Emerging pattern
The behaviour has shown up in more than one type of context. Enough for preliminary feedback, stated as preliminary - with the limited confidence written on the report, not buried in a footnote.
Validated pattern
Repeated across several distinctly different contexts, including difficult scenarios rather than only comfortable ones. This is the point where the profile starts to be worth planning development around.
Stable under pressure
The behaviour holds in high-tension conversations and in the face of active resistance from the avatar, without quality falling away. The highest status the model assigns, and the rarest.
Why we do not take an average
Two layers, and a hard boundary
What this page does and does not claim
What this shows
- It describes the assessment logic: what qualifies as evidence, how quality is judged, how confidence is assigned and why aggregation is not an average.
- It shows the boundaries the architecture imposes on itself - no personality diagnosis, no static labels, practice excluded from assessment.
- It explains why the profile openly reports low confidence instead of producing a number for everything.
What this does not show
- It does not show that the AI assessment agrees with expert human judgement. That comparison is a separate piece of work and is still in progress.
- It is not evidence of learning effectiveness or of changed behaviour at work. Assessment methodology and training outcomes are different questions.
- It is not a basis for hiring, promotion or any other employment decision, and the system is deliberately built to stay out of those.
Limitations
This is a description of how the assessment is designed, written before the validation results are in. Both halves of that sentence matter.
- The exact thresholds behind the four confidence statuses - how many samples and how many contexts each one requires - are internal. Publishing them would let people optimise for the threshold instead of the behaviour.
- Prompt structures and the list of limiting behaviours stay internal for the same reason.
- Assessment is based on a conversation transcript. It does not capture body language, tone of voice or anything else outside the words, and it is not intended to.
- Large language models can be inconsistent. The design counters this with structured evidence, required quotations and cross-conversation aggregation, but it does not make the problem disappear.
- Validation against independent human assessors is in progress. Until it reports, the confidence statuses describe how much evidence there is, not proven predictive accuracy.
Source
Metaskills
2026-08-24
Methodology note prepared by Metaskills, with an external team of psychometrics and leadership-assessment specialists. This is an internal methodology note, not a peer-reviewed publication.
Digital Twin methodology
Related evidence
Validation methodology
How we are testing AI-generated assessment against independent human assessors, and what has not happened yet.
Validation methodologyDigital Twin
The competency profile this logic produces, and what it looks like for a learner and a team.
Digital TwinUsability of avatar-based VR training
Forty-seven participants rated the training experience. Usability evidence, clearly labelled as such.
Usability of avatar-based VR training