Metaskills
RESEARCH · DIGITAL TWIN · METHODOLOGY

From the Interview to the Profile: Facts, Not Impressions

A competency profile is only valuable if you can identify its constituent elements. This page describes what the assessment module considers evidence of a given competency, what it does not consider evidence, and why a single successful conversation is not enough to influence the profile.

Metaskills assessment methodology, 2026

MethodologyMetaskillsMethodology noteMethodology note - part of the Digital Twin series

What counts as an evidence event

The unit the engine works with is not an overall impression of the conversation. It is an evidence event: a specific moment in the transcript where something the person did demonstrates a competency in action, quotable and locatable. One rule sits underneath everything else on this page: the mere occurrence of a behaviour is not enough. Somebody can say the words of an apology, ask a checking question or name an emotion and none of it counts if it arrived at the wrong moment, missed the person in front of them or came only after being pushed into it. That is why every evidence event goes through a quality assessment rather than a tick in a box.

Four things we ask about every piece of evidence

Spontaneity

Did the behaviour appear on its own initiative, or only after the avatar applied explicit pressure? Both are recorded. They do not mean the same thing.

Timing

Did it land at the stage of the conversation where it could still change the outcome? The right sentence three minutes too late is a different result from the right sentence on time.

Fit to the other person

Was the style calibrated to who was actually there - what they needed emotionally and how much detail they could take? A well-formed explanation aimed at the wrong person is not good communication.

Proportionality

Was there too much of it? Excessive empathy undermines clarity as reliably as excessive assertiveness escalates tension. More of a good behaviour is not automatically a better score.

Some behaviours cap the result

Not everything in a conversation adds up. A defined set of destructive behaviours acts as a limiter rather than a deduction - behaviour that damages the other person or the working relationship, rather than merely failing to help. Which behaviours belong to that set is part of the model and stays internal; the principle is the part worth publishing. If one of them lands at a decisive moment in the scenario, it caps the score achievable for the related competency, and positive behaviour elsewhere in the same conversation does not buy it back. This is deliberate. In a real clinical or managerial conversation, one damaging move does not get averaged away by five good ones, and an assessment that pretended otherwise would be teaching the wrong lesson.

How confident the profile is, stated out loud

1

Insufficient data

Too few samples, or everything from one and the same kind of situation. The competency is shown as not yet assessable. We would rather show a gap than a number nobody should rely on.

2

Emerging pattern

The behaviour has shown up in more than one type of context. Enough for preliminary feedback, stated as preliminary - with the limited confidence written on the report, not buried in a footnote.

3

Validated pattern

Repeated across several distinctly different contexts, including difficult scenarios rather than only comfortable ones. This is the point where the profile starts to be worth planning development around.

4

Stable under pressure

The behaviour holds in high-tension conversations and in the face of active resistance from the avatar, without quality falling away. The highest status the model assigns, and the rarest.

Why we do not take an average

The engine is explicitly prohibited from computing a simple arithmetic mean of session scores. An average of a strong performance in an easy conversation and a weak one under pressure produces a middling number that describes neither, and it hides the only thing worth knowing: which situations this person handles and which ones they do not. Instead, aggregation looks at the pattern - how much evidence there is, how varied the contexts are, how difficult the scenarios were, and whether quality holds when the conversation gets hard. Context diversity is a requirement, not a bonus: a behaviour confirmed only in one relationship type and one kind of counterpart does not become a confirmed pattern however often it repeats there. One more filter runs before any of this. Guided practice, where the learner gets prompts during the conversation, is kept out of the assessment data. Practice is for learning; it would flatter the profile if it counted.

Two layers, and a hard boundary

Assessment runs in two separate layers. The first analyses one completed conversation and does nothing else: it reads the transcript against the scenario and the avatar's profile and returns structured evidence events. It is not allowed to draw conclusions about the person. The second layer works across many conversations. It looks for patterns, including interference between competencies - the case where strong active listening is quietly masking a real deficit in assertiveness or decision-making - and it is this layer that assigns the confidence status and writes the justification for a competency level. The boundary between them is the point. Judgement about a person may only be made from evidence gathered across many situations, never from one conversation. And it stops there: the architecture prohibits diagnosing stable personality traits and forbids static labels such as you are an empathetic person. Feedback describes behaviour in a context, states the confidence behind it, and names the next scenario worth practising.

Source

Metaskills

2026-08-24

Methodology note prepared by Metaskills, with an external team of psychometrics and leadership-assessment specialists. This is an internal methodology note, not a peer-reviewed publication.

Related evidence

Methodology

Validation methodology

How we are testing AI-generated assessment against independent human assessors, and what has not happened yet.

Validation methodology
Product

Digital Twin

The competency profile this logic produces, and what it looks like for a learner and a team.

Digital Twin
Study

Usability of avatar-based VR training

Forty-seven participants rated the training experience. Usability evidence, clearly labelled as such.

Usability of avatar-based VR training