Metaskills
RESEARCH · DIGITAL TWIN · VALIDATION PROCESS UNDERWAY

A Comparison of AI-Based Assessment Results with Those Conducted by Humans

We are publishing the study plan before the results are announced. The study involves comparing the conclusions drawn by artificial intelligence about a given conversation with the conclusions of independent human evaluators regarding the same conversation. The methodology has already been developed, and we are now moving into the testing phase-this page describes what we intend to do, and after the results are published, this page will be updated at this address.

Metaskills validation study - design

ValidationMetaskillsOngoing researchStudy design published in advance - trials not yet completed

Why publish a design before the results

Any vendor can announce that their AI agrees with expert judgement once the numbers are in and the awkward variants have quietly been dropped. Publishing the method first is a way of removing that option from ourselves. So, in plain terms: the validation trials described below have not been carried out yet. Nothing on this page is a result. What exists today is a finished study design, a defined scope, a defined participant group and the assessment tooling the human assessors will use.

How is it constructed? comparison

1

The same scenario, twice

The participant engages in a structured conversation with an AI-based avatar and, at the same time, goes through the same scenario with a real actor acting according to identical parameters. Only one variable changes; all others remain constant.

2

Two independent evaluators from a group of people

Qualified evaluators analyze the transcripts without access to the results generated by artificial intelligence. That is precisely the point of independence-an evaluator who has seen the machine-generated answer is no longer a control group.

3

Observation Sheet Templates

Assessors use special rubrics designed for the competencies being assessed, guided by the same dimensions, criteria, and proficiency levels specified by the model-ensuring that both people and machines answer the same question.

4

Compare, and then prepare a report

The results generated by artificial intelligence and the interpretations of behavior are compared with human evaluations of the same conversations. The results, including instances of discrepancies, will be published at this address.

Six competencies in the first round

Adaptability

Shifting approach when the situation moves, without losing the outcome that matters.

Active listening

Catching, checking and reflecting back what was actually said.

Assertiveness

Holding a position and setting boundaries without escalating the conversation.

Empathy

Registering what the other person is feeling and responding to it.

Evidence-based approach

Grounding a position in evidence rather than status or volume.

Insightful clarity

Making a complex thing understandable to the person actually in the room.

Who takes part, and why it matters

The participant group consists of healthcare professionals and clinical leaders who have agreed to take part in the trials. This is a deliberate constraint rather than convenience. Communication competencies are context-bound. A model validated on graduate assessment-centre exercises tells you very little about a consultant explaining a treatment decision to a frightened family, or a ward manager reassigning duties mid-shift. If we intend the Digital Twin to be used in healthcare, the validation has to happen with people who actually have those conversations. A preliminary comparison, run on a single competency during model calibration, informed the design of this study. It was too small to constitute evidence of accuracy and we do not present it as such - it told us where the study had to be tighter.

What this page does and does not claim

What this shows

  • It shows a completed study design: comparison against human assessors, the competencies in scope, the participant group and the assessment tooling.
  • It shows what we consider a fair test of an AI assessment - and it is on record before the results are.

What this does not show

  • It does not show that AI assessment agrees with human judgement. The trials have not been run yet.
  • There are no results on this page. Any number here would describe the design, not an outcome.
  • It is not psychometric validation of the whole competency model. The first round covers six of the eighteen competencies.
  • It says nothing about learning effectiveness or about behaviour change at work. Those are separate questions and need separate studies.

Limitations we already know about

Some of these we can fix in the next round. Others are inherent to this kind of study, and pretending otherwise would defeat the purpose of publishing the design.

  • The trials have not been carried out. Everything here is a plan.
  • The first round covers six competencies of eighteen. It will not tell us whether the remaining twelve are assessed reliably.
  • Human assessors are not ground truth. They disagree with each other too, and agreement between AI and assessor is evidence of consistency, not of correctness.
  • The participant group is healthcare professionals. Results will not automatically transfer to other sectors.
  • A human actor playing a scenario is not the same thing as a real conversation with real consequences. The comparison isolates the assessment, not the stakes.
  • This is internal research, not a peer-reviewed study, and we will label the results the same way when they arrive.

Source

Metaskills

2026-08-24

Study design prepared by Metaskills with an external team of psychometrics and leadership-assessment specialists. Design published in advance of the trials; this page will be updated with results at the same address.

Related evidence

Methodology

The competency model

The 18 competencies, four domains and the structure being validated here.

The competency model
Study

Usability of avatar-based VR training

A study that has reported: 47 participants rating the training experience, with its own limits stated.

Usability of avatar-based VR training
Product

Digital Twin

The competency profile this logic produces, and what it looks like for a learner and a team.

Digital Twin
Validating AI assessment against humans | Metaskills