Metaskills
RESEARCH · DIGITAL TWIN · PROJECT FRAMEWORKS

Privacy, Bias, and the Human Factor

When artificial intelligence analyzes how a person handled a difficult conversation, the design decisions related to that analysis are just as important as the analysis itself. This page describes the framework underlying the “Digital Twin” tool: what data the assessment tool can take into account, how known biases are addressed, and at what points in the process humans are involved.

Metaskills design framework, 2026

MethodologyMetaskillsMethodology noteDesign framework - see the security page for what runs in production today

What this page is, and what it is not

This page describes a design framework: the principles and constraints the Digital Twin is being built to. It is not a description of the production system and should not be read as one. If you need to know what is running today - where data sits, who can reach it, what the contracts say - that is the security page, and it is the only page we treat as the source of truth on the current state. Keeping the two apart is the point: a design principle and a shipped control are different things, and blurring them is how vendors end up promising infrastructure they do not have.

Four principles the assessment is designed around

1

Minimise what the model sees

The assessment engine needs the substance of a conversation, not the identity of the person who had it. The design principle is that identifying details are stripped from what goes to the model, leaving a depersonalised behavioural abstract.

2

Keep identity apart from analytics

Who somebody is and what their competency profile says are designed to live in separate places, joined by a reference rather than stored together. One breach should not hand somebody both halves.

3

No biometrics, by design

The design excludes biometric identification and biometric profiling, and with them eye tracking, facial micro-expression analysis and video from the headset cameras: the assessment reads what was said, not the body that said it. What that means for the product as it stands today is on the security page, which is the one place we keep that answer.

4

Formative use only

The Digital Twin is designed as a development tool. It is not built to feed hiring, promotion or disciplinary decisions, and the whole assessment logic - behaviour in context, stated confidence, no personality diagnosis - follows from that choice.

The questions that come before any of this

Two things get asked before anything else on this page: whether client data is used to train models, and where the data actually sits. Both are answered on the security page, and we deliberately do not repeat the answers here. A commitment written down in two places, in four languages, is a commitment that will eventually say two different things - so there is one page that carries them and one page to update when they change. The framework described here is about how the assessment is designed to work. It is not a substitute for the technical and contractual documentation, and where the two appear to disagree, the security page is right.

Four kinds of bias we design against

Eloquence bias

Language models reward fluent speakers. Someone who stammers, works in a second language or speaks colloquially can be marked down for how they said it rather than what they did.

Gender and stereotype bias

The same firmness read as decisive in one speaker and abrasive in another. Training data carries these patterns; an assessment engine will reproduce them unless it is stopped from doing so.

Cultural bias

Communication norms differ. A model trained mostly on direct, low-context communication will systematically undervalue people whose professional culture is more indirect.

Evaluative hallucination

The failure mode specific to this application: a confident, well-written assessment of something that did not happen in the conversation at all.

The countermeasure that does most of the work

Every one of those biases is countered the same way: the assessment is tied to what was said, not to how it sounded or to who appeared to be saying it. The engine is instructed to disregard names, pronouns and demographic markers if they happen to be spoken, and to look past grammatical errors, hesitation and colloquialism to the intent underneath. That is the design answer to eloquence and stereotype bias. Against hallucination the rule is stricter. An assessment must quote the exact words from the transcript that justify it. If the quote is not there, the evaluation is rejected rather than reported - a claim about somebody's behaviour with no quotable moment behind it is not a finding, it is a guess with good grammar. None of this eliminates bias. It removes the easiest routes to it and makes the remainder visible, which is why the validation study runs against human assessors rather than against our own expectations.

Where a human stays in the loop

An assessment model left alone drifts. Criteria loosen, the same behaviour scores differently in March than it did in January, and nobody notices because there is nothing to notice against. The framework answers this with human-in-the-loop monitoring: trained behavioural assessors review anonymised samples of transcripts blind and compare their evaluations with what the system produced. Where they diverge, the criteria are recalibrated. It is the same comparison that sits at the centre of the validation study, run as a continuing check rather than a one-off. A learner can also ask for human review of feedback they disagree with. Automated feedback that cannot be questioned is not feedback - it is a verdict, and this system is not designed to hand down verdicts.

Source

Metaskills

2026-08-24

Design framework prepared by Metaskills. It describes principles the Digital Twin is built to, not the state of the production system - for that, see the security page.

Related pages

Security

Security and data protection

What runs in production today: where data is held, who can reach it and what the contracts cover.

Security and data protection
Methodology

How behaviour becomes evidence

The assessment logic these principles constrain - evidence events, confidence and the limits on inference.

How behaviour becomes evidence
Research

Research & Evidence

Everything we have published, with sources and limitations stated.

Research & Evidence