Skip to main content
BENCHProfessional Twin

BenchMark Research

Can we infer what someone is capable of from how they work, not just what they produce?

Capability is judged through CVs, interviews, manager reviews and finished work. AI makes those signals weaker, because a polished output says less about the reasoning, verification and judgment behind it.

BenchMark Research is the open programme behind BENCH. Its subject is capability inference: what can be known about professional ability from the process of the work rather than its product. The notes below are working material, and they are published as they are written.

Capability is a latent state.

Capability is a latent state. We cannot observe it directly. We see fragments of it through what someone notices, checks, ignores, prioritises, revises, escalates and decides under changing conditions.

Capability ontology

A structured map of the capabilities being inferred and the relationships between them: reasoning, verification, prioritisation, judgment, communication, adaptability, independence.

Professional world model

The meaning of an action depends on context, so the system represents the task, the actors, the information available, the goals, the constraints and the consequences around each decision.

Behavioural evidence

Signals from the path through a task, not only the final answer: what someone checks, what they miss, how they revise, and how they respond when the situation changes.

Capability inference

The engine that combines those observations and updates the Professional Twin, estimating the capability underneath performance rather than scoring the output.

The two hard problems.

Neither is solved. They are set down in the state the work is actually in, with what each one would have to overcome, so that anyone thinking about the same question can see where it stands and where it breaks.

Capability as topology

Professional capability is probably not a row of independent scores. Verification can constrain judgment. Poor prioritisation can make strong reasoning ineffective. Some capabilities may act as bridges while others become bottlenecks.

Graph-based representations, inspired by brain topology. This is a modelling analogy, not a claim that BENCH maps biological neural activity.

Predictive representations

JEPA-style representation learning: rather than reconstructing every observable action, the system should learn an internal representation that captures enough of the underlying state to predict what matters next.

The practical test is whether evidence from prior tasks produces a representation that predicts how someone performs on a new one.

Notes.

The work is written down as it is done, in full, including the parts that have not resolved. Read it, argue with it, cite it.

  1. Research note

    Inferring Professional Capability from Behaviour

    The founding note. Why a finished document has stopped carrying evidence of the judgment behind it, and what a system would have to model, from the process of work rather than its product, to recover it.

The BENCH Brief

New notes as they are published, and occasional field notes on legal capability, AI-era supervision and how professional judgement develops.