Research lines

Four lines. The evidence behind each differs considerably, and the difference is stated rather than smoothed over.

Construct validity in machine judgement

Whether an evaluation instrument measures the construct it declares, or the surface statistics that accompany it. The Cryptotype probe is the worked case: a controlled design removed my own headline result, and what remained was a different and smaller effect.

Status: one published dataset with confidence intervals; analyses exploratory; blind expert validation not yet run.

Evidence governance and denominator authorship

Who authors the reference a system is graded against, and whether the graded system could see it. Where a coverage metric derives its denominator from the system's own declaration, an omission leaves numerator and denominator together and the score does not move.

Status: engineering evidence across three independent axes with mutation tests; one real field case; reference independence is procedural rather than structural, which is the open residual.

Human and machine decision authority

Which stages of a decision should be model-authored, mechanically derived, or human-confirmed — and what a legal standard of care presupposes about the person whose conduct it measures. Article 14 of the EU AI Act enumerates the capacities that standard always assumed, and relies on capacities the regulated deployment erodes.

Status: legal and regulatory analysis resting on published empirical work by others.

Roman law and the standard of care

The bonus et diligens pater familias is not a fixed measure but a semantic field, graded by the internal logic of the legal relation. The tradition matters here because it separated by name two objects a learned system conflates: an idealised standard, and a record of what people actually did.

Status: completed thesis work; derived paper in preparation.