What the index cannot do
The project anchors itself in Lincoln and Guba's criteria of trustworthiness, and dependability and confirmability are established by saying plainly where a method stops. This page is that statement.
Known limitations
The index describes design at adoption, not governance today
This follows directly from rule 2, and it is the limitation most likely to be mistaken for an error. A framework that has since been amended, deferred or quietly ignored still scores as it was written on the day it was agreed. That is a defensible scope — it is what makes the five frameworks comparable — but it is a scope, not a claim about the present.
Design is not outcome
Rule 1 keeps reception out of the scores, which means the index cannot distinguish a framework whose enforcement powers are used from one whose identical powers sit unused. Two frameworks can score alike and behave nothing alike. The index measures the instrument that was built, not the governance that followed.
The scores are expert judgements, not measurements
Every score is assigned by reading primary documents against the published anchors. Nothing has been coded blind, and nothing has been coded twice, so no inter-coder reliability has been established. With twelve features and five frameworks, double-coding even two frameworks would materially strengthen the dependability claim the project makes for itself. The anchors exist so that a reader who disagrees can say exactly which anchor we applied wrongly.
Lukes' third face is not measured
The power axis captures decision-making and agenda-setting, the first two faces of power. It does not capture the third dimension — the shaping of what parties come to want. In AI governance that dimension is substantial: norm diffusion and third-country adoption of regulatory templates are exactly this kind of power, and the index currently has nowhere to record them.
External leverage is not measured
Under rule 4, resource and technical concentration measure how capability is distributed among the parties to a framework, not how much leverage its sponsor holds over outsiders. The Brussels effect — setting global rules through market access rather than through membership — is therefore invisible to the index as designed. This sits alongside the absence of Lukes' third face: both concern power exercised over those outside the room.
Efficiency is the axis that discriminates least
Across the coded sample the efficiency scores span 21 points, against 55 on power concentration and 47 on reversibility, and most of the spread rests on a single feature — binding implementation. Reversibility has the mirror problem: it separates binding from non-binding instruments sharply and then tells you little about how the non-binding ones differ from each other. Both are properties of the current design, not findings about the frameworks.
Version history
Index restructured from fifteen indicators to twelve features, four per axis. Power concentration rebuilt around the faces of power plus material and technical resources. Reversibility and efficiency restored as axis names, with 2.2 and 2.4 reverse-coded so that all four features in the reversibility axis count in the same direction.
The four coding rules were settled on 21 September 2026 and the anchors for 1.3, 1.4, 2.4, 3.2, 3.3 and 3.4 were tightened to match them — most substantially 2.4, whose top anchor had included a condition about adoption by non-parties that rule 1 excludes, and which would otherwise have been unreachable by construction.
Corrections
Where a score changes after publication, the change is recorded here with the date and the reason. Nothing is amended silently.
Nothing has been published as final, so nothing has been amended.