← The Rulebook Part XVII

Part XVII, The Scorecard (Self-Review)

This is the "self review" the model demands of itself. Per §0.4, the whole system is scored against the published rubric, honestly, including weaknesses and open questions. A design that scores itself 10/10 on every count is lying (§0.6.5), and lying is a failure mode. This scorecard is candid: it states what the design earns at the design stage, what can only be earned empirically once running, and what it commits to measuring forever. It now covers the complete model, Parts 0–XIX; it applies the numeric cap of §0.4; and it is published alongside an independent score (§0.4.1) rather than resting on the authors' self-grade.

XVII.1 Method and an important distinction

We score against the twelve weighted criteria of §0.4. Two kinds of score exist:

Honesty requires saying: a brand-new design cannot claim a verified 10/10, because several criteria are empirical. What it can claim is a design built to reach 10/10 that commits to being measured against this very rubric, in public, forever.

And it does not mark its own homework alone. Per §0.4.1, this design-stage self-score is published alongside an independent score — a rotating external-expert panel + a sortition citizens' jury + a published red-team — and alongside the score's sensitivity to alternative weightings. The honest headline is therefore "here is our self-score, here is the independent score, here is how both move under different weights, and here is the live measured score once it runs", never a single self-asserted number. "Critically low" is now a numeric rule: any criterion below 5/10 caps the composite at that value (§0.4). Two criteria are marked "unearned until run" (‡): they can be designed well but only earned in operation, and the honest reader should discount them accordingly.

XVII.2 The design-stage scorecard (full model)

#CriterionWtScore /10Rationale (incl. honest limits)
1Legitimacy & consent ‡12%9STV+ + sortition + bounded referenda, now disaggregated at ratification with a founding bar ≥ amendment bar (§XV.2a); the individual franchise entrenched in the core (§I.9.2); informed consent defined by process, no retrospective voiding (§0.2); resident path + diaspora taper (§XIII). Design 9; unearned until real participation is measured — cannot be 10 on paper.
2Outcome quality12%8Strong mechanism: bounded expertise + evidence standards + measurement + correction (IV-VI). But outcome quality is empirical, provable only once run. Honest 8 by design.
3Rights protection12%9Codified, entrenched, justiciable; tiered Class A with intra-tier ordering (torture the top lock) + legitimate-distinction test (§I.3.2); franchise + free-expression/press/information floor in the core; surveillance notification + any-person scope + bulk-data warrants; measurable social minimum with a mandatory-order remedy; children/family/disability (§I.4a); ECHR incorporate-and-exceed (§I.7a). Residual: rests on Court integrity — now super-entrenched against packing (§I.8.5, §IX.4).
4Accountability10%9Four-master accountability, named owners, removability, radical transparency; war powers to the legislature; independent fiscal authorities; intelligence oversight (IV-VI, X, XII).
5Capture & corruption resistance12%9Defence-in-depth across every branch, now hardened where it was weakest: the Integrity Assembly split so no body concentrates record + measure + audit + message (§VI.1); verifiable sortition closes the common-mode lot (§VI.3a); the Router made advisory, affirmed by a screened citizen panel of the Sortition Chamber (§XIX.3); the judiciary structurally super-entrenched against packing (§I.8.5); multi-party identity issuance (§VIII.2); the appointment criteria/briefing gate closed (§IV.4); guardian anti-collusion controls (§VI.1, XVI Scenario J); the emergency cumulative cap (§I.6.3a). Still not 10: no design fully stops a determined faction with a passive public (IDEA); slow elite convergence and the boundary rule-set stay residual (§XVI.6).
6Transparency & verifiability10%9Open-source mandate, transparency ledger, E2E-verifiable voting, open reasoning, machine-readable budget (VI, VIII, X); strengthened by pre-deployment algorithmic impact assessments and sovereign, inspectable governing AI (§V.6, §VIII.5). Residual: remote voting deliberately limited.
7Representation & proportionality8%9STV+ ends manufactured majorities; sortition adds a representative cross-section; the nations represented at the centre (III, IX, XI).
8Resilience8%9Crisis doctrine; deterministic reconstitution automaton + dispersed reserve verifiers (§VII.5); suppression-resistant anti-coup brake (§XIX.4); ledger BFT + omission detection (§VIII.4); funded security programme + trusting-trust defence + scoped formal verification (§VIII.6); honest paper-fallback with pre-positioned capacity (§VII.6). Caveat: a sophisticated digital state is itself attack/operating surface, mitigated by fallback, open audit, and modularity.
9Adaptability & self-correction6%9Review·Pause·Correct, sunsets, error register, continuous scoring, lawful amendment (V, VI, I).
10Intergenerational fairness5%9The Future-Generations mandate is now binding on irreversible/long-horizon decisions, proceeding against a negative assessment needs a published supermajority override (§IX.8), alongside the long-run-weighted objective, the debt rule, intergenerational accounting, the long-term fund, and an explicit environmental-stewardship duty (§I.4, 0.2, IX, X). Raised from 8: a brake the present must consciously release is a real mechanism, not an exhortation. Representing the unborn is still inherently imperfect.
11Simplicity & usability3%7Still the lowest score, and honestly so. Raised from 6 because the design now separates system complexity from user complexity: the citizen-facing surface (§II.0) is genuinely simple, rank a ballot, occasionally vote or deliberate, verify if you wish, sitting over a one-page Citizen's Charter (§I.1) and a consolidated, minimal institutional set (§IX.1). The machinery is still larger than FPTP's, so not a 9; but the demand on a citizen is smaller than today's.
12Inclusiveness2%9Offline-equivalent paths, accessibility, automatic registration, age-16 entry, universal (non-citizen) rights floor (II, XIII).

Weighted design-stage self-score: ≈ 8.8 / 10 (88/100), now computed over the whole model (Parts 0–XIX). The extensive hardening of this revision raises the mechanism quality of the capture-resistance, rights, resilience, and accountability criteria, while the honest ceiling is unchanged — because the criterion that caps it, outcome quality (‡), is provable only in operation and cannot be raised by writing more.

The number is presented honestly, not defended as precise. Under plausible alternative weightings — raising outcome-quality's weight toward its §0.2 primacy, or discounting the two ‡ criteria (legitimacy, outcomes) as unearned until run — the design-stage figure moves in a band of roughly 84–89. The authoritative figure is the independent panel's (§0.4.1), published alongside this self-score together with the full weight-sensitivity table; a self-grade is evidence of what the authors believe, not proof of the score. Only a rising live score in operation (§XVII.5) can carry the design toward a measured 10.

No criterion scores below the §0.4 numeric cap (any criterion < 5/10 would cap the whole composite at that value): the weakest, simplicity at 7, is a manageable trade-off honestly costed in docs/COSTING.md, not a fatal flaw such as "rights: capturable" would be.

For comparison, BIG's existing matrix scores the electoral system alone: STV+ 81/100, FPTP 39/100. The whole-of-governance design scoring ~88 at the design stage, before any empirical outcome credit, is the self-assessed headline — which the design does not ask anyone to take on trust (§XVII.2a).

XVII.2a The independent score — and where it disagrees with us

Per §0.4.1 the self-score above is published beside an independent score, not in place of one. That independent assessment has now been run as a high-fidelity proxy (docs/INDEPENDENT_SCORE.md): six expert-persona review panels (constitutional law, security & intelligence, cryptography, public finance, democratic theory, civil liberties) and a simulated demographically-stratified citizens' jury, each scoring the model from the primary text without sight of this scorecard, each red-teaming it. It is a proxy, not a substitute for the real external panel and real citizens' jury, which remain to be convened — but it is deliberately adversarial, and it disagrees with us in a consistent direction.

The independent proxy composite is ≈ 74/100, against this self-score of ≈ 88 — every criterion scored lower. The band is ≈ 72–75 under reweighting; individual panels ranged 65 (cryptography) to 83 (constitutional law), the technical panels and the lay jury scoring hardest. The sharpest result was the numeric cap: the panel initially put simplicity & usability at 4.9 — mean at the §0.4 cap line, three of seven reviewers scoring it 4 — which under a strict reading of the model's own cap rule would cap the composite near 5/10 (≈49/100), not 74.

That finding was then acted on, and re-checked (§0.4.1's feedback duty). A comprehension layer was added — the one-page whole-model map (§0.8, the architecture as 1·5·4·4·3), a counted sixteen-concept citizen budget with an item-by-item institutional minimality ledger (docs/COMPLEXITY.md), and a name-collision fix. The three panels that scored simplicity lowest re-assessed it on that new material: citizens' jury 4→6, democratic theory 5→7, constitutional law 6→7. Blended with the unchanged panels (the cryptography and security 4s concern operator/execution complexity the citizen map does not claim to touch), simplicity moves ≈ 4.9 → ≈ 5.6 — above the cap line — so the composite stands at ≈ 74 uncapped and the ≈49 cap risk is resolved. What we do not claim: none of the re-checked panels moved it past 7, and all three named the same residual — the map aids orientation, not comprehension; the operator/auditor machinery stays irreducibly dense. Simplicity is now on an honest footing, off the cap, not solved — the true state of a whole-of-government design.

The proxy's substantive critique was fed back into the design in the same pass (§0.4.1's whole point): the platform's electorate figure was corrected to the real ~53m franchise; the emergency domestic/external-threat classification — an escape hatch two panels found independently — was made independently refereed, published and challengeable (§I.6.3a); the fiscal engine was hardened (debt-rule conformance on continuity, an anti-serial-invocation bar on the escape clause); and the cryptographic prototypes' residuals (beacon last-revealer bias; the nullifier's reliance on the identity layer's zero-knowledge proof) are now documented rather than glossed. What the panel could not close, and we therefore still carry as residual, is the comprehension burden of a comprehensive model, the contestability of the value/means boundary, the rights that rest on institutional integrity, and — above all — that outcome quality and legitimacy are earned only in operation.

So the honest headline is a spread, not a point: self-assessed ≈ 88; independent proxy ≈ 74 (≈ 49 if the simplicity cap bites); real panel, jury, and live-operation score still to come. A self-grade is evidence of what the authors believe; the independent figure is the more skeptical and, per §0.4.1, the more authoritative of the two available today.

XVII.3 Honest weaknesses

  1. Outcome quality is unproven. The strongest mechanism is a hypothesis until it runs. Mitigated by piloting (Part XV) and measuring (Part VI), but unearned until produced.
  2. Appointment is the residual capture surface. Now harder, a sortition citizen-panel sits inside the process (membership unknowable in advance) and completed appointments are audited after the fact (§IV.4, §VI.3), but sustained, coordinated, long-horizon capture remains the hardest residual (§XVI.6). The evidence is clear that no mechanism closes it entirely.
  3. Remote voting is constrained because coercion-resistant, malware-proof internet voting at scale is unsolved (§VIII.3), a real limit on convenience, accepted to protect integrity.
  4. Complexity is the binding constraint — and the independent panel scores it at the cap. The model is comprehensive, so its machinery is more elaborate than what it replaces (Criterion 11). Mitigated by separating user-facing simplicity from system complexity (§II.0), the one-page Citizen's Charter (§I.1), and institutional consolidation (§IX.1) — the second completion pass consolidated further (a six-body integrity family to three, a standing convention to a convened procedure, a standing "boundary" panel to a walled materiality-affirmation function, four fiscal brakes to one). The independent proxy panel (§XVII.2a) initially scored simplicity at the §0.4 numeric-cap boundary (4.9), the single most consistent finding across every reviewer and the citizens' jury. That was answered directly: a comprehension layer — the one-page whole-model map (§0.8, 1·5·4·4·3), a counted sixteen-concept citizen budget with an item-by-item institutional minimality ledger (docs/COMPLEXITY.md), and a name-collision fix — moved the three lowest-scoring panels up (jury 4→6, democratic theory 5→7, constitutional law 6→7) and simplicity off the cap to ≈ 5.6. It is now on an honest footing, not solved: a comprehensive constitution is still harder to hold in mind than a single reform, the operator/auditor machinery stays irreducibly dense, and the map aids orientation rather than achieving comprehension. This is the honest ceiling — raised off the cap by genuine work, but a whole-of-government design has an inherent simplicity limit below 10, and pretending otherwise would itself breach §0.6.5.
  5. Transition is hard, slow, expensive, and can fail at the ballot (Part XV), and failing at the ballot is a legitimate outcome the design must accept.
  6. The militant-democracy balance is two-sided (§XIV.3): leaning toward freedom risks tolerating a genuine threat. The design accepts this lesser risk and relies on structure over bans.
  7. The ultimate backstop is the people themselves. No design survives a sustained popular movement to end it, nor should it (Axiom 1). Enforcement-of-last-resort ultimately rests on oaths, distributed loyalty, and an engaged citizenry (§XIV.4), which is a strength but also the final, irreducible dependency.
  8. The operating layer is now tested, and its residuals named. Scoring Parts XV–XIX (which an earlier scorecard did not cover) surfaces honest new residuals: the Router's rule-set can carry a subtle value tilt; the lot has participation/opt-out bias; guardian convergence is bounded not abolished; and the citizen levers trade responsiveness against obstruction (§XVI.6.7–9). Each is defended and measured, none is eliminated.

XVII.4 Design choices, all resolved in Part XVIII

The framework deliberately left the following choices open; Part XVIII (the Recommended Settlement) now resolves every one with a decisive recommendation that Referendum 2 ratifies. In summary, the items resolved are:

  1. The exact domain map of the Expert Layer (§IV.2) → ten domains (§XVIII D1).
  2. Head of state, reformed ceremonial monarchy vs elected non-executive presidency (§IX.3).
  3. Identity technology specifics, the precise cryptographic stack (§VIII.2), to be designed, audited, piloted.
  4. The wellbeing composite's exact indicators and weights (§0.2, §VI.7).
  5. District magnitudes and boundary criteria in detail (§III.5).
  6. The territorial structure's exact form, degree of federalism, the English tier's shape, and the second chamber's composition (sortition + territorial representation balance) (§XI.4, IX.2).
  7. Fiscal rule calibration, the specific debt path, escape-clause thresholds, and the long-term fund's rules (§X.4, X.6).
  8. Immigration policy is value-laden and left to democratic choice; only the process and rights floors are fixed (§XIII.4).
  9. Costing, a full, honest cost and capability plan for the transition (§XV.7).
  10. Rubric weights themselves (§0.4) are an editorial choice and should be put to public/expert deliberation.

XVII.5 The live score, a permanent commitment

Once implemented, the system scores itself continuously against this rubric (§VI.7) and publishes the result on the outcomes ledger. This is the design's central act of integrity:

XVII.6 Conclusion

This rulebook sets out a clean-slate, first-principles design for governing the United Kingdom: the people sovereign and consenting; their mandate formed honestly through a reformed electoral system; competent experts executing that mandate within bounded, reversible scope; a constitutionalised public-finance system that cannot loot the future; a codified, consented territorial settlement; the instruments of force bound to the constitution and the people; a civic, rights-bound definition of who belongs; an integrity system that polices everyone, including itself; resilience against crisis and coup; a verifiable technological substrate; a closed web of checks with no apex; a lawful, consensual path to get there; and a full adversarial defence.

It does not claim perfection. It claims something more defensible: a design built to the highest standard its authors could reach, honest about where it falls short, and committed to measuring itself against its own published rubric for as long as it runs, with the United Kingdom as the first country to put it into practice.

"Replace opinion with evidence. Mandate from the people, means from the experts, integrity from the institutions, verifiability from the technology, and reversibility always."

Part XVII ends. The rulebook continues with the resolved settlement (Part XVIII) and the operating model (Part XIX); this scorecard covers all of it (Parts 0–XIX), applies the numeric cap (§0.4), and is published alongside an independent score (§0.4.1). The design is rendered on the public site and enforced as code in /platform.