# Evaluation Covenant v0.1

**Evaluation Covenant · Phase 0 · Draft · Unsealed · Unregistered · No Run**

**State: Draft · Unsealed · Unregistered · No Run · No Result · Not Approved**

## Design the learning before the result exists.

Three synthetic covenants bind the UK Government Transformation Engine’s held dossiers. They make questions, contrasts, measures, guardrails, distribution, missingness, affected-party obligations, interpretation and publication rules explicit and receipt-visible—without collecting data, running an evaluation, selecting an option or creating authority.

**Truth statement:** Independent assisted design publication; exact text not separately human-reviewed. Not government, a Trial Chamber plan, registered protocol, ethics approval, analytical assurance, evaluation, result, advice, decision, adoption or permission. The system has no authority.

The publication is `DRAFT`, `UNSEALED`, `UNREGISTERED`, `NO-RUN`, `NO-RESULT`, and `NOT-APPROVED`. Each design repeats those states. There is no pass state, score, traffic light, progress state, winner, preferred option, or pipeline.

## Purpose

The Evaluation Covenant records in advance what a possible future evaluation would need to ask and what it must not infer. It is a design contract, not evidence that an intervention works. It contains no live intervention, evaluation registration, recruitment, data collection, participant record, result, estimate, amount, threshold, causal claim, economic conclusion, recommendation, adoption, deployment, or use authority.

Three principles govern the draft:

- **Questions before outcomes.** A desired outcome inherited from the Transformation Engine is context for a question, not a result.
- **Affected people before averages.** Missing groups and subgroup harm cannot be averaged away or replaced by a weighted-only view.
- **Nulls, harms and deviations stay publishable.** Outcome direction alone can never justify withholding.

## Frozen lineage

The sole admitted upstream design-context publication is the frozen UK Government Transformation Engine v0.1 snapshot dated `2026-08-21`:

- 83,915 bytes
- byte receipt `sha256:fef3ca1719cb7fd02c09c558db6f239cad29a173b921763a10a59e042a4cdf00`
- canonical receipt `sha256:7d9df6d5b6292b46abfbc076847f7e033b8d889be7a868dcbeac5d901b75c715`
- producer tag `https://github.com/cambridgetcg/kingdom-government-transformation-engine/tree/v0.1.0`

The exact Transformation schema and guide bytes are also pinned. The base Government Engine is transitive validation context only:

- 196,005 bytes
- byte receipt `sha256:1751b2ac331d6a1283cbf57b5b6961d9071f528298f97589fcbd7cf36407e139`
- canonical receipt `sha256:cc2c22e356cef0160ed5da1e2fd756bad5801f574c4cb064dec157b9207b9739`

Neither upstream publication is used as evidence of an intervention effect, estimand value, identification strategy, result, authority, or preferred option. Exact dossier, proposition, option, desired-outcome, assumption, negative-pathway, evaluation-need, and incomplete human-gate IDs are receipt-bound as synthetic design context only.

The Reality Influence research map from KINGDOM-OS PR #75 is untrusted related research. It is not imported, bound, projected, evidence or authority. J-space is not imported or depended on. Neither appears in the public snapshot, source ledger, lineage, theory, design authority, affected-party claims, or discovery summary.

## The three synthetic designs

### Four-nation information orientation

`design:four-nation-information-orientation` binds the held four-nation public-service-information dossier and all three of its options: business as usual, jurisdiction branches, and a comparison matrix. No option is selected, assigned, implemented, ranked, or preferred.

The process question concerns whether four territory branches and their holds could be found and comprehended. The impact question asks whether either alternative, compared with business as usual, could reduce unsupported UK-wide equivalence classification. It does not answer legislative or executive competence and does not establish equivalence. The economic question keeps future authoring, review, accessibility, translation, maintenance, user-burden, and distribution categories distinct; every amount is null.

England, Scotland, Wales, and Northern Ireland remain exact, separate distribution territories. Access needs, digital exclusion, language, and prior familiarity remain planned dimensions. Missing a territory stops interpretation.

### HMRC notice-route orientation

`design:hmrc-notice-route-orientation` binds the held synthetic HMRC notice-route dossier and its business-as-usual, layered-content, and route-map options. No option is selected, assigned, implemented, ranked, or preferred.

The process question concerns visible route separation and accessibility. The impact question covers explicit contrasts for both alternatives against business as usual across route-conflation and unsupported-deadline-assurance measures. It cannot identify a correct route, deadline, eligibility, liability, legal effect, or personal-case answer. The economic question preserves future authoring, maintenance, staff-burden, user-burden, distribution, and uncertainty categories, all without amounts or conclusions.

Any inference framed as a recommendation, deadline assurance, eligibility or liability determination, tax or legal advice, or handling of a personal case stops interpretation.

### Public-money category orientation

`design:public-money-category-orientation` binds the held public-money transparency dossier and its business-as-usual, glossary, and linked-map options. No option is selected, assigned, implemented, ranked, or preferred.

The process question concerns category and hold findability. The impact question covers explicit contrasts for both alternatives against business as usual across category-confusion and labelled-pound-inference measures. It creates no authority chain, spending authority, transaction, trace, or labelled-pound result. The economic question contains only future authoring, maintenance, source, user-burden, distribution, uncertainty, and unmonetisable benefit/harm categories; all amounts and conclusions remain null.

Plan is not estimate. Estimate is not appropriation. Appropriation is not credit or issue. Issue is not outturn. Payment or settlement is not liability or spending power. A representation cannot establish a labelled-pound trace.

## Questions, theory and contrasts

Every design has exactly one process, one impact, and one economic question. Each is `NOT-ANSWERED`, has `result:null`, and has no authority effect.

Theory bindings contain only exact upstream IDs and directions. They do not copy an upstream narrative and do not claim causal truth, testing, or validation. Monitoring is not evaluation; output is not outcome; change is not attribution.

Each design contains exactly one business-as-usual comparator and two candidate alternatives, one for each Transformation option. Every comparator is unselected, unassigned, unimplemented, and unexposed. For each candidate alternative and every outcome-role measure named by the impact question, an explicit estimand is specified against business as usual. Identification is `NOT-ASSESSED`; unit, horizon, and interference remain unassessed or undesigned; estimate and uncertainty are null; causal status is `NOT-CLAIMED`.

## Measures, guardrails and Goodhart risk

Measures are separated into process, outcome, economic, and qualitative roles. Every measure is used by at least one compatible question. No measure is a target or total/composite score. Every value is null, sources are empty, numerator and denominator are null, and definitions remain explicitly unsealed (`definitionFrozen:false`). The dated draft makes definitions receipt-visible; it does not formally freeze, seal, or register them.

Every outcome measure names a gaming hazard and at least one guardrail. Every bound negative pathway is covered by a guardrail. Guardrail thresholds and results are null, breach states are `NOT-OBSERVED`, and guardrails stop interpretation. Changing a denominator, population, window, or definition changes the design receipt.

## Distribution and missingness

Distribution plans name affected-party classes and planned dimensions for access needs, digital exclusion, language, and prior familiarity. Protected-characteristic and intersectional analyses are not designed. Aggregate results cannot override harm. Weighted results cannot replace unweighted results. Small cells are `SUPPRESS-NOT-INFER`. No distribution result exists.

Expected missingness kinds are non-response, attrition, not-covered, suppressed-for-privacy, measurement failure, and unknown. Missing is not zero and is not no effect. Complete-case analysis is not a default. Imputation and sensitivity analysis are not designed. Failure to observe a critical class stops interpretation. No missingness result exists.

## Process, impact and economic separation

Process evaluation is limited to implementation, fidelity, reach, accessibility, and comprehension. It cannot establish impact or causality.

Impact evaluation contains candidate methods, but none is selected. Comparator and assignment feasibility are unknown. Spillover, interference, and identification assumptions are not assessed. It cannot claim attribution.

Economic evaluation separates resource use, user burden, maintenance, distribution, uncertainty, and unmonetisable benefit and harm. Every amount, benefit-cost ratio, net present value, and result is null. It cannot establish value for money or spending authority. Impact is not value for money.

## Data, privacy, ethics and affected parties

No dataset is referenced or acquired. Administrative, monitoring, survey, qualitative, collection, and linkage states are false. Retention is not designed. Any future data proposal requires a new major contract and separate authority.

The draft contains no personal, special-category, or criminal-offence data and no human participants. Lawful basis and purpose compatibility are not assessed. A DPIA is not started. Ethics review is required but not completed. Consent is not applicable to Phase 0. Reidentification risk is not assessed. Raw feedback and public microdata are false.

Affected parties are role classes, never named people, handles, contacts, case identifiers, testimony, or protected values. Each class is mapped-not-engaged; accessibility, language, burden, and compensation remain unassessed; dissent is preserved; personal records are false. Participation is not representation, consent, or endorsement. Silence is not acceptance.

## Quality assurance and human gates

Quality assurance is not performed. Future analyst and independent assurer roles are distinct and unassigned. Reproducibility planning is not designed; peer review and sign-off are not performed. QA, peer review, registration, sealing, or a completed review cannot by themselves establish truth, approval, adoption, or authority.

Each design has six required, incomplete, outside-system human gates: independent design review; analytical QA by an assurer distinct from the analyst; privacy, ethics, and rights review; affected-party, accessibility, and distribution review; and the two exact inherited Transformation domain gates. No gate is completed by v0.1.

## Trial Chamber separation

The Trial Chamber asks how a bounded artifact might someday be tested. The Evaluation Covenant asks what learning about a synthetic intervention would mean, for whom, under which contrast and publication obligations. Neither runs.

The Covenant duplicates no Trial Chamber plan. `trialChamberPlanRefs` is exactly empty and compatibility is `not-applicable`. Referrals, evaluation requests, evaluation plans, fixtures, criterion rules, resource envelopes, passing effects, preflight, runners, artifacts, artifact pins, execution authority, trial authority, and trial runs are forbidden fields.

## Stops, corrections and deviations

Stop rules are design obligations, not automatic controls. Each is unobserved, non-automatic, owned outside the system, and creates no authority effect. A future human process would have to assess correction, pause, withdrawal, or republication under separate authority.

Deviations are exactly empty. A future post-seal amendment model would append—not overwrite—the target receipt, changed fields, timing, reason, author and assurer roles, and a classification such as prespecified, amended before data, or post-hoc exploratory. That model is not activated in Phase 0.

## Nulls and outcome-neutral publication

No result is not a null result. A null result is not no effect, failure, success, or equivalence. An equivalence claim requires a prespecified margin and a suitable design. Inconclusive status must be preserved and interpretation requires human review.

Null, negative and inconclusive findings remain in publication scope; outcome direction alone can never justify withholding. Selective outcome reporting is forbidden by the design. Any future restriction must be separately justified, minimised and reviewed under privacy, data protection, rights, licensing, legal, confidentiality, safety or security obligations.

Allowed grounds are finite: privacy and data protection, rights and licensing, legal restriction, security, confidentiality, and risk of harm. A future reason and scope must be recorded; minimum necessary withholding applies; redaction, aggregation, or partial publication must be considered. Restricted material cannot be replaced by a misleading positive summary. Decision authority remains outside the system. Phase 0 has no raw data, result, publication decision, or withholding decision.

## Method source ledger

The public snapshot contains exactly ten official primary method or governance links:

1. Magenta Book: Central Government guidance on evaluation — HM Treasury and Evaluation Task Force.
2. Transparency in Government Evaluation Research (TIGER) — HM Treasury and Evaluation Task Force.
3. Quality in policy impact evaluation (QPIE) — HM Treasury and Evaluation Task Force.
4. The Green Book (2026) — HM Treasury.
5. The AQuA Book — Government Analysis Function.
6. Government Social Research Publication Protocol — Government Social Research Profession.
7. Data and AI Ethics Framework — Government Digital Service.
8. Consultation Principles 2018 — Cabinet Office.
9. Guidance on using the Evaluation Registry — Evaluation Task Force, Cabinet Office and HM Treasury.
10. Data protection by design and by default — Information Commissioner’s Office.

These are link-only orientation. No source body, body hash, quoted text, or reuse licence is retained. A URL observation and `checkedOn` date do not prove exact-text review, currentness, legal effect, endorsement, licence, or rights. Method guidance may orient how an estimand is specified; it supplies no empirical estimand value, identification result, observed dossier outcome, intervention-result finding, or authority.

## Deterministic validation

The zero-dependency Node.js producer requires Node 20 or later. It uses strict fatal UTF-8 decoding, duplicate-key rejection, safe-number and structural bounds, symlink rejection, null-prototype canonicalisation, exact input receipts, one closed Draft 2020-12 schema, semantic cross-reference gates, forbidden Trial-field and personal-record-shape scans, and an atomic fixed-path writer.

```sh
npm test
npm run validate
npm run build
npm run build:check
npm run check
npm run overview
```

Only `build` writes, and only to `public/evaluation-covenants.json`. The producer does not fetch, serve, submit, deploy, acquire data, read a clock, invoke an LLM, expose MCP, or accept runtime/free-text input.

## Publication and freshness

- Human route: `https://thekingdom.dev/evaluation-covenant/`
- Mutable alias: `https://thekingdom.dev/evaluation-covenants.json`
- Immutable dated snapshot: `https://thekingdom.dev/evaluation/uk-government-transformation-evaluation-covenants-2026-08-23.json`
- Schema: `https://thekingdom.dev/evaluation/schemas/v0.1/public-snapshot.schema.json`
- Guide: `https://thekingdom.dev/EVALUATION-COVENANT-v0.1.md`
- Intended tag: `https://github.com/cambridgetcg/kingdom-evaluation-covenant/tree/v0.1.0`

The next review is due `2026-09-21`. After that deadline the immutable bytes are historical design only. The mutable alias must be replaced or retired under a separate release gate. The reader’s clock is not embedded and no runtime enforcement claim is made.

## Rights

The producer is `UNLICENSED`. See `NOTICE.md`. Official links do not imply permission to copy their contents, and no rights or licence are inferred from a URL, publisher name, or public availability.
