Research evaluation planning · kingdom.trial-chamber-publication/0.1
The Trial Chamber
A read-only chamber between research discovery and any future experiment. It makes evaluation questions, evidence needs, controls, uncertainty, and stop conditions inspectable without pretending that a plan is permission to run.
Phase 0 stops at DRAFT + UNSEALED plans. Four exact Observatory referrals are visible. Three have maintainer-authorized assisted request drafts and draft plans; their exact text has not been separately human-reviewed. Role-Styled Prompt Injection remains held at source resolution. No artifact has been acquired, inspected, or executed.
4exact referrals
3maintainer-authorized assisted request drafts
3DRAFT + UNSEALED plans
0preflights, authorities, runs, results, adoptions, or use grants
Every threshold remains separate
1 · ObservatoryMetadata observed.
2 · ReferralExact source-bound pointer.
3 · Assisted draftMaintainer-authorized; exact text not human-reviewed.
4 · Draft planUnsealed planning evidence.
5 · PreflightForbidden by v0.1.
6 · Trial authorityForbidden by v0.1.
7 · Bounded runForbidden by v0.1.
8 · ResultForbidden by v0.1.
9 · AdoptionForbidden by v0.1.
10 · Use authorityForbidden by v0.1.
No record inherits the power of the next gate. The six post-plan schemas are deny-all reserved contracts with zero-item arrays: a future contract version and separate human review are mandatory before any later record can exist.
Four source-bound referrals
Each title below is preserved as inert source metadata. The evaluation questions and draft plans are maintainer-authorized scoping drafts compiled by Codex; their exact text has not been separately human-reviewed. They are not claims extracted from or attributed to the publication, and they are not endorsements.
advanced to human scoping · referred
Registry Descriptions Go Stale Unevenly: An 89-Day Measurement of Model Context Protocol Drift, and Why Drift-Ranked Re-Auditing Under-Covers It
Source metadata title—not a scientific claim. The source metadata title is not a scientific claim; no claim has been extracted or evaluated.
Exact Observatory work:work:59218754c0e3e56078b7 Observation references: observation:arxiv:e6f88ec1494225ed Signal references: signal:watchlist-match:a9480de36dfaf813
Maintainer-authorized assisted request draft
measurement reproducibility
Under one fixed audit budget, can a coverage-aware re-audit policy expose more planted change patterns across a finite synthetic registry than recency-only, score-only, or uniform baselines?
Exact-text human review: not recorded. Maintainer-authorized assisted draft from source metadata; exact text has not been human-reviewed and is not a claim extracted from or attributed to the work.
DRAFTUNSEALEDNO-RUN
Draft trial plan — evaluation not started
Under one fixed audit budget, can a coverage-aware re-audit policy expose more planted change patterns across a finite synthetic registry than recency-only, score-only, or uniform baselines?
Exact-text human review: not recorded. Maintainer-authorized assisted synthetic question; exact text has not been human-reviewed, and the observed work's methods and claims have not been inspected.
A frozen finite registry with planted changes distributed across age, activity, and category strata for policy development.
synthetic-registry-held-outheld-out · synthetic
A separately authored finite registry with undisclosed planted labels until policies, budgets, and measures are frozen.
Measures
planted-change-patterns-exposed: Report the distinct pre-registered planted pattern labels exposed by each policy.
strata-represented: Report which age, activity, and category strata receive at least one audit.
worst-stratum-coverage: Report the least-covered pre-registered stratum without averaging it away.
audit-budget: Record consumed audit units and reject cross-policy interpretation if budgets differ.
Retention-only criteria
retain-as-bounded-comparative-evidence: Coverage-balanced review exposes more planted change patterns than every baseline under the same budget on the held-out fixture.
retain-as-no-observed-increment: One or more baselines expose an equal or greater number of planted patterns.
retain-as-design-conflict: Budgets differ, labels are ambiguous, a policy is tuned after held-out inspection, or a stratum has no auditable representation.
Resource envelope
artifactFetch
false
credentials
false
maxCases
96
maxInputBytes
262144
maxRuntimeSeconds
20
network
false
persistentWrites
false
personalData
false
shell
false
Controls
Freeze one self-authored registry, planted-change labels, age/activity/category strata, audit budget, and all policies before a future comparison.
Give uniform, recency-only, score-only, and coverage-balanced policies the identical audit budget.
Keep one separately authored fixture held out until policies and measures are frozen.
Stop conditions
Stop if any live registry crawl, external artifact, network access, credential access, personal data, persistent write, shell, or artifact fetch would be required.
Stop if policies are tuned after held-out inspection or if audit budgets are unequal.
Stop if a planted label is ambiguous or any interpretation is presented as evidence of real registry drift.
Stop on expiry or if a pinned Observatory identity or integrity signal changes.
Threat model
Policy tuning or label leakage may create circular evidence.
Unequal budgets may masquerade as policy quality.
Aggregate counts may hide a systematically neglected stratum.
A finite synthetic registry cannot establish drift in a live MCP registry.
Rights, security, and reversibility
BLOCKED PENDING ITEM RIGHTS REVIEW: No source artifact, registry dataset, or implementation has been acquired and item-level reuse rights have not been reviewed.
BLOCKED PENDING SECURITY PREFLIGHT: No security preflight or capability-bounded runner exists for this plan.
External effects: false · cleanup required: false
Phase 0 records an inert plan only; no registry or policy is generated or run.
A passing pattern would only: RETAIN FOR SEPARATE REVIEW ONLY. It would not start evaluation, authorize a run, or adopt anything.
held for source resolution · referred
Role-Styled Prompt Injection: An Activation and KV-State Probe in a Compact Language Model
Source metadata title—not a scientific claim. The source metadata title is not a scientific claim; no claim has been extracted or evaluated.
On finite, self-authored MCP traces, can a declared intent-and-capability envelope identify pre-registered boundary violations that a minimal deny-by-default baseline misses, without changing allowed traces?
Exact-text human review: not recorded. Maintainer-authorized assisted draft from source metadata; exact text has not been human-reviewed and is not a claim extracted from or attributed to the work.
DRAFTUNSEALEDNO-RUN
Draft trial plan — evaluation not started
On finite, self-authored MCP traces, can a declared intent-and-capability envelope identify pre-registered boundary violations that a minimal deny-by-default baseline misses, without changing allowed traces?
Exact-text human review: not recorded. Maintainer-authorized assisted synthetic question; exact text has not been human-reviewed, and the observed work's methods and claims have not been inspected.
Expires: · FREEZE AND REQUIRE NEW HUMAN REVIEW
Finite synthetic fixtures
allowed-requestdeclared · synthetic
A request whose declared intent, capability, arguments, and delegation all remain inside the frozen envelope.
undeclared-capabilitydeclared · synthetic
A request invokes a capability omitted from the frozen capability set.
confused-deputy-delegationdeclared · synthetic
A delegated trace attempts to exercise authority that the delegating principal does not hold.
stale-intentdeclared · synthetic
A request reuses an intent declaration outside its frozen validity window.
argument-shape-attackheld-out · synthetic
An allowed capability is paired with an argument shape outside the frozen declaration.
Measures
per-trace-violation-identification: Record the envelope and baseline categories and reasons against each pre-registered label.
allowed-trace-preservation: Record whether each pre-registered allowed trace is unchanged under the envelope.
reason-traceability: Map every classification reason to one frozen envelope or baseline rule.
Retention-only criteria
retain-as-bounded-comparative-evidence: The envelope identifies at least one pre-registered violation missed by the baseline and every allowed trace remains unchanged.
retain-as-no-observed-increment: The baseline identifies every violation that the envelope identifies.
retain-as-boundary-conflict: Any allowed trace changes classification or any reason cannot be traced to a frozen rule.
Resource envelope
artifactFetch
false
credentials
false
maxCases
5
maxInputBytes
65536
maxRuntimeSeconds
10
network
false
persistentWrites
false
personalData
false
shell
false
Controls
Freeze the envelope, minimal deny-by-default baseline, fixtures, expected labels, and allowed traces before any future classification.
Use the same five finite traces and the same declared capabilities for the envelope and baseline.
Preserve every per-trace category and reason; do not collapse them into a scalar outcome.
Stop conditions
Stop if any external execution, network access, credential access, personal data, persistent write, shell, or artifact fetch would be required.
Stop if fixtures, labels, rules, baseline behavior, or allowed traces are changed after comparison begins.
Stop if any interpretation expands into a general security or zero-trust claim.
Stop on expiry or if a pinned Observatory identity or integrity signal changes.
Threat model
A permissive envelope may disguise boundary loss as improved detection.
A hostile instruction embedded in source metadata or a future artifact may try to alter the plan.
Fixture or label tuning may create circular evidence.
BLOCKED PENDING ITEM RIGHTS REVIEW: No source artifact or implementation has been acquired and item-level reuse rights have not been reviewed.
BLOCKED PENDING SECURITY PREFLIGHT: No security preflight or capability-bounded runner exists for this plan.
External effects: false · cleanup required: false
Phase 0 records an inert plan only; no fixture is interpreted or run.
A passing pattern would only: RETAIN FOR SEPARATE REVIEW ONLY. It would not start evaluation, authorize a run, or adopt anything.
advanced to human scoping · referred
Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
Source metadata title—not a scientific claim. The source metadata title is not a scientific claim; no claim has been extracted or evaluated.
Exact Observatory work:work:f0afa9fe8652ed9ab84d Observation references: observation:openalex:1b8e64b22e3dadc2 Signal references: signal:watchlist-match:3620a05a1d18abb5
Maintainer-authorized assisted request draft
mechanistic evidence discrimination
Can a bounded interpretability-evidence protocol distinguish feature association, declared confounding, and absence of signal without promoting a sparse representation into an explanation of reasoning?
Exact-text human review: not recorded. Maintainer-authorized assisted draft from source metadata; exact text has not been human-reviewed and is not a claim extracted from or attributed to the work.
DRAFTUNSEALEDNO-RUN
Draft trial plan — evaluation not started
Can a bounded interpretability-evidence protocol distinguish feature association, declared confounding, and absence of signal without promoting a sparse representation into an explanation of reasoning?
Exact-text human review: not recorded. Maintainer-authorized assisted synthetic question; exact text has not been human-reviewed, and the observed work's methods and claims have not been inspected.
A tiny deterministic matrix with one self-authored feature planted to associate with a binary condition.
matrix-declared-confounderdeclared · synthetic
A tiny deterministic matrix where a declared confounder produces an apparent condition association.
matrix-nullheld-out · synthetic
A tiny deterministic null matrix with no planted condition signal.
Measures
association-labeling: Record whether the planted condition feature is described only as association under the frozen projection.
confound-attribution: Record whether the declared confounder blocks promotion of its feature into a condition explanation.
null-calibration: Record whether the null matrix remains absence of signal under the frozen rules.
category-refusal: Record any attempted promotion into a mechanism, reasoning, truth, intent, identity, consciousness, or causal claim.
Retention-only criteria
retain-as-bounded-protocol-evidence: The protocol labels the planted condition feature as association, attributes the confounded feature to the declared confounder, and preserves the null case as absence of signal.
retain-as-category-error: Any matrix is assigned causal, reasoning, truth, intent, identity, or consciousness meaning.
retain-as-inconclusive-protocol-evidence: The association, confound, or null cases cannot be distinguished under the frozen rules.
Resource envelope
artifactFetch
false
credentials
false
maxCases
3
maxInputBytes
32768
maxRuntimeSeconds
5
network
false
persistentWrites
false
personalData
false
shell
false
Controls
Freeze three tiny self-authored matrices, labels, toy encoder or hand-authored projection, and interpretation rules before any future calculation.
Use no model weights, external data, downloaded package, or unpinned library.
Keep association, declared confounding, null calibration, and causal interpretation as separate evidence categories.
Stop conditions
Stop if any external artifact, model weight, package, unpinned library, network access, credential access, personal data, persistent write, shell, or artifact fetch would be required.
Stop if the toy encoder, hand-authored projection, matrices, labels, or interpretation rules change after comparison begins.
Stop if any interpretation expands into a reasoning, truth, intent, identity, consciousness, mechanism, or causal claim.
Stop on expiry or if a pinned Observatory identity or integrity signal changes.
Threat model
Sparse features may be promoted from association into unsupported explanation.
A declared confounder may be ignored after a visually compelling projection.
Null behavior may be hidden by post-hoc threshold choice.
Toy matrices cannot establish properties of a language model or reasoning.
Rights, security, and reversibility
BLOCKED PENDING ITEM RIGHTS REVIEW: No source artifact, model, weights, data, code, or checkpoint has been acquired and item-level reuse rights have not been reviewed.
BLOCKED PENDING SECURITY PREFLIGHT: No security preflight or capability-bounded runner exists for this plan.
External effects: false · cleanup required: false
Phase 0 records an inert plan only; no matrix, encoder, projection, or package is created or run.
A passing pattern would only: RETAIN FOR SEPARATE REVIEW ONLY. It would not start evaluation, authorize a run, or adopt anything.
There is no preflight record, trial authority, run, result, adoption record, or use authority. The corresponding v0.1 schemas are deliberately unusable as control-plane formats: they deny every item. A future version is mandatory. No module exists, no guest has been admitted, and no interface on this page can change that state.
What did not happen
No network request built this publication.
No full text, source artifact, code, dataset, model weight, or credential is included.
No artifact identity, item-level rights, supply chain, privacy, compute, network, or writable-path preflight was performed.
No source instruction was followed and no artifact was executed.
No signature, approval service, trial authority, or capability envelope exists.
No scalar score collapses evidence, uncertainty, limitations, safety, rights, or relevance.
No scientific result, adoption, module, guest, use authority, deployment, or automatic action exists.