{
  "schemaVersion": "1.0",
  "page": "https://evidencepress.org/observatory/assurance/",
  "generated": "2026-08-04",
  "scoring": {
    "P": "Probability the project delivers its stated resolution criterion within its horizon, given a small human team directing capable AI agents that execute the mechanical share of the work at 2026 capability. Median of three prompt-isolated adversarial scoring passes (calibration-focused forecaster; sceptical methodologist; research-strategy realist focused on supersession). Disclosed limitation: all three passes ran on the same frontier-model substrate, so by the essay's own dependence argument their medians carry the weight of fewer than three independent judgements — structured expert judgement, not a calibrated forecast record.",
    "payoff": "Ordinal 0-1 judgement, conditional on delivery, of how much the result advances assurance infrastructure in the world where it lands, supersession-discounted: 0 = practice unchanged, 1 = a binding constraint removed across the field. An expert scale, not monetary expected value.",
    "doctrine": "Agents execute the mechanical share of every project now (corpus assembly, annotation at scale, statute encoding, proof search, simulation, implementation), with horizons and costs revised on that assumption; payoff is discounted for supersession, so measurements of current-model behaviour are valued below instruments, protocols, theorems, and standards that outlive model turnover.",
    "caveat": "Resolution criteria are drafts that pre-registration review should tighten. Two identified defects are recorded in the essay: the risk-limiting-audit project's drift sub-criterion (re-derived after review found it below an information-theoretic floor) and its sequential-efficiency margin (which must name its comparison class: sequential designs beat matched fixed two-point acceptance plans when true quality is inside the acceptable level; they cannot beat the zero-defect rule-of-three bound on its own worst-case ground).",
    "assessedOn": "2026-08-04",
    "reviewBy": "2026-11-04",
    "expiry": "P and payoff are judgements about the model landscape and institutional context of the assessment date. On the essay's own doctrine they have a short half-life: treat them as stale after the reviewBy date, and re-run the panel rather than reading them forward."
  },
  "instrumentPortfolio": [
    {
      "id": "P03",
      "title": "The curated-corpus attack: sampled assurance is unsound without cryptographic commitment",
      "description": "Formalise sampled auditing as a game in which the operator may re-generate, silently repair, reorder or selectively present items after learning the sample. Derive the maximum true error rate an operator can sustain while passing a 3/p audit as a function of post-hoc freedom, then build the attack end-to-end on a 50,000-item consultation-categorisation corpus with planted errors: a live demonstration in which roughly 300 clean sampled checks certify 'error rate below 1%' against a true rate an order of magnitude higher. Then implement the countermeasure — pre-published Merkle root over corpus and outputs, sample indices drawn from a public randomness beacon, inclusion proof served with every audited item — and show the attack defeated. Publish a two-page protocol addendum suitable for insertion into any government analytical quality or reproducible-pipeline guidance, plus a reference implementation. A short round of practitioner interviews with statistical-function staff precedes the build to establish whether real practice already commits implicitly; a loud null is an equally acceptable outcome.",
      "resolution_criterion": "By month 4, either: a working demonstration in which an operator passes a standard 3/p sampled audit (≥300 clean checks, 95% confidence, claimed p ≤ 0.01) on a ≥50,000-item corpus while the true error rate is ≥10× the certified bound, AND the same operator strategy fails under the commitment protocol at ≥99% detection across ≥100 trials; OR a documented negative result showing achievable undetected inflation is below 2× under realistically constrained operator freedom. Gates unchanged.",
      "horizon_months": 4,
      "cost_band": "~£35k — one human researcher at ~0.5 FTE for 4 months plus agent execution; compute under £2k",
      "crux": "Whether any real-world operator actually has post-hoc freedom over the audited set. If organisations already freeze outputs to an immutable store before quality assurance as routine practice, the attack is a paper tiger and the contribution collapses to a footnote. This is settled by interviews, which are calendar-bound and set the floor on the horizon.",
      "agent_execution": "Agents do the game formalisation and the derivation, build the 50,000-item planted-error corpus, implement the attack, the Merkle/beacon commitment protocol and the inclusion-proof server, and run the 100-trial evaluation. The human critical path is scheduling and conducting the practitioner interviews and carrying the protocol addendum into a standards or guidance process. Interview scheduling, not the build, is now the binding constraint.",
      "durability": "The most durable project in the portfolio. The result is a theorem about an adversary model plus a commitment protocol; neither depends on which model generated the corpus, and both remain correct through arbitrary model turnover. Nothing here is a measurement of a current model.",
      "judge_probs": [
        0.88,
        0.88,
        0.88
      ],
      "judge_payoffs": [
        0.68,
        0.62,
        0.68
      ],
      "p_med": 0.88,
      "pay_med": 0.68,
      "rationales": [
        "Disjunctive criterion (working attack OR documented negative), a four-month horizon, and a build that agents can complete in days: formalisation, a 50k planted-error corpus, Merkle commitment, beacon-derived indices, inclusion-proof server and a 100-trial evaluation. The only uncompressible element is scheduling practitioner interviews, and even those feed a crux that resolves either way rather than gating delivery. Payoff rises above its original 0.50 under the supersession doctrine precisely because nothing in it is a model measurement: a theorem about an adversary model and a commitment protocol remain correct through arbitrary model turnover, and it establishes a soundness precondition every other sampled-audit project silently assumes.",
        "Reference class: adversarial demonstrations of a known cryptographic principle (cut-and-choose without commitment), where the attack is essentially a theorem and the countermeasure works by construction — a Merkle root plus beacon-drawn indices defeats post-hoc curation with near-certainty, and agents can build the 50,000-item planted corpus, the attack and the 100-trial evaluation in weeks; the criterion is disjunctive, and the calendar-bound practitioner interviews inform relevance rather than gate resolution. Payoff is the portfolio's most durable — a result about an adversary model and a commitment protocol contains no measurement of any model — but discounted for intrinsic novelty, since election RLAs already require committed manifests and the contribution is transport plus a two-page guidance addendum. It is load-bearing precisely because it decides whether every other sampled-assurance project in the list is sound.",
        "Reference class: adversarial constructions against a sampling protocol with no commitment, which essentially always succeed — repairing the ~300 sampled items is trivial, and Merkle-plus-beacon defeating it is a theorem with an implementation. The disjunctive criterion means either branch resolves, and the interviews that set the calendar floor gate relevance rather than resolution, so agent execution collapses the build to weeks against a 4-month clock. Payoff is the portfolio's most durable: an adversary model plus a commitment protocol remains correct through arbitrary model turnover, and a two-page addendum inserted into analytical quality guidance is genuine infrastructure rather than a measurement."
      ],
      "essay_rank": 1,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P15",
      "title": "Replay feasibility audit: can any published AI-assisted government analysis be independently re-derived?",
      "description": "Identify published government or regulator analytical outputs that state AI assistance — sampling frames exist in national algorithmic-transparency registers, federal AI use-case inventories, evaluation registries, and departmental publications — and attempt independent third-party replay of a defined subset of their quantitative claims using only public materials. Score each on a graded ladder: sources identifiable, corpus reconstructible, model and version stated, parameters stated, deterministic replay possible, numbers reproduced within tolerance. The ladder, not the binary, is the product: which rung outputs fall off determines whether the fix is retention-and-hashing, a procurement field for model versioning, or something else. Then draft model contract clauses (replay, assurance receipt, third-party audit access), stress-test them against actual central procurement framework terms and standard AI services contracts, and cost retention obligations against data-protection constraints for personal data in consultation corpora.",
      "resolution_criterion": "By month 3: a published graded replay-feasibility score for ≥50 published AI-assisted analytical outputs drawn from ≥2 jurisdictions, adjudicating the pre-registered claim that fewer than 20% reach the 'numbers reproduced within tolerance' rung from public materials, with the per-rung failure distribution reported; plus drafted clauses reviewed by at least one commercial-legal practitioner for insertion into existing procurement vehicles. CHANGED: the sample rises from ≥20 outputs in one jurisdiction to ≥50 across two. Attempting replay was the entire cost, and it is now cheap; since the product is the rung distribution rather than the headline rate, a larger multi-jurisdiction sample makes that distribution properly readable and cross-nationally comparable.",
      "horizon_months": 3,
      "cost_band": "~£30k — one human researcher at ~0.5 FTE for 3 months plus ~£8k legal review; desk-based, needs no permissions",
      "crux": "Decision-relevance rather than mere embarrassment. Everyone expects near-zero replay; the project earns its keep only if the ladder localises the blocker, because corpus availability, model versioning and parameter disclosure imply very different interventions. A flat 'nothing replays' is cheap, near-certain, and tells a buyer nothing about what to change — a risk that grows, not shrinks, when the audit is cheap enough to run carelessly.",
      "agent_execution": "Agents do essentially the whole audit: harvest the sampling frames, reconstruct corpora where possible, attempt replay of every scored claim, and populate the ladder. The human critical path is the commercial-legal practitioner review of drafted clauses and the judgement calls on what counts as 'reproduced within tolerance'. This is the portfolio's largest compression, from 8 months to 3, because attempted replay is exactly the labour agents absorb.",
      "durability": "The rung ladder is a durable instrument and portable to any jurisdiction. The measured rung distribution is a snapshot of institutional practice, which decays on an institutional clock of years rather than a model clock of months — and decays only if procurement actually improves, which is the project's own objective. Cheap enough to re-run annually as a transparency indicator.",
      "judge_probs": [
        0.8,
        0.85,
        0.8
      ],
      "judge_payoffs": [
        0.55,
        0.6,
        0.5
      ],
      "p_med": 0.8,
      "pay_med": 0.55,
      "rationales": [
        "The portfolio's largest compression, 8 months to 3, because attempted replay is exactly the labour agents absorb: harvesting sampling frames, reconstructing corpora, attempting replay of every scored claim and populating the ladder across two jurisdictions. It is desk-based, needs no permissions, and the pre-registered claim adjudicates either way, leaving legal-practitioner review of the drafted clauses as the only human path; the residual risk is the raised bar itself, since registers list use cases and not always analytical outputs carrying replayable quantitative claims, so assembling 50 qualifying outputs is a real sampling-frame constraint. Payoff is moderate and mildly discounted: the rung ladder is a durable, portable instrument, but the measured distribution is a snapshot of institutional practice, and the crux is decision-relevance rather than delivery, a risk that grows when the audit is cheap enough to run carelessly.",
        "Reference class: desk-based transparency audits with published sampling frames, which deliver reliably and need no permissions — and attempted replay is exactly the labour agents absorb wholesale, which is why the sample could rise from 20 in one jurisdiction to 50 across two while the horizon fell from 8 months to 3; the only human critical path is one commercial-legal practitioner reviewing drafted clauses and the tolerance judgement calls, both easily accommodated. The pre-registered claim adjudicates either way. Payoff is capped by the project's own crux: everyone already expects near-zero replay, so value rests entirely on the rung distribution localising whether the fix is retention, hashing or a procurement versioning field — a risk that grows when the audit is cheap enough to run carelessly. The ladder is a durable, jurisdiction-portable instrument, cheap to re-run annually.",
        "Reference class: desk-based reproducibility audits, which almost always deliver their ladder — attempted replay is exactly the labour agents absorb, so raising the sample from 20 to 50 across two jurisdictions costs little and legal review of the clauses fits a 3-month clock. The residual risk is supply rather than execution: transparency registers and use-case inventories enumerate systems, not published analytical outputs carrying replayable quantitative claims, so assembling 50 qualifying outputs may require broadening the frame. Payoff is capped by the near-certainty of the headline direction — everyone expects nothing replays — leaving the rung distribution and the stress-tested procurement clauses as the real product; durable and annually re-runnable, but pointing mostly at interventions already suspected."
      ],
      "essay_rank": 2,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P11",
      "title": "A metamorphic relation library and mutation-adequacy benchmark for government analytical pipelines",
      "description": "Curate a public library of metamorphic relations covering the recurring shapes of policy analysis — stock-flow accounting, transition and duration models, weighted survey estimation, deflation and indexation, microsimulation, small-area estimation — with reference implementations wired into standard reproducible-analytical-pipeline templates so adoption costs an import, not a project. Then build the missing measurement: a mutation-adequacy benchmark. Compile a catalogue of realistic defects from documented analytical incidents (off-by-one period alignment, wrong deflator base, dropped subgroup weights, denominator mismatch, silent NA coercion, data-vintage slip, timezone boundaries), inject them blind into real open-source analytical pipelines, and measure the kill rate under four regimes: existing unit tests, deterministic replay, the metamorphic library, and agent code review.",
      "resolution_criterion": "By month 5: ≥40 relations published with working implementations; a catalogue of ≥60 realistic defect types applied blind to ≥6 real open analytical pipelines; and a published kill-rate table across all four regimes, adjudicating the pre-registered prediction that metamorphic relations kill at least 2× as many injected defects as deterministic replay (refuted if the ratio falls below 1.3). Gates unchanged — the counts were never the binding constraint, defect realism was — with one addition: the agent-code-review regime must be reported as a versioned column, stamped with model identifier and date, because it is the one column that decays.",
      "horizon_months": 5,
      "cost_band": "~£45k — one human researcher at ~0.5 FTE for 5 months directing agent implementation; negligible compute",
      "crux": "Whether realistic defect injection is achievable without rigging. Mutation benchmarks are gameable in both directions, and the only defence is blinding the injector from the library author and sourcing defects from documented incidents. Agent execution sharpens this: the injector and the library author should be separate model families under sealed briefs, or the blinding is nominal. Secondary risk: agent code review dominating all mechanical regimes, which would sit awkwardly with the correlated-error argument and needs independent replication before being believed.",
      "agent_execution": "Agents write the relations and their reference implementations, wire them into pipeline templates, build the defect catalogue, perform blind injection across the pipelines, and run all four evaluation regimes. The human critical path is sourcing defects from documented analytical incidents — which requires access to incident records held by analytical organisations — and enforcing the blinding discipline between injector and library author.",
      "durability": "The relation library and the mutation benchmark are durable instruments: metamorphic relations for stock-flow accounting or deflation are mathematics and will not expire. Three of the four kill-rate columns are properties of code rather than models and are stable. Only the agent-code-review column decays, and re-running that single column against a new model is a day's work — which is the argument for shipping the benchmark rather than the table.",
      "judge_probs": [
        0.72,
        0.68,
        0.65
      ],
      "judge_payoffs": [
        0.7,
        0.72,
        0.7
      ],
      "p_med": 0.68,
      "pay_med": 0.7,
      "rationales": [
        "The deliverables are counts and a table (40 relations, 60 defect types, 6 pipelines, a four-regime kill-rate comparison) with the headline ratio adjudicating in either direction, and agents write the relations, wire them into templates, perform blind injection and run all four regimes; the counts were never the binding constraint. Residual risk is defect realism and the integrity of blinding, which now requires the injector and library author to be separate model families under sealed briefs, plus human access to documented incident records. Payoff is durable because metamorphic relations for stock-flow accounting or deflation are mathematics and three of four kill-rate columns are properties of code rather than models; only the agent-code-review column decays, and re-running one column is a day's work, which is the argument for shipping the benchmark rather than the table.",
        "Reference class: mutation-adequacy benchmarking, which is well-trodden in software engineering; the headline prediction is close to structurally guaranteed because deterministic replay reproduces an injected logic defect faithfully every time and therefore kills near zero, so a ratio of 2× is far from the 1.3 refutation boundary and far from any dead zone. Agents write the 40 relations, wire them into pipeline templates, build the 60-defect catalogue and run all four regimes; the real risks are access to documented incident records held by analytical organisations and enforcing sealed-brief blinding between the injector and the library author, which must now be separate model families or the blinding is nominal. Payoff is durable instrument value — metamorphic relations for stock-flow accounting and deflation are mathematics, three of four kill-rate columns are properties of code not models, and adoption costs an import rather than a project.",
        "Reference class: mutation-adequacy benchmarks, which reliably ship their artefacts — the relation library, defect catalogue and four-regime kill table are all agent-executable, and both directions of the 2x prediction adjudicate. Two real risks remain: sourcing 60 defects from documented incident records held by analytical organisations is a human access path, and the 1.3-2.0 ratio band leaves the headline prediction unadjudicated; blinding injector from library author across separate model families is enforceable but must be designed in, not asserted. Payoff is high and mostly durable — metamorphic relations for stock-flow accounting and deflation are mathematics — with only the agent-code-review column decaying, and that column re-runs in a day, which is the argument for shipping the benchmark rather than the table."
      ],
      "essay_rank": 3,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P02",
      "title": "Risk-limiting audits for AI-generated evidence: sequential certification benchmarked against the rule of three",
      "description": "Port the SHANGRLA/ALPHA risk-limiting audit stack from elections to evidence products. Define assorters for the machine-checkable rungs of the verifiability stack (citation resolves; recomputed statistic matches; coding matches adjudicated label), implement betting-supermartingale sequential tests with a defined escalation branch, and benchmark against matched fixed-n rule-of-three designs on five corpora with adjudicated ground truth — a completed consultation, parliamentary or congressional record claim extraction, local land-use plans, public-body annual accounts, planning or permitting applications. A second workstream adds recertification under model drift: a conformal test martingale plus CUSUM monitor with designed average run length, tested against injected model-version regressions. Deliverables: an open implementation, a methods note in the register format used by government analytical quality guidance, and the operating-characteristic curves an analytical team needs to choose a design.",
      "resolution_criterion": "By month 6, pre-registered before data collection: (a) across five corpora and three tolerances (p = 1%, 0.3%, 0.1%), the sequential design certifies at risk limit α = 0.05 with a median ≥40% fewer human-adjudicated items than the matched fixed-n rule-of-three plan; (b) empirical risk-limit violation ≤5% over 10,000 simulation replicates per configuration; (c) the drift monitor detects an injected error-rate step from 0.5% to 2% within a pre-registered detection budget derived from the information-theoretic floor for the specified monitor at a false-alarm rate ≤1 per 10,000. CHANGED: the original criterion fixed that budget at 250 items, which independent calculation during scoring showed sits roughly threefold below the achievable floor; agent execution does not repeal an information bound, so the sub-goal is re-derived rather than dropped. Failure on any of (a)–(c) resolves negatively and is publishable as such.",
      "horizon_months": 6,
      "cost_band": "~£90k — one human statistician at ~0.6 FTE for 6 months directing agent implementation, of which ~£40k is adjudication; compute under £5k",
      "crux": "Whether item-level defects in policy evidence are adjudicable as bounded per-item quantities. If a material share of defects are contestable rather than binary, the assorter formalism fails and ground-truth adjudication costs more than sampling saves.",
      "agent_execution": "Agents implement the assorters, the betting-supermartingale tests and the escalation logic, assemble the five corpora, run the full simulation grid, generate the OC curves, and draft the methods note. The human critical path is adjudicating ground truth on the five corpora and re-deriving the drift criterion before pre-registration. Statistical design review stays human because the failure mode is a criterion that cannot be met by construction.",
      "durability": "Almost entirely durable: assorter definitions, the sequential test implementation, and the OC curves are mathematics plus code, indifferent to which model generated the evidence. Only the measured baseline defect rates of current pipelines decay, and those are inputs to the design rather than the deliverable.",
      "judge_probs": [
        0.58,
        0.62,
        0.6
      ],
      "judge_payoffs": [
        0.68,
        0.68,
        0.62
      ],
      "p_med": 0.6,
      "pay_med": 0.68,
      "rationales": [
        "Porting SHANGRLA/ALPHA is mature statistics with a strong prior of transfer, and the single largest reason the original scored 0.42 (a drift sub-goal set roughly threefold below the information-theoretic floor) has been repaired by re-derivation before pre-registration, which converts an unmeetable-by-construction gate into a real one. Agents write the assorters, the betting supermartingales, the 10,000-replicate grid and the OC curves; residual risk is the three-way conjunction holding at the tightest tolerance (p=0.1%, where sequential savings compress) and the adjudicability of policy defects as bounded per-item quantities. Almost nothing here decays: assorters, sequential tests and operating characteristics are mathematics plus code, so the payoff lands undiscounted, held below the top only because it is a port of known machinery rather than a new result.",
        "Reference class: ports of mature statistical machinery (SHANGRLA/ALPHA) into a new domain, which transfer reliably; agents absorb the assorters, betting supermartingales, escalation logic and the entire 10,000-replicate simulation grid, and the re-derivation of the drift criterion to the information-theoretic floor converts the original criterion's impossible sub-goal into an attainable one. Residual risk sits almost entirely in gate (a) — a median 40% item saving across five corpora and three tolerances holds when true defect rates are well below tolerance but collapses when they are not — and in whether policy defects are bounded per-item quantities at all. Payoff is close to fully durable: assorter definitions, sequential tests and OC curves are mathematics plus code that survive arbitrary model turnover, with only the input defect rates decaying.",
        "Two of three components are near-mathematical (risk-limit validity holds by construction; the drift budget is now re-derived to its own information floor so it is satisfiable by design), but component (a) is the real risk: the rule of three is already the exact zero-error binomial bound, so a median 40% saving depends on assorters being continuous comparison-style rather than binary pass/fail — precisely what the crux doubts, and agents cannot make a defect adjudicable if it is contestable. Payoff is high-durability: assorters, sequential tests and OC curves are mathematics plus code, indifferent to model turnover, though the contribution is a port of mature machinery rather than a new result."
      ],
      "essay_rank": 4,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P05",
      "title": "Prediction-powered consultation statistics: valid intervals from machine labels plus a small human sample",
      "description": "Apply prediction-powered and active inference to the actual published estimands of consultation analysis — share of respondents raising a theme, share opposing, subgroup breakdowns — combining machine labels on the full corpus with human labels on a small random subsample to produce intervals valid regardless of model bias. Deliver documented R and Python implementations, a methods note shaped for a national statistical or analytical function and its statistics regulator, and an empirical effective-sample-size multiplier on at least four completed public consultations spanning 10^3 to 10^5 responses; item-level response corpora are published by several sectoral regulators and, at very large scale, on the US regulations.gov docket system. Includes active sampling where the machine is uncertain, the finite-population correction, and explicit stress tests under the non-exchangeability consultations really have: organised campaigns, near-duplicates, late-window bursts.",
      "resolution_criterion": "By month 4: on ≥3 of 4 consultations, the prediction-powered estimator attains the same CI width for the primary estimand with ≥2× fewer human-labelled items than the human-only estimator, with coverage over 10,000 resampling replicates in [0.94, 0.96]; and the released package reproduces the published headline figures of a large deployed consultation-analysis exercise with valid intervals attached. STRENGTHENED: because agent execution makes the full pipeline cheap to re-run, the multiplier must additionally be reported as a curve across a ladder of at least three labeller qualities rather than a single number, so the result remains readable when labeller quality moves. Falling short of the 2× multiplier at the top of that ladder resolves negatively.",
      "horizon_months": 4,
      "cost_band": "~£45k — one human researcher at ~0.5 FTE plus ~0.2 FTE statistician for 4 months; human labelling budget and compute both under £8k",
      "crux": "Whether the published estimand is a population functional. If the analytical product is 'the set of distinct arguments raised' — a coverage estimand, not a mean — prediction-powered inference does not apply and the project must pivot to capture–recapture, a different and harder problem. Worth resolving in week one, before any implementation.",
      "agent_execution": "Agents implement the estimators in both languages, machine-label the full corpora, run active sampling, the finite-population correction, the non-exchangeability stress tests and the 10,000-replicate coverage simulations, and draft the methods note. The human critical path is the small random human-labelled subsample — which is small by design and is the whole point of the method — plus engagement with the statistical function and regulator to get the note recognised.",
      "durability": "The estimator, the package and the methods note are durable; they are correct regardless of labeller quality, which is the method's defining property. The measured effective-sample-size multiplier decays, but favourably: better labellers raise it, so a measured value stands as a lower bound rather than becoming wrong. This is the rare project whose decaying quantity decays in the useful direction.",
      "judge_probs": [
        0.6,
        0.68,
        0.6
      ],
      "judge_payoffs": [
        0.78,
        0.7,
        0.72
      ],
      "p_med": 0.6,
      "pay_med": 0.72,
      "rationales": [
        "Prediction-powered inference is off-the-shelf and its coverage guarantee is a theorem rather than an empirical hope, so the [0.94, 0.96] band and the 2x multiplier at the top of a labeller-quality ladder are both well inside what 2026-quality machine labels support; agents implement both languages, the active sampling, the FPC, the non-exchangeability stress tests and the 10,000-replicate study inside the four months. Residual risk concentrates in two places agents cannot touch: whether the published estimand is a population functional at all (a week-one crux that kills the project if it is coverage), and obtaining the corpus of a large deployed exercise to reproduce its headline figures. Durability is unusually favourable because the decaying quantity decays upward: better labellers raise the multiplier, so a measured value stands as a lower bound rather than becoming wrong, and the package and methods note are correct regardless.",
        "Reference class: applications of established prediction-powered/active inference machinery to a new estimand, which deliver reliably when machine labels are decent — and consultation theme coding at 2026 capability is well past the threshold where the 2× effective-sample multiplier is comfortably exceeded at the top of the labeller ladder; coverage in [0.94, 0.96] is a property of a correct implementation, and agents run the full corpora, active sampling and 10,000-replicate simulations cheaply. Residual risk is the week-one crux (if the published estimand is coverage of distinct arguments rather than a population share the method does not apply) and access to a large deployed exercise's underlying data for the reproduction sub-goal. Payoff is durable and, unusually, decays favourably: the estimator and methods note stay correct as labellers improve, so the measured multiplier stands as a lower bound rather than becoming wrong.",
        "Reference class: PPI/PPI++ efficiency demonstrations, which reliably clear 2x effective-sample multipliers when machine labels correlate strongly with human ones — plausible for binary theme coding at 2026 quality, and the strengthened labeller-quality ladder requires 2x only at the top rung, its easiest point. The residual risks are the two-sided coverage band, where finite-population correction and active sampling can push empirical coverage conservatively above 0.96 and technically fail, and the week-one crux that the published estimand may be coverage rather than a mean. Payoff is durable and, unusually, decays in the useful direction: the estimator is valid regardless of labeller quality, so a measured multiplier stands as a lower bound as models improve."
      ],
      "essay_rank": 5,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P07",
      "title": "The assurance coverage census: what fraction of real policy claims any mechanical oracle can touch",
      "description": "A stratified sample of roughly 60 real evidence products from at least two jurisdictions and a recent three-year window — impact assessments, statistical bulletins, written evidence to legislative committees, appraisal-guidance business cases, consultation-response analyses. Segment every assertive sentence into a seven-class obligation taxonomy (attribution, transcribed arithmetic, estimated statistic, rule mechanics, identification, model projection, framing/value) with two independent annotators — one human, one agent — and blind adjudication. Elicit decision weight by asking whether the product's recommendation would change if the claim were false. Then run the cheap oracles over the corpus: identifier resolution, hash identity against national web archives and point-in-time legislation XML, span containment, statcheck/GRIM-class recomputation, unit and stock-flow identity checks. Corpus, taxonomy, protocol, agreement statistics and oracle implementations released under an open licence. This is the project that prices every other project on the list.",
      "resolution_criterion": "By month 7: a public dataset of ≥50 products and ≥10,000 classified claims with Cohen's κ ≥ 0.7 between independent annotators on the seven-class label, plus per-class oracle catch rates with binomial CIs; and adjudication of the pre-registered prediction that mechanically dischargeable claims exceed 55% by count but fall below 45% by decision weight. Resolves NO if κ < 0.6 — the taxonomy cannot be applied reliably to policy prose at all. Gates unchanged, but the κ ≥ 0.7 gate is now materially more attainable: agents can iterate the annotation protocol across dozens of drafts on pilot text before the expensive human round begins, which is exactly the refinement that comparable discourse-annotation efforts could not afford on the first protocol.",
      "horizon_months": 7,
      "cost_band": "~£70k — one human researcher full-time for 7 months plus agent annotation and oracle-build budget; inference under £5k",
      "crux": "Whether decision weight is elicitable at all. If competent analysts systematically disagree about which claims are load-bearing, the count/weight distinction — the central quantitative claim — dissolves into an artefact of the elicitation protocol. Secondary: assertion segmentation stability in prose that fuses attribution, number and inference into one sentence.",
      "agent_execution": "Agents assemble the stratified corpus, serve as one of the two independent annotators, implement all six oracle families, and — critically — run many cheap protocol-iteration cycles on held-out pilot text before human annotation starts. The human critical path is the second independent annotator, blind adjudication of disagreements, and decision-weight elicitation from competent analysts, where the human judgement is the measurement rather than an input to it.",
      "durability": "Strongly durable. The taxonomy, protocol, open corpus and oracle implementations are reference infrastructure; the per-class oracle catch rates depend on the properties of source substrates and of policy prose, not on which model wrote the text, so they age on a decade clock rather than a model clock. The one model-dependent element is the agent annotator's contribution to κ, which is cheap to re-establish.",
      "judge_probs": [
        0.58,
        0.57,
        0.52
      ],
      "judge_payoffs": [
        0.84,
        0.8,
        0.82
      ],
      "p_med": 0.57,
      "pay_med": 0.82,
      "rationales": [
        "The binding gate is kappa >= 0.7 on a seven-class taxonomy over prose that fuses attribution, number and inference, and comparable discourse-annotation efforts land at 0.5-0.65 on a first protocol; the decisive change is that agents can iterate the protocol across dozens of drafts on held-out pilot text before the expensive human round, which is exactly the refinement those efforts could not afford, and the count/weight prediction adjudicates either way rather than gating. Human instrumentation remains for the second annotator, blind adjudication, and decision-weight elicitation, which is the crux and not compressible. Payoff is the portfolio's highest because the taxonomy, protocol, open corpus and oracle implementations are reference infrastructure that prices every other project, and per-class catch rates depend on source substrates and policy prose rather than on which model wrote the text, so they age on a decade clock.",
        "Reference class: multi-class discourse annotation over policy prose, which historically lands at kappa 0.5–0.65 on a first protocol — but that reference class is materially outdated here, because the reason protocols underperform is that iteration on pilot text was unaffordable, and agents can now run dozens of protocol revisions before the expensive human round begins, which is exactly the intervention that moves kappa. The remaining risk is real: decision-weight elicitation may not be stable across competent analysts, and assertion segmentation in prose fusing attribution, number and inference is genuinely hard. Payoff is near the top because taxonomy, protocol, open corpus and six oracle families are reference infrastructure whose catch rates depend on source substrates and policy prose rather than on which model wrote the text — they age on a decade clock — and because this project prices every other project on the list.",
        "Reference class: seven-class discourse annotation over fused prose, which typically lands at kappa 0.5-0.65 on a first protocol — agent-run protocol iteration across dozens of drafts on pilot text is a real and unusual lift on exactly the hard subproblem, which is why I put this above a coin flip rather than at the essay's 0.48. But the gate is a human-versus-agent kappa, not human-human, and the agent annotator is tuned by the very iteration that makes the gate attainable, so a passing number carries less validity than it appears to; decision-weight elicitation remains genuinely human-instrument work. Payoff is near the top of the portfolio and strongly durable: the taxonomy, protocol, open corpus and oracle catch rates are properties of policy prose and source substrates, ageing on a decade clock, and this is the project that prices every other one."
      ],
      "essay_rank": 6,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04",
      "editorial_note": "One judge rationale below references a 0.48 figure from the pre-doctrine scoring vintage; the published essay reflects the current 0.57 median. Prior-vintage scores will appear in a history field at the next panel re-run (reviewBy date)."
    },
    {
      "id": "P10",
      "title": "Exhaustive rule-mechanics verification: SMT over published tax-benefit claims",
      "description": "Assemble published claims about rule mechanics from impact assessments, ministerial and agency statements, and think-tank analyses covering a small number of recent reforms, and give each a formal statement over a declared household space. Encode the rules against an executable reference microsimulation from the OpenFisca or PolicyEngine families, in Catala-style literate form pinning each block to a cited statutory provision — so the encoding, which is the actual trust boundary, is reviewable by a lawyer and not only by a programmer. Run an SMT and bounded-model-checking harness that returns, per claim, a proof over the stated domain or a concrete counterexample household. Publish the encoding, the harness and the verdict table. The design ports to any jurisdiction with an executable rules model and a published statute.",
      "resolution_criterion": "By month 7: ≥60 published quantitative rule claims adjudicated to a machine verdict (verified over a stated domain / refuted with a concrete counterexample / not-formalisable with stated reason), with ≥1 counterexample confirmed valid by the original publisher or an independent tax-benefit analyst, and the full statute-pinned encoding published. CHANGED: the claim floor rises from 30 to 60. Statute encoding was the binding constraint and historically consumed person-years; with that removed, 30 claims no longer represents a meaningful stretch. Resolves NO if fewer than 20 sampled claims can be given a precise formal statement at all — the formalisability floor is unchanged, because it measures the quality of public claims rather than the capacity of the team.",
      "horizon_months": 7,
      "cost_band": "~£70k — one human tax-benefit analyst at ~0.5 FTE for 7 months plus ~£15k legal review; encoding and verification agent-executed; negligible compute",
      "crux": "Whether published claims are stated precisely enough to formalise. 'Most families will be better off' conceals both quantifier and population; the 20-claim floor exists precisely to keep the formalisability finding distinguishable from the verification finding. Agent execution raises how many claims can be attempted but does not make vague claims precise.",
      "agent_execution": "Agents produce the Catala-style literate encoding pinned to statutory provisions, build the SMT/BMC harness, attempt formalisation of every candidate claim, and search for counterexample households. The human critical path is lawyer review of the encoding — which is the trust boundary and the reason the literate form exists — and independent analyst or original-publisher confirmation of at least one counterexample. Both are irreducible: the point of the design is that a human legal reader can check the encoding.",
      "durability": "Durable on a statutory clock rather than a model clock. A statute-pinned encoding and a verdict table remain valid until the law changes, and the versioning is explicit. Nothing in the deliverable is a property of a model generation. The harness is reusable against every subsequent reform.",
      "judge_probs": [
        0.55,
        0.58,
        0.55
      ],
      "judge_payoffs": [
        0.65,
        0.68,
        0.62
      ],
      "p_med": 0.55,
      "pay_med": 0.65,
      "rationales": [
        "Statute encoding was the binding constraint and historically the reason for an 18-month horizon; with Catala-style literate encoding and the SMT/BMC harness agent-executed, raising the floor from 30 to 60 claims is close to free because 'not-formalisable with stated reason' counts as a machine verdict, leaving the real gates as the unchanged 20-claim formalisability floor and one counterexample confirmed by a publisher or independent analyst. The uncompressible human path is lawyer review of the encoding, which is the whole point of the literate form since the encoding is the trust boundary. Payoff is durable on a statutory rather than a model clock: a statute-pinned encoding and verdict table remain valid until the law changes, the harness is reusable against every subsequent reform, and even the failure mode is a publishable finding that public rule claims are unfalsifiable as written.",
        "Reference class: formal verification against an executable reference model, where statute encoding was the historic bottleneck and is now agent work against OpenFisca/PolicyEngine substrates — which is why the claim floor could sensibly rise from 30 to 60 and still be attainable, given a large candidate pool of impact-assessment and think-tank claims stated with explicit household arithmetic. Human critical path is genuine but bounded: lawyer review of the literate encoding (the actual trust boundary, and the reason the Catala-style form exists) and one counterexample confirmed by an independent analyst, which the 'or original publisher' alternative makes easy. Payoff is durable on a statutory rather than a model clock — a statute-pinned encoding and verdict table stay valid until the law changes, the harness is reusable against every subsequent reform, and nothing in the deliverable is a property of a model generation.",
        "Raising the floor from 30 to 60 claims costs little once agents do the Catala-style statute-pinned encoding and the SMT/BMC search, and 'not-formalisable with stated reason' counts as a verdict, so the 60-claim gate is largely mechanical; the binding gates are lawyer review of the encoding, which is the trust boundary and deliberately irreducible, and external confirmation of at least one counterexample — softened considerably by allowing an independent analyst rather than the original publisher. Encoding piecewise tapers into decidable form remains fiddly but tractable over a declared household space. Payoff is durable on a statutory rather than a model clock: the encoding, verdict table and harness stay valid until the law changes and are reusable against every subsequent reform."
      ],
      "essay_rank": 7,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P12",
      "title": "Contamination forensics and fragility bounds for published consultation responses",
      "description": "Assemble every public consultation in a defined window that publishes individual raw responses — government departments plus the sectoral regulators that routinely do so, with the US regulations.gov docket system as the largest single publisher and home of the FCC precedent. Pre-register a detection protocol (near-duplicate clustering, stylometry, perplexity-curvature, submission-timing burstiness where metadata survive), calibrate the false-positive rate on the near-certainly-human pre-November-2022 stratum, disaggregate FPR by respondent type and non-native-speaker proxies, and estimate post-2022 synthetic prevalence as the excess over baseline, with capture–recapture across two methodologically independent detectors bounding what detection missed. For each consultation, compute partial-identification bounds on the reported conclusions and the tipping fraction γ* at which the reported direction flips — a fragility index analogous to the E-value — published alongside the detection-adjusted contamination estimate.",
      "resolution_criterion": "By month 9: screen ≥100 consultations covering ≥20,000 individual responses; for ≥20 with obtainable item-level data, publish reproducible γ* values and detection-adjusted prevalence lower bounds, with detectors validated on held-out known-synthetic submissions at stated sensitivity and FPR calibrated on the pre-2022 stratum. Adjudicate two pre-registered claims: (a) median γ* across the twenty is below 5% — conclusions flip under contamination an order of magnitude smaller than the ~80% seen in the FCC proceeding; (b) at least one consultation's estimated synthetic share has a 95% lower bound above 5%. Both resolve in either direction. Gates unchanged; the prevalence estimate must additionally be stamped with the generator families against which detectors were validated.",
      "horizon_months": 9,
      "cost_band": "~£85k — one human researcher full-time for 9 months, mostly data-access negotiation and disclosure handling; compute under £5k",
      "crux": "Publication practice and estimand fuzziness. If most consultations release only summaries, the corpus does not exist at scale — which is why the largest docket system, publisher of nearly everything, is the natural first target. And legitimate AI-assisted human responses are not contamination, so 'synthetic share' conflates authorship with origination unless the protocol can separate them, which may not be possible from text alone and becomes less possible every year.",
      "agent_execution": "Agents assemble and normalise the corpus across 100+ consultations, implement all four detector families, run FPR calibration on the pre-2022 stratum, execute the capture–recapture bounding, and compute γ* across every published theme table. The human critical path is negotiating item-level data access with consultation-owning bodies, which is calendar-bound, plus the disclosure and ethics judgement around publicly characterising submissions as synthetic. Data access, not analysis, sets the horizon.",
      "durability": "The project splits cleanly. γ* and the partial-identification bounds are arithmetic over published theme tables — model-independent, permanently valid, and the durable half. Detection-adjusted prevalence has the shortest half-life in the portfolio: detectors calibrated on 2026 generators fail on later ones, and the estimand itself degrades as AI-assisted human writing becomes universal. Fund it, but expect the publication weight to sit on the fragility half.",
      "judge_probs": [
        0.52,
        0.45,
        0.58
      ],
      "judge_payoffs": [
        0.48,
        0.42,
        0.45
      ],
      "p_med": 0.52,
      "pay_med": 0.45,
      "rationales": [
        "The project splits into a half that will deliver (gamma* and partial-identification bounds are arithmetic over published theme tables, and agents compute them across every consultation) and a half that carries the risk (detector validation and FPR calibration on the pre-2022 stratum), with corpus availability the historic constraint that a fully public, agent-harvestable docket system largely dissolves. Both pre-registered claims resolve either way, so the gate is really assembly plus validation, with human time going to data-access negotiation and the disclosure judgement around publicly calling submissions synthetic. Payoff is the most heavily discounted in the portfolio: detection-adjusted prevalence has the shortest half-life here, since detectors calibrated on 2026 generators fail on later ones and the estimand itself degrades as AI-assisted human writing becomes universal, leaving the durable weight on the fragility half alone.",
        "Reference class: AI-text-detection prevalence estimation, which has a poor track record of producing defensible numbers — agents can assemble 100+ consultations and run four detector families and the capture–recapture bounding cheaply, but false-positive calibration on the pre-2022 stratum, disaggregated by respondent type and non-native-speaker proxies, is exactly where such studies fail, and data-access negotiation with consultation-owning bodies is calendar-bound rather than labour-bound. The fragility half will deliver; the prevalence half is the drag on the conjunction. Payoff is the portfolio's most sharply split: gamma* and the partial-identification bounds are permanently valid arithmetic over published theme tables, but detection-adjusted prevalence has the shortest half-life in the list, and the estimand itself degrades toward meaninglessness as AI-assisted human writing becomes universal, so I discount the composite heavily.",
        "The gamma* half is arithmetic over published theme tables and will deliver; both pre-registered claims resolve in either direction, so the binding constraints are item-level data access — largely solved by targeting the docket system that publishes nearly everything — and detector validation at stated sensitivity with FPR calibrated on the pre-2022 stratum, which agents can execute at scale. My scepticism is that the prevalence gate is satisfiable while being close to hollow: a stated FPR does not rescue an estimand that conflates AI-assisted human authorship with synthetic origination. Payoff is the portfolio's most supersession-exposed — detectors calibrated on 2026 generators fail on later ones and the estimand itself degrades yearly — leaving the durable weight on a fragility index that is simple enough for others to reproduce."
      ],
      "essay_rank": 8,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P09",
      "title": "Identification-gate certificates: witness formats and independent checkers for a stage-5 identification gate",
      "description": "For each identification method an evidence-review gate admits — rank, null-space, observability, causal-equivalence, optimisation, symbolic, countermodel — define a witness format and a small independent checker in exact arithmetic: exact rational null-space vectors for linear non-identification with an interval-arithmetic variant for floating-point inputs, explicit observationally-equivalent parameter pairs as countermodels, hedge output re-derived by a second implementation, autobounds primal DGPs at interval endpoints with epsilon-sharpness bounds, and LRAT proofs validated by an independent proof checker for SAT-encoded infeasibility. Retro-fit to an existing published non-identification result and two new live cases. The output is a portable certificate format: a reader who trusts only the checker, and not the analysis that produced the claim, can confirm the identification verdict.",
      "resolution_criterion": "By month 8: three completed cases ship with witnesses that a checker of under 2,000 lines — written from the published manifest alone, without production-code access, by an implementer sharing no model family with the producing system — validates successfully; and at least one witness format kills a planted defect in a blinded mutation test that existing deterministic-replay machinery passes. Resolves NO if any of the three cases' central identification result admits no expressible witness. CHANGED: 'a different team' becomes 'no shared model family', for the same reason as P08 — agent authorship makes team independence cheap and model independence the substantive condition. The conjunction is otherwise preserved in full, including the 2,000-line ceiling.",
      "horizon_months": 8,
      "cost_band": "~£60k — one human econometrician at ~0.5 FTE for 8 months, with formal-methods implementation agent-executed",
      "crux": "Whether load-bearing identification results in real policy cases are of a certifiable form at all. Many genuine non-identification arguments are semi-parametric and essentially verbal; if the certifiable fraction is small, this avenue is much narrower than the method inventory suggests. That is a fact about econometric practice, not about execution capacity, so it is untouched by agent execution.",
      "agent_execution": "Agents do essentially all of the formal work: witness formats, exact-rational null-space computation, interval-arithmetic variants, countermodel construction, autobounds endpoint derivations, SAT encodings with LRAT proof emission, the under-2,000-line independent checker built from the manifest alone, and the blinded mutation suite. The human critical path is selecting the three cases and judging whether each central identification argument is of certifiable form. Formal-methods labour, historically the reason this project cost 18 months, is the part that compresses hardest.",
      "durability": "The most durable class of result in the portfolio alongside P03. A witness format is a definition, a checker is a program in exact arithmetic, and a validated proof stays validated regardless of what models exist later. Nothing here is a measurement of a model. Under the supersession doctrine this project rises sharply in relative attractiveness: it was ranked fifteenth largely on the cost of formal-methods labour, which is precisely the cost agents remove.",
      "judge_probs": [
        0.38,
        0.52,
        0.5
      ],
      "judge_payoffs": [
        0.62,
        0.7,
        0.6
      ],
      "p_med": 0.5,
      "pay_med": 0.62,
      "rationales": [
        "This is the portfolio's largest re-rating: it sat at 0.18 principally because formal-methods labour (exact rational null-spaces, interval variants, autobounds endpoints, SAT encodings with LRAT emission, an under-2,000-line independent checker, a blinded mutation suite) consumed person-years, and that is precisely the cost agents remove. What survives the compression is the severe conjunction and the structural risk that many genuine non-identification arguments in real policy cases are semi-parametric and essentially verbal, so no witness exists to express; case selection gives the team some latitude but the criterion resolves NO if any of three cases fails. Payoff is durable in the strongest sense alongside P03 (a witness format is a definition, a checker is a program in exact arithmetic, a validated proof stays validated), held short of the top only because certifiable identification claims are a narrow slice of policy evidence.",
        "Reference class: formal-methods certificate work, which historically cost person-years of specialist labour — and that cost is precisely what agents remove, since exact rational null-space computation, interval-arithmetic variants, countermodel construction, SAT encoding with LRAT emission and a sub-2,000-line independent checker are all squarely within 2026 agent capability; this is the largest justified re-rating in the portfolio, from 0.18 upward. The surviving risk is not execution but subject matter: many genuine non-identification arguments in real policy cases are semi-parametric and essentially verbal, admitting no expressible witness, and the criterion resolves NO if any one of the three central results is of that form. Payoff sits with P03 as the most durable class in the list — a witness format is a definition and a validated proof stays validated regardless of what models exist later — discounted only for the narrowness of identification-gated evidence.",
        "This is the portfolio's largest genuine re-rating: formal-methods labour — witness formats, exact rational null-spaces, interval arithmetic, LRAT emission, the sub-2,000-line checker from manifest alone — is precisely what agents absorb, and the historical 18-month cost was almost entirely that labour. The structural crux survives (many real non-identification arguments are verbal and admit no witness), but that branch resolves NO rather than failing to resolve, and human case selection can steer toward certifiable cases; the residual risk is needing two live cases inside 8 months and an unambiguous mutation kill. Payoff is durable in the strongest sense — a witness format is a definition and a checker is a program in exact arithmetic — but reach is narrow if the certifiable fraction of real policy identification claims is small."
      ],
      "essay_rank": 9,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P01",
      "title": "PolicyCAPA: cross-model error correlation and the generator–verifier miss rate on policy tasks",
      "description": "A fully crossed measurement of ρ, the correlation of errors across model families, on four policy-analytic task families: consultation-response coding, verification of statistical claims against official statistical tables, extraction of regulatory obligations from versioned legislation substrates, and appraisal of identification strategies. Roughly 1,000 items per family, eight model families including at least two open-weight, three prompt replicates, with human-adjudicated gold on a double-adjudicated 15% subsample and latent-class treatment of rater disagreement. GLMM variance components yield CAPA, error consistency, the induced Kish ceiling 1/ρ, and the conditional quantity P(verifier passes | generator wrong), which governs how far machine pre-screening can safely displace human hours. A two-parameter (c, β) scaling law is fitted on pilots of five or fewer agents and tested by extrapolation to twenty. Task families are specified against generic substrates (a national statistical institute's published tables, a point-in-time legislation register, a published consultation corpus) so the design ports without modification.",
      "resolution_criterion": "By month 7, under a pre-registered analysis plan: publish ρ per task family with 95% CI half-width ≤ 0.05 and the generator–verifier conditional miss correlation with a stated interval; and adjudicate the headline test that the (c, β) law fitted on ≤5-agent pilots predicts measured 20-agent effective diversity inside its stated prediction interval for ≥3 of 4 task families. Either direction resolves; extrapolation failure invalidates the cheap-pilot recommendation now circulating and is the more valuable result. Gates unchanged: agent execution makes item counts and model crossings nearly free, but the precision on ρ is bounded by the size of the human gold subsample, which is not compressible.",
      "horizon_months": 7,
      "cost_band": "~£150k — roughly 1.5 human FTE over 7 months directing agent execution, of which ~£60k is expert adjudication; inference under £20k",
      "crux": "Gold-standard construction. If expert–expert agreement resembles the 55% seen in deployed consultation-analysis evaluations, measured 'shared model error' is partly shared item ambiguity, and ρ is identifiable only under conditional-independence assumptions between raters that are probably false. Agent execution does not touch this: the humans are the instrument.",
      "agent_execution": "Agents build the four task families, assemble and clean the corpora, run the entire crossed design (eight families × three replicates × ~4,000 items), implement the GLMM variance-component estimation, and fit and cross-validate the scaling law. The human critical path is the double-adjudicated expert gold subsample, the latent-class identification review, and pre-registration sign-off. Wall-clock is now set by how fast expert adjudicators can be recruited and worked through, not by inference or analysis labour.",
      "durability": "ρ measured on named model versions is the portfolio's fastest-decaying quantity — it is a property of a model landscape that turns over inside a year. What survives is the crossed design, the adjudicated gold corpus (which retains reference value indefinitely), and a re-runnable harness; the project is worth funding only if it ships as a standing instrument re-run each model generation rather than as a single published number.",
      "judge_probs": [
        0.42,
        0.45,
        0.42
      ],
      "judge_payoffs": [
        0.55,
        0.55,
        0.55
      ],
      "p_med": 0.42,
      "pay_med": 0.55,
      "rationales": [
        "Reference class is fully-crossed multi-rater measurement studies with a tight pre-registered precision gate plus a conjunctive out-of-sample extrapolation test; those deliver both halves maybe two times in five. Agents make the eight-family crossing, the GLMM, and the scaling-law fit essentially free, but CI half-width on rho is bounded by the double-adjudicated expert subsample and by latent-class identifiability under rater assumptions that are probably false, and compressing 12 to 7 months tightens the recruitment path rather than loosening it. Payoff falls hardest of any project under supersession: rho on named model versions is the portfolio's fastest-decaying number, so what lands in the world of month seven is the crossed design, the gold corpus, and the re-runnable harness, not the headline coefficient.",
        "Reference class: large crossed measurement studies with a tight pre-registered precision gate; agents make the 8-family × 3-replicate × 4,000-item design nearly free, so the binding constraint is the 15% double-adjudicated expert subsample, which caps precision on rho and cannot be compressed — a CI half-width of 0.05 off ~150 gold items per family is a genuine stretch, and the crux (shared model error vs shared item ambiguity) is an identification problem agents cannot dissolve. The scaling-law sub-goal resolves either way, which helps. Payoff is heavily supersession-discounted: rho on named 2026 model versions is the portfolio's fastest-decaying quantity, and only the adjudicated gold corpus, the crossed design and the re-runnable harness survive the next model generation.",
        "Reference class: crossed multi-rater reliability studies with a tight pre-registered precision gate, which historically miss the gate rather than the design. Agents make the 8x3x4,000 crossing and the GLMM essentially free, but the CI half-width of 0.05 on rho is bounded by the double-adjudicated expert subsample and by latent-class identifiability under rater dependence — the exact place agents cannot help, and 7 months is tight for recruit-adjudicate-analyse-preregister. Payoff is heavily supersession-discounted: rho on named model versions is the portfolio's fastest-decaying quantity, and only the adjudicated gold corpus and re-runnable harness survive the next model generation."
      ],
      "essay_rank": 10,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P06",
      "title": "Certified auto-accept: conformal risk control for claim-level triage and its coverage–risk cost curve",
      "description": "Score every machine-generated claim with candidate nonconformity measures (source-hash resolvability, entailment against the retrieved passage, self-consistency across samples, retrieval agreement, ensemble disagreement), calibrate thresholds with conformal risk control and Learn-then-Test on a human-labelled set, and route only sub-threshold claims to humans. The central deliverable is the coverage–risk curve — auto-accept fraction as a function of certified claim-error rate, per rung of the verifiability stack — which is the cost curve any analytical organisation needs before designing an assurance budget. Transfer is tested explicitly: calibrate on one policy domain, deploy on three held-out domains, and quantify the realised coverage gap against the beyond-exchangeability bound. The deliverable is packaged as a re-runnable calibration harness, not a single published curve.",
      "resolution_criterion": "By month 6: at a certified 1% claim-error level with 95% confidence, ≥50% of claims are auto-accepted, and on three held-out policy-domain corpora the empirical risk exceeds nominal in <10% of 1,000 calibration re-draws. An auto-accept fraction below 20% at the 1% level also resolves — negatively but informatively, establishing that per-item certified triage does not pay at policy-grade tolerances with the scoring functions available at the time of the run. ADDED: every reported curve must be stamped with model identifiers and run date, since the curve is a property of a model generation rather than of the method. Gates otherwise unchanged.",
      "horizon_months": 6,
      "cost_band": "~£70k — one human researcher at ~0.5 FTE for 6 months directing agent execution, including ~£12k compute for multi-sample generation",
      "crux": "Score separability. If the best nonconformity score discriminates at AUC ≈ 0.65, auto-accept at strict risk levels collapses toward zero. The deeper danger from the correlated-error literature is that machine scorers fail precisely on items the generator got confidently wrong, so the score is least informative exactly where it is needed. The most likely landing point remains the criterion's dead zone between 20% and 50%.",
      "agent_execution": "Essentially the whole build is agent-executable: implementing and sweeping the nonconformity scores, the CRC and Learn-then-Test calibration, the multi-sample generation, the cross-domain transfer evaluation and the 1,000-draw re-calibration study. The human critical path is the labelled calibration set and the decision about what counts as a claim error at policy-grade tolerance. Compute, not labour, is now the dominant marginal cost.",
      "durability": "The coverage–risk curve is a current-model measurement with a short half-life — it moves with every capability release, generally upward. The durable contribution is the calibration harness, the inventory of nonconformity scores, and the beyond-exchangeability transfer protocol, all of which are re-runnable in hours once built. Fund it as an instrument; treat any published curve as a dated observation.",
      "judge_probs": [
        0.42,
        0.45,
        0.35
      ],
      "judge_payoffs": [
        0.55,
        0.48,
        0.55
      ],
      "p_med": 0.42,
      "pay_med": 0.55,
      "rationales": [
        "The structural problem is not execution but the criterion's dead zone: CRC and Learn-then-Test are mature, agents can sweep every nonconformity score and run the cross-domain transfer study in weeks, and the most likely landing point remains an auto-accept fraction between 20 and 50 per cent at a certified 1 per cent claim-error level, which resolves neither way. The correlated-error mechanism sharpens this, since machine scorers plausibly fail precisely on items the generator got confidently wrong. Payoff is discounted heavily from its original 0.80 because a coverage-risk curve is a dated observation of one model generation with a half-life measured in months; what lands durably is the calibration harness, the score inventory, and the beyond-exchangeability transfer protocol, all re-runnable in hours.",
        "Reference class: conformal risk control applied to a new scoring problem — the calibration machinery is off-the-shelf and agent execution lets the whole nonconformity-score space be swept, but the criterion demands landing outside the 20–50% dead zone, and the correlated-error literature predicts precisely the dead-zone outcome because machine scorers are least informative exactly where the generator was confidently wrong. Score separability, not execution capacity, decides it, and no amount of agent labour raises an AUC. Payoff is the weakest in the portfolio on supersession grounds: the coverage–risk curve is a dated property of one model generation, moving upward with every capability release, so what survives is only the calibration harness, the score inventory and the beyond-exchangeability transfer protocol — real but modest instrument value.",
        "The criterion has an undefined 20-50% band between its positive and negative branches, and the project's own analysis names that band as the most likely landing point; agents make the score sweep, the CRC/Learn-then-Test calibration and the 1,000-draw transfer study nearly free without moving where the curve lands, which is the one thing that decides resolution. The deeper methodological worry is untouched by acceleration: correlated-error arguments predict the nonconformity score is least informative exactly on items the generator got confidently wrong. Payoff sits mid-scale because the coverage-risk curve is a dated observation with a short half-life; the durable residue is the calibration harness, the score inventory and the beyond-exchangeability transfer protocol."
      ],
      "essay_rank": 11,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P13",
      "title": "A no-cooperation checker for inbound evidence, benchmarked on theme-table distortion",
      "description": "Build a checker requiring nothing from the submitter — statcheck-class recomputation of test statistics, GRIM/GRIMMER attainability, shares summing, stock-flow identities, per-capita reconciliation against published denominators, citation resolution and hash identity against archived public substrates, near-duplicate and template detection — shipped as an open tool a consultation team can run over 50,000 responses in under an hour on commodity hardware. Evaluate on a purpose-built contamination benchmark: a closed, fully published consultation as clean base, attacks at four sophistication tiers (duplicate/template, naive zero-shot, persona-conditioned style transfer, retrieval-grounded with genuine local specifics plus light human editing) planted at prevalences from 0.1% to 50%, plus a paid red-team round designed to evade the checker with evasion cost documented. Score defences — the checker at fixed FPR, near-duplicate removal, argument-level clustering with per-cluster capping, provenance weighting — not on classification accuracy but on distortion of the theme table: L1 distance between clean and contaminated theme shares, and rank inversions in the top ten. Harness and results published; the attack generator withheld.",
      "resolution_criterion": "By month 8, pre-registered: (a) ≥80% detection of arithmetic and attribution defects at ≤5% false-positive rate measured on genuine submissions from non-expert members of the public (resolves NO if FPR exceeds 15% at any detection rate above 50% — the tool would then be a tax on legitimate participation); (b) at 10% tier-3 contamination, at least one evaluated defence reduces theme-share L1 distortion by ≥50% relative to no defence, with the comparison against detectors at FPR ≤1% reported; (c) a published evasion-cost characterisation from the red-team round, whatever it shows. Gates unchanged, with the attack tiers required to be versioned by generator family and date.",
      "horizon_months": 8,
      "cost_band": "~£95k — one human engineer at ~0.5 FTE for 8 months plus ~£25k contracted red team and ~£10k compute; the red-team line does not compress",
      "crux": "The false-positive rate on genuine public submissions, which routinely contain loose arithmetic, misremembered figures and informal citation. If legitimate submissions trip the checker at anything like the rate fabricated ones do, precision collapses and the tool harms the process it protects. Secondary: whether capping is institutionally acceptable — it makes argument-weighting visible and contestable for the first time, so a policy-side partner must be recruited at the outset, not at the end.",
      "agent_execution": "Agents build the entire checker, construct the four-tier contamination benchmark at all prevalences, implement and score the four defence strategies, and compute the L1 and rank-inversion metrics. The human critical path is the contracted red team — an adversary with a budget and an incentive is not substitutable by the same agents that built the defence — and the policy-side partner needed to make capping a live option. Red-team cost arguably rises rather than falls, since a competent red team now works with agents too.",
      "durability": "The checker, the theme-table-distortion scoring methodology and the defence comparison framework are durable; a GRIM check does not expire. The attack tiers decay fast — tier-3 and tier-4 attacks that cost effort in 2026 become trivial soon after — so the benchmark must be published with generator provenance and re-populated periodically. The evasion-cost figure is a dated observation, not a constant.",
      "judge_probs": [
        0.4,
        0.38,
        0.4
      ],
      "judge_payoffs": [
        0.62,
        0.52,
        0.58
      ],
      "p_med": 0.4,
      "pay_med": 0.58,
      "rationales": [
        "Everything mechanical compresses (the checker, the four-tier benchmark across all prevalences, the four defence strategies, L1 and rank-inversion scoring), but delivery hinges on an empirical number agents cannot move: 80 per cent detection of arithmetic and attribution defects at 5 per cent false-positive rate on genuine submissions from members of the public, who legitimately write loose arithmetic and informal citation. The contracted red team is the other uncompressed line and arguably gets more expensive, not less, since a competent adversary now works with agents too, and a policy-side partner must be recruited at the outset for capping to be a live option. Payoff holds up moderately: the checker and the theme-table-distortion scoring methodology are durable (a GRIM check does not expire), while the attack tiers and the evasion-cost figure are dated observations requiring periodic re-population.",
        "Reference class: precision-constrained automated screening deployed against genuine public prose; parts (b) and (c) are near-free — capping and near-duplicate removal mechanically cut theme-share L1 distortion at 10% contamination, and the evasion-cost characterisation resolves whatever it shows — but gate (a) demands 80% detection at ≤5% false positives on submissions from non-expert members of the public who legitimately write loose arithmetic and informal citations, which is where GRIM- and statcheck-class checks characteristically bleed precision. The contracted red team is an irreducible human line item whose cost arguably rises, since a competent adversary now works with agents too. Payoff is moderate: the checker and the theme-table-distortion scoring framework endure, but the four attack tiers are dated observations that need periodic re-population, and the tool's reach is confined to consultation-style inbound evidence.",
        "Component (b) should clear — near-duplicate removal and per-cluster capping plausibly halve theme-share L1 distortion at 10% contamination — and (c) resolves whatever it shows, but (a) carries a 5-15% FPR dead zone on genuine public submissions whose arithmetic is sparse and loose, which is exactly the crux and exactly what agent scale does not touch. The contracted red team is the one line item that may cost more, not less, under agent acceleration, since a competent adversary now works with agents too. Payoff is moderate: the checker and the distortion-scoring reframing are durable, but the four attack tiers decay fast enough that the benchmark needs periodic re-population, and capping's institutional acceptability is unresolved."
      ],
      "essay_rank": 12,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P08",
      "title": "Citation receipts with inclusion proofs: a verifiable evidence-bundle profile for versioned public sources",
      "description": "Specify a citation-receipt format — resolver, retrieval timestamp, SHA-256 of the canonicalised object, canonicalisation spec version, span byte offsets, normalised quoted text, archive fallback via a national web archive or Memento — targeting the five families of public source that have versioned substrates: point-in-time legislation XML, national statistical dataset editions, legislative debate records, government web publications via archive snapshots, and DOI-bearing academic sources. Package receipts as an in-toto/SLSA-style provenance statement over a content-addressed frozen corpus, each citation carrying an inclusion proof against a published Merkle root, using RO-Crate packaging and COSE signatures; invent nothing. Ship an offline checker, apply the profile to a re-run at deployed-system scale (≥50,000 objects), commission an independently written second verifier, and measure two things: how much of the documented 39–77% claim-support gap the mechanical checks convert into mechanical failures, and expert quote-checking time on bundled versus unbundled output.",
      "resolution_criterion": "By month 7: (a) an open spec plus offline checker running against ≥5 versioned source families; (b) on ≥500 citations, hash-identity and span-containment checking flags ≥60% of citations a blind three-person panel independently rules 'unsupported' (resolves NO below 40%, establishing citation assurance for policy prose as irreducibly semantic); (c) a published bundle over a ≥50,000-object corpus in which ≥99% of citation claims verify offline against the published root, confirmed by a second verifier that shares no code with the reference implementation AND was written by a different model family. CHANGED: agent execution makes an independently written second verifier cheap, which weakens the original independence test — two agent-written verifiers from the same family share failure modes. Model-family independence replaces team independence as the binding condition.",
      "horizon_months": 7,
      "cost_band": "~£90k — one human engineer-researcher for 7 months directing agent implementation, including ~£8k adjudication panel and ~£4k inference",
      "crux": "Whether policy prose quotes and decomposes enough for span anchoring to bite. If agent-generated analysis is predominantly paraphrase and synthesis ('respondents were broadly concerned about X'), the mechanical oracles have little surface and the profile looks like ceremony. Mundane but real: PDF text extraction is non-deterministic across tool versions, so canonicalisation must be pinned hard enough that hashes are stable across independent implementations — a requirement agent-written implementations make easier to test and easier to accidentally violate.",
      "agent_execution": "Agents draft the spec, build the offline checker, pin the canonicalisation, construct and verify the 50,000-object bundle, and write the second verifier under a different model family with manifest-only access. The human critical path is the blind three-person adjudication panel ruling on semantic support — the humans are the instrument for 'unsupported' — plus carrying the profile into a standards process. Two of the three deliverables are ordinary engineering now measured in weeks.",
      "durability": "The spec, the canonicalisation pinning and the offline checker are standards artefacts and durable indefinitely; a receipt verified in 2026 verifies in 2036. The measured conversion of the claim-support gap into mechanical failures is a current-model property and decays. The profile is worth building even if part (b) fails, because it ships the mechanical layer that policy evidence work currently lacks entirely.",
      "judge_probs": [
        0.33,
        0.3,
        0.5
      ],
      "judge_payoffs": [
        0.78,
        0.78,
        0.75
      ],
      "p_med": 0.33,
      "pay_med": 0.78,
      "rationales": [
        "Two of the three parts (spec plus offline checker, and a 50,000-object bundle verifying at 99 per cent against a published root with a second verifier from a different model family) become ordinary agent engineering measured in weeks, but delivery requires all three, and part (b) runs straight into the essay's own diagnosis that the 39-77 per cent claim-support gap is dominated by semantic non-support: hash identity and span containment plausibly flag well under the 40 per cent floor. Substituting model-family independence for team independence is the right repair, since two agent-written verifiers from one family share failure modes, and it costs little probability. Payoff stays high and undiscounted because a receipt verified in 2026 verifies in 2036: the spec, the pinned canonicalisation and the checker are standards artefacts, and the profile ships the mechanical layer policy evidence work currently lacks entirely even when part (b) fails.",
        "Reference class: standards-and-checker engineering, where parts (a) and (c) are now ordinary weeks-long agent work — spec, offline checker, pinned canonicalisation, a 50,000-object verified bundle and a second verifier from a different model family are all high-confidence deliveries. The conjunction fails on (b): the documented claim-support gap is dominated by semantic non-support, where the citation resolves and the span exists and the source still does not bear the claim, so mechanical hash-identity and span-containment flagging 60% of panel-ruled-unsupported citations is unlikely, and the 40–60% band resolves neither way. Payoff conditional on full delivery is high and near-permanently durable — a receipt verified in 2026 verifies in 2036, and the spec plus checker are standards artefacts indifferent to model turnover — with only the measured gap-conversion rate decaying.",
        "Parts (a) and (c) become ordinary agent-weeks of engineering, and swapping team independence for model-family independence keeps the second verifier a substantive test rather than a formality; the conjunction still turns on part (b), which has an undefined 40-60% band, and the documented claim-support gap is dominated by semantic non-support that hash-identity and span containment structurally cannot reach. Canonicalisation stability for PDF extraction across independent implementations is the mundane risk agents make both easier to test and easier to violate silently. Payoff is high and durable regardless: a receipt verified now verifies in a decade, and the profile ships the mechanical layer policy evidence work currently lacks entirely, so partial delivery still leaves standing infrastructure."
      ],
      "essay_rank": 13,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04",
      "source_note": "The 39–77% claim-support vs 94–100% link-validity range cited in this spec traces to \"Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents\" (2026), arXiv:2605.06635."
    },
    {
      "id": "P16",
      "title": "Coverage estimation and stopping rules for census-class evidence workloads",
      "description": "On a real census-class corpus — a 20,000–50,000-response consultation, or every local land-use plan bearing on one question — run three or more agent passes built on deliberately different architectures and model families. Treat distinct themes or extracted claims as capture events and apply capture–recapture estimators (Chao1, coverage-based rarefaction) to estimate how many themes no pass captured, adjudicating theme identity by embedding-match plus human review on a subsample. Deliver the marginal-valid-evidence curve, a coverage-based stopping rule, a direct estimate of pairwise inter-pass agreement locating the workload in the correlation-regime table, and the sampled-assurance cost curve as N varies over a ≥20× range — a direct empirical test of the rule-of-three prediction that assurance cost per certified error bound is flat in corpus size.",
      "resolution_criterion": "By month 7: on ≥3 held-out corpora with a human-exhaustive gold standard on a subsample, the coverage estimator predicts the missed-theme count to within ±25% of the observed shortfall; and the measured sampled-assurance cost per certified error bound varies by less than 2× across a 20× range of corpus size. Either component failing resolves negatively — the first would undermine census stopping rules, the second the flat-cost claim at the heart of the census-class argument. Gates unchanged; the design must now additionally document how architectural independence between passes was established, since converging model families make 'deliberately different architectures' harder to satisfy than when the design was written.",
      "horizon_months": 7,
      "cost_band": "~£70k — one human researcher at ~0.5 FTE for 7 months, of which the majority is human-exhaustive gold-standard adjudication; compute ~£10k",
      "crux": "Theme individuation. Capture–recapture needs a stable identity criterion for the units recaptured, and themes are not species: two passes naming the same theme differently inflate richness, an over-aggressive matcher deflates it, and the human gold standard needed to validate matching is itself the expensive thing the method is meant to avoid. Agent execution makes the passes free but not the gold standard, so the cost concentrates entirely in the crux.",
      "agent_execution": "Agents run all the passes, implement the capture–recapture estimators and embedding-based theme matching, sweep the corpus-size cost curve across its 20× range, and compute pairwise inter-pass agreement. The human critical path is the exhaustive gold standard on a subsample — establishing that no pass captured a theme requires a human who has read everything — plus adjudicating theme identity where matching is contested. Nearly all remaining cost is that adjudication.",
      "durability": "The coverage estimators and the stopping rule are durable method, and the flat-cost claim is structural arithmetic from the rule of three, so it does not decay at all. The measured inter-pass agreement is a current-model quantity with a short half-life, and it is threatened from a second direction: as model families converge, architecturally independent passes become harder to construct, which erodes the design's central assumption rather than merely its numbers.",
      "judge_probs": [
        0.26,
        0.28,
        0.25
      ],
      "judge_payoffs": [
        0.58,
        0.55,
        0.55
      ],
      "p_med": 0.26,
      "pay_med": 0.55,
      "rationales": [
        "Agents make the passes, the capture-recapture estimators, the embedding matcher and the 20x cost sweep free, so essentially all remaining cost concentrates in the one thing they cannot do: a human-exhaustive gold standard, which requires someone who has read everything to establish that no pass captured a theme. The gates are asymmetric, with the flat-cost component near-certain because it is structural arithmetic from the rule of three, and the +/-25 per cent accuracy gate on a Chao1-class estimator over themes genuinely hard, since themes are not species and richness estimators are badly behaved under unstable identity criteria. Payoff is middling and eroding from an unusual direction: the estimators and stopping rule are durable method, but as model families converge, architecturally independent passes get harder to construct, which attacks the design's central assumption rather than merely its numbers.",
        "Reference class: capture–recapture applied to units without a stable identity criterion, which is where such estimators characteristically fail — agents make the three-plus passes and the 20× cost sweep free, but the crux concentrates all remaining cost in the one thing they cannot supply: a human-exhaustive gold standard on three corpora, requiring someone who has read everything to establish that no pass captured a theme, and adjudication where embedding-matching is contested. The ±25% accuracy gate on missed-theme count is demanding under theme individuation instability. The flat-cost component is near-structural arithmetic and will pass. Payoff is mid-tier: the stopping rule and coverage estimators are durable method, but the inter-pass agreement measurement decays fast and is threatened from a second direction, since converging model families erode the architectural-independence assumption the design rests on rather than merely dating its numbers.",
        "The flat-cost component is close to arithmetic from the rule of three and should hold, but the coverage component is hard for reasons acceleration does not touch: capture-recapture with heterogeneous capture probabilities characteristically underestimates richness by more than the +/-25% tolerance, theme individuation has no stable identity criterion, and the human-exhaustive gold standard is both the validating instrument and the expensive thing the method exists to avoid. Converging model families additionally erode the architectural independence the estimator assumes, so the design's central premise is weakening as it is tested. Payoff is mid-scale: the stopping rule and estimators are durable method and the flat-cost claim does not decay at all, but the measured inter-pass agreement is a short-half-life quantity threatened from two directions."
      ],
      "essay_rank": 14,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P14",
      "title": "Replay conformance classes for agentic evidence pipelines, with a measured divergence dataset",
      "description": "Re-run three to five real large-scale analytical pipelines 40–60 times each, varying seed, batch composition, provider endpoint and model minor version across a span of months. Measure divergence at three levels — raw tokens, extracted claim set, substantive conclusion — and use the data to define and calibrate three conformance classes: C1 deterministic (non-LLM stages bit-exact), C2 verifier-stable (checks reproduce every pass/fail verdict), C3 distributional (claim-level metrics within pre-registered tolerance). The central hypothesis is that C2 is stable even where C3 is not, which is what would let assurance be anchored to the verifier rather than the generator. The divergence dataset, which does not exist for policy-analytic workloads, is published and reusable well beyond this question.",
      "resolution_criterion": "By month 8: a published dataset of ≥200 re-runs across ≥3 pipelines; a conformance specification that classifies run pairs in agreement with blinded, external human adjudication of 'substantive conclusion unchanged' on ≥95% of a held-out set of ≥200 pairs, with the rubric pre-registered before adjudication; and a reported C2 verdict-stability rate with confidence interval. ADDED: the inter-human agreement ceiling on the same adjudication task must be measured and published alongside, since the 95% gate may exceed it and a result at 88% against an 89% ceiling means something very different from 88% against a 99% ceiling. The 95% gate itself is unchanged — agent execution does not raise a human agreement ceiling.",
      "horizon_months": 8,
      "cost_band": "~£60k — one human researcher at ~0.4 FTE over 8 months plus ~£15k inference; horizon set by the required drift span, not by labour",
      "crux": "Operationalising 'substantive conclusion unchanged' without circularity. If the adjudication rubric is written by the people who design the tolerance bands, the 95% agreement figure is manufactured; the honest risk is landing at 80% agreement, a much less useful result. Publishing the human ceiling is the defence against reading that number wrongly.",
      "agent_execution": "Agents execute all 200+ re-runs, instrument divergence at the three levels, draft and calibrate the conformance specification, and classify run pairs. The human critical path is the blinded external adjudication panel — humans are the instrument for 'substantive conclusion unchanged' — and, unusually, wall-clock: the design requires model minor-version drift to actually occur across a span of months, which no amount of parallelism produces. That calendar floor, not labour, is why this compresses only from 12 months to 8.",
      "durability": "The C1/C2/C3 conformance-class definitions are a durable standard and the intended deliverable. The divergence dataset is a snapshot of a 2026 model landscape; it decays as a live measurement but retains reference value as the first baseline of its kind. The structural claim that verifier-stability outlasts distributional stability is the sort of finding likely to survive model turnover, which is what makes the project worth running despite the decaying dataset.",
      "judge_probs": [
        0.3,
        0.25,
        0.25
      ],
      "judge_payoffs": [
        0.6,
        0.62,
        0.6
      ],
      "p_med": 0.25,
      "pay_med": 0.6,
      "rationales": [
        "Agents execute all 200-plus re-runs and the three-level divergence instrumentation, but two constraints are untouched: a calendar floor, because the design requires real model minor-version drift to occur across months and no parallelism produces that, and the 95 per cent agreement gate against blinded human adjudication of 'substantive conclusion unchanged', which may sit above the inter-human ceiling itself. Publishing that ceiling alongside is the right addition but explicitly does not relax the gate, so the honest modal outcome remains around 80 per cent agreement, a much less useful result that fails the criterion. Payoff rests on the C1/C2/C3 conformance-class definitions as a durable standard and on the structural claim that verifier-stability outlasts distributional stability, which is the kind of finding likely to survive model turnover; the divergence dataset decays as a live measurement but keeps reference value as the first baseline of its kind.",
        "Reference class: mechanical-vs-human agreement studies, which routinely land at 80–90% against blinded adjudication of a semantic judgement; the 95% gate is the binding problem and may sit above the inter-human ceiling itself, and publishing that ceiling alongside — a genuine improvement to the design — clarifies the result without relaxing the gate. Agents execute all 200+ re-runs and instrument divergence, but wall-clock is irreducible here in an unusual way: the design requires real model minor-version drift to occur across months, which parallelism cannot manufacture. Payoff conditional on delivery is solid and durable: the C1/C2/C3 conformance classes are a standard, and the structural claim that verifier-stability outlasts distributional stability is the kind of finding that survives model turnover, even as the divergence dataset itself ages into a first-of-its-kind baseline rather than a live measurement.",
        "The 200 re-runs and the divergence instrumentation are free under agent execution, but the gate is 95% agreement with blinded human adjudication of 'substantive conclusion unchanged', which plausibly exceeds the inter-human ceiling on that judgement — publishing the ceiling makes an 85% result interpretable without making it a pass, and the design explicitly declines to relax the bar. Wall-clock is also irreducible: the study needs real model minor-version drift to occur, which no parallelism produces. Payoff is solid because the C1/C2/C3 conformance classes are a standard rather than a measurement, and the structural claim that verifier-stability outlasts distributional stability is the kind of finding that survives model turnover even as the dataset ages into a historical baseline."
      ],
      "essay_rank": 15,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "P04",
      "title": "Blind proficiency testing of evidence assurers: planted-defect sensitivity and a corrected rule of three",
      "description": "An external-quality-assessment scheme measuring the number every sampled-audit certification silently assumes: auditor sensitivity s. Build a validated planted-error generator from a mutation-operator taxonomy (citation substitution, figure transposition, denominator swap, sign reversal, over-generalised causal verb, silent scope change, right number wrong period) calibrated against real adjudicated agent errors. Recruit 40–60 assurers spanning in-house analysts and accredited commercial assurance providers, with informed consent to unannounced test items; run announced-audit and blind-queue arms across three defect severity tiers, injecting via a shadow queue with hard interception so no planted defect reaches a live product. Estimate s by defect class, assurer and time; run a parallel-reviewer capture–recapture design as an independent second estimator; publish the corrected certification bound p ≤ 3/(n·s), per-auditor sensitivity control charts as the anti-rubber-stamping instrument, and a rotating item bank so the scheme persists as standing infrastructure.",
      "resolution_criterion": "By month 15, pre-registered: (i) realism — expert raters shown mixed planted and real errors classify provenance at ≤60% accuracy; (ii) convergent validity — planted-error and capture–recapture estimates of s overlap within 95% CIs in ≥3 of 4 defect classes; (iii) across ≥40 assurers and ≥200 episodes, report s per class with CI, adjudicate whether the uncorrected rule of three understates the certified error rate by ≥1.5× (s ≤ 0.67) in any class, and whether blind-arm detection on the highest-severity tier falls ≥15pp below announced-arm detection with a 95% interval excluding zero. All resolve either way; near-perfect auditors is a substantive negative. Gates unchanged.",
      "horizon_months": 15,
      "cost_band": "~£190k — roughly 1 human FTE over 15 months, of which ~£100k is participant and auditor time, which does not compress at all",
      "crux": "Institutional hosting. Without a public body or a national assurance competency programme routing blind items through real queues with real stamps at stake, the study degrades into a lab exercise measuring attentiveness rather than the rubber-stamping dynamic. Covert testing of professionals also needs an ethics route with no precedent in this sector. Neither is an execution problem an agent can absorb.",
      "agent_execution": "Agents build the mutation taxonomy and the planted-error generator, calibrate it against real adjudicated errors, implement the shadow queue with interception, the capture–recapture estimator, the control charts and the rotating item bank — the entire instrument can be ready before recruitment closes. The critical path is human and institutional throughout: recruitment and consent of 40–60 professionals, ethics approval, a host willing to expose live queues, and the calendar time needed to accumulate ≥200 episodes at realistic assurance throughput. This is the project that compresses least in the portfolio, from 18 months to 15.",
      "durability": "Highly durable, because the deliverable is a standing scheme and a corrected bound, not a measurement. p ≤ 3/(n·s) is arithmetic; the item bank, control charts and shadow-queue mechanism outlive any model generation. The measured value of s will drift — particularly as assurers themselves adopt agents — which is an argument for the scheme, since a standing instrument tracks that drift rather than being invalidated by it.",
      "judge_probs": [
        0.17,
        0.15,
        0.16
      ],
      "judge_payoffs": [
        0.82,
        0.82,
        0.85
      ],
      "p_med": 0.16,
      "pay_med": 0.82,
      "rationales": [
        "The reference class is covert proficiency-testing schemes for professionals, which have essentially only succeeded where a regulator or accrediting body mandated participation; a small external team achieves this rarely, and the criterion is a three-part conjunction on top. Agents can have the entire instrument (mutation taxonomy, generator, shadow queue with interception, capture-recapture estimator, control charts, item bank) built before recruitment even opens, which is why the horizon moves 18 to 15 months and the probability barely moves at all: ethics approval with no sector precedent, 40-60 consenting professionals, a host exposing live queues, and 200 episodes at real assurance throughput are all calendar- and institution-bound. Payoff stays at the top of the portfolio because the deliverable is a standing scheme and the corrected bound p <= 3/(n*s), both of which outlive any model generation, and the drift in measured sensitivity is an argument for the instrument rather than against it.",
        "Reference class: covert external-quality-assessment schemes in pathology and clinical labs, which have only ever succeeded under a regulator or accrediting body with mandated participation; the entire critical path — ethics approval with no sector precedent, an institutional host willing to route planted items through live queues with real stamps at stake, consent from 40–60 professionals, and calendar time to accumulate 200 episodes at real throughput — is human and institutional, so agents building the mutation generator, shadow queue and control charts move the horizon only from 18 months to 15. The three-part conjunctive criterion (realism, convergent validity, the headline sensitivity estimate) compounds the risk. Payoff is near the top of the portfolio because the deliverable is a standing scheme plus the corrected bound p ≤ 3/(n·s) — arithmetic and institutional machinery that outlive model turnover, and that track drift in s rather than being invalidated by it.",
        "Reference class: covert proficiency testing of professionals, which has only ever run under a regulator or accrediting body — a small external team landing institutional hosting, an ethics route with no sector precedent, 40-60 consenting assurers and 200 real episodes inside 15 months is a compound low-probability event, and the convergent-validity gate between planted-error and capture-recapture estimates of s is an independent empirical coin flip. Agents can have the whole instrument built before recruitment even closes, which is why this compresses least in the portfolio: the critical path is entirely human and institutional. Payoff is the highest here because s is the multiplier every sampled certification silently assumes, and p <= 3/(n*s), the item bank and the control charts are arithmetic and mechanism, not a model-generation measurement."
      ],
      "essay_rank": 16,
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    }
  ],
  "theoreticalCore": [
    {
      "id": "T3",
      "title": "Anytime-Valid Auditing of Evidence Censuses: Stopping Certificates and Graded Release",
      "description": "Extends e-value/e-process adaptive auditing (Zhou et al. 2026) from the streaming setting to documentary censuses: finite populations, stratified and clustered corpora, fallible human labels, repeated audits across document versions, adversaries who learn the sampling policy, graded error severity, several claims per audited item, and simultaneous guarantees across subgroups. The central output is a single object — the audit stopping certificate — asserting that under the declared population, loss function, sampling history and policy, the probability of falsely certifying an error rate below p is at most alpha, and that this holds under optional stopping and adaptive targeting rather than only at a pre-planned sample size. On top of it sits a graded release ladder (certified / certified-within-conditions / human-review / unresolved / rejected) built from conformal selective prediction with general risk control (Bai & Jin 2026), so that an audit can decline to certify part of a corpus without discarding the rest. Delivered as a theorem set, a machine-checked core, an open implementation, and a specified certificate format.",
      "resolution_criterion": "PASS requires all five. (1) Theory: prove sup over all stopping times of P(certify error rate ≤ p | true rate > p) ≤ alpha for the finite-population stratified/clustered construction with bounded label noise, with the supermartingale property and the optional-stopping/Ville step machine-checked in Lean or Isabelle, zero sorries, and the label-noise robustness stated as an explicit bounded-contamination condition rather than an assumption of correct labels. (2) Monte Carlo validity: at least 10^5 runs per configuration across at least 20 pre-registered configurations (population size, stratification, cluster structure, loss severity, adversary type), with realised false-certification rate satisfying a Clopper-Pearson upper bound ≤ alpha + 0.005 at every single configuration — one configuration exceeding it fails the gate. (3) Adversarial evasion: against an adaptive adversary given full knowledge of the sampling policy and a budget to place errors so as to maximise false certification, realised false-certification rate still ≤ alpha + 0.005; and against a misspecification adversary that violates the declared label-noise bound by 50%, the degradation is bounded and quantified rather than unbounded. (4) Power, so validity is not achieved by never stopping: median stopping time ≤ 1.5x an oracle fixed-sample benchmark at matched alpha and matched power, across all 20 configurations. (5) Graded release: distribution-free risk control proved for each of the five states; on at least three real corpora (differing in language coverage and document type) empirical risk for every state is within its guarantee, simultaneous validity holds across at least 50 pre-specified subgroups with FWER ≤ alpha, and the human-review state captures at least 80% of planted severe errors that the certified state would otherwise have absorbed.",
      "horizon_months": 12,
      "cost_band": "Medium: $290k–$430k. About 2 human FTE-years (one sequential-testing/e-process statistician, one audit-practice lead with real census-auditing experience), $40k for gold-standard human labelling of the three validation corpora including a double-labelled subsample for the noise bound, $30k simulation compute, $25k external statistical review. (converted: ≈£230k–£340k at $1.27/£; the essay quotes the midpoint)",
      "crux": "The certificate is conditional on a declared population and loss function, and that declaration is exactly what an interested party will manipulate. An audit that is anytime-valid over a corpus boundary drawn to exclude the inconvenient material is technically impeccable and substantively empty — the guarantee is real but the object it quantifies over was chosen by the party being audited. Formalisation cannot close this; it can only make the declaration explicit and machine-readable so that the manipulation is visible, and whether that is enough is a judgement about institutions, not about martingales. Second risk: fallible human labels break the conditions the e-process needs unless label error is itself bounded, and bounding it requires a gold subsample whose cost may exceed the audit it is meant to license, which would make the whole construction economically inert. Third: power. Anytime validity is cheap to achieve by stopping late; gate 4 is the one most likely to fail quietly.",
      "agent_execution": "High, roughly 80%. Agents do: reconstruction of the e-process, betting-martingale, finite-population sampling and survey-audit literatures (including the substantial pre-existing statistical-audit and election-audit work on risk-limiting procedures, which is the closest prior art and must be positioned against, not rediscovered); derivation of candidate e-processes for the stratified/clustered finite-population case; the Lean/Isabelle formalisation of the supermartingale and optional-stopping steps; the full simulation harness and the adaptive-adversary implementation; the conformal selective-prediction layer and its risk-control proofs; and cross-model adversarial review aimed at finding a stopping rule that breaks validity. Human critical path: specifying the loss function and severity grading so they mean what auditors mean, designing the declaration format so boundary manipulation is detectable, obtaining and curating the three real corpora, and independent statistical review by someone outside the e-value community.",
      "durability": "High. A proved anytime-valid stopping certificate for finite populations under bounded label noise is a permanent result, and the certificate format is the kind of artefact that outlives the systems it audits. Two decay channels: the adversary model is specified against evasion strategies imaginable now, and a genuinely novel evasion class would require an extension (though not a retraction); and the graded-release calibration on the three corpora is a measurement that ages. The theory itself is model-agnostic, which is what makes it durable — nothing in it depends on which systems produced the documents being audited.",
      "judge_probs": [
        0.33,
        0.21,
        0.38
      ],
      "judge_payoffs": [
        0.61,
        0.64,
        0.55
      ],
      "p_med": 0.33,
      "pay_med": 0.61,
      "rationales": [
        "Probability: highest of the five, because most of the machinery already exists. Risk-limiting election audits, betting martingales and sampling-without-replacement e-processes cover finite populations and stratification directly; the genuinely new work is bounded label contamination, the five-state graded release layer, and simultaneous subgroup validity. That is extension rather than invention, and the Lean supermartingale/Ville step is a well-trodden formalisation target. The real hazard is the collision between gate 3 and gate 4: contamination-robust e-processes buy their robustness with conservatism, and median stopping within 1.5x an oracle fixed-sample benchmark at every one of 20 configurations is a demanding power requirement once a 50%-misspecification adversary must also be survived. Gate 4 is correctly identified as the one that fails quietly. Second schedule risk: three real corpora spanning languages and document types, with a double-labelled gold subsample sized to certify the noise bound, inside 12 months on a $40k labelling budget — the crux's own observation that the gold subsample may cost more than the audit it licenses is a live economic threat to gate 5. Payoff: solid but the most incremental of the set. Anytime-valid certificates for documentary censuses are adoptable — auditors already understand risk-limiting procedures, so the adoption path is extension of an accepted practice rather than a new framework, which is a real advantage. Against that: substantial overlap with existing RLA literature caps the citable novelty, and the honest limitation is structural and unfixable by theorem — a certificate quantified over a population boundary drawn by the audited party is impeccable and empty. Making that declaration machine-readable is worth something, but it is an institutional hedge, not a result. The theory is model-agnostic, so it does not decay with model generations; the corpus calibration does.",
        "Substantial directly reusable prior art — risk-limiting election audits, ALPHA/RiLACS betting martingales, finite-population sequential testing — which the spec correctly flags must be positioned against rather than rediscovered. That prior art raises probability relative to T1 but also caps the novelty. Gate 1 is achievable (Ville plus optional stopping is formalisable in mathlib-adjacent territory, and bounded contamination is the right honest framing). Gate 2 follows from a correct construction. The two real risks are gate 4 and gate 5. Gate 4 is the quiet-failure gate the spec names: median stopping within 1.5x oracle fixed-sample in EVERY one of 20 configurations, where the label-noise robustness margin and cluster-level variance inflation both buy validity by stopping later — one adverse configuration out of twenty fails the whole gate, and the conservatism needed for gate 3's misspecification adversary directly costs power here. Gate 5 stacks distribution-free risk control for five states, three corpora differing in language coverage, FWER<=alpha across 50 subgroups, and 80% severe-error capture; simultaneous validity over 50 subgroups while retaining useful power is the tightest constraint in the whole spec, and the $40k gold double-labelled subsample is a 12-month schedule risk that agents cannot compress because it is human labelling. The crux is also the most institutionally unfixable of the five: a certificate quantified over a population the audited party declared is impeccable and empty, and no gate tests for boundary manipulation. Payoff solid but mid-pack: the stopping certificate format is a durable artefact and the theory is model-agnostic, but the marginal advance over existing risk-limiting-audit machinery is an extension rather than a new object, and the graded-release calibration is a measurement that ages.",
        "Technically the most standing-on-shoulders of the statistical three: risk-limiting election audits already supply anytime-valid finite-population betting martingales for sampling without replacement, and conformal risk control supplies the graded layer, so gates 1-3 are extension rather than invention and the Ville/optional-stopping formalisation has partial mathlib scaffolding. That cuts both ways — it raises probability and shrinks the novelty payoff. Gate 4 is the quiet killer the spec correctly identifies: median stopping time within 1.5x an oracle fixed-sample benchmark at matched alpha AND matched power, uniformly across all 20 pre-registered configurations, is demanding once bounded-contamination robustness is bolted on, because contamination-robust e-processes pay their premium exactly in the adversarial and clustered configurations. Gate 5 compounds it — FWER control across 50 pre-specified subgroups costs power on the same corpora where gate 4 needs efficiency, and it requires three real multi-language corpora with gold labels plus a double-labelled subsample, which is the largest irreducible human-and-money item in the spec and the most likely source of a 12-month overrun. On payoff, the certificate format and the finite-population theorem are durable and reusable, but the crux concedes the operative failure: anytime-validity over a corpus boundary drawn by the audited party is impeccable and empty, and the criterion cannot detect that failure — it can only make the declaration machine-readable and hope institutions notice. Whether the gold-subsample cost exceeds the audit it licenses is also unresolved, and if it does the construction is economically inert regardless of every gate passing."
      ],
      "essay_rank": "T1",
      "workflow_id": "T3",
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "T2",
      "title": "Dependence-Robust Verification Portfolios: Partial-Identification Bounds and Assurance Ceilings",
      "description": "Verification stacks are currently justified by an implicit independence assumption that measured error correlation across 350+ models (Kim, Garg, Peng & Garg 2025) refutes. This project replaces the single pairwise-correlation summary with a theory of state-dependent common-cause failure across verifiers that share training data, benchmarks, retrieval corpora, architecture, serving infrastructure, prompt scaffolds and human reviewers. The central deliverables are sharp partial-identification bounds on P(all verifiers wrong) from marginal accuracies plus limited agreement data — with the upper bound, not a point estimate, governing certification — a verifier-diversity certificate format that records which common causes were and were not broken, a minimax allocation rule for spending an assurance budget across model verifiers, deterministic recomputation, formal checkers, fresh sources, human experts and direct measurement, and impossibility theorems stating when no voting or debate procedure can exceed a given assurance level. The output is the arithmetic a certification body needs to stop treating five correlated verifiers as five verifiers.",
      "resolution_criterion": "PASS requires all five. (1) Sharpness: for the identified-set characterisation given K marginals plus a specified agreement functional, prove both directions — every distribution in the set satisfies the bound, and the bound is attained by an explicitly constructed joint distribution — with the LP/dual characterisation verified numerically to relative error below 1e-9 on at least 10,000 random instances, and the core supremum argument machine-checked in Lean or Isabelle with zero sorries. (2) Simulation cross-validation: across at least 10^6 synthetic dependence structures spanning latent-common-cause, mixture and adversarially-coupled generators, zero violations of the derived upper bound (any single violation is a proof bug and fails the gate until resolved). (3) Empirical coverage on real correlated-error data: using the 350+ model error dataset with held-out verifier subsets, the upper bound covers the realised joint-failure rate in at least 95% of 1,000 bootstrap resamples at nominal 95%. (4) Non-vacuity, the make-or-break gate: the upper bound must be strictly below the trivial min-of-marginals bound by at least a factor of two in at least 60% of realistic parameter regimes, where 'realistic' is pre-registered from the empirical correlation distribution before the bounds are derived — a bound that collapses to triviality exactly when verifiers are correlated has failed. (5) Allocation and ceiling: the minimax portfolio rule achieves lower worst-case P(all wrong) than both equal-weight and accuracy-greedy allocation at matched budget on the empirical data, with the gap significant at p<0.01 under a pre-registered test; and at least one impossibility theorem is proved with a tight constant plus a matching construction attaining the ceiling, establishing that the limit is real and not an artefact of a loose proof.",
      "horizon_months": 12,
      "cost_band": "Medium: $260k–$400k. About 1.8 human FTE-years (one partial-identification/econometric theorist, one empirical ML lead), $35k compute for the 10^6-structure simulation grid and bootstrap work, $30k external review by a partial-identification specialist outside the AI-assurance community, $20k for access/replication of correlated-error measurement infrastructure. (converted: ≈£205k–£315k at $1.27/£; the essay quotes the midpoint)",
      "crux": "Whether the observable data identifies enough of the dependence structure to say anything useful. Marginal accuracies plus pairwise agreement may pin the joint-failure probability only within an interval whose upper end is min(marginals) — precisely in the high-common-cause regime that motivates the project. The bounds would then be correct, sharp, and operationally worthless, which is why gate 4 is pre-registered rather than chosen after seeing the results. The deeper problem is identification of shared-training-data effects without interventional variation: two verifiers' overlap in pretraining corpora is not observable to the certifier, so the diversity certificate may end up recording declared rather than actual independence — the same trust-relocation failure that threatens T1. A third risk is construct drift: 'all verifiers wrong' is well defined only relative to a ground truth that, for causal and normative claims, is contested.",
      "agent_execution": "High, roughly 80%. Agents do: reconstruction of the partial-identification, Fréchet-bound, common-cause-failure (nuclear/avionics reliability) and ensemble-diversity literatures, including the pre-2010 software-reliability work on correlated failures that the current ML literature largely rediscovers; derivation and algebraic verification of candidate bounds; LP formulation and dual verification; the full simulation grid; the Lean/Isabelle formalisation of the supremum argument; the minimax solver; and adversarial construction of dependence structures designed to break each candidate bound. Human critical path: choosing which agreement functionals are actually observable to a certifier (this choice determines whether gate 4 can be passed at all), judging construct validity of the common-cause taxonomy, and independent review by someone who does partial identification for a living and has no stake in the assurance framing.",
      "durability": "High but not maximal. Sharp bounds and impossibility theorems are permanent — a proved ceiling on what voting can achieve under common-cause error stays true. The perishable component is the empirical calibration: the specific correlation structure across today's 350+ models is a measurement of current systems and will decay within two to three years as architectures and training corpora shift. The design hedge is to keep the theory parameterised by observable functionals so re-calibration is a data refresh rather than a re-derivation, and to treat the empirical dataset as validation of the bounds rather than as an input to them.",
      "judge_probs": [
        0.28,
        0.32,
        0.4
      ],
      "judge_payoffs": [
        0.7,
        0.72,
        0.72
      ],
      "p_med": 0.32,
      "pay_med": 0.72,
      "rationales": [
        "Probability: the mathematics is classical and well-mapped — Fréchet-Hoeffding, Boole-Bonferroni, LP duality, common-cause failure work from nuclear/avionics reliability — which makes the sharpness proof, the dual characterisation, the 1e-9 numerics, the 10^6-structure grid and a Lean supremum argument all high-confidence agent work. Two gates carry the risk. Gate 4 is the pre-registered non-vacuity test and is roughly a coin flip on substance: with K marginals plus pairwise agreement the sharp bound is essentially the min over pairs of joint-error probability, which beats min-of-marginals by a healthy factor at moderate correlation and collapses toward it exactly in the high-common-cause regime that motivates the project. Pre-registering 'realistic' from the empirical correlation distribution before deriving bounds is methodologically honest and therefore genuinely dangerous. A second, under-flagged risk sits in gate 3: the bound must be estimated from finite agreement data, and 95% bootstrap coverage at nominal 95% is not automatic for a plug-in upper bound unless finite-sample inflation is built in — the spec does not obviously fund that. Gate 5's tight-constant impossibility plus matching construction is the kind of result that either falls out in a month or not at all. Payoff: the most immediately usable output of the five. 'Five correlated verifiers are not five verifiers' is arithmetic that drops into existing certification practice with no new standard, no schema adoption, no committee — the lowest adoption friction here, which matters more than durability rankings suggest. Impossibility ceilings on voting and debate under common-cause error are permanent and are the kind of result that constrains lab-side assurance claims directly. Discounted because the empirical calibration on today's 350+ models is a two-to-three-year artefact, and because the diversity certificate records declared rather than observable training-corpus overlap — the same trust relocation as T1, with no schema-level hedge.",
        "Highest probability of the five because it is the closest to known mathematics: Frechet-Hoeffding bounds, LP duality, and the pre-2010 common-cause-failure reliability literature give agents a well-mapped derivation path, and the Lean/Isabelle obligation is a single supremum argument rather than a six-theorem package. Gates 1-3 are near-mechanical once the bound is right (sharpness both directions with an attaining construction is a standard LP-dual exercise; the 10^6 simulation grid and bootstrap coverage follow automatically from a correct proof — they are bug detectors, not independent hurdles). The project therefore turns almost entirely on gate 4, which is honestly pre-registered and roughly a coin flip: from marginals plus an agreement functional the identified set's upper end can collapse toward min-of-marginals precisely in the high-common-cause regime, but for K=5 verifiers at realistic accuracies, pairwise agreement does rule out heavily nested error sets, so a 2x improvement in 60% of regimes is plausible rather than fanciful. Gate 5's minimax-beats-baselines-at-p<0.01 is the secondary risk (accuracy-greedy is a stronger baseline than it looks when correlations are moderate). The identification-of-shared-pretraining problem is real but does not block the gates — it degrades what the diversity certificate means, which is a payoff problem, not a probability problem. Payoff high but below T1: sharp bounds and a proved voting ceiling are permanent and immediately actionable arithmetic for a certification body, and the failure mode is graceful (a correct-but-loose bound still publishes). Discounted for the perishable empirical calibration on today's 350+ models, for the certificate recording declared rather than actual independence, and because 'all verifiers wrong' needs a ground truth that is contested exactly for the causal and normative claims that matter most.",
        "The most decision-relevant of the five and the one whose make-or-break gate is honestly pre-registered. Gates 2 and 3 are close to automatic if the theory is right — an upper bound will cover a realised rate in a bootstrap almost by construction, and zero violations across 10^6 synthetic structures tests the implementation, not the mathematics. Gate 4 is less lethal than the crux fears: once pairwise joint-error rates are among the observable functionals, P(all wrong) <= min over pairs of P(both wrong) already beats min-of-marginals by a factor of two wherever pairwise co-failure sits meaningfully below the marginals, which the empirical correlation data suggests is common. The genuine risk is that pre-registering 'realistic regimes' from the empirical distribution loads the sample toward the high-common-cause tail, exactly where the sharp bound does collapse to triviality; call that gate 55-65%. The likelier failure is gate 1's Lean/Isabelle requirement: LP duality and a sharpness supremum argument are far more painful to formalise than a discrete type system, mathlib support is thin, and 'zero sorries' on the core supremum is where a 12-month schedule with 1.8 FTE breaks. Gate 5's tight impossibility constant with a matching construction is usually feasible when pursued jointly. Payoff highest of the five: a sharp ceiling on what voting or debate can achieve under common-cause error is used by anyone who needs the number, with no ecosystem adoption required — unlike a schema, arithmetic needs no interoperating second implementation. It also directly refutes a scaling story (add more verifiers) that assurance stacks currently rely on. Discounted for the same trust-relocation flaw as T1 — shared-pretraining overlap is not observable to a certifier, so the diversity certificate may record declared independence — and for the empirical calibration decaying in two to three years, though the design deliberately keeps that a data refresh."
      ],
      "essay_rank": "T2",
      "workflow_id": "T2",
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "T4",
      "title": "Temporal Assurance and Minimal Revalidation: An Invalidation Calculus for Versioned Certificates",
      "description": "Treats assurance certificates as versioned conditional objects rather than one-off approvals, which is the direction NIST's monitoring-fragmentation findings and its guardrail-impossibility result both point. The project builds bitemporal claim graphs separating valid time from transaction time, an invalidation calculus that propagates source corrections, statutory changes, statistical-series revisions and silent model replacement through the dependency structure, a minimal-revalidation theorem identifying the smallest set of claims and tests that must be re-run to restore an assurance grade after a change, and sequential change-point monitoring with controlled false alarms over indefinite operation. The point is to make continuous assurance affordable: if every upstream correction forces full recomputation, continuous monitoring is a slogan, and if propagation is unsound, stale certificates survive corrections silently. Delivered as a semantics with soundness and completeness results, a machine-checked propagation core, a complexity characterisation with an optimal algorithm for the tractable subclass, and differential-tested reference code.",
      "resolution_criterion": "PASS requires all five. (1) Semantics: prove soundness (no certificate survives a change on which it depends) and completeness (no certificate is invalidated that does not depend on the change) for the bitemporal invalidation calculus, with the propagation soundness theorem machine-checked in Lean or Isabelle, zero sorries, and the declaration-dependence of completeness stated as an explicit hypothesis rather than buried. (2) Minimality with teeth: prove that the computed revalidation set is minimum-cost under a stated cost model for the tractable subclass, give the complexity classification of the general problem including a hardness proof, and give an approximation algorithm with a proved ratio; the ratio must be verified empirically to hold on 10^4 instances with no violations. (3) Differential testing: on at least 10^4 randomly generated claim graphs up to 200 nodes, the implementation's invalidation set agrees exactly with exhaustive brute-force recomputation in 100% of cases — zero tolerance, since a single missed propagation is a stale certificate. (4) Real revision histories: on at least three genuine change episodes (a national statistical series revision, a statutory change, and a model-version replacement), recall of truly-affected claims is 100% against ground truth established by full recomputation, and the revalidation set is at most 40% the size of full recomputation — if minimal revalidation is not materially smaller than doing everything again, the theorem is correct and the project has failed its purpose. (5) Monitoring: proved false-alarm control over unbounded operation, with simulated ARL0 within 10% of the pre-registered target across 10^5 indefinite-operation runs and detection delay within 1.5x an oracle change-point benchmark.",
      "horizon_months": 15,
      "cost_band": "Medium-high: $340k–$500k. Around 2.3 human FTE-years (one formal-methods/database-theory lead, one lead with statistical-revision and regulatory-change domain knowledge), $30k compute for differential testing and monitoring simulation, $35k external review split between a bitemporal-database specialist and a change-point statistician, $25k for curating the three real revision histories including provenance reconstruction. (converted: ≈£270k–£395k at $1.27/£; the essay quotes the midpoint)",
      "crux": "Silent model replacement is the motivating case and is the one the calculus structurally cannot detect. Propagation only moves changes it has been told about, so completeness holds conditional on declaration discipline, and a provider who swaps a model without emitting a version event leaves every downstream certificate looking fresh. The calculus can make this failure mode nameable and auditable but cannot make it detectable from inside — that requires an external attestation channel, which is an institutional artefact, not a theorem. Second risk, and it is the one gate 4 is built to catch: real dependency graphs may be dense enough that almost everything depends on almost everything, in which case the minimal set is nearly the full set and the theorem is true but economically pointless. Third: the cost model that minimality is defined against is a modelling choice, and minimising the wrong cost produces confidently optimal revalidation plans that miss what actually matters.",
      "agent_execution": "High, roughly 80%. Agents do: reconstruction of bitemporal database theory, truth maintenance and belief-revision systems, incremental/self-adjusting computation, build-system minimal-rebuild theory and the change-point-detection literature — this project has more directly reusable prior art than the other four, and the main agent value is finding it rather than reinventing it; formulation of the propagation semantics and candidate minimality theorems; the hardness proof and approximation analysis; the Lean/Isabelle formalisation; the differential-testing harness and brute-force oracle; the monitoring simulation; and adversarial search for propagation counterexamples. Human critical path: choosing the cost model, deciding what counts as a declarable change event (this choice fixes the completeness hypothesis and therefore the project's honesty), reconstructing the three real revision histories, and independent review.",
      "durability": "High. An invalidation calculus with proved soundness and a minimal-revalidation result is permanent and reusable well beyond AI assurance — the same machinery applies to statistical production, financial restatement and regulatory-change management, which broadens the survival base considerably. The perishable parts are the cost model's parameters and the specific change-event taxonomy, both of which are refreshable. The most likely long-run outcome is that the theorems are cited and the specific implementation is replaced, which is the normal and acceptable fate of infrastructure theory.",
      "judge_probs": [
        0.3,
        0.23,
        0.5
      ],
      "judge_payoffs": [
        0.66,
        0.7,
        0.5
      ],
      "p_med": 0.3,
      "pay_med": 0.66,
      "rationales": [
        "Probability: the most reusable prior art of the five (bitemporal databases, truth maintenance, self-adjusting computation, minimal-rebuild theory, change-point detection), and the spec correctly identifies that the agent value is finding it rather than reinventing it. Soundness/completeness for a propagation semantics, NP-hardness for the general minimal-revalidation problem, an approximation ratio, differential testing against a brute-force oracle on 10^4 graphs, and CUSUM-style ARL0 control are all high-confidence deliverables — gate 3's 100% agreement is engineering discipline, and gate 4's 100% recall is close to implied by gate 3. The project turns on the ≤40% size requirement in gate 4, which is an empirical fact about real dependency graphs, not something the team controls: a headline statistical-series revision touches nearly everything downstream, and if real graphs are dense the theorem is correct and the purpose is defeated. Some selection latitude in choosing three episodes helps but also invites a weaker result. Reconstructing provenance for three genuine change histories on $25k is the schedule risk in a 15-month window. The cost model is a free modelling choice, and minimising the wrong cost yields confidently optimal plans that miss what matters — unfalsifiable within the project. Payoff: broader survival base than the AI-assurance framing suggests, since invalidation and minimal revalidation apply to financial restatement, statistical production and regulatory-change management — three constituencies that outlive any model generation, which materially raises the durability floor. Continuous-assurance regimes (ISO 42001-style monitoring) need this arithmetic or their monitoring claims are slogans. Discounted on two counts: the motivating case, silent model replacement, is structurally undetectable from inside the calculus and requires an external attestation channel that no theorem supplies; and the likeliest outcome is theorems cited, implementation replaced — normal for infrastructure theory but still a discount against the 'infrastructure' claim.",
        "The most favourable prior-art-to-novelty ratio of the four theory projects: bitemporal databases, truth-maintenance systems, self-adjusting computation and minimal-rebuild theory supply most of the machinery, and the spec candidly says the agent value is finding it rather than inventing it. That plus 15 months makes gates 1-3 comparatively safe — soundness/completeness with declaration-dependence as an explicit hypothesis is an honest and provable formulation; gate 3's 100% agreement with a brute-force oracle on 10^4 graphs is an implementation-debugging task with an oracle available, which is exactly what agents are best at. Two things hold probability down. Gate 4: 100% recall is easy (a superset always recalls), but <=40% of full recomputation on all three real episodes runs straight into the density crux — a national statistical series revision plausibly touches nearly every downstream claim, and reconstructing three genuine change episodes with provenance and a full-recomputation ground truth is the kind of archival work that consumes the schedule. Gate 5 carries a mild internal tension: proved false-alarm control over UNBOUNDED operation (alpha spent once, forever) is not obviously reconcilable with hitting a finite ARL0 target within 10%, and squaring those may require two different monitors, one of which then misses the delay bound. The motivating case — silent model replacement — is structurally undetectable by the calculus, which the spec admits; that is a payoff ceiling, not a gate failure. Payoff is second-highest: an invalidation calculus with a minimality result transfers cleanly to statistical production, financial restatement and regulatory-change management, which widens the survival base well beyond AI assurance, and the cost model plus event taxonomy are refreshable parameters rather than load-bearing assumptions. Discounted because the likely long-run reading is 'careful synthesis of known incremental-computation theory for a new domain' and because minimising a mis-specified cost yields confidently optimal plans that miss what matters.",
        "Highest probability of the five, because most of the machinery already exists and one gate is purely mechanical. Soundness and completeness for bitemporal invalidation is essentially reachability over a dependency graph — shallow to prove and to formalise; minimum-cost revalidation for a tractable subclass plus a hardness classification plus a proved approximation ratio is standard combinatorial optimisation with dense reusable prior art in truth-maintenance, self-adjusting computation and minimal-rebuild theory; gate 3's differential test against a brute-force oracle on 10^4 graphs under 200 nodes is a harness, and it passes or exposes a bug in week one. Gate 5's ARL0 within 10% is routine change-point calibration. The soft spot is gate 4, and it is soft in a way that inflates rather than deflates the pass rate: no deployed corpus of assurance certificates with real dependency graphs exists, so the team must reconstruct three change episodes and author the graph structure itself — and graph density is precisely what determines whether the revalidation set clears the 40%-of-full-recomputation threshold. A team that builds the corpus controls the number the gate tests. That is the same self-authored-evidence problem as T1's gate 5, and here it sits on the gate that decides whether the whole minimality result is economically meaningful. Payoff is the reason not to rank this higher despite the probability: the motivating case, silent model replacement, is the one the calculus structurally cannot detect, since propagation only moves declared change events — so the deliverable names a failure mode it cannot close and needs an external attestation channel that is institutional, not mathematical. Offsetting that, the machinery transfers cleanly to statistical production and financial restatement, which broadens the survival base beyond AI assurance and makes citation likely even if this implementation is replaced."
      ],
      "essay_rank": "T3",
      "workflow_id": "T4",
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "T1",
      "title": "Compositional Assurance Calculus: A Typed Algebra for Evidential Information",
      "description": "Build the missing formal core of evidence-based assurance: a type system and algebra over typed claim graphs whose nodes carry an evidential kind (source quotation, measurement, accounting identity, deterministic computation, statistical estimate, causal/identification claim, forecast, recommendation, normative judgement, authorisation) together with provenance, verification artefacts, declared dependence, temporal validity, authority, and open defeaters. The deliverable is a six-theorem package — sound composition, no upcasting, no double-counting, defeater monotonicity, authority separation, honest abstention — proved on paper and machine-checked in Lean, plus a machine-readable assurance-object schema and a reference checker that returns certified / conditional / abstain / rejected and never anything else. The work extends Assurance 2.0 (Bloomfield & Rushby) from an argumentation methodology into a calculus with proof obligations, and takes the tool-attested kernel-proof discipline (Ren 2026) as the standard for the trusted core. The intended output is infrastructure others build on: a small kernel a regulator or standards body can read in an afternoon, and a schema two independent teams can implement without talking to each other.",
      "resolution_criterion": "PASS requires all five gates. (1) Lean formalisation of all six theorems compiles against a pinned mathlib with zero `sorry` and zero `axiom` declarations beyond Lean's three standard axioms, verified by `#print axioms` on every top-level theorem in CI; the trusted kernel (definitions plus checker specification) is under 1,500 significant lines and is reviewed line-by-line by two external proof engineers who each submit a signed statement of what the theorems do and do not establish. (2) Necessity of hypotheses: for each of the six theorems, an SMT/model-finder-produced countermodel showing that dropping any one stated hypothesis makes the theorem false, machine-checked, so the axiom set is minimal rather than merely sufficient. (3) Independent-implementation interoperability: a second team, given only the published schema and English specification and firewalled from the reference checker's source, builds a checker; the two agree on the verdict for at least 99% of 2,000 held-out assurance objects, every disagreement is adjudicated to a specific specification ambiguity, the specification is amended, and a re-run reaches 100% agreement on the same corpus. (4) Adversarial discrimination: on a suite of at least 300 planted objects, the checker rejects or abstains on 100% of objects that violate a theorem's precondition (zero false certifications — this gate is absolute, one false certification fails the project) and certifies at least 95% of valid objects, so soundness is not bought with universal abstention. (5) Non-vacuity: on a corpus of at least 50 real assurance cases drawn from published audit, statistical and regulatory documentation, the checker returns `certified` or `conditional` on at least 40% — a calculus that abstains on nearly all real evidence has failed even if every theorem holds.",
      "horizon_months": 14,
      "cost_band": "Medium-high: $380k–$550k. Roughly 2.5 human FTE-years (one type theorist, one assurance-domain lead, part-time standards liaison), $50k external kernel review by two independent proof engineers, $45k for the firewalled second-implementation team, $25k model/compute for proof search and adversarial case synthesis. (converted: ≈£300k–£435k at $1.27/£; the essay quotes the midpoint)",
      "crux": "That the ten-way evidential type lattice carves the problem the way real disputes do. If practitioners' actual disagreements sit inside a type rather than between types — two analysts contesting whether an identification assumption holds, both agreeing it is a causal claim — the calculus is a well-formed answer to a question nobody asked, and the theorems are true but idle. Second and nearly as dangerous: no-double-counting requires a dependence model, and if dependence must be declared honestly by the submitter, the kernel has relocated trust from the argument to the metadata rather than reducing it — the checker becomes sound-given-honest-annotation, which is the assumption that fails in practice. Third: the abstention design is what makes soundness cheap; gate 5 exists because the most likely failure is a checker that is provably correct and returns `abstain` on everything real.",
      "agent_execution": "High mechanical share, roughly 75–80%. Agents do: reconstruction of the Assurance 2.0, Toulmin-argumentation, defeasible-logic and provenance-model literatures into a comparison table of primitives; enumeration and stress-testing of candidate axiomatisations (expect 10–20 discarded before one survives); the full Lean development including proof repair against mathlib churn; SMT/model-finder search for the gate-2 countermodels; schema design and the reference checker; synthesis of the 300+ adversarial objects with planting recipes; and cross-model adversarial review of each axiom set framed to refute. Human critical path, and it is the whole risk: choosing the type lattice and the dependence primitive, judging whether the formalisation preserves what practitioners mean, the independent kernel review, and engagement with OMG SACM / ISO-IEC 42001 style bodies where a schema either becomes a standard or becomes a paper.",
      "durability": "Highest of the five. A sound calculus with a machine-checked kernel and a published schema does not decay — the no-upcasting and authority-separation results are true about evidence, not about any model generation, and remain citable and reusable if every current system is replaced. The realistic decay path is not falsification but disuse: calculi that nobody implements survive as literature rather than infrastructure. Gate 3 (two interoperating implementations) and standards-body contact are the hedges that convert a durable theorem into durable infrastructure.",
      "judge_probs": [
        0.22,
        0.13,
        0.44
      ],
      "judge_payoffs": [
        0.81,
        0.88,
        0.6
      ],
      "p_med": 0.22,
      "pay_med": 0.81,
      "rationales": [
        "Probability: the proof burden is the least of it — six composition/monotonicity/separation theorems over a typed algebra are definitional-depth results, and Lean development plus mathlib proof repair plus SMT countermodel search for gate 2 is exactly what agents now do reliably. The conjunction is what kills it. Gate 5 (non-vacuity on 50 real audit/regulatory documents) is the binding constraint and the spec knows it: published audit and statistical documentation does not carry declared dependence, verification artefacts or temporal validity, so encoding is a human re-authoring step, and the honest version of that step is where abstention swallows the corpus. The 'conditional' verdict is a real escape valve that makes 40% reachable, but only if the type lattice and dependence primitive were chosen right the first time, and the crux is precisely that they may not be. Gate 4's absolute zero-false-certification against gate 4's 95% certify-valid is a genuine two-sided squeeze, though planted-recipe objects are more tractable than real ones. Gate 3 is survivable because the amend-and-rerun loop is written into the criterion. 14 months with 2.5 FTE is tight for kernel review by two external proof engineers plus a firewalled second implementation plus standards liaison. Payoff: highest of the five conditional on delivery. A <1,500-line machine-checked kernel with a schema and two interoperating checkers is the one artefact here that other programmes (T2's diversity certificate, T3's stopping certificate, T4's invalidation events) would plug into, which is a rare structural advantage — it makes the others cheaper rather than competing with them. No-upcasting and authority-separation are claims about evidence, not about model generations, so frontier-lab progress turbo-charges rather than moots them. Discounted for the real adoption hazard: OMG SACM is the cautionary precedent (formally reasonable, barely implemented), and W3C VC/PROV and ISO 42001 occupy adjacent ground with more institutional momentum. Trust relocation into submitter-declared dependence is the honest cap on what the kernel achieves, but naming and machine-checking that boundary is itself durable.",
        "Five conjunctive gates, each with an absolute clause, plus the hardest version of the doctrine's central risk: this is a NEW calculus, so the gap between 'consistent formal system' and 'right formal system' is the whole project. Agents will almost certainly deliver a sorry-free Lean development of six theorems and a <1,500-line kernel — that part is now routine-ish. The failures cluster elsewhere. Gate 2 requires a countermodel for dropping ANY hypothesis of ANY of six theorems; in practice several hypotheses will be well-formedness conditions whose removal makes the statement ill-typed rather than false, which is unfalsifiable-by-countermodel and forces axiom-set surgery late. Gate 5 is the killer and the spec knows it: real audit/statistical/regulatory documents do not carry declared dependence, verification artefacts or authority metadata, so encoding 50 real cases is a large manual effort whose output is exactly the annotation the checker needs to abstain on. 'Conditional' being a passing verdict is a real concession that lifts this from near-zero. Gate 3 also imports two independent human organisations (two proof engineers with signed statements, a firewalled second team) — coordination risk that pure agent throughput cannot compress. And the crux is fatal if true: if practitioner disputes live inside a type rather than between types, the ten-way lattice is a well-formed answer to no question, and no gate detects that. Payoff conditional on delivery is the highest of the five: a machine-checked kernel plus a schema with two interoperating implementations plus standards-body contact is the one deliverable here that is genuinely assurance INFRASTRUCTURE rather than a result about it; no-upcasting and authority-separation are facts about evidence, not about model generations. Discounted slightly for the disuse path (calculi nobody implements) and for the trust-relocation objection, which would survive even a full PASS.",
        "Binding gates are 3, 4 and 5, and two of them have escape hatches that make PASS easier than it looks while making PASS mean less. The six theorems are bookkeeping soundness results about a type system the team itself designs — closer to a small type-safety development than to deep mathematics — so the Lean work under a 1,500-line kernel cap is credible in 14 months with agent-driven proof repair, and gate 2's countermodels are mechanical model-finder work. Gate 3 permits amend-the-spec-and-re-run to 100% on the SAME corpus, which converts a genuine interoperability test into an iterated debugging loop; gate 5 never says who encodes the 50 real audit documents into assurance objects, and if the team encodes them, a 40% certified/conditional rate is close to self-fulfilling. That is the smuggling: gate 5 was inserted to catch universal abstention, but it cannot distinguish a calculus that fits real evidence from a team that fits real evidence to the calculus. Gate 4's absolute zero-false-certification clause is the one real hazard, since a single planted object slipping through fails the project outright, and 300 adversarial objects against a permissive-enough checker to clear gate 5 is a genuine tension. Human coordination load (two signing external proof engineers, a firewalled second team, 10-20 discarded axiomatisations before one survives) is the schedule risk rather than the proofs. Payoff is high but not top: this is the most infrastructural of the five and a kernel plus schema plus two implementations is a durable asset, yet the crux is right that no gate detects a mis-carved type lattice, and no-double-counting resting on submitter-declared dependence means the kernel relocates trust into metadata rather than reducing it. Sound-given-honest-annotation is worth much less to a regulator than sound, and adoption by a standards body is outside the criterion entirely."
      ],
      "essay_rank": "T4",
      "workflow_id": "T1",
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    },
    {
      "id": "T5",
      "title": "Assurance Measurement Science and the Adversarial Planted-Case Benchmark",
      "description": "Builds the measurement layer the other four programmes need, and does so as a vector rather than a score: corpus coverage, claim coverage, residual-risk bound, dependence exposure, replayability, formal verifiability, freshness, subgroup coverage, abstention behaviour, audit burden, correction latency, and authorised scope, reported as twelve numbers with a formal argument for why no scalar aggregation of them can be both faithful and non-gameable. Alongside it, a public benchmark of planted cases across nine failure families — correct computation from false sources, invalid causal inference from correct data, duplicated evidence disguised as independent, correlated model consensus, stale evidence, strategic submissions, missing minority-language material, legitimate-uncertainty cases, and normative conclusions presented as factual deduction — with a scoring rule that pays for correct refusal and correct scope limitation rather than accuracy alone. The benchmark's job is to make the failure modes T1–T4 formalise empirically detectable, and to make gaming visible when it happens.",
      "resolution_criterion": "PASS requires all six. (1) Corpus: at least 1,200 planted cases with at least 100 per failure family, each with machine-checkable ground truth and a published planting recipe, plus a held-out private split of at least 300 cases never released. (2) Scoring rule properties, pre-registered before any system is evaluated: prove monotonicity/strategy-proofness — no strategy improves the reported vector without improving at least one component in the intended direction — and prove the anti-single-score result, i.e. that no scalar aggregation of the twelve components retains this property, which is the formal backbone of the vector-not-score design. (3) Red-team gaming trial: at least three independent teams, paid a fixed budget and given the public split and full scoring specification, attempt to inflate scores without genuine capability improvement; PASS requires that no team achieves more than a 5-percentile gain on the private split (pre-registered threshold), and any team that exceeds it must be diagnosed into a specific scoring-rule defect that is then repaired and re-tested. (4) Human-judged components reach Krippendorff alpha ≥ 0.8 across at least three raters, or are removed from the vector. (5) Construct validity, pre-registered with a stated falsification condition: named vector components predict downstream failure rates on at least two external held-out task sets at pre-specified effect sizes — and the report publishes every component that fails to predict, including the case where abstention behaviour turns out not to track real-world error, since the value of this gate depends entirely on the negative results being reported. (6) Reproducibility: byte-identical scores from released code and data on a clean environment, verified by an external party.",
      "horizon_months": 15,
      "cost_band": "Medium: $300k–$460k. About 2 human FTE-years (one measurement/psychometrics lead, one adversarial-evaluation engineer), $75k for the red-team gaming trial across three independent teams, $45k for expert case authoring and multi-rater reliability work including minority-language material, $30k evaluation compute, $20k external reproducibility audit. (converted: ≈£235k–£360k at $1.27/£; the essay quotes the midpoint)",
      "crux": "Non-gameability cannot be certified, only pressure-tested. Once the nine planting recipes are public they become training targets, so what the benchmark measures thereafter is response to known families rather than general robustness — the private split delays this rather than preventing it, and the honest ceiling on the claim is 'resistant to three funded red teams for one cycle', not 'non-gameable'. The deeper crux is construct validity: whether 'dependence exposure' as operationalised actually tracks joint failure, or whether 'abstention' as scored rewards well-calibrated refusal rather than strategic vagueness, are human judgements that no proof settles, and getting them wrong produces a benchmark that measures something real but not the thing it claims. There is also a reflexive risk specific to this programme: publishing a twelve-component vector invites exactly the scalar collapse the anti-single-score theorem forbids — someone will average it, and the theorem will not stop them.",
      "agent_execution": "High for construction, low for validation — roughly 70% overall, the lowest of the five. Agents do: reconstruction of the measurement-validity, psychometrics, benchmark-gaming and Goodhart-resistance literatures; case authoring at scale with planting recipes and machine-checkable ground truth across all nine families; the scoring-rule formalisation including the monotonicity and anti-aggregation proofs; the harness, the reproducibility pipeline, and the automated red-team baselines; and adversarial self-attack on each candidate scoring rule. The human critical path is heavier here than in T1–T4 and is not compressible: choosing what the twelve components are supposed to mean, judging construct validity, recruiting and adjudicating three genuinely independent red teams, multi-rater reliability work, minority-language case authoring by native speakers, and pre-registration discipline that survives contact with disappointing results.",
      "durability": "Lowest of the five, by design and unavoidably. Measurements of current-system behaviour decay fastest — planted-case difficulty is calibrated against today's systems and erodes as they improve or as the recipes leak into training. What endures is the smaller and more valuable part: the scoring-rule properties, the anti-single-score impossibility result, the nine-family failure taxonomy, and the planting methodology, all of which remain applicable when every case in the corpus has been memorised. Budget expectation accordingly: the theorems and the taxonomy are ten-year artefacts, the leaderboard is an eighteen-month artefact, and the project should be written up so the durable half is separable from the perishable half rather than bundled with it.",
      "judge_probs": [
        0.2,
        0.15,
        0.3
      ],
      "judge_payoffs": [
        0.52,
        0.55,
        0.45
      ],
      "p_med": 0.2,
      "pay_med": 0.52,
      "rationales": [
        "Probability: lowest, and the reason is not proof difficulty but that three gates depend on things the team cannot make happen by working harder. Corpus construction (1,200 planted cases, nine families, machine-checkable ground truth, planting recipes, private split) is exactly what agents do well and is close to a solved problem here. The scoring-rule formalisation is plausible: strategy-proofness in the stated one-directional sense is provable, and the anti-single-score result is Arrow-flavoured and likely falls out. But gate 5 is a genuine empirical wager with a pre-registered falsification condition — requiring named components to predict downstream failure rates at pre-specified effect sizes on two external task sets is the honest design, and honest designs fail. 'Dependence exposure' and 'abstention behaviour' are the two components most likely not to predict, and they are the two the whole vector is built around. Gate 3 requires recruiting three genuinely independent funded red teams and adjudicating them; gate 4 requires alpha >= 0.8 on human-judged components across three raters, which is where measurement projects usually lose components. Fifteen months with 2 FTE and the heaviest incompressible human path of the five. Payoff: benchmarks get adopted faster than calculi — that is a real and underrated advantage, and this is the layer that makes T1-T4's failure modes empirically visible rather than merely formalised. But the perishable share is the largest here by design: once nine planting recipes are public they are training targets, the private split delays rather than prevents leakage, and the honest ceiling on the claim is resistance to three funded teams for one cycle. What endures is the smaller half — the anti-aggregation impossibility, the nine-family taxonomy, the planting methodology — and the reflexive risk is real: publishing twelve numbers invites the scalar collapse the theorem forbids, and the theorem will not stop anyone. Scored on the assumption the write-up separates the ten-year half from the eighteen-month half, as the spec promises.",
        "Six gates, the lowest agent share of the five (70%), and the two hardest gates are precisely the incompressible human ones. Construction is tractable: 1,200 planted cases with published recipes, a private split, the harness and byte-identical reproducibility are all things agents plus a competent eval engineer will deliver. Then it hits three walls. Gate 3: three funded independent red teams, none exceeding a 5-percentile private-split gain. The base rate for benchmarks surviving funded gaming attempts for one cycle is not good, and the repair-and-retest clause requires effectively a second trial with credible teams, which the budget and the 15-month horizon do not obviously fund. Gate 4: Krippendorff alpha >= 0.8 across three raters on constructs like 'normative conclusion presented as factual deduction' and 'strategic vagueness' — typical alphas for judgement of that kind land 0.6-0.75, and the escape hatch (remove the component) directly weakens gate 5, since the removed components are the ones most central to what the vector claims to measure. Gate 5 then asks named components to predict downstream failure at pre-specified effect sizes on two external task sets; the gate is admirably written to make negative results publishable, but PASS still requires the predictions to land. The anti-single-score impossibility result is provable in some formal model, and the risk there is a true theorem about a strawman aggregation class. Payoff is lowest and the spec says so honestly: the leaderboard is an eighteen-month artefact, planted-case difficulty decays as recipes leak into training, and the durable residue is the nine-family taxonomy, the planting methodology and the scoring-rule properties. Kept above 0.5 because that residue is genuinely the measurement layer the other four programmes need to be empirically checkable at all — T1's gate 4 and T2's construct-validity worry both want something like this to exist — and because the instruction to write it so the durable half is separable from the perishable half is the right hedge.",
        "Lowest probability and lowest durability, and the two are linked. Six gates, three of them gated on independent human parties: three funded red teams, three-rater reliability at Krippendorff alpha >= 0.8, an external reproducibility audit. Gate 3's 5-percentile threshold is the sort of bar funded red teams routinely clear on a published scoring specification, and the diagnose-repair-retest escape hatch turns each breach into another cycle inside a 15-month window — one or two iterations is survivable, three is not. Gate 5 is where I expect the honest failure: requiring named vector components to predict downstream failure rates at pre-specified effect sizes on two external task sets is the classic construct-validity wall, and 'dependence exposure' and 'abstention' are exactly the components most likely to fail to predict. The spec deserves credit for pre-registering the falsification condition and committing to publish negatives, but that makes PASS harder, not easier, and the gate is ambiguous about whether a component failing to predict fails the project. Gate 1 also overpromises: machine-checkable ground truth is straightforward for duplicated evidence or stale sources, and much closer to circular for invalid causal inference from correct data or normative conclusions dressed as deduction — the ground truth there is the judgement the benchmark claims to test. Gate 6's byte-identical scores are achievable only over cached model outputs. Payoff modest but not negligible: the leaderboard is an eighteen-month artefact whose planting recipes become training targets the moment they publish, so 'resistant to three funded red teams for one cycle' is the honest ceiling on the central claim. What survives is the anti-single-score impossibility result, the nine-family taxonomy and the planting methodology — genuinely durable and genuinely useful, since T1-T4 are otherwise largely unfalsifiable empirically. I bump payoff slightly for that enabling role, and dock it for the reflexive risk the spec names: publishing twelve components invites the scalar collapse the theorem forbids, and the theorem will not stop anyone averaging them."
      ],
      "essay_rank": "T5",
      "workflow_id": "T5",
      "assessedOn": "2026-08-04",
      "reviewBy": "2026-11-04"
    }
  ]
}
