E Evidence Press

Press release · 6 September 2026 · version 0.1.0-candidate

A five-vertex counterexample to ferromagnetic Potts censoring

An exact five-vertex example shows that omitting one predetermined update can leave a ferromagnetic three-colour Potts model closer to equilibrium at a fixed horizon.

Listen to this briefingNarrated summary · OpenAI API synthetic voice (fable) · MP3 · download
Watch this briefingVideo summary · YouTube · watch · play on this page

Summary

Can skipping an update leave a random system closer to equilibrium? For an important class of ordered models, a censoring theorem says no when the system starts at its top state. This unrefereed candidate shows that the analogous rule fails for a three-colour ferromagnetic Potts model started all one colour.

The example is small: five vertices, six edges, and nine scheduled opportunities. Omit one predetermined update and the final distribution is strictly closer to equilibrium. The difference is tiny but exactly positive. This is a counterexample to a universal finite-horizon rule—not evidence of a useful simulation speed-up.

Summary for specialists

Take the zero-field three-state Potts model on edges $AB,AC,AD,BC,BE,DE$, with Gibbs weights $30^{M(\sigma)}$, where $M$ counts monochromatic edges. Start at $00000$. Compare the word $(C,E,B,C,B,A,E,B,E)$ with its subsequence obtained by deleting the seventh opportunity, at $E$.

Writing $\mu$ and $\nu$ for the full and censored final laws, respectively, the exact gap is

$$\|\mu-\pi\|_{\mathrm{TV}}-\|\nu-\pi\|_{\mathrm{TV}} =\frac{7905357280856578194954129502105}{766036711510586802141859485820665762204}>0.$$

Both laws are compared with the full unconditioned Gibbs distribution at the same nine-opportunity horizon. Censorship is a fixed no-operation, not rejection based on realised colours.

Technical account

Numerical search found the object; exact arithmetic supplies the evidence. Three producer-side implementations reconstruct the same result: direct conditional transitions, integer masses over Gibbs fibres, and explicit marginal formulas.

The short word leaves $D$ untouched. The final update at $E$ gives both evolving laws the same conditional factor as equilibrium. Summing out that factor reduces the proof from 243 states to 27 marginal entries. The stationary mass on $D=0$ is one third; it is not renormalised. The other two thirds are included in the distance calculation.

Every equilibrium-preserving update contracts each chain's own distance. What can reverse is the ordering between two different chains' distances after a common suffix. The extra update initially helps: the full-minus-censored gaps after opportunities seven, eight and nine are approximately $-1.04\times10^{-8}$, $-2.04\times10^{-6}$ and $+1.03\times10^{-8}$. Exact rational receipts retain all three comparisons.

Evidence, assurance and limitations

The package supplies the complete finite argument, all final probabilities, marginal sign certificate, source code, negative controls and fresh-extraction replay. A longer activity-100 witness updates every vertex in both words and also has a strictly positive exact gap.

These are producer-coordinated checks. Internal AI editorial review, hosted CI, hashes and a DOI do not establish independent mathematical validation. Unaffiliated reproduction, formal verification, authenticated specialist review and journal peer review remain unestablished. Priority checking is targeted rather than exhaustive.

No smallest graph, shortest word, counterexample for every larger colour count, random-scan failure, high-temperature theorem, mixing-time improvement or practical benefit is claimed. The original AIM endpoint was inaccessible; its dataset transcription is substantively corroborated by Holroyd's primary paper.

Relationship to earlier work

Peres and Winkler prove the positive censoring theorem for monotone systems. Holroyd gives different failures, including an antiferromagnetic Potts example, and asks about the ferromagnetic constant-start case. The present witness has positive ferromagnetic activity and a monochromatic start.

Fill and Kahn explain comparison orders preserved under suitable monotonicity assumptions. Shared equilibrium and reversible individual updates alone do not imply those orders. Recent mean-field systematic-scan cutoff results address asymptotic convergence on complete graphs, not this finite-word deletion question.

Who should care, and why

AudiencePotential useRequired caution
Probability and statistical-mechanics researchersA compact obstruction to a universal Potts censoring extension.Preserve the finite-word and initial-state assumptions.
Exact-computation researchersInspect the full-state to marginal proof interface.Specialised checkers are not general input validators.
Sampling practitionersRecognise the limit of “more updates must be better” comparisons.The result supplies no practical scheduling recommendation or speed-up.

Why the problem matters

Censoring arguments let researchers simplify an update schedule while controlling convergence. Knowing where their hypotheses are essential prevents an attractive but invalid extension. A single exact positive gap has full logical force against a universal inequality, irrespective of its small decimal size.

How to inspect or reproduce the recorded checks

Download and extract the versioned evidence ZIP. With Python 3.10 or later, run python3 verify.py counterexample.json, python3 verify_integer.py, python3 verify_reduced.py, and python3 verify_supplement.py. Then run python3 test_verify.py and python3 -O test_verify.py.

Exact replay requires only the standard library. NumPy is needed solely for optional discovery. The README describes each checker's scope; the manifest and receipts identify the retained files. PDF builds need the separately documented TeX toolchain.

The most valuable next projects

  • Obtain an unaffiliated reconstruction of the written reduction and exact arithmetic.
  • Investigate smaller graphs or shorter schedules without treating this search as exhaustive.
  • Determine explicit activity intervals and what happens for other colour counts.
  • Identify structural conditions that prevent or allow the relative-distance reversal.

What is in the evidence package

The GitHub release and Zenodo archive contain the formatted paper, accessible Markdown, graph, exact laws and certificates, three checkers and supplemental replay, tests, discovery provenance, source comparisons, review response, internal editorial reports and component licences. A separate frozen editorial submission preserves what the reviewers saw. Communication art, audio and the prepared YouTube thumbnail add no mathematical evidence.

Media

The audio briefing is provided in the header above. Download the MP3 briefing · read the transcript.

The Rule and the Anomaly · Watch on YouTube

Open directions for follow-up research

Also available in machine-readable form for research agents and follow-up projects.

  1. Obtain an unaffiliated reconstruction and authenticated specialist review.
  2. Determine smaller graphs or shorter schedules; no minimality was proved.
  3. Find explicit activity intervals and assess other colour counts.
  4. Identify structural conditions controlling relative-distance reversal without inferring mixing-time or practical gains.

Research process, metrics and reusable methods

Prospective process metadata under the Evidence Press operating model and research-metrics policy. It records the intended handoff, measured scope and claim boundary; it is not evidence that the method accelerated this work.

Work ID
ep-work:potts-censoring-counterexample
Attempt and metric receipts
  • ep-attempt:potts-censoring-counterexample-assurance-publication — published / positive

    Measurement scope
    assurance-through-publication — Prospective supplemental assurance and publication only. The exact discovery, supplied review, and initial wording/documentation edits predate registration and are excluded; no historical discovery clock is reconstructed.
    Frozen target
    Supplemental exact checks, complete package and PDF, five-role internal editorial gate, public GitHub and Zenodo archives, media and two guarded publication cycles.
    Fermi active-time forecast
    120 minutes; plausible interval 80–180; expected unattended wait 20. Reference class: Evidence Press full-candidate procedure (n=1) — Procedural reference only; not a calibrated comparative sample..
    • Exact supplements, packaging and PDF: 1 × 20/30/45 minutes (low/central/high) — Existing written proof and three exact checkers.
    • Five-role editorial gate: 1 × 15/25/40 minutes (low/central/high) — Frozen target and bounded differentiated roles.
    • Immutable identities, page and media: 1 × 25/35/50 minutes (low/central/high) — Established generators and authenticated services.
    • Two deployment cycles and closeout: 1 × 20/30/45 minutes (low/central/high) — Required new-slug A/B and C/D cycles.
    Tractability forecast
    Within 180 active minutes: positive signal 0.95; target closure 0.85. Stop rule: Fail closed on scientific, rights, CI or public-byte-integrity failures; no expansion of the mathematical claim.
    Observed clocks
    5 active-agent; unknown active-human; unknown substantive-compute; 4 unattended-wait; 0 blocked; 1 rework minutes. Calendar elapsed: 31 minutes.
    Research search
    Cycles: 0 positive, 0 negative, 0 inconclusive. Falsification gates: 1. Candidate architectures: 0 tested, 0 rejected.
    Agent and review load
    6 agent runs; maximum parallelism 4; 6 model turns; 12260070 deduplicated model tokens; 1 substantive review rounds; P0/P1 findings 0/0; pre-publication claim corrections 0.
    Result and calibration
    positive-signal — Exact supplemental assurance, five internal Accept reports, immutable GitHub/Zenodo identity and the first canonical Evidence Press release passed. No new discovery cycle is claimed. The frozen target also includes the second ledger-sealing deployment, still pending at this immutable measurement cut; final Git/CI/deployment receipts establish that later event separately. Positive signal: true; target reached: false. Active-time error -115 minutes; actual/forecast 0.041666666666666664; inside interval: false. Brier score: positive signal 0.0025; target closure 0.7225. Variance: Generation and context-compaction wait are observed lower bounds, not complete labour or wait totals. Rework is an upper attribution bound on coordinator generation during the first failed site-build to repaired-content build window. Token scope uses response-completion timestamps; a boundary-crossing response may include some pre-registration input. Cached input is included in raw tokens and separated below. No acceleration inference. This first-publication cut is narrower than the forecast full closeout; no performance or forecast-success conclusion follows.
    Missing telemetry
    activeHumanMinutes — No contemporaneous human active-time telemetry.; computeMinutes — No complete resource-time accounting for all local and hosted checks.
    Measurement corrections
    • measurement.reworkMinutes -> metrics.outcome.reworkMinutes — Retain zero in the historical snapshot; terminal measured rework is 1 rounded minutes, an upper attribution bound on instrumented coordinator generation during the repair window, not complete labour. Reason: The first snapshot retained zero before site preflight identified stale derived Atlas counts and missing catalogue audio metadata.
    • measurement.missingnessReason -> metrics.outcome.deduplicatedModelTokens and metrics.outcome.uncachedInputTokens — Terminal token totals deduplicate response_id, match each source file own thread identity, and include cached input in the raw total. No cumulative token_count or inherited fork counters are summed; boundary-completion timing and incomplete active-generation coverage are disclosed. Reason: Runtime token_usage_record events were subsequently identified; the initial unavailability statement remains a historical snapshot.
Prospective work ledger · metrics policy
Intended aims
science
Artifact roles
research-output, evidence-assessment, communication
Decision object
counterexample — Explicit finite ferromagnetic Potts heat-bath update word and exact positive distance gap. Scope: q=3, five vertices, activity 30, monochromatic start and one fixed nine-opportunity comparison.
Reusable methods
Certificate-first, proof-carrying research (certificate-first); Structural compression (structural-compression); Adversarial scientific controls (adversarial-controls); Counterexample- and proxy-first analysis (counterexample-proxy-first); Assurance as a vector (assurance-vector); Agent-readable research objects (agent-readable-research-object) · registry
Targeted clocks
assurance, publication
Semantic bridge
explicit — Neighbour-count heat-bath probabilities, full Gibbs-fibre reconstruction and conditional-factor marginal reduction target the same finite state laws and predetermined deletion. Remaining risks: Same-producer implementation and semantic errors remain possible.; No proof-assistant certification.; Inaccessible original AIM endpoint; bounded priority search.; No broader mixing-time or practical consequence established..
Human judgement gates
  • Check the source-to-state-law and unconditioned-equilibrium bridge.
  • Preserve finite-horizon scope and distinguish contraction from relative ordering.
  • Assess priority separately from exact correctness.
  • Preserve rights, Anonymous authorship and assurance separation.
Next assurance action
Seek unaffiliated exact reconstruction and specialist assessment; treat extensions as new research. Claim ceiling: Unrefereed exact finite-horizon q=3 counterexample candidate. No mixing-time, random-scan, every-q, minimality, practical speed-up, exhaustive novelty, priority, unaffiliated validation, formal proof or impact claim.
Aim-scoped impact evidence
  • science: NO_IMPACT_EVIDENCE — Reusable exact censoring obstruction in Producer-coordinated mathematical candidate publication. Design: none; comparator: No matched workflow comparator.; estimand: No discovery, effort, reuse or impact effect estimated.. No real-world effect evidence is asserted.
Parent handoffs
  • depends-on-claim https://arxiv.org/abs/1101.4690 — inherited claim: Primary text poses the ferromagnetic constant-start censoring question.; inherited ceiling: Historical question framing, not a proof or current priority clearance.

Verification status

Unrefereed exact finite-horizon q=3 counterexample candidate. No mixing-time, random-scan, every-q, minimality, practical speed-up, exhaustive novelty, priority, unaffiliated validation, formal proof or impact claim.

Cite

Anonymous. (2026). A five-vertex counterexample to ferromagnetic Potts censoring (Version 0.1.0-candidate) [Unrefereed candidate and exact evidence package]. Evidence Press. https://doi.org/10.5281/zenodo.22546547
BibTeX
@misc{pottscensoringcounterexample2026,
  title        = {A five-vertex counterexample to ferromagnetic Potts censoring},
  author       = {Anonymous},
  year         = {2026},
  doi          = {10.5281/zenodo.22546547},
  url          = {https://doi.org/10.5281/zenodo.22546547},
  version      = {0.1.0-candidate},
  howpublished = {Zenodo},
  note         = {Unrefereed; internally replayed evidence package. Press page: https://evidencepress.org/releases/potts-censoring-counterexample/}
}

Also: cite.bib · paper.json · this page as Markdown