---
title: "A five-vertex counterexample to ferromagnetic Potts censoring"
date: 2026-09-06
version: "0.1.0-candidate"
doi: 10.5281/zenodo.22546547
pdf: https://github.com/ipitchford/potts-censoring-counterexample/releases/download/v0.1.0-candidate/potts-censoring-counterexample-0.1.0-candidate.pdf
repository: https://github.com/ipitchford/potts-censoring-counterexample
archive: https://zenodo.org/records/22546547
license: CC0-1.0
status: unrefereed (internally replayed; not peer reviewed, not independently reproduced, not formally verified)
---

# A five-vertex counterexample to ferromagnetic Potts censoring

## Summary

Can skipping an update leave a random system closer to equilibrium? For an important class of ordered models, a censoring theorem says no when the system starts at its top state. This unrefereed candidate shows that the analogous rule fails for a three-colour ferromagnetic Potts model started all one colour.

The example is small: five vertices, six edges, and nine scheduled opportunities. Omit one predetermined update and the final distribution is strictly closer to equilibrium. The difference is tiny but exactly positive. This is a counterexample to a universal finite-horizon rule—not evidence of a useful simulation speed-up.

## Summary for specialists

Take the zero-field three-state Potts model on edges $AB,AC,AD,BC,BE,DE$, with Gibbs weights $30^{M(\sigma)}$, where $M$ counts monochromatic edges. Start at $00000$. Compare the word $(C,E,B,C,B,A,E,B,E)$ with its subsequence obtained by deleting the seventh opportunity, at $E$.

Writing $\mu$ and $\nu$ for the full and censored final laws, respectively, the exact gap is

$$
\|\mu-\pi\|_{\mathrm{TV}}-\|\nu-\pi\|_{\mathrm{TV}}
=\frac{7905357280856578194954129502105}{766036711510586802141859485820665762204}>0.
$$

Both laws are compared with the full unconditioned Gibbs distribution at the same nine-opportunity horizon. Censorship is a fixed no-operation, not rejection based on realised colours.

## Technical account

Numerical search found the object; exact arithmetic supplies the evidence. Three producer-side implementations reconstruct the same result: direct conditional transitions, integer masses over Gibbs fibres, and explicit marginal formulas.

The short word leaves $D$ untouched. The final update at $E$ gives both evolving laws the same conditional factor as equilibrium. Summing out that factor reduces the proof from 243 states to 27 marginal entries. The stationary mass on $D=0$ is one third; it is not renormalised. The other two thirds are included in the distance calculation.

Every equilibrium-preserving update contracts each chain's own distance. What can reverse is the ordering between two different chains' distances after a common suffix. The extra update initially helps: the full-minus-censored gaps after opportunities seven, eight and nine are approximately $-1.04\times10^{-8}$, $-2.04\times10^{-6}$ and $+1.03\times10^{-8}$. Exact rational receipts retain all three comparisons.

## Evidence, assurance and limitations

The package supplies the complete finite argument, all final probabilities, marginal sign certificate, source code, negative controls and fresh-extraction replay. A longer activity-100 witness updates every vertex in both words and also has a strictly positive exact gap.

These are producer-coordinated checks. Internal AI editorial review, hosted CI, hashes and a DOI do not establish independent mathematical validation. Unaffiliated reproduction, formal verification, authenticated specialist review and journal peer review remain unestablished. Priority checking is targeted rather than exhaustive.

No smallest graph, shortest word, counterexample for every larger colour count, random-scan failure, high-temperature theorem, mixing-time improvement or practical benefit is claimed. The original AIM endpoint was inaccessible; its dataset transcription is substantively corroborated by Holroyd's primary paper.

## Relationship to earlier work

Peres and Winkler prove the positive censoring theorem for monotone systems. Holroyd gives different failures, including an antiferromagnetic Potts example, and asks about the ferromagnetic constant-start case. The present witness has positive ferromagnetic activity and a monochromatic start.

Fill and Kahn explain comparison orders preserved under suitable monotonicity assumptions. Shared equilibrium and reversible individual updates alone do not imply those orders. Recent mean-field systematic-scan cutoff results address asymptotic convergence on complete graphs, not this finite-word deletion question.

## Who should care, and why

| Audience | Potential use | Required caution |
|---|---|---|
| Probability and statistical-mechanics researchers | A compact obstruction to a universal Potts censoring extension. | Preserve the finite-word and initial-state assumptions. |
| Exact-computation researchers | Inspect the full-state to marginal proof interface. | Specialised checkers are not general input validators. |
| Sampling practitioners | Recognise the limit of “more updates must be better” comparisons. | The result supplies no practical scheduling recommendation or speed-up. |

## Why the problem matters

Censoring arguments let researchers simplify an update schedule while controlling convergence. Knowing where their hypotheses are essential prevents an attractive but invalid extension. A single exact positive gap has full logical force against a universal inequality, irrespective of its small decimal size.

## How to inspect or reproduce the recorded checks

Download and extract the versioned evidence ZIP. With Python 3.10 or later, run `python3 verify.py counterexample.json`, `python3 verify_integer.py`, `python3 verify_reduced.py`, and `python3 verify_supplement.py`. Then run `python3 test_verify.py` and `python3 -O test_verify.py`.

Exact replay requires only the standard library. NumPy is needed solely for optional discovery. The README describes each checker's scope; the manifest and receipts identify the retained files. PDF builds need the separately documented TeX toolchain.

## The most valuable next projects

- Obtain an unaffiliated reconstruction of the written reduction and exact arithmetic.
- Investigate smaller graphs or shorter schedules without treating this search as exhaustive.
- Determine explicit activity intervals and what happens for other colour counts.
- Identify structural conditions that prevent or allow the relative-distance reversal.

## What is in the evidence package

The GitHub release and Zenodo archive contain the formatted paper, accessible Markdown, graph, exact laws and certificates, three checkers and supplemental replay, tests, discovery provenance, source comparisons, review response, internal editorial reports and component licences. A separate frozen editorial submission preserves what the reviewers saw. Communication art, audio and the prepared YouTube thumbnail add no mathematical evidence.




## Open directions for follow-up research

- Obtain an unaffiliated reconstruction and authenticated specialist review.
- Determine smaller graphs or shorter schedules; no minimality was proved.
- Find explicit activity intervals and assess other colour counts.
- Identify structural conditions controlling relative-distance reversal without inferring mixing-time or practical gains.

## Research process, metrics and reusable methods

This is prospective process metadata under the Evidence Press operating model and research-metrics policy. It records the intended handoff, measured scope and claim boundary; it is not evidence that the method accelerated this work.

- Work ID: ep-work:potts-censoring-counterexample
- Attempt and metric receipts: ep-attempt:potts-censoring-counterexample-assurance-publication: published / positive; scope assurance-through-publication; target Supplemental exact checks, complete package and PDF, five-role internal editorial gate, public GitHub and Zenodo archives, media and two guarded publication cycles.; active forecast 120 minutes (80-180); Fermi components Exact supplements, packaging and PDF: 1 x 20/30/45 minutes low/central/high (Existing written proof and three exact checkers.); Five-role editorial gate: 1 x 15/25/40 minutes low/central/high (Frozen target and bounded differentiated roles.); Immutable identities, page and media: 1 x 25/35/50 minutes low/central/high (Established generators and authenticated services.); Two deployment cycles and closeout: 1 x 20/30/45 minutes low/central/high (Required new-slug A/B and C/D cycles.); positive-signal/closure probabilities 0.95/0.85 within 180 active minutes; observed active-agent/human/compute/wait/blocked/rework minutes 5/unknown/unknown/4/0/1; cycles positive/negative/inconclusive 0/0/0; falsification gates 1; architectures tested/rejected 0/0; result positive-signal; target reached false; forecast error -115 minutes; ratio 0.041666666666666664; inside interval false; positive-signal/target-closure Brier scores 0.0025/0.7225; missing telemetry activeHumanMinutes: No contemporaneous human active-time telemetry.; computeMinutes: No complete resource-time accounting for all local and hosted checks.; appended measurement corrections measurement.reworkMinutes -> metrics.outcome.reworkMinutes: Retain zero in the historical snapshot; terminal measured rework is 1 rounded minutes, an upper attribution bound on instrumented coordinator generation during the repair window, not complete labour. (reason: The first snapshot retained zero before site preflight identified stale derived Atlas counts and missing catalogue audio metadata.); measurement.missingnessReason -> metrics.outcome.deduplicatedModelTokens and metrics.outcome.uncachedInputTokens: Terminal token totals deduplicate response_id, match each source file own thread identity, and include cached input in the raw total. No cumulative token_count or inherited fork counters are summed; boundary-completion timing and incomplete active-generation coverage are disclosed. (reason: Runtime token_usage_record events were subsequently identified; the initial unavailability statement remains a historical snapshot.). Work ledger: https://evidencepress.org/api/work-ledger.json. Metrics policy: https://evidencepress.org/api/research-metrics-policy.json
- Intended aims: science
- Artifact roles: research-output, evidence-assessment, communication
- Decision object: counterexample — Explicit finite ferromagnetic Potts heat-bath update word and exact positive distance gap. Scope: q=3, five vertices, activity 30, monochromatic start and one fixed nine-opportunity comparison.
- Reusable methods: Certificate-first, proof-carrying research (certificate-first); Structural compression (structural-compression); Adversarial scientific controls (adversarial-controls); Counterexample- and proxy-first analysis (counterexample-proxy-first); Assurance as a vector (assurance-vector); Agent-readable research objects (agent-readable-research-object). Registry: https://evidencepress.org/api/method-registry.json
- Targeted clocks: assurance, publication
- Semantic bridge: explicit — Neighbour-count heat-bath probabilities, full Gibbs-fibre reconstruction and conditional-factor marginal reduction target the same finite state laws and predetermined deletion. Remaining risks: Same-producer implementation and semantic errors remain possible.; No proof-assistant certification.; Inaccessible original AIM endpoint; bounded priority search.; No broader mixing-time or practical consequence established..
- Human judgement gates: Check the source-to-state-law and unconditioned-equilibrium bridge.; Preserve finite-horizon scope and distinguish contraction from relative ordering.; Assess priority separately from exact correctness.; Preserve rights, Anonymous authorship and assurance separation.
- Next assurance action: Seek unaffiliated exact reconstruction and specialist assessment; treat extensions as new research.
- Claim ceiling: Unrefereed exact finite-horizon q=3 counterexample candidate. No mixing-time, random-scan, every-q, minimality, practical speed-up, exhaustive novelty, priority, unaffiliated validation, formal proof or impact claim.
- Aim-scoped impact evidence:
  - science: NO_IMPACT_EVIDENCE — Reusable exact censoring obstruction in Producer-coordinated mathematical candidate publication; design none; comparator No matched workflow comparator.; estimand No discovery, effort, reuse or impact effect estimated.; no real-world effect evidence asserted
- Parent handoffs: depends-on-claim https://arxiv.org/abs/1101.4690; inherited claim: Primary text poses the ferromagnetic constant-start censoring question.; inherited ceiling: Historical question framing, not a proof or current priority clearance.



## Verification status

Unrefereed exact finite-horizon q=3 counterexample candidate. No mixing-time, random-scan, every-q, minimality, practical speed-up, exhaustive novelty, priority, unaffiliated validation, formal proof or impact claim.

## References

1. Peres and Winkler (2013). Can extra updates delay mixing? Monotone-system positive theorem. <https://doi.org/10.1007/s00220-013-1776-0>
2. Holroyd (2011). Some circumstances where extra updates can delay mixing. Antiferromagnetic failure and ferromagnetic constant-start question. <https://doi.org/10.1007/s10955-011-0365-x>
3. Fill and Kahn (2013). Comparison inequalities and fastest-mixing Markov chains. <https://doi.org/10.1214/12-AAP886>
4. Blanca and Tahmidur (2026). Mixing and cutoff for the systematic scan dynamics of the mean-field ferromagnetic Potts model. Different asymptotic setting. <https://arxiv.org/abs/2607.09841>
