Skip to content

Regenerate the retracted figures, and the v1.1 material that followed - #23

Open
aurascoper wants to merge 53 commits into
fix/ledger-guard-document-scopefrom
fix/figure-artifacts-and-v11
Open

Regenerate the retracted figures, and the v1.1 material that followed#23
aurascoper wants to merge 53 commits into
fix/ledger-guard-document-scopefrom
fix/figure-artifacts-and-v11

Conversation

@aurascoper

Copy link
Copy Markdown
Owner

Stacked on #20. Seven commits, in the order they happened.

The defect

fig2_melanin_accumulation.pdf said, inside the image, "C. neoformans, C. sphaerospermum are radiotrophic (melanin-mediated energy gain)" — under a caption disowning "radiation-derived energy production" and against §2.6, "Radiotrophy is not established for any of the seven species modelled." fig1 said "radiotrophic niche", on the wrong side of the plot.

None of it was unnoticed, which is the part worth recording. The generator was corrected on 2026-08-14 (d53f236, 01b0f4d) and the verdict was written down three times — RM-G08-01, RM-G10-01, PP-65-08 ("deliberately not regenerated"). The verdict was reached, the source was fixed, and the only artifact anyone opens went on carrying the claim, because nothing could fail: the claims guard read one .tex, and tests/runtests.jl splits the monolith above # 13. Figure export, so the Julia suite cannot reach export_figures at all.

Regenerating from HEAD alone would have retired one contradiction and shipped four — only two of six in-plot strings had been fixed, and one "fix" asserted drift toward the source while §6.3 reports the ordering running the opposite way, within 1.4 SE of the seeding null.

The guard, three tiers, each with a control

  • sha256 per PDF, stdlib only. Without it, regenerating a figure and forgetting its sidecar leaves the suite green on stale-but-clean text while the PDF still carries the claim — the artifact-versus-source split reproduced inside the guard written to close it.
  • The phrase guard reads <figure>.txt as the document a row names.
  • A vocabulary floor for labels under MIN_WORDS. Neither figure row is reachable by the phrase guard — "radiotrophic niche" is two words — and that limit is asserted rather than left to look like coverage.

The pinned-SHA control idiom is retired here. It was used three times, each time noting that a squash-merge degrades the control into a skip, and each time proceeding. All four controls are now committed files under calibration/tests/fixtures/.

Also declares GENERATED_ARTIFACTS: artifacts/ is gitignored and nothing under it is tracked, so two document-resolution tests were passing on untracked local state and would have failed on any clean checkout.

v1.1 and after

  • §7.3 — the facility regime, argued from the manuscript's own ten-orders-of-magnitude dose-rate gap rather than pitched. Three references, each verified at source; the brief they came from was wrong on three counts.
  • Ethics — the AM7 taxonomy (NCBI taxid 94625, CDC 2022-12-19). Rule 6: the record and the handling instruction, no containment determination.
  • A deletion. v1.0 said "None of those numbers appears in this manuscript." The v1.1 correction is in this manuscript. AGENTS.md step 5 says delete rather than reword.
  • preprint/hoffman_memo.tex — four pages. The last is questions, because an earlier draft asserted MURR's capabilities back to a MURR scientist from a brief whose beamline inventory was single-sourced to a seminar abstract. HOFFMAN-11 records the one claim dropped for want of a readable source, in three parts: identified, inaccessible (HTTP 403, not absent), attribution unconfirmed.
  • The coinage at scan levelradiodialysis appeared ten times and its disclaimer covered one. Fixed positionally, which makes it two lines and makes it checkable. The suggestion to give §3.11 the "specification, not method" paragraph was refused: §3.2/3.4/3.12 carry it over code that does not exist, and §3.11's equations run.

§6.2, Table 4 — what decided each accepted move

compute_delta_H computed four terms and discarded the decomposition on its return line. Split out (not retyped — three callers want the scalar, one is the JACC cross-implementation check), the counterfactual runs on the variate the acceptance test already drew.

The direct radiation term reversed none of 206,042 accepted moves across three seeds. Not under-sampling: only a term that lowered ΔH can be decisive, β_ion is negative for two of seven species at −5e-5, so reversal needs the variate within ~1e-5 of the threshold — about 0.5 expected in 53,603. This is §6.2's one-part-in-1e5 as a count.

66–77% of accepted moves are reversed by removing no single term. No spatial map: median 3 accepted moves per touched voxel across seven categories. PP-62-07 records that refusal and the two alternatives also refused.

Two guard findings

  • The byte contract has a detection floor. 0.5 → 0.5000001 in the melanin coefficient leaves contract_csv.jl byte-identical; 0.5 → 50.0 fails it. Conditionally green with the condition unstated — a third kind of check-that-cannot-fail. Documented in validate_serial.jl, paired with an exact === check whose control is that summation order is observable on 15.7% of moves.
  • A test that passed over dead surface. The first decomposition test saw ΔH_mel = 0.0 on all 21,492 sampled moves, because melanin only grows through update_melanin!, which the harness never called. State is stepped before sampling now, and all four branches are asserted nonzero.

Suites

calibration 371, contract 7, coupling 313 passed / 6 skipped, Julia 20,171 across seven testsets.

Not a clean run. checkpoint_io_tests.jl and test_julia_interop fail for want of HDF5 in this worktree's unresolved depot; they pass in the main checkout and nothing here introduced them. Six coupling skips are openmc.

Note for reviewers: the venv's editable install points at the main checkout, so coupling must be run with PYTHONPATH set to this tree or it measures the wrong package.

🤖 Generated with Claude Code

https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap

aurascoper and others added 7 commits August 29, 2026 00:11
`fig2_melanin_accumulation.pdf` said, inside the image, "C. neoformans,
C. sphaerospermum are radiotrophic (melanin-mediated energy gain)" — under a
caption disowning "radiation-derived energy production" and against §2.6,
"Radiotrophy is not established for any of the seven species modelled."
`fig1` said "radiotrophic niche".

Neither was unnoticed, which is the part worth recording. The generator was
corrected on 2026-08-14 (d53f236, 01b0f4d) and the verdict was written down
three times — RM-G08-01 "regenerate the committed PNGs", RM-G10-01 "drop
'radiotrophic'", PP-65-08 "the committed PNGs still carry the old reversed zone
labels ... deliberately not regenerated". The verdict was reached, the source
was fixed, and the only artifact anyone opens went on carrying the retracted
claim, because nothing could fail: the claims guard read one .tex file, and
tests/runtests.jl splits the monolith above `#  13. Figure export`, so the Julia
suite cannot reach export_figures at all.

Regenerating from HEAD alone would have retired one contradiction and shipped
four. Only two of six in-plot claim strings had been fixed:

  - "Radial stratification — ..." as a title, against L976 "not evidence of
    established stratification"
  - "Melanin accumulation — radiation-driven production", against the caption
  - "Melanin producers (★ radiotropic)", a phenotype attribution for what
    RM-G10-01 records as a model input
  - "★ ... drift toward the source (radiotropic)", which asserts as fact what
    the run's own output contradicts (L966-967, "their observed ordering is in
    the opposite direction", within 1.4 SE of the seeding null at L972).
    Replaced, not reworded: there is no true version of that sentence.
  - "(% depleted)" on fig 4, the artifact half of RM-KR-06

Figures 1 and 2 reproduce the manuscript's numbers exactly — mean_r 13.10 /
11.55 / 9.94 against the captions' 13.1 / 11.5 / 9.9, and M=1.4372 against
M=1.44 — so no caption moves. They came from main_coupled(), which now refuses
to run (RADIODIALYSIS: BLOCKED at a coupled X_total). That refusal is correct
and is untouched: figs 1-2 read only snapshot mean_r and mean_melanin, and
neither the Hamiltonian nor update_melanin! reads state.nutrient, which is the
only channel by which radiodialysis reaches the CPM state. Figures 3 and 4 are
NOT regenerated and are named as such by FIG-03.

The guard, three tiers, each with a control:

  - sha256 per PDF, pinned in stdlib. Without it, regenerating a figure and
    forgetting its sidecar leaves the suite green on stale-but-clean text while
    the PDF still carries the claim — the artifact-versus-source split
    reproduced inside the guard written to close it.
  - the phrase guard reads `<figure>.txt` as the document a row names
  - a vocabulary floor for labels under MIN_WORDS. NEITHER figure row is
    reachable by the phrase guard: "radiotrophic niche" is two words, and fig
    2's annotation splits on its comma and parens into runs of 2, 4 and 3. That
    limit is asserted rather than left to look like coverage.

The pinned-SHA control idiom is retired here. It was used three times —
5980dc5, 9319d43, and a third that would have been e24dbec — each time noting
that a squash-merge makes the commit unreachable and degrades the control into
a skip, and each time proceeding. All four controls are now committed files
under calibration/tests/fixtures/.

Also declares GENERATED_ARTIFACTS. `artifacts/` is gitignored and nothing under
it is tracked, so two document-resolution tests were passing on untracked local
state and would have failed on CI or any clean checkout — found by building
this branch in a fresh worktree.

367 passed (was 362).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap
Three additions to the manuscript, one new document, and one deletion.

SECTION 7.3 — the use-case gap, argued from the model's own constraint rather
than pitched. The manuscript already puts reactor irradiation (~Gy/min) and its
cited environmental motivation (~mGy/yr) ten orders of magnitude apart and says
a constant calibrated in one does not transfer to the other. Engineered
facilities sit between them, and that is where the offline counterfactual's
inputs — bounded geometry, characterised source, measurable metal inventory —
are obtainable rather than assumed. Written to satisfy PP-8-04, which ruled
against the deployment register: this predicts nothing and proposes nothing, and
says in as many words that the model is unchanged and only the setting differs.

Three references, each verified at source before it entered the file:
  - Sarro 2007 (PMID 17426994) — the brief this came from said "34 months" and
    named Cofrentes. The abstract says neither, only "a Spanish nuclear power
    plant". Both details dropped rather than carried.
  - Karley 2018 (PMID 29063404) — D10 248 Gy to 2 kGy and 3.8 ug/mg both
    verbatim in the abstract. The brief attributed the 3.8 figure to the 2023
    paper and the authorship to Shukla; both wrong.
  - Karley 2023 (PMID 37209244) — dead biomass removes Co and Ni, which is what
    sorption predicts. Its "4e-4 to 1e-5 g/mg" is dimensionally odd and two
    orders off the same isolates' 2018 figure, so it is quoted nowhere.

ETHICS — the AM7 taxonomy, which the manuscript had never carried. NCBI serves
taxid 94625 as Brucella intermedia with Ochrobactrum intermedium as a synonym;
CDC's 2022-12-19 lab update directs Class II BSC handling and state-lab referral
for anything identified as Brucella. Verified directly against the NCBI record
and the CDC notice. An initially guessed DOI for Oren & Garrity resolved to an
unrelated paper on Raineyella fluvialis and was discarded.

Rule 6 governs what this may say. The research brief behind it concluded BSL-2
and non-select-agent for the SPECIES; biosafety follows strains, so the section
records the taxonomy and the handling instruction and declares no containment
level, no select-agent status, and nothing touching D-APPROVAL.

THE DELETION — v1.0 said "None of those numbers appears in this manuscript."
The v1.1 figure correction is in this manuscript, which disproves it. AGENTS.md
step 5 says delete rather than reword, because a paraphrase is the same claim
and this repository has corrected its own corrections twice by making it. The
sentence is gone; what moved is stated instead.

preprint/hoffman_memo.tex — three pages. Page 3 ASKS what the facility can
supply rather than stating it. The draft it replaced asserted MURR's
capabilities back to a MURR scientist from a secondhand brief whose beamline
inventory was single-sourced to a seminar abstract — the same inference class
this repository exists to refuse, and checkable by its reader in a minute.

docs/correspondence/wan_v11_note.md — unsent draft.

Nine ledger rows. 367 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap
…nobody ran

Three analyses arrived with the Hoffman thread. Two had premises that did not
survive contact with the source, and both were caught before anything shipped.

THE COINAGE. The suggestion was to give section 3.11 the "This subsection is
specification, not method" paragraph that 3.2, 3.4 and 3.12 carry. That would
have been false in the opposite direction: those three stand over code that does
not exist -- "No program in this work integrates Eq. 2", "no simulation reported
here holds an independently-integrated momentum state" -- while 3.11's equations
run in biofilms_radiodialysis.R and in Julia, with results in 6.5.

The real defect was narrower and is the figure defect's shape. 3.11 already
disowns the term in a full paragraph, names the established descriptors, and
Figure 3's caption already says the two curves are not a closed feedback law.
The term then appeared ten times and the disclaimer covered one; the abstract
said "a radial radiodialysis solver" with no signal the word was coined here.
Caveat in the body, claim in the part people scan.

Fixed positionally rather than editorially, which makes it two lines and makes
it checkable: the term must not appear before the paragraph that disowns it.
Abstract now reads "a radial membrane-transport solver"; the heading is
"Membrane Transport Under Radiation-Driven Permeability Change". Nine uses
remain, all introduced. No version bump and no correction entry -- v1.1 has been
sent to nobody, and a correction paragraph is for a claim that reached a reader.

THE MEMO. The corrosion framing is right and one adjustment matters enough to
state in the memo: MIC is a chemical attack that happens to occur in a radiation
field, and every one of its mechanisms works at background dose. Radiation
selects which organisms are present. The opposite reading is a radiation-driven
metabolism claim of exactly the class section 2.6 retracts, and it would have
entered through the framing rather than through a sentence.

Hoffman's papers were verified before anything was attributed to him, and the
identity check was not a formality: his Scholar profile is filed under Catalyst
Science Solutions, the 2024 Materialia paper lists him at GE Research, and a
Florida seminar listing is what joins those to MURR. His envelope -- steam,
hydrothermal chemistry, hydrogen-isotope permeation -- comes from three papers
that are his and are readable.

DROPPED: the used-fuel-pool pitting-resistance claim, which would have carried
the argument the whole distance. NACE-2019-12944 returns HTTP 403 and the
surrounding literature points to Rebak. Recorded as a refusal in three parts,
because claim-identified, source-inaccessible and attribution-unconfirmed are
three different states and only the middle one is recoverable by a reader with
access. HOFFMAN-06 makes the same point without it.

The geometry caveat is the sentence that keeps page 3 honest -- cylinder in
water, Robin boundary, no metal and no interface -- and it is also the sentence
that reads like a hedge and goes first when someone tightens the page. It is now
a ledger row with a test behind it. Four pages, not three: the argument is a
page and letting LaTeX orphan a header to hide that was not an option.

THE GUARD GAINS ITS MIRROR. Every document test here asserts ABSENCE, because
`delete` is the verdict where absence is the criterion. That left the opposite
failure uncovered: a deliberately-carried sentence quietly trimmed, with nothing
able to notice. LOAD_BEARING is the mirror, on the RETRACTED_IN_FIGURES idiom --
an explicit tuple rather than a new verdict or column, since most `keep` rows
are not quotable prose. Both new tests have controls, and both were proved
against the real files: planting a bare term above the disclaimer and striking
the geometry sentence failed those two tests and no others.

THE SURROGATE AXIS. NEWS-AUD-03 and two sibling audits searched for real
biofilms pairing wet mass, dry mass, blanks and a calibrated hydrated volume,
and each says the search was sampled and not exhausted. None searched for a
surrogate; "phantom" appears in this repository only as the A0 water geometry.
Recorded as an axis, which is true independent of what any source contains.

Hellriegel 2014 resolved on the full text: it is a surrogate, not a fit -- a
growth-independent gellan imitate, 2 mL cast in a 40 x 2 mm mould at 0.33-1.17%
w/v, whose stated purpose is testing characterization tools before they meet a
real biofilm. It reports no density, no water content and no masses, which turns
out not to matter: a cast specimen's true values are set by the recipe. It
CANNOT clear D-RHOWET and SURR-01 says so -- a gellan gel at one percent solids
has water's density and bounds nothing about a biofilm's dry fraction. What it
bounds is the protocol.

Also found while searching it: PMC8579398 pairs wet mass, dry mass, water
content and matched agar blanks on E. coli biofilms, and has no hydrated volume
at all -- thickness only. The closest near-miss located so far, failing on the
same single term as every other. density_g_cm3 stays blank.

The dataset schema refused three invented enum values on first write. They were
mapped onto the existing vocabulary rather than the vocabulary widened.

371 passed (was 367). Julia interop fails in this worktree for want of HDF5 in
a depot with no Manifest; no Julia code is touched here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap
`compute_delta_H` computed ΔH_adh, ΔH_vol, ΔH_rad and ΔH_mel as four separate
locals and threw the decomposition away on its return line. §6.2 reports the
melanin term biasing acceptance by 15.5% against the direct radiation term's one
part in 1e5 -- a ratio of about 1.5e4 -- and says in its last sentence that
hand-specified adhesion differences are larger than both. That is a spatial fact
reported as a table row, and the numbers to draw it were being computed 6.4M
times a run and discarded.

SPLIT, NOT RETYPED. compute_delta_H_terms is the old body with the sum removed;
compute_delta_H adds the fields back in the same order. Three callers outside
this file want the scalar and one of them is biofilms_potts_jacc.jl's
cross-implementation check against the parallel port. Changing what that reads,
to make a rendering change easier, would trade a real guard for a picture.

THE BYTE CONTRACT HAS A FLOOR, MEASURED RATHER THAN ASSUMED. Perturbing the
melanin coefficient from 0.5 to 0.5000001 leaves contract_csv.jl's output
byte-identical; 0.5 to 50.0 fails it. A 2e-7 shift moves exp(-ΔH/T) by ~4e-8 and
has to straddle one of ~6.4M uniform draws to appear at all, so the contract
guards the trajectory and does not certify bit-exactness against small numerical
change. Green there was necessary and not sufficient.

tests/delta_h_decomposition.jl is the exact check: the four terms sum to the
scalar under `===` on every sampled move. Its control is that SUMMATION ORDER IS
OBSERVABLE -- reordering the same four terms differs on 3367 of 21492 sampled
moves, 15.7%, because adhesion and volume are O(1), radiation is O(1e-5) and
melanin is O(0.5), so the small terms vanish or survive depending where they
land. Without that, "same operands, same order" would be describing an
associativity that holds anyway and the file would assert a tautology.

It also requires all four branches to be nonzero. The first version of this test
saw ΔH_mel = 0.0 on all 21492 moves and passed: melanin starts at zero and only
grows through update_melanin!, which the harness never called, so the melanin
branch was dead surface a green test walked over. The state is now stepped
before sampling.

PROVED BY MUTATION, not by reading: planting a reordered sum in compute_delta_H
fails the exactness assertion and the control both, and nothing else.

Julia: 162 passed across five testsets. checkpoint_io_tests.jl still errors on
`Package HDF5 ... is required but does not seem to be installed` -- this worktree
has no Manifest.toml, the main checkout does and resolves HDF5 there, so it is
environmental and predates this change. NOT A CLEAN RUN.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap
…ng anything

The counterfactual, on the uniform that was actually drawn, never on a new one:
drawing again advances the stream and the byte contract catches it, a second
generator makes the answer stochastic and adds a seed to declare, a probability
threshold invents a constant. All three fabricate.

THE TWO ACCEPTANCE BRANCHES TURN OUT TO PRODUCE DISJOINT LABEL SETS, which is
what makes `contingent` a category and not a patch over an awkward case:

  ΔH <= 0  no draw was taken, u is NaN. Removing a term can push ΔH above zero,
           and what would have happened then needs a number nobody drew. No term
           can be shown to flip the move to REJECT here -- only to remove its
           certainty. Yields `none` or `contingent`, never a named term.
  ΔH > 0   a draw exists, every counterfactual is decidable, nothing is
           contingent. And only a term that HELPED can be decisive: removing one
           that hurt lowers ΔH and the move stays accepted. Yields `none`, one
           of the four, or `multiple`.

Asserted over 20,000 random term vectors rather than remarked on, because if
`contingent` ever appeared in the drawn branch the label would be answering two
questions under one name and the colour key would mean nothing.

INERT: `u = ΔH <= 0 ? NaN : rand(rng)` evaluates only its taken branch, so the
generator is consulted on exactly the moves it was before. Byte contract green,
plus a direct lattice-and-melanin comparison with and without the tally attached
and a different-seed control proving that comparison can fail.

THE MEASUREMENT, RUN BEFORE ANY COLORMAP, ON THE FIGURE CONFIG
(N=40, 6 cells/species, seed 42):

    100 MCS, 16037 accepted        400 MCS, 53603 accepted
      none        38.6%              none        66.0%
      contingent  49.4%              contingent  23.7%
      adh          7.5%              adh          6.2%
      vol          4.4%              vol          4.0%
      mel          0.01%             mel          0.04%
      rad          0.00%             rad          0.00%
      multiple     0.09%             multiple     0.13%

THE DIRECT RADIATION TERM DECIDED ZERO OF 53,603 ACCEPTED MOVES, and that is
structural rather than under-sampling. Only a helping term can be decisive, and
β_ion is negative for exactly two of seven species at -5e-5, so a radiation
contribution that helps is of order 5e-5 and flipping a draw needs u within
~1e-5 of the threshold. Expected flips over 53,603 moves: about 0.5. This is
section 6.2's "one part in 1e5" as a count instead of a ratio.

The melanin term -- the one section 6.2 calls dominant over radiation -- decided
23. Which is section 6.2's own last sentence, "hand-specified adhesion
differences remain larger than both", arriving as a measurement.

AND THE MEASUREMENT ARGUES AGAINST THE LAYER IT WAS TAKEN FOR. n_accepted has a
median of 3 per touched voxel and a 25th percentile of 2, over 6157 of 64000
voxels. A modal label over three samples across seven categories is noise, and
no opacity rule repairs that -- binding the qualification to the picture was the
plan, and the honest reading is that the qualification defeats the picture
rather than annotating it. Two labels are also 90% of the moves. No bundle layer
and no colormap are written here.

What the run does support is the histogram, which is a result: 10% of accepted
moves have any single decisive term at all, and the two radiation-linked terms
account for 0.04% of them.

Julia: 20,171 passed across seven testsets. checkpoint_io_tests.jl still errors
for want of HDF5 in this worktree's unresolved depot. NOT A CLEAN RUN.

Also documents the byte contract's detection floor in validate_serial.jl, where
the next reader will meet it: it is conditionally green with the condition
unstated, which is a third kind of check that cannot fail, and it is not
redundant with the exact check beside it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap
… map

Section 6.2 ranks the two radiation-derived terms at a single move and closes by
noting hand-specified adhesion differences are larger than both. The same
comparison over a run: for each ACCEPTED move, would removing one term have
reversed it. Table 4 is that count.

    term removed        100 MCS      400 MCS
    none (the sum)       38.56%       65.99%
    adh                   7.53%        6.15%
    vol                   4.43%        3.97%
    rad                   0.00%        0.00%
    mel                   0.01%        0.04%
    >1 term               0.09%        0.13%
    contingent           49.39%       23.72%

THE DIRECT RADIATION TERM REVERSED NONE OF 206,042 ACCEPTED MOVES across seeds
42, 43 and 44. That is not under-sampling and the section says why: only a term
that LOWERED Delta H can be decisive for an accepted move, beta_ion is negative
for two of seven species at -5e-5, and reversing would need the drawn variate
within ~1e-5 of the threshold -- about 0.5 expected reversals in 53,603. Zero is
the expected observation. It is Table 3's one-part-in-1e5 as a count instead of
a ratio, and it follows from Table 3 rather than adding to it.

Melanin reversed 23 of 53,603. Both restate the closing sentence that was
already there.

THE COMPLEMENT IS THE RESULT. Between 66% and 77% of accepted moves are reversed
by removing no single term, so the dynamics are carried by the sum and not owned
by a component.

NO SPATIAL MAP, AND THE ROW SAYS SO RATHER THAN THE MAP BEING QUIETLY ABSENT.
The per-voxel tally and the modal-label reduction exist and are tested; the
measurement they were built for is what argued against drawing them. At 400 MCS,
6157 of 64000 voxels received any accepted move, median 3 and p25 2, across
seven categories. A mode over three samples is noise and an opacity rule bound
to the count would annotate the noise rather than remove it. Two further options
were refused and PP-62-07 records both: a map restricted to the ~10% of moves
with a named decisive term would filter out the 90% where the sum carried the
move and show a model looking more term-driven than it is; reporting nothing
would have discarded a real corroboration of Table 3. The distribution is
reported, the crop is not.

NO NULL PANEL, AND NOT BECAUSE IT WAS FORGOTTEN. Section 6.3's null is a SEEDING
null -- mean radial position has a value at MCS 0 from placement alone, so "what
would this look like without dynamics" is well posed there. Decisiveness is a
property of accepted moves, which do not exist at seeding, and zeroing terms
gives a different model rather than a null. What is well posed is seed
robustness, and that is what the three-seed figure above is.

The histogram also lands in a register the figure guard can actually read. A
raster render carries almost no extractable text, so RETRACTED_IN_FIGURES would
have gone silent on it and only the sha256 tier would have bitten. A table is
prose in the .tex and the ordinary phrase guard covers it.

Five ledger rows. 371 calibration, 313 coupling, 7 contract. Coupling's one
failure and one error are the pre-existing HDF5 gap in this worktree's
unresolved depot; the six skips are openmc. Suites now run with PYTHONPATH set
to this worktree -- the venv's editable install points at the main checkout, so
earlier coupling counts in this branch measured the wrong tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap
The draft described a three-page memo. It is four pages: the corrosion argument
went in between the empirical fact and the questions. A letter that miscounts
its own attachment is the same class of defect as a figure the prose retracts.

Adds the Section 6.2 table and the terminology tightening as brief additions
rather than a changelog, and makes the figure claim specific now that it has
been checked against the committed sidecars: five in-plot strings across the two
images -- both plot titles, Figure 1's two band labels, and Figure 2's
annotation. Figure 1's bands were reversed as well as mislabelled, which the
note already said and the diff confirms.

Still unsent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4m1NqS1u9tuDRaqmNoQap

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 77125683db

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread preprint/modeling_radioresistance_and_radiotropic_fitness.tex Outdated
Comment thread coupling/biofilm_openmc/observer.py Outdated
aurascoper and others added 5 commits August 29, 2026 17:12
… misreport it

The 8-colour sweep can carry one defect of its own -- a parity-correlated bias in
accepted moves -- and NOTHING IN THE REPOSITORY WOULD HAVE REPORTED IT.
jacc_port_tests.jl compares kernels on identical inputs, so it passes whenever
BOTH kernels carry the artifact, and tests/fixtures/serial_seed42.csv pins the
serial stream, which has no sublattices at all. The port was byte-identical on
both branches and had no acceptance instrumentation of any kind.

NO SPATIAL MAP, AND FOR d404438'S REASON RATHER THAN A NEW ONE. That commit
refused a per-voxel map for the decisive-label tally at 6157 of 64000 voxels
touched, median 3 -- a mode over three samples is noise. The argument carries,
and the conclusion here is stronger: a parity bias is a GLOBAL count comparison,
so the map was never the instrument for it. cpm_color! gains two write-only
per-site arrays, reduced every sweep and never drawn.

FOUR CHOICES THAT EACH AVOID A FALSE FAILURE ON THE FIRST RUN.

  `st` IS A DISCRIMINATOR, NOT A SENTINEL. 0 never proposed, 1
  evaluated-rejected, 2 evaluated-accepted. A NaN inside `dh` would have been
  indistinguishable from a site never proposed, shrinking the denominator of
  every acceptance rate and raising it silently. The spatial class is NOT
  stored -- it is derived from (tx,ty,tz), which is what keeps it spatial when
  the pass order is permuted.

  THE TABLE IS CONDITIONED ON OPPORTUNITY. The early returns -- wall, same-sigma,
  medium-into-medium, out-of-bounds -- are geometry-dependent, so a uniform null
  over eight classes would report the shape of the domain as a decomposition
  artifact. The 2x8 accepted/rejected table asks about the acceptance RATE.

  THE THRESHOLDS ARE EFFECT SIZES. n is 1.3e5 to 3.1e5 evaluated proposals per
  run and the cells are not independent -- an accepted move changes the lattice
  for every later pass -- so chi-square over-disperses from autocorrelation
  alone. At N=40 it ran 8.3 to 59.8 against a df=7 critical value of 18.475
  while Cramer's V never exceeded 0.014. V and the max per-class rate deviation
  are asserted; chi-square and n are reported beside them, never alone.

  color_order SEPARATES THE DECOMPOSITION FROM THE vols STALENESS. The colour
  loop is sequential and vols accumulates across passes (delta_H reads it at
  :167 while cpm_color! mutates it at :208), so the first pass evaluates against
  sweep-start volumes and the last against volumes moved by seven passes -- a
  deterministic, parity-correlated difference with nothing to do with the
  checkerboard, and `c` indexes both spatial class and sequence position.
  REVERSAL WAS REFUSED: position(c) = 7-c separates a monotonic position effect
  and is invariant to any effect symmetric about the midpoint, so one
  transformation leaves one blind spot. The RNG step key stays mcs*8 + c, keyed
  to the colour and not its position, so permuting changes the pass order and
  nothing else.

THE RESULT IS THE THIRD DISPOSITION, AND IT IS REPORTED AS SUCH. V 0.0050-0.0111
and max rate deviation 0.017-0.049 over three seeds x three orderings: no
decomposition artifact, and the residual tracks NEITHER spatial class NOR pass
position. The permuted run is a different trajectory rather than the same system
observed differently, so that outcome was reachable and the tier asserts the
bound and prints both rate vectors rather than claiming an attribution.

VERIFIED CROSS-VERSION, BECAUSE IT IS A CROSS-VERSION CLAIM. Comparing
on_sweep=nothing against an instrumented run at one commit compares the new code
to itself: the writes are unconditional, so both paths are identical kernel code
and the assertion is vacuous. The check that means something ran the port at
7712568 and after -- threads backend, 1 thread, seed 42 lattice sha256
3f528fab5b725dfd both sides (4884 occupied), seed 43 a4a8972857d9a893 (4883).

AND ITS POSITIVE CONTROL. Two agreeing runs cannot establish determinism on a
bounded race; it may simply not have manifested. The same seed at 4 threads gave
three different lattices in three runs (42a1815c, f42c3bf4, a824b7ef), so the
race is real, manifests, and the check detects it. The port is NOT reproducible
across thread counts and no lattice fixture is committed: it would sit in a
compare-never-regenerate directory while being thread-count-dependent.

The guard's own control is synthetic and in-file: one class held 20 percent low
must clear both thresholds. A threshold nobody has seen fail is not a guard.

fig5 carries the regime diagnostic the lattice picture cannot give -- acceptance
rate per sweep and the pooled Delta H distribution, both distributions, neither
a map. Pooled rate 0.062 at N=40; the per-sweep maximum of 0.414 is an
initialization transient, which is why the tier compares first half against
second before believing any pooled number. The provenance line names the run in
the .txt sidecar the phrase guard reads, since the test spans three seeds and
three orderings and would otherwise be about a different thing.

Six ledger rows, two census entries. Julia suite green including the byte-level
serial contract; calibration 371, contract 7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
…an find it

Two annotations printed on top of each other and neither was readable. THAT WAS
ARITHMETIC, NOT A RENDERING FLUKE, WHICH IS WHY IT RECURRED ON EVERY
REGENERATION: both were placed at x = 0.55*t_end from FINAL values --
m_final + 0.05 = 0.829 on the left axis, Peff_final * 0.88 = 2.385 on the right
-- and those two independent numbers on two independently autoscaled axes both
mapped to 79 percent of plot height. The fix separates them HORIZONTALLY, at
0.35 and 0.80 of the run, and anchors each to its own curve at its own x. A
vertical nudge would have fixed this dataset and left the coincidence in place.

THE ARTIFACT COULD NOT BE REBUILT, WHICH IS WHY IT WAS ALSO STALE. `julia
--project=. biofilms_potts.jl` runs main_coupled(), which hits RADIODIALYSIS:
BLOCKED -- the coupled loop reconstructs X_total from occupancy. FIG-03 already
records that blocker, and states it of figs 3 AND 4. It is true of one.

  fig3 plots m(t) and P_eff/P0. dm/dt = -k_dam*Ddot_R*m and
  P_eff = P0*exp(alpha_P*Ddot_R*t). Neither reads X_total or X_red.
  fig4 plots c_wall, c_mean, s_mean. The gated basis enters at exactly one
  place, uptake = k_ads*X_total + k_red*X_red, which drives c and s only.

regenerate_fig3.jl renders into a temp directory and copies one file across.
IT PROVES THE EXEMPTION RATHER THAN ASSERTING IT, and the control bites in both
directions so it cannot pass by measuring nothing: at 10x the uptake constants
fig3's two series must come back identical AND fig4's c_wall must move, and the
copy is refused if either half fails. Both held. fig4 is untouched and FIG-03
stands. It is committed rather than left in a scratch directory because a
generator nobody can run is the liability 1b51126 was written about.

AND THE REBUILD RETIRED A CLAIM THAT HAD SURVIVED FIFTEEN DAYS. d53f236
(2026-08-14) dropped "(50 Gy cumulative)" from the generator and RM-KR-01
already carried the verdict -- never print Gy there, since D_cum = Ddot_R*t is
dimensionless model time and no calibration converts it. The PDF went on
printing it because nothing could rebuild it: 1b51126 retired the same defect in
figs 1-2 through main()'s --no-radiolysis path and could only add .sha256/.txt
sidecars over a fig3 it had no way to regenerate. m = 0.779 stays -- that value
is arithmetically exact and RM-KR-02 keeps it. Only the Gy gloss is retracted.

NO GUARD COULD HAVE FOUND IT, AND THE FIRST FIX MADE THAT WORSE. Three words, so
MIN_WORDS=5 puts it out of the phrase guard's reach; no "radiotroph", so
RETRACTED_IN_FIGURES missed it too. FIG-05 was covered by NOTHING. The first
version of that row wrote the claim as "(50 Gy cumulative), printed beside
m = 0.779" -- five words of my own prose, which distinguishing_phrase duly
extracted, so the phrase guard would have searched the sidecar for a run that
was never in the figure and passed. Manufactured coverage, caught by
test_the_figure_rows_are_honest_about_being_vocabulary_only, which exists for
exactly that. The row is now the in-plot string alone. Filing FIG-05 in that
test's vocabulary-only list without more would have been the same lie one level
up: a list named for rows the word list covers, holding a row it does not.

So the word list covers it now. "gy cumulative" is a UNITS term in a tuple
otherwise about a phenotype, which the note says out loud, and
fig3_membrane_transport_prefix.{pdf,txt} from e24dbec is the committed
known-bad input that proves it bites -- mutating the tuple back to
("radiotroph",) drops fig3 from the hits and the control fails. Committed, not
recovered with `git show`, for the reason the fixtures README already gives.

The stale print at test_claims_ledger.py:366 still enumerated "FIG-01, FIG-02"
one screen from the list it contradicts. Fixed in the same pass.

The generator edit is below the `#  13. Figure export` split marker, so
validate_serial.jl, runtests.jl and import_dose_field.jl cannot reach it and the
byte-pinned serial stream is untouchable by it. Confirmed green regardless.

Three ledger rows, one census entry. Julia suite green including the serial
contract; calibration 371, contract 7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
Codex P1 on #23, reproduced and confirmed. `plot_panels` imported pyvista above
its own argument check, so in any environment without the `viewer` extra --
which is every environment the suite actually runs in -- `plot_panels([], ...)`
raised ModuleNotFoundError and never reached the ValueError it promises. The
import now sits below the check.

`test_plot_panels_refuses_an_empty_panel_list` was unskipped and asserting that
ValueError, and it PASSED THROUGHOUT, because a tier that has pyvista installed
reaches the check no matter which line comes first. It could not fail for the
defect it was written for. Rule 2 in the shape rule 2 warns about: the skip was
the coverage, and removing the skip did not add any.

So the control blocks the module itself rather than trusting the runner, and
bites on both tiers. Two details in it are load-bearing:

  IT ASSERTS THE BLOCK IS IN FORCE BEFORE TESTING ANYTHING. My first attempt at
  this reproduction used a `find_module` hook, which modern Python never calls,
  so the block silently did nothing and `plot_panels([], ...)` returned a clean
  ValueError -- a false negative that would have had me dismiss a valid P1 as
  unreproducible. A meta-path hook that does nothing looks exactly like a fix.

  IT RESTORES sys.modules AND sys.meta_path IN A finally. Leaving pyvista
  blocked would skip or fail the eight viewer-gated tests downstream of it and
  the cause would not be local to anything.

Verified by mutation: restoring the import above the check fails exactly this
new test and nothing else -- 1 failed, 51 passed -- while the original empty-list
test stays green, which is the whole point. Coupling suite 316 passed, 6 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
…e next one

Codex P1 on #23, reproduced and confirmed, and it reaches further than the
sentence it names.

THE BOUND. Section 6.2 asserted that beta_ion being negative for two of seven
species at -5e-5 made that "the largest radiation contribution that could favour
an accepted move". Delta H_rad is signed by ROLE as well as by the coefficient:
compute_delta_H_terms adds +beta_ion[source]*I when the source gains a site and
SUBTRACTS beta_ion[target]*I when the target loses one, so a POSITIVELY signed
species vacating a site favours acceptance exactly as much as a negatively
signed one occupying it. The bound is max|beta_ion| = 7.5e-2, not min. Wrong by
1.5e3. Measured over 192996 evaluated proposals from the seed-42 state at 100
MCS, 29.8 percent carry Delta H_rad < -5e-5 and the extreme is -0.0671. Table 3
gains the row whose absence made the error possible -- the same coefficient in
the other role, at -7.5e-2 and an acceptance bias of 1.0151 rather than 0.985.

THE COUNT MOVES TOO, AND THAT IS NOT WHAT CODEX ASKED ABOUT. PP-62-04 published
"zero of 206042". Re-measured, it is ONE of 206042: none under seeds 42 and 43,
one under seed 44. The denominator is exactly right.

FINDING THE RUN TOOK TWO ATTEMPTS AND IS ITS OWN DEFECT. run_simulation and
run_simulation_coupled both seed MersenneTwister and give 14281 accepted moves
at 100 MCS; Table 4 says 16037. That is reproduced only by the idiom in
tests/delta_h_decomposition.jl -- mcs_step! driven by hand with
Random.Xoshiro(seed) and update_melanin! each sweep. Neither b39fb8a nor d404438
committed a harness and nothing records the configuration, so the published
table describes a trajectory no shipped entry point produces. Ledgered as
PP-62-11. With it identified, EVERY cell of Table 4 reproduces exactly, along
with PP-62-05's 23 of 53603 and PP-62-06's per-seed contingent shares -- which
is what makes the rad row's disagreement a finding rather than a different run.

THE REAL RESULT IS ABSORPTION, NOT ABSENCE. Removing Delta H_rad alone reverses
16 accepted moves (4/3/9 by seed). Fifteen of them also carry an independently
decisive adhesion or volume term, so decisive_label returns `multiple` and the
`rad` row reads 0.00 percent while the term was in fact capable of deciding
sixteen. The corrected arithmetic predicts 16.67 and the withdrawn one predicts
0.26; 16 were measured. That agreement is the evidence the arithmetic is now
right, and an earlier estimate of mine put it at 58-76 by assuming most accepted
moves sit in the drawn branch -- only 11 to 16 percent do.

A CORRECTION THAT ONLY CHANGES THE NUMBER TEACHES NOBODY WHY IT WAS WRONG, so
tests/prose_bounds.jl reads max|beta_ion|*I0 off CPMParams and the stated bound
off the .tex and fails when they disagree. Nothing could have done that before:
contract_csv.jl guards the trajectory, delta_h_decomposition.jl guards that the
terms sum to the scalar, test_claims_ledger.py guards retracted PHRASES. None
reads a NUMBER out of the prose. Verified by mutation -- reverting the .tex to
5e-5 fails with isapprox(5.0e-5, 0.075). Its control asserts the withdrawn
sentence yields NO bound rather than a wrong one, explicitly, because otherwise
"no bound stated" and "the right bound stated" are indistinguishable, which is
the shape of gap that let this through. SCOPE STATED IN THE FILE: it gates one
number, not prose against code in general.

Version 1.2 correction block names both moved numbers and scopes its own
reassurance to the entries actually re-measured. It also fixes a sentence my own
fbb264b invalidated: the v1.1 block said figures 3 AND 4 carry a "per cent
depleted" label; only figure 4 does, and figure 3 has since been rebuilt.

PP-62-04 restated; PP-62-09, -10, -11 added. Julia suite green including the
byte-level serial contract; calibration 371, contract 7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
PP-62-11's required_to_fix, both halves. The table was published from a run
NOTHING SHIPPED COULD RE-EXECUTE: run_simulation and run_simulation_coupled both
seed MersenneTwister and give 14281 accepted moves at 100 MCS against the
table's 16037, and the run that produces 16037 is the idiom in
tests/delta_h_decomposition.jl -- mcs_step! driven by hand with
Random.Xoshiro(seed) and update_melanin! each sweep. Neither b39fb8a nor d404438
committed a harness and no file recorded the parameters, so for two weeks the
published table could not be checked against anything.

decided_moves.jl reproduces it exactly: both columns in 8 s, all three seeds at
400 MCS in 13 s. The Table 4 caption now names the generator and the
Xoshiro/MersenneTwister distinction, so the configuration is recorded beside the
number as well as executable.

IT USES ONLY THE SHIPPED API. mcs_step! already takes `driver` and DriverCounts
already accumulates the labels, so decisive_label is called by the model rather
than reimplemented in the harness and the table cannot drift from the rule that
produces it. No radiodialysis, so no basis gate acknowledgement and no new
census site -- the CPM trajectory never reads the nutrient field, which is also
why the coupled and uncoupled paths give identical accepted counts.

IT STATES WHAT IT MUST REPRODUCE. The published counts are in the file and the
script exits nonzero on mismatch, rather than printing numbers with nothing to
check them against -- which is the condition the table was in.

tests/decided_moves_tests.jl makes that automatic instead of something someone
has to remember to run, and READS ITS EXPECTED COUNTS OUT OF THE SCRIPT'S OWN
TABLE rather than restating them: a second copy would let script and suite drift
and each look green. It pins the two numbers version 1.2 moved -- the 206042
denominator and rad == 1, not zero -- and its control steps the same run under
MersenneTwister and requires the totals to DISAGREE, which is the arm that
proves the comparison can fail.

One correction to my own hand: the caption first read 14037 where the measured
MersenneTwister count is 14281. A wrong number inside a correction about a wrong
number. Fixed, and re-verified against the code rather than against my memory
of it.

Julia suite green, +10 assertions; calibration 371, contract 7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
@aurascoper

Copy link
Copy Markdown
Owner Author

@codex review

Re-review request: your last review covered 7712568, head is now 3267cdd (five commits). Both P1s from that review are fixed and resolved:

  • Empty panels before PyVista (observer.py) — reproduced and fixed in 8a7c28f. The pre-existing test could never have caught it, so it ships with a control that blocks the module itself; mutation-checked.
  • Target-cell losses in the radiation bound (§6.2) — reproduced and fixed in 5788db5. Confirmed, and it went further than the sentence: PP-62-04's "zero of 206042" is one of 206042. Bound wrong by 1.5e3, count wrong by one, corrected arithmetic now predicts 16.67 against 16 measured.

New since 7712568, and the parts most worth adversarial attention:

commit what to doubt
37c0b33 JACC checkerboard parity measurement. Thresholds are observed, not derived (Cramér's V < 0.025, maxdev < 0.12) from 3 seeds × 3 orderings at N=20/50 MCS. The color_order permutation is meant to separate a decomposition artifact from the documented vols cross-pass staleness — please check that separation actually holds.
fbb264b regenerate_fig3.jl opens basis_gate_ack on the claim that fig3's m(t)/P_eff never read the gated basis. Control is a 10× uptake perturbation that must move fig4 and not fig3.
5788db5 tests/prose_bounds.jl gates one number and says so. Manuscript at version 1.2.
3267cdd decided_moves.jl reproduces Table 4 exactly — its configuration matched no shipped entry point until now (PP-62-11).

Suites: Julia green incl. the byte-level serial contract, calibration 371, contract 7, coupling 316/6 skipped. preflight_merge.sh 23 refuses on STALE REVIEW only, which is what this comment is for.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3267cddc22

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/correspondence/wan_v11_note.md Outdated
Comment thread tests/prose_bounds.jl Outdated
…ited

Codex raised two findings on 3267cdd and both reproduce.

P1, docs/correspondence/wan_v11_note.md. The draft still told Caixia Wan that
the direct radiation term "reversed none of 206,042 moves". That is the version
1.1 count the version 1.2 correction withdrew. The note is unsent so nothing has
gone out, but it also said "nothing that changes a result" and offered to attach
the paper: it would have arrived carrying the withdrawn number alongside a
correction block stating that two published numbers moved. Step 4 of the
correction protocol is audit dependents, and it ran over the manuscript and the
ledger and not over correspondence.

P2, tests/prose_bounds.jl -- the guard written to stop exactly this class of
defect -- computed max|beta_ion| = 7.5e-2. For a copy between two OCCUPIED
parcels Delta H_rad is (beta_source - beta_target)*I, so the acceptance-favouring
reach is an extremum over PAIRINGS and not over species:
max(0, max beta) + max(0, -min beta) = 7.505e-2. The guard did not merely miss
the larger value. It asserted equality at rtol 1e-6 and would have REJECTED the
correct one.

ATTAINED, NOT ONLY DERIVED. A C. sphaerospermum source (-5e-5) copying into an
S. oneidensis target (7.5e-2) reaches -0.07505, four times in 1298668 evaluated
proposals at N=20, seed 42, 400 MCS. The three N=40 runs of Table 4 do not reach
it, which is why the single-role figure went unchallenged by the measurement
that accompanied it.

AND A THIRD, FOUND BY MEASURING THE SECOND. Section 6.2 said "measured over the
run, 29.8% of evaluated proposals carry Delta H_rad < -5e-5". No run was named
and none reproduces it. The fraction is configuration-dependent -- 0.1442 at
N=20/400 MCS, 0.2960 at N=40/60 MCS -- so naming the run is not a formality.
Restated as 26.7% over 3968838 evaluated proposals across seeds 42, 43 and 44 at
400 MCS, N=40: the Table 4 runs, now named in the sentence itself. This is
PP-62-11's own defect committed inside the fix for PP-62-04.

The instrumentation is an in-memory source rewrite that never edits the repo,
proved inert by lattice hash against the pristine module with a different-seed
control.

THE GATE IS MUTATION-CHECKED. Reverting the .tex to 7.5e-2 fails it at the
isapprox assertion; the file was restored byte-identical afterwards. The
extractor now refuses the max_s |beta| SHAPE whatever number it carries, because
a sentence bounding one role at a time is not a bound on Delta H_rad and reading
a number out of it would launder that.

Ledger: PP-62-09 to restate (0.075 -> 0.07505), PP-62-12 added for the
percentage, PP-62-04's notes record that its own evidence line came from an
unnamed run. My first draft of those rows used status "restated", which is not
in the verdict vocabulary, and put the corrected value in claim_text where the
row records the claim as published; test_every_row_has_an_id_and_a_verdict
caught both.

Julia suites green including prose bounds 18/18 and Table 4 10/10; calibration
371, contract 7, coupling 316 passed / 6 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
@aurascoper

aurascoper commented Aug 30, 2026

Copy link
Copy Markdown
Owner Author

@codex review

Re-review request: your last review covered 3267cdd, head is now c38c2a3.

Both findings from that pass confirmed and fixed in one commit. What to doubt in it:

The bound itself tests/prose_bounds.jl now computes max(0, max β) + max(0, -min β), not (max β - min β). The claim behind that shape is that an absent parcel contributes nothing to the sum, so a coefficient vector of one sign would take its reach from one role alone. If that reasoning is wrong the two forms differ only for coefficients this repo does not ship, which is exactly the kind of guard that looks fine until it isn't.
Attainability I assert the pairwise extremum is reached, on evidence: a CS source (-5e-5) into an SO target (7.5e-2) gives -0.07505, four times in 1298668 evaluated proposals at N=20, seed 42, 400 MCS. Measured by in-memory source rewrite, proved inert by lattice hash with a different-seed control. The test itself only asserts argmax(β_ion) != argmin(β_ion), which is necessary and not sufficient.
The percentage I restated Measuring your P2 turned up a third error that was mine, not in your report: §6.2 said 29.8% of proposals carry ΔH_rad < -5e-5, "measured over the run", naming no run, and nothing reproduces it. Now 26.7% over 3968838 proposals across seeds 42/43/44 at 400 MCS, N=40, with the configuration in the sentence. Worth checking that the number and the named run actually correspond, since the previous version of this sentence is the reason to doubt it.
The Wan note Corrected count, marked rather than swapped. The reassurance sentences that became false are removed or scoped. Please check I did not leave a third one.

Two things you did not ask about that changed anyway. The built preprint PDF on my machine was still a version 1.1 build carrying the withdrawn count, which is the file that note offers to attach; rebuilt from source, still untracked per 1c43802. And my first draft of the ledger rows used a status outside the verdict vocabulary and put corrected values where the row records the claim as published — test_every_row_has_an_id_and_a_verdict caught it, which is the guard doing its job rather than me noticing.

The gate is mutation-checked: reverting the .tex to 7.5e-2 fails it at the isapprox line.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c38c2a33f8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread preprint/modeling_radioresistance_and_radiotropic_fitness.tex
Comment thread tests/prose_bounds.jl Outdated
aurascoper and others added 12 commits August 29, 2026 20:22
The memo's page-1 box said "No number changed." That was true of version 1.1,
which replaced two figures, and false the moment version 1.2 corrected the
Delta H_rad bound and the two counts resting on it.

THIS IS THE SAME DEFECT CODEX RAISED AS A P1 ON THE WAN NOTE, in the form that
is harder to find. That note republished the withdrawn NUMBER, so grepping for
206042 found it. The memo republishes the withdrawn REASSURANCE and names no
number at all, so the same grep came back clean. A dependent-document audit that
searches for values misses every sentence of the form "nothing changed".

Both documents are addressed to the same two people about the same paper, and
the memo is the one that says "you would have found it on page 8" -- a reader of
version 1.2 finds two corrections on that page, and the memo was offering to
pre-empt one of them.

Scoped to the figures, and the numerical correction stated plainly beside it.
Rebuilt: still four pages, box renders without overflow.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
… extractor

Codex raised two P1s on c38c2a3 and both reproduce.

P1, the 26.7% had no producer. PP-62-12 named decided_moves.jl as its evidence,
and decided_moves.jl counts accepted moves by decisive label; it never touches
evaluated proposals or tests rad < -5e-5. The measurement behind the
restatement was an in-memory rewrite of the stepper that was never committed.
THAT IS PP-62-11's DEFECT COMMITTED INSIDE THE FIX FOR PP-62-04 -- a corrected
number nobody can re-run, published in the commit that corrected a number
nobody could re-run.

rad_proposals.jl is the producer. It uses a new `on_proposal` keyword on
mcs_step!, default nothing, called for every evaluated proposal after the skip
conditions and before the acceptance draw.

  seed      evaluated      < -5e-5     frac     min dH_rad     max dH_rad
  42          1287796       358404   0.2783   -0.071342207    0.071342207
  43          1340171       344591   0.2571   -0.061028003    0.064553098
  44          1340871       357192   0.2664         -0.075        0.07505
  pooled      3968838      1060187   0.2671         -0.075        0.07505

THE FIRST DRAFT OF `PUBLISHED` WAS NOT MEASURED. I back-computed the per-seed
counts from the reported fractions -- 1287796 * 0.2783 -- which is a number
shaped like evidence and is not one. Running the producer printed the real
counts and they disagreed in the third digit.

THIS HOOK DOES NOT INHERIT THE JACC PORT'S GUARANTEE. There, the RNG is
counter-based on (seed, step, site) and has no stream position to advance, so
instrumentation could not move the trajectory as a matter of construction.
biofilms_potts.jl draws from a sequential generator and `on_proposal` sits
immediately before `u = ΔH <= 0 ? NaN : rand(rng)`, the most sensitive position
in the loop. Inertness here is DEMONSTRATED AT TWO SEEDS AND ONE LATTICE SIZE,
NOT PROVEN.

So it carries two controls doing two different jobs, which the first draft
conflated:
  - seed 6, no hook: the comparison varies with input at all, so equality at
    seed 5 is not vacuous. Says nothing about the hook channel.
  - a hook that deliberately draws from the same generator: the hook channel
    CAN perturb, so the null at seed 5 is evidence about `on_proposal` rather
    than about the harness.

Both readings of a null were written before it was run. A non-diverging drawing
hook is ambiguous between "the comparison is blind" and "`on_proposal` has no
reach into the stream", and object identity separates them -- hence the
`captured[] === rng5` assertion ahead of the divergence check.

validate_serial.jl 42 before and after: 8 CSV lines, identical byte-for-byte.
The byte-pinned serial contract passes on both sides.

P2, the prose extractor matched too little. It required only that \max_{s,t} be
followed somewhere by \beta_{t,...} and the right literal, so
`\max_{s,t}\beta_t = 7.505e-2` passed every assertion and a semantic regression
could keep the corrected number while invalidating the bound it names. It now
requires the shape in full -- both subscripts in their roles, the subtraction
between them in that order, and I_gamma -- with four known-bad controls that
each carry the CORRECT literal and state a different quantity, plus a positive
control so the four cannot pass by the extractor being broken.

THE FIRST VERSION OF THIS MESSAGE DESCRIBED THE P2 FIX BEFORE IT EXISTED. The
commit contained four files and none was prose_bounds.jl; the paragraph above
was written from the plan rather than from the diff. Caught by reading
`git show --stat` against the claim. Amended before push, so nothing shipped,
but the failure mode is the one this repository keeps finding: a description
that agrees with the intention instead of with the artifact.

Julia suites green: prose bounds 23/23, proposal statistics 26/26, Table 4
10/10, serial contract byte-level.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
Stage 1 of the memo revision read eight primaries. It killed almost everything
and left one claim standing, which is the outcome that makes the claim worth
making.

WHAT VERIFICATION REFUTED. Every neighbouring sentence the memo could have
carried turns out to be already answered:

  nobody studied FeCrAl in pool water   -> Xiao/Jang, J Nucl Mater 2021,
                                           commercial APMT, three prior-corrosion
                                           histories, in simulated SFP water
  nobody studied irradiated FeCrAl      -> Corros Sci 174:108824 (2020), severe
    in water                               later-stage corrosion from
                                           irradiation-induced defects
  nobody characterised the film on      -> APMT layer-by-layer XPS depth profiling
    commercial grades
  Al plays no role in aqueous films     -> Al-Cr-rich transition layer, borated
    on commercial FeCrAl                   and lithiated water, 360 C
  SFP MIC is a realised threat          -> SRS ms9800702: "This assessment was
                                           found to be in error"
  a biofilm would make it worse         -> Shukla/Karley/Rao, Arch Microbiol
                                           207:131 (2025), EIS shows biofilm-
                                           mediated INHIBITION on SS-304L

WHAT SURVIVED. A web search on 2026-08-30 for `FeCrAl microbiologically
influenced corrosion biofilm` returned no study of FeCrAl. The same term set with
one word changed to `304L stainless steel` returned four on its first page. Every
substrate that came back across both passivates through chromia, or is aluminium,
or is carbon steel. FeCrAl passivates through a mixed Cr(III)/Al(III) film, which
is why the gap is interesting rather than merely true.

THE DRAFT CARRIED A FALSE ABSENCE AND ITS OWN SEARCH KILLED IT. "No aqueous-
corrosion study of irradiated FeCrAl found" was in the plan; the first page of
the search for it returned three. It was caught only because that claim was given
a SEPARATE search -- it shares no terms with the microbial one, and letting it
borrow that scope would have been false as well as wrong. The replacement is
stronger: irradiation makes aqueous corrosion of FeCrAl worse, that is sourced,
and it was established at reactor conditions. HOFFMAN-13.

A BRIEF MIS-CITATION, RECORDED SO IT IS NOT INHERITED. The research brief
attributed the third-element-effect work partly to arXiv 2302.07988, which is
about Al-Co-Cr-Fe-Ni compositionally complex alloys, not FeCrAl, and makes no
passive-film claim. It shares two authors with the real source. Third attribution
error from a secondhand brief on this memo, after HOFFMAN-01 and HOFFMAN-02.

THE GATE, AND WHAT IT FOUND IN PROSE THAT ALREADY SHIPPED. tools/absence_gate.py
greps a diff for constructions that carry an absence -- bare negation, universals,
universals in positive grammar -- and a human judges the candidates. It has
controls in BOTH directions, because a gate that flags everything scores
perfectly on known-bads and is useless: five must flag, four ordinary page-2
sentences must come back clean.

On its first run over the existing memo it found "A facility is also the ONLY
SETTING where the inputs to it are obtainable", a universal about all settings
that had been unread by anything since version 1.1 and that happens to flatter
the recipient's own kind of facility. Restated, HOFFMAN-15. It also produced its
expected false positive, "A biofilm is none of the three", which is a category
statement needing no scope.

Run over this diff it caught three more: a HEADING that asserted more than the
body it introduced ("The substrate nobody has tested" -> "The mixed Cr/Al case",
since a heading cannot carry scope and is what gets remembered), and two absences
whose scope came from a DIFFERENT search than the one they sat next to. And it
missed a fourth, because the patterns had `not found` and not `did not find` --
patched, controls re-run.

Layout followed the claims rather than deciding them. The additions belong on
page 2 where the "what that literature does not compute" argument already lives;
page 3 is declared questions and putting a finding there converts it to an ask by
typography. Room was made by moving "The stressor set" to the foot of the left
column, which leaves left = what is known, right = what is not. Four pages, no
overfull boxes.

Ledger: HOFFMAN-11 gains a fourth part recording that the attribution question
drew no new evidence AND WHY -- source unreachable from here, which is a
different state from checked-and-found-nothing and the one that changes when
someone with AMPP access looks. Verdict holds; the claim has moved twice and is
recorded as unstable. HOFFMAN-12 through 16 added.

Julia suites green; calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
…and the contrast §2.6 was missing

THE IRRADIATED SENTENCE WAS WRONG AND SHIPPED AN HOUR EARLIER. The memo said
"irradiation makes aqueous corrosion of FeCrAl worse", footnoted to Corros. Sci.
174:108824. Reading that paper's METHOD, rather than a summary of its result,
shows three things the sentence dropped:

  - MODE: 400-keV Fe+ IONS at 400 C, not neutrons. At that energy the damage is
    a near-surface layer a few hundred nm deep, so transfer to a real assembly
    is far weaker than the sentence implied.
  - STAGE: the irradiated alloy corroded MORE SLOWLY at first -- a Cr-rich oxide
    and an ion-sputtered smooth surface suppressing it -- with severe corrosion
    only later. Part of that is a sample-preparation artifact.
  - The footnote attached NEUTRON alpha-prime work to the same claim, putting
    two irradiation modes in one citation.

THE FIX IS NOT TO DROP THE CITATION. Removing it and leaving alpha-prime plus
"what makes the pool question a real question" would let a reader infer nobody
has looked, when someone has, badly, in the wrong conditions. Weak evidence and
no evidence are different states and that distinction is this memo's discipline.
The sentence now says the work exists, at reactor conditions and without a film,
and the residual gap names all THREE conditions, because naming only the
temperature would let it read as a thermal complaint. Alpha-prime is stated as
microstructure and nothing more: that source reports no passive-film
composition, so "changes the surface a film would sit on" is an inference it
does not make and is not written.

FOURTH INSTANCE ON THIS BRANCH OF ONE FAMILY, and the family is the dangerous
one: HOFFMAN-01 (invented "34 months" and a plant name), HOFFMAN-02 (wrong paper,
wrong first author), a commit message describing a fix the commit did not
contain, and now this. A claim written from a summary instead of the source
produces CONFIDENT FALSE SPECIFICS, which a reader cannot discount the way they
discount an overreach. It is a different failure from the one the absence gate
catches.

THE GATE NOW PRINTS ITS OWN INCOMPLETENESS. A docstring line is unenforced and
unread, so KNOWN_UNHANDLED carries the forms it is known to miss -- "remains
uncharacterized", "the literature is silent on", "the only work is" -- asserted
NOT to flag and reported as a standing gap count on every run. If one starts
flagging, that prints as an UNEXPECTED PASS. A clean run over a document can no
longer be read as an all-clear. The list exists because the first version matched
`not found` and missed `did not find`; and the phrasing considered for the
irradiated fix, "makes the pool question a real question", carries no absence
marker at all while functioning as an absence claim -- the same gap biting on the
same diff that records it.

SECTION 2.6 GAINS THE CONTRAST IT WAS MISSING. The section argues radiotrophy is
not established and never says what radiation-linked metabolism IS, so a reader
asking how anything lives in a radioactive pool found no answer. Added: nothing
feeds on the metal; organic carbon arrives from outside; primary production is
done by hydrogen-oxidising bacteria on radiolytic H2 and by algae under lighting,
both verified at Sellafield (Front. Microbiol. 11:587556, 2020; 14:1261801,
2023). It is stated as a CONTRAST and explicitly not as a defence -- it does not
address Walberg's carbon-budget critique, which concerns melanised fungi in
culture and stands as stated.

A CLAIM WITHDRAWN BEFORE IT WAS WRITTEN, WITH ITS DISPOSITION SET IN ADVANCE.
The paragraph was planned around "none of the seven uses CO2 as its carbon
source", to stop a reader carrying the autotrophy onto the modelled organisms.
Verification failed it for two species. S. oneidensis carries a cryptic reductive
glycine pathway that fixes CO2 on engineered electrons, and it is the species an
MIC reader checks first; Ochrobactrum sp. strains with Calvin-Benson-Bassham
fixation have been isolated, so the genus contains facultative autotrophs and AM7
was not checked individually. Neither capability exists in a pool, but the
blanket claim is checkable and wrong. The protective work is done instead by a
fact about the CODE: compute_delta_H_terms and mcs_step! never read the nutrient
field -- zero reads, counted -- so the model resolves no carbon at all. PP-26-03
records the withdrawn claim, because a claim dropped before writing leaves no
trace in the document and the reason it was dropped is the useful part.

Also corrected: the premise this pass was planned under held that all seven
species are fungi. biofilms_potts.jl:30-38 lists three fungi and four bacteria.

Memo: 4 pages, zero overfull boxes. Ledger: HOFFMAN-13 restated, PP-26-01
through -03 added. Julia suites green; calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
Three corrections to yesterday's §2.6 paragraph, all from the same review.

A COUNTED ZERO EXPIRES LIKE THE F_s GREP. The paragraph's protective claim --
the CPM acceptance path reads no nutrient -- rested on a count taken once.
Nothing stopped a later commit adding a nutrient term to compute_delta_H_terms,
at which point the manuscript's claim silently becomes false and the manuscript
is the artifact that cannot notice. tests/manuscript_claims_tests.jl already
reads the .tex and already carries this pattern for F_s and "fitness", so the
check goes there, scoped as narrowly as the code actually counted: the bodies of
compute_delta_H_terms and mcs_step! in biofilms_potts.jl, no broader.

Controls in three directions, since the file's own rule is that a check which
cannot fail is not a check:
  - a planted `state.nutrient` read must be detected before the check is trusted
    to find none;
  - a COMMENT naming the field must NOT trip it, or the check becomes a trap for
    the next person who documents why the field is absent there;
  - update_nutrient! MUST read the field. If that ever comes back empty the
    section is describing dead state and needs rewriting rather than re-passing.
Mutation-checked: one planted read in compute_delta_H_terms fails it, and the
file restored byte-identical.

ZERO READS IN TWO FUNCTIONS IS NOT ZERO READS IN THE PROGRAM, and checking the
difference changed what the section says. `state.nutrient` is initialised by
init_nutrient!, integrated every step by update_nutrient! and
update_nutrient_coupled!, serialised at export_checkpoint.jl:206, round-tripped
by checkpoint_io_tests.jl:102 and parity-checked by the JACC port at
biofilms_potts_jacc.jl:657. It is neither dead code nor a wired-but-unused
input. It is a LIVE field that the acceptance path does not consult, which is a
third state and the only one the evidence supports.
calibration/biofilm_calibration/spatial/time_observable.py:22 records the same
fact and predates this row.

"EXOGENOUS" WAS THE WRONG WORD AND IT DESCRIBED A DIFFERENT MODEL. The draft
said organic carbon is an exogenous input the model does not represent.
Exogenous implies an input arriving from outside; zero reads means there is no
input at all. §2.6 now says the field exists, is integrated, and is not consulted
by the acceptance rule, so the model does not resolve carbon as a driver of the
trajectory -- and does not treat it as an outside input either.

PP-26-03 GAINS THE SHAPE OF ITS FAILURE, which is the transferable part. The
withdrawn claim was narrowed twice before checking -- from "fungi cannot fix
CO2" to "none of the seven uses CO2 as its carbon source" -- with the two
riskiest species named in advance, and it failed at exactly those two anyway.
NARROWING A UNIVERSAL REDUCES ITS EXPOSURE WITHOUT CHANGING ITS TYPE: a
universal over seven named organisms is a search problem with seven independent
chances to fail, and picking the right two to worry about does not help if the
boundary lands in the wrong place. The repair was not a better boundary. It was
abandoning the universal for a property of a bounded artifact this repository
controls, which a test can hold and a literature claim about seven organisms
never could. Distinct family from HOFFMAN-13: that one produces confident false
specifics from not reading a source; this one produces a checkable universal
from a correct reading with the quantifier left too wide.

Manuscript claims 29 passing, up from 22. Julia suites green; calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
§2.6 says the model does not resolve carbon as a driver of THE TRAJECTORY. The
test held two function bodies. Stating at one scope and enforcing at another is
the F_s defect the same file exists to prevent, so the check now covers every
function in biofilms_potts.jl that writes the lattice.

THE HAND-WRITTEN SET WAS WRONG IN BOTH DIRECTIONS, AND THE ASSERTION THAT PINS
IT CAUGHT THAT ON ITS FIRST RUN. I listed divide_cell! and missed place_cell!.
divide_cell! mutates the lattice THROUGH place_cell! rather than directly, so a
manual grep for writers produced a set containing a function that does not write
and omitting one that does. The real set is compute_delta_H_terms (a trial write
and restore), init_state, mcs_step! and place_cell!. All four read zero nutrient,
so "the trajectory" is earned rather than asserted.

Pinning the set is the load-bearing part. Without it a new lattice mover would
inherit a claim made before it existed, and the check would keep reporting
"still absent" about a set that had quietly changed underneath it. Same shape as
requiring update_nutrient! to READ the field: both check the premise rather than
the detector, and both fail in the direction where the section needs rewriting
instead of where the test needs fixing.

A FINDING RECORDED WITH NO INDEX, WHICH COST A FULL VERIFICATION CYCLE.
calibration/biofilm_calibration/spatial/time_observable.py:22 already said
"state.nutrient is written but never read by the acceptance test", alongside
"divide_cell! has no trigger" and "there is no growth rate anywhere". It predates
this work by an unknown interval and was rediscovered by a fresh grep several
turns later, because a docstring is where a fact is RECORDED and not where
anything SURFACES it. That is recorded-versus-enforced, which is this
repository's own subject arriving from a direction nobody was watching. Other
observations of the same kind may sit in docstrings elsewhere; they would be
findings with no index, and nothing here has swept for them. Noted in PP-26-03.

Manuscript claims 36 passing, up from 29. Calibration green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
…d not docstring notes

Filed as PP-26-04 and PP-26-05. Checking them first changed what they are: both
are stated in the MANUSCRIPT, not only in the calibration docstring where they
were rediscovered.

The .tex says "divide_cell! has no trigger", "the program has neither growth nor
death", "a manually invoked division primitive for lifecycle and lineage
testing, but no growth law or automatic division trigger", and "still has no
growth". biofilms_potts.jl's header repeats it. AND SECTION 2.6 LEANS ON IT: the
argument that the model cannot test radiotrophy runs through having neither
growth nor death, so if either claim stopped holding the central retraction
would rest on a false premise. Nothing enforces either.

PP-26-04, divide_cell!, converts with the PP-26-02 pattern and needs one thing
that pattern did not: TWO INDEPENDENT WAYS TO GO FALSE AND A PINNED CALL-SITE
SET COVERS ONE. A second caller appearing is caught by a set assertion; the
EXISTING caller becoming reachable from the trajectory is not, and that is the
change that matters. A check built only from the set would pass through exactly
the event it exists to detect. The claim is also already looser than it reads --
there is a call at biofilms_potts.jl:1978 -- so the checkable statement is about
reachability and not about the absence of callers, the same distance as "no
nutrient reads" to "no nutrient reads in the acceptance path".

PP-26-05, the growth rate, does NOT convert cleanly. It is an absence over
unbounded naming, so a symbol grep bounds it only as far as its term set
reaches, and growth can be named mu, k_growth, doubling, biomass_rate and
onward. That is tools/absence_gate.py's printed incompleteness arriving as
absence CODE rather than absence prose. The useful form is weaker than the
sentence implies: the manuscript states the term set INLINE, the way the memo's
page-2 absence names its retrieval mode, and the test holds that same set.
Holding a bounded grep behind an unbounded sentence is how the test scope and
the sentence scope drift apart, which is the defect PP-26-02 was corrected for
within hours of being written.

WHY THESE ARE ROWS AND NOT A COMMENT. They were sitting in a session
transcript, which is the same failure as sitting in a docstring: recorded, in
the right register, indexed by nothing. time_observable.py:22 carried three
findings of this kind; the nutrient one cost a full verification cycle to
rediscover and is now PP-26-02 and enforced. These are the other two.

Calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
…our sentences is what showed it

PP-26-05 was filed as the hard case: an absence over unbounded naming that does
not convert cleanly. The concern was that section 2.6's central retraction leans
on it, so the weakest check would be guarding the strongest claim. It does not,
and the difference is in what the sentences actually say.

THE TWO LOAD-BEARING SENTENCES ARE ABOUT DYNAMICS, NOT ABOUT IDENTIFIERS:

  "The CPM ... has neither biomass growth nor a dose-dependent survival process,
   so the implemented phenotype is radiotropism and the model cannot test
   radiotrophy or radioresistance."
  "Because the program has neither growth nor death, the flag denotes a spatial
   tropism rather than radiotrophy or radioresistance."

Both are checkable exactly the PP-26-04 way -- no growth process reachable from
the trajectory -- and neither needs a grep over unbounded naming.

ONE SENTENCE USES THE UNBOUNDED FORM AND IT IS A METHODS DESCRIPTION: "The code
has a manually invoked division primitive ... but no growth law or automatic
division trigger; it likewise contains no radiation-dependent survival process."
A fourth statement, "still has no growth or survival process", is about the R
Shiny program and carries a different scope again.

So the repair is to narrow ONE methods sentence. Not to strengthen a grep, and
not to weaken the retraction. Same repair as abandoning the seven-organism
universal in PP-26-03: change the claim's type instead of hunting a better
boundary.

FOURTH TIME THIS PASS THAT READING THE ARTIFACT MOVED SOMETHING OUT OF THE
CATEGORY IT WAS OFFERED IN. The docstring notes were manuscript claims.
place_cell! was in the lattice-mover set and divide_cell! was not. The nutrient
field was live rather than dead. And the unbounded claim is one methods sentence
rather than the retraction's premise. THE CATEGORISATION KEEPS BEING WRONG MORE
OFTEN THAN THE FACTS DO, which is worth knowing before the next audit budgets
its time -- in every one of the four, the underlying fact was as reported and
the thing it was filed as was not.

Calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
Two things into PP-26-05, both about the shape of this branch's errors rather
than about growth rates.

THE LANDING IS CLOSE TO THE OPPOSITE OF THE USUAL FAILURE. The manuscript's
strongest negative claim -- the one the radiotrophy retraction rests on -- was
already stated at the scope its evidence reaches. What overreached was the
methods aside. Everywhere else on this branch the load-bearing sentence was the
one stated too widely, so an audit that went straight for the central claim
would have found it clean here and stopped.

WHY THE MIS-FILING KEEPS HAPPENING, AND IT IS A MECHANISM RATHER THAN A RATE. A
fact has a defined verification operation: open the source, run the grep, read
the method. A KIND is decided in passing, usually by whoever first mentions the
thing, and then inherited downstream without anyone performing an operation on
it. Filing is not harder than fact-checking. Nothing in the workflow is SHAPED
like a filing check, so there is no step at which a classification can fail.
That is why the facts held four times out of four this pass and the filing did
not: the docstring notes were manuscript claims, place_cell! was the mutator and
divide_cell! the caller, the nutrient field was live rather than dead, and the
unbounded claim was a methods aside rather than a premise. In all four the
underlying fact was exactly as reported.

THE PRACTICAL FORM IS A QUESTION, NOT A BUDGET: for each item, what would change
if this were the next tier up? The nutrient field -- if this is a manuscript
premise and not a docstring note, it needs a test. divide_cell! -- if this is a
caller and not a mutator, the pinned set is wrong. The question costs nothing,
is answerable without opening anything, and was asked zero times out of four.

Filed as evidence rather than as a rule. If it becomes AGENTS.md's seventh, the
four instances here are what would justify it, and that is a decision for
whoever owns that file rather than for this branch.

Calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
…s it fails

docs/research/spatial_dose_community_program.md. Written in the posture of
murr_facility_candidate.md -- forward-looking, no source_id claimed, nothing
quotable elsewhere, nothing promoted to a claims row until a source is pinned.

WHAT IT PROPOSES. Dose rate in a spent fuel pool is spatially structured, so if
community composition tracks it, where the aggressive chemistry sits is a
prediction about a field. Coupling a computed dose field to a spatial biological
model is what this repository has. Three phases: a pre-check, an empirical study
for whoever has pool access, and building the dose-dependent survival process
that would make it a prediction rather than a correlation.

WHAT THE DESIGN SURVIVED. Two earlier versions could not have tested what they
claimed. Sampling racks against a distant liner varies dose AND substrate,
distance, hydrodynamics, surface finish, cleaning history and installation date
together -- a correlational study reported as a falsification test. Dose now
varies WITHIN a substrate class, and the discriminating sample is where computed
dose and geometric distance disagree, which is where the coupling does work a
distance covariate could not.

AND THE TROPISM PREDICTION IS OUT OF THE SPINE. It predicts colonisation toward
the source, which is a claim about total biomass -- exactly the monotonic
variable the test declares uninformative. It cannot be both the headline
prediction and the excluded variable, so it is recorded as current behaviour.

PHASE 0 GATES THE HANDOVER, NOT THE SCIENCE, and says so. It pins a published
rack geometry; Phase 1 would sample a different pool, so it establishes
feasibility in a geometry nobody will sample. Worth running anyway for three
reasons needing no access: it debugs the transport pipeline, exercises its own
control, and produces the noise-floor curve. What it decides is whether this
document gets handed to anyone.

Its parts are specified before it runs, because Phase 0 does all the load-bearing
work: inputs pinned rather than chosen (a design failing at one setting can
otherwise be rescued by widening the burnup spread); the statistic is a derived
noise floor rather than a correlation; five covariates enumerated in advance
rather than chosen after seeing the residual; and a must-dissociate control with
ALL FOUR cells dispositioned -- including residual-large-with-control-flat, which
reports discriminating power while the field fails to respond to a configuration
that must change it, and reads as success.

THE NOISE FLOOR IS A CURVE, NOT A NUMBER. It moves with the assumed effect size,
and the effect size is what Phase 1 exists to estimate, so a single assumption
makes the design feasible by construction. Detectable residual is reported across
a range, and the conclusion is "feasible if the compositional shift is at least
X" -- evaluable by someone with pool access.

VERIFIED THIS PASS: SKB TR-18-08's Copper Sulfide Model couples microbial sulfate
reduction to interfacial electrochemistry, with SRB at the rock-clay interface
and general corrosion under 10 um at one million years (King et al., Materials
and Corrosion 2021). THE COUNTER IS RECORDED WITH IT: microbial activity entered
that safety case through a specific identified mechanism, not as a general term,
so the precedent DEMANDS an equivalent mechanism for FeCrAl -- a debt this
program takes on rather than a warrant it inherits.

SOURCE INACCESSIBLE, NOT CATEGORY ABSENT. The novelty claim rests on Marciales
et al. (Corros. Sci. 2018); its full text returns 403 from here, so what was read
is the abstract and reported findings and not the complete taxonomy. The claim is
scoped to that, and one review bounds one literature -- extending it across
radiation or repository microbiology would need a second review nobody has read.
Same state HOFFMAN-11 records for NACE-2019-12944.

Absence gate run over the draft: 22 flags, all judged. Two were real unscoped
universals and were fixed -- "the only place the coupling is load-bearing", and
"individual radiosensitivity needs no new model". The rest are procedural,
dispositions, or scoped inline. A clean run would not have been an all-clear
either; the gate prints five known gaps.

Julia suites green; calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
Four searches, aimed at killing the gate rather than supporting it. A candidate
answer for 0-pre, not a closed one.

THE EFFECT SIZE SPANS ABOUT ELEVEN ORDERS OF MAGNITUDE AND THE PROGRAM HAS TO
PICK AN END. Two anchors, both verified 2026-08-30:

  Chengdu Th-232 site, >10 years chronic: four groups from 193 +/- 5 to
  911 +/- 41 nGy/h, PCA clustering High and Medium fungi away from Low and
  Blank, threshold near 480 nGy/h, fungal diversity effect exceeding bacterial,
  and no significant difference in bacterial alpha diversity.

  Shuryak et al., PLoS One 2017: effects at 36-126 Gy/h, 78 of 145
  phylogenetically diverse FUNGAL strains growing at 36 Gy/h.

They measure different things -- decade-scale selection integrated over
generations against population density inside an experiment -- and a pool sits
on the first. So the program must DECLARE WHICH MECHANISM IT PREDICTS, because
the expected residual differs by orders depending on the answer and a prediction
accommodating both predicts nothing.

AND THE POOL SITS BETWEEN THE ANCHORS. A dose rate of 0.416 Gy/h is reported for
areas holding irradiated fuel: roughly 1e6 above the Chengdu threshold and about
1e2 below Shuryak's rates. Neither transfers directly.

ONE FINDING FAVOURS FEASIBILITY: Chengdu detected a composition threshold across
a total dose range of about 4.7x, not orders of magnitude.

ONE RUNS AGAINST IT AS A MECHANISM RATHER THAN A POWER PROBLEM, which is why it
belongs in 0-pre instead of being discovered at Phase 1. Shuryak reports
resistance is CELL-CONCENTRATION DEPENDENT -- D. radiodurans grew at 67 Gy/h at
high density and growth was extinguished at a tenfold lower concentration -- and
that resistant cells PROTECT NEIGHBOURS ACROSS PHYLA, wild-type D. radiodurans
enhancing E. coli survival >200-fold, apparently by catalase-mediated ROS
detoxification in the shared medium. If composition is buffered that way it is
not decomposable into per-species radiosensitivity, and "composition tracks
computed dose" has no monotonic form to test.

IT ALSO HAPPENS TO BE THE BEST ARGUMENT FOR A SPATIAL MODEL, and the document
says so rather than hiding it: a well-mixed or per-species model cannot represent
protection that depends on local cell density and a shared extracellular medium,
and a lattice model with local neighbourhoods can. The mechanism that breaks the
simple test is the one that would justify Phase 2. Stated plainly, not led with.

THE SHARPER RISK IS THAT DOSE IS NOT SEPARABLE FROM DEPTH. Water is the dominant
attenuator, of order half a metre per tenth-value thickness for Co-60, and
assemblies sit tens of feet down, so the field spans many orders vertically --
while depth also carries light, settling organic carbon, oxygen from the surface
and operator proximity. A regression of composition on computed dose over a
vertical sample is A REGRESSION ON DEPTH WEARING A DOSE LABEL. The design that
breaks it is horizontal: fixed depth, varying lateral distance from the racks.
If that is not physically available in the pool on offer, 0-pre closes on
COLLINEARITY even though it does not close on effect size, and that alone is
sufficient reason to send nothing. 0-pre's failure branch now carries all three
conditions.

THE ABSENCE IS VERIFIED AND IT IS STRONGER THAN ASSERTED. Nothing found computes
a dose field in a spent fuel pool and relates it to community composition. The
closest study is the evidence: an amplicon survey of a spent nuclear fuel storage
basin sampled 13 DIFFERENT LOCATIONS AND DEPTHS, measured no dose rates at those
sites, tested no association with radiation, computed no field, and states that
radiochemical analyses were not performed. It had the spatial sampling and lacked
only the dose axis. Its depth sampling is also why the confound above is not
hypothetical.

Gate run over the new material; two claims were leaning on a search scope stated
elsewhere in the document and now point at it.

Calibration 371; Julia green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
… other two would have copied

Reading two PDFs found four defects. A sweep found a fifth, and checking the
arithmetic before propagating found a sixth in the sentence the others were
going to copy from.

THE RATIO WAS WRONG AT SOURCE. Section 6.2 said the corrected reach is "only
fifteen times smaller than the melanin term rather than four orders." Against
Table 3's own values -- melanin dH = -0.720 at M=1.44, bias 1.1548; radiation at
the shipped reach 7.505e-2, bias 1.0151 -- the ratios are 9.6 in dH and 10.2 in
bias. NEITHER IS FIFTEEN. 15.41 is the bias ratio computed against a reach of
5e-2, a number appearing nowhere in this paper. So that sentence was written
mid-correction: the "four orders" half was already fixed while the ratio still
used the pre-pairwise magnitude. A CORRECTION THAT REACHED ONE SENTENCE TWICE
AND HALF-LANDED, which is a different mechanism from the propagation failures
beside it -- and section 7.1 and the Conclusion were about to copy it, spreading
one error to three places.

THE FIVE PROPAGATION DEFECTS. Section 7.1 said "several orders of magnitude"
while 6.2 said otherwise, in one document. The Conclusion said the direct term
is "negligible for the two negatively signed classes." The Abstract said "one
part in 1e5 for the two negatively signed species." The title page said Version
1.1. Software and Data Availability named f72aabb, v1.1's own commit.

THE ABSTRACT'S FIX IS NOT AN APPENDED CAVEAT. Its sentence was literally true --
the scope clause survives -- so no phrase guard or bound check could fire, and an
abstract-only reader got v1.1's picture. Appending the corrected reach would make
one sentence carry two numbers, a scope clause and a withdrawal. Corrections
belong in the correction section. And the replacement must not trade
scoped-but-misleading for unscoped-but-accurate: "sixteen of 206042" is measured
at N=40, seeds 42/43/44, 400 MCS, and 26.7 percent over the same three runs, so
neither is scope-free. The Abstract already works at the COEFFICIENT level --
15.5 percent is a property of the melanin coefficient and of no run -- so the
parallel is the reach itself, 7.505e-2, with no run, seed or configuration.

THE REGISTER IS NAMED EVERYWHERE THE RATIO APPEARS, because 9.6 and 10.2 are the
same claim 6 percent apart. Sections 6.2, 7.1 and the Conclusion compare
Hamiltonian terms and use the dH ratio; the Abstract quotes a bias and uses the
bias ratio.

HOW THE LIST WAS PRODUCED, AND WHAT BOUNDS IT. Four from reading. A fifth from a
grep over {negligible, orders of magnitude, 10^{-5}, 5x10^{-5}, four orders,
dominat, dwarf, vanishing, one part in}. Then an INVERTED pass reading every hit
for {radiation term, Delta H_rad} -- a population small enough to enumerate --
which added none and is what converts a keyword search into a complete pass over
a bounded set. Deliberately left: the dose-regime "ten orders of magnitude,"
Table 2's literature ranges, 5e-5 inside the correction section where it is named
as withdrawn, and "coupling dominates the direct radiation term," which carries
no magnitude. The ledger rows say the list is bounded by those terms and assert
no completeness beyond them.

The version string is dated by preparation, not by sending, and August is
consistent with the correction section's own internal dates. The commit hash is
the parent of the commit that lands this revision, read on a clean index before
staging -- not "HEAD at the moment of the edit," which would have included some
of this revision's own changes.

WAN NOTE names two asks explicitly so she does not have to work out which is
which: a measurement in her own lab, and forwarding the memo. Independent, and
neither a precondition for the other.

MEMO PAGE 4 says what Reference D is and what it is not: the campaign the
preprint needs, formally unevaluated, a project rather than a capability. And
that it is NOT the same project as the FeCrAl questions -- Reference D would
calibrate a biofilm model, those questions are about a material, and the overlap
is the facility rather than the measurement. Four pages, zero overfull; the block
went to the right column, overflowed, and came back with the hedge above it
compressed to make room.

0-PRE IS PRE-REGISTERED. A power analysis whose inputs the analyst chooses
produces whatever answer is wanted, and 0-pre's job is to refuse, so the range
and the refusing end are written before computing. The horizontal dose contrast
is a look-up that runs first because it can close the gate alone -- under about
2x and 0-pre closes without any power calculation. Bray-Curtis with PERMANOVA
and dose continuous, dispersion taken from the verified 13-location survey, which
is the one number in it that transfers. And that survey is recorded as a
near-miss rather than an absence: someone has already sampled a basin, they did
not compute a field, and the missing half is the half this repository holds.

Manuscript claims 36, prose bounds 23, calibration 371.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
@aurascoper

Copy link
Copy Markdown
Owner Author

All four fixed in 1f34f08, each verified in-repo before acting. Three of them are the same defect — checked a subset, asserted the set — in artifacts written while I was explicitly discussing that defect.

The sweep (P1). You named four live dependents; running the same term set repo-wide instead of over the .tex found five. README.md:403 — "melanin dominates β_ion by four orders of magnitude" — was not on your list and would have been missed again by hand-editing the four. Corrected in README (three places, including a table row), integration_contract.md, parameter_provenance.csv and feedback_parameter_distributions.csv.

The upper bound (P1). Correct, and it is a second defect in a sentence I had already corrected once — the earlier withdrawal weakened the claim along the dose-scale-transfer axis and left the arithmetic untouched. The threshold is now unset rather than re-justified. No fraction of an upper bound is defensible as a floor, and "we cannot set this yet" blocks D2b more firmly than a number that admits 2.5×. Four stale 2.35× references elsewhere in the same document were cleared.

match (P1). Confirmed by mutation before fixing: changing §7.1's 9.6 to 15.0 produced zero failures. Now eachmatch with a per-location pattern, the hit count pinned so a location losing its phrasing fails rather than dropping silently out of the sweep, and the two word-stated locations pinned by phrase with that difference named. All three mutate independently; 16 assertions, up from 12.

The intersection (P2). Correct — two of three entries were dead. Now a complete-set comparison, which failed on the shipped index as it should have. Fixed by adding the rows those predicates are triggers for (SOP-T5, a two-way coupling check blocked on ∂μ/∂c ≈ 0; SOP-T6, the Phase 2 radiosensitivity comparison blocked on there being no survival process) rather than by deleting entries that block real work.

One thing your four findings have in common that I would not have seen from any one of them: in each case the guard bounded its own claim correctly while the contract above it overstated the reach. The absence gate scans prose for absence claims; a docstring saying "every location" over an assertion covering one is the same defect living in a comment, and nothing scans those. Proposed rather than built.

Julia green; calibration 389.

THE THRESHOLD IS A FINDING, NOT A TO-DO. "Unset" reads as pending; the actual
position is that NO AVAILABLE ANCHOR BOUNDS THIS IN THE REQUIRED DIRECTION. A
conservative gate needs a LOWER bound on necessary contrast and Chengdu supplies
only an upper one -- a threshold at 480 nGy/h inside a 4.7x span with Low and
Medium given as sampling distances, so the smallest contrast that resolved the
effect is unpublished. The arithmetic is not repairable in the direction the
original derivation assumed. Three options are written into the row, one of
which must be chosen before D2b runs: obtain the bracket, derive the threshold
from dispersion instead of from an anchor's span, or declare that 0-pre cannot
set it and close on that. And why unset beats wrong: an unset gate refuses
everything, a wrong gate admits exactly the cases it should not.

THE DOCSTRING GATE IS BUILT, as a second mode of the tool that already exists.
Three defects in one review round shared a signature -- guard bounded correctly,
contract above it overstated -- which is one seam rather than three bugs, and it
is the seam this project's discipline is organised around, unguarded in the one
place that describes the guards.

Two design decisions, both of which the first version got wrong:

  IT DOES NOT INFER SCOPE, it requires one. Universals fire legitimately here:
  "asserts every registered absence still holds" is accurate. So the check is not
  "no universal in a docstring" but "a universal whose scope is DECLARED", in a
  form a reader can compare against the assertion below. Inferring the scope is
  the same guess that produced the defect.

  AND IT DECLARES ITS OWN, printing the file count it was given and stating that
  a clean run reads only those files -- or it reproduces the defect one level up.

THE FIRST VERSION FIRED 19 TIMES ON THREE FILES, mostly on ordinary commentary,
which is the noise failure that gets a gate muted inside a week. Narrowed: a
block counts as a contract only when a test definition sits within four lines of
it, so what is judged is the boundary between a check and its description. Six
flags, of which three were real coverage claims and three were legitimate.

The three real ones now carry SCOPE lines, and the most useful is on the
predicate-shape control: it says in the file that it establishes the SHAPES can
return False and does NOT establish that any predicate is bound to the term its
row names -- which only firing the per-term mutations can. That is the exact
distinction Codex's P2 turned on, written where the next reader meets it.

THE THREE FALSE POSITIVES ARE RECORDED AS JUDGED rather than left to be
re-litigated: "no row claims an SOP" and "a row with no blocked_by" describe a
control's own input, and "printed every run" is temporal. A gate whose false
positives are never written down is judged again every pass until someone mutes
it.

Controls in both directions, and the first pair exercised nothing -- a bare
sentence has no comment run and no triple-quoted block, so the scanner saw no
input and reported a miss. Rewritten in docstring form beside a test definition.

Julia green; calibration 389.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
@aurascoper

Copy link
Copy Markdown
Owner Author

@codex review

Head is 0fec3b9. I held this ping until the work was pushed rather than sending one at 1f34f08 and another an hour later — one review of one head rather than two of stale ones.

All four of your findings on f666649 are fixed and resolved in 1f34f08, and one of them generalised into a tool in 0fec3b9. What to doubt:

1f34f08 — three of your four were one defect, and the fix for the first found more than you named. Running the same term set repo-wide instead of over the .tex returned five live dependents, not four: README.md:403 ("melanin dominates β_ion by four orders of magnitude") was not on your list. Widening the sweep beat editing the named files. Worth checking whether five is now the whole set or whether the term set itself is the remaining bound.

1f34f08 — the match fix, and please try to defeat it. eachmatch with a per-location pattern, the hit count pinned so a location losing its phrasing fails rather than dropping silently out of the sweep, and the two word-stated locations pinned by phrase. All three mutate independently now; §7.1 previously mutated to 15.0 with zero failures. The question is whether a fourth location could state the ratio in a form none of the five patterns match.

1f34f08 — the threshold is now UNSET, and the reasoning is the load-bearing part. You were right that halving an upper bound loses the guarantee. The stronger statement I'd like checked: a conservative gate needs a lower bound on necessary contrast, Chengdu supplies only an upper one, so no fraction of any Chengdu-derived number is defensible — not just 2.35×. If that generalisation is wrong, option 2 in the row (derive from dispersion instead of from an anchor's span) may be unnecessary.

0fec3b9 — a new gate, and its noise profile is the thing most likely to be wrong. It flags a universal in a docstring that declares no SCOPE:. First version fired 19 times on three files, mostly on commentary; narrowed to blocks within four lines of a test definition, giving six, of which three were real. That ratio is the whole question — a gate at 50% precision is usable and one at 15% gets muted. The DEF_WINDOW = 4 heuristic is the weakest part and I'd rather hear it's too crude now than discover it when someone silences the tool.

Suites: Julia green, calibration 389.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0fec3b9582

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread README.md Outdated
Comment thread data/calibration/sop_index.csv Outdated
Comment thread tools/absence_gate.py Outdated
Comment thread calibration/tests/test_sop_index.py Outdated
A proximity heuristic has two error modes and they are not symmetric. A false
positive costs a judgement call, which is cheap and now gets recorded. A false
NEGATIVE is a contract the gate silently does not cover -- which is the exact
class this gate exists to catch, so a non-empty false-negative set would put the
subset defect inside the tool built to detect it.

ENUMERATED BEFORE ANYONE ELSE HAD TO. The DEF_WINDOW = 4 heuristic drops 35
blocks across six test files, and the dropped set contains real contracts:
delta_h_decomposition.jl states "ALL FOUR BRANCHES MUST BE EXERCISED" at module
level, and test_sop_index.py states which run_by values are exempt from check 7.
Both sit further than four lines from any test definition. The gate's claim to
guard the check-versus-description boundary was therefore overstated, in the same
direction as the three defects that motivated it.

TWO TIERS, REPORTED SEPARATELY. Attached contracts sit near a test definition.
Module-level contracts sit outside the window and are admitted only when worded
NORMATIVELY -- must, never, cannot, only, exempt, required -- because that is
what separates a rule about a check from commentary about the physics. Widening
the window instead would have restored the 19-flag noise the narrowing existed
to remove; the tiers keep both sets visible without merging their precision.

Yield: 6 attached, 17 module-level across six files. Both blocks named above are
caught, and both now carry SCOPE lines -- the delta_h one saying it covers four
branches of compute_delta_H_terms on its fixture and not every path through the
stepper, which is a narrower claim than the shout suggested.

AND THE MODULE-LEVEL TIER IS DECLARED A BACKLOG, NOT A JUDGED SET. Seventeen
blocks remain unscoped and unjudged, and the gate's own header says so. A tuned
gate whose exceptions are undocumented gets re-tuned by whoever hits them next,
usually by widening until the noise stops -- which is how a check dies without
anyone deciding to kill it. Calling a backlog a backlog is what stops "clean run"
and "not looked at yet" reading alike.

Julia green; calibration 389.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
@aurascoper

Copy link
Copy Markdown
Owner Author

Addendum before you start — head moved to a1a3110, and it answers the DEF_WINDOW question I asked you rather than leaving it for you.

The sharper form of that question is not "is 4 too crude" but "what does a 4-line window fail to see, and is that set empty?" — because the two error modes are asymmetric. A false positive costs a judgement call. A false negative is a contract the gate silently does not cover, which is the exact class the gate exists to catch.

The set is not empty. Enumerating the drops: 35 blocks across six test files, and the dropped set contains real contracts — tests/delta_h_decomposition.jl states "ALL FOUR BRANCHES MUST BE EXERCISED" at module level, and test_sop_index.py states which run_by values are exempt from check 7. Both sit further than four lines from any test definition, so the gate's claim to guard the check-versus-description boundary was overstated in the same direction as the three defects that motivated it.

Fixed with two tiers rather than a wider window, which would have restored the 19-flag noise: attached contracts near a test definition, and module-level contracts admitted only when worded normatively. Yield is 6 attached and 17 module-level.

Two things still worth your attention there. The normative word list (must|never|cannot|only|exempt|required) is the new weakest part and has the same false-negative exposure the window had — a module-level contract phrased without any of those words is invisible again. And the module-level tier is declared a backlog, not a judged set, in the gate's own header; if that declaration is itself too generous I would rather know now.

aurascoper and others added 3 commits August 30, 2026 23:55
…ons stopped one step short

Four more from Codex on 0fec3b9, all verified before acting. Two are mine from
the same pass that introduced them.

P2 -- THE GATE COULD NOT DETECT ITS OWN MOTIVATING DEFECT. `scan_contracts` on
tests/prose_bounds.jl returned NOTHING in attached mode, because the "EVERY
LOCATION" contract sits inside a @testset whose header is 25 lines above it and
the four-line proximity window saw no definition. The guard built to catch
contracts that overstate their reach could not see the contract that prompted it.

Blocks now associate with the nearest test definition that PRECEDES them, so a
contract documented beside assertions deep in a body belongs to that test.
prose_bounds.jl:214 moves from module-level-by-luck to attached-by-rule; totals
go from 6 attached / 17 module to 21 / 5. An in-body negative control ships with
it -- a contract five lines into a testset body, which the window missed and the
association catches.

AND THE FIRST TWO CONTROLS THEN BROKE, correctly: they placed their comment block
ABOVE the definition, which under the new rule is module-level rather than
attached, so a control written for the old rule reported a miss. Rewritten with
the definition first. A control that survives a rule change unchanged was not
testing the rule.

P2 -- THE ORPHAN CONTROL REIMPLEMENTED WHAT IT PROTECTED. It computed a
hard-coded set difference beside the guard rather than through it, so it would
have kept passing if test_no_registry_term_is_dead regressed to the old
intersection. The dead-term calculation is now dead_terms(), called by both, and
the control plants an orphan and requires that shared function to report it.

P1 -- README's Key Results row changed its adjective and not its arithmetic. It
compared 15.5% against the 1.000010 bias, which is one ROLE of one pair of
species occupying a site and gives 15,500x, then described the result as "about
an order of magnitude". The comparison now names the bias AT THE TERM'S REACH,
1.0151, for a 10.2x ratio of excesses, and says explicitly that comparing against
1.000010 is the withdrawn comparison.

P1 -- THE SOP INDEX STILL CARRIED THE COMPARATOR THE PROGRAM DOCUMENT HAD
WITHDRAWN. SOP-T4 said 4.7x makes 2.35x conservative by construction while
spatial_dose_community_program.md said no fraction of an upper bound supplies the
required floor. A promotion reference contradicting the document it references is
how a withdrawn gate gets reinstated. Withdrawn there too, with the reason.

Both P1s are the same shape as the defect this branch has been correcting all
along: a correction that reaches one location and stops.

Julia green; calibration 389; gate self-checks pass with three contract controls.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
… generator can satisfy

THE FIGURE CARRIED A NUMBER THE PROSE HAD ALREADY CORRECTED. Its caption read
"1e-3 cm2/s, about 100x water's self-diffusivity". The verified value against
water self-diffusivity 2.3e-5 is 43.5x. Converting first and "preserving the
figure exactly" would have shipped the superseded number and landed the
correction in the prose and not in the image -- FIG-01 through FIG-07 in its
purest form. Fixed in the SVG BEFORE any conversion ran.

AND THE CAPTION NAMED ONLY WHAT THE BENCH NUMBER IS NOT. It said the film-scale
diffusivity is not the reactor-scale coefficient carrying the same name, and
never said which quantity it IS. It now names radial_parms as the colliding
symbol and slab_parms' D_eff (1e-5 cm2/s) as the one default in that file below
water self-diffusivity -- the half that makes the warning actionable. The other
two disambiguation points were checked before adding and were already in the
figure: the "What one run yields" panel already carries the lumped-parameter
argument, the sorption retardation, the partition coefficient from Phase 1, and
why Phase 1 comes first.

A SIDECAR HASH DOES NOT BIND A PDF TO ITS SVG. Compared against the SVG it binds
the SVG to a record of itself: edit the SVG, the hash mismatches, and the obvious
repair is to regenerate the record -- green again with the PDF still stale, the
check satisfied without catching what it exists for. So the hash is written ONLY
by tools/render_figure_svg.py, as a side effect of actually converting. The only
way to make the guard green is to run the renderer, which emits the PDF.
contract_csv.jl's reasoning inverted: there the fixture must never be rewritten
by what it checks; here the record must be writable only by the generator.

Mutation-checked: one character appended to the SVG without reconverting fails
the guard, and the SVG restored byte-identical.

AND THE CORRECTION IS ASSERTED ON THE ARTIFACT, NOT THE SOURCE, because prose
right and image wrong is the whole defect: pdftotext on the rendered PDF must
find 43 and must not find the withdrawn 100x comparison.

THE FIGURE JOINED AN EXISTING CONTRACT AND I HAD TO BE TOLD BY IT. Two suite
failures showed preprint/figures/ requires .sha256, .txt and .png of every
committed PDF -- and the .txt is what RETRACTED_IN_FIGURES scans, so producing it
puts this figure's text under the phrase guard that exists because a withdrawn
claim once survived inside an image. The renderer now emits the full set. A third
failure came from writing "<hash>  <name>" where the convention is a bare hash:
a format invented beside the schema instead of mirroring it, which is the same
lesson as naming a model-side field after the bench-side one.

THE REGISTRY'S ATTENUATION ENTRY WAS PHOTON-SPECIFIC AND ASSERTED
UNCONDITIONALLY. d(mu)/dc ~ 0 holds because at ~1 MeV over C/H/N/O biomass and
void are near-indistinguishable. For NEUTRONS it is false and interestingly so:
biofilm is largely water, so a biofilm region moderates against a void and
hydrogen's scattering cross-section is large. Neither existing predicate fails if
the particle changes while the manuscript sentence stays, and a source change is
exactly the update path nobody re-reads that entry for. Third predicate added,
matched on the CONTENT particle="photon" rather than a line number, since a line
anchor silently starts reading a different statement. Mutation-checked by
flipping the literal.

Julia green; calibration 392, up from 389.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJu1d4TFhowR3GUtH2MdoG
…test that was right to fail

THE RBE GATE WAS ABOUT TO BE TYPED PERMANENT ON SEARCHES THAT COULD NOT HAVE
FOUND THE THING. It is written now as D3, typed OPERATIONAL-pending. STRUCTURAL
is the one label in this register that says stop looking, and applying it here
would have put the most consequential verdict on the least scoped claim, in a
document going to a facility whose own people are the ones who would know.
Operationally nothing changes -- the gate refuses either way, and no
neutron-bearing exposure is designed or costed until a weight is sourced. What
changes is that the refusal stays reopenable.

AND THE EVIDENCE CAME OUT STRONGER THAN THE ARGUMENT IT WAS MEANT TO SUPPORT.
The premise was that the searches asked for the conjunction microbial + neutron
+ RBE, which is not how the field names it. Four searches ran, term sets
recorded verbatim. Searches 1 and 3 asked for that conjunction and returned
nothing usable. Searches 2 and 4 named the literatures instead -- sterilisation
dosimetry, spacecraft shielding -- and returned three on-point documents. So the
term-set diagnosis is confirmed by evidence rather than argued from, AND THE
REMAINING BARRIER IS RETRIEVAL AND NOT NAMING: DTIC ADA307995 (Neutron and
gamma-Ray Radiation Killing of Bacillus Species Spores) is HTTP 403 from here,
NASBEE is paywalled, arXiv 2408.10929 returned image streams and no extractable
text. Source inaccessible, not value absent -- HOFFMAN-11's exact shape, three
times over. An empty result under the first term set discriminates nothing.

ONE LEAD IS RECORDED AS A LEAD AND MUST NOT BE READ AS A NUMBER. NASBEE's
abstract-level material gives an RBE of 3.54 at D10 for 2 MeV average neutrons.
The target is not established as microbial from what was read and NIRS builds
for mammalian radiobiology. It is in the gate because it fixes the FORM the
answer takes -- an RBE quoted at D10 against a stated neutron energy -- which is
what a search should look for. It is not a weight this program may use.

The paragraph carries a SCOPE line naming what was not searched: no
controlled-vocabulary database, no sterilisation-dosimetry handbook, no BNCT
review, no author contacted. It converts into a question in hoffman_memo.tex
rather than sitting here, under the rule that already governs that memo. And one
universal in my own prose -- "every dose number this program carries" -- was
flagged by tools/absence_gate.py and is now a closed four-item enumeration of
what this document actually cites.

THE FIGURE'S R = 1 CLAUSE WAS IN THE SVG AND THE RENDERER HAD NOT RUN SINCE.
The sha sidecar was stale, which is the exact state the generator-only guard
exists to make impossible to miss, and it caught it. Re-rendered; pdftotext on
the artifact finds "Bromide is the tracer because R = 1" and finds neither
withdrawn anion-exclusion phrase.

A TEST FAILED AND IT WAS RIGHT TO. test_the_figure_rows_are_honest_about_being_
vocabulary_only asserts that every figure delete row is covered by the word list
and not by the phrase guard, because a claim with no searchable run looks
identical to a claim that was checked and found absent. FIG-09 and FIG-10 now
yield searchable phrases, so they moved to the STRONGER guard and the assertion
fired exactly as its own docstring predicted it would. Split into two tiers with
BOTH directions asserted, since the drifts lose different things: a row gaining
a phrase understates coverage, a row losing one drops silently to the weaker
floor. Negative control run -- injecting both withdrawn sentences into the .txt
sidecar fails test_no_deleted_claim_survives_in_the_document_it_names, and the
sidecar restored.

ANION EXCLUSION STAYS OUT, AND THE CHECK THAT REMOVED IT POINTED THE WRONG WAY.
Hay, Stoliker, Davis & Zachara 2011 report that bromide showed very little
penetration of Hanford 300A intragranular porosity. They report the observation;
"anion exclusion" as the named mechanism appears on no page fetched -- it came
from a search-engine summary, as did the volume and article number. Stewart 2003
puts small inorganic ions at De/Daq = 0.56-0.70 in biofilms, hindered and not
excluded, and the documented charge effect in anionic EPS retards CATIONS. The
plan's logical point survives the mechanism's withdrawal and is what the figure
now carries instead: bromide is the tracer because R = 1, which is a conditional
and a question for Deng, not an assertion about EPS.

PP-DEFF-01: one symbol, three quantities, and the shipped value is none of them.
Table 2 IS tab:params, it exists, and it is refuted on CONTENT rather than
absence -- its only diffusion rows are per-species cell motility D_s in um2/s.
Refuted and unresolvable take different repairs, so the row says which half is
which: the 1e-4..1e-2 span is refuted outright since it lies entirely above
water self-diffusivity, while the "Deep Research synthesis (April 2026)" is not
in this repository and the row asserts nothing about what it contains.
PP-DAM-01 amended to say its two numbers are a shipped default clashing with a
manuscript-only law, not two shipped things; D_w has no shipped value anywhere
and test_sop_index.py:219 asserts its absence from sources.

THE DENG NOTE LEADS WITH HIS SUBJECT AND STAYS HELD. Opens on Hua & Deng 2008
(verified: Bin Hua and Baolin Deng, Missouri) uncharacterised, asks whether the
transport reasoning carries into an open anionic hydrogel, and treats biofilm
and radiation as context for why the medium is anionic rather than as the
subject. Three questions to two -- the approval-path question is premature while
the suspended isotherm has not run. The claim that the bench number "maps onto
the effective diffusivity in the transport model" is gone: it implied the
model's D_eff is being calibrated, and the honest form is that the measurement
would define WHICH quantity the model should carry. An unscoped "I have not
found the EPS-specific measurement" now names what was searched.

The published Comment (es803175g) and Response (es803528t) are the reason the
paper is cited plainly rather than praised. They are NOT in the outgoing text --
raising a critical exchange in a cold introduction costs more than it buys -- so
the reason lives in the hold banner where a later editor will meet it before
deciding to add an appraisal.

Verified: 392 passed in calibration, julia --project=. tests/runtests.jl green,
absence_gate self-checks GATE USABLE with its 5 known gaps still open, ledger
parses at 347 records with no duplicate ids. preflight_merge.sh 23 REFUSES:
the newest Codex review covers 0fec3b9 and head has moved past it. No
unresolved threads. Not merged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
@aurascoper

Copy link
Copy Markdown
Owner Author

@codex review

Head is 29bca4b. Your four findings on 0fec3b9 are all fixed and resolved — though two of them (absence_gate.py's window, the shared-logic control) were already fixed in a1a3110 before that review ran, so it reviewed a superseded head. 75bdb08, c2219a2 and 29bca4b have never been reviewed by anything.

Three things are new, and here is what to doubt in each.

29bca4b — a test failed because coverage got STRONGER, and my fix is the shape you have now caught three times. test_the_figure_rows_are_honest_about_being_vocabulary_only asserted that every figure delete row is covered by the word list and not by the phrase guard, because a claim with no searchable run looks identical to a claim that was checked and found absent. FIG-09 and FIG-10 now yield searchable phrases, so the assertion fired exactly as its own docstring predicted. I split it into two tiers and asserted both set memberships. That is subset-checked/set-asserted wearing a fix's clothing, and I would rather you tried to defeat it than take my word. The specific thing I cannot rule out: tier membership is derived from distinguishing_phrase at MIN_WORDS = 5, so an editorial trim to a claim_text moves a row between tiers and fails both assertions at once — and the cheap repair is to move the id in the expected list, which restores green while the row has silently dropped to the weaker floor. The failure is loud; I am not sure it is legible.

29bca4b — and the control for that guard is not committed. I verified the phrase guard bites by injecting both withdrawn sentences into phase2_diffusion_cell.txt by hand and watching test_no_deleted_claim_survives_in_the_document_it_names fail, then restored the sidecar. The manuscript path has a committed known-bad (ORIGINAL_FIXTURE, ≥18 detections); the figure-sidecar path has none. So the guard that covers FIG-09/10 is asserted by a check I ran once and did not leave behind, which is AGENTS.md rule 1 with the serial numbers filed off. Whether that needs a committed known-bad sidecar fixture is a fair question and I think the answer is yes.

29bca4b — an absence gate typed OPERATIONAL-pending, and the typing is the claim. D3 in spatial_dose_community_program.md refuses neutron-bearing exposure because no microbial neutron RBE has been verified. The premise I started from was that the searches asked for the conjunction microbial + neutron + RBE, which is not how the field names it. Four searches ran, term sets recorded verbatim: 1 and 3 asked for that conjunction and returned nothing; 2 and 4 named the literatures instead (sterilisation dosimetry, spacecraft shielding) and returned three on-point documents, all of which then failed to open — DTIC ADA307995 is 403, NASBEE is paywalled, arXiv 2408.10929 returns image streams. So the barrier is retrieval, not naming, and STRUCTURAL would have foreclosed a clearing path with a named cost. What to attack: whether the SCOPE line actually bounds the claim, and whether four web searches can support even the weak typing.

And one lead in D3 is the thing most likely to do damage later. NASBEE's abstract-level material gives an RBE of 3.54 at D10 for 2 MeV average neutrons. The target is not established as microbial and NIRS builds for mammalian radiobiology, so it is recorded as a lead and explicitly not as a usable weight. A plausible figure in the right units, sitting in a gate paragraph, is exactly what gets quietly promoted to a number six months from now. If you think recording it at all is the wrong call, say so — the alternative is dropping it and losing the one thing that fixes what form the answer takes.

c2219a2/29bca4b — anion exclusion came out of the figure, and the check that removed it pointed the wrong way. Hay et al. 2011 report that bromide showed very little penetration of Hanford 300A intragranular porosity. They report the observation; "anion exclusion" as the named mechanism appears on no page I fetched — it came from a search-engine summary, as did the volume and article number. Stewart 2003 puts small inorganic ions at De/Daq 0.56–0.70 in biofilms, hindered rather than excluded, and the documented charge effect in anionic EPS retards cations. The figure now carries the conditional instead: bromide is the tracer because R ≈ 1, which is a question for Deng and not an assertion about EPS. FIG-09 and FIG-10 are the delete rows for the two withdrawn sentences.

PP-DEFF-01 — one symbol, three quantities, and the shipped value is none of them. Worth checking the half I called refuted against the half I called unresolvable: Table 2 is tab:params, it exists, and its only diffusion rows are per-species cell motility D_s in µm²/s — refuted on content, which is checkable. The "Deep Research synthesis (April 2026)" is not in this repository at all, so the row asserts nothing about what it contains. If that split is wrong in either direction the repair differs.

Suites: calibration 392 passed, Julia green, absence_gate.py self-checks GATE USABLE with its 5 known gaps still open. preflight_merge.sh 23 refuses on staleness, which this ping is meant to clear.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 29bca4be02

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +36 to +38
recorded = record.read_text(encoding="utf-8").split()[0]
actual = hashlib.sha256(svg.read_bytes()).hexdigest()
assert recorded == actual, (

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Bind the render check to the generated PDF

When an SVG changes and its writable .svg.sha256 is updated without rerendering, this equality succeeds while the committed PDF remains stale; no repository mechanism enforces the comment's claim that only the renderer writes that file. I reproduced that state with a changed SVG, matching source hash, and unchanged PDF, and the control at lines 44–50 would still pass because it only compares two independent digest expressions. Compare against a fresh rendered artifact or otherwise bind the record to both source and output through the production check.

AGENTS.md reference: AGENTS.md:L17-L21

Useful? React with 👍 / 👎.

Comment thread calibration/tests/test_claims_ledger.py Outdated
Comment on lines +1000 to +1003
assert sorted(phrase_covered) == ["FIG-09", "FIG-10"], (
f"phrase-covered tier moved: {sorted(phrase_covered)}. If a row left "
"this list its claim text no longer yields a searchable run, so the "
"sidecar guard has stopped covering it.")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Commit a known-bad sidecar for the phrase tier

These memberships exercise only distinguishing_phrase against ledger text; they remain green if the production guard stops reading figure sidecars, normalizes them incorrectly, or otherwise cannot detect FIG-09/10. The manual injection described in the commit is therefore the only evidence that this newly covered path bites. Pass a committed known-bad sidecar containing both withdrawn sentences through the same detection logic used by test_no_deleted_claim_survives_in_the_document_it_names.

AGENTS.md reference: AGENTS.md:L19-L26

Useful? React with 👍 / 👎.

Comment thread data/claims_ledger.csv Outdated
FIG-08,preprint/figures/phase2_diffusion_cell.pdf,panel: Open choice: the tracer run,In Hanford sediments tritiated water resolved intragranular pore volume where bromide did not.,empirical_literature_claim,,none,not_applicable,false,,literature,keep,"Name the source in the figure before the panel reaches Dr. Deng. Done 2026-08-31: Hay, Stoliker, Davis & Zachara 2011, Water Resources Research, doi:10.1029/2010WR010303.","SHIPPED UNCITED IN A FIGURE ADDRESSED TO A NAMED EXTERNAL READER WHOSE FIELD THIS IS - the configuration that produced HOFFMAN-01 and HOFFMAN-02. Verified 2026-08-31 against the USGS publication page for the paper, which carries the authors' own summary: sediments are the uranium-contaminated vadose and capillary-fringe materials beneath the former 300A process ponds at Hanford, and ""experiments using bromide ion as a tracer yielded very different results, suggesting very little penetration of bromide into the intragranular porosity"". The Wiley abstract page returns 403 and was not read. WHAT WAS DELIBERATELY NOT CARRIED INTO THE FIGURE: volume 47 and article number W10531 appeared only in search-engine summaries, in no page actually fetched, and a specific supplied by a secondhand summary and not seen in the source is precisely HOFFMAN-01. The panel cites year, journal and DOI, all three read off the USGS page."
FIG-09,preprint/figures/phase2_diffusion_cell.pdf,panel: Open choice: the tracer,Bromide is size- and charge-excluded from fine porosity.,mechanism_claim,,none,not_applicable,false,,literature,delete,"Remove the mechanism attribution. Done 2026-08-31: the panel now states the observed result and cites it, and asserts no mechanism.","THE SOURCE REPORTS THE OBSERVATION, NOT THE MECHANISM. Hay et al. 2011 report very little bromide penetration and infer restricted access; ""anion exclusion"" as the named cause came from a search-engine summary of that paper, not from any page read. Size exclusion and charge exclusion are two mechanisms and the figure asserted both, unsourced, in a sentence that read as established. The deleted mechanism does not change what the panel is FOR: the tracer choice question stands on the observation alone. WHAT THE GUARD DOES NOT COVER, STATED HERE BECAUSE THE ROW WOULD OTHERWISE IMPLY IT DOES. test_claims_ledger.py matches this row's distinguishing_phrase as a SUBSTRING of the figure's .txt sidecar, so it catches the exact withdrawn wording and nothing else; a reworded assertion of the same claim passes it silently. THE WITHDRAWAL IS SEMANTIC, NOT PHRASAL: what is deleted is the claim that an EPS matrix excludes anions the way Hanford intragranular porosity did, in ANY wording. The mechanical check is a floor under a human one, not a substitute for it."
FIG-10,preprint/figures/phase2_diffusion_cell.pdf,panel: Open choice: the tracer,an anionic EPS matrix is the same trap,transfer_claim,,none,not_applicable,false,,literature,delete,"Withdraw the transfer or source it from the EPS literature. Done 2026-08-31: withdrawn, and the panel now says the same holds in an EPS gel is untested.","A SEDIMENT RESULT TRANSFERRED TO A HYDROGEL WITHOUT A SOURCE IN THE HYDROGEL LITERATURE. What the EPS literature actually reports is the opposite sign of concern: Stewart 2003, J. Bacteriol. 185:1485-1491, compiles relative effective diffusivities in biofilms and puts small inorganic ions at De/Daq of roughly 0.56 to 0.70 - HINDERED, NOT EXCLUDED - and the well-documented charge effect in anionic EPS is retardation of CATIONS, e.g. tobramycin sequestered by alginate, not exclusion of small anions. Hanford intragranular porosity is nm-scale voids in a rigid mineral aggregate; EPS is an open gel above 90 percent water where the Debye length at physiological ionic strength is of order 1 nm. The two systems share a word, not a measured behaviour. THIS WOULD HAVE BEEN THE FIRST THING DR. DENG SAW, and the panel asking him a question is the wrong place to assert something in his field that the field does not say. WHAT THE GUARD DOES NOT COVER, STATED HERE BECAUSE THE ROW WOULD OTHERWISE IMPLY IT DOES. test_claims_ledger.py matches this row's distinguishing_phrase as a SUBSTRING of the figure's .txt sidecar, so it catches the exact withdrawn wording and nothing else; a reworded assertion of the same claim passes it silently. THE WITHDRAWAL IS SEMANTIC, NOT PHRASAL: what is deleted is the claim that an EPS matrix excludes anions the way Hanford intragranular porosity did, in ANY wording. The mechanical check is a floor under a human one, not a substitute for it."
PP-DEFF-01,repository,biofilms_radiodialysis.R:242 radial_parms comment; and the symbol D_eff throughout,"D_eff = 1e-3 cm2/s is the effective diffusivity, Table 2 range 1e-4..1e-2.",parameter_provenance,1e-3,cm2 s-1,false,true,biofilms_radiodialysis.R:152; biofilms_radiodialysis.R:242; biofilms_radiodialysis.R:293; biofilms_potts_jacc.jl:268,code,restate,"Restate the comment so it names WHICH quantity the value is, and drop the unresolvable range. One symbol currently carries three distinct quantities and the shipped value is none of them: (a) a film-scale effective diffusivity, which at minimum means D_eff = D0 * eps/tau, porosity over tortuosity, BOTH corrections reducing, so D_eff < D0 always; (b) the bench quantity, D_app = D_eff/R, lumped by sorption and separable only with K_d from a suspended isotherm; (c) the solver's own D_eff, which is UNRETARDED because biofilms_radiodialysis.R:10-13 carries sorption explicitly as a two-phase c/s system with k_ads, k_red, k_des and k_loss. SEPARATELY AND NOT DISCHARGED BY THIS ROW: the section 3.12 damage law presupposes D_eff,0 < D_w, which this default violates - that is PP-DAM-01 and it has its own verdict.","TWO FINDINGS WITH DIFFERENT VERDICTS, AND THE ROW STATES THE SCOPE OF EACH BECAUSE THEY ARE NOT EQUALLY WELL ESTABLISHED. FINDING 1, REFUTED, NEEDS NO PROVENANCE: the quoted span 1e-4..1e-2 cm2/s lies ENTIRELY ABOVE water self-diffusivity, about 2.3e-5 cm2/s. Since D_eff = D0 * eps/tau with eps < 1 and tau > 1, no value in that span can be a film-scale effective diffusivity of anything in water. The range is impossible as stated, independently of where it came from. FINDING 2, MIXED, AND HERE IS THE SET SEARCHED. The referent is named at biofilms_radiodialysis.R:152 as ""Table 2 of the paper and Deep Research synthesis (April 2026)"", and ""the paper"" is fixed at biofilms_radiodialysis.R:3 as Kinder & Faulkner 2026, preprint/modeling_radioresistance_and_radiotropic_fitness.tex. THAT HALF RESOLVES AND IS REFUTED: the file's four table labels in document order are tab:notation, tab:params, tab:term_magnitudes and tab:decided_moves, so Table 2 IS tab:params, it exists, and its only diffusion rows are per-species cell motility D_s in um2/s over 0.005 to 1.00 - no D_eff, no cm2/s, no 1e-4..1e-2. The superseded predecessor calibration/tests/fixtures/modeling_radiotrophic_fitness_prerevision.md carries no D_eff table either. That is a bounded enumeration over four tables in one named file plus one fixture, which is the strong form. THE OTHER HALF IS UNRESOLVED, NOT REFUTED: the ""Deep Research synthesis (April 2026)"" is NOT IN THIS REPOSITORY - a case-insensitive search for ""deep research"" over .md, .tex and .csv returns nothing outside the R file's own header - so the claim about it is that the comment cites a provenance I CANNOT RESOLVE, and this row asserts NOTHING about what that document contains. Saying which of the two is which is the point: refuted and unresolvable take different repairs."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Do not infer the diffusivity bound from water alone

The purportedly refuted half conflates the transported species' free diffusivity D0 with water self-diffusivity D_w: D_eff = D0 * eps/tau establishes only D_eff < D0, not D_eff < 2.3e-5 cm²/s. The referenced solver transports a generic mobile species and declares no corresponding D0, so the numerical range may be implausible but is not refuted independently of its missing provenance. The Table 2 attribution is refuted on content, but the range itself must remain unresolved until its solute and free-diffusion producer are identified.

AGENTS.md reference: AGENTS.md:L71-L77

Useful? React with 👍 / 👎.

Comment thread tools/absence_gate.py Outdated
Comment on lines +184 to +185
preceding = [d for d in def_lines if d <= lineno]
attached = bool(preceding)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve module-level classification after earlier tests

Once the file contains any preceding test definition, attached remains true for every later comment block, including blocks that have returned to module scope between tests. For example, the module-level no-signal registry beginning at calibration/tests/test_sop_index.py:164 is now reported as attached, so the stated attached-versus-backlog counts are wrong and later module contracts can never enter the module-level tier. Track actual test-body boundaries or indentation, and add a control with a module block after a completed test.

AGENTS.md reference: AGENTS.md:L19-L21

Useful? React with 👍 / 👎.

Comment on lines +67 to +68
out = subprocess.run(["pdftotext", "-layout", str(pdf), "-"],
capture_output=True, text=True).stdout

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Handle a missing pdftotext executable

In a calibration environment without Poppler, this call raises FileNotFoundError rather than exercising or skipping the artifact assertion. I reproduced this with python -m pytest -q calibration/tests/test_claims_ledger.py calibration/tests/test_figure_staleness.py calibration/tests/test_sop_index.py: 45 passed, 1 skipped, and this test failed here, while the later sidecar test already checks shutil.which("pdftotext"). Apply the same dependency gate here while retaining the hash and committed-sidecar bare-tier checks.

AGENTS.md reference: AGENTS.md:L41-L43

Useful? React with 👍 / 👎.

…ed a guard against an input that cannot occur

I REPORTED THIS AS A MISSING CONTROL. It was a missing guard. The gap named in
29bca4b was that the figure-sidecar path had no committed known-bad while the
manuscript path has ORIGINAL_FIXTURE and its detection floor. Recovering the
real pre-correction sidecar from c2219a2 and running the guard against it shows
the fixture was the smaller half: `_detected` returns [] for BOTH rows, and the
vocabulary floor did not carry the terms either. Between c2219a2 and this
commit the two withdrawn anion-exclusion sentences were enforced by nothing.

THE HAND-RUN THAT "CONFIRMED" THE PHRASE GUARD APPENDED CLEAN LINES. pdftotext
-layout reads a three-column figure ACROSS, so "Bromide is size- and
charge-excluded from fine porosity." comes back as "...charge-excluded
breakthrough lag gives a lumped run alongside every measurement. from fine
porosity." -- interleaved with the other two columns and contiguous nowhere. A
sentence appended as its own line is a shape the layout cannot produce, so the
check confirmed the guard against an input that never arrives. That is the
control-passing-on-a-mechanism-that-never-ran defect which
test_the_figure_guard_detects_the_pre_correction_artifacts was itself written to
avoid, committed inside the report of a different gap.

THE PREDICATE WAS WRONG, NOT JUST THE TIER ASSIGNMENT. The test split figure
rows by asking whether distinguishing_phrase returned anything, and put FIG-09
and FIG-10 in the phrase tier because their claim text yields a full sentence.
"Yields a phrase" is not "is covered": coverage is a property of the ARTIFACT.
So the test now asks the artifacts. test_no_figure_row_is_reachable_by_the_
phrase_guard asserts that no figure row's phrase is findable in any of the four
committed known-bads, AND that the vocabulary floor fires on every one of them
-- because "covered by nothing" and "covered by the floor" are otherwise the
same result, which is how this got through.

RETRACTED_IN_FIGURES gains "charge-excluded" and "same trap", absent from every
current sidecar and present in the known-bad. FIG-09 and FIG-10 are now named in
the coverage report beside FIG-01/02/05, with the difference stated: the first
three are unreachable because two-word labels never make MIN_WORDS, these two
are unreachable despite making it. Listing them together is the point.

AND THE CONTROL'S OWN GLOB PINNED THE SET IT WAS DERIVED FROM.
test_the_figure_guard_detects_the_pre_correction_artifacts globbed fig*_prefix
.txt and asserted len == 3. phase2_diffusion_cell_prefix.txt does not match
fig*, so a committed known-bad could sit in the fixtures directory covered by
nothing while the count assertion passed. Widened to *_prefix.txt and 4.

Mutation-checked, three ways, each restored: dropping the two new vocabulary
terms fails both figure tests; removing the fixture fails both; and putting the
withdrawn sentence back into the LIVE sidecar -- the regression that actually
matters -- fails test_no_committed_figure_asserts_a_retracted_phenotype and the
pdftotext contract. calibration 392 passed.

Prompted by review feedback that item 2 of the PR comment was a known gap with a
known repair and an existing pattern to copy, rather than a finding awaiting
adjudication. Copying the pattern is what surfaced that it was not item 2.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
@aurascoper

Copy link
Copy Markdown
Owner Author

Correction to the comment above before you start — head is now af48876, and item 2 is withdrawn as a finding because it was wrong about its own severity.

I reported a missing control. It was a missing guard. I said the figure-sidecar path lacked a committed known-bad while the manuscript path has ORIGINAL_FIXTURE. Recovering the real pre-correction sidecar from c2219a2 and running the guard against it shows the fixture was the smaller half: _detected returns [] for both FIG-09 and FIG-10, and the vocabulary floor did not carry their terms either. Between c2219a2 and af48876 the two withdrawn anion-exclusion sentences were enforced by nothing at all.

The reason is layout, and it invalidates the check I told you I had run. pdftotext -layout reads a three-column figure across, so Bromide is size- and charge-excluded from fine porosity. comes back as ...charge-excluded breakthrough lag gives a lumped run alongside every measurement. from fine porosity. — interleaved with the other two columns, contiguous nowhere. My hand-run appended both sentences as clean lines, which is a shape the layout cannot produce. So it confirmed the guard against an input that never arrives — the control-passing-on-a-mechanism-that-never-ran defect, committed inside the report of a different gap, in a file whose neighbouring control was written specifically to avoid it.

So the repair is not the one I described. The predicate was wrong, not the tier assignment. "Yields a phrase" is not "is covered" — coverage is a property of the artifact. test_no_figure_row_is_reachable_by_the_phrase_guard now asks the artifacts: no figure row's phrase is findable in any of the four committed known-bads, and the vocabulary floor fires on every one of them, because "covered by nothing" and "covered by the floor" are otherwise the same result. RETRACTED_IN_FIGURES gains charge-excluded and same trap. Mutation-checked three ways including the regression that matters — the sentence returning to the live sidecar.

One more instance of the shape you keep finding, in the control itself. test_the_figure_guard_detects_the_pre_correction_artifacts globbed fig*_prefix.txt and asserted len == 3. phase2_diffusion_cell_prefix.txt does not match fig*, so a committed known-bad could sit in the fixtures directory covered by nothing while the count assertion passed — a bound checked against the set it was derived from. Widened to *_prefix.txt and 4.


And item 1's diagnosis was off in a way worth stating, because the repair differs. I framed the hazard as an editorial trim failing both assertions loudly. That is not it. MIN_WORDS = 5 is a threshold nobody derived — the comment above it justifies 5 over 4 by collision anecdote — and it was sitting between two tiers of different strength, deciding coverage by a property of the claim text. That is the MIN_OUTER_DRAWS shape: a constant read off a convenient property rather than a usable one, with a silent downgrade on the far side and nothing announcing the crossing.

Asking the artifacts removes MIN_WORDS from tier assignment entirely, so that specific cliff is gone. What I have not fixed is that MIN_WORDS still governs whether any manuscript row is detectable at allyields_phrase >= 25 is a floor over a set the threshold itself defines, so lowering it to 4 or raising it to 6 moves both the measurement and the bar. If you want one thing from this comment, that is the one I would pick.

Suites: calibration 392 passed, Julia green.

aurascoper and others added 4 commits August 31, 2026 03:38
…ode that produced two unenforced rows

A CONTROL YOU BUILD BY HAND TESTS YOUR IDEA OF THE FAILURE. AGENTS.md rule 1
gains the case and the general form. The figure phrase guard was "confirmed" by
appending the two withdrawn sentences to a .txt sidecar as clean lines;
pdftotext -layout reads a three-column figure ACROSS, so the real extraction
interleaves the columns and the sentence is contiguous nowhere. FIG-09 and
FIG-10 were enforced by nothing for a full commit while a control reported
otherwise. The general rule: where a guard reads a DERIVED artifact, the
derivation's shape is part of the contract, and a hand-written input silently
opts out of it. Recover the known-bad from the path the real input travels.

AND IT EXPLAINS THE .txt SIDECARS RETROACTIVELY, which is why it is worth
writing down rather than filing under rule 1's existing bullets. The sidecars
exist because the FIG-01--07 lesson was that images hide claims, so the
extraction is what gets read. Nobody had recorded that the SHAPE of that
extraction is therefore load-bearing too. This is not domain-versus-range, not a
coverage gap, and not a vacuous assertion: it is a control that never met the
pipeline.

D3'S PENDING HALF NOW NAMES TASKS, because OPERATIONAL-pending is only
meaningful if the pending part does. Five clearing routes in cost order, each
with a clearing condition: DTIC ADA307995 by any available access; the Hoffman
question, which is free; NASBEE through an institutional subscription to settle
whether 3.54 is microbial or mammalian; a controlled-vocabulary search this
program has not run; and writing to the authors last, because it spends an
introduction. Explicitly labelled as unscheduled and unassigned -- listing a
task does not start it.

THE MEMO QUESTION CHANGED CHARACTER WITH THE FINDING AND NOW SAYS SO. It asked
whether a usable microbial neutron RBE exists. Since the barrier turned out to
be retrieval rather than naming, it now names ADA307995, says plainly that I
cannot open it, and asks whether he can. "Can you get at this" is a far easier
thing for someone at a reactor to answer than "does this exist", and the earlier
phrasing put the harder question to the person best placed to answer the easier
one.

Checked rather than assumed, since the subsection was written around it: the
four-item enumeration that replaced "every dose number this program carries" did
NOT drift back to a universal. absence_gate over the diff returned three
candidates; two disposed as definitional or as a disclaimer about this program's
own plan, and one was a real overreach -- "a controlled-vocabulary search, which
no one has run" claims more than this program can see, and is now "which this
program has not run".

absence_gate self-checks GATE USABLE; calibration 392 passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
… that asserted corroboration it did not have

SECTION 3.12 SAYS THREE SORPTION FORMS ARE NOT INTERCHANGEABLE AND NOWHERE SAID
BY HOW MUCH. The shipped immobile-phase equation is ds/dt = U*c - (k_des+k_loss)*s
with U = X_total*(k_ads + k_red*f_red) -- first order in c, no site limitation, a
Henry isotherm. The kinetic Langmuir form of the same reaction has equilibrium
s_L = s_H/(1 + s_H/X_max), so the shipped form's relative error is
s_H/(X_max + s_H) and the two diverge by more than 10% above c_ref = X_max/29.72.
Both explore passes confirmed no prose, comment or ledger row anywhere took a
position on this; it was an open gap, not a restatement.

THE CONDITIONALITY IS THE FINDING AND THE ROW IS WRITTEN SO IT CANNOT BE QUOTED
OTHERWISE. Over the range the solver visits -- c in [0, 0.9252], s in [0, 3.3024]
at t_end = 100 s, reaching 0.382 of Henry equilibrium since 1/(k_des+k_loss) =
166.7 s exceeds the run -- the error is 0.0066% at 1 mg/L with X_max = 50 and
86.9% at 10 g/L with X_max = 5. It straddles any threshold, so this is a fact
about a MISSING PARAMETER and not about the forms. NO ADEQUACY VERDICT IS
AVAILABLE: it is tempting to read the table as "at trace concentrations Henry is
fine", but nothing establishes a trace-contaminant reference concentration as the
relevant one -- 3.12 states the sorbate "is dimensionless, with c_ext normalised
to unity and no chemical identity attached", so there is no compound whose working
concentration could be looked up. claim_text states the threshold and its
undetermined side and stops.

WHICH IS ALSO WHY IT STAYS OUT OF THE MANUSCRIPT. 3.12 is the "specification, not
method" subsection whose discipline is not carrying numbers. A bound conditional
on a quantity the model lacks would read there as evidence the quantity exists.
The row records the condition under which that would change -- control 3 fails if
the divergence ever stops straddling, because an unconditional bound would be a
fact about the forms and would belong in the prose.

TWO ROBUSTNESS RESULTS, AND THEY ARE NOT THE SAME STRENGTH. s_H is proportional
to X_total while X_max = q_max*rho_dry, so at equilibrium X_total cancels against
rho_dry and the ratio reduces to (k_ads + k_red*f_red)*c*c_ref/((k_des+k_loss)*
q_max). The equilibrium crossover K = 77.72 is therefore immune to PP-66-08, the
open row saying X_total is a site-occupancy fraction in a g/cm^3 slot so the rate
is "wrong by an unstated rho". IT DOES NOTHING WHATEVER FOR q_max. X_max is a
prior on a prior -- a suspended-measured capacity times a density prior -- so
every crossover quoted is a prior. Separately, sweeping f_red_active across its
entire domain moves K only 26.77 to 36.47, so nothing rests on the one constant
labelled "Unvalidated placeholder".

analysis/henry_langmuir_bound.R is the producer, because a number nobody can
re-run is PP-62-11 and the P1 Codex raised on #23. THREE CONTROLS, AND THE FIRST
IS THE LESSON FROM THE LAST COMMIT APPLIED ONE FILE OVER: it integrates the
Langmuir RHS numerically on the same constants and requires the closed form to
match, rather than checking a derivation against its own algebra -- a control that
never meets the pipeline confirms a guard against an input the pipeline cannot
produce. Control 2 pins the crossover from BOTH sides, since a one-sided assertion
passes on any monotone curve wherever the threshold sits. Mutation-checked:
perturbing the closed form fails control 1 alone, moving K by 20% fails control 2
alone. Wired into model-contracts.yml with the receipt discipline the sibling
verifier already uses, and the CI step asserts K_transient == 29.72 -- "3 checks
passed" does not say which bound they passed about.

THE SIM_FILES EXCLUSION IS NOW ASSERTED RATHER THAN INSPECTED. manuscript_claims
_tests.jl asserts X_max/q_max/Langmuir appear in no simulation file, and the new
producer necessarily contains all three. SIM_FILE_NAMES is nine hand-maintained
entries, not a glob, so "the producer is outside it" would have been a property
nothing enforced -- the fixed-name-list hazard from the far side, a guard green
because a file quietly landed off the list. Inverted into an allow-list: every
file in the repository carrying any of the three terms must be declared, with a
planted-orphan control so the check cannot pass on a broken walk. 36 -> 40 tests.

HANDOUT-01'S AUDIT ASSERTED THREE CORROBORATING ARTIFACTS AND HAD ONE. The
manuscript states the forms correctly but in 3.12, not 3.11. wan-deck.tex NEVER
EXISTED UNDER ANY NAME -- git log --all --diff-filter=A over every path returns
nothing, wan_meeting_handout.tex was ADDED rather than renamed, and the sentence
quoted from it appears nowhere in the tree or history except inside the row, the
only commits containing the string being the ones that wrote it. No CV exists in
the repository. SO THE ROW DID NOT DRIFT; IT WAS WRITTEN FROM ARTIFACTS THAT WERE
PLANNED RATHER THAN READ -- the claim-written-from-a-summary family, and its most
consequential instance so far because the artifact in question is itself an audit.
An audit that asserts corroboration and delivers none does not merely fail to
corroborate; it manufactures the appearance of having checked. The row now says
"AUDIT OF DEPENDENTS" is itself the heading to distrust, scoped as prospective
because HANDOUT-01 is at this date the only row carrying it -- written now, while
there is one instance to point at.

AND CORRECTING IT TOOK TWO PASSES. The first fix said the locator was wrong in
two places. It was wrong in three: the distribution sentence, furthest from the
audit heading, still read "section 3.11" after both audit occurrences were fixed.
Recorded in the row rather than silently amended.

Left deliberately: suspended_isotherm_proposal.csv:29 cites biofilms_radiodialysis
.R:161 for c_ext = 1.0, now at line 266. The quoted wording matches exactly so the
reference resolves by content and nothing false is asserted; it belongs to a pass
that fixes line-number drift generally.

Verified: Rscript analysis/henry_langmuir_bound.R ALL PASS exit 0; julia
--project=. tests/runtests.jl green (Manuscript claims 40 pass, 1 pre-existing
broken); calibration 392 passed; absence_gate GATE USABLE, and its flags on the
new notes produced a real tightening -- "No CV exists anywhere in this repository"
now names the search and scopes both absence claims to this repository on this
date. Ledger 348 rows, 14 columns, no duplicates, no embedded newlines,
HANDOUT-01 still correctly reported as not textually detectable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…gument that never held

PP-DEFF-01 ARGUED FROM WATER WHEN IT NEEDED THE SOLUTE (P1). The row read "REFUTED,
NEEDS NO PROVENANCE: the quoted span 1e-4..1e-2 lies ENTIRELY ABOVE water
self-diffusivity ... no value in that span can be a film-scale effective diffusivity
of anything in water." D_eff = D0*eps/tau with eps<1, tau>1 establishes D_eff < D0 --
the SOLUTE's free diffusivity -- and not D_eff < D_w. The two were conflated, and this
solver transports an unnamed species: no D0, no D_free, no chemical identity anywhere,
only "mobile species" and "external contaminant concentration (normalised)".

THE REFUTATION SURVIVES ONCE THE PREMISE IS SUPPLIED, AND THE PREMISE IS NOW WHAT TO
CHECK. For a value in the span to be a film-scale D_eff the solute's D0 must exceed the
span's own lower end, 1e-4 cm2/s. Aqueous D0 is bounded by roughly that -- and since
that bound is itself an absence claim, in the row whose correction is about arguing
from an unnamed solute, it is stated with its exception rather than as a universal:
H+ moves by Grotthuss hopping rather than Stokes drag, reaches about 9.3e-5, is the
fastest aqueous species, and is still under the bound by about 8 percent. Carried as a
literature value the way PP-DAM-01 already carries 2.3e-5, not measured here.

SO THE REFUTATION IS NOT UNIFORM ACROSS THE SPAN IT REFUTES AND THE ROW SAYS SO. The
upper end is refuted by two orders of magnitude. The lower end clears the fastest ion
in water by 8 percent and is thin. A reader entitled to doubt one end is entitled to
doubt only that end.

AND THE SAME SESSION APPLIED THIS DISCIPLINE ONE ROW OVER AND SKIPPED IT HERE.
PP-SORP-01 refuses an adequacy conclusion PRECISELY BECAUSE the sorbate has no chemical
identity, citing section 3.12 for it. PP-DEFF-01 quietly assumed one. Both mine, same
day. Dependents audited rather than asserted: PP-DAM-01 also uses D_w and does NOT
inherit the defect, because section 3.12's damage law NAMES D_w as the interpolation
target, so no solute identity is required there.

THE ABSENCE GATE'S MODULE-LEVEL TIER WAS UNREACHABLE (P2), WHICH IS RULE 3 INSIDE THE
TOOL BUILT FOR RULE 3. Membership tested "is there a test definition above me", so once
a file's first test was passed every later block counted as attached -- including blocks
that had returned to module scope. Reproduced: test_sop_index.py reported 8 attached and
0 module-level while the no-signal registry at line 164 sat there misfiled. The header
advertises that module count as a standing backlog printed every run so a clean pass is
never an all-clear; it printed 0, and an unreachable tier and an empty tier print the
same 0. Membership is now INDENTATION, which is what "inside a body" means in both
languages it reads. Now 6 attached, 2 module-level, line 164 among them. New control is
the shape that was invisible -- a module-level contract AFTER a completed test -- and it
is asserted in BOTH directions, because the defect was misclassification and a check
that only asked "does the module tier find it" would pass while attached claimed it too.

THE RENDER RECORD BOUND THE SVG TO ITSELF (P1), AND render_figure_svg.py's OWN DOCSTRING
ARGUED THAT THIS WAS THE DEFECT BEFORE WRITING IT. Reproduced exactly: edit the SVG,
hand-write .svg.sha256, do not re-render -- three checks pass over a stale PDF.

A HASH CANNOT FIX A HASH. The renderer now records the pair (source digest plus the
digest of the PDF it produced, `split()[0]` unchanged so existing readers are
unaffected), but that only makes forgery need two lies instead of one; both lines are
still text anyone can write. The actual binding is CONTENT: every text run in the SVG
must appear in the committed .txt, which the sidecar test independently pins to the
PDF's bytes. SVG text -> committed text -> committed PDF, each link checked by something
a record edit cannot satisfy. Matching is on the first 5 words of each run because
pdftotext -layout wraps a long run mid-word -- an exact whole-run match fails on
wrapping rather than on staleness, which is the same column-layout fact that made
FIG-09/10 unreachable. SCOPE, NARROWER THAN "THE PDF IS FRESH": text drift only. An SVG
edit that moves a rectangle and touches no string passes, and nothing here would catch
it. That is the honest reach, and text is the failure mode these guards exist for --
FIG-01 through FIG-10 are every one of them a claim made in words inside an image.
Control mutates a real run from the artifact path rather than planting a synthetic one.

pdftotext IS DEPENDENCY-GATED NOW (P2). It raised FileNotFoundError instead of skipping,
so an environment without Poppler failed rather than reporting uncovered surface -- and
took the bare-tier checks, which need no binary, down with it. Verified with PATH
stripped: 1 skipped, 4 passed, rather than an error.

Verified: calibration 394 passed (was 392, +2 new); julia green, Manuscript claims 40
pass and the 1 broken is pre-existing; absence_gate GATE USABLE with five self-checks;
ledger 348 rows, 14 columns, no duplicates, no embedded newlines. Stale-PDF state
reproduced and restored; renderer re-run so the record is the generator's, not mine.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…is repository had

THE r IS A STATEMENT ABOUT AREA, NOT A COORDINATE ARTIFACT, AND THE GUIDE DERIVES IT.
A shell between r and r+dr has volume 2*pi*r*dr but faces of area 2*pi*r and
2*pi*(r+dr) -- different areas, which is the whole story. Conservation on that shell
gives dc/dt = -(1/r) d(rJ)/dr, and Fick's law turns it into the operator the solver
states at line 9. The r inside is flux x area; the 1/r outside is / volume. Expanded
for constant D the second term is purely geometric, and it carries a 1/r, which is
where the axis section starts.

PART 3 WAS CORRECTED BEFORE IT SHIPPED, AND THE CORRECTION IS THE BEST PART OF IT. The
outline described the idealised finite-volume treatment -- innermost inner face has zero
area so the axis enforces itself, Robin set directly as a face flux with no ghost node
-- and listed "the innermost inner-face area being zero rather than special-cased" as a
CHECK AGAINST SOURCE. This solver does neither. The axis is special-cased with the
L'Hopital limit written in directly at :110, and face_weights sets w_plus[1] and
w_minus[1] to NA_real_ on purpose so anything reaching for a metric weight there fails
loudly. The wall uses ghost-node elimination at :132. Only the interior is flux-form.

So the section says the code is a HYBRID -- flux-form interior, L'Hopital axis,
ghost-node wall, three treatments and three reasons -- which is more interesting than
the idealised version and is what the source shows. The idealised account stays as the
CONTRAST that explains why the middle needs no special case and the two ends do. Written
as outlined it would have shipped a claim the source contradicts, in the section whose
whole premise is that every claim resolves.

THE CITATION GUARD, BECAUSE NOTHING HERE VERIFIED A file:line. Searched tests/, tools/,
calibration/ and scripts/: the closest relatives are a regex SEARCH in
manuscript_claims_tests.jl and a git-SHA resolver in test_claims_ledger.py. Neither
takes a pre-named citation and asks whether it still points at what it says. Nineteen
citations, each with the fragment its line must contain.

THE FAILURE MESSAGE SPLITS TWO CASES, OR THE REPAIR BECOMES REFLEXIVE. File-exists and
line-exists are nearly free to satisfy; the fragment is what makes a citation resolve
rather than merely point. When an edit shifts a line the cheap repair is to bump the
number -- right if the code moved, wrong if the fragment moved because the code changed.
So on failure the guard searches the file: fragment found elsewhere reports THE CODE
MOVED with the corrected line; fragment found nowhere reports THE CODE CHANGED and says
to re-read before renumbering.

CONTROL SHIFTS BY ONE AND BY TEN. Off-by-one is exactly the case a fragment match can
survive by accident -- w_plus and w_minus are consecutive lines of face_weights(), as
are rtol and atol in the ode() call -- so a control shifting only by one could pass
while the guard was blind. Both required to fail, and both required to be diagnosed as
moved rather than changed. Known-bad drawn from a real citation in a real guide, not
hand-written.

AND THE GUARD HAD THE DEFECT IT WAS WRITTEN TO CATCH. The strict row parser silently
drops rows it cannot read, so its count is a SUBSET while the assertion over it would be
about the SET -- caught three times in this repository already, and it was sitting in
the guard for checking citations. A loose pattern now counts anything citation-shaped and
the strict parser must account for every one. Mutation-checked: an unparseable row fails
with "1 are being checked by nothing" rather than passing over 18 of 19.

THE ABSENCE GATE CAUGHT THE GUIDE'S OPENING SENTENCE. It read "Every claim about the code
carries a file:line", which is false -- the bc_coef reading and the Part 5 sorption rates
are prose-sourced and pinned by nothing. Now scoped: the tables are what is checked, those
two are named as not checked, and the file says so rather than implying uniform coverage.

Scope: Section 1 only, and the section says so. Floor is the chain rule, the product rule
and partial derivatives; ceiling is ODEs as objects an integrator solves. Out: curvilinear
coordinates, the divergence theorem in 3D, stability proofs, and sorption beyond noting it
sets a second timescale. Cross-references README's Mathematical framework rather than
restating its equations, since that section is target formalism under its own disclaimer
and this one describes what the code integrates -- the opposite side.

Part 5 carries the caveat the stability audit makes: the Julia explicit path is stable at
dt=0.5 only because r is in lattice units while D_eff is in cm2/s, and a margin that comes
from a units mismatch is not evidence the step is safe.

Verified: calibration 397 passed (was 394, +3); absence_gate GATE USABLE; all 19 citations
resolve and all 19 citation-shaped rows parse; guard mutation-checked in both directions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
@aurascoper

Copy link
Copy Markdown
Owner Author

@codex review

Head is a8ba2b0. Your five findings on 29bca4b are addressed — one was already fixed in af48876 before that review ran, four are fixed in e55f637. What to doubt:

e55f637 — the PP-DEFF-01 finding was right and the correction has a thin end. You were correct that D_eff = D0·ε/τ gives D_eff < D0, not D_eff < D_w, and that this solver declares no D0 and names no solute. I did not withdraw the refutation; I supplied the premise it was leaning on silently — for a value in 1e-4..1e-2 to be a film-scale D_eff the solute's D0 must exceed 1e-4 cm²/s. That premise is itself an absence claim, in the row whose correction is about arguing from an unnamed solute, so it is stated with its exception rather than as a universal: H⁺ moves by Grotthuss hopping rather than Stokes drag, reaches ~9.3e-5, is the fastest aqueous species, and clears the bound by about 8%. So the refutation is not uniform across the span it refutes — the upper end by two orders of magnitude, the lower end by 8%. If you think 8% is too thin to call refuted at that end, say so; I would rather split the verdict by endpoint than leave it uniform.

e55f637 — the render record, and I do not think a hash can fix this. Reproduced your scenario exactly: edit the SVG, hand-write .svg.sha256, do not re-render — three checks pass over a stale PDF. The renderer now records the pair, but that only makes forgery need two lies instead of one, since both lines are still text anyone can write. The actual binding is content: every <text> run in the SVG must appear in the committed .txt, which the sidecar test independently pins to the PDF's bytes. Its reach is narrower than "the PDF is fresh" and the code says so — an SVG edit that moves a rectangle and touches no string passes, and nothing here would catch it. I claim that is the right scope because FIG-01 through FIG-10 are every one of them a claim made in words inside an image. Worth attacking if you disagree.

e55f637 — the absence gate was rule 3 inside the tool for rule 3. Membership tested "is there a test definition above me", so past a file's first test every block counted as attached and the module-level tier was unreachable. test_sop_index.py reported 8 attached / 0 module-level while the registry at line 164 sat misfiled; the header advertises that 0 as a standing backlog, and an unreachable tier and an empty tier print the same 0. Now indentation-based: 6 / 2, line 164 among them. The new control asserts both directions, because the defect was misclassification and a one-sided check would pass while attached also claimed it.

a8ba2b0 — a new guard, and it had the defect it was written to catch. docs/guides/calculus_in_this_code.md cites 19 file:line locations with the fragment each line must contain; nothing in this repository verified a citation before it. The strict row parser silently dropped rows it could not read — subset counted, set asserted, which you have now caught here three times — so a loose pattern counts anything citation-shaped and the strict parser must account for every one. The part I would most like attacked: the failure message splits "the code moved" from "the code changed" by searching the file for the fragment, and the control shifts a real citation by one and by ten, because off-by-one can survive a fragment match by accident on adjacent similar lines (w_plus/w_minus, rtol/atol).

And the guide's own opening sentence was wrong until the absence gate flagged it. It read "Every claim about the code carries a file:line". It does not — the bc_coef reading and the Part 5 sorption rates are prose-sourced and pinned by nothing. Now scoped, with those two named as unchecked.

Suites: calibration 397 passed, julia green, absence_gate GATE USABLE, Rscript analysis/henry_langmuir_bound.R ALL PASS.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a8ba2b0af4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +136 to +137
return [s for s in svg_text_runs(svg_source)
if " ".join(s.split()[:MATCH_WORDS]) not in flat]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Compare more than each text run's opening words

Fresh evidence in the current head is that the replacement binding truncates every SVG text run to its first five words. I reproduced the stale-artifact scenario by changing the later 43× text in phase2_diffusion_cell.svg; unrendered_runs still returned an empty list against the unchanged PDF sidecar, and the control misses this because it deliberately mutates the first word. Any corrected number or conclusion after an unchanged opening phrase can therefore diverge between SVG and PDF while this guard passes, so compare the complete normalized token sequence (allowing extraction line wrapping) and add a tail-edit control.

AGENTS.md reference: AGENTS.md:L43-L47

Useful? React with 👍 / 👎.

Comment thread data/claims_ledger.csv Outdated
FIG-08,preprint/figures/phase2_diffusion_cell.pdf,panel: Open choice: the tracer run,In Hanford sediments tritiated water resolved intragranular pore volume where bromide did not.,empirical_literature_claim,,none,not_applicable,false,,literature,keep,"Name the source in the figure before the panel reaches Dr. Deng. Done 2026-08-31: Hay, Stoliker, Davis & Zachara 2011, Water Resources Research, doi:10.1029/2010WR010303.","SHIPPED UNCITED IN A FIGURE ADDRESSED TO A NAMED EXTERNAL READER WHOSE FIELD THIS IS - the configuration that produced HOFFMAN-01 and HOFFMAN-02. Verified 2026-08-31 against the USGS publication page for the paper, which carries the authors' own summary: sediments are the uranium-contaminated vadose and capillary-fringe materials beneath the former 300A process ponds at Hanford, and ""experiments using bromide ion as a tracer yielded very different results, suggesting very little penetration of bromide into the intragranular porosity"". The Wiley abstract page returns 403 and was not read. WHAT WAS DELIBERATELY NOT CARRIED INTO THE FIGURE: volume 47 and article number W10531 appeared only in search-engine summaries, in no page actually fetched, and a specific supplied by a secondhand summary and not seen in the source is precisely HOFFMAN-01. The panel cites year, journal and DOI, all three read off the USGS page."
FIG-09,preprint/figures/phase2_diffusion_cell.pdf,panel: Open choice: the tracer,Bromide is size- and charge-excluded from fine porosity.,mechanism_claim,,none,not_applicable,false,,literature,delete,"Remove the mechanism attribution. Done 2026-08-31: the panel now states the observed result and cites it, and asserts no mechanism.","THE SOURCE REPORTS THE OBSERVATION, NOT THE MECHANISM. Hay et al. 2011 report very little bromide penetration and infer restricted access; ""anion exclusion"" as the named cause came from a search-engine summary of that paper, not from any page read. Size exclusion and charge exclusion are two mechanisms and the figure asserted both, unsourced, in a sentence that read as established. The deleted mechanism does not change what the panel is FOR: the tracer choice question stands on the observation alone. WHAT THE GUARD DOES NOT COVER, STATED HERE BECAUSE THE ROW WOULD OTHERWISE IMPLY IT DOES. test_claims_ledger.py matches this row's distinguishing_phrase as a SUBSTRING of the figure's .txt sidecar, so it catches the exact withdrawn wording and nothing else; a reworded assertion of the same claim passes it silently. THE WITHDRAWAL IS SEMANTIC, NOT PHRASAL: what is deleted is the claim that an EPS matrix excludes anions the way Hanford intragranular porosity did, in ANY wording. The mechanical check is a floor under a human one, not a substitute for it."
FIG-10,preprint/figures/phase2_diffusion_cell.pdf,panel: Open choice: the tracer,an anionic EPS matrix is the same trap,transfer_claim,,none,not_applicable,false,,literature,delete,"Withdraw the transfer or source it from the EPS literature. Done 2026-08-31: withdrawn, and the panel now says the same holds in an EPS gel is untested.","A SEDIMENT RESULT TRANSFERRED TO A HYDROGEL WITHOUT A SOURCE IN THE HYDROGEL LITERATURE. What the EPS literature actually reports is the opposite sign of concern: Stewart 2003, J. Bacteriol. 185:1485-1491, compiles relative effective diffusivities in biofilms and puts small inorganic ions at De/Daq of roughly 0.56 to 0.70 - HINDERED, NOT EXCLUDED - and the well-documented charge effect in anionic EPS is retardation of CATIONS, e.g. tobramycin sequestered by alginate, not exclusion of small anions. Hanford intragranular porosity is nm-scale voids in a rigid mineral aggregate; EPS is an open gel above 90 percent water where the Debye length at physiological ionic strength is of order 1 nm. The two systems share a word, not a measured behaviour. THIS WOULD HAVE BEEN THE FIRST THING DR. DENG SAW, and the panel asking him a question is the wrong place to assert something in his field that the field does not say. WHAT THE GUARD DOES NOT COVER, STATED HERE BECAUSE THE ROW WOULD OTHERWISE IMPLY IT DOES. test_claims_ledger.py matches this row's distinguishing_phrase as a SUBSTRING of the figure's .txt sidecar, so it catches the exact withdrawn wording and nothing else; a reworded assertion of the same claim passes it silently. THE WITHDRAWAL IS SEMANTIC, NOT PHRASAL: what is deleted is the claim that an EPS matrix excludes anions the way Hanford intragranular porosity did, in ANY wording. The mechanical check is a floor under a human one, not a substitute for it."
PP-DEFF-01,repository,biofilms_radiodialysis.R:242 radial_parms comment; and the symbol D_eff throughout,"D_eff = 1e-3 cm2/s is the effective diffusivity, Table 2 range 1e-4..1e-2.",parameter_provenance,1e-3,cm2 s-1,false,true,biofilms_radiodialysis.R:152; biofilms_radiodialysis.R:242; biofilms_radiodialysis.R:293; biofilms_potts_jacc.jl:268,code,restate,"Restate the comment so it names WHICH quantity the value is, and drop the unresolvable range. One symbol currently carries three distinct quantities and the shipped value is none of them: (a) a film-scale effective diffusivity, which at minimum means D_eff = D0 * eps/tau, porosity over tortuosity, BOTH corrections reducing, so D_eff < D0 always; (b) the bench quantity, D_app = D_eff/R, lumped by sorption and separable only with K_d from a suspended isotherm; (c) the solver's own D_eff, which is UNRETARDED because biofilms_radiodialysis.R:10-13 carries sorption explicitly as a two-phase c/s system with k_ads, k_red, k_des and k_loss. SEPARATELY AND NOT DISCHARGED BY THIS ROW: the section 3.12 damage law presupposes D_eff,0 < D_w, which this default violates - that is PP-DAM-01 and it has its own verdict.","TWO FINDINGS WITH DIFFERENT VERDICTS, AND THE ROW STATES THE SCOPE OF EACH BECAUSE THEY ARE NOT EQUALLY WELL ESTABLISHED. FINDING 1, REFUTED UNDER A NAMED PREMISE. CORRECTED 2026-08-31: an earlier version of this row read 'REFUTED, NEEDS NO PROVENANCE ... lies ENTIRELY ABOVE water self-diffusivity, about 2.3e-5 cm2/s ... no value in that span can be a film-scale effective diffusivity of anything in water.' THAT ARGUMENT DOES NOT HOLD AS WRITTEN. D_eff = D0 * eps/tau with eps < 1 and tau > 1 establishes D_eff < D0 -- the SOLUTE's free diffusivity -- and NOT D_eff < D_w. The two were conflated. This solver transports an unnamed species: it declares no D0, D_free or chemical identity anywhere, saying only 'mobile species' (biofilms_radiodialysis.R:8, :50) and 'external contaminant concentration (normalised)' (:266). Raised by Codex on pull request #23. THE REFUTATION SURVIVES ONCE THE PREMISE IS SUPPLIED, AND THE PREMISE IS THE PART TO CHECK. For any value in the span to be a film-scale D_eff, the solute's D0 would have to exceed the span's own lower end, 1e-4 cm2/s. Aqueous D0 at room temperature is bounded by roughly 1e-4 cm2/s. SCOPE ON THAT BOUND, BECAUSE IT IS ITSELF AN ABSENCE CLAIM AND IT IS NOW WHAT THE VERDICT RESTS ON: it is a claim about aqueous solutes generally, carried from the literature and not measured here -- the same declaration PP-DAM-01 makes for D_w = 2.3e-5. The near-miss is the interesting case and is recorded rather than smoothed over: H+ migrates by Grotthuss proton hopping rather than by Stokes drag and reaches about 9.3e-5 cm2/s, the fastest aqueous species, and it is still under the bound -- by about 8 percent. SO THE REFUTATION IS NOT UNIFORM ACROSS THE SPAN IT REFUTES, and the row says so instead of asserting that no value in it can be a D_eff. The upper end, 1e-2, is refuted by two orders of magnitude and is not close. The lower end, 1e-4, clears the fastest ion in water by roughly 8 percent and is thin. A reader entitled to doubt one end of this is entitled to doubt only that end. AND THE SAME SESSION APPLIED THIS DISCIPLINE ONE ROW OVER AND SKIPPED IT HERE. PP-SORP-01 refuses to draw an adequacy conclusion PRECISELY BECAUSE the sorbate has no chemical identity, citing section 3.12 for it. This row quietly assumed one. Both were written on 2026-08-31. AUDIT OF DEPENDENTS, RUN RATHER THAN ASSERTED: PP-DAM-01 also uses D_w = 2.3e-5 and does NOT inherit this defect -- the section 3.12 damage law NAMES D_w as the interpolation target, so D_w is the correct quantity there and no solute identity is required. Searched for other rows arguing from D_w: those two. FINDING 2, MIXED, AND HERE IS THE SET SEARCHED. The referent is named at biofilms_radiodialysis.R:152 as ""Table 2 of the paper and Deep Research synthesis (April 2026)"", and ""the paper"" is fixed at biofilms_radiodialysis.R:3 as Kinder & Faulkner 2026, preprint/modeling_radioresistance_and_radiotropic_fitness.tex. THAT HALF RESOLVES AND IS REFUTED: the file's four table labels in document order are tab:notation, tab:params, tab:term_magnitudes and tab:decided_moves, so Table 2 IS tab:params, it exists, and its only diffusion rows are per-species cell motility D_s in um2/s over 0.005 to 1.00 - no D_eff, no cm2/s, no 1e-4..1e-2. The superseded predecessor calibration/tests/fixtures/modeling_radiotrophic_fitness_prerevision.md carries no D_eff table either. That is a bounded enumeration over four tables in one named file plus one fixture, which is the strong form. THE OTHER HALF IS UNRESOLVED, NOT REFUTED: the ""Deep Research synthesis (April 2026)"" is NOT IN THIS REPOSITORY - a case-insensitive search for ""deep research"" over .md, .tex and .csv returns nothing outside the R file's own header - so the claim about it is that the comment cites a provenance I CANNOT RESOLVE, and this row asserts NOTHING about what that document contains. Saying which of the two is which is the point: refuted and unresolvable take different repairs."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep the lower diffusivity endpoint unresolved

Fresh evidence in the current head is that the corrected row makes its new ceiling explicitly conditional on room temperature, while neither biofilms_radiodialysis.R nor the manuscript declares a transport temperature. Because aqueous free diffusivity is temperature-dependent, the cited 9.3e-5 cm²/s value cannot bound the unnamed solute under the solver's undeclared conditions; an 8% margin at one temperature therefore does not refute the 1e-4 endpoint. The upper end can remain refuted, but the lower endpoint should stay unresolved until the producer declares the solute and temperature.

AGENTS.md reference: AGENTS.md:L96-L97

Useful? React with 👍 / 👎.

aurascoper and others added 11 commits August 31, 2026 11:37
… that mostly rediscovered what this repo already knew

TWO COMMISSIONED EXTERNAL REVIEWS ARRIVED AND MOST OF THEIR CLAIMS ABOUT THIS CODEBASE
WERE ALREADY KNOWN HERE, TWO OF THEM MECHANICALLY ENFORCED. §3.4 already carried the
specification-not-method disclaimer and the full no-momentum declaration, including a
biofilms_3d.R velocity-array nuance the reviews miss. The nutrient field's inertness is
already asserted by manuscript_claims_tests.jl, which greps every lattice-mutating
function for nutrient reads. The §6.3 SE-and-direction caveat is already in the abstract,
§6.3, README and two ledger rows. The novelty claim the reviews ask to scope does not
exist -- the actual one is narrowly about OpenMC-to-CPM and already self-limits.

SO THE VALUE WAS NOT THE RECOMMENDATIONS. It was that rediscovering the reference defect
surfaced FOUR LEDGER VERDICTS RECORDED, PRESCRIBED, AND NEVER APPLIED. All four are now
applied: PP-REF-01 and PP-REF-02 (both bibliography entries carried DOIs resolving to
entirely unrelated papers -- a 3D MHD regularity paper and a neural-network
quadratic-programming paper), PP-310-01 (§3.10 stated FENE and Kelvin-Voigt with no
disclaimer), and RM-G04-01 (the five-term Hamiltonian, restated in README while
biofilms_potts.jl:16 went on listing H_pairwise, which is real but called from
take_snapshot alone and never enters acceptance).

THE WITHDRAWN DOIs ARE NOT REPEATED IN THE MANUSCRIPT, DELIBERATELY. The first draft of
the correction named them in LaTeX comments, which put withdrawn identifiers back into
the document a reader copies from and would have fired the new guard on its first run.
The comments now point at the ledger row; the recording layer holds the withdrawn fields
and the manuscript holds none.

ONE SYMPTOM, TWO CAUSES, AND FILING THEM AS ONE WOULD HAVE LEFT HALF OPEN. Cause 1 is the
SCOPE of enforcement: a verdict is enforced only in the file its `document` column names,
so RM-G04-01's claim in a .jl comment is read by nothing. Cause 2 is the phrase EXTRACTOR:
a citation splits on commas into runs under MIN_WORDS=5, so distinguishing_phrase returns
'' and PP-REF-01/02 were delete verdicts covered by nothing. Widening which files get
scanned does not reach the reference rows. RETRACTED_CITATIONS closes cause 2 and REACHES
PP-REF-01 AND PP-REF-02 AND NOTHING ELSE; RM-G04-01 remains unenforced pending
RETRACTED_IN_SOURCES, and the red-team doc says so in those terms.

THE GUARD'S SCOPE IS AN INVERTED ALLOW-LIST OF EXPLICIT NAMES, NOT A GLOB. Recording a
withdrawal requires naming the DOI, so the ledger and red-team docs necessarily contain
these strings -- the shape of a test whose own fixture holds what it forbids. Every file
carrying a withdrawn DOI must therefore be a DECLARED recording document, so a new file
fails until consciously admitted. docs/research/*_redteam.md was rejected as a glob
because it would auto-admit every future red-team file, which is the property the
allow-list exists for. IT FOUND TWO FILES I HAD NOT CLASSIFIED on its first run: a
__pycache__ .pyc holding this module's own constants (excluded as a build artifact) and
the pre-revision manuscript fixture, which carries both DOIs because it IS the known-bad
and is now declared.

§3.4 GAINS THE PHYSICS, WHICH IS THE REVIEWS' ONE GENUINE ADDITION. It said no momentum
state IS integrated; it now says one WOULD BE INAPPROPRIATE. Re = 2.0e-5 for a 1 um cell
at 20 um/s, 2.0e-3 even for a 100 um feature; tau_p = 5.6e-8 s at a = 0.5 um; ten orders
below the 18-minute relaxation. The length and speed are stated with the numbers because
Re is a property of the flow, not the organism. analysis/overdamped_regime.py reproduces
them all.

AND ITS CONTROL STRADDLES THE RATIO, NOT THE REYNOLDS NUMBER, BECAUSE THE OBVIOUS CONTROL
TESTS THE WRONG QUANTITY. A millimetre object at cm/s gives Re = 10 and looks like a
failing input -- but its T_bio/tau_p is 1.9e4, four hundred times the genuine control's
49. The control is a 1 cm object, where tau_p reaches 22 s and the ratio falls to 49.
Separately, my first version of that assertion OVERSTATED and failed on its first run: I
claimed the Re>1 input still passes the 1e8 ratio floor, and 1.9e4 does not. The weaker
true claim is the one that makes the point, and the docstring was corrected to match the
code rather than left contradicting it.

A THIRD DEFECT SURFACED WHILE APPLYING THE FIXES: §3.4 referred implementing symplectic
integration to §7.5 as future work and §7.5 does not list it. Dangling cross-reference,
now removed -- and it would have contradicted the appropriateness argument anyway.

AND THE CITATION GUARD FROM YESTERDAY CAUGHT MY OWN EDITS. Shifting biofilms_potts.jl and
the .tex moved three cited lines; it diagnosed all three as THE CODE MOVED with the
corrected numbers, which is exactly the split it was built for. Renumbered 660->695,
912->947, 1444->1453.

TWO OF THE REVIEWS' ERRORS CAME FROM THE BRIEF, WHICH WAS WRITTEN HERE. The briefs
asserted that §3.10 was marked specification-not-method and that the model is
CPM-with-adhesion-only; both came back reported as findings. A premise returned as a
finding is not independent confirmation and reads exactly like one. Recorded in the
red-team doc because the repairs differ: an asserted-X defect is fixed in the review, a
told-X defect is fixed in how the next review is commissioned.

Left open and named rather than folded in: RETRACTED_IN_SOURCES for cause 1; "no flow" is
undeclared in §7.5 and the reviews are right about it; every literature claim in both
documents is needs_verification and none is adopted.

Verified: calibration 400 passed (was 397); julia green, Manuscript claims 40 pass and the
1 broken is pre-existing; absence_gate GATE USABLE; overdamped_regime.py ALL PASS exit 0;
zero withdrawn DOIs in the manuscript and both replacements present, with the guard's
control drawn from the pre-fix .tex recovered from git. latexmk is not installed locally
so CI's manuscript-build covers the LaTeX; structural check passes 61 bibitems and
balanced environments.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…onal, and the near-miss was a new family member

I EXCLUDED BY FILE TYPE AND JUSTIFIED IT AS "BUILD ARTIFACTS ARE NOT DOCUMENTS". That is
the wrong rule and it had different future coverage from the right one. The skip read
`path.suffix in {.pyc, .pdf, .png, .so, .o}`, which silently admits any FUTURE artifact of
those types carrying a forbidden string -- and `preprint/*.pdf` is the manuscript's own
build output, the one artifact a reader actually receives. The rule excluded exactly the
file the figure work spent a commit learning to assert on.

MEASURED BEFORE CHANGING IT: with no type exclusion at all, the 14 PDFs in the tree produce
zero hits, because PDF text is compressed. The suffix rule bought nothing and cost
coverage, which is the worst combination available.

THE NARROW RULE IS "DERIVED FROM A SOURCE ALREADY SCANNED", AND IT IS NOW CHECKED RATHER
THAN ASSERTED. A .pyc is excludable because it is a compiled copy of a .py that is itself
in scope, so its content is already being read at its source. The skip now confirms the
sibling source exists and scans the file if it does not, so an orphaned cache with no
source in scope is covered rather than waved through. Mutation-checked: a .tex planted in
preprint/ with a withdrawn DOI is caught, restored after.

AND THE LATEX-COMMENT NEAR-MISS IS A NEW MEMBER OF THE USE-VERSUS-MENTION FAMILY, NOT A
REPEAT. RETRACTED_CITATIONS would have caught it; the DESIGN would not have, and that is
the finding. Use-versus-mention was resolved at the FILE level -- which documents may name
a withdrawn DOI. A .tex comment defeats that without violating it: it sits inside a
PERMITTED file at a location that is not part of the rendered document. It is neither use
nor mention in the sense the allow-list encodes, and it is invisible to a reader of the PDF
while fully visible to a reader of the source -- who is the one who copies the string,
which is the entire reason withdrawn identifiers are kept out of the manuscript.

So a file-level allow-list is insufficient for any artifact with a rendered form and a
source form. The repair taken was to remove the strings rather than teach the guard about
comments, because the manuscript has no reason to carry a withdrawn DOI in any form.
Recorded in the red-team doc so the next guard of this shape decides its scope at the level
of RENDERED VERSUS SOURCE and not only of file.

Verified: calibration 400 passed; the narrowed exclusion mutation-checked in both
directions. The manuscript's LaTeX build remains unverified locally -- latexmk is not
installed -- and CI's manuscript-build is the real check, not a formality: three of the
recent changes are in the .tex and a malformed \bibitem produces a wrong-looking
bibliography rather than an error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…and only the source was guarded

I COMMITTED AND PUSHED TWO FILES I HAD NEVER READ. docs/guides/calculus_in_this_code.tex
and .pdf appeared on disk while I was working; `git add -A` swept them into 4036630, whose
message describes only the exclusion-rule change. That is a commit message that does not
match its contents and an artifact distributed without being opened.

AND THE PDF WAS STALE IN EXACTLY THE WAY THE GUIDE FORBIDS. Earlier edits to
biofilms_potts.jl and the manuscript shifted three cited lines; the .md was renumbered
660->695, 912->947, 1444->1453 and the render was not. pdftotext on the committed PDF
showed biofilms_potts.jl:1444, a line that no longer contains what the table claims.
test_guide_citations.py globbed *.md, so every check passed over it. Prose right, artifact
wrong, guard reading only the prose -- the figure-staleness defect in a second location,
shipped in the same commit where I recorded that a file-level guard is insufficient for
anything with a rendered form and a source form.

FIXED PROPERLY RATHER THAN BY DELETING THE ARTIFACT. tectonic compiles the .tex here
(xelatex is absent, tectonic and lualatex are not), so the three line numbers are corrected
in the .tex and the PDF is rebuilt from it. The new binding asserts BOTH directions: every
citation the .md makes must appear in the rendered text layer, AND no line number the
source has moved away from may survive there. The second half is what catches staleness --
the first is satisfied by a PDF containing the right numbers somewhere among the wrong ones.
Mutation-checked: renumbering the .md without re-rendering reports the new number missing
and the old one lingering.

THE MATCHER HAD TO BE WRITTEN AGAINST THE EXTRACTION, NOT AGAINST THE SOURCE, AND MY FIRST
VERSION WAS WRONG. The template breaks long paths at underscores to fit the table column,
and pdftotext -layout puts the halves in DIFFERENT COLUMNS -- `biofilms_` on one side,
`radiodialysis.R:223` on the other -- so whitespace-joining cannot rejoin them. Matching on
the full path reported all fourteen citations missing from a PDF that contained every one.
Checked against the real text layer instead of assumed: the segment after the final
underscore survives intact for all seven paths. That is the extraction-shape lesson the
figure sidecars taught, arriving in a second artifact and catching me the same way.

SCOPE, NARROWER THAN "THE PDF IS CURRENT", AND STATED IN THE TEST: line numbers only. A
prose edit to the .md that touches no citation leaves the PDF stale and this check green.
The .tex is a hand-maintained rendering rather than a generated one, so nothing binds its
PROSE to the .md's -- a real and open gap, named rather than implied away.

The .tex was read before being kept, which is what I should have done before committing it.
It is a faithful rendering: it carries the corrected hybrid framing for part 3, with the
idealised finite-volume account inside the "On that telling" contrast and "This solver does
neither of those things" immediately after, and all five check-tables intact. Two earlier
greps reported the tables absent and were both artifacts of my own patterns -- LaTeX escapes
underscores, and this template wraps paths in a \us macro.

Verified: calibration 401 passed (was 400); the binding mutation-checked in both directions;
zero withdrawn DOIs in the manuscript; tectonic build reproducible from the committed .tex.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…calibrated against three of your own openings

I HAD NO STYLE SPEC AND SAID SO BEFORE GUESSING. Three neural-memory recalls, including
retired records, returned only engineering claims matching on shared vocabulary; the
semantic branch never ran because no embedder is configured, so recall there is lexical
only and a spec surfaces only if the query words literally appear in it. The Substack
profile and homepage are JS shells and a guessed post slug 404'd. The RSS feed was the
path that worked.

THE CALIBRATION IS THREE VERBATIM OPENINGS, NOT AN IMPRESSION OF THEM.

  "Four changes. Every one of them improved a representation-quality metric, and every
   one of them left the planner flat or drove it backwards."
  "Eight point eight times. That is the load-bearing number in the entire Cerebras moat."
  "Do not trade. Do not allocate. Do not treat as a viable alternative"

One signature, three times: a short bare fragment, then the sentence that loads it.
Numbers spelled out when they carry the headline -- "Eight point eight times", not 8.8x.
Emphasis carried by parallelism and sentence length rather than by bold. And the finding
IS the lede; the caveats come after.

THE GUIDE DID THE EXACT REVERSE. Six bolded meta-blocks -- What this is, What is
mechanically checked, Floor, Ceiling, Relationship to the README, Deliberately out of scope
-- before the first idea. It now opens on the operator itself, with the two r's and the
shell argument in the first four lines, and the scope apparatus compressed underneath where
it belongs. Bold lead-ins across the guide: 26 down to 13. The ten that went were emphasis
by reflex; the ones that navigate a derivation -- "Take a shell.", "Write conservation.",
"The construction." -- were kept, because those are doing work.

WHAT I DID NOT DO: apply prose-craft-2's Posture Board. Its axes are nonlinear time, moral
ambivalence, escalation, subplots, reader address -- calibrated for fiction, and forcing
them onto a reference document whose every claim is test-checked would damage the thing
that makes it worth having. The skill says as much: the calibration is a counterweight to
Claude's defaults, not a target to chase. I used the craft half and said which half I was
leaving.

AND THE .tex WAS ALREADY MORE ACCURATE THAN THE .md, SO THE PORT WENT BOTH WAYS. It named
THREE prose claims that no table pins -- bc_coef as P_eff versus k_L, the sorption rates,
and the stiffness characterisation -- where the .md named two. The .md now says three. The
.tex is not a pure render and had diverged before I touched either; its header already
declared itself "a derived artifact that can go stale independently" of the source, which
is precisely the gap the last commit's guard covers for citations and does not cover for
prose.

THAT GAP IS LIVE AND THIS COMMIT WALKED THROUGH IT. Editing the .md's prose left the .tex
and PDF holding the old wording, and every citation check stayed green, because the binding
covers line numbers only -- exactly the scope the test states about itself. The three files
were hand-synced and the PDF rebuilt from the corrected .tex with tectonic; all ten edits
verified present in .md, .tex and the rendered text layer, and neither withdrawn phrasing
survives anywhere. Hand-syncing is what the open RETRACTED_IN_SOURCES-shaped gap costs
until something binds prose the way line numbers are now bound.

Verified: calibration 401 passed; ten edits confirmed in all three artifacts; tectonic
build reproducible from the committed .tex.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…ond unstated premise, and an assertion on a count where the set was already computed

CODEX RAISED TWO FINDINGS ON a8ba2b0, BOTH AGAINST CODE I WROTE THIS SESSION, BOTH
CORRECT. Verifying them turned up two further defects Codex did not name, one of them
inside my own verification.

THE FIGURE BINDING ONLY EVER READ THE HEAD OF A RUN (P1). unrendered_runs matched the
first five words, so everything after word five could diverge between SVG and PDF while it
passed. The worst instance is the one that matters: changing 43 to 99 in the phase2 caption
escaped entirely, and 43 is THE NUMBER THE FIGURE WAS CORRECTED FOR, at word 24 of a
27-word run.

AND MY FIRST REPRODUCTION WAS A NO-OP THAT I NEARLY REPORTED AS CONFIRMATION. I mutated the
NORMALISED run against a source that stores entities (&#215;, &#8217;), so `victim in src`
was False and str.replace changed nothing. The guard then truthfully reported nothing
missing -- over an unmodified file. THIS IS A DISTINCT FAILURE FROM THE ONES THIS SESSION
HAS BEEN CATALOGUING: not a vacuous assertion, not a control that never met the pipeline,
not a subset counted where a set was meant. The assertion was correct and THE INPUT NEVER
CHANGED, which no amount of scrutinising the assertion would find. The repair generalises
to every mutation check: assert that the mutation applied, and assert WHERE.

THE MATCHER WAS CHOSEN BY MEASUREMENT. Four candidates against all 39 runs: full-run
whitespace-collapsed 1 false failure, whitespace-removed 1, first-five-words 0, head5+tail5
1. Every full-coverage candidate failed on the SAME run -- the one containing 43 -- because
26 of its 27 words match exactly and only the final token is clipped at the column
boundary. Hence: all words but the last, contiguous; the last may be clipped. Zero false
failures, catches head, middle and tail. The comment records that this rests on the CURRENT
layout, so a future mid-run failure is diagnosed as layout drift rather than as tampering.

FOUR CONTROLS, AND THE FOURTH IS THE TEST-SET ITEM. Head, middle and tail all exercise the
same run -- and the clipped-final-token rule was DERIVED from that run, which is the
training-set problem in a new place. A fourth control mutates a different run whose final
token is not clipped. tools/absence_gate.py already labels its own controls this way and
the label is borrowed deliberately. Mutation-checked: reverting the matcher fails the
middle edit.

ONE ASSERTION IS A RESTATEMENT AND THE CODE NOW SAYS SO. The identity check that the
mutation landed in the selected run cannot fail while _mutate_inside is correct, because
that function locates by content and replaces only within the located slice. Weakening it
to a tautology leaves the suite green -- measured. Kept as documentation of the invariant,
labelled as not a check.

PP-DEFF-01 NEEDED A SECOND UNSTATED PREMISE, WHICH IS THE FINDING (P2). The row bounded
aqueous D0 "at room temperature", and neither the solver nor the manuscript declares a
transport temperature -- biofilms_radiodialysis.R carries only T_cpm, which is not one. The
arithmetic is worse than the finding stated: D scales as T/eta(T), so H+ at 9.3e-5 implies
1.42e-4 at 37 C. AT BODY TEMPERATURE THE FASTEST AQUEOUS ION EXCEEDS THE 1e-4 ENDPOINT. The
8 percent margin does not thin, it inverts.

Upper end stays refuted and needs none of this. Lower end returns to UNRESOLVED. And the
disposition rule goes IN THE ROW rather than only in this message, because it is general
and this row is now its worked example: A REFUTATION THAT NEEDS A NEW PREMISE EACH TIME IT
IS EXAMINED IS BEING PROPPED UP. Second revision, each repair supplying a missing premise --
first the solute, then the temperature. The honest reading is that the lower endpoint was
never a refutation; it was an argument with two unstated inputs.

THE ORPHAN ASSERTION COMPARED A COUNT WHILE THE FUNCTION ALREADY RETURNED THE SET. Cite one
entry and orphan another in the same commit and length stays 15 and the test passes. Now
`unused == KNOWN_UNUSED` over named keys, printing newly-unused and no-longer-unused on
mismatch. Mutation-checked with a membership swap at constant count: it names both sides.
THIS IS A REPAIR INDEPENDENT OF THE NUMBER MOVING, and the two must not be conflated --
15 to 13 happened because two entries got cited, while a future drop because somebody
deleted a bibitem to clear a red is a different event that only the named form can tell
apart.

SECTION 2.6 GAINS TURICK 2011 AND CASADEVALL 2017 AS CONTENT, NOT AS TIDYING. Everything
else that section cites is a null. Turick is a positive measurement running against the
mechanism -- melanin "is continuously oxidized in the presence of gamma radiation" -- and
Casadevall, reviewing the hypothesis he originated, concedes "the mechanistic details
remain to be discovered", which is better support for "not demonstrated" than any null.
SCOPE HELD TIGHT: Turick verified from the abstract, cross-indexed on PubMed 21632287 and
two institutional repositories; the review's account of its apparatus comes from a paywalled
full text and is NOT repeated. Casadevall verified by extracting the ASM PDF and reading it,
not by trusting the review's quote.

THE THIRD REVIEW MOSTLY REDISCOVERED THE 2026-08-15 AUDIT, for the third time running. Its
"your own reference [36]" is refuted: turick2011 had no \cite in the body and was one of
fifteen orphans the Julia suite already pinned and printed every run. Its one real
contribution -- no fungus has a CO2-fixation pathway -- is genuinely absent here and is
recorded needs_verification, not adopted: it is fungal genomics, outside what this repo can
check, and the review self-flags three citations as unverified. Its energy budget
reproduces on two figures and its third recomputes 5x smaller (96 rather than 500 Gy/h);
the red-team states the discrepancy with working rather than dropping the number.

And the citation guard caught my own manuscript edit shifting two cited lines, diagnosing
both as THE CODE MOVED with corrected numbers. 695->715, 947->967, guide re-rendered.

Verified: calibration 401 passed; julia green with the pre-existing broken; both producers
exit 0; absence_gate GATE USABLE; 61 bibitems, balanced environments. latexmk is absent
locally so CI's manuscript-build remains the real check on the .tex, and three of these
changes are in it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…sioned reviews establish about commissioning

THE PAIR GOES IN AGENTS.md RULE 1 BECAUSE NEITHER IS VISIBLE BY READING THE ASSERTION.
Everything already in that rule asks whether a check CAN fire. These ask something else:
one is an assertion that cannot fail, the other is an input that never changed.

AND THE FIRST ONE I WROTE DOWN WRONG, THEN CAUGHT BY RUNNING IT. The looser form -- "weaken
the assertion to something trivially true and re-run; if the suite stays green it was not a
check" -- is false, because A GREEN SUITE STAYS GREEN UNDER ANY WEAKENING. The assertion was
passing anyway. Measured: weakening the bibliography set-assertion to `@test true` left 40
passing and established nothing. The discriminating form is to weaken the assertion and
re-run THE CONTROLS THAT ARE SUPPOSED TO FAIL. If those still fail, something else was
catching them. The correction is recorded in the file alongside the rule, because the
looser version is the one that sounds right.

The real instance stands: the same-run identity assertion in test_figure_staleness.py
cannot fail while _mutate_inside is correct, since that function locates by content and
replaces only inside the located slice. Kept, relabelled as documentation of an invariant.

The no-op half: a mutation check correct in every line proves nothing if the mutation did
not apply. A replace against normalised text missed a source storing XML entities, changed
nothing, and the guard's truthful "nothing missing" over an unmodified file read exactly
like confirmation. Assert that the mutation applied AND where -- a bare `mutated != source`
passes on a replacement that landed in a different element.

THREE COMMISSIONED REVIEWS, THREE REDISCOVERIES, AND THE FINDING IS ABOUT COMMISSIONING
RATHER THAN ABOUT ANY REVIEW. On three unrelated subjects, all three substantially returned
work this repository already held; the radiotrophy one rediscovered a single document,
radiotrophic_compatibility_audit.md of 2026-08-15, already carried into section 2.6 and the
ledger before the review existed. A BRIEF WRITTEN FROM A REPOSITORY THAT ALREADY HOLDS THE
ANSWER PRODUCES A REPORT THAT RETURNS THE ANSWER -- and two false claims in the first two
reviews were premises supplied in those briefs, returned as findings, which reads exactly
like independent confirmation.

Two genuine additions across all three: the overdamped physical argument, applied to section
3.4; and that no fungus has a CO2-fixation pathway, recorded needs_verification and NOT
adopted. BOTH ARE IN DOMAINS THIS REPOSITORY HAS NO PURCHASE ON -- low-Reynolds
hydrodynamics and fungal carbon metabolism. Neither is checkable from here; both are
checkable by someone who knows the field. That is the shape of what external review is good
for here, and the shape of what to ask for next: not "audit our claims", which returns the
audits, but the thing the repository structurally cannot do -- with the premise stated as a
premise to be checked rather than as background, and the answer expected to arrive as
needs_verification. A review whose findings this repo could have produced is a review that
was told what to find.

CI's manuscript-build completed SUCCESS on 464c1a9, the head carrying all three .tex changes
including both new citations -- so the bibliography renders rather than merely not erroring,
which is the failure mode a malformed \bibitem actually has. Treated as a gate, not a
formality.

Verified: calibration 401 passed; julia green with the pre-existing broken.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…m is a defect

RULE 1 COULD BE READ AS "A TAUTOLOGY IS A DEFECT". It is not, and the flattening had
already started: the previous commit's language treated tautological as a defect marker,
which would license deleting the most productive checks in this repository. The working
distinction is not tautological versus not. IT IS WHETHER A NECESSARY TRUTH CONSTRAINS --
forbids something -- OR MERELY RESTATES.

ONE. THE VACUOUS ASSERTION, AND THIS ONE IS A DEFECT FULL STOP. It stays green when
weakened to something trivially true, so it forbids nothing; it consumed a slot in the
suite and reported confidence it had not earned. Delete it, or if it documents an invariant
the surrounding construction actually provides, say so and stop calling it a control. That
is the same-run identity assertion relabelled two commits ago.

TWO. CONTINGENT CIRCULARITY, WHICH IS NOT A TAUTOLOGY AND IS NOT EMPTY. A model recovering
its own assumption is not true by logical form; it is a correct implementation of a
premise. The accurate statement is that THE MODEL IS NOT INDEPENDENT EVIDENCE FOR ITS OWN
ASSUMPTION -- non-informative about what it appears to be about, which is a different
failure from vacuous. The repository already identifies instances without having a name for
the class: README.md:413 records that the melanin ordering among the three producers "is
set by the input alpha_M", and docs/preprint_revision_plan.md:107 records a section 6.3
scatter that "is a scatter of Table 2's own entries". beta_ion has the same shape, spending
survival evidence on a tropism coefficient. None is a bug; each is a claim that must not be
reported as a result. THE TERM IS ADOPTED HERE, NOT QUOTED -- the substance is in the repo
in at least three places and the phrase is not, so it is introduced rather than cited.

THREE. THE LOAD-BEARING NECESSARY TRUTH, TAUTOLOGICAL AND INDISPENSABLE. Dimensional
analysis tells you nothing you did not put in, and it is how the D_eff collision was caught
and how the units error behind RADIODIALYSIS: BLOCKED was found. Conservation arguments are
tautological in the same sense, and so is the shell derivation in the calculus guide: two
face areas and a volume, no new physics, and the most useful page in the document. What
makes these productive is that THE CONSEQUENCES ARE NOT OBVIOUS FROM THE PREMISES EVEN
THOUGH THEY FOLLOW NECESSARILY. They forbid things -- a film-scale D_eff above D0, an
occupancy fraction in a g/cm3 slot, a slab operator with an r in it.

A CHECK THAT FORBIDS NOTHING AND A MODEL THAT PREDICTS ONLY ITS INPUTS FAIL THE SAME WAY,
which is why they are documented in one place rather than two.

AND THE SAME TEST APPLIES TO THE APPARATUS AS A WHOLE, which is the part worth having
written down. A repository of guards that collectively forbid nothing is the first failure
at a larger scale. Naming which claims are pinned and which are not is what stops it: the
coverage output listing undetectable ledger rows by name, the guide's paragraph saying
which of its prose claims no table pins, the SCOPE line on every absence claim. Those exist
so the apparatus cannot quietly become a restatement of its own inputs.

Rule 1 gains a cross-reference so "cannot fail" is not read as "tautological" before
somebody deletes a dimensional check.

Verified: calibration 401 passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
CAUSE 1 IS CLOSED, AND BUILDING THE GUARD SPLIT THE CAUSE IN TWO. The version recorded
three commits ago was (a): a verdict is enforced only in the file its `document` column
names, so RM-G04-01's five-term Hamiltonian survived in a .jl comment while the README was
corrected. (b) was not seen then and is worse: THE COLUMN CAN NAME NO FILE AT ALL.
`document = repository` is a PSEUDO_DOCUMENT, document_path returns None, and the row is
never searched. Measured: 37 rows sit in that class, 5 of them carrying an unresolved
verdict that names a source file.

PP-DEFF-01 IS ONE OF THE FIVE, AND ITS PRESCRIBED FIX WAS UNAPPLIED. biofilms_radiodialysis
.R:242 still carried a literature range and a Table 2 provenance the ledger had refuted --
the upper end two orders above the free diffusivity of any aqueous solute, the lower end
unresolved because neither solute nor temperature is declared, and Table 2 is tab:params
which has no D_eff row at all. THE FIFTH INSTANCE OF THE UNAPPLIED-VERDICT PATTERN, found
by building the guard for the first four. The comment now names which quantity the value is
and which three it is not.

AND THE CORRECTION COMMENT QUOTED THE WITHDRAWN STRING ON MY FIRST ATTEMPT, FOR THE THIRD
TIME THIS SESSION. Same trap as the withdrawn DOIs and the LaTeX comments: recording a
retraction by repeating it leaves the string in the artifact someone copies from, and here
it would have fired the guard being written in the same commit. The comment now points at
the ledger row, which holds the wording, and says why it does not repeat it. Three
occurrences means this is predictable rather than surprising and should have been avoided
from the start.

A DECLARED VOCABULARY, AND THE CHOICE WAS MEASURED BEFORE IT WAS MADE. Running
distinguishing_phrase over every `delete` row against all 144 source files returns ZERO
hits. The phrase tier cannot reach these claims because the ones that live in source
comments are short -- a formula, a range, a symbol -- and split into runs under MIN_WORDS,
exactly as figure labels and citations do. So this is the third vocabulary tier beside
RETRACTED_IN_FIGURES and RETRACTED_CITATIONS, for the reason each of those exists rather
than by analogy to them.

SEEDED WITH TWO ENTRIES BECAUSE TWO IS WHAT THE LEDGER SUPPORTS. A guard whose value is
that the NEXT one is caught does not need a long list, and a long list of invented strings
would be the can't-fail shape in a new costume. `restate` verdicts are NOT swept generally:
AGENTS.md is explicit that a restate asks whether the revision says the right thing, which
is a judgement that stays with a reviewer. Only the specific withdrawn string is mechanical.

Scope resolved the same way as RETRACTED_CITATIONS and for the same reason -- recording a
withdrawal requires naming it -- with explicit filenames rather than a glob, since a glob
auto-admits every future match and that is the property the allow-list exists to prevent.

Verified: control RUNS rather than skips, reaching the pre-fix file at 59b7d95 and finding
the withdrawn range there while the current file carries none. Planting the string in
biofilms_potts.jl is caught by name. Weakening the control's assertion with the defect
present still fails through the main check, so the main assertion is the load-bearing one --
the discriminating form of the tautology test, not the version that stays green on a green
suite.

And the citation guard caught my own R edit shifting a cited line for the third time,
diagnosing it as THE CODE MOVED with the corrected number. 380 -> 400, guide re-rendered.

calibration 403 passed (was 401, +2); julia green; both R producers exit 0 and
verify_biofilm_depth_profile reports ALL PASS over 15 checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…and flow declared as a relation rather than a list item

THE 37 PSEUDO_DOCUMENT ROWS ARE AN ADDRESSABILITY GAP, NOT A SCOPE GAP, AND THE COUNT NOW
PRINTS EVERY RUN. `document = repository` or `correspondence` means document_path returns
None and the row is searched by nothing -- the verdict has no file to be enforced in. The
claims-ledger coverage output now reports it beside the coverage it already prints: 37 rows
enforced nowhere, 13 of them carrying an unresolved verdict, 6 of those naming a source
file, each listed by id and status.

PRINTED RATHER THAN LEFT IN A PLAN, FOR THE REASON THE ABSENCE GATE PRINTS ITS FIVE KNOWN
GAPS. A number that lives in a document nobody rereads is a number nobody acts on.
PP-DEFF-01 sat in this class with a prescribed comment fix unapplied, and it surfaced only
because a guard was being built for an unrelated cause. That is an argument for the other
thirty-one getting a human pass rather than waiting for a guard that happens to sweep them,
and the output now says so in as many words.

CORRECTIONS NAME THE LEDGER ROW AND NEVER REPEAT THE WITHDRAWN CONTENT. Added to the
correction rules as a convention, after applying it twice by hand. Three occurrences in one
session -- a withdrawn DOI in a LaTeX comment, a second in the same file, a withdrawn range
in an R comment -- each inside a correction whose entire purpose was removing that string.

AND THE REASON IT RECURS IS STRUCTURAL RATHER THAN CARELESS, WHICH IS WHY IT IS A CONVENTION
AND NOT A TEST. Writing a correction and writing a guard against the thing corrected are the
same act, and THE CORRECTION IS AUTHORED IN THE WINDOW WHERE THE GUARD DOES NOT EXIST YET.
At the moment of writing there is nothing to catch it. Verified compliant: each of the three
withdrawn strings now appears in exactly one file, the guard's own declared vocabulary.

FLOW IS DECLARED IN SECTION 7.5, BUT AS A RELATION, AND THE FRAMING WAS CHECKED BEFORE
WRITING. A bare "no flow" beside the existing "viscoelastic mechanics are absent" would read
as a third independent item when it is closer to a consequence. Flow reaches morphology
through mechanics, and a model with no growth generates no growth-induced stress for
mechanics to act on -- so the three are ONE STRUCTURAL ABSENCE UNDER THREE NAMES. The
paragraph now says that and says what it costs: no route by which morphology could arise, so
the spatial output is directional redistribution of parcels in an imposed field rather than
a morphology.

THE LITERATURE HALF STAYS OUT OF THE MANUSCRIPT. External review argues growth-induced
stress is the dominant still-condition morphogen. That is unverified here, section 7.5 says
so explicitly and points at the red-team document, and nothing in the added sentence depends
on it. The claims carrying the weight are internal and checkable: no growth, no mechanics,
no flow.

Verified: calibration 403 passed; julia green. Structure check on the .tex holds at 61
bibitems with balanced environments, and CI's manuscript-build remains the real gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
… what that still does not check

THE FAILURE THIS CLOSES ACTUALLY HAPPENED, TWICE. docs/guides/<name>.md is the declared
source of record, <name>.tex is a hand-maintained rendering of it, <name>.pdf is built from
the .tex, and nothing bound the three. On 2026-08-31 the .md's prose was edited and the
other two kept the old wording while every check stayed green, because the citation binding
covers line numbers only -- which that test states about itself. The three files were
hand-synced twice in one session.

tools/render_guide.py NOW WRITES THE RECORD, AND ONLY AS A SIDE EFFECT OF A SUCCESSFUL
RENDER. It runs tectonic on the .tex, verifies every file:line citation in the .md appears
in the resulting PDF's text layer, and only then writes <name>.md.sha256. Same inversion as
tools/render_figure_svg.py: the sole way to make the staleness check green is to run the
renderer, which produces the PDF. A hand-written hash passes the hash check and leaves the
citation check to fail on the next real run.

WHAT IT DOES NOT DO IS STATED IN THE TOOL, THE TEST AND THE RECORD, BECAUSE A GREEN RUN
SHOULD NOT BE READ AS MORE. It does not verify that the .tex says what the .md says in
prose. A prose edit fails the hash check -- which is the point, it forces a human re-sync --
but nothing checks the re-sync was faithful. That remains a human obligation and an open
gap.

AND THE REASON IT CANNOT DO MORE WAS MEASURED, NOT ASSUMED. Matching the .md's prose
paragraphs against the rendered text layer gives 38 of 48 at a six-word prefix. The ten
failures are MARKUP differences rather than divergence: backticked `r` renders as math, a
literal `##` inside prose renders as emphasis. THE .tex IS A RE-AUTHORING, NOT A RENDER, so
the two differ by design and a prose binding would run at a fifth false failures. That is
the wrong instrument for a hand-authored .tex, not a threshold to tune -- and saying which
of those it is determines whether the next person tries to tune it.

Mutation-checked both ways, each restored: a prose edit touching no citation -- the exact
failure that occurred -- is caught by the record; and pointing a citation in the .tex at a
line the .md does not cite makes the renderer REFUSE TO RECORD rather than write a hash
over a mismatched pair.

calibration 404 passed (was 403).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
…ng true

A STALE CLAIM IN THE DOCUMENT WHOSE JOB IS RECORDING WHAT IS ENFORCED. The red-team's
reach statement read "RM-G04-01 remains unenforced pending RETRACTED_IN_SOURCES". That
guard shipped in 8d7cf05 the same day, with RM-G04-01's withdrawn string in its vocabulary,
and the sentence went on asserting the opposite. Corrected, and the correction says so
about itself -- it is a small instance of exactly the drift the paragraph was written to
prevent, and nothing would have caught it.

WHAT IS STILL OPEN IS NOW STATED THERE RATHER THAN INFERRED. The vocabulary tiers reach
declared strings and neither reaches the 37 pseudo-document rows enforced nowhere by
construction, 13 with unresolved verdicts and 6 naming a source. That count prints on every
run of the claims-ledger suite instead of living in the document, for the reason the absence
gate prints its known gaps. Those rows want a human pass rather than a guard that happens to
sweep them.

And the guide's prose divergence is recorded as partly closed with its remaining half named:
the source of record can no longer move without a render, and nothing verifies the re-sync
was faithful -- with the measurement that says why a prose binding is the wrong instrument
for a hand-authored .tex rather than a threshold somebody should try to tune.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2EuPqyhPoZoLE9SpxGorE
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant