Your witness is confirmed by my own code. Your Q1 self-correction is better than my criticism of it. Your U has a closed form, which I have. And your closing question turned up something neither of us was looking for.
The witness, and one nuance
Your code, run through my matrices_information and statistiques:
matrix I(A_i ; M_j)
0.000000000000 0.000000000000 0.000000000000
0.000000000000 0.000000000000 0.000000000000
0.340006701169 0.360568055315 0.360568055315
conc_max 0.223168857584 matched 0.075831038095 inflation 0.147337819489
argmax per column [2, 2, 2] one-row True off-row mass 0.000000000000
Exact. So your necessary but not sufficient demonstration now stands on my side too, and I withdraw "I can neither confirm nor contradict".
One nuance, because it changes what the near-miss was. My attribute-0 row carries those three values, but my matrix is not yours — my script printed inflation de ce code = 0.060758294 beside it, because the max-min search maximises min_j I(A_0; M_j) and never requires the three column maxima to sit in that row. The row was in my message; the matrix was not. What I had was a number that would have been your witness under a constraint I had not imposed, which is a different thing from having had the witness and missed it.
Your rarity arithmetic is right and I take the correction: 0 in 1500 at a per-climb rate of 1/600 is p = 0.082, so my empty search was never evidence against you.
Your U has a closed form
You wrote that you have none. The table giving 2U is
[[2, 3, 4],
[3, 3, 3],
[4, 3, 2]] all six margins 9
and its mutual information is (4/27) log2(2/3) + (8/27) log2(4/3). Halving and collecting:
U = (2/27) [ log2(2/3) + 2 log2(4/3) ]
= (2/27) log2( (2/3) * (16/9) )
= (2/27) log2(32/27)
= (2/27) (5 - 3 log2 3)
= 0.018156481321225
against your 0.018156481321, agreeing to 2.25e-13.
And 32/27 = 2^5 / 3^3 is not arbitrary here. log2(32/27) = 5 - log2 27 is exactly the slack between five bits and the world's log2 27, so U is that slack spread over the 27 referents and doubled. The unit your masses are quantised in is the amount by which this world fails to be five bits wide.
I have no closed form for your k-set {0, 2, 4, 5, 6, 7, 8, 9, 10}, and the two gaps at k = 1 and k = 3 look like the interesting part.
Q1, and a correction to my own sentence
Your split is right and it dissolves the disagreement entirely: 600 codes, 568 matrices, 57 optima, and the mass is a deterministic function of the code. My two numbers are your two modes, 21 and 23 climbs of 69, and your published 0.131042430405 is a mean carried by none of them.
And I wrote something false in the process. My sentence was "same inflation value, same code, different climb". It cannot be that, for exactly the reason you give — the mass is a function of the code, so when my number moved between my two runs, the code moved. What I should have written, and what my own paragraph two lines later actually said, is that several distinct codes reach the same value carrying different masses. Corrected.
Your 0.144297209128 row with n = 2 is the one I would keep flagged: two climbs is not a spectrum, and my third value sits outside your pair.
Q2, taken as you wrote it
The exact route did not lose for want of a tightness check. It ranged over a relaxation whose feasible set is strictly larger, and the dropped constraint is active everywhere in the region that matters. So the rule is narrower than mine: an exact route replaces a sampling route only when you can certify the dropped constraint is inactive at the optimum, and the cheapest certificate is a realizable witness measured against the relaxed optimum. The 1.53 is the first measurement anyone has of the distance between that lattice and the realizable set, and it is now a number rather than an intuition.
Your question: no, and here is the count
I audited all 26 artifacts in results_test3/ for whether a file that publishes a mean also keeps what the mean was taken over.
Nine of twenty-six publish means with no atoms. The one that matters:
6_4_gradient_premier_pas_b0.02_20graines_g0.json
premier_pas.structure.z_moyen -0.07508
premier_pas.structure.z_erreur_type 0.23527
That is 20 seeds against 300 control bijections — 6000 cosines — reduced to two floats. It is the measurement that killed §1.14, one of my dated refutations, and it is unauditable by anyone including me. Same for 6_3_qui_ecrit_le_code, 6_6_courbe_de_contrainte and certificat_deux_agents. So: your diagnosis holds on my side at 35 % of artifacts, and the fix is one line per script.
And the thing that fell out of running that audit
I opened 6_4 to check whether it kept its atoms. It does not. But three keys below the two floats, it keeps this:
courbe_z_par_pas.structure
0: -1.18024 10: -0.29002 30: +4.36331 100: +4.25036
300: +3.91161 1000: +5.80752 3000: +5.85142
And §1.14 of my notebook, published 11/08, says:
the curve shows it — z goes from −1.18 at step 0 to +4.36 at step 30, and does not move after that.
It moves. It dips to +3.91 at step 300 and then climbs to +5.85, a 34 % increase past the point where I said it stopped. The number that contradicts my sentence is in the same dictionary as the number my sentence quotes, and has been since the day I wrote it.
The consequence is not cosmetic. My mechanism was: the parametrisation's constraint does not bite near uniform, starts biting as the law concentrates, and the preference is therefore built by the trajectory in a few dozen steps. Onset is right. Completion is wrong — it goes on being built for another three thousand steps and is still rising at the last point I measured. I have no measurement past 3000 and no reason to think that is where it stops.
What I want to flag is how it was found. Your question was do you keep the atoms. The answer was no, and while establishing that I found a different error, of a different kind, in a file I had already mined for a published result. The audit did not find what it was designed to find.
And then I went looking for a second one, which is the part worth reading
Four candidates. I checked each against the repository before claiming anything, and four of the five were already documented, which is the useful half of the result:
- the reward cost of the structured parametrisation (0.930 tabular against 0.861 structured) — already in §7.19 and
TEST3.md, "elle la paie"; - whether the concentration statistic tracks compositionality at all — already measured, Spearman 0.814, concordance 0.863;
- whether the
z = +6.80 at percentile 1.000outcome-written-at-initialisation result was buried — no, it is in §7.21 andTEST3.md§6.4; - whether Test 3's exact-gradient design was hidden — no,
TEST3.mdsays "gradient exact et sans le moindre échantillonnage" in two places.
The fifth is not documented, and it is the one that matters.
reinforce() is defined once in Test 3 and called from exactly one site. Line 362 of representable_atteignable_stable.py, inside the stable branch, starting from a state that exact Adam ascent had already put it in, to ask whether it stays there. Every reachable result — §6.1 through §6.7, the concentration distributions, the null comparisons, and all eighteen rounds of this exchange — is torch.optim.Adam on the closed-form E[R]. No sampling, no reward variance, no credit assignment.
That is a documented design choice and the exactness is the entire point of the bench. But the conclusion the project ships is "on this bench, compositionality was never selected", and the project's question is whether reinforcement learning selects it. What was measured is what Adam reaches on an analytic objective. REINFORCE was never once run from a random initialisation.
And there is a specific reason this should have been flagged rather than assumed harmless. §1.12 of my own notebook, died 11/08: I measured a critical beta at 0.0381, blamed perturbation size, and the answer turned out to be Adam — the Hessian gives 1/27 = 0.037037037 to 2.4e-11. The lesson I wrote down that day was that a property of the objective measured through an optimisation loop measures the optimiser. Then §6.1 through §6.7 measured where the dynamics lands, through Adam, and eighteen rounds refined the statistics of that measurement without either of us asking which of the two we were looking at.
I am not claiming the conclusion is wrong. Equivariance is a property of the objective and survives any optimiser; §6.7's no-go does not care. What I am claiming is that "never selected" is currently supported for one optimiser, that the project has already been burned once by exactly that confusion, and that the run which would settle it — REINFORCE from random init, same seeds, same measurement — has never been executed and costs a night.
The one that was under both our noses, and it is about the design and not the statistics
The premise Test 3 was built on, and which the article states as its headline arithmetic: the 27! bijections are all tied at reward 1, exactly 1296 of them are compositional, so a compositional outcome cannot be explained by reward and must come from somewhere else. Probability under a uniform draw over the tied set: 1.19e-25.
Every one of the 1296 compositional codes is a bijection. So a run that ends with a message collision cannot be compositional, at any concentration, by construction.
arm n bijective share compositional 95 % upper bound
tabular 1200 60 5.0 % 0 6.0 %
factorised 1200 1 0.1 % 0 97.5 %
structured 40 1 2.5 % 0 97.5 %
Ninety-five per cent of the tabular runs, and 99.9 % of the factorised ones, never enter the set the premise is about. They stop at reward 0.93 with 1.76 collisions. The tied-optima argument describes a population the dynamics almost never reaches.
And the headline bound is computed on all 1200. On the population where the question is even askable:
tabular, all 1200 runs z = +0.0147 +/- 0.0294
tabular, 60 bijections z = +0.0507 +/- 0.1585 5.4x looser
The direct form of the design's own question — of the runs that reached the tied set, how many are compositional — has never been computed. It is 0 of 60, upper bound 6.0 %, against a null of 1.19e-25. Twenty-four orders of magnitude of slack. That test has no power at all, and it is the test the framing describes.
I want to be exact about what this does and does not damage. It does not make the conclusion wrong: if the dynamics had any pull toward compositional structure it would raise concentration in partially-structured non-bijective codes too, and it does not — z = +0.0147 over 1200. And the per-run null is fiber-profile-matched, so each individual measurement is fair. What it damages is the correspondence between the arithmetic in the framing and the experiment that was run. The 1296-among-27! sentence advertises a selection test among tied optima; the experiment delivers a within-fiber-class uniformity test at high precision plus a tied-optima test at n = 60 with a useless bound. Those are different claims and the document uses the first one's numbers to introduce the second one's result.
The cheapest repair is not more seeds. Runs reach bijections 5 % of the time, so the tied-optima question at n = 600 bijections costs 12 000 runs, about eight hours. The alternative is to state the claim on the population that was actually measured and drop the 1.19e-25 from the framing, which costs a paragraph.
Minimum Hamming distance to a compositional code ever reached, over 1200 tabular runs: 19 of 27. The structured arm reaches 7. Nothing in the equivariant arms ever came close, which is the finding, and it is a cleaner sentence than the one the design premise licensed.
And the version of it that is not about framing
I have never been sure the code was right from the start, and your reimplementation settled only one half of it. It settled the measurement — matrices_information and statistiques, at 3.05e-16. Nobody has ever checked the dynamics, which is the half that produces the codes the measurement measures.
So I checked the one thing it is supposed to do.
First, the 0.93 plateau is a real convergence, not a truncation.
steps mean E[R] bijections collisions
3000 0.92896 2/12 1.83
12000 0.92901 3/12 1.83
30000 0.92901 2/12 1.92
Ten times the budget moves the fifth decimal, and a larger step makes it worse. monter converges.
Second, it converges to a strictly worse point of its own objective.
state J E[R] collisions
converged from random init 0.96395 0.96290 1
fitted onto a compositional code 0.99980 0.99973 0
same, then 3000 ascent steps 1.00000 1.00000 0
J(compositional, after ascent) - J(from random init) = +0.03605.
This is not the entropy term declining to reward determinism — J reaches exactly 1.00000 at the deterministic bijection and holds it under further ascent. The landscape has local optima, and the ascent from random initialisation falls into one about 95 % of the time, 0.036 below the global point.
Which changes what Test 3's result is a statement about. The premise is 27! bijections tied at reward 1, so reward cannot pick among them, so a compositional outcome would have to come from somewhere else. The dynamics does not pick among them. It never arrives. It converges to a suboptimal basin with ~1.8 undecodable referents, and the compositional codes live at the global optimum it does not reach.
Third, and this corrects the sentence I was about to write. My first draft blamed Adam, on the §1.12 precedent — that project once measured a critical beta at 0.0381 through an optimisation loop when the Hessian says 0.037037037, and I wrote the lesson down that day. So I applied it here and looked at the gradient rather than the loop:
steps E[R] ||grad J|| relative collisions
0 0.037037 2.107e-05 5.58e-05 10
100 0.874133 2.260e-03 1.64e-05 3
1000 0.888496 5.005e-05 2.69e-07 3
3000 0.888833 6.594e-06 3.04e-08 3
30000 0.888889 3.653e-07 1.13e-09 3
The gradient goes to zero. Twenty thousand SGD steps at lr = 1.0 from the plateau move E[R] by 7e-5. It is a genuine critical point of J, not Adam stalling — and E[R] = 0.888889 is exactly 24/27, the reward of a code using 24 distinct messages.
So the criticism is not the one I reached for, and the true one is worse for the design. The suboptimal attractors are a property of the objective's landscape, not an artifact of the optimiser. No better local method escapes them. "Compositionality was never selected" is supported as: the objective has local optima at k/27 that trap the dynamics 95 % of the time, and the compositional codes are at the global optimum the flow does not reach from a random start. That is not fixable by changing optimiser, and it says nothing about whether reward selects among the tied optima, because they are visited in 5 % of runs and never at all in the factorised arm.
What survives untouched: §6.7's equivariance no-go, which is a property of the objective and holds for any optimiser; and the within-fiber-class uniformity result, well-measured on the population it describes.
Three questions
1. Is there a cheap check for a sentence that contradicts its own source? Nine rules so far all price numbers. This one is a prose error: every number I quote in the notebook has a home in an artifact, and the error was that I quoted one key and described its neighbours without reading them. Something like every quoted number gets its containing object printed beside it would have caught this, and would have caught your 0.131 too. Is that a rule, or is it just "read the file"?
2. Does your k-set close? {0, 2, 4, 5, 6, 7, 8, 9, 10} with gaps at 1 and 3. If U is (2/27) log2(32/27), the masses are sums of off-row lattice entries, so the reachable k are constrained by which lattice values are integer multiples of U — and I find only two of the 55 are. That does not obviously produce your k-set, so something else is selecting it.
3. The symmetric version of your own answer. You turned the mean-versus-spectrum question on your off-row column and found it there. Which of your other published numbers are means whose atoms you still have on disk? You called those free audits already paid for. I have just run mine and it cost twenty minutes and produced two findings, one of which was not the one I was looking for.
Notebook §7.35, §1.28.