Petty's Notebook
ArticlesPapersnfieldAbout
Get notified when new posts are published. No spam, just math.
Alexander S. Petty  |  ©2009-2026
← Back
Collision Capacity

The Bias Beneath the Secondary Term

August 9, 202610 min read
Companion paper: The Stationary-to-Diagonal Transition in Collision Capacity →
The blue rings suggest repeating remainder clocks. The gold boundary cuts through unfinished cycles, evoking the bias created when the observation grows with the resolution.
The blue rings suggest repeating remainder clocks. The gold boundary cuts through unfinished cycles, evoking the bias created when the observation grows with the resolution.

Long division generates repeating digits and remainder patterns. Counting agreements between digits gives a collision statistic. The squared size of its centered response is the collision energy studied in this sequence of papers.

The cubic law explains how collision energy grows. The secondary term explains its first systematic departure from that growth. Even after those terms have been accounted for, the values keep moving.

That remaining movement is worth understanding. A smooth formula can describe the overall shape of a table while leaving its finer arithmetic almost untouched. A slowly shifting average can also be mistaken for part of the fluctuation. We need to separate the trend from the variation before we can read either one correctly.

This paper makes that separation for collision capacity, the continuous energy of the remainder waves underlying the finite digit-collision table. It finds the next two terms in the smooth profile and proves a law for the values left around it.

There is a complication. Increasing the resolution changes the collection of periodic waves being measured. Their cycles are balanced, but each observation includes cycles that have only just begun. The measurement acquires a bias from its own moving boundary.

Once that bias is removed, a stable arithmetic distribution emerges. Reading only at the resolutions associated with prime bases gives a different distribution. Freezing the boundary at a certain scale exposes another contribution with the shape of a simple parabola.

The purpose is to tell these effects apart. They would otherwise all sit inside the same remainder.

The unfinished cycles

The capacity formula rewrites the remainder waves as clocks. A period-five clock reads the remainder after division by five. Its five positions repeat forever.

Subtract the midpoint, and its readings become minus two, minus one, zero, one and two. Together they add to zero.

The finite calculation first includes this clock at resolution five, where it reads minus two. At resolution seven, its three admitted readings add to minus three. Balance arrives at nine, when the full cycle has been read.

By then, more clocks have entered. Each begins at the low end of its cycle.

Four rows of remainder clocks enter at resolutions five, seven, nine and eleven. Gold marks negative readings and blue marks positive readings. The period-five prefix reaches minus three before returning to zero.
Hollow positions precede a clock’s admission. Larger dots indicate readings farther from its midpoint. Each full cycle balances, while an unfinished cycle can leave a negative total.

This is an exact feature of the capacity formula. Each period carries a positive weight fixed by the collision construction, and the weights add to one. Complete cycles cancel in the cumulative centered reading. Every unfinished cycle leaves a nonpositive contribution.

The construction therefore explains the direction of the drift before the asymptotic calculation determines its size. A collection of individually balanced periods can produce a biased sequence when new periods continually enter it.

Separating the drift

The stationary model describes a fixed collection of clocks observed through all its compatible phases. The finite capacity calculation keeps enlarging that collection. At resolution nnn, it reads all periods through nnn at the address nnn itself. This simultaneous choice is the diagonal in the paper’s title.

Call the resulting clock reading yny_nyn​. The theorem determines a negative logarithmic drift and a constant correction, both from the exact collision weights. Subtracting them leaves the centered reading ZnZ_nZn​.

The logarithm grows slowly, so the drift is easy to overlook over a short range. Across the five windows below, the uncorrected average moves steadily downward. The corrected average stays near zero.

Five complete integer windows show the raw diagonal mean drifting downward. After logarithmic centering the means stay near zero, with a comparable fluctuation scale. Shading displays the observed centered standard deviation.
These are finite nfield calculations through two million. The lines show observed means. The shaded width is the centered standard deviation, translated to each mean for comparison. The limit is established by the proof.

Now choose an integer uniformly between X+1X+1X+1 and 2X2X2X. As this interval moves outward, the distribution of ZnZ_nZn​ approaches a fixed law. Its average square approaches a fixed positive value as well.

That is a statement about the whole spread of readings. Individual values keep changing. The limiting law describes how frequently different values occur, and its variance measures their average squared size.

The law is the one predicted by the stationary collision potential. The work of the proof is to justify using that model when the collection of clocks changes at every step. Holding a finite collection fixed would have been a different observation.

Reading the capacity remainder

Return to the capacity itself. Write J(n)J(n)J(n) for its continuous energy and Dn=n−J(n)D_n=n-J(n)Dn​=n−J(n) for its deficit below the leading linear profile. The secondary-term calculation identified the leading deficit π−2log⁡2n\pi^{-2}\log^2 nπ−2log2n.

The present result resolves the next layer,

Dn=log⁡2nπ2+Blog⁡n+C+Rn,D_n=\frac{\log^2 n}{\pi^2}+B\log n+C+\mathcal R_n,Dn​=π2log2n​+Blogn+C+Rn​,

where B≈0.604898B\approx0.604898B≈0.604898 and C≈0.377541C\approx0.377541C≈0.377541.

The distinction matters when interpreting a graph. Removing only the quadratic-logarithmic term still leaves a systematic contribution Blog⁡n+CB\log n+CBlogn+C. At a resolution of one million, that contribution is about 8.78.78.7. A graph of the partially centered remainder would continue to drift.

After the full centering, Rn\mathcal R_nRn​ has the same limiting distribution and second moment as the stationary potential. The proof identifies how much of the movement belongs to the smooth profile before describing the fluctuation around it.

A stable distribution still permits rare large readings. The paper constructs increasingly large residuals of both signs. Some remain large even after the contributions from periods dividing either neighboring endpoint are removed. The largest possible fluctuation at every resolution remains an open question.

This separates two tasks. A distribution describes the frequency of values. A pointwise bound must also account for the most exceptional ones.

Prime bases change the sample

The original collision energy is often studied in a prime base ppp. Its associated continuous capacity is evaluated at p−1p-1p−1. Selecting those resolutions also selects particular clock positions.

Modulo six, every prime greater than three is one or five. Its predecessor is therefore zero or four. Those two positions average to two. All six positions average to two and a half.

Two six-position clocks compare all integer phases with prime predecessors at zero and four. Their means are 2.5 and 2. The half-position shift, weighted by conserved mass and the readout factor, produces minus one third.
Prime predecessors sit half a position below the full-cycle mean when averaged over the unit residues of a fixed period. Summing with the collision weights gives the mean shift minus one third.

This half-position shift occurs at every fixed period when prime predecessors are averaged over the admissible residue classes. The clock readout has a factor of two thirds, and the period weights add to one. The combined mean shift is therefore exactly minus one third.

Passing from this finite residue calculation to growing primes requires a separate proof using prime-distribution estimates. The resulting residual law has mean −1/3-1/3−1/3. It is symmetric about that mean and has a positive finite variance, at least 1/3241/3241/324.

Thus the choice of prime bases remains visible after the smooth trend has been removed. Their residuals have their own distribution. Correcting the mean alone does not identify that distribution with the all-integer law.

This theorem concerns continuous capacity at p−1p-1p−1. Transferring this finer law to the original finite digit-grid energy still requires control of the grid-sampling defect. That distinction preserves the exact scope of the result.

The shape of the observation boundary

There is a third observation to test. Choose a period cutoff MMM, keep those clocks fixed while reading the integers from X+1X+1X+1 through 2X2X2X, then enlarge the cutoff for the next interval.

A slowly growing cutoff recovers the stationary law. Increasing it too quickly leaves enough unfinished boundary to change the answer. The paper locates the transition through the ratio Mlog⁡M/XM\log M/XMlogM/X.

Within the regime where MMM tends to infinity and M=o(X)M=o(X)M=o(X), the stationary law and second moment persist precisely when that ratio tends to zero. At MMM of order X/log⁡XX/\log XX/logX, the boundary contributes a fluctuation of its own.

Its shape is the second Bernoulli polynomial,

B2(u)=u2−u+16,B_2(u)=u^2-u+\frac16,B2​(u)=u2−u+61​,

where uuu records position within the cutoff cycle.

Blue finite cutoff discrepancies follow the gold Bernoulli parabola across the phase interval from zero to one. The comparison uses 401 resolutions between one and two million, with cutoff 72382 and no fitted parameters.
The blue points show the scaled difference between the centered diagonal and a frozen cutoff. The gold curve is the boundary term proved in the paper. This finite comparison uses the explicit constants without fitting.

The gold curve is predicted by the proof. The blue points come from finite capacity calculations with the stated constants and no fitting. Once the known amplitude has been removed, the discrepancy follows the boundary curve.

If Mlog⁡M/XM\log M/XMlogM/X tends to a positive value λ\lambdaλ, the boundary adds exactly λ2/(90π4)\lambda^2/(90\pi^4)λ2/(90π4) to the limiting second moment. The proof also establishes that this boundary term and the stationary fluctuation become independent in the limit.

This gives the observation error a specific origin, shape and size. It also explains why a fixed-cutoff calculation can have the wrong limiting variance even when its cutoff is a vanishing fraction of the observation scale.

Interpreting the remaining variation

The starting problem was the movement left after the main and secondary terms. It now has a more precise description.

A further smooth correction removes the drift. The centered integer readings approach a stable arithmetic law. Prime predecessors select a different law, and a cutoff near the critical scale adds a calculable boundary fluctuation.

These distinctions make the capacity model usable. A changing average need not be mistaken for irregularity, and a boundary contribution need not be folded into the arithmetic fluctuation we wanted to study.

The remainder has become something we can interpret. Its drift, its typical variation and its dependence on the sampling rule have separate mathematical explanations.

Read the paper.

Companion paper: The Stationary-to-Diagonal Transition in Collision Capacity →
Share

Discussion

Sign in to join the discussion.

← All articlesRead the paper →
← Previous: The Clocks Beneath Collision Energy
Next: Repetend Rigidity at the Golden Scale →