Fix a base b\geq2 and a prime p>b+1. The nonzero residues modulo p are partitioned by the leading base-b digit of r/p. At lag one, let C_{p,b}(b) count the residues that remain in the same digit bin after multiplication by b, and subtract the exact mean collision count over the constructive nonidentity multipliers. The resulting fluctuation \Delta_b(p) is bounded but not centered across the primes.
For p>b^2, the integer deviation S_b(p) in the decomposition of \Delta_b(p) is determined by p\bmod b^2. Reflection of that finite table forces its mean to be -1/2. The rational correction from the constructive mean contributes 1/2-1/b. Mertens’ theorem in arithmetic progressions therefore gives the exact asymptotic \sum_{b+1<p\leq x}\frac{\Delta_b(p)}p =-\frac{b-1}{b}\log\log x+\kappa_b+o(1). In base ten the coefficient is -9/10. Exact finite tables reproduce the class means and show how the additive constant remains visible through the primes below 10^7.
Every prime greater than ten ends in 1, 3, 7, or 9. In the lag-one collision count, none of these four streams is centered. Each carries a negative bias. Reciprocal-prime weighting accumulates those finite biases with the exact leading term -\frac9{10}\log\log x.
A bounded local fluctuation therefore produces an unbounded prime sum. The coefficient is not fitted from prime data. It is forced by two finite means. Reflection contributes -1/2, and the constructive-mean correction contributes another -(1/2-1/b). In every base b, their sum is -(b-1)/b.
The first 1{,}000 base-ten terms total -1.367764. After 664{,}574 terms the sum is -1.895354. The motion is slow because \log\log x is slow, but the source of the drift is already present in a finite table modulo b^2.
The construction below uses the complete nonzero residue system. It does not require the base to generate the full multiplicative group. When the repetend of 1/p occupies only one of several multiplicative orbits, the count includes the other orbits as well.
When the base is primitive modulo p, the lag-one count is coordinatewise agreement between a reciprocal digit word and its one-place cyclic shift. This is Hamming correlation. Lempel and Greenberger place the observable in the general theory of periodic sequence correlation [3]. Kak and Chatterjee study Hamming distance between prime-reciprocal digit sequences and their cyclic shifts [4]. A separate arithmetic line due to Girstmair expresses the variance of digit values along a full reciprocal period through a Dedekind sum [5].
The statistic here is an equality count on the complete nonzero residue system, varied with the prime. Its collision-specific input is the explicit table modulo b^2 and the reflection that fixes the table mean. Once those finite facts are known, the passage to a prime harmonic asymptotic is a direct application of classical Mertens theory in arithmetic progressions.
Let b\geq2 and let p>b be prime. For a nonzero residue, write [a]_p for its least positive representative. Define the digit map \delta_{p,b}(r)=\left\lfloor\frac{br}{p}\right\rfloor, \qquad 1\leq r<p, and the collision count C_{p,b}(g) =\#\left\{r\in\{1,\ldots,p-1\}\ \middle|\ \delta_{p,b}(r)=\delta_{p,b}([gr]_p)\right\}. A nonidentity multiplier is constructive when its collision count is positive.
Write p-1=bQ+R, \qquad 0\leq R<b. The first structural fact identifies the zero-collision multipliers and also explains why every collision count is even. The zero gate and constructive mean are established in Bin Derangements and the Gate Width Theorem [1] and used in Silent Primes and the Variance of the Collision Count [2]. Both are derived here in the form needed for the prime sum.
Lemma 1 (The zero gate). For a nonzero multiplier g\neq1, put c(g)=\bigl[b(1-g)^{-1}\bigr]_p. Then C_{p,b}(g) =2\#\left\{m\in\{1,\ldots,Q\}\ \middle|\ [mc(g)]_p>mb\right\}. The collision count vanishes exactly when c(g)\in\{1,\ldots,b-1\}. Hence there are b-1 zero-collision multipliers, and C_{p,b}(g) is always even.
Proof. Euclidean division gives br=p\delta_{p,b}(r)+[br]_p. Since p is invertible modulo b, two residues have the same digit exactly when their images under multiplication by b agree modulo b. After permuting the residues by that multiplication, a collision is a solution of [gx]_p=x+mb with a nonzero integer m satisfying |m|\leq Q.
The identity c(g)(1-g)\equiv b\pmod p turns the last relation into x+mc(g)\equiv0\pmod p. For m>0, the solution lies in the allowed range exactly when [mc(g)]_p>mb. The index -m supplies a second solution under the same condition. This proves (4).
The value c(g)=b cannot occur. If c(g)<b, no index contributes. If c(g)>b, the index m=1 contributes. The map g\mapsto c(g) is a bijection from the nonidentity multipliers onto \{1,\ldots,p-1\}\setminus\{b\}, which proves the remaining claims. ◻
Assume from now on that p>b+1. There are N=p-b-1 constructive nonidentity multipliers. Their exact mean comes directly from the digit-bin populations.
Proposition 2 (Constructive mean). The mean collision count over the constructive nonidentity multipliers is \overline C_{p,b} =\frac{Q\bigl(b(Q-1)+2R\bigr)}{p-b-1} =Q+\frac{QR}{p-b-1}.
Proof. Let n_d be the size of digit bin d. Exactly R bins have size Q+1, and the other b-R bins have size Q. For fixed r, the values [gr]_p run through every residue other than r as g runs through the nonidentity multipliers. Reversing the order of summation therefore gives \sum_{g\neq1}C_{p,b}(g) =\sum_{d=0}^{b-1}n_d(n_d-1) =Q\bigl(b(Q-1)+2R\bigr). The zero-collision multipliers contribute nothing. Lemma 1 shows that the remaining denominator is p-b-1. The second expression in (5) follows from p-1=bQ+R. ◻
Definition 3. The lag-one integer deviation and collision fluctuation are \begin{aligned} S_b(p)&=C_{p,b}(b)-Q,\\ \Delta_b(p)&=C_{p,b}(b)-\overline C_{p,b}. \end{aligned}
Proposition 2 gives the exact decomposition \Delta_b(p)=S_b(p)-\frac{QR}{p-b-1}. The correction has the exact expansion \frac{QR}{p-b-1} =\frac Rb+\frac{R(b-R)}{b(p-b-1)}. The first term depends only on p\bmod b. It is integral exactly when R=0. The second is the finite-size remainder and tends to zero as p grows with b fixed. Both terms vanish when the digit bins have equal size. Lemma 1 also shows that S_b(p) has the same parity as Q.
Set M=b^2 and let G_b=\{(b+1)d\mid 0\leq d<b\}. These are exactly the two-digit base-b words whose first and last digits agree.
Lemma 4 (Finite determination). Let p>M be prime and write p=Mt+a with 1\leq a<M. Then a is a unit modulo M, and S_b(p)=F_b(a), where F_b(a)= -1-\left\lfloor\frac ab\right\rfloor +\sum_{n\in G_b} \left( \left\lfloor\frac{(n+1)a}{M}\right\rfloor -\left\lfloor\frac{na}{M}\right\rfloor \right). Thus the integer deviation depends only on p\bmod b^2.
Proof. For 1\leq r<p, put n=\lfloor Mr/p\rfloor. The first digit of n is \lfloor n/b\rfloor, and \delta_{p,b}(r)=\left\lfloor\frac nb\right\rfloor. Write br=p\delta_{p,b}(r)+s with s=[br]_p. Then \delta_{p,b}(s) =\left\lfloor\frac{bs}{p}\right\rfloor =n-b\left\lfloor\frac nb\right\rfloor =n\bmod b. A collision occurs exactly when these two digits agree, which is equivalent to n\in G_b.
The number of positive residues in slice n is \left\lfloor\frac{(n+1)p}{M}\right\rfloor -\left\lfloor\frac{np}{M}\right\rfloor, for every interior slice. The relation \gcd(p,M)=1 makes its two endpoints nonintegral. The initial slice starts at zero and already counts only the positive residues. The terminal floor difference includes the endpoint p, which is not a residue in the count. The terminal index M-1 belongs to G_b, so C_{p,b}(b) =-1+\sum_{n\in G_b} \left( \left\lfloor\frac{(n+1)p}{M}\right\rfloor -\left\lfloor\frac{np}{M}\right\rfloor \right).
Substituting p=Mt+a contributes t from each of the b diagonal slices. Since a is a unit modulo b, Q=bt+\left\lfloor\frac ab\right\rfloor. Subtracting Q from (13) gives (12). ◻
Corollary 5. For fixed b, both S_b(p) and \Delta_b(p) are uniformly bounded as p ranges over the primes greater than b+1. Consequently, \sum_{p>b+1}\frac{\Delta_b(p)}{p^s} converges absolutely when \operatorname{Re}(s)>1.
Proof. Apart from finitely many primes, S_b(p) is drawn from the finite table in (12). The expansion (9) shows that the correction in (8) is bounded for fixed b. Comparison with the prime subseries of \sum n^{-\operatorname{Re}(s)} proves the convergence claim. ◻
The table has a symmetry that fixes its mean without evaluating its entries one by one.
Lemma 6 (Reflection). For every unit a modulo M=b^2, F_b(a)+F_b(M-a)=-1.
Proof. Write D_n(a)= \left\lfloor\frac{(n+1)a}{M}\right\rfloor -\left\lfloor\frac{na}{M}\right\rfloor. For n\in G_b with 0<n<M-1, neither endpoint is integral and D_n(a)+D_n(M-a)=1. At the two endpoint indices, the combined contributions are 0 for n=0 and 2 for n=M-1. Since G_b has b elements, \sum_{n\in G_b}\bigl(D_n(a)+D_n(M-a)\bigr)=b. Write a=bq+u with 1\leq u<b. Then \left\lfloor\frac ab\right\rfloor +\left\lfloor\frac{M-a}{b}\right\rfloor=b-1. Substitution of (15) and (16) into (12) proves (14). ◻
Corollary 7 (Grand mean). The exact mean of the finite table over the units modulo b^2 is \frac{1}{\varphi(b^2)} \sum_{a\in(\mathbb Z/b^2\mathbb Z)^\times}F_b(a)=-\frac12.
Proof. Negation pairs the units modulo b^2 without a fixed point. Every pair has sum -1 by Lemma 6. ◻
The rational term in (8) has a second exact mean.
Lemma 8 (Remainder mean). If u runs over the least positive representatives of the reduced residue classes modulo b and R=u-1, then \frac{1}{\varphi(b)} \sum_{u\in(\mathbb Z/b\mathbb Z)^\times}\frac{u-1}{b} =\frac12-\frac1b.
Proof. The reduced residues pair as u and b-u, so their mean is b/2. The same identity holds for b=2. Subtracting one and dividing by b gives (18). ◻
For x>b+1 and s\in\mathbb C, define \Phi_b(x;s)=\sum_{b+1<p\leq x}\frac{\Delta_b(p)}{p^s}.
Theorem 9 (Collision fluctuation law). For every fixed integer base b\geq2, there is a real constant \kappa_b such that \Phi_b(x;1) =-\frac{b-1}{b}\log\log x+\kappa_b+o(1) as x\to\infty. In particular, the lag-one collision fluctuation sum diverges to -\infty at the prime harmonic boundary.
Proof. Mertens’ theorem in arithmetic progressions gives, for every reduced class a modulo a fixed modulus q, \sum_{\substack{p\leq x\\p\equiv a\, (\mathrm{mod}\ q)}}\frac1p =\frac{1}{\varphi(q)}\log\log x+M(q,a)+o(1), where M(q,a) is a class-dependent constant [6, 8].
Apply (21) with q=b^2. Lemma 4 and Corollary 7 give \sum_{b+1<p\leq x}\frac{S_b(p)}p =-\frac12\log\log x+\kappa_b^{(S)}+o(1). The finitely many primes at most b^2 change only the constant.
The correction in (8) has the expansion (9). After division by p, its second term is O_b(p^{-2}) and therefore contributes an absolutely convergent constant. Reduction from units modulo b^2 to units modulo b has equal fibers. Equations (21) and (18) now give \sum_{b+1<p\leq x}\frac1p\frac{QR}{p-b-1} =\left(\frac12-\frac1b\right)\log\log x +\kappa_b^{(R)}+o(1). Subtracting (24) from (23) proves (20). ◻
Corollary 10 (Removal of the drift). The limit \lim_{x\to\infty} \sum_{b+1<p\leq x} \frac{\Delta_b(p)+(b-1)/b}{p} exists.
Proof. The classical relation \sum_{p\leq x}\frac1p=\log\log x+B_1+o(1) multiplied by (b-1)/b cancels the leading term in Theorem 9 [7, 8]. The finitely omitted primes only change the limit. ◻
The centering in Corollary 10 removes one global number. It does not describe the individual residue classes inside the finite table. Those classes retain a nontrivial pattern even though their grand mean is fixed by reflection.
For b=10, Theorem 9 gives \Phi_{10}(x;1)=-\frac9{10}\log\log x+\kappa_{10}+o(1). Four finite cutoffs show the exact coefficient together with its additive remainder. For compactness, write A_x=\frac{\Phi_{10}(x;1)}{\log\log p_{\max}}, \qquad B_x=\Phi_{10}(x;1)+0.9\log\log p_{\max}.
| primes | p_{\max} | \Phi_{10}(x;1) | A_x | B_x |
|---|---|---|---|---|
| 1{,}000 | 7{,}951 | -1.367764 | -0.623094 | 0.607842 |
| 10{,}000 | 104{,}779 | -1.596580 | -0.652326 | 0.606185 |
| 100{,}000 | 1{,}299{,}811 | -1.773479 | -0.670605 | 0.606655 |
| 664{,}574 | 9{,}999{,}991 | -1.895354 | -0.681796 | 0.606595 |
At the largest cutoff the raw ratio is -0.681796. The stable B_x column shows why. The additive constant is still visible after division by the slowly growing \log\log p_{\max}.
The forty units modulo 100 split into ten lifts of each possible final digit. Exact evaluation of (12) gives the finite-table means below. The prime number theorem in arithmetic progressions makes the ten lifts equiprobable, while (9) tends to R/10. This gives the limiting fluctuation means. The observed column uses every prime up to 10^7.
| p\bmod10 | mean S_{10} | R/10 | limiting mean \Delta_{10} | observed mean \Delta_{10} |
|---|---|---|---|---|
| 1 | -17/10 | 0 | -17/10 | -1.705526 |
| 3 | -9/10 | 1/5 | -11/10 | -1.100334 |
| 7 | -1/10 | 3/5 | -7/10 | -0.701248 |
| 9 | +7/10 | 4/5 | -1/10 | -0.103091 |
The four limiting means average to -9/10, as required. Through 10^7, the overall mean is -0.902640. The fluctuation is negative for 70.0294 percent of the included primes, positive for 19.9871 percent, and zero for 9.9835 percent. These proportions describe the finite range.
The finite calculations in this section were performed with nfield [9]. They reproduce the finite table, class means, direct collision counts, and every value in Tables 1 and 2. They also check the finite reflection identity.
The theorem has a clean division of labor. Collision arithmetic creates the finite table and fixes its mean by reflection. Classical Mertens theory transports that mean into the prime harmonic sum. The drift is the sum of two finite effects. Reflection contributes -1/2, and the unequal-bin correction contributes the additional -(1/2-1/b). The primes sample those finite residue classes evenly at leading order.
Removing the constant part exposes variation among the residue classes inside the table. At lag one the table lives modulo b^2, while its coarse means live modulo b. The same separation can be made at higher lags. The drift law does not determine which arithmetic channels survive after those coarse means are removed.
A digit partition creates an exact finite bias, and prime harmonic weighting turns that bias into a logarithmic drift. Once the bias is removed, which channels carry the remaining signal?
[1]A. S. Petty, Bin Derangements and the Gate Width Theorem, July 2022. https://doi.org/10.5281/zenodo.21850917.
[2]A. S. Petty, Silent Primes and the Variance of the Collision Count, January 2023. https://doi.org/10.5281/zenodo.21851955.
[3]A. Lempel and H. Greenberger, Families of sequences with optimal Hamming-correlation properties, IEEE Trans. Inform. Theory 20 (1974), no. 1, 90–94. https://doi.org/10.1109/TIT.1974.1055169.
[4]S. C. Kak and A. Chatterjee, On decimal sequences, IEEE Trans. Inform. Theory 27 (1981), no. 5, 647–652. https://doi.org/10.1109/TIT.1981.1056394.
[5]K. Girstmair, Digit variance and Dedekind sums, J. Number Theory 65 (1997), no. 2, 197–205. https://doi.org/10.1006/jnth.1997.2149.
[6]K. S. Williams, Mertens' theorem for arithmetic progressions, J. Number Theory 6 (1974), no. 5, 353–359. https://doi.org/10.1016/0022-314X(74)90032-8.
[7]F. Mertens, Ein Beitrag zur analytischen Zahlentheorie, J. Reine Angew. Math. 78 (1874), 46–62. https://doi.org/10.1515/crll.1874.78.46.
[8]H. Davenport, Multiplicative Number Theory, 3rd ed., Springer, 2000.
[9]A. S. Petty, nfield, software repository. https://github.com/alexspetty/nfield
Discussion
Sign in to join the discussion.