← Blog post

The Secondary Term in Digit-Collision Energy

Alexander S. Petty

Abstract

For a positive integer N, we determine the secondary term in the classical square-box GCD and LCM sum \sum_{r,s\le N}\frac{(r,s)}{[r,s]} =3N-\frac{3}{\pi^2}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr). This determines the leading correction within a published O((\log N)^2) remainder.

For an odd prime p, the finite lag-one collision energy is a p^2-point sample of the Bernoulli boundary response underlying that sum. Sampling could introduce an alias at the same secondary scale and change the coefficient. We prove that it does not. The sampling defect satisfies \mathcal S_p \ll \frac{(\log(2p))^2}{\log\log(3p)}, and consequently E_p =p^3-\frac1{\pi^2}p^2(\log p)^2 +O\left(\frac{p^2(\log(2p))^2}{\log\log(3p)}\right).

The proof gives two exact descriptions of the finite defect. One is a boundary-prefix aliasing law on the carrier grid. The other is the surviving lower-modulus Dedekind term in a Rademacher decomposition. Opening that term into centered inverse-residue series supplies the strict saving. An exact switching ledger also separates local spikes from their aggregate signed contribution. The stronger limit \mathcal S_p\to1 remains open.

June 2026 (revised August 2026)
2020 Mathematics Subject Classification: Primary 11K38; Secondary 11A25, 11F20, 11L03, 42A16

The finite secondary problem

The lag-one digit-collision energy studied here is the square mass of a class-centered response on prime-square residues. Its leading order is cubic, while its first correction sits two powers lower and carries two logarithms. The definitions and exact finite identity are reconstructed below. The continuous response predicts the coefficient -1/\pi^2, but the prime-square carrier grid could still contribute at exactly the same scale. The question is whether the correction -\pi^{-2}p^2(\log p)^2 survives finite sampling unchanged.

The distinction is structural. A coefficient altered by the grid would belong partly to the sampling scheme. A coefficient that survives belongs to the collision law itself. The floor weights retain the entire discrepancy. Each downward rounding selects a prefix of an inverse-residue sequence. A single prefix can spike, and a single carrier can retain one sign. The proof keeps the additive phase within each descended carrier, then uses a bounded mean for the divisor burden across carriers.

The continuous response has a classical form. Put A(N)=\sum_{r,s\le N}\frac{(r,s)}{[r,s]}, where (r,s) and [r,s] denote the greatest common divisor and least common multiple. Since (r,s)/[r,s]=(r,s)^2/(rs), the Bernoulli square capacity below is A(N)/3. Hilberdink, Luca, and Tóth proved A(N)=3N+O\bigl((\log N)^2\bigr) in [3]. The continuous argument here determines the leading correction at that scale and gives A(N)=3N-\frac{3}{\pi^2}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr). The finite question is then precise. One must show that sampling the Bernoulli response on the prime-square carrier grid does not create another quadratic logarithm. The strict all-prime saving proved below places the sampling defect beneath that scale and transfers the coefficient unchanged to the digit-collision energy.

The surrounding analytic tools are classical. Saffari and Vaughan provide the distribution input for fractional parts of N/k [8]. Rademacher and Grosswald supply the reciprocity framework [7]; Vardi and Conrey, Fransen, Klein, and Scott place Dedekind sums in Kloosterman and mean-value theory [9, 1]; Lemke Oliver and Soundararajan study the relevant additive Fourier transform at prime moduli [6]; and Korolev supplies the incomplete inverse-residue estimate used after modulus descent [5]. The results proved here are the explicit coefficient left unresolved by the published square-box error term, two exact descriptions of the carrier-grid defect, and a uniform saving that proves the grid does not change that coefficient.

Fix an odd prime p, put N=p-1 and m=p^2, and define the lag-one collision diagonal G_p=\{d(p+1):0\le d<p\}. For 0\le a<m, let C_p(a) =\sum_{n\in G_p} \left( \left\lfloor\frac{(n+1)a}{m}\right\rfloor -\left\lfloor\frac{na}{m}\right\rfloor \right), \qquad F_p(a)=C_p(a)-\frac ap. The finite collision energy is E_p=\sum_{a=0}^{m-1}F_p(a)^2. Define \mathop{\mathrm{saw}}(x)= \begin{cases} \{x\}-\tfrac12,&x\notin\mathbb{Z},\\ 0,&x\in\mathbb{Z}, \end{cases} \qquad \mathcal{B}_p(x)=2\sum_{r=1}^{N}\mathop{\mathrm{saw}}(rx).

Lemma 1 (Floor-to-fractional response). For every 0\le a<m, F_p(a) =\sum_{n\in G_p} \left( \left\{\frac{na}{m}\right\} -\left\{\frac{(n+1)a}{m}\right\} \right).

Proof. For each n, \left\lfloor\frac{(n+1)a}{m}\right\rfloor -\left\lfloor\frac{na}{m}\right\rfloor =\frac am +\left\{\frac{na}{m}\right\} -\left\{\frac{(n+1)a}{m}\right\}. There are p collision points and m=p^2. Summing the constant term therefore gives a/p, which disappears after centering. ◻

Theorem 2 (Collision-to-Bernoulli straightening). Put c=1-p, and let [x]_m denote the least residue of x modulo m. Then c is a unit modulo m, and for every integer A, F_p([cA]_m)=\mathcal{B}_p(A/m). Consequently, E_p=\sum_{A=0}^{m-1}\mathcal{B}_p(A/m)^2.

Proof. Since c\equiv1\pmod p, one has (c,m)=1. For n=d(p+1)\in G_p, cn=d(1-p^2)\equiv d\pmod m, \qquad c(n+1)\equiv d+1-p\pmod m. Fractional parts are unchanged by replacing [cA]_m with cA. Lemma 1 therefore gives F_p([cA]_m) =\sum_{d=0}^{p-1}\left\{\frac{dA}{m}\right\} -\sum_{d=0}^{p-1}\left\{\frac{(d+1-p)A}{m}\right\}. Put j=p-1-d in the second sum. The two zero-index terms vanish, so F_p([cA]_m) =\sum_{r=1}^{p-1} \left( \left\{\frac{rA}{m}\right\} -\left\{-\frac{rA}{m}\right\} \right) =2\sum_{r=1}^{p-1}\mathop{\mathrm{saw}}(rA/m). This proves (1). Multiplication by c permutes the residues modulo m, which proves (2). ◻

To identify the finite object with the digit-collision energy, write a unit as a=pq+s with 1\le s<p. Its class-centered response is C_p(a)-q-\frac sp=C_p(a)-\frac ap=F_p(a). If a=pc, then the defining floor sum telescopes to C_p(pc) =\sum_{d=0}^{p-1} \left( \left\lfloor\frac{(d+1)c}{p}\right\rfloor -\left\lfloor\frac{dc}{p}\right\rfloor \right) =c, so F_p(pc)=0. Thus E_p is exactly the square mass of the class-centered finite response defined above.

The normalized lag-one energy is therefore exactly the carrier-grid average \mathcal{D}_p =\frac1m\sum_{A=0}^{m-1}\mathcal{B}_p(A/m)^2 =\frac{E_p}{p^2}. The L^2 Fourier series of the centered sawtooth gives \int_0^1\mathop{\mathrm{saw}}(rx)\mathop{\mathrm{saw}}(sx)\,dx =\frac{(r,s)^2}{12rs} \qquad (r,s\ge1). Summing this identity gives the continuous square capacity \mathcal{J}(N) =\int_0^1\mathcal{B}_p(x)^2\,dx =\frac13\sum_{1\le r,s\le N}\frac{(r,s)^2}{rs}. The GCD sum on the right defines \mathcal{J}(N) for every positive integer N. We study \mathcal{S}_p=\mathcal{J}(N)-\mathcal{D}_p.

The defect first appears in two exact forms. The geometric form uses the rational jump phases of the straightened boundary. The arithmetic form uses Rademacher reciprocity. Their equality identifies the quotient-shell reservoir with the sampling alias. The additive transform then opens every carrier into inverse-residue series and supplies the strict saving needed for finite transfer.

The continuous secondary coefficient

The leading estimate \mathcal{J}(N)=N+O((\log N)^2) is equivalent to the published GCD and LCM asymptotic of Hilberdink, Luca, and Tóth [3]. The argument below determines the first coefficient by retaining the fractional-floor term that their error bound leaves unresolved.

For k\ge2, put H_k^*=\sum_{\substack{1\le a<k\\(a,k)=1}}\frac1a, \qquad H_1^*=0.

Lemma 3 (Reduced-ratio capacity). For every positive integer N, 3\mathcal{J}(N) =N+2\sum_{k=2}^{N}\frac{\lfloor N/k\rfloor}{k}H_k^*.

Proof. The diagonal of the GCD sum in (4) contributes N. For an off-diagonal pair, divide by its common divisor and call the larger reduced coordinate k and the smaller one a. The conditions are 1\le a<k and (a,k)=1. The common scale has \lfloor N/k\rfloor choices, and the two orders of the pair give the factor two. ◻

Lemma 4 (Coprime Euler sum). \sum_{k=2}^{\infty}\frac{H_k^*}{k^2}=1.

Proof. Möbius inversion gives \sum_{k=2}^{\infty}\frac{H_k^*}{k^2} =\frac1{\zeta(3)}\sum_{k=2}^{\infty}\frac{h_{k-1}}{k^2}, \qquad h_j=\sum_{n=1}^{j}\frac1n. The classical Euler sum \sum_{k\ge1}h_k/k^2=2\zeta(3) [2] implies \sum_{k\ge2}h_{k-1}/k^2=\zeta(3). ◻

Proposition 5 (Exact continuous-defect ledger). For every N\ge1, N-\mathcal{J}(N) =\frac23N\sum_{k>N}\frac{H_k^*}{k^2} +\frac23\sum_{k\le N} \left\{\frac Nk\right\}\frac{H_k^*}{k}.

Proof. Insert \lfloor N/k\rfloor=N/k-\{N/k\} into (6) and complete the truncated sum with Lemma 4. ◻

Both terms in (8) are nonnegative. The tail is O(\log(2N)), since H_k^*\le h_{k-1}\ll\log(2k). The squared logarithm comes from the fractional-floor term.

The only external asymptotic input in this continuous calculation is the logarithmically weighted fractional-part theorem of Saffari and Vaughan [8]. For 2\le Y\le X, it gives, uniformly in 0\le\alpha\le1, \frac1{\log Y}\sum_{n\le Y}\frac1n \mathbf 1_{[0,\alpha)}\!\left(\left\{\frac Xn\right\}\right) =\alpha+O\left(\frac{(\log X)^{2/3}}{\log Y}\right).

Lemma 6 (Weighted fractional parts). For real X\ge3, W(X):=\sum_{n\le X}\frac{h_{n-1}}n\left\{\frac Xn\right\} =\frac14(\log X)^2+O\bigl((\log X)^{5/3}\bigr).

Proof. Integrate (9) in \alpha. Since \int_0^1\mathbf 1_{[0,\alpha)}(u)\,d\alpha=1-u, one first obtains h_{\lfloor Y\rfloor}-A_X(Y) =\frac12\log Y+O\bigl((\log X)^{2/3}\bigr). Since h_{\lfloor Y\rfloor}=\log Y+O(1), rearrangement gives, uniformly for 2\le Y\le X, A_X(Y):=\sum_{n\le Y}\frac1n\left\{\frac Xn\right\} =\frac12\log Y+O\bigl((\log X)^{2/3}\bigr). Partial summation gives \sum_{n\le X}\frac{\log n}{n}\left\{\frac Xn\right\} =\frac14(\log X)^2+O\bigl((\log X)^{5/3}\bigr). Finally, h_{n-1}=\log n+\gamma+O(1/n). The constant term contributes O(\log X) by (11), and the final error is O(1). ◻

Lemma 7 (Möbius reduction). Put R(N)=\sum_{k\le N}\left\{\frac Nk\right\}\frac{H_k^*}{k}. Then R(N)=\sum_{d\le N}\frac{\mu(d)}{d^2}W(N/d).

Proof. Use H_k^*=\sum_{d\mid k}\frac{\mu(d)}dh_{k/d-1} and write k=dn. ◻

Proposition 8 (Fractional-floor asymptotic). As N\to\infty, R(N)=\frac1{4\zeta(2)}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr).

Proof. For d\le N/3, insert Lemma 6 into (12) and sum the errors absolutely. For 1\le X<3, use the same finite-sum definition of W(X), which is O(1). The range d>N/3 therefore contributes O(1/N). Absolute convergence of \sum_{d\ge1}\frac{1+|\log d|^2}{d^2} then gives \begin{aligned} R(N) &=\frac14\sum_{d\le N}\frac{\mu(d)}{d^2} \left(\log\frac Nd\right)^2 +O\bigl((\log N)^{5/3}\bigr)\\ &=\frac1{4\zeta(2)}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr). \end{aligned} Extending the main d-sum from N/3 to N costs another O(1/N). The linear and constant logarithmic moments are absorbed by the stated error. ◻

Theorem 9 (Continuous secondary coefficient). As N\to\infty, \mathcal{J}(N) =N-\frac1{\pi^2}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr).

Proof. The first term in (8) is O(\log(2N)). Proposition 8 gives N-\mathcal{J}(N) =\frac23\cdot\frac1{4\zeta(2)}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr). Since \zeta(2)=\pi^2/6, the coefficient is 1/\pi^2. ◻

Corollary 10 (Secondary term for the GCD and LCM sum). Let F(n)=\sum_{k=1}^{n}\frac{(k,n)}{[k,n]}. As N\to\infty through the positive integers, \sum_{r,s\le N}\frac{(r,s)}{[r,s]} =3N-\frac{3}{\pi^2}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr), and \sum_{n\le N}F(n) =2N-\frac{3}{2\pi^2}(\log N)^2 +O\bigl((\log N)^{5/3}\bigr). This identifies the first correction at the logarithm-squared scale left by [3].

Proof. The identity \frac{(r,s)}{[r,s]}=\frac{(r,s)^2}{rs} and the definition of \mathcal{J}(N) give \sum_{r,s\le N}\frac{(r,s)}{[r,s]}=3\mathcal{J}(N). Theorem 9 proves (15). Symmetry gives \sum_{r,s\le N}\frac{(r,s)}{[r,s]} =2\sum_{n\le N}F(n)-N, because every diagonal term equals one. Solving for the one-sided sum proves (16). ◻

The origin of the quadratic logarithmic coefficient is now explicit. The finite question is whether sampling on the p^2-carrier adds another term of that size. Theorem 27 proves that it does not.

The jump geometry of the straightened response

Write B_1(x)=\mathop{\mathrm{saw}}(x), \qquad B_2(x)=\{x\}^2-\{x\}+\frac16. Away from its jumps, \mathcal{B}_p has the constant slope L=2\sum_{r=1}^{N}r=N(N+1)=pN. Every nonzero jump has a unique reduced location x=\frac ak, \qquad 2\le k\le N, \qquad 1\le a<k, \qquad (a,k)=1. Put M_k=\left\lfloor\frac Nk\right\rfloor, \qquad s_k=N-kM_k, and define the cyclic prefix response T_{k,s}(a)=2\sum_{r=1}^{s}\mathop{\mathrm{saw}}\left(\frac{ra}{k}\right).

Lemma 11 (Jump data). At the reduced point x=a/k, the midpoint value and the one-sided jump of \mathcal{B}_p are \mathcal{B}_p(a/k)=T_{k,s_k}(a), \qquad [\mathcal{B}_p]_{a/k}=-2M_k. For f=\mathcal{B}_p^2 one has [f]_{a/k}=-4M_kT_{k,s_k}(a), \qquad [f']_{a/k}=-4LM_k. At the origin, the one-sided values of \mathcal{B}_p are N and -N. Hence f has equal one-sided value N^2, its chosen value is f(0)=0, and [f]_0=0, \qquad [f']_0=-4LN.

Proof. Exactly M_k indices among 1\le r\le N are divisible by k, and each corresponding sawtooth drops by 1. The factor 2 in \mathcal{B}_p gives the jump -2M_k. Split the index interval into M_k complete blocks of length k and one terminal block of length s_k. Since multiplication by a permutes the residues modulo k, \sum_{r=1}^{k}\mathop{\mathrm{saw}}(ra/k)=0. Only the terminal block remains, giving the midpoint value in (19). If the midpoint is T and the half-jump is M_k, the one-sided values are T+M_k and T-M_k. Squaring gives the first formula in (20). Since the classical derivative is L on both sides, f'=2L\mathcal{B}_p away from jumps, which gives the second formula. The origin statements follow directly from the one-sided values of the sawtooth. ◻

Exact carrier-grid aliasing

The carrier grid does not meet any nonzero jump. Indeed, if a reduced fraction a/k with k\le N<p were equal to A/p^2, then k would divide p^2, which is impossible for k\ge2.

We use the following periodic aliasing identity. The statement is specialized to the present piecewise-quadratic function in order to keep all endpoint conventions explicit.

Lemma 12 (Periodic jump aliasing). Let f=\mathcal{B}_p^2, and let the jump sums include the origin. Then \begin{aligned} \frac1m\sum_{A=0}^{m-1}f(A/m)-\int_0^1f(x)\,dx &=-\frac{N^2}{m} +\frac1m\sum_x[f]_x B_1(mx)\\ &\quad-\frac1{2m^2}\sum_x[f']_xB_2(mx). \end{aligned}

Proof. Point values do not affect the Fourier coefficients. In the distributional sense, f''=2L^2+\sum_x[f']_x\delta_x+\sum_x[f]_x\delta'_x. For every nonzero integer n this gives \widehat f(n) =\frac1{2\pi i n}\sum_x[f]_x\mathrm{e}^{-2\pi i nx} +\frac1{(2\pi i n)^2}\sum_x[f']_x\mathrm{e}^{-2\pi i nx}. Interpret the first frequency series by symmetric Abel summation. For 0<\rho<1, multiply the term with index j by \rho^{|j|}. The jump sums are finite, so they may be interchanged with the regularized series. Letting \rho\uparrow1 gives \sum_{j\ne0}\frac{\mathrm{e}^{-2\pi i jy}}{2\pi i j}=B_1(y), \qquad \sum_{j\ne0}\frac{\mathrm{e}^{-2\pi i jy}}{(2\pi i j)^2} =-\frac12B_2(y), where the first identity has the midpoint value B_1(y)=0 at integers. The second series is absolutely convergent. These identities produce the two jump sums in (22). The regulated Fourier value of f at the origin is the common one-sided value N^2, while the actual carrier sample is f(0)=0. This supplies the first term -N^2/m. ◻

Define \mathcal{A}_p =\sum_{k=2}^{N}M_k \sum_{\substack{1\le a<k\\(a,k)=1}} T_{k,s_k}(a)\mathop{\mathrm{saw}}\left(\frac{ma}{k}\right) and \mathcal{C}_p =\frac N6+ \sum_{k=2}^{N}M_k \sum_{\substack{1\le a<k\\(a,k)=1}} B_2\left(\frac{ma}{k}\right).

Theorem 13 (Exact endpoint-reservoir alias formula). For every odd prime p, \boxed{ \mathcal{S}_p =\frac{N^2}{m} +\frac4m\mathcal{A}_p -\frac{2L}{m^2}\mathcal{C}_p. }

Proof. Insert the jump data from Lemma 11 into Lemma 12, and use \mathcal{S}_p=\int f-m^{-1}\sum f(A/m). ◻

The final term is elementary and lower order.

Proposition 14 (Bernoulli-curvature evaluation). One has \boxed{ \mathcal{C}_p =\frac N6+ \sum_{k=2}^{N}\frac{M_k}{6k} \prod_{\varpi\mid k}(1-\varpi), } where the product is over the distinct prime divisors of k. Consequently, \mathcal{C}_p\ll N\log(2N), \qquad \frac{2L}{m^2}\mathcal{C}_p\ll\frac{\log(2p)}p.

Proof. Since (m,k)=1, multiplication by m permutes the reduced residues modulo k. Möbius inversion and the Bernoulli multiplication formula give \sum_{a\in(\mathbb{Z}/k\mathbb{Z})^\times}B_2(a/k) =\frac1{6k}\sum_{d\mid k}\mu(d)d =\frac1{6k}\prod_{\varpi\mid k}(1-\varpi). The last expression has absolute value at most 1/6. We also have \sum_{k\le N}M_k\ll N\log(2N). Together with (17), this finishes the proof. ◻

Thus the nontrivial alias is the single signed jump-phase correlation \mathcal{A}_p. \mathcal{S}_p =\frac{(p-1)^2}{p^2} +\frac4{p^2}\mathcal{A}_p +O\left(\frac{\log p}{p}\right). This is the collision-native form of the sampling problem.

Exact Rademacher-reservoir bridge

We now identify the same residual in the reduced-ratio coordinates of Dedekind reciprocity. For coprime h and q, write s(h,q)=\sum_{x=1}^{q-1}\mathop{\mathrm{saw}}(x/q)\mathop{\mathrm{saw}}(hx/q) for the classical Dedekind sum.

We shall also use its inversion symmetry. If (h,q)=1, then s(h,q)=s(\overline h,q). Indeed, the substitution y=hx\bmod q in the defining sum exchanges the two sawtooth factors.

For a>k\ge1, (a,k)=1, put w_p(a)=\left\lfloor\frac Na\right\rfloor and define X_p= \sum_{\substack{2\le a\le N\\1\le k<a\\(a,k)=1}} w_p(a)\left(\frac ak+\frac ka\right), \mathfrak R_p(k) =\sum_{\substack{k<a\le N\\(a,k)=1}} w_p(a)\,s\bigl(m\overline a^{(k)},k\bigr), where \overline a^{(k)} is the inverse of a modulo k. Then R_p^{(k)}=-8\sum_{k=2}^{N-1}\mathfrak R_p(k).

Lemma 15 (Discrete pair correlation). For 1\le r,s\le N, \frac1m\sum_{A=0}^{m-1} \mathop{\mathrm{saw}}(rA/m)\mathop{\mathrm{saw}}(sA/m) =\frac1m s(s\overline r^{(m)},m).

Proof. Multiplication by r permutes the residues modulo m. Substitute x=rA\pmod m in the finite sum. ◻

Theorem 16 (Exact sampling-reservoir bridge). For every odd prime p, \boxed{ \mathcal{S}_p =\frac{N^2}{p^2} -\frac{2}{3p^4}\bigl(N+X_p\bigr) -\frac{R_p^{(k)}}{p^2}. } Moreover, X_p\ll p^2\log(2p), so \boxed{ \mathcal{S}_p =\frac{(p-1)^2}{p^2} -\frac{R_p^{(k)}}{p^2} +O\left(\frac{\log p}{p^2}\right). }

Proof. Expand (3). Lemma 15 then gives the weight-free Dedekind representation of the sampled energy. The diagonal is \Delta_p=4N s(1,m) =\frac{N(m-1)(m-2)}{3m}. For the off-diagonal, divide a pair by its gcd and choose the orientation with larger reduced coordinate a. The two orientations agree by (32). We then apply three-term reciprocity [7]. \begin{aligned} s(k^{-1}a\bmod m,m) &=\frac{m}{12ka} +\frac1{12}\left(\frac{a}{km}+\frac{k}{am}\right)-\frac14\\ &\quad-s(k\overline m^{(a)},a)-s(m\overline a^{(k)},k). \end{aligned} The first term is \operatorname{Main}_p =\frac{2m}{3} \sum_{\substack{a>k\ge1\\(a,k)=1\\a\le N}} \frac{w_p(a)}{ka}. The a-modulus Dedekind sum vanishes after summing a complete reduced residue system in k, by s(-u,a)=-s(u,a). The k-modulus sum is exactly R_p^{(k)}. The elementary remainder is R_p^{(0)} =\frac{2}{3m}X_p-N(N-1), because \sum_{\substack{a>k\ge1\\(a,k)=1\\a\le N}}w_p(a) =\binom N2. Thus E_p=\Delta_p+\operatorname{Main}_p+R_p^{(0)}+R_p^{(k)}. On the continuous side, the reduced-ratio form of (4) is \mathcal{J}(N)=\frac N3+\frac{\operatorname{Main}_p}{m}. Subtracting E_p/m cancels the complete rational main term and yields \mathcal{S}_p =\frac N3-\frac{\Delta_p}{m} -\frac{R_p^{(0)}}m-\frac{R_p^{(k)}}m. Now \frac N3-\frac{\Delta_p}{m} =\frac Nm-\frac{2N}{3m^2}, and insertion of R_p^{(0)} gives (41).

Finally, \begin{aligned} X_p &\le \sum_{a=2}^{N}\frac Na \left(aH_{a-1}+\frac1a\sum_{k<a}k\right) \ll N\sum_{a\le N}\log(2a) \ll N^2\log(2N), \end{aligned} which proves (42) and (43). ◻

Remark 17 (Exact identification of the reservoir). The quantity R_p^{(k)} is the surviving k-modulus Rademacher carrier in the reduced-ratio decomposition. Formula (43) identifies it with the carrier-grid sampling defect in Dedekind-reciprocity coordinates after the endpoint reservoir and explicit curvature correction are removed.

The absolute-moment baseline

The exact bridge converts every improvement in the signed k-modulus reservoir directly into a sampling theorem.

Lemma 18 (First absolute moment). For k\ge2, \sum_{u\in(\mathbb{Z}/k\mathbb{Z})^\times}|s(u,k)| \ll k(\log(2k))^2.

Proof. Use the cotangent formula [7] s(u,k)=\frac1{4k}\sum_{n=1}^{k-1} \cot\frac{\pi n}{k}\cot\frac{\pi nu}{k}. Put C(k)=\sum_{v=1}^{k-1}|\cot(\pi v/k)|. Splitting at k/2 and using |\cot x|\ll1/x near zero gives C(k)\ll k\log(2k). Fix n, write \delta=(n,k) and q=k/\delta. As u runs over the units modulo k, the residues nu\bmod k are nonzero multiples \delta r with (r,q)=1, each with multiplicity at most \delta. Hence \sum_{u\in(\mathbb{Z}/k\mathbb{Z})^\times} \left|\cot\frac{\pi nu}{k}\right| \le \delta C(q)\ll k\log(2k). Summing the cotangent formula absolutely over u and then n proves (44). ◻

Lemma 19 (One-carrier bound). For 2\le k<p, \boxed{ |\mathfrak R_p(k)|\ll p\bigl(\log(2k)\bigr)^2. } The implied constant is absolute.

Proof. Put \beta=m\bmod k and f_k(a)=s(\beta\overline a,k), \qquad g(a)=\left\lfloor\frac Na\right\rfloor. The function f_k is periodic modulo k and has zero mean on a complete reduced residue system. Every interval sum of f_k is therefore bounded by twice the first absolute moment in (44). Abel summation on k<a\le N gives |\mathfrak R_p(k)| \ll \frac pk\,k\bigl(\log(2k)\bigr)^2. ◻

Proposition 20 (Unconditional sampling bound). For odd primes p, \boxed{ \mathcal{S}_p\ll(\log p)^2. }

Proof. Sum Lemma 19 over k<p, use (39), and apply Theorem 16. ◻

Inverse-residue damping

The absolute moment discards the phase of each Dedekind sum. To go below the quadratic logarithmic scale, we keep that phase and open it additively. Classical work relates Dedekind sums to Kloosterman sums and studies their mean values [9, 1]. The estimate needed here must remain uniform after every composite carrier is descended to its lower moduli and the resulting divisor costs are summed. For prime modulus, Lemke Oliver and Soundararajan identify the normalized additive Fourier transform of the Dedekind sum with an inverse-residue sawtooth series and establish its limiting distribution and large-value behavior [6]. The bound required here has a different uniformity requirement. The carrier moduli are generally composite, and their Fourier arguments may be nonunits. The exact transform and numerator descent isolate the true lower moduli before the carrier family is summed. For q\ge2, put \mathrm{e}_q(x)=\exp(2\pi i x/q) and extend the Dedekind sum to every residue h\bmod q by \widetilde s(h,q) =\sum_{x\bmod q}\mathop{\mathrm{saw}}(x/q)\mathop{\mathrm{saw}}(hx/q). This agrees with s(h,q) when (h,q)=1. Define the unrestricted and unit-supported additive transforms \begin{aligned} F_q(t) &=\sum_{h\bmod q}\widetilde s(h,q)\mathrm{e}_q(-th), \\ D_q(t) &=\sum_{h\in(\mathbb{Z}/q\mathbb{Z})^\times}s(h,q)\mathrm{e}_q(-th). \end{aligned}

The transform is governed by one centered inverse-residue series. For q\ge2 and r\in\mathbb{Z}, let B_q(r) =\sum_{\substack{n\ge1\\(n,q)=1}} \frac1n\mathop{\mathrm{saw}}\left(\frac{r\overline n^{(q)}}q\right). The coefficient sequence is periodic and has mean zero. The series therefore converges by summation by parts.

Lemma 21 (Exact additive transform). For K\ge2 and t\in\mathbb{Z}, \boxed{ \frac{F_K(t)}K =-\frac1{\pi i} \sum_{\substack{g\mid(K,t)\\K/g\ge2}} \frac1g B_{K/g}(t/g). } Moreover, \boxed{ D_k(t)=\sum_{d\mid k}\mu(d)F_{k/d}(t), } where F_1=0.

Proof. The Bernoulli multiplication formula gives \widetilde s(db,dK)=\widetilde s(b,K). Indeed, write x=y+jK and sum first over 0\le j<d. Then \sum_{j=0}^{d-1} \mathop{\mathrm{saw}}\left(\frac{y+jK}{dK}\right) =\mathop{\mathrm{saw}}(y/K). Insert 1_{(h,k)=1}=\sum_{d\mid(h,k)}\mu(d) into (49), use (55), and write h=db. This proves (54).

For the first identity, use the exact finite Fourier expansion of the sawtooth. Its nonzero Fourier coefficients are \sum_{x\bmod K}\mathop{\mathrm{saw}}(x/K)\mathrm{e}_K(-nx) =\frac i2\cot\frac{\pi n}{K}. Expanding the second sawtooth in (47), summing over h, and solving nx\equiv t\pmod K gives F_K(t) =\frac i2 \sum_{\substack{g\mid(K,t)\\K/g\ge2}} \sum_{u\in(\mathbb{Z}/(K/g)\mathbb{Z})^\times} \cot\frac{\pi u}{K/g} \mathop{\mathrm{saw}}\left( \frac{(t/g)\overline u^{(K/g)}}{K/g} \right). To justify the solution sum, write n=gu and K=gq. The g solutions of nx\equiv t\pmod K have the form x=x_0+jq, and the same Bernoulli multiplication formula reduces their sawtooth sum to \mathop{\mathrm{saw}}(x_0/q).

Pair u with q-u. Abel regularization is needed only in passing from the finite cotangent sum to the following infinite series. The partial fraction identity gives \sum_{\substack{u\bmod q\\(u,q)=1}} \cot\frac{\pi u}{q} \mathop{\mathrm{saw}}\left(\frac{r\overline u^{(q)}}q\right) =\frac{2q}{\pi}B_q(r). One may obtain the identity first with an Abel factor. The limit is valid because the coefficient sequence in (50) is periodic with mean zero. Substitution proves (52). ◻

Nonunit numerators in (50) descend to their true modulus before any analytic estimate is applied.

Lemma 22 (Exact numerator descent). Let c=(r,q) and Q=q/c. If Q=1, then B_q(r)=0. If Q\ge2, put P(q,Q)=\prod_{\substack{\ell\mid q\\\ell\nmid Q}}\ell. Then \boxed{ B_q(r) =\sum_{d\mid P(q,Q)}\frac{\mu(d)}d B_Q\left(\frac rc\,\overline d^{(Q)}\right). } Consequently, |B_q(r)| \le \sigma_{-1}(q) \max_{u\in(\mathbb{Z}/Q\mathbb{Z})^\times}|B_Q(u)|. Here \sigma_{-1}(q)=\sum_{d\mid q}d^{-1}.

Proof. For (n,q)=1, reduction of \overline n^{(q)} modulo Q is \overline n^{(Q)}, and \mathop{\mathrm{saw}}\left(\frac{r\overline n^{(q)}}q\right) =\mathop{\mathrm{saw}}\left(\frac{(r/c)\overline n^{(Q)}}Q\right). The condition (n,q)=1 now consists of (n,Q)=1 together with the exclusion of the primes dividing P(q,Q). Inclusion and exclusion followed by n=dm proves (56). Taking absolute values gives (57). ◻

The completion estimate needed beyond Korolev’s range is recorded separately.

Lemma 23 (Completed inverse prefixes). For q\ge2, (r,q)=1, and 1\le X\le q, put A_{q,r}(X) = \sum_{\substack{n\le X\\(n,q)=1}} \mathop{\mathrm{saw}}\left(\frac{r\overline n^{(q)}}q\right). Then A_{q,r}(X) \ll q^{1/2}\tau(q)\sigma_{-1/2}(q)(\log(2q))^2, \qquad \sigma_{-1/2}(q)=\sum_{d\mid q}d^{-1/2}. Here \tau(q) is the divisor function. In particular, A_{q,r}(X)=q^{1/2+o(1)} uniformly in r and X.

Proof. Let b_q(y)=\mathop{\mathrm{saw}}(y/q). Its finite Fourier transform satisfies \widehat b_q(0)=0, \qquad \widehat b_q(j) =\frac i2\cot\frac{\pi j}{q} \quad (j\not\equiv0\pmod q). Fourier inversion gives A_{q,r}(X) = \frac1q \sum_{j=1}^{q-1} \widehat b_q(j) \sum_{\substack{n\le X\\(n,q)=1}} \mathrm{e}_q(jr\overline n^{(q)}). Standard completion followed by the Weil and Estermann bound for complete Kloosterman sums [4] gives \sum_{\substack{n\le X\\(n,q)=1}} \mathrm{e}_q(jr\overline n^{(q)}) \ll q^{1/2}\tau(q)(j,q)^{1/2}\log(2q). Since \left|\cot\frac{\pi j}{q}\right| \ll \frac{q}{\min(j,q-j)}, grouping the frequencies by divisors of (j,q) gives \sum_{j=1}^{q-1} |\widehat b_q(j)|(j,q)^{1/2} \ll q\log(2q)\sigma_{-1/2}(q). Substitution proves (58). The final assertion follows from the standard divisor bounds [4]. ◻

The logarithmic saving enters at the descended modulus.

Lemma 24 (Inverse-residue harmonic bound). Uniformly for q\ge2 and (r,q)=1, \boxed{ |B_q(r)| \ll\frac{\log(2q)}{\log\log(3q)}. }

Proof. Put A_{q,r}(x) =\sum_{\substack{n\le x\\(n,q)=1}} \mathop{\mathrm{saw}}\left(\frac{r\overline n^{(q)}}q\right). Korolev’s arbitrary-modulus inverse-fraction estimate [5] applies uniformly in the unit r. For sufficiently large q, set Y_q =\exp\left((\log q)^{4/5}(\log\log q)^{73/5}\right). When Y_q\le x\le q^{4/7}, it gives A_{q,r}(x) \ll x(\log x)^{-5/24}(\log q)^{1/6} (\log\log q)^{49/24}. Korolev states the corresponding estimate for the fractional-part sum. Because r and n are units, r\overline n^{(q)} is never zero modulo q. Subtracting one half from every term therefore gives (60) for the centered sawtooth. At x=Y_q, the factor following x is O(1/\log\log q). It decreases with x. Summation by parts on [Y_q,q^{4/7}] therefore contributes O(\log q/\log\log q) to (50). The initial range costs O(\log Y_q) =O\left((\log q)^{4/5}(\log\log q)^{73/5}\right) =o\left(\frac{\log q}{\log\log q}\right).

For q^{4/7}\le x\le q, Lemma 23 gives A_{q,r}(x)=q^{1/2+o(1)}. Summation by parts over this range contributes q^{1/2+o(1)}q^{-4/7}=q^{-1/14+o(1)}. The coefficient sequence has period q and zero mean. Its partial sums beyond the first period are therefore bounded by the largest partial sum within one period. A final summation by parts contributes O(1) beyond q. These estimates prove (59); the finitely many remaining moduli are absorbed into the implied constant. ◻

Proposition 25 (Uniform additive transform bound). For every k\ge2, \boxed{ \max_{t\bmod k}\frac{|D_k(t)|}{k} \ll \frac{\log(2k)}{\log\log(3k)}\sigma_{-1}(k)^3. }

Proof. Apply Lemma 22 inside (52), followed by Lemma 24. We use \sum_{g\mid K}g^{-1}=\sigma_{-1}(K). It follows that \frac{|F_K(t)|}{K} \ll \frac{\log(2K)}{\log\log(3K)}\sigma_{-1}(K)^2. Now use (54). The outer divisor sum costs at most one more factor of \sigma_{-1}(k). ◻

The divisor cost can be large at selected carriers, but no fixed power of it has a growing mean.

Lemma 26 (Bounded reciprocal-divisor moments). For every fixed positive integer j, \sum_{k\le x}\sigma_{-1}(k)^j\ll_j x.

Proof. It is enough to display the argument for j=3. Expanding the cube gives \sum_{k\le x}\sigma_{-1}(k)^3 \le x \sum_{d_1,d_2,d_3\ge1} \frac1{d_1d_2d_3[d_1,d_2,d_3]}. The last sum is an Euler product. Its local factor at a prime \ell is \sum_{a,b,c\ge0} \ell^{-a-b-c-\max(a,b,c)} =1+O(\ell^{-2}). The product converges. The same argument works for every fixed j. ◻

Theorem 27 (All-prime carrier damping). For every odd prime p, \begin{aligned} |R_p^{(k)}| &\ll \frac{p^2(\log(2p))^2}{\log\log(3p)}, \\ |\mathcal{S}_p| &\ll \frac{(\log(2p))^2}{\log\log(3p)}. \end{aligned} In particular, \mathcal{S}_p=o((\log p)^2).

Proof. Fix k<p and put \beta=p^2\bmod k. On the units modulo k, let f_{k,\beta}(a)=s(\beta\overline a^{(k)},k), and set it equal to zero on the nonunits. Dedekind inversion gives f_{k,\beta}(a)=s(a\overline\beta^{(k)},k). Its additive Fourier coefficient at n is therefore D_k(n\beta). Oddness gives D_k(0)=\sum_{u\in(\mathbb{Z}/k\mathbb{Z})^\times}s(u,k)=0. Fourier inversion and the geometric-sum bound therefore give, for every integer interval I, \begin{aligned} \left|\sum_{a\in I}f_{k,\beta}(a)\right| &\ll \max_{t\bmod k}|D_k(t)| \sum_{n=1}^{k-1}\frac1{\min(n,k-n)}\notag\\ &\ll \frac{k(\log(2k))^2}{\log\log(3k)} \sigma_{-1}(k)^3. \end{aligned} Abel summation with the decreasing weight \lfloor N/a\rfloor now yields |\mathfrak R_p(k)| \ll \frac{p(\log(2k))^2}{\log\log(3k)} \sigma_{-1}(k)^3.

Sum over k<p. For k\le\sqrt p, Lemma 19 gives a total contribution of O(p^{3/2}(\log(2p))^2). For \sqrt p<k<p, use Lemma 26 with j=3. On this range \log\log(3k)\asymp\log\log(3p), so the large-carrier contribution has the bound displayed below. The small-carrier contribution is absorbed into it. Thus \sum_{k=2}^{p-1}|\mathfrak R_p(k)| \ll \frac{p^2(\log(2p))^2}{\log\log(3p)}. Equation (39) proves (63), and Theorem 16 proves (64). ◻

The inverse-residue saving is obtained within each descended carrier, but a highly composite carrier may pay a large divisor cost after descent. Lemma 26 shows that these costs have bounded mean, so summing the carrierwise absolute bounds cannot rebuild a quadratic logarithm.

Small square-lock displacements cannot carry the secondary term

For 2\le k<p, let c_p(k) be the least signed representative of p^2 modulo k, chosen so that -\frac k2<c_p(k)\le\frac k2. We call c_p(k) the square-lock displacement of the carrier. It is an ordinary signed residue, not a character twist. It is nonzero because p is prime and k<p. Put \Delta_{k,p}(X) =\sum_{\substack{1\le a\le X\\(a,k)=1}} s\bigl(p^2\overline a^{(k)},k\bigr). Complete reduced residue blocks have sum zero. Expanding the floor weight in (37) therefore gives the exact collar decomposition \mathfrak R_p(k) =\sum_{h\le N/(k+1)} \Delta_{k,p}\left(\left\lfloor\frac Nh\right\rfloor\right). For an integer C with 1\le C\le p/2, define the normalized absolute collar mass \mathcal M_p(C) =\frac8{p^2} \sum_{\substack{2\le k<p\\0<|c_p(k)|\le C}} \sum_{h\le N/(k+1)} \left| \Delta_{k,p}\left(\left\lfloor\frac Nh\right\rfloor\right) \right|. The corresponding restricted part of the normalized reservoir satisfies \frac8{p^2} \left| \sum_{\substack{2\le k<p\\0<|c_p(k)|\le C}} \mathfrak R_p(k) \right| \le \mathcal M_p(C).

Theorem 28 (Square-lock localization). Uniformly for 1\le C\le p/2, \boxed{ \mathcal M_p(C) \ll \frac Cp\bigl(\log(2p)\bigr)^3 +p^{-1/3}\bigl(\log(2p)\bigr)^3. } Thus \mathcal M_p(C)=o(1) whenever C=o(p/(\log p)^3). It is o((\log p)^2) whenever C=o(p/\log p).

Proof. The signed representative satisfies the exact divisor lock k\mid p^2-c_p(k). Hence \#\{2\le k<p:0<|c_p(k)|\le C\} \le \sum_{1\le|c|\le C}\tau(p^2-c). Write D(x)=\sum_{n\le x}\tau(n). Voronoi’s classical divisor estimate [10] gives D(x)=x\log x+(2\gamma-1)x+O(x^{1/3}\log(2x)). Taking the difference across the two intervals adjacent to p^2 gives \sum_{1\le|c|\le C}\tau(p^2-c) \ll C\log(2p)+p^{2/3}\log(2p).

Every prefix in (71) is bounded by the first absolute moment in (44). A fixed carrier has at most p/k floor levels. Its absolute collar mass is therefore \sum_{h\le N/(k+1)} \left| \Delta_{k,p}\left(\left\lfloor\frac Nh\right\rfloor\right) \right| \ll \frac pk\,k\bigl(\log(2k)\bigr)^2 \ll p\bigl(\log(2p)\bigr)^2. Combine this estimate with (77) and divide by p^2. The two final assertions follow directly. ◻

Corollary 29 (Two localization scales). For all sufficiently large primes p, put C_4(p)=\left\lfloor\frac{p}{(\log p)^4}\right\rfloor, \qquad C_2(p)=\left\lfloor\frac{p}{(\log p)^2}\right\rfloor. Then \mathcal M_p(C_4(p))=o(1), \qquad \mathcal M_p(C_2(p))=O(\log p)=o((\log p)^2). After the coefficient-negligible C_2(p) sector is removed, every remaining carrier satisfies k>2C_2(p), \qquad |c_p(k)|>C_2(p), and the weight \lfloor N/a\rfloor with a>k assumes only O((\log p)^2) positive values.

Proof. The two mass estimates follow from (75). Since |c_p(k)|\le k/2, every carrier outside the C_2(p) sector has k>2C_2(p). Its largest positive floor value is at most \left\lfloor\frac{N}{k+1}\right\rfloor \ll(\log p)^2. ◻

Large local spikes and global asymptotic weight are different. Fixed-displacement square locks can violate a uniform pointwise bound, but their total normalized collar mass is o(1). They cannot carry the quadratic logarithmic secondary term. More generally, the entire C_2(p) sector is too small at the coefficient scale. Only the large-displacement, short-cadence sector isolated by Corollary 29 survives this absolute localization. Theorem 27 controls that sector after all carrier phases are assembled.

The floor-state switching ledger

Square-lock localization identifies where the residual mass can live. The floor machine identifies how that mass moves. Fix a set of carriers \mathcal L\subseteq\{2,\ldots,N-1\}. Set B_{p,\mathcal L}(1)=0. For 2\le X\le N, define the state burden and its fresh increment by \begin{aligned} B_{p,\mathcal L}(X) &=\sum_{\substack{k\in\mathcal L\\k<X}}\Delta_{k,p}(X), \\ F_{p,\mathcal L}(X) &=B_{p,\mathcal L}(X)-B_{p,\mathcal L}(X-1). \end{aligned} Put H_X=\left\lfloor\frac NX\right\rfloor, \qquad d_X=H_X-H_{X+1}. The integer d_X counts the positive integers h for which \lfloor N/h\rfloor=X. The integer H_X counts those for which \lfloor N/h\rfloor\ge X.

Proposition 30 (Exact switching ledger). For every carrier set \mathcal L, one has F_{p,\mathcal L}(X) =\sum_{\substack{k\in\mathcal L\\2\le k<X\\(k,X)=1}} s\bigl(p^2\overline X^{(k)},k\bigr). Moreover, \boxed{ \sum_{k\in\mathcal L}\mathfrak R_p(k) =\sum_{X=3}^{N}d_XB_{p,\mathcal L}(X) =\sum_{X=3}^{N}H_XF_{p,\mathcal L}(X). } Thus every fresh collision increment at X is counted once for every sampled floor state at or above X.

There is also an exact quadratic ledger. Suppress the subscripts and put Q_{p,\mathcal L}=\sum_{X=3}^{N}d_XB(X)^2. Then \boxed{ Q_{p,\mathcal L} =\sum_{X=3}^{N}H_X \bigl(F(X)^2+2B(X-1)F(X)\bigr). } If J_X=B(X-1)F(X)+\frac12F(X)^2, \qquad J_X^{\pm}=\max\{\pm J_X,0\}, then \boxed{ \sum_{X=3}^{N}H_XJ_X^+ =\sum_{X=3}^{N}H_XJ_X^-+\frac12Q_{p,\mathcal L}. } A negative J_X occurs exactly when the fresh increment opposes the stored burden and has magnitude less than twice the magnitude of that burden.

Proof. The complete reduced-residue sum in every carrier is zero. We may therefore adjoin the vanished term \Delta_{X-1,p}(X-1) when X-1\in\mathcal L. The difference between the two consecutive prefixes is the new term at X when (k,X)=1, and is zero otherwise. This proves (82).

In the collar decomposition, the value X=\lfloor N/h\rfloor occurs exactly d_X times. Regrouping first by X gives the first equality in (85). Discrete summation by parts, using B(2)=0 and H_{N+1}=0, gives the second.

Apply the same summation by parts to B(X)^2 and use B(X)^2-B(X-1)^2=F(X)^2+2B(X-1)F(X). This proves (86). The definition of J_X makes the last display equal to 2J_X. Splitting its positive and negative parts gives (87). The final criterion follows by factoring J_X=F(X)(2B(X-1)+F(X))/2. ◻

For p\ge5, the energy ledger also separates coherent signed area from fluctuation. Put D=H_3, \qquad \overline B=\frac1D\sum_{X=3}^{N}d_XB(X). Then \frac1D\left(\sum_{X=3}^{N}d_XB(X)\right)^2 +\sum_{X=3}^{N}d_X\bigl(B(X)-\overline B\bigr)^2 =Q_{p,\mathcal L}. This is the exact point at which a tall local spike and a persistent signed area become different quantities.

Three-term reciprocity rewrites the full fresh sum on the prime-square carrier. Let \begin{aligned} \mathcal H^{\ast}(X) &=\sum_{\substack{1\le k<X\\(k,X)=1}}\frac1k, \\ \mathcal P_p(X) &=\sum_{\substack{1\le k<X/2\\(k,X)=1}} s\bigl(k\overline{(X-k)}^{(m)},m\bigr). \end{aligned} The packet \mathcal P_p(X) lies on the anti-diagonal k+(X-k)=X of the prime-square Dedekind table.

Proposition 31 (Complementary-partition identity). Let \mathcal L contain every carrier from 2 through N-1, and abbreviate F_{p,\mathcal L}(X) to F_p(X). For 3\le X\le N, \boxed{ F_p(X) =\frac{m}{12X}\mathcal H^{\ast}(X) +\frac{X\mathcal H^{\ast}(X)-\varphi(X)}{12m} -\frac{\varphi(X)}8 -\mathcal P_p(X). }

Proof. The units k modulo X split into pairs k and l=X-k. There is no fixed point for X\ge3. The harmless modulus-one term allows the pair containing k=1 to be included. Inversion symmetry and three-term reciprocity give \begin{aligned} &s\bigl(m\overline X^{(k)},k\bigr) +s\bigl(m\overline X^{(l)},l\bigr)\\ &\qquad= \frac{m}{12kl} +\frac1{12m}\left(\frac kl+\frac lk\right) -\frac14 -s\bigl(k\overline l^{(m)},m\bigr). \end{aligned} Choose the representative k<l in every pair. The elementary sums are \sum_{k<l}\frac1{kl}=\frac{\mathcal H^{\ast}(X)}X, \qquad \sum_{k<l}\left(\frac kl+\frac lk\right) =X\mathcal H^{\ast}(X)-\varphi(X). There are \varphi(X)/2 pairs. Summing the reciprocity identity proves (93). ◻

The anti-diagonal packet is the primitive part of an additive convolution. For integers a,b coprime to m, put D_m(a,b) =\sum_{y\bmod m}\mathop{\mathrm{saw}}(ay/m)\mathop{\mathrm{saw}}(by/m). Changing variables gives D_m(a,b)=s\bigl(a\overline b^{(m)},m\bigr).

Corollary 32 (Primitive additive convolution). Set C_m(1)=A_m(1)=0. For 2\le X\le N, define \begin{aligned} C_m(X) &=\sum_{\substack{1\le k<X\\(k,X)=1}}D_m(k,X-k),\\ A_m(X) &=\sum_{1\le k<X}D_m(k,X-k). \end{aligned} For 3\le X\le N, \boxed{ C_m(X)=2\mathcal P_p(X). } For every 1\le X\le N, \boxed{ A_m(X)=\sum_{d\mid X}C_m(d), \qquad C_m(X)=\sum_{d\mid X}\mu(d)A_m(X/d). } Consequently, (93) expresses F_p(X) as an explicit elementary term minus one half of the primitive additive convolution C_m(X).

Proof. Symmetry of D_m pairs k with X-k, proving (95). Group the terms of A_m(X) by g=(k,X). Since every divisor of X is a unit modulo m, simultaneous multiplication by g leaves D_m unchanged. The group with X/g=d is therefore C_m(d). Summing over divisors gives the first identity in (96), and Möbius inversion gives the second. ◻

The lifetime-weighted packets combine into one triangular sum. Put \mathcal T_m(N) =\sum_{\substack{a,b\ge1\\a+b\le N}}D_m(a,b) and let \mathcal E_p(X) =\frac{m}{12X}\mathcal H^{\ast}(X) +\frac{X\mathcal H^{\ast}(X)-\varphi(X)}{12m} -\frac{\varphi(X)}8. Then \boxed{ \sum_{k=2}^{N-1}\mathfrak R_p(k) =\sum_{X=3}^{N}H_X\mathcal E_p(X) -\frac12\mathcal T_m(N) +\frac{H_2}{2}s(1,m). } Indeed, \sum_{X=2}^{N}H_XC_m(X) =\sum_{n=2}^{N}A_m(n) =\mathcal T_m(N), because H_X counts the multiples of X not exceeding N. Now use C_m(2)=D_m(1,1)=s(1,m), Proposition 30, and (93).

One state X is one integer, and k runs through its two-part divisions X=k+(X-k). The kernel D_m measures the overlap of the two parts on the ambient prime-square carrier. Möbius inversion removes the common scale shared by the two parts and leaves their primitive divisions. Formula (97) then places the whole weighted fresh sum in one triangular overlap packet.

The difference between the explicit partition term and this prime-square packet, together with the displayed depth-two correction, is another exact coordinate for the residual. Reciprocity fixes its sign geometry. Complementary members map to inverse prime-square ratios, and Dedekind sums are invariant under inversion. They reinforce rather than cancel.

The half packet contains no opposite prime-square partner supplied by oddness. If two terms in (92) had opposite prime-square ratios, then m\mid X(k+k')-2kk'. For 1\le k,k'<X/2, the integer on the right lies strictly between 0 and X^2<m. This is impossible. The complementary-partition identity exposes the aggregate controlled by Theorem 27. The damping enters through the additive modes of the descended carriers.

One carrier can retain a single sign even at large displacement. At p=101 and k=14, the square-lock displacement is -5, while C_2(101)=4. The six floor levels have endpoint residues 2,\ 8,\ 5,\ 11,\ 6,\ 2 and exact collar values -\frac3{14},\ -\frac{13}{14},\ -\frac{13}{14},\ -\frac3{14},\ -\frac{13}{14},\ -\frac3{14}. They are all negative and sum to -24/7. Thus no universal carrierwise sign-reversal argument exists. The coefficient theorem assembles the additive modes within each carrier before taking absolute values, then controls the resulting divisor costs on average across carrier moduli.

The live sector and sharper targets

Let \mathcal L_p be the large-displacement sector \mathcal L_p =\{2\le k\le N-1:|c_p(k)|>C_2(p)\}. The all-prime damping theorem and square-lock localization give a proved common-state estimate on this sector.

Proposition 33 (Live-sector damping). \boxed{ \sum_{X=3}^{N}H_XF_{p,\mathcal L_p}(X) \ll\frac{p^2(\log(2p))^2}{\log\log(3p)}. }

Proof. The full carrier family has the exact complementary-partition coordinate. The difference between its fresh sum and the live fresh sum is \sum_{X=3}^{N}H_X \bigl(F_p(X)-F_{p,\mathcal L_p}(X)\bigr) =\sum_{\substack{2\le k\le N-1\\|c_p(k)|\le C_2(p)}} \mathfrak R_p(k) =O(p^2\log p). The first equality is exact, and the last bound follows from Corollary 29. Apply Theorem 27 to the full sum. ◻

This is lifetime-weighted cancellation of fresh collision increments. No pointwise bound on spike height is required. The next questions concern a stronger scale.

A stronger local estimate would remove a full logarithm from the sampling bound.

Problem 34 (Signed carrier damping). Prove, uniformly for k\ge2, units \beta\bmod k, and integer intervals I, \boxed{ \left| \sum_{\substack{a\in I\\(a,k)=1}} s(\beta\overline a,k) \right| \ll k\log(2k). }

Theorem 35 (Sampling theorem conditional on signed damping). If (99) holds, then \boxed{ \mathcal{S}_p=O(\log p) } uniformly over odd primes p.

Proof. The same Abel summation as above now bounds the k-slice by O(p\log(2k)). Summing over k<p yields R_p^{(k)}=O(p^2\log p). Apply (43). ◻

There is an equally native jump-correlation criterion. Put V_{k,s}(c) =\sum_{a\in(\mathbb{Z}/k\mathbb{Z})^\times} T_{k,s}(a)\mathop{\mathrm{saw}}(ca/k). The alias formula uses V_{k,s_k}(m\bmod k).

Proposition 36 (Cyclic correlation ladder). The following implications hold.

  1. If |V_{k,s}(c)|\ll k(\log(2k))^2 uniformly, then \mathcal{S}_p\ll(\log p)^2.

  2. If |V_{k,s}(c)|\ll k\log(2k) uniformly, then \mathcal{S}_p=O(\log p).

  3. If |V_{k,s}(c)|\ll k uniformly, then \mathcal{S}_p=O(1).

Proof. Use (28), Proposition 14, and M_k\le N/k. The three assumptions give respectively \mathcal{A}_p\ll N^2(\log N)^2, \mathcal{A}_p\ll N^2\log N, and \mathcal{A}_p\ll N^2. ◻

Remark 37 (Limit conjecture). The data suggest the substantially sharper statement \boxed{\mathcal{S}_p\longrightarrow1.} By (43), this is equivalent to R_p^{(k)}=o(p^2). Boundedness of the sampling defect is equivalent to R_p^{(k)}=O(p^2). Neither assertion is proved here.

The finite secondary term

The exact decomposition is \frac{E_p}{p^2}=\mathcal{J}(p-1)-\mathcal{S}_p. Combining it with (14) and Theorem 27 closes the finite transfer.

Corollary 38 (Finite secondary coefficient). As p\to\infty through the odd primes, \boxed{ E_p =p^3-\frac1{\pi^2}p^2(\log p)^2 +O\left( \frac{p^2(\log(2p))^2}{\log\log(3p)} \right). } Under the signed damping hypothesis (99), one has the stronger sampling bound \mathcal{S}_p=O(\log p) and hence E_p =p^3-\frac1{\pi^2}p^2(\log p)^2 +O\bigl(p^2(\log p)^{5/3}\bigr). If the continuous-capacity remainder is separately sharpened to O(\log p), then the same signed damping theorem gives the sharper error O(p^2\log p).

Proof. Insert (14) into (104). The differences between p and p-1, and between \log p and \log(p-1), are lower order at the displayed scale. Equation (64) supplies the stated error. ◻

The coefficient 1/\pi^2 enters through continuous capacity. The inverse-residue theorem proves that finite carrier sampling preserves it.

Exact computation

The continuous capacity is evaluated by the Jordan-totient identity 3\mathcal{J}(N) =\sum_{d\le N}\frac{J_2(d)}{d^2} H_{\lfloor N/d\rfloor}^{\,2}. The energy is computed by exact integer floor sums. The values below were obtained with nfield [11]. At 26 declared prime checkpoints from 3 through 10007, nfield evaluates the continuous capacity and the finite energy independently. At the 21 declared checkpoints through 2003, it also evaluates both exact bridge formulas. Their discrepancies are below 2\mathbin{\cdot}10^{-15} in long-double arithmetic. The tabulated values provide independent finite checks of the formulas.

Sampling defect, regulated endpoint reservoir, and their signed difference.
p \mathcal{S}_p (p-1)^2/p^2 difference
31 0.951176 0.936524 0.014652
101 1.026972 0.980296 0.046676
199 1.058311 0.989975 0.068336
401 1.011478 0.995019 0.016459
809 1.013187 0.997529 0.015658
1201 0.992167 0.998335 -0.006169
2003 1.010416 0.999002 0.011414
3001 1.008431 0.999334 0.009097
4001 0.999951 0.999500 0.000451
5003 1.002754 0.999600 0.003154
7001 1.006215 0.999714 0.006501
10007 1.008630 0.999800 0.008830

The selected values in Table 1 show the arithmetic oscillation. They are consistent with bounded sampling and the endpoint limit (103). Both assertions remain open.

Relation to the endpoint reservoir

The finite secondary correction also has a lifted endpoint-reservoir form. Its quotient-shell and visible-minus-shadow coordinates lead to the same surviving lower-modulus carrier. Theorem 16 fixes its relation to the Bernoulli normal form. \boxed{ \begin{aligned} \text{carrier-grid alias} &=\text{regulated endpoint term}-\frac{R_p^{(k)}}{p^2}\\ &\quad+\text{explicit curvature}. \end{aligned} } Quotient-shell damping is a second coordinate for the same signed sum. The inverse-residue argument proves enough damping to preserve the coefficient. The quotient-shell coordinate remains useful for sharper estimates because it records where the cancellation occurs rather than only its final size.

This also explains the observed scale separation. The endpoint term tends to 1. The algebraic curvature is O(\log p/p^2) in the reciprocity coordinates. Any additional quadratic logarithm had to come from R_p^{(k)}. Theorem 27 rules that out by saving a factor of \log\log p. The stronger signed damping problem asks whether a full logarithm can be removed.

The coefficient survives the carrier grid

The finite sampling defect has two exact descriptions. The first is a collision-native correlation between cyclic boundary prefixes and carrier-grid jump phases. The second is the surviving lower-modulus Dedekind reservoir from Rademacher reciprocity. Each description includes its own explicit endpoint and curvature terms. Together they identify the same signed obstruction.

Opening the Dedekind carrier additively reveals the decisive cancellation. Every nonunit phase descends to its true modulus. At that modulus the centered inverse-residue series saves a factor of \log\log p. The divisor burden of lifting back through the carrier grid has bounded mean, so it cannot restore the lost factor. This proves \mathcal{S}_p=o((\log p)^2) for every odd prime. The continuous coefficient 1/\pi^2 therefore survives unchanged in the finite collision energy.

The square-lock and floor-state coordinates supply the local picture behind the global estimate. Small displacement carries too little absolute mass to reach the coefficient scale. Large displacement can produce a sustained one-sided spike inside a single carrier, so carrierwise sign reversal is unavailable. The required saving comes from additive cancellation within each descended carrier, together with the bounded mean of the divisor-lift cost. The spikes remain real local features, but the sum of the resulting carrierwise bounds is too small to alter the secondary term.

The finite secondary coefficient is settled. An O(\log p) sampling bound would remove a full logarithm, an O(1) bound would reach the endpoint scale, and the observed limit \mathcal{S}_p\to1 remains the strongest target. Those questions begin below the coefficient proved here. The surviving correction is -\pi^{-2}p^2(\log p)^2. This is The Secondary Term in Digit-Collision Energy.

References

[1]J. B. Conrey, E. Fransen, R. Klein, and C. Scott, Mean values of Dedekind sums, J. Number Theory 56 (1996), no. 2, 214–226.

[2]P. Flajolet and B. Salvy, Euler sums and contour integral representations, Experiment. Math. 7 (1998), no. 1, 15–35.

[3]T. Hilberdink, F. Luca, and L. Tóth, On certain sums concerning the gcd's and lcm's of k positive integers, Int. J. Number Theory 16 (2020), no. 1, 77–90, https://doi.org/10.1142/S1793042120500049.

[4]H. Iwaniec and E. Kowalski, Analytic Number Theory, American Mathematical Society Colloquium Publications 53, American Mathematical Society, Providence, RI, 2004.

[5]M. A. Korolev, Incomplete Kloosterman sums and their applications, Izv. Math. 64 (2000), no. 6, 1129–1152, https://doi.org/10.1070/IM2000v064n06ABEH000311.

[6]R. J. Lemke Oliver and K. Soundararajan, The distribution of consecutive prime biases and sums of sawtooth random variables, Math. Proc. Cambridge Philos. Soc. 168 (2020), no. 1, 149–169, https://doi.org/10.1017/S0305004118000592.

[7]H. Rademacher and E. Grosswald, Dedekind Sums, Carus Mathematical Monographs 16, Mathematical Association of America, Washington, DC, 1972.

[8]B. Saffari and R. C. Vaughan, On the fractional parts of x/n and related sequences. II, Ann. Inst. Fourier (Grenoble) 27 (1977), no. 2, 1–30, https://doi.org/10.5802/aif.649.

[9]I. Vardi, A relation between Dedekind sums and Kloosterman sums, Duke Math. J. 55 (1987), no. 1, 189–197.

[10]E. C. Titchmarsh, The Theory of the Riemann Zeta-Function, 2nd ed., revised by D. R. Heath-Brown, Oxford University Press, 1986.

[11]A. S. Petty, nfield, software repository, https://github.com/alexspetty/nfield.