\chapter{Introduction}\label{sec:introduction} \section{Background}\label{sec:intro-background} We consider a problem that arises in the study of higher-order influence function estimators \cite{RobinsEtAl2008} and semiparametric minimax lower bounds \cite{RobinsEtAl2009EJS} for nonlinear functionals. We observe independent copies of $(X,Y)$, where $X\in[0,1]^d$ has an unknown density $p$ bounded above and away from zero, and let \[ f(x)=\E(Y\mid X=x),\qquad \Var(Y\mid X=x)=V. \] The regression function belongs to a fixed $s$-H\"older ball. Let $Q_x$ denote the conditional law of $\varepsilon=Y-f(X)$ given $X=x$. It has mean zero and variance $V$, and its fourth moment is uniformly bounded, but it may otherwise depend on $x$. No smoothness is imposed on $p$. We aim to estimate $V$, which satisfies \begin{equation}\label{eq:variance-quadratic-identity} V=\E(Y^2)-\int_{[0,1]^d}f(x)^2p(x)\,\mathrm dx. \end{equation} Under the uniform H\"older-norm and conditional fourth-moment bounds, the empirical mean of $Y^2$ estimates $\E(Y^2)$ at the root-$n$ rate. Thus the only potentially non-root-$n$ component in \eqref{eq:variance-quadratic-identity} is the quadratic regression functional. For $04s$, Robins et al.\ \cite[Section~4]{RobinsEtAl2008} give a close-pair construction with convergence rate $n^{-4s/(d+4s)}$, which uses no smoothness of $p$. They divide $[0,1]^d$ into small cubes, select one pair of observations in each cube that contains at least two, and use the identity \[ \E\left[\frac{(Y_i-Y_j)^2}{2}\,\middle|\,X_i=x,X_j=y\right] =V+\frac{(f(x)-f(y))^2}{2}. \] By this identity, it is natural to average the squared differences over pairs whose covariates are within distance $h$. The expected number of such pairs is of order $n^2h^d$. Averaging the corresponding squared differences gives a stochastic fluctuation bounded in root mean square by $C\{n^{-1/2}+(n^2h^d)^{-1/2}\}$, while $s$-H\"older smoothness gives bias $O(h^{2s})$. Balancing the bias with the pair term gives $h\asymp n^{-2/(d+4s)}$; since $d>4s$, the root-$n$ term is smaller, and the stated rate follows. Robins et al.\ introduced their construction to show that a random design can give a faster rate than an equally spaced design even when the design density has no smoothness \cite[p.~338]{RobinsEtAl2008}. For $s$-H\"older regression on an equally spaced $d$-dimensional grid, the corresponding minimax rate for the root-mean-square risk is $n^{-2s/d}\vee n^{-1/2}$. Robins et al.\ attribute the rate $n^{-2s/d}$, for $s1$. For \(s>1\), the same lower bound follows by using rescaled \(C^\infty\) bumps that equal one at the grid points and are supported in balls of radius less than half the grid spacing. Taking heights \(cn^{-s/d}\) with a sufficiently small fixed \(c>0\) gives uniformly bounded \(C^s\) norms, while leaving the prescribed values at the design points, and hence the testing calculation, unchanged.} Richardson and Rotnitzky \cite[Section~6]{RichardsonRotnitzky2014} record the following question about the differentiable case, attributed to James Robins: when $s>1$ but $s/d<1/4$, does there still exist an estimator attaining $n^{-4s/(d+4s)}$? They write that analogy with other nonparametric estimation problems would suggest that the answer is ``yes''. As they also note \cite[Section~6, footnote~50]{RichardsonRotnitzky2014}, the squared-difference estimator of Robins et al.\ does not attain this rate for $s>1$, because it does not exploit the differentiability of the regression function. For $s>1$, its worst-case bias is generally of order $h^2$, rather than $h^{2s}$. This shows only that the estimator fails to exploit the additional smoothness; it does not rule out that another estimator may attain a better rate. Their question concerns rates of convergence; we study its minimax formulation, which requires a uniform bound on the root-mean-square risk over the model class specified below. For $s>1$ and $d>4s$, we determine the minimax root-mean-square risk up to a power of $\log n$. Let \begin{equation}\label{eq:lambda} \lambda=\frac{2(s+1)}{d+4}. \end{equation} For all sufficiently large $n$, the minimax root-mean-square risk lies between \(c\Psi_n\) and \(C\Psi_n(\log n)^{\Gamma}\), where \[ \Psi_n=n^{-\lambda}e^{-\kappa\sqrt{\log n}}(\log n)^{(s-1)/(d+4)} \] and the constants $\kappa>0$ and $\Gamma>0$ are given explicitly in \cref{thm:main}. These bounds identify $\lambda$ as the minimax root-mean-square exponent. The minimax risk is smaller than $n^{-\lambda}$ by the factor $e^{-\kappa\sqrt{\log n}}$, up to powers of $\log n$, with the same constant $\kappa$ in the upper and lower bounds, and the ratio of the upper and lower bounds is at most $C(\log n)^\Gamma$. Thus the rate \(n^{-4s/(d+4s)}\) is not uniformly attainable over the stated model class, which gives a negative answer to the minimax formulation of Robins's question. The upper bound also improves on the fixed-design minimax rate. For $00$ and $d\le4s$, ordinary least-squares residuals on a fixed partition attain the parametric rate $n^{-1/2}$. Some of these results appear in earlier versions of this paper, and related results were obtained in concurrent work; we describe both in \cref{sec:intro-related}. \section{Main results}\label{sec:main-result} Fix \(s>0\) and an integer \(d\ge 1\). Fix an open set \(\mathcal U\subset\mathbb R^d\) containing \([0,1]^d\), and constants \begin{equation}\label{e:fix-constants} 00,\qquad 00\), and \(Q=(Q_x)_{x\in[0,1]^d}\) is a measurable probability kernel from \([0,1]^d\) to \(\mathbb R\). Define the regression error by \(\varepsilon_i=Y_i-f(X_i)\). \begin{enumerate} \item[\textup{(i)}] Each \(X_i\) has Lebesgue density \(p\) on \([0,1]^d\), and \(p\) satisfies \[ p_-\le p(x)\le p_+ \quad\text{for Lebesgue-a.e. }x\in[0,1]^d. \] \item[\textup{(ii)}] The function \(f\) has an extension \(\widetilde f\) to \(\mathcal U\) such that \[ \norm{\widetilde f}_{C^s(\mathcal U)}\le L_s, \] where the norm is defined in \eqref{eq:holder-norm}. \item[\textup{(iii)}] For \(p(x)\,\mathrm dx\)-almost every \(x\in[0,1]^d\), the conditional distribution of \(\varepsilon_i\) given \(X_i=x\) is \(Q_x\), and \[ \int_{\mathbb R}u\,Q_x(\mathrm du)=0,\qquad \int_{\mathbb R}u^2\,Q_x(\mathrm du)=V,\qquad \int_{\mathbb R}u^4\,Q_x(\mathrm du)\le C_4, \] with \[ v_-\le V\le v_+. \] The conditional variance \(V\) is constant in \(x\), but the conditional distribution \(Q_x\) may otherwise depend arbitrarily on \(x\). \end{enumerate} Since \(V^2\le C_4\) by conditional Jensen's inequality, the effective variance interval is \[ \bigl[v_-,\min\{v_+,\sqrt{C_4}\}\bigr]. \] The assumption \(v_-^21$ and let $d>4s$ be an integer. Let $\lambda$ be as in \eqref{eq:lambda} and $\ell=\lceil s\rceil-1$ as in \eqref{eq:holder-indices}, and put \begin{equation}\label{eq:main-constants} \begin{gathered} \tau=\log\frac{\sqrt{p_+}+\sqrt{p_-}}{\sqrt{p_+}-\sqrt{p_-}}, \qquad \kappa=4\sqrt{\frac{2\tau(s-1)(d-4s)}{(d+4)^3}},\\ q_s=\binom{d+\ell}{d}, \qquad \Gamma=\frac{d\bigl(q_s-\frac12\bigr)}{d+4}. \end{gathered} \end{equation} There exist $c,C>0$ such that, for all sufficiently large $n$, \begin{equation}\label{eq:main-bracket} c\,\Psi_n \le\mathfrak r_n\le C\,\Psi_n(\log n)^{\Gamma}, \qquad \Psi_n=n^{-\lambda}e^{-\kappa\sqrt{\log n}} (\log n)^{\frac{s-1}{d+4}}. \end{equation} \end{theorem} The upper bound in \eqref{eq:main-bracket} is \cref{thm:upper-sharp}, proved in \cref{sec:upper-sharp}, and the lower bound is \cref{thm:lower}, proved in \cref{sec:lower}. The two bounds in \eqref{eq:main-bracket} differ by the factor $(\log n)^\Gamma$; we do not know the correct power of $\log n$. The estimator of \cref{sec:upper-sharp} is nearly minimax in the sense that, for all sufficiently large $n$, its root-mean-square risk is within a factor $C(\log n)^\Gamma$ of the minimax risk. The bounds agree in the exponent $\lambda$ and in the constant $\kappa$ of the stretched-exponential factor. In particular, \[ \log\mathfrak r_n=-\lambda\log n-\kappa\sqrt{\log n}+O(\log\log n). \] The constants $\tau$ and $\kappa$ depend on the density bounds only through the ratio $p_+/p_-$. Here $q_s$ is the dimension of the space of polynomials of degree at most $\ell$ in $d$ variables. In \cref{thm:main}, the constants $c$ and $C$ and the threshold on $n$ depend only on $(s,d,p_-,p_+,L_s,v_-,v_+,C_4)$, and not on $\mathcal U$. Neither proof gives an explicit threshold on $n$, and apart from $\kappa$ we have made no attempt to optimize the constants. The upper-bound proof requires $\sqrt{\log n}$ eventually to dominate fixed multiples of $\log\log n$, with constants depending on $q_s$. Our analysis is asymptotic and does not establish the estimator's performance at moderate sample sizes. These bounds determine the minimax exponent. We state this as a corollary. \begin{corollary}\label{thm:exponent} Let $s>1$ and let $d>4s$ be an integer. For every $\epsilon>0$, there exist $c_\epsilon,C>0$ such that, for all sufficiently large $n$, \[ c_\epsilon n^{-\lambda-\epsilon}\le\mathfrak r_n\le Cn^{-\lambda}. \] The minimax root-mean-square exponent satisfies \[ \lim_{n\to\infty}\frac{-\log\mathfrak r_n}{\log n} =\frac{2(s+1)}{d+4}. \] \end{corollary} \begin{proof} Since $e^{-\kappa\sqrt{\log n}}(\log n)^{(s-1)/(d+4)+\Gamma}\to0$ as $n\to\infty$, the upper bound in \eqref{eq:main-bracket} implies that $\mathfrak r_n\le Cn^{-\lambda}$ for all sufficiently large $n$. For every $\epsilon>0$, we have $e^{-\kappa\sqrt{\log n}}(\log n)^{(s-1)/(d+4)}\ge n^{-\epsilon}$ for all sufficiently large $n$, and then the lower bound in \eqref{eq:main-bracket} gives $\mathfrak r_n\ge cn^{-\lambda-\epsilon}$ for all sufficiently large $n$. Taking logarithms in these two bounds, dividing by $\log n$, and letting first $n\to\infty$ and then $\epsilon\downarrow0$, we obtain the limit. \end{proof} The separation from the exponent in the question recorded in \cite[Section~6]{RichardsonRotnitzky2014} is \begin{equation}\label{eq:intro-gap} \frac{4s}{d+4s}-\lambda =\frac{2(s-1)(d-4s)}{(d+4s)(d+4)}>0. \end{equation} The lower bound rules out a uniform bound of order \(n^{-4s/(d+4s)}\) on the root-mean-square risk, even after multiplication by any fixed power of \(\log n\). The exponent also satisfies \[ \lambda-\frac{2s}{d}=\frac{2(d-4s)}{d(d+4)}>0, \] which means that the upper bound is faster than the fixed-design minimax rate. In the other direction, the minimax risk is not of the exact order $n^{-\lambda}$, since the upper bound in \eqref{eq:main-bracket} implies that $n^{\lambda}\mathfrak r_n\to0$. In particular, the weaker bound $Cn^{-\lambda}$ recorded in \cref{sec:upper} is not the minimax rate. The bound $Cn^{-\lambda}$ was proved in the third arXiv version of this paper~\cite{AronowLopatto2026v3} and, in concurrent work, by Dobriban et al.~\cite[Proposition~3.4]{DobribanEtAl2026}; see \cref{sec:intro-related}. For \(01\); see \cref{sec:small-s}. \begin{theorem}\label{thm:main2} Let $00$ such that, for every $n\ge1$, \begin{equation}\label{eq:main-rate} C^{-1}\bigl(n^{-1/2}\vee n^{-4s/(d+4s)}\bigr) \le\mathfrak r_n\le C\bigl(n^{-1/2}\vee n^{-4s/(d+4s)}\bigr). \end{equation} \end{theorem} The residual-projection estimator of \cref{lem:projection-upper} attains the parametric rate when $d\le4s$. Together with the lower bound in \cref{sec:prelim-parametric}, it yields the following theorem. \begin{theorem}\label{thm:parametric} Let $s>1$ and let $d\le4s$ be a positive integer. There exists $C>0$ such that, for every $n\ge1$, \begin{equation}\label{eq:parametric-rate} C^{-1}n^{-1/2}\le\mathfrak r_n\le Cn^{-1/2}. \end{equation} \end{theorem} At $s=2$, the threshold $d\le4s$ of \cref{thm:parametric} has a precedent in~\cite[Section~3.1]{Spokoiny2002}. For regression functions with bounded second derivatives, i.i.d.\ Gaussian errors and an i.i.d.\ design with a continuous positive density on a compact support, root-$n$ accuracy is obtained there when $d\le8$, with bounds that hold conditionally on the design on events of probability arbitrarily close to one. \section{Ideas of the proofs}\label{sec:intro-ideas} \emph{The common approximation problem.} Both bounds are governed by the same problem of polynomial approximation. Let $\tau$ be as in \eqref{eq:main-constants}. On $[p_-,p_+]$, the reciprocal $1/z$ is approximated by polynomials of degree $m$ with uniform error of order $e^{-\tau m}$. This classical fact follows from the Chebyshev expansion of the reciprocal \cite[equation~(2.4)]{Mathar2006}, and the geometric factor $e^{-\tau}$ cannot be improved \cite[Corollary~2.2]{KrausEtAl2012}. Dually, a signed measure on $[p_-,p_+]$ that represents the value at the exterior point zero of every polynomial of degree at most $m$ has total variation at least of order $e^{\tau m}$, by the growth of Chebyshev polynomials outside the interval \cite[Section~2.7, equation~(2.37)]{Rivlin1990}, and the rule at the Chebyshev nodes attains this order (\cref{lem:finite-rules}). The upper bound uses the first statement and the lower bound uses the second, and this is the source of the common constant $\kappa$. \emph{The upper bound.} We modify the close-pair construction of Robins et al. For $s>1$, the bias of half the squared difference of two responses at distance $h$ comes from the regression increment $f(x')-f(x)$, which is of order $h$. We subtract a prediction of this increment from the difference of the responses, and we use two predictions computed from independent samples, one in each factor of the square. As in double cross-fitting \cite[Section~2]{NeweyRobins2018}, the mean of the product then contains the product of the two mean prediction errors, and the fluctuations of the predictions contribute no bias. As in \cite[Section~3]{McCleanEtAl2024}, their covariances across evaluation points remain in the variance, which we control by averaging over many pairs (\cref{prop:pair-risk}). If the predictions resolve $f$ at $K$ cells, the bias is of order $h^2K^{-2(s-1)/d}$, while the pair statistic fluctuates at order $n^{-1}h^{-d/2}$. At the sampling scale $K=n$, balancing these two terms gives the rate $n^{-\lambda}$. The two-scale estimator of Dobriban et al.\ \cite{DobribanEtAl2026}, which projects out the local polynomial trend in cells of diameter of order $n^{-1/d}$ and uses one close pair per retained cell, attains the same rate; see \cref{sec:intro-related}. To improve on $n^{-\lambda}$, we resolve $f$ at $K=ne^T$ cells with $T>0$ tending to infinity, so that most cells contain no observation and a local fit cannot be computed in them. As in Park~\cite[Section~3.2, equations~(9)--(14)]{Park2026} and in our companion paper with Kallus~\cite[Sections~4.1--4.3]{AronowKallusLopatto2026}, we replace the inverse of each local Gram matrix by a polynomial in that matrix. The population fit is then a polynomial in local moments of $p$ and $pf$. Each monomial in these moments can be estimated without bias by a U-statistic \cite{Hoeffding1948}, that is, by an average over distinct observations, even when most cells are empty; for cell counts, this is the falling-factorial estimate of powers of cell probabilities \cite[equation~(21)]{JiaoEtAl2015}. This allows resolution below the sampling scale (\cref{sec:U-factorial}). Related estimates of functionals of an inverse Gram or covariance matrix appear in the higher-order corrections of Robins et al., which use U-statistics of increasing order with inverse Gram matrices computed under initial nuisance estimates \cite[Theorems~3.14--3.15]{RobinsEtAl2008}, and in the polynomial estimates of Kong and Valiant \cite[Section~2.2]{KongValiant2018} and of Chen, Liu and Mukherjee \cite[Section~3.2]{ChenLiuMukherjee2024}, who estimate quadratic forms involving matrix powers by U-statistics. We express the prediction as a sum of increments of anchored local fits over nested dyadic cells (\cref{sec:U-anchored}). Since consecutive fits reproduce the same Taylor polynomial of $f$, each increment is small. To keep this Taylor cancellation, we clear the determinant of the parent Gram matrix before we approximate the remaining inverse (\cref{sec:U-inverse}). After a known preconditioning, each local Gram matrix has its spectrum in $[p_-,p_+]$, and an approximation of degree $m$ gives an increment error of relative order $e^{-\tau m}$, up to a factor polynomial in $m$ (\cref{lem:U5}). The variance of the unbiased lift grows factorially in the degree when the cells hold few observations, and this limits the usable degree. We therefore lower the degree as the resolution increases, along a quadratic profile in the logarithm of the number of cells per observation (\cref{sec:U-estimator}). With this profile, we may take $T$ of order $\sqrt{\log n}$, that is, a resolution gain of $e^{O(\sqrt{\log n})}$ cells per observation, and we gain the factor $e^{-2(s-1)T/(d+4)}$, which is $e^{-\kappa\sqrt{\log n}}$ up to powers of $\log n$ (\cref{sec:U-allocation}). Our companion paper \cite[Sections~4.1--4.4]{AronowKallusLopatto2026} and Park \cite[Sections~3.2--3.4]{Park2026} also keep the cancellation in projection increments over nested partitions before approximating the inverses of Gram matrices, and they also lower the degree on finer partitions. Here the increments are anchored at a design point, evaluated on close pairs, and normalized by an estimate of the noise coefficient. \emph{The lower bound.} We build on the ternary law of \cref{sec:prelim-ternary}, which already appears in the third version of this paper \cite{AronowLopatto2026v3}. Its probabilities are quadratic in the regression value. By \eqref{eq:LB-heat-one}, we see that a centered perturbation of $f$ with variance $\eta^2$ changes the law of one response exactly as an increase of $V$ by $\eta^2$ does. Hiding a change of the variance by randomizing the regression mean is the principle of the lower-bound proof in~\cite[Section~4.5]{Spokoiny2002}, which places Gaussian regression values on smooth bumps, and of the fixed-design lower bounds of Wang et al.\ \cite[Section~6.2]{WangBrownCaiLevine2008}, which match finitely many Gaussian moments. With the ternary law, only observations that share a local perturbation of the mean distinguish the two hypotheses. For $01$, the same balance would give only $n^{-4s/(d+4s)}$, which is smaller than $n^{-\lambda}$ by a power of $n$, and we must reduce the contribution of these pairs. The likelihood of the observations in a patch contains the product of the values of the design density at their positions. Robins et al.\ use priors in which the design density and the regression function vary together \cite[Section~3]{RobinsEtAl2009EJS}. Here this correlation has two roles. First, a joint update of the density and of $f$ acts on the observations in a patch through a function of their positions, provided that function has a separated representation (\cref{sec:interp}). As in the third version, we choose this function by cardinal interpolation, so that the covariance of the perturbation of $f$, evaluated at the observed points, is diagonal (\cref{lem:interp}); here count two uses a separate rule (\cref{sec:pair-new}). Second, the density factor selects the number of observations in a patch. We integrate the updates of the density against a signed measure that represents evaluation at zero, as in the dual problem above. Among counts at most $m$, an update then acts only on patches with a prescribed number $r\le m$ of observations; larger counts can give aliases. The leading dependence of exterior evaluation on the cutoff $m$ is $e^{\tau m}$; the full signed update also has factors that depend on the selected count, the degree of the derivative rule and the separated representation (\cref{lem:finite-rules,lem:local-operator}). Signed measures dual to polynomial approximation were used to construct priors with matching moments by Lepski, Nemirovski and Spokoiny \cite[Section~4.4]{LepskiEtAl1999}, and for the reciprocal by Wu and Yang \cite[Appendix~A]{WuYang2019}. The diagonal covariance imitates an increase of the variance at each observation, which we cancel by decreasing $V$. What remains are the defects of the interpolation, the remainders of the fixed Taylor rule and of the rule for count two, and the aliases at counts beyond the exactness range of the selectors (\cref{sec:LB-rows,sec:LB-alias}). These updates are signed. We realize them by a path of probability priors: we give polynomial weights to finite histories of updates, which we organize as heaps of pieces in the sense of Viennot \cite{Viennot1986}. Conditional annihilation \eqref{eq:LB-annihilation} keeps the law of the design density fixed along the path, and the centering of the reference marks (\cref{lem:centered-row}) makes the mean of its total mass equal to one; the method of bounded differences then concentrates this mass near one. The weights are positive by a deletion induction, which adapts the argument for hard-core partition functions credited to Dobrushin and presented by Scott and Sokal \cite[Section~5.1]{ScottSokal2005} (\cref{sec:LB-heaps}). A tilt by the $n$th power of the total mass of the density cancels the normalization of the likelihood at a fixed sample size. The third version Poissonized the sample instead and reweighted the priors to cancel the random mass in the Poisson likelihood \cite[Section~1.2, Step~3]{AronowLopatto2026v3}. Removing the last update computes the score of the path on every data observable, and the local cancellation makes this score small (\cref{prop:local}). Using the score-to-risk comparison of \cref{lem:path-risk}, we then obtain the lower bound (\cref{sec:LB-fisher}). This comparison is related to the information inequality \cite[Chapter~2, Theorem~5.10]{LehmannCasella1998} and to the comparison of mixtures of Cai and Low \cite[Section~2]{CaiLow2011}; we integrate the derivative of the mean of an estimator along the path and bound the two endpoint biases. \emph{The common balance.} Both proofs balance the logarithm $x$ of the number of cells, or patches, per observation against the degree $m$ of the approximation, or the largest selected count. Relative to $n^{-\lambda}$, a larger $x$ improves the upper bound and weakens the lower bound by the factor $e^{-2(s-1)x/(d+4)}$, while the approximation problem enters through $e^{\tau m}$. In the upper bound, $mx$ must be at most a multiple of $\log n$, to leading order, to control the factorial variance of the unbiased lifts. In the lower bound, it must be at least a multiple of $\log n$, to leading order, to make patches with more than $m$ observations rare. In each bound, the optimal choice takes $x$ and $m$ of order $\sqrt{\log n}$, and in both bounds it gives the factor $e^{-\kappa\sqrt{\log n}}$ with the same constant $\kappa$. The third version already used a balance of this kind in its lower bound, with an unspecified constant. Park~\cite[Sections~2 and~3.4]{Park2026} and our companion paper~\cite[Section~1.3]{AronowKallusLopatto2026} use the same balance in both of their bounds. The constructions below explain these mechanisms where they are used, including the logarithmic corrections. \section{Related work}\label{sec:intro-related} \emph{Variance estimation in nonparametric regression.} Difference-based estimators of the variance go back to the first-difference estimator that Brown and Levine \cite[Section~1]{BrownLevine2007} attribute to von Neumann and, later, Rice \cite{Rice1984}. Gasser, Sroka and Jennen-Steinmetz \cite{GasserEtAl1986} estimate the residual variance from local linear pseudo-residuals, Hall, Kay and Titterington \cite{HallEtAl1990} construct difference sequences that are asymptotically optimal, and Dette, Munk and Wagner \cite{DetteEtAl1998} compare quadratic-form estimators and their sensitivity to the regression mean. Munk et al.\ \cite{MunkEtAl2005} study the estimation of the noise variance by differences in multivariate regression with fixed covariates, and Brown and Levine \cite{BrownLevine2007} and Wang et al.\ \cite{WangBrownCaiLevine2008} study the estimation of a variance function by differences under fixed design. Residual-based estimators form a second family. Hall and Marron \cite{HallMarron1990} estimate the variance from squared nonparametric residuals, and Fan and Yao \cite{FanYao1998} estimate a conditional variance function by local linear smoothing of squared residuals. The residual projection of \cref{lem:projection-upper} belongs to this family, and variance estimators based on pseudo-residuals that annihilate local linear trends were studied by Spokoiny \cite[Sections~2.1--2.2]{Spokoiny2002}. Under random design, M\"uller, Schick and Wefelmeyer \cite{MullerEtAl2003} estimate the error variance by a covariate-matched U-statistic of squared differences. Du and Schick \cite{DuSchick2009} give a covariate-matched version of the pseudo-residual estimator of Gasser et al., which attains the asymptotic performance of residual-based estimators under their conditions. The close-pair construction of Robins et al.\ recalled in \cref{sec:intro-background} does not use smoothness of the design density. Shen et al.\ \cite{ShenGaoWittenHan2020} study the estimation of the variance under random design, when the standardized errors have a common law independent of $X$. In the multivariate case, they estimate the variance by a ratio of two U-statistics of squared differences weighted by a product kernel, and they prove that the minimax mean squared error is of order $n^{-8\alpha/(4\alpha+d)}\vee n^{-1}$ when every coordinatewise H\"older index lies in $(0,1]$, where $\alpha$ is the harmonic mean of these indices \cite[Section~4.1]{ShenGaoWittenHan2020}. For isotropic smoothness $s\le1$, this is the square of the rate in \cref{thm:main2}, and the estimator of \cref{prop:elementary-upper} for $00$, Robins et al.\ conjecture that the rate in their equation~(1.3) is minimax, up to logarithmic factors, when $\beta_g$ is too small for their estimators to attain $n^{-4s/(d+4s)}$. For $\E\{\Var(Y\mid X)\}$, this rate tends to $n^{-2s/d}$, up to a logarithmic factor, as $\beta_g\to0$ \cite[p.~338]{RobinsEtAl2008}. \emph{Functional estimation under rough design.} For doubly robust functionals such as the expected conditional covariance, let $s$ be the average H\"older smoothness of the two nuisance functions. Robins et al.\ conjecture that their higher-order estimators, built from an initial estimate of the density that is uniformly consistent, attain the rate $n^{-2s/d}$ in probability, up to a power of $\log n$, when $s1$, when $d>4s$. For this model when $s>1$ and $d>4s$, with a design density of H\"older smoothness above an explicit threshold, Dobriban et al.\ show that a clipped higher-order influence function estimator of Robins et al.\ has mean squared error at most $Cn^{-8s/(d+4s)}$ if, in addition, the unclipped construction satisfies the uniform integrated bounds in their equation~(50) on squared conditional bias and conditional variance \cite[Theorem~2.1]{DobribanEtAl2026}. In our lower-bound priors for $s>1$, the design densities are updated patch by patch and can be discontinuous, so these priors are not supported on a fixed H\"older ball. Our lower bounds therefore do not apply to the H\"older-density classes of Dobriban et al., where their conditional bound is faster; they note this limitation for an earlier version of this paper \cite[Section~4]{DobribanEtAl2026}. \emph{Polynomial approximation and moment matching.} Lower bounds for nonsmooth functionals are often built from two priors whose moments agree up to some order, through the duality between moment matching and best polynomial approximation; see Lepski, Nemirovski and Spokoiny \cite[Section~4.4]{LepskiEtAl1999}, Cai and Low \cite{CaiLow2011} and Wu and Yang \cite{WuYang2019}. Wu and Yang use best polynomial approximation of the reciprocal on an interval bounded away from zero \cite[Appendix~A]{WuYang2019}. The corresponding estimators replace the functional by a polynomial and estimate its terms without bias, as in \cite[Section~4.1]{CaiLow2011} and \cite{JiaoEtAl2015}. In the present lower bound, the signed measures act on the values of the design density, and in the upper bound, polynomial approximation replaces the inverses of local Gram matrices. Related polynomial estimators of functionals of inverse covariance matrices appear in \cite{KongValiant2018,ChenLiuMukherjee2024}. \emph{Concurrent work.} Dobriban, Mukherjee, Robins and Wang \cite{DobribanEtAl2026} study the present model with the additional assumption that the design density is H\"older of some positive order. For $s>1$ and $d>4s$, they construct a two-scale estimator. It partitions $[0,1]^d$ into cells of diameter of order $n^{-1/d}$, projects out the local polynomial trend of degree $\lceil s\rceil-1$ in suitable cells, and averages the squared normalized contrasts of one close pair per cell. They prove that its mean squared error is at most $Cn^{-4(s+1)/(d+4)}$, that is, that its root-mean-square risk is at most $Cn^{-\lambda}$ \cite[Section~3.3, Proposition~3.4]{DobribanEtAl2026}. Their proof of this bound uses only the upper and lower bounds on the design density \cite[Appendix~A.5]{DobribanEtAl2026}. Their paper was posted on 8 September 2026, one day before the third version of this paper, which contained an estimator with the same bound. \Cref{thm:upper-sharp} improves the bound $Cn^{-\lambda}$ by the factor $e^{-\kappa\sqrt{\log n}}$, up to powers of $\log n$, and \cref{thm:lower} shows that the improved bound is optimal over $\Theta$ up to a power of $\log n$. Park \cite{Park2026}, in a paper posted on 4 October 2026, studies the expected conditional covariance $\E\{\Cov(A,Y\mid X)\}$ for responses bounded by one in absolute value, conditional means $\E(A\mid X)$ and $\E(Y\mid X)$ in H\"older classes of orders $\alpha$ and $\beta$, and an unknown design density between known bounds $01$ and $d>4s$, our exponent $\lambda$ exceeds $2s/d$ by $2(d-4s)/\{d(d+4)\}$, and our coefficient is $\kappa$. The difference reflects the constant conditional variance of the present model, which gives the identity for squared differences recalled in \cref{sec:intro-background}. Park also reports a lower bound for the average conditional variance of a single binary response \cite[Section~2]{Park2026}; in that model, the conditional variance may depend on $X$, and the lower bound does not transfer to the present model. The two works share the threshold $d=4s$, the constant $\tau$, and the identification of the coefficient of $\sqrt{\log n}$ in the upper and lower bounds. Both works use polynomial approximation on the density interval. In both, the approximation of the reciprocal controls the bias of the estimator, and in Park's lower bound it controls the separation; our lower bound uses signed rules for the evaluation of polynomials at an exterior point. Park's estimator starts from a double cross-fit average of products of local polynomial residuals and subtracts bias corrections on nested partitions. The corrections replace the inverses of the population Gram matrices by Chebyshev matrix polynomials, they keep the cancellation of the parent polynomial trends, they are estimated from distinct observations, and their degrees decrease on finer partitions \cite[Section~3]{Park2026}. Our estimator uses polynomial estimates of the regression increment inside a pair statistic, and it relies on the constant conditional variance. Park's main text describes his lower bound through two mixtures with fixed cube probabilities, which give the same distribution to the observations in each cube that contains at most $r+1$ observations, where $r\ge1$ is the number of matched moments, and whose targets differ by a signed reciprocal moment \cite[Section~2]{Park2026}. Our lower bound selects counts by signed local updates, which we realize by positive priors at a fixed sample size. Park cites the third version and describes its bounds, which have an unspecified constant in the stretched-exponential factor \cite[Sections~1 and~8]{Park2026}; the present version identifies this constant as $\kappa$ in both bounds. Our comparison is based on the main text of arXiv:2610.05006v1. Its proofs are deferred to a supplement that is not part of the arXiv PDF or source files, as checked on 8 October 2026. In our companion paper with Kallus \cite{AronowKallusLopatto2026}, we study missing-at-random means, treatment effects and expected conditional covariances in fully nonparametric models under rough design. For two nuisance functions with average H\"older smoothness $s$, we obtain the exponent $2s/d$ below $s=d/4$, and we identify the coefficient of $\sqrt{\log n}$ in the upper and lower bounds \cite[Theorem~2.3]{AronowKallusLopatto2026}. For the expected conditional covariance, the exponent and the coefficient agree with those of Park. The estimator for the main theorem of the companion also approximates the inverses of local Gram matrices by polynomials on the interval that contains their eigenvalues and estimates the polynomial terms without bias, but it estimates the target directly, without regression fits or close pairs of observations. The lower bound of the companion follows the strategy of the third version, with a Poissonized sample, but its priors give the design density a random phase, and the two mixtures are separated through a Fourier coefficient of the reciprocal of the density. \emph{Earlier versions of this paper.} Versions~1 and~2 of arXiv:2607.13170 \cite{AronowLopatto2026v1,AronowLopatto2026v2}, posted on 14 July and 17 August 2026 under the title ``On Rates Attainable under Random Design: A Negative Answer to a Problem of Robins'', proved that $\mathfrak r_n\ge cn^{-\beta}$ for all $n\ge1$ when $s>1$ and $d>4s$, where $\beta=\{d(3s+1)+8s\}/\{(d+2s)(d+4)\}$. Since $\beta<4s/(d+4s)$, this lower bound already shows that the rate in the question recorded by Richardson and Rotnitzky is not uniformly attainable. Version~2 also established the exact minimax rate of \cref{thm:main2} for $01$ and $d>4s$, with an unspecified constant $C$ in the stretched-exponential factor. These bounds identified the exponent $\lambda$, and version~3 also proved \cref{thm:parametric}. Its upper bound used an interpolation estimator built from corrected response differences, and its lower bound used a Poissonized sample and approximate matching of local moments. The present version identifies the constant $\kappa$ in both bounds, so that the upper bound improves to $C\Psi_n(\log n)^\Gamma$ and the two bounds differ only by the factor $(\log n)^\Gamma$. It gives different proofs of \cref{thm:main2,thm:parametric}, and its lower bound for $s>1$ uses signed local updates and positive priors at a fixed sample size. \section{Organization and conventions}\label{sec:intro-organization} \Cref{sec:upper} proves the parametric upper bounds by residual projection and the upper bounds for low smoothness by pair differences. Its common risk bound for pair statistics is also used in \cref{sec:upper-sharp}, which proves the refined upper bound for $s>1$ and $d>4s$. \Cref{sec:lower-prelim} collects the inputs common to the lower bounds: the ternary law, one score-to-risk comparison, and the smooth weights; it also contains the proof of the parametric lower bound. \Cref{sec:small-s} proves \cref{thm:main2}, introducing the local cancellation and score argument in its elementary form. \Cref{sec:lower} develops this argument into the lower bound in \cref{thm:main}, using interpolation, count selection and positive priors built from signed local updates. \Cref{thm:exponent} is proved immediately after its statement. All logarithms are natural. Constants may increase between occurrences. For a finite signed measure $\mu$, $\norm{\mu}$ is the total mass of $|\mu|$, that is, the total variation of $\mu$ without a factor $1/2$. The constants of \cref{thm:main} are recalled in \cref{thm:upper-sharp}. Apart from these and from the model of \cref{sec:main-result}, each of \cref{sec:upper,sec:upper-sharp,sec:small-s,sec:lower} fixes its own notation, and several letters, among them $h$, $m$, $q$, $M$, $N$, $\rho$, $\mu$ and $\Gamma$, are also used there with local meanings, declared where they occur. \chapter{Pair differences and elementary upper bounds} \label{sec:upper} Least-squares residuals directly give all of the parametric upper bounds. For low smoothness, we obtain the remaining upper bounds from half the squared difference of two nearby responses. We first record the resulting rates, which include a consequence of the upper-bound construction for higher smoothness. \begin{proposition}\label{prop:elementary-upper} Let $s>0$, $d\ge1$, and $t=\min\{s,1\}$. There exists a measurable estimator based on $n$ observations such that, for every $n\ge1$, \begin{equation}\label{eq:upper-simple} \sup_{\theta\in\Theta} \bigl\{\E_\theta(\widehat V_n-V)^2\bigr\}^{1/2} \le C\left(n^{-1/2}\vee n^{-2(s+t)/(d+4t)}\right). \end{equation} \end{proposition} For $s>1$, this proves the upper bound in \cref{thm:parametric} and the simpler bound $Cn^{-\lambda}$ when $d>4s$. For $01$ and $d>4s$, \cref{thm:upper-sharp} implies the bound $Cn^{-\lambda}$ for all sufficiently large $n$, since its stretched exponential dominates its logarithmic factor. The constant estimator $v_-$ covers the remaining sample sizes. We write $\mathbb O=[0,1]^d\times\mathbb R$ for the observation space, equipped with its Borel $\sigma$-field, and $\mu_\theta$ for the law of one observation under $\theta$. \section{Parametric upper bounds by residual projection} \label{sec:projection-upper} When we fit cellwise polynomials, a fixed fraction of the observations remains available as residual degrees of freedom. The approximation error is small even if the design points are irregular or the local fit is rank deficient. Related variance estimators based on normalized local linear pseudo-residuals are studied in~\cite[Sections~2.1--2.2]{Spokoiny2002}, for fixed design and conditionally on an i.i.d.\ random design. \begin{lemma}\label{lem:projection-upper} For every $s>0$ and $d\ge1$, there exists a Borel estimator satisfying \begin{equation}\label{eq:projection-risk} \sup_{\theta\in\Theta}\|\widehat V_n-V\|_{L^2} \le C\bigl(n^{-1/2}+n^{-2s/d}\bigr),\qquad n\ge1. \end{equation} In particular it attains $Cn^{-1/2}$ whenever $d\le4s$. \end{lemma} \begin{proof} Let $\ell=\lceil s\rceil-1$ and $q=\binom{d+\ell}{d}$. For $n\ge2q$, we set $k=\lfloor(n/(2q))^{1/d}\rfloor$ and partition $[0,1]^d$ into $K=k^d$ cubes, assigning their faces by a fixed convention. Then $K\asymp n$ and $qK\le n/2$. Let $B_X$ evaluate a fixed basis of the cellwise degree-$\ell$ polynomials at $X_1,\ldots,X_n$, and let $P_X$ be the orthogonal projection onto its column space. We define \[ A_X=I_n-P_X,\qquad d_X=\operatorname{tr}A_X\ge n-qK\ge n/2,\qquad \widetilde V_n=\frac{Y^\top A_XY}{d_X}. \] The projection is Borel even at changes of rank, since $P_X=\lim_{\epsilon\downarrow0} B_X(B_X^\top B_X+\epsilon I_{qK})^{-1}B_X^\top$. Let $f_X=(f(X_i))_{i\le n}$. On each cube, the Taylor polynomial centered at a fixed corner of that cube approximates $f$ uniformly to within $CK^{-s/d}$. For $\ell=0$, this is the H\"older bound; otherwise, it follows by integrating the order-$\ell$ H\"older remainder along a segment in the cube. Since these polynomials give a vector in the range of $P_X$, we have \[ \|A_Xf_X\|^2\le CnK^{-2s/d},\qquad \E(\widetilde V_n\mid X)-V=\frac{\|A_Xf_X\|^2}{d_X} \le CK^{-2s/d}. \] Conditional on $X$, the errors are independent and centered, with variance $V$ and fourth moments at most $C_4$. Expanding the quadratic form of a symmetric matrix $A$, we find \[ \Var(\varepsilon^\top A\varepsilon\mid X) =\sum_i A_{ii}^2\Var(\varepsilon_i^2\mid X) +4V^2\sum_{i0\}$, we denote points by $\mathsf w=(x,x')$ and let $M_h(d\mathsf w)=k_h(x,x')p(x)p(x')\,dx\,dx'$. Then \begin{equation}\label{eq:U7-mass} 1\le\int dM_h\le p_+,\qquad M_h\{(x,x'):x\in A\}\le p_+^2|A| \quad(A\subset[0,1)^d\text{ Borel}). \end{equation} Indeed, writing $p_I=\Pp(X\in I)$, we see that the total mass is $h^{-d}\sum_Ip_I^2$. By Cauchy--Schwarz and $\sum_Ip_I=1$, we obtain the lower bound, and $p_I\le p_+h^d$ gives the upper bound. The second assertion follows by restricting the first integration to $A$. A vector field $c(\mathsf w)=(1,c_2(\mathsf w),c_3(\mathsf w))^\top$ represents the affine pair residual $y'+c_2(\mathsf w)y+c_3(\mathsf w)$. We allow random coefficients to cover predicted increments; the constant field $(1,-1,0)^\top$ gives the ordinary response difference. For this field, the ratio $\widehat N/\widehat D$ of the statistics defined below is the average of half the squared response differences over pairs of observations in a common cube. M\"uller, Schick and Wefelmeyer \cite{MullerEtAl2003} estimate the variance by a covariate-matched U-statistic of this kind, and Shen et al.\ \cite[Section~4.1, equation~(19)]{ShenGaoWittenHan2020} use the same ratio form with a product kernel. Independent pilots are useful because the mean of their product contains the product of their mean prediction errors, without a pilot variance term. The fluctuations of the pilots still matter, but the test-function bound below controls them once many pairs are averaged. In the double cross-fitting of Newey and Robins \cite[Section~2]{NeweyRobins2018}, two nuisance functions are likewise estimated from independent samples, which removes the additional nonlinear bias that remains under single cross-fitting. McClean et al.\ \cite[Section~3, Lemma~1]{McCleanEtAl2024} bound the resulting bias by products of mean nuisance errors, and in their analysis the covariances between evaluations of the same nuisance estimates at different points remain in the variance. When the observed response at the first design point of the ordered pair is used in place of $f(x)$, the noise coefficient changes, and the denominator estimates exactly that normalization. \begin{proposition}[A common pair-risk bound]\label{prop:pair-risk} Let $c_{\mathcal A},c_{\mathcal B}$ be independent, identically distributed jointly Borel random vector fields on the pair set, with first coordinate one, independent of an evaluation sample of $m\ge2$ observations. Suppose $m\ge c_*n$ for a fixed $c_*>0$. Put $\bar c=\E c_{\mathcal A}$ and $\zeta_{\mathcal S}=c_{\mathcal S}-\bar c$ for $\mathcal S\in\{\mathcal A,\mathcal B\}$. Suppose that, uniformly over $\theta\in\Theta$, for deterministic $b,W\ge0$ and a fixed $C_p<\infty$, \begin{equation}\label{eq:pair-pilot-bounds} \begin{gathered} \sup_{\mathsf w}|\bar c(\mathsf w)|\le C_p,\qquad \sup_{\mathsf w}\E|\zeta_{\mathcal S}(\mathsf w)|^2\le C_pW,\\ \left\|\int\nu^\top\zeta_{\mathcal S}\,dM_h\right\|_{L^2} \le C_p\sqrt{W/n}\,\|\nu\|_{L^2(M_h)} \quad\text{for every Borel }\nu\in L^2(M_h;\mathbb R^3),\\ r_h=\bar c^\top\mathbf f,\qquad \mathbf f(\mathsf w)=(f(x'),f(x),1)^\top,\qquad \|r_h\|_\infty\le b. \end{gathered} \end{equation} For $O_i=(X_i,Y_i)$, let $\mathsf w_{ii'}=(X_i,X_{i'})$, $z_{ii'}=(Y_{i'},Y_i,1)^\top$, $P_0=\operatorname{diag}(1,1,0)$, and $(m)_2=m(m-1)$. Define \begin{equation}\label{eq:pair-statistics} \begin{aligned} \widehat N&=\frac1{2(m)_2}\sum_{i\ne i'}k_h(X_i,X_{i'}) c_{\mathcal A}(\mathsf w_{ii'})^\top z_{ii'}z_{ii'}^\top c_{\mathcal B}(\mathsf w_{ii'}),\\ \widehat D&=\frac1{2(m)_2}\sum_{i\ne i'}k_h(X_i,X_{i'}) c_{\mathcal A}(\mathsf w_{ii'})^\top P_0c_{\mathcal B}(\mathsf w_{ii'}), \end{aligned} \end{equation} reading a summand as zero when $k_h=0$. These statistics are square integrable and $\E\widehat D\ge1/2$. If a fixed $a>0$ satisfies $\E_\theta\widehat D\ge2a$ for every $\theta$, then, with clipping denoting projection onto the displayed interval, \begin{equation}\label{eq:pair-risk} \begin{gathered} \widehat V_n=\operatorname{clip}_{[v_-,v_+]} \left(\frac{\widehat N}{\max\{\widehat D,a\}}\right),\\ \sup_{\theta\in\Theta}\|\widehat V_n-V\|_{L^2} \le C\left[b^2+(n^{-1/2}+n^{-1}h^{-d/2})(1+W)\right]. \end{gathered} \end{equation} Here $C$ depends only on the model parameters and $(c_*,C_p,a)$. In particular, $a=1/4$ always works. \end{proposition} \begin{proof} Fix $\theta$, and write $\mu=\mu_\theta$. Using the assumptions and \eqref{eq:U7-mass}, we obtain $\sup_{\mathsf w}\E|c_{\mathcal S}(\mathsf w)|^2\le C(1+W)$ and, by Tonelli's theorem, $\E\|\zeta_{\mathcal S}\|_{L^2(M_h)}^2\le CW$. Then the random test functions used below belong to $L^2(M_h)$ almost surely. For an evaluation pair $o=(x,y),o'=(x',y')$, let $z=(y',y,1)^\top$. We take either $\mathsf A=zz^\top-VP_0$ or $\mathsf A=P_0$, and define $F=c_{\mathcal A}^\top\mathsf A c_{\mathcal B}$ and $T_F=[2(m)_2]^{-1}\sum_{i\ne i'}k_h(X_i,X_{i'})F(O_i,O_{i'})$. The two resulting statistics are $\widehat N-V\widehat D$ and $\widehat D$, respectively. The assumption on the conditional fourth moments and $|f|\le L_s$ imply that $\E[\|\mathsf A\|^2\mid X,X']\le C$. Since the two pilots and the evaluation sample are independent, we obtain \begin{equation}\label{eq:U8-G2} B_F:=\iint k_hF^2\,d\mu^2,\qquad \E B_F\le C(1+W)^2. \end{equation} Note that this bound uses only one second moment from each pilot. By finite mass and Cauchy--Schwarz, we also have $\E\iint k_h|F|\,d\mu^2<\infty$. Since $\E[zz^\top\mid X,X']=\mathbf f\mathbf f^\top+VP_0$, we have \begin{equation}\label{eq:U8-cond} \E[T_F\mid\mathcal A,\mathcal B] =L_B:=\frac12\int c_{\mathcal A}^\top Bc_{\mathcal B}\,dM_h, \qquad B=\mathbf f\mathbf f^\top\ \text{or}\ P_0. \end{equation} Note that this calculation of the mean does not require independence between pairs, since each ordered pair has law $\mu^2$. Averaging over the pilots, we obtain \begin{align} \E(\widehat N-V\widehat D) &=\frac12\int r_h^2\,dM_h\le Cb^2,\label{eq:U8-bias}\\ \E\widehat D &=\frac12\int\bar c^\top P_0\bar c\,dM_h\ge\frac12. \label{eq:U8-denominator} \end{align} In the last inequality, we used that $\bar c_1=1$ and the lower bound in \eqref{eq:U7-mass}. Both $B$ and $\bar c$ are uniformly bounded. Expanding the pilots, we find \[ L_B-\E L_B=\frac12\int \{\zeta_{\mathcal A}^\top B\bar c+ \bar c^\top B\zeta_{\mathcal B}+ \zeta_{\mathcal A}^\top B\zeta_{\mathcal B}\}\,dM_h. \] By the test-function bound, each linear term has $L^2$ norm at most $C\sqrt{W/n}$. For the bilinear term, we condition on $\mathcal B$ and apply the same bound with the independent test function $B\zeta_{\mathcal B}$. Then the squared $L^2$ norm of the bilinear term is at most $(CW/n)\E\|B\zeta_{\mathcal B}\|_{L^2(M_h)}^2\le CW^2/n$. Since $\sqrt W\le1+W$, we conclude that \begin{equation}\label{eq:U11-pilot-fluctuation} \|L_B-\E L_B\|_{L^2}\le Cn^{-1/2}(1+W). \end{equation} For completeness, we include the order-two variance bound, which is elementary. For a symmetric $H\in L^2(\mu^2)$, let $H_1(o)=\int H(o,o')\mu(do')$ and $\bar H=\int H\,d\mu^2$. We decompose $H=\bar H+h_1(o)+h_1(o')+h_2(o,o')$, where $h_1=H_1-\bar H$ and $h_2$ has mean zero in either argument. The two orders are orthogonal; within each order, only equal sets of observation indices contribute. Subtracting conditional means is an orthogonal projection. The order-two Hoeffding decomposition \cite{Hoeffding1948} then gives \begin{equation}\label{eq:U6-order-two} \Var\left(\frac1{(m)_2}\sum_{i\ne i'}H(O_i,O_{i'})\right) =\frac4m\|h_1\|_2^2+\frac2{(m)_2}\|h_2\|_2^2 \le\frac4m\int H_1^2d\mu+\frac2{(m)_2}\int H^2d\mu^2. \end{equation} We now fix pilots with $B_F<\infty$, symmetrize $\overline F(o,o')=(F(o,o')+F(o',o))/2$, and set $H=k_h\overline F/2$. This symmetrization preserves $T_F$. By Jensen's inequality and $k_h^2=h^{-d}k_h$, we have $\int H^2d\mu^2\le h^{-d}B_F/4$. A weighted Cauchy--Schwarz inequality gives \[ H_1(o)^2\le\frac14\left(\int k_h\,d\mu(o')\right) \left(\int k_h\overline F(o,o')^2\,d\mu(o')\right), \qquad \int H_1^2d\mu\le p_+B_F/4. \] Then \eqref{eq:U6-order-two} implies that \begin{equation}\label{eq:U11-pair-variance} \Var(T_F\mid\mathcal A,\mathcal B) \le\left(\frac{p_+}{m}+\frac1{2(m)_2h^d}\right)B_F. \end{equation} By Cauchy--Schwarz, we also have $|\E[T_F\mid\mathcal A,\mathcal B]|^2\le p_+B_F/4$, and \eqref{eq:U8-G2} then proves square integrability. Since $(m)_2\ge m^2/2$ and $m\ge c_*n$, taking expectations yields \begin{equation}\label{eq:U11-pair-fluctuation} \|T_F-L_B\|_{L^2} \le C(n^{-1/2}+n^{-1}h^{-d/2})(1+W). \end{equation} Finally, we note that clipping is $1$-Lipschitz and fixes $V$, and that \[ \frac{\widehat N}{\max\{\widehat D,a\}}-V =\frac{\widehat N-V\widehat D-V(a-\widehat D)_+} {\max\{\widehat D,a\}}. \] Since $\E\widehat D\ge2a$, we have $(a-\widehat D)_+\le|\widehat D-\E\widehat D|$, and it follows that \begin{equation}\label{eq:U14-stab} |\widehat V_n-V|\le a^{-1} \{|\widehat N-V\widehat D|+v_+|\widehat D-\E\widehat D|\}. \end{equation} Combining this bound with the two fluctuation bounds and \eqref{eq:U8-bias}, we obtain \eqref{eq:pair-risk}. The reduction \eqref{eq:U14-stab} is similar to that of Shen et al.\ \cite[Appendix~A.1, equation~(28)]{ShenGaoWittenHan2020}, who control the error of their ratio through the numerator minus the target times the denominator, on an event where the denominator is bounded below. \end{proof} \begin{proof}[Proof of \cref{prop:elementary-upper} for $01$ and that $d>4s$ is an integer. We construct one estimator from two predictions of a regression increment and an independent sample of nearby pairs. By resolving the regression function below the sampling scale, we obtain the following improvement. \begin{theorem}\label{thm:upper-sharp} Let $s>1$ and let $d>4s$ be an integer. Put $L=\log n$ and \[ \begin{gathered} \lambda=\frac{2(s+1)}{d+4}, \qquad \tau=\log\frac{\sqrt{p_+}+\sqrt{p_-}}{\sqrt{p_+}-\sqrt{p_-}}, \\ \kappa=4\sqrt{\frac{2\tau(s-1)(d-4s)}{(d+4)^3}}, \qquad q_s=\binom{d+\lceil s\rceil-1}{d}. \end{gathered} \] There exist $C<\infty$ and $n_0\in\mathbb N$, depending only on $(s,d,p_-,p_+,L_s,v_-,v_+,C_4)$, such that for every integer $n\ge n_0$ \begin{equation}\label{eq:upper-sharp} \mathfrak r_n \le C\,n^{-\lambda}e^{-\kappa\sqrt L}\, L^{\frac{s-1}{d+4}+\frac{d(q_s-1/2)}{d+4}} . \end{equation} \end{theorem} Recalling that $\Psi_n=n^{-\lambda}e^{-\kappa\sqrt L}L^{(s-1)/(d+4)}$ and $\Gamma=d(q_s-\tfrac12)/(d+4)$, we see that the bound \eqref{eq:upper-sharp} reads $\mathfrak r_n\le C\,\Psi_nL^{\Gamma}$. The constants do not depend on the extension domain $\mathcal U$; the only property of the domain used in bounding the Taylor remainder is the convexity of $[0,1]^d$. We complete the proof in \cref{sec:U-final}. \section{Setting and fixed quantities}\label{sec:U-setup} For pairs in one cube of side $h$, the estimator uses products of two corrected response differences, with corrections given by independent predictions of the increment $f(x')-f(x)$. Resolving the regression function at $K$ cells yields a bias of order $h^2K^{-2(s-1)/d}$ and a pair error $n^{-1}h^{-d/2}$. When $K=ne^T$, balancing the two improves the rate $n^{-\lambda}$ at the sampling scale by the factor $e^{-2(s-1)T/(d+4)}$. We achieve $T\asymp\sqrt{\log n}$ using unbiased polynomials of local moments, whose degrees are chosen to control their factorial variances. Kong and Valiant \cite[Section~2.2]{KongValiant2018} approximate a scalar functional of an unknown covariance matrix by a polynomial in its powers, and they estimate the resulting quadratic forms without bias by sums over distinct observations. Chen, Liu and Mukherjee \cite[Section~3.2]{ChenLiuMukherjee2024} approximate quadratic forms in an inverse covariance matrix by Chebyshev polynomials and estimate the resulting quadratic forms in matrix powers by U-statistics. The higher-order influence function estimators of Robins et al.\ \cite[Theorems~3.14--3.15]{RobinsEtAl2008} use inverse Gram matrices computed under initial nuisance estimates, together with U-statistics of increasing order, to correct the bias of an estimator of a functional. We write $|\cdot|$ for the Euclidean norm of a vector and $\norm{\cdot}$ for the operator norm of a matrix; the matrix blocks of a moment tuple are measured in the Frobenius norm. Let $\ell=\lceil s\rceil-1$ and \begin{equation}\label{eq:U-setup-dims} q_s=\binom{d+\ell}{d},\qquad r=q_s-1,\qquad L=\log n. \end{equation} The constants for the approximation of the inverse and for the allocation are \begin{equation}\label{eq:U-setup-rho} \tau=\log\frac{\sqrt{p_+}+\sqrt{p_-}}{\sqrt{p_+}-\sqrt{p_-}}, \qquad \rho=e^{-\tau}\in(0,1), \end{equation} \[ \vartheta=(s-1)/d,\qquad \beta=\vartheta/\tau, \] \begin{equation}\label{eq:U-setup-exponents} \varpi=\frac{d-4s}{d(d+4)},\qquad \eta=\frac{2(d+4s)}{d(d+4)}, \end{equation} \begin{equation}\label{eq:U-setup-allocation} a_*=\sqrt{8\varpi/\beta}. \end{equation} In particular, we have \begin{equation}\label{eq:U-setup-rate} \lambda=\frac{2(s+1)}{d+4},\qquad \kappa=\frac{2(s-1)}{d+4}a_* =4\sqrt{\frac{2\tau(s-1)(d-4s)}{(d+4)^3}}. \end{equation} Below, we write $C,c$ for fixed positive constants. They depend only on $\mathfrak p=(s,d,p_-,p_+,L_s,v_+,C_4)$, and their values may change between lines. Together with the model moment bounds, the feature constant $C_\psi$ of \cref{lem:U1} and the analytic constants $M_1,C_1$ of \cref{lem:U5} determine the variance constant $C_0\ge1$, which is fixed sufficiently large in \cref{lem:U7}. \subsection*{Model bounds} Let $\mu$ be the law of $O=(X,Y)$, and let $\varepsilon=Y-f(X)$. By the assumptions of the model, we have \begin{equation}\label{eq:U-M0} \sup_{[0,1]^d}|f|\le L_s, \end{equation} \begin{equation}\label{eq:U-M2} \E[(1+|Y|)^2\mid X]\le M_2:=2(1+L_s^2+v_+). \end{equation} For the second bound, note that $\E[Y^2\mid X]=f(X)^2+V$ and $(1+|Y|)^2\le2(1+Y^2)$. By Taylor's formula, for all $x,z\in[0,1]^d$, we have \begin{equation}\label{eq:U-Tay} \left|f(z)-\sum_{|\gamma|\le\ell} \frac{\partial^\gamma f(x)}{\gamma!}(z-x)^\gamma\right| \le C_{\rm Tay}|z-x|^s,\qquad C_{\rm Tay}=d^\ell L_s/\ell!. \end{equation} To verify the endpoint convention for integer $s$, we apply the integral form of Taylor's formula in one dimension to $F(t)=\widetilde f(x+t(z-x))$. The remainder after degree $\ell$ is \[ \int_0^1\frac{(1-t)^{\ell-1}}{(\ell-1)!} \{F^{(\ell)}(t)-F^{(\ell)}(0)\}\,dt. \] By the H\"older bound and the identity $\sum_{|\gamma|=\ell}\ell!/\gamma!=d^\ell$, its integrand is bounded by $d^\ell L_s|z-x|^s$, and the weight has mass $1/\ell!$. Since the segment stays in the convex cube, the extension domain does not enter these constants. Finally, since $V^2\le C_4$, we have \begin{equation}\label{eq:U-setup-IV} V\in I_V:=[v_-,\min\{v_+,\sqrt{C_4}\}]. \end{equation} \subsection*{Cells and fixed samples} For $j\ge0$, let $\mathcal D_j$ be the set of half-open dyadic cubes of side $2^{-j}$ in $[0,1)^d$, let $z_I$ be a cell's corner, and set $K_j=2^{jd}=|I|^{-1}$. We denote the cell containing $x$ by $I_j(x)$. Any two cells are either disjoint or nested, and \begin{equation}\label{eq:U-setup-cellprob} p_-K_j^{-1}\le\Pp(X\in I)\le p_+K_j^{-1} \qquad(I\in\mathcal D_j). \end{equation} We divide the first $3n_b$ observations into independent blocks $\mathcal A,\mathcal B,\mathcal P$, where \begin{equation}\label{eq:U-setup-Lambda} n_b=\lfloor n/3\rfloor,\qquad \Lambda=n/6. \end{equation} The first two blocks are used for the pilots, and the third for the pairs; the remaining observations are discarded. For a block $\mathcal S$ of size $m$, we define the factorial sums \begin{equation}\label{eq:U-setup-Sigma} \Sigma_k(\phi;\mathcal S) =\sum_{\substack{1\le i_1,\ldots,i_k\le m\\ i_1,\ldots,i_k\text{ distinct}}} \phi(O_{i_1},\ldots,O_{i_k}), \qquad (m)_k=m!/(m-k)!. \end{equation} The sum is zero when $k>m$, and $\Sigma_0(c;\mathcal S)=c$. The admissibility condition below ensures that the denominators are positive. \section{Anchored monomials and local fits}\label{sec:U-anchored} We index coordinates by $\mathcal I=\{\gamma\in\mathbb N_0^d:1\le|\gamma|\le\ell\}$, taken in a fixed order, and let $q(v)=(v^\gamma)_{\gamma\in\mathcal I}$. For $I\in\mathcal D_j$ and $x\in I$, we define \begin{equation}\label{eq:Udef-features} \psi_{I,x}(z)=q(2^j(z-x)),\qquad B_I(x)=\int_{[0,1]^d}q(w-u)q(w-u)^\top dw, \quad u=2^j(x-z_I). \end{equation} We abbreviate $\psi_{j,x}=\psi_{I_j(x),x}$ and $B_j(x)=B_{I_j(x)}(x)$. Note that the matrix $B_I(x)$ is known; its $(\gamma,\delta)$ entry equals \[ \prod_{i=1}^d \frac{(1-u_i)^{\gamma_i+\delta_i+1} -(-u_i)^{\gamma_i+\delta_i+1}} {\gamma_i+\delta_i+1}. \] The coordinate transport is the fixed diagonal matrix \begin{equation}\label{eq:U-setup-transport} \mathcal T_j=\operatorname{diag}_{\gamma\in\mathcal I}(2^{-|\gamma|}) \quad(j\ge1),\qquad \mathcal T_0=0. \end{equation} In part~(3) of the next lemma, we normalize the local Gram matrix by its Lebesgue counterpart $B_I(x)$, so that the density bounds place its spectrum in $[p_-,p_+]$. Park~\cite[Section~3.2, equation~(9)]{Park2026} obtains the corresponding spectral interval with a block basis orthonormal for Lebesgue measure. Our companion paper with Kallus~\cite[Section~4.2, proof of Lemma~4.3]{AronowKallusLopatto2026} uses a basis orthonormal for normalized Lebesgue measure on each cell and includes the reciprocal cell volume in its Gram matrix. \begin{lemma}\label{lem:U1} There are fixed $c_B,C_B>0$ and $C_\psi\ge1$, depending only on $(d,s)$, such that the following holds at all cells and anchors. \begin{enumerate} \item[\textup{(1)}] The components of $\psi_{I,x}$ form a basis of the degree-at-most-$\ell$ polynomials vanishing at $x$, and \begin{equation}\label{eq:U1-Gram} K_j\int_I\psi_{I,x}\psi_{I,x}^\top dz=B_I(x), \qquad c_BI_r\preceq B_I(x)\preceq C_BI_r. \end{equation} \item[\textup{(2)}] We have $\sqrt r,\sqrt{C_B/c_B}\le C_\psi$, $\sup_I|\psi_{I,x}|\le C_\psi$, \begin{equation}\label{eq:U1-lipschitz} |\psi_{I,x}(x')|\le C_\psi2^j|x'-x|\quad(x'\in I), \end{equation} and on $I$, $|B_I(x)^{-1}\psi_{I,x}|, \|B_I(x)^{-1}\psi_{I,x}\psi_{I,x}^\top\|_F\le C_\psi^2$. \item[\textup{(3)}] If $p_-\le p\le p_+$ almost everywhere on $I$ and $H=K_j\int_Ip\psi_{I,x}\psi_{I,x}^\top dz$, then \begin{equation}\label{eq:U1-spectrum} p_-B_I(x)\preceq H\preceq p_+B_I(x). \end{equation} Thus $B_I(x)^{-1}H$ is similar to a symmetric matrix with spectrum in $[p_-,p_+]$, with similarity condition number at most $\chi:=\sqrt{C_B/c_B}$. \item[\textup{(4)}] On all of $\mathbb R^d$, for $j\ge1$, \begin{equation}\label{eq:U1-transport} \psi_{j-1,x}=\mathcal T_j^\top\psi_{j,x},\qquad \norm{\mathcal T_j}\le1/2. \end{equation} \item[\textup{(5)}] The features and known inverse $B_I(x)^{-1}$ are continuous in their arguments on each fixed cell and globally Borel; the transports are constant. \end{enumerate} \end{lemma} \begin{proof} The centered monomials form a basis of the space of anchored polynomials, and the Gram identity follows from a change of variables on the cell. If a coefficient vector is nonzero, then the corresponding polynomial is nonzero, and its squared integral is positive. Since $u$ ranges over the compact cube $[0,1]^d$, the constants $c_B,C_B$ can be chosen uniformly by continuity, and $B_I(x)^{-1}$ is continuous. Part (2) holds after we increase the single fixed constant $C_\psi$, since the monomials, their gradients and this inverse are bounded. Part (3) follows from the density bounds: writing $B=B_I(x)$, we see that $S=B^{-1/2}HB^{-1/2}$ has spectrum in $[p_-,p_+]$ and that $B^{-1}H=B^{-1/2}SB^{1/2}$. Scaling each centered monomial gives (4), because $|\gamma|\ge1$. Finally, (5) follows from the explicit formulas on each cell and the fact that the cells of each level form a finite Borel partition. \end{proof} \subsection*{Preconditioned moments and the anchored fit} For a parameter $\theta=(p,f,V,Q)\in\Theta$, we define \begin{equation}\label{eq:Udef-moments} \begin{aligned} H_j(x)&=K_j\int_{I_j(x)}p(z)\psi_{j,x}(z)\psi_{j,x}(z)^\top dz, &G_j(x)&=B_j(x)^{-1}H_j(x),\\ \beta_j(x)&=B_j(x)^{-1}K_j\int_{I_j(x)}p(z)\psi_{j,x}(z)f(z)dz, &\mathrm m_j(x)&=B_j(x)^{-1}K_j\int_{I_j(x)}p(z)\psi_{j,x}(z)dz. \end{aligned} \end{equation} These quantities are bounded uniformly in the level and the anchor, and they are Borel by dominated convergence on each cell. Using the similarity in \cref{lem:U1}, we obtain \begin{equation}\label{eq:U-setup-Gspec} \norm{G_j}\le\chi p_+,\qquad \norm{G_j^{-1}}\le\chi p_-^{-1},\qquad p_-^r\le\det G_j\le p_+^r. \end{equation} For a formal anchor value $y\in\mathbb R$, we set \begin{equation}\label{eq:Udef-b} b_j(x;y)=\beta_j(x)-y\mathrm m_j(x) =B_j(x)^{-1}K_j\int_{I_j(x)}p(z)\psi_{j,x}(z)(f(z)-y)dz. \end{equation} At the root, the parent values are $G_{-1}=I_r$ and $\beta_{-1}=\mathrm m_{-1}=0$. Since $p_-<10$, depending only on $(d,s,p_-,p_+,L_s)$, with the following properties at every population tuple, level and anchor, for arbitrary nonnegative degrees $m_j$. \emph{(1)} For every real $y$, \begin{equation}\label{eq:U5-identity} G^{-1}\det(G^-)^{-1}\mathcal N_j(x;y) =G^{-1}b_j(x;y)-\mathcal T_j(G^-)^{-1}b_{j-1}(x;y), \end{equation} and hence \begin{equation}\label{eq:U5-N} |\mathcal N_j(x;f(x))| \le\chi p_+^{r+1}|\delta_j(x)|\le C_{\rm inc}2^{-js}. \end{equation} \emph{(2)} The approximation error satisfies: \begin{equation}\label{eq:U5-error} |F_j(x;f(x))-\delta_j(x)| \le C_{\rm inc}2^{-js}(m_j+1)^r\rho^{m_j}. \end{equation} \emph{(3)} The polynomial $F_{j,x}(\theta;y)$ has degree at most $D_j=m_j+r+1$. Every monomial contains one vector-moment coordinate, to first power, and has coefficient constant or linear in $y$. Thus each observation's response enters its factorial lift at most to first power. \emph{(4)} Uniformly for $|y_0|\le L_s$, $\Phi\in\{F_{j,x}(\cdot;0),-\partial_yF_{j,x}, F_{j,x}(\cdot;y_0)\}$, $k\ge0$ and all directions in $E$, \begin{equation}\label{eq:U5-deriv} |D^k\Phi(\theta_j(x))[\gamma_1,\ldots,\gamma_k]| \le M_1 k!C_1^k\prod_i|\gamma_i|. \end{equation} The constants $M_1,C_1$ depend only on $(r,p_-,p_+,C_\psi,L_s)$, not on the level or degrees. \end{lemma} \begin{proof} For $z\in[p_-,p_+]$, let $\phi$ be such that $t(z)=\cos\phi$. Summing two geometric series, we obtain the Chebyshev identity \[ g_\zeta(z)=\kappa_0\{1+2\sum_{k\ge1}(-\rho\zeta)^k\cos(k\phi)\}, \qquad |\zeta|<\rho^{-1}. \] Its coefficients have modulus at most $2\kappa_0\rho^k$, and a direct substitution of $\rho$ shows that $g_1(z)=z^{-1}$. The identity is the weighted generating function of the Chebyshev polynomials \cite[Eq.~18.12.7]{DLMF}; for generating functions, see also \cite[Section~5.2.1]{MasonHandscomb2003} and \cite[Section~1.5]{Rivlin1990}. At $\zeta=1$, after an affine change of variables, it is the classical Chebyshev expansion of the reciprocal \cite[equation~(5.14)]{MasonHandscomb2003}, \cite[equation~(2.4)]{Mathar2006}. By \cref{lem:U1}(3), each population Gram matrix is similar to a symmetric matrix with spectrum in this interval, through a similarity with condition number at most $\chi$. The bound on its matrix coefficients then gains only the fixed factor $\chi$. Since the determinant is invariant under similarity, its $r$ scalar factors satisfy the original bounds. Convolving these bounds, we find \[ \|[\zeta^k]\{g_\zeta(A)\det g_\zeta(C)\}\| \le\chi(2\kappa_0)^{r+1}\binom{k+r}{r}\rho^k. \] For these population matrices, the series at one equals $A^{-1}\det(C)^{-1}$. Using $\binom{m+1+i+r}{r}\le(r+1)(m+1)^r\binom{i+r}{r}$ and $\sum_{i\ge0}\binom{i+r}{r}\rho^i=(1-\rho)^{-r-1}$, we obtain \begin{equation}\label{eq:U4-approx} \norm{\mathsf P_m(A,C)-A^{-1}\det(C)^{-1}} \le C(m+1)^r\rho^m. \end{equation} We obtain \eqref{eq:U5-identity} from the adjugate identity, and at $y=f(x)$ this identity gives $\mathcal N_j=\det(G^-)G\delta_j$. Then \eqref{eq:U-setup-Gspec}, together with \eqref{eq:U2-increment}, implies (1). For (2), we multiply the bound on the small numerator by \eqref{eq:U4-approx}. Each monomial of the numerator contains at most $r$ matrix coordinates and one vector coordinate. Since $\mathsf P_m$ contains only matrix coordinates and has degree at most $m$, this proves (3), including the substitution at the root. For (4), let $R=(1+\rho^{-1})/2$, and consider the set \[ \mathcal G=\{B^{-1}H:\ B=B^\top,\ H=H^\top,\quad I_r\preceq B\preceq C_\psi^2I_r,\quad p_-B\preceq H\preceq p_+B\} \] This set is compact. To see that it contains every population Gram matrix, we divide the known Lebesgue Gram and its density Gram by $c_B$. It also contains the root parent $I_r$. Since its members are similar to symmetric matrices with spectrum in $[p_-,p_+]$, the matrix $I_r+2\rho\zeta t(A)+\rho^2\zeta^2I_r$ is invertible on $\mathcal G\times\{|\zeta|\le R\}$. Recall that the vector moments are uniformly bounded and that $\|\mathcal T_j\|\le1/2$, including $\mathcal T_0=0$. By compactness and the continuity of inversion, there exist fixed $\epsilon,M>0$ such that the $y=0$, $y=y_0$ and $-\partial_y$ versions of $g_\zeta(G)\det g_\zeta(G^-)\mathcal N_j(\theta;y)$ are bounded by $M$ for $|\zeta|\le R$, $|\theta-\theta_j(x)|\le\epsilon$ and $|y_0|\le L_s$, and are holomorphic on a neighborhood of this set. We choose $\epsilon$ small enough that the closed neighborhood lies in this holomorphic domain. Within this neighborhood, perturbations of the matrices may be complex and need not be symmetric. By Cauchy's formula, each $\zeta$ coefficient is bounded by $MR^{-l}$, where $l$ is its order in the series. Summing the geometric series, we find that all partial sums are bounded by $M'=M/(1-R^{-1})$, independently of $m_j$. For nonzero directions, the Cauchy estimate in each direction on $|t_i|=\epsilon/(k|\gamma_i|)$ gives \[ |D^k\Phi(\theta_j)[\gamma_1,\ldots,\gamma_k]| \le M'(k/\epsilon)^k\prod_i|\gamma_i| \le M'k!(e/\epsilon)^k\prod_i|\gamma_i|. \] In the last inequality, we used the fact that $k!\ge(k/e)^k$. The cases of zero directions and $k=0$ are immediate. All compact sets and bounds above depend only on $(r,p_-,p_+,C_\psi,L_s)$, and this proves (4). \end{proof} \section{Unbiased polynomials from a fixed sample}\label{sec:U-factorial} Let $\mathcal S=(O_1,\dots,O_{n_b})$ consist of independent observations with common law $\mu$. A polynomial $\Phi:\mathbb R^N\to\mathbb R^r$ of degree $D\le n_b$ has a unique homogeneous expansion $\Phi(\theta)=\sum_{k=0}^D\Phi_{(k)}[\theta,\dots,\theta]$, where $\Phi_{(k)}$ is symmetric and $k$-linear. Using the factorial sums of \eqref{eq:U-setup-Sigma}, we define its unbiased lift by \begin{equation}\label{eq:U6-hat} \widehat\Phi(\mathcal S;g) =\sum_{k=0}^D\frac{\Sigma_k( \Phi_{(k)}[g(\cdot),\dots,g(\cdot)];\mathcal S)}{(n_b)_k}. \end{equation} The lift is linear in $\Phi$. If the coefficients of the polynomial are affine in a formal variable $y$, we lift the coefficients and leave $y$ unchanged. For the indicator of a cell and a monomial of degree $k$, the lift is the falling factorial of the cell count divided by $(n_b)_k$, the unbiased estimator of a power of a cell probability used in \cite[equation~(21)]{JiaoEtAl2015}. To compute the variance, we expand around the mean $\theta=\int g\,d\mu$. This mean enters the analysis, while \eqref{eq:U6-hat} is computed from the known coefficients of the polynomial. For scalar kernels on $\mathbb O^k$, let $E_i$ denote integration of the $i$th argument against $\mu$, and let $\Pi_k=\prod_{i=1}^k(I-E_i)$. Since the operators $E_i$ commute and are orthogonal projections in $L^2(\mu^k)$, we have $\|\Pi_k H\|_2\le\|H\|_2$. \begin{lemma}[Centered Taylor identity]\label{lem:U6} Let $g\in L^2(\mu;\mathbb R^N)$, $\theta=\int g\,d\mu$ and $Z(o)=g(o)-\theta$. The lift is square integrable, has mean $\Phi(\theta)$, and satisfies \begin{equation}\label{eq:U6-centered} \widehat\Phi-\Phi(\theta) =\sum_{k=1}^D\frac{1}{k!(n_b)_k} \Sigma_k\bigl(D^k\Phi(\theta)[Z(\cdot),\dots,Z(\cdot)];\mathcal S\bigr). \end{equation} For $a\in\mathbb R^r$, put \begin{equation}\label{eq:U6-poly-kernel} \mathcal K_k(o_1,\dots,o_k) =a\cdot D^k\Phi(\theta)[g(o_1),\dots,g(o_k)]. \end{equation} If $0<\Lambda\le n_b-D+1$, then \begin{equation}\label{eq:U6-comparison} \Var(a\cdot\widehat\Phi) =\sum_{k=1}^D\frac{\|\Pi_k\mathcal K_k\|_{L^2(\mu^k)}^2}{k!(n_b)_k} \le\sum_{k=1}^D\frac{\|\mathcal K_k\|_{L^2(\mu^k)}^2}{k!\Lambda^k}. \end{equation} For a measurable field of scalar lifts, write $Z(w)$ for the centered lift and $\mathcal K_{w,k}$ for its raw kernels, padding with zero up to a common degree $D$. Let $M$ be a finite measure and suppose these kernels have uniformly bounded $L^2(\mu^k)$ norms. For $\nu\in L^2(M)$, if a finite partition into blocks $B$ has disjoint observation supports for raw kernels from distinct blocks at every order $k\ge1$, then \begin{equation}\label{eq:U6-localized} \begin{aligned} \left\|\int\nu(w)Z(w)\,dM(w)\right\|_{L^2}^2 &\le\sum_{k=1}^D\frac{\left\|\int\nu(w)\mathcal K_{w,k}\,dM(w)\right\|_2^2} {k!\Lambda^k}\\ &\le\max_B M(B)\int\nu(w)^2 \sum_{k=1}^D\frac{\|\mathcal K_{w,k}\|_2^2}{k!\Lambda^k}\,dM(w). \end{aligned} \end{equation} The first inequality also holds without the support assumption. \end{lemma} \begin{proof} By multilinearity and independence, every summand of the lift is integrable, and the lift has the claimed mean. Each summand is square integrable because $g\in L^2$. We expand each homogeneous term using $g=\theta+Z$. For a term of degree $l$ and a choice of $k$ centered arguments, summing over the unused observation indices gives the factor $(n_b-k)_{l-k}$. Since $(n_b)_l=(n_b)_k(n_b-k)_{l-k}$, its coefficient after averaging is $\binom lk\Phi_{(l)}[Z,\dots,Z,\theta,\dots,\theta]$. Summing over $l\ge k$, we obtain $D^k\Phi(\theta)[Z,\dots,Z]/k!$, which proves \eqref{eq:U6-centered}. Observe that each centered kernel integrates to zero in any one of its arguments. Two summands in \eqref{eq:U6-centered} are therefore orthogonal unless their tuples contain exactly the same observation indices, since an index that appears only once can be integrated first. At order $k$, there are $(n_b)_k$ choices of the first tuple and $k!$ reorderings for the second. Further, we have $\Pi_k\mathcal K_k=a\cdot D^k\Phi(\theta)[Z,\dots,Z]$, and this proves the variance identity. The inequality follows from the fact that $\Pi_k$ is a contraction and from $(n_b)_k\ge(n_b-D+1)^k\ge\Lambda^k$. For the integrated version, since $M$ is finite and the kernel bounds are uniform, we have $\int|\nu|\|\mathcal K_{w,k}\|_2\,dM<\infty$. By Minkowski's inequality and Fubini's theorem, we may then integrate the centered Taylor identity and commute the integral with $\Pi_k$. We apply the variance identity and the contraction property of the projections to the integrated kernels. At each order, the block integrals are orthogonal because the raw supports are disjoint, and Cauchy--Schwarz on each block gives \[ \left\|\int\nu\mathcal K_{w,k}\,dM\right\|_2^2 =\sum_B\left\|\int_B\nu\mathcal K_{w,k}\,dM\right\|_2^2 \le\sum_B M(B)\int_B\nu^2\|\mathcal K_{w,k}\|_2^2\,dM. \] Summing over $k$ completes the proof of \eqref{eq:U6-localized}. We note that the localization takes place before centering; the actual covariances between different blocks for the fixed sample need not vanish. \end{proof} Since each centered kernel integrates to zero in any one of its arguments, the centered Taylor identity is the Hoeffding decomposition \cite{Hoeffding1948} of the lift. Our companion paper defines the same lift and proves the same covariance identity \cite[Section~4.3, Lemma~4.4]{AronowKallusLopatto2026}. A variance identity of the same form, for polynomial combinations of complete U-statistics estimating trace moments, appears in \cite[Theorem~6.2]{Song2026}. \section{Definition of the estimator}\label{sec:U-estimator} The estimator uses the independent pilot and evaluation blocks and the comparison parameter $\Lambda$ of \eqref{eq:U-setup-Lambda}. Its parameters are a final level $J\ge0$, degrees $m_0,\dots,m_J\ge0$ and a pair level $J_h\ge J$, all of which are integers. We set \[ h=2^{-J_h},\quad K_*=K_J=2^{dJ},\quad D_j=m_j+q_s\ (0\le j\le J). \] We say that the parameters are \emph{admissible} if \begin{equation}\label{eq:Udef-admissible} \begin{gathered} J\ge0,\quad m_j\ge0\ (0\le j\le J),\quad J_h\ge J,\\ \max_{0\le j\le J}D_j\le n_b-\Lambda+1,\qquad n_b-1\ge\Lambda. \end{gathered} \end{equation} The condition on the degrees ensures that the unbiased polynomial estimates below are defined and that their variances satisfy \cref{lem:U6}. Here $\widehat V_n$ denotes the estimator with whichever admissible parameters have been specified. The bounds in \cref{sec:U-variance,sec:U-equation} apply to every such choice; the allocation below yields the bound in \cref{thm:upper-sharp}. \subsection*{Raw features and pilots} For an observation $o=(z,y_o)$, we define the raw features \begin{equation}\label{eq:Udef-rawfeatures} (\mathcal G_{j,x},\beta^{\rm f}_{j,x},\mathrm m^{\rm f}_{j,x})(o) =K_j\1_{I_j(x)}(z) \bigl(B_j(x)^{-1}\psi_{j,x}(z)\psi_{j,x}(z)^\top, B_j(x)^{-1}\psi_{j,x}(z)y_o, B_j(x)^{-1}\psi_{j,x}(z)\bigr). \end{equation} Since $\E[Y\mid X]=f(X)$, their means are $(G_j,\beta_j,\mathrm m_j)(x)$. For the fixed parent of the root, we use the features $(\mathcal G_{-1,x},\beta^{\rm f}_{-1,x},\mathrm m^{\rm f}_{-1,x})(o) =(I_r,0,0)\1_{[0,1)^d}(z)$. These have the required means, because $X\in[0,1)^d$ almost surely. We collect the current and parent features in the same order as the moment tuple: \begin{equation}\label{eq:Udef-g} g_{j,x}=(\mathcal G_{j,x},\mathcal G_{j-1,x}, \beta^{\rm f}_{j,x},\beta^{\rm f}_{j-1,x}, \mathrm m^{\rm f}_{j,x},\mathrm m^{\rm f}_{j-1,x}),\qquad j\ge0. \end{equation} By \eqref{eq:U-M2} and \cref{lem:U1}, these known features are square integrable. Further, $\E g_{j,x}=\theta_j(x)$. Let $g_x=(g_{0,x},\ldots,g_{J,x})$. For each independent pilot block $\mathcal S\in\{\mathcal A,\mathcal B\}$, we apply the lift \eqref{eq:U6-hat} directly to \eqref{eq:U-affine-polynomial}, keeping $y$ as a formal variable: \begin{equation}\label{eq:Udef-pilot} \widehat{\mathcal P}^{\mathcal S}_{x,x'}(y) =\widehat{\mathcal P_{x,x'}}(\mathcal S;g_x)(y) =\sum_{j=0}^J\psi_{j,x}(x')^\top \widehat{F_{j,x}(\cdot;y)}(\mathcal S;g_{j,x}). \end{equation} The second equality follows from linearity, so only degrees up to $\max_jD_j$ are required. The pilot is affine in $y$, with coefficients given by \begin{equation}\label{eq:U-affine-coefficients} u_{\mathcal S}=\widehat{\mathcal P}^{\mathcal S}_{x,x'}(0),\qquad v_{\mathcal S}=-\partial_y\widehat{\mathcal P}^{\mathcal S}_{x,x'}(y), \qquad \widehat{\mathcal P}^{\mathcal S}_{x,x'}(y)=u_{\mathcal S}-yv_{\mathcal S}. \end{equation} \subsection*{Pair statistic} The pair statistic uses the kernel $k_h$ of \cref{sec:pair-risk}, with $h=2^{-J_h}$. Since $J_h\ge J$, its pair points satisfy $x'\in I_J(x)$. On the pair set, we define the coefficient vector $c_{\mathcal S}=(1,v_{\mathcal S}-1,-u_{\mathcal S})^\top$, so that \[ c_{\mathcal S}^\top(y',y,1)^\top =y'-y-\widehat{\mathcal P}^{\mathcal S}_{x,x'}(y). \] Let $\widehat N,\widehat D$ be given by \eqref{eq:pair-statistics}, with $m=n_b$ and the evaluation block $\mathcal P$. The estimator is \begin{equation}\label{eq:Udef-estimator} \widehat V_n=\operatorname{clip}_{I_V} \left(\frac{\widehat N}{\max\{\widehat D,1/4\}}\right). \end{equation} As in \eqref{eq:pair-statistics}, a summand is zero off the pair set. The denominator bound in \cref{prop:pair-risk} allows us to use the threshold $1/4$, once we verify the pilot hypotheses of that proposition in \cref{lem:U7}. \subsection*{Parameters for the refined bound} Recall that $L=\log n$. Approximation at the fine levels requires roughly $\tau m_j\ge\vartheta(T-t_j)$, while the factorial variance costs $\exp\{m_j(t_j+\log m_j)\}$. At the required degree, the leading logarithmic variance cost is largest near $t_j=T/2$. The quadratic profile below keeps the factorial variances within the budget specified below and gives a Gaussian bound for the approximation errors. Park \cite[Sections~3.3--3.4]{Park2026} also lowers the degree on finer partitions and balances the geometric approximation error against a factorial variance, and our companion paper uses a related balance \cite[Section~4.4]{AronowKallusLopatto2026}. For $n\ge3$, we set \begin{equation}\label{eq:Udef-S} C_S=1+\frac{2(\eta+2)}{\beta},\qquad S=a_*\sqrt L-C_S. \end{equation} Since $\beta a_*^2/4=2\varpi$, we eventually have $S\ge1$ and \begin{equation}\label{eq:Udef-budget} \beta S^2/4+(\eta+2)S\le2\varpi L. \end{equation} When $S\ge1$, we define \begin{equation}\label{eq:Udef-degree-constants} C_D=\max\{1,4\beta\},\quad \ell_S=\log(C_0C_DS),\quad \gamma_*=(r+1/2)/\vartheta,\quad T^\circ=S-\ell_S-\gamma_*\log S, \end{equation} where $C_0\ge1$ is the constant fixed in \eqref{eq:U7-constants}, and the quadratic profile \begin{equation}\label{eq:Udef-degree-profile} \mathcal Q_S(z)=S/4+(S-z)^2/S. \end{equation} We will need the identities \begin{align} z+\mathcal Q_S(z)-S&=(z-S/2)^2/S,\label{eq:Udef-profile-slack}\\ z\mathcal Q_S(z)&=S^2/4-(S-z)(z-S/2)^2/S\le S^2/4 \quad(z\le S).\label{eq:Udef-profile-variance} \end{align} The candidate parameters are \begin{equation}\label{eq:Udef-params} \begin{aligned} J&=\left\lfloor\frac{L+T^\circ}{d\log2}\right\rfloor, \quad K_*=2^{dJ},\quad T=\log(K_*/n),\\ t_j&=\log(K_j/n),\quad z_j=t_j+\ell_S,\quad u_j=z_j-S/2,\\ D_j&=\lfloor\beta\mathcal Q_S(z_j)\rfloor, \quad m_j=D_j-q_s\quad(0\le j\le J),\\ h^\circ&=n^{-2/(d+4)}K_*^{4\vartheta/(d+4)}, \quad J_h=\lceil\log_2(1/h^\circ)\rceil,\quad h=2^{-J_h},\\ R_*&=(h^\circ)^2K_*^{-2\vartheta}=n^{-1}(h^\circ)^{-d/2}. \end{aligned} \end{equation} Level-dependent quantities are used only if $J\ge0$. Rounding gives \begin{equation}\label{eq:Udef-params-rounding} T^\circ-d\log20$, \begin{equation}\label{eq:U13-sums} W\le1,\qquad n^{-1/2}\le n^{-c}R_*, \end{equation} and \begin{equation}\label{eq:U13-pairs} n^{-1}h^{-d/2}\le2^{d/2}R_*. \end{equation} The resulting risk scale is \begin{equation}\label{eq:Urate-R} R_*\asymp n^{-\lambda}e^{-\kappa\sqrt L} L^{\frac{s-1}{d+4}+\frac{d(q_s-1/2)}{d+4}}. \end{equation} \end{lemma} \begin{proof} By the choice of $S$, we have $S\asymp\sqrt L$ and, eventually, $\beta S^2/4+(\eta+2)S\le2\varpi L$. The floor in $J$ gives $T=T^\circ+O(1)>0$ and \begin{equation}\label{eq:Urate-range} -L+\ell_S\le z_j\le T+\ell_S \le S-\gamma_*\log S\le S. \end{equation} Since $\mathcal Q_S(z)=S/4+(S-z)^2/S$, all degrees satisfy $D_j\ge\beta S/4-1$ and $\max_jD_j=O(L^{3/2})=o(n)$. Then $J\ge0$, $m_j\ge0$, and the sample-size conditions in \eqref{eq:Udef-admissible} hold eventually. In addition, we have \begin{equation}\label{eq:U13-reserve} h^2K_*^{2/d}\le e^{-2\varpi L+\eta T} \le e^{-\beta S^2/4-2S}<1. \end{equation} Since $hK_*^{1/d}=2^{J-J_h}$, this proves $J_h\ge J$ and hence admissibility. For the bias, recall that $u_j=z_j-S/2$. Using the identity $z+\mathcal Q_S(z)-S=(z-S/2)^2/S$, the degree floor, and $S-\ell_S-T\ge\gamma_*\log S$, we obtain \begin{equation}\label{eq:U12-geom} K_j^{-\vartheta}\rho^{m_j} \le CK_*^{-\vartheta}S^{-r-1/2}e^{-\vartheta u_j^2/S}. \end{equation} Here we also used that $\tau\beta=\vartheta$ and $m_j\ge\beta\mathcal Q_S(z_j)-q_s-1$. For $S\ge1$, we have $\mathcal Q_S(z_j)\le S+3u_j^2/(2S)$, which implies that $(m_j+1)^r\le CS^r(1+u_j^2/S)^r$. Absorbing this polynomial factor into half of the Gaussian exponent, we find \begin{equation}\label{eq:U12-term} K_j^{-\vartheta}(m_j+1)^r\rho^{m_j} \le CK_*^{-\vartheta}S^{-1/2}e^{-\vartheta u_j^2/(2S)}. \end{equation} Since the $u_j$ have spacing $d\log2$, their Gaussian sum is at most $C\sqrt S$. This proves \eqref{eq:U12-sum}, and \eqref{eq:U12-bias} follows from $h\le h^\circ$ and \eqref{eq:Urate-h}. For the variance, we abbreviate $x_j=C_0K_j/n=e^{z_j}/(C_DS)$. When $z\le0$ and $S\ge2$, differentiation shows that $e^z\mathcal Q_S(z)\le5S/4$. This implies that $D_jx_j\le5\beta/(4C_D)\le5/16<1/2$ for $z_j\le0$. The factorial series then has ratios of successive terms at most $1/2$, which gives $w_j\le2x_j$ and $w_j\max\{1,n/K_j\}\le2C_0$. For $00$. The lower bound in \cref{thm:main} is proved in \cref{sec:lower}, and the lower bound in \cref{thm:main2} is proved in \cref{sec:small-s}. \section{The ternary law}\label{sec:prelim-ternary} Every construction for the lower bounds uses the following response law, which already appears in the third version of this paper \cite[Section~3, equation~(3.1)]{AronowLopatto2026v3}. Fix constants \[ v_-0. \] Such a choice is possible because $v_-^20$ such that, whenever $|f|\le\rho$ and $|V-v|\le\rho$, we have $V\in(v_-,v_+)$, $q_y(f,V)\ge c$ for every response value, and $\E(Y-f)^4\le C_4$. This means that the response constraints hold uniformly on this neighborhood. Note also that the probabilities are quadratic in $f$ and affine in $V$, with $\partial_f^2q_y=2\partial_Vq_y$. Randomizing the regression mean in order to hide a change of the variance is also the device of the lower-bound proof in~\cite[Section~4.5]{Spokoiny2002}, with Gaussian regression values on smooth bumps, and of the fixed-design lower bounds of Wang et al.\ \cite[Section~6.2]{WangBrownCaiLevine2008}, who use finite Gaussian moment matching; for the ternary law, the last identity makes this matching exact for one response. \section{One score-to-risk comparison}\label{sec:prelim-risk} An accurate estimator must follow the target along any path of mixtures. The following lemma bounds how quickly its mean can move. The random variable $R_t$ in the lemma need only represent derivatives of data observables; it need not be the score of the joint law of the latent parameter and the data. Let $J_V=[v_-,\min\{v_+,\sqrt{C_4}\}]$ and $D_V=|J_V|$. The fourth-moment bound implies that every variance in $\Theta$ belongs to $J_V$. \begin{lemma}[Score-to-risk comparison]\label{lem:path-risk} For $0\le t\le\delta$, let $\mathbb E_t$ describe a mixture of $n$-observation laws with deterministic variance $V_t\in J_V$. Suppose each mixture assigns probability at most $\epsilon$ to parameters outside $\Theta$. Suppose also that, for every bounded data-only statistic $T$, $t\mapsto\mathbb E_tT$ is absolutely continuous and, almost everywhere, \[ \frac d{dt}\mathbb E_tT=\mathbb E_t[TR_t],\qquad \mathbb E_tR_t=0. \] If $B=\int_0^\delta\|R_t\|_{L^2(\mathbb E_t)}\,dt<\infty$ and $\Delta=|V_\delta-V_0|$, then \begin{equation}\label{eq:path-risk} \mathfrak r_n^2\ge\frac{\Delta^2}{(2+B)^2}-D_V^2\epsilon. \end{equation} \end{lemma} \begin{proof} We clip $T$ to $J_V$, which cannot increase its risk on $\Theta$, and set $r_T^2=\sup_{\theta\in\Theta}\mathbb E_\theta(T-V)^2+D_V^2\epsilon$. Since the exceptional states contribute at most $D_V^2\epsilon$, we have $\mathbb E_t(T-V_t)^2\le r_T^2$. Then, for $m(t)=\mathbb E_tT$, we have \[ |m(t)-V_t|\le r_T,\qquad |m'(t)|=|\mathbb E_t[(T-V_t)R_t]| \le r_T\|R_t\|_{L^2(\mathbb E_t)}. \] Bounding the two endpoint errors and integrating the derivative, we find that $\Delta\le(2+B)r_T$. It remains to square this bound and take the infimum over clipped estimators. \end{proof} The lemma combines the covariance inequality behind the information bound, which controls the derivative of the mean of a statistic \cite[Chapter~2, Theorem~5.10]{LehmannCasella1998}, with a comparison of the means of an estimator under two mixtures, as in the constrained risk inequality for mixtures that Cai and Low \cite[Section~2, Theorem~2 and Corollary~1]{CaiLow2011} develop from that of Brown and Low \cite{BrownLow1996}. Here the comparison integrates the derivative along a path of mixtures. \section{The parametric lower bound}\label{sec:prelim-parametric} \begin{proof}[Proof of the lower bound in \cref{thm:parametric}] We use $p\equiv1$, $f=0$, and the ternary law with $V_t=v-a_nt$, $0\le t\le1$, where $a_n=c_0n^{-1/2}$. If the fixed constant $c_0>0$ is sufficiently small, then the whole path is admissible for every $n$. The ordinary likelihood score of this path is \[ R_t=-a_n\sum_{i=1}^n \frac{\partial_Vq_{Y_i}(0,V_t)}{q_{Y_i}(0,V_t)},\qquad \mathbb E_tR_t^2= \frac{na_n^2}{V_t(a_Y^2-V_t)}\le Cc_0^2. \] The summands are independent and centered, and the equality follows directly by summing over the three outcomes. By decreasing $c_0$ if necessary, we can ensure that $B\le1$. Applying \cref{lem:path-risk} with $\epsilon=0$ and $\Delta=a_n$, we obtain \begin{equation}\label{eq:parametric-lower} \mathfrak r_n\ge cn^{-1/2}. \end{equation} Since the zero function belongs to every H\"older class, this bound holds for all $s>0$ and $d\ge1$. \end{proof} \section{Smooth weights}\label{sec:prelim-windows} Both nonparametric lower bounds use one family of smooth weights. As in the third version of this paper \cite[Section~1.2, Step~3]{AronowLopatto2026v3}, their squares sum to one at every point. Let $Q=(-1,1)^d$, and fix a nonnegative $\psi\in C_c^\infty(Q)$ that is positive on $[-1/2,1/2]^d$. Then $\sum_{z\in\mathbb Z^d}\psi(\cdot-z)^2$ is smooth, $1$-periodic and bounded below by a positive constant, and \begin{equation}\label{eq:lattice-windows} w=\psi\Bigl(\sum_{z\in\mathbb Z^d}\psi(\cdot-z)^2\Bigr)^{-1/2} \in C_c^\infty(Q) \quad\text{satisfies}\quad \sum_{z\in\mathbb Z^d}w(u-z)^2=1,\qquad u\in\mathbb R^d. \end{equation} We write $\mathbb T^d=\mathbb R^d/\mathbb Z^d$ for the torus and $\|x-x'\|_\infty$ for the sup-distance on it. Let $h=1/n_h$ with $n_h\ge4$ an integer, and let $\mathcal J=(\mathbb Z/n_h\mathbb Z)^d$. For $j\in\mathcal J$, the \emph{patch} $Q_j$ is the set of $x\in\mathbb T^d$ with $\|x-hj\|_\infty4s$. We use the smooth weights \eqref{eq:smooth-windows} at scale $h=1/n_h$ with $n_h\ge4$, and set \[ \eta=bh^s,\qquad \xi_j\ \hbox{independent with law}\quad \nu_t=(1-t)\delta_0+\tfrac t2(\delta_{-1}+\delta_1), \qquad 0\le t\le1. \] Let $p\equiv1$, $f_\xi=\eta\sum_jw_j\xi_j$ and $V_t=v-\eta^2t$, and let the responses follow the ternary law. By bounded overlap, we have \[ \|f_\xi\|_\infty\le C\eta,\qquad |f_\xi(x)-f_\xi(x')| \le C\eta\min\{1,\|x-x'\|/h\} \le Cb\|x-x'\|^s. \] For a sufficiently small fixed $b>0$, every state is then admissible, including at $s=1$. The periodic weights give the required extension beyond the design cube. We will apply \cref{lem:path-risk} with $\epsilon=0$ and separation $\Delta=\eta^2$. Next, we represent derivatives of expectations of bounded statistics of the data along this path, using the latent coefficients. We write $\xi^{j,a}$ for $\xi$ with coordinate $j$ set to $a$, and define \[ D_jF(\xi)=\tfrac12F(\xi^{j,1})+ \tfrac12F(\xi^{j,-1})-F(\xi^{j,0}). \] Since $\nu_t'=(\delta_{-1}+\delta_1)/2-\delta_0$, differentiating the finite product prior replaces its $j$th factor by exactly this signed measure. For a fixed sample, let $L_\xi=\prod_{i=1}^nq_{Y_i}(f_\xi(X_i),V_t)$ and \begin{equation}\label{eq:small-s-score} R_t=\sum_j r_j,\qquad r_j=\frac{D_jL_\xi- \eta^2\sum_i w_j(X_i)^2\partial_{V_i}L_\xi}{L_\xi}, \end{equation} where $\partial_{V_i}$ differentiates only the $i$th factor. Differentiating the prior and the likelihood, and then using $\sum_jw_j^2=1$, we obtain $\frac d{dt}\mathbb E_tT=\mathbb E_t[TR_t]$ for all bounded $T$. Since all sums are finite and the likelihood is uniformly positive for fixed $n$, the derivative and the integration interchange directly. Only observations in the support patch $Q_j$ enter $r_j$, because the other likelihood factors cancel. After this cancellation, summing the remaining numerator over all configurations of the responses in the patch gives zero, since $D_j1=0$ and every response probability sums to one. This implies that $\mathbb E_t[r_j\mid\xi,X_1,\ldots,X_n]=0$, and that scores from disjoint patches are conditionally orthogonal. Further, the quadratic ternary probabilities satisfy the exact identity \[ \tfrac12q_y(g+\eta w,V)+\tfrac12q_y(g-\eta w,V)-q_y(g,V) =\eta^2w^2\partial_Vq_y(g,V). \] The definition gives $r_j=0$ when $Q_j$ contains no observation. For one observation $X_i\in Q_j$, this identity, with $w=w_j(X_i)$ and with $g$ excluding the $j$th field, gives the same conclusion. Indeed, $\partial_Vq_y$ does not depend on the mean, so the variance term in \eqref{eq:small-s-score} also equals $\eta^2w^2\partial_Vq_y(g,V)$, times the likelihood factors of the other observations. This is the cancellation that makes the rate possible. Independent local perturbations of the mean also mask a change of the variance, through finite Gaussian moment matching, in the lower bound of Shen et al.\ \cite[Sections~2.2 and~5]{ShenGaoWittenHan2020}, whose design density vanishes between the bumps. Here the ternary law and the identity $\sum_jw_j^2=1$ make this cancellation exact for one observation with a design density bounded away from zero. For the remaining counts, we need only a second-order Taylor estimate, which we include for completeness. Given $k$ observations in $Q_j$, their response product is $H(a)=\prod_{i=1}^kq_{y_i}(g_i+\eta w_i a,V_t)$, where $g_i$ excludes the $j$th field. Along $-1\le a\le1$, all probabilities are at most one, and their first two derivatives in $a$ are bounded by $C\eta$ and $C\eta^2$, respectively. Then $\sup|H''|\le Ck^2\eta^2$, which gives $|\{H(1)+H(-1)\}/2-H(0)|\le Ck^2\eta^2$. The variance derivative in \eqref{eq:small-s-score} is at most $Ck\eta^2$. Since the denominator is at least $c^k$, this yields $|r_j|^2\le C^k\eta^4$ for $k\ge2$, after increasing $C$. For $\mu=nh^d\le1$, the binomial count and $|Q_j|\le Ch^d$ give \[ \mathbb E_t r_j^2\le \eta^4\sum_{k\ge2}\frac{(C\mu)^k}{k!} \le C\eta^4\mu^2. \] There are $O(h^{-d})$ patches, each meeting only a bounded number of others. By conditional orthogonality and $2ab\le a^2+b^2$, we have \[ \mathbb E_tR_t^2\le C h^{-d}\eta^4\mu^2 =Cn^2\eta^4h^d. \] We take $h=\lceil4n^{2/(d+4s)}\rceil^{-1}$. Then $\mu\le1$ because $d>4s$, and $B\le Cb^2n h^{(d+4s)/2}\le Cb^2$. We then decrease $b$ so that $B\le1$. Applying the score-to-risk comparison, we obtain \[ \mathfrak r_n\ge\eta^2/3\asymp n^{-4s/(d+4s)}. \] Together with \eqref{eq:parametric-lower} and \cref{prop:elementary-upper}, this proves \cref{thm:main2}, including the boundary $d=4s$ covered at the outset. \end{proof} \chapter{The lower bound}\label{sec:lower} Throughout this section, we assume that $s>1$ and that $d>4s$ is an integer, and the model is that of \cref{sec:main-result}; in particular, $p_-<11$ and let $d>4s$ be an integer. Let $\lambda$, $\tau$ and $\kappa$ be as in \eqref{eq:lambda} and \eqref{eq:main-constants}. There are $c>0$ and $n_0$, depending only on $(s,d,p_-,p_+,L_s,v_-,v_+,C_4)$, such that \[ \mathfrak r_n\ge c\,n^{-\lambda}e^{-\kappa\sqrt{\log n}} (\log n)^{(s-1)/(d+4)} \qquad\text{for all }n\ge n_0 . \] \end{theorem} The constants below depend only on these fixed model parameters and on the design constants in \cref{sec:LB-proof}; they never depend on the sample, the state, $n,h,M,D,N$ or the nuisance values. They also do not depend on $\mathcal U$, since the constructed means extend periodically to $\mathbb R^d$. We write $L=\log n$. The ternary law depends on the mean through $f$ and $f^2+V$. A centered perturbation of the mean imitates an increase of the variance through its diagonal covariance. We correlate the mean with the design density in order to control this covariance at the observed points, and then cancel its diagonal part by decreasing $V$. Priors under which the covariate density and the regression functions are perturbed together appear in the missing-data lower bound of Robins et al.\ \cite[Section~3]{RobinsEtAl2009EJS}; here the perturbations of the density also select the number of observations in a patch. The third version of this paper used cardinal interpolation at every count at least two and absorbed signed corrections into a common prior for the design density \cite[Section~1.2, Steps~1--2, and Appendix~B]{AronowLopatto2026v3}. Here we construct the local updates for counts at least three by cardinal interpolation. At count two, the unit and coarse kernels use ordinary selectors, while the fine kernel uses a three-point selector with bounded cell signs. We realize the signed generator through positive weights on update histories, and an $n$th-power mass tilt cancels the normalization of the likelihood. The risk criterion depends on the activity and the local score energy of the construction, and balancing these two quantities gives the stated rate. \section{The sample, patches and raw states}\label{sec:LB-setup} We use the ternary law \eqref{eq:ternary} with its constants $v,a_Y,\rho$. Let $c_q$ denote its uniform positivity constant, and set $\mathcal Y=\{-a_Y,0,a_Y\}$ and $\eta_0=\rho/2^{d+1}$. Let $\mathsf P(\mathbf f,\mathbf V;\mathbf y)=\prod_{i\le k}q_{y_i}(f_i,V_i)$, with $\mathsf P=1$ when $k=0$. When all variance coordinates equal $V$, we denote the product by $\mathsf P(\mathbf f,V)$, and the derivative $\partial_{V_i}$ is taken before this substitution. The ternary formulas give \begin{equation}\label{eq:LB-heat-one} \partial_f^2q_y=2\partial_Vq_y,\qquad \sum_{y\in\mathcal Y}\partial_Vq_y=0. \end{equation} All mark spaces are standard Borel, and densities are Borel versions of their almost-everywhere classes. We identify $[0,1)^d$ with $\mathbb T^d$ and use the sample space $\mathfrak X=(\mathbb T^d\times\mathcal Y)^n$ with the reference measure $\Lambda_n=(dx\otimes\#_{\mathcal Y})^{\otimes n}$. For a positive bounded density function $p$ and a mean $f$, we define \begin{equation}\label{eq:LB-likelihood} \begin{gathered} \mathsf U(p,f,V;\xi)=\prod_{i\le n}p(x_i)q_{y_i}(f(x_i),V),\\ \mathsf L(p,f,V;\xi)=m(p)^{-n}\mathsf U(p,f,V;\xi),\quad m(p)=\int p. \end{gathered} \end{equation} When $|f|\le\rho$ and $|V-v|\le\rho$, this is the likelihood of $n$ independent observations with design density $p/m(p)$ and the ternary law, and $P_{p,f,V}$ and $\E_{p,f,V}$ denote the corresponding law and expectation. For a patch $A$, $\xi_A$ consists of the observations that lie in the patch, kept in their original order. We take the smooth weights \eqref{eq:smooth-windows} at $h=1/n_h$ with $n_h\ge4$, and set $\mu=nh^d$ and $\omega=\log(1/\mu)$. The chart $\chi_j:Q_j\to Q=(-1,1)^d$, $x\mapsto(x-hj)/h$, has Jacobian $h^{-d}$, and $w_j=w\circ\chi_j$ on $Q_j$. Recall that each point lies in at most $2^d$ patches. For distinct labels, we write $j\sim j'$ when $Q_j\cap Q_{j'}\ne\emptyset$. Each label has exactly $3^d-1$ neighbors, given by the coordinate offsets $\{-1,0,1\}^d\setminus\{0\}$ modulo $n_h$. For integers $M\ge M_1=\lceil4/(p_+-p_-)\rceil$, let $a_M=p_-+1/M$ and $b_M=p_+-1/M$. The interval between them has width at least $(p_+-p_-)/2$, and \begin{equation}\label{eq:LB-tauM} \tau_M=\operatorname{arcosh}\frac{b_M+a_M}{b_M-a_M} =\log\frac{\sqrt{b_M}+\sqrt{a_M}}{\sqrt{b_M}-\sqrt{a_M}}, \qquad \tau\le\tau_M\le\tau+C_\tau/M. \end{equation} Here $C_\tau$ is a fixed constant, since differentiating the displayed logarithm gives $\partial_a=\sqrt b/[\sqrt a(b-a)]$ and $\partial_b=-\sqrt a/[\sqrt b(b-a)]$, and these derivatives are bounded on the segment from $(p_-,p_+)$ to $(a_M,b_M)$. For $D\ge1$, define \begin{equation}\label{eq:frame} I_D=\{\beta\in\mathbb N^d:|\beta|\le D-1\},\qquad \phi_\beta(u)=(u/8)^\beta,\quad \phi=(\phi_\beta)_{\beta\in I_D}, \end{equation} where $\mathbb N=\{0,1,\ldots\}$ and $\phi_0=1$. The fixed constant \[ C_{\rm fr}=\max\{1,\sup_\beta\|w(\cdot)(\cdot/8)^\beta\|_{C^s(\mathbb R^d)}\} \] is finite, because on the compact support of $w$, the derivatives through order $\ell+1$ are bounded by $C(1+|\beta|)^{\ell+1}8^{-|\beta|}$, and these derivative bounds control the H\"older seminorm. Let \begin{equation}\label{eq:LB-coefficient-ball} \mathbb B=\{c\in\mathbb R^{I_D}:\|c\|_1\le C_{\rm fr}^{-1}\}, \qquad F_c=\phi^\top c \end{equation} Then $|F_c|\le1$ and $\|G_c\|_{C^s(\mathbb R^d)}\le1$ for $G_c=wF_c$, extended by zero outside $Q$. A raw state $\vartheta=(p,(c_j)_j)\in\Theta_{\rm raw}$ consists of a Borel function $p:\mathbb T^d\to[a_M,b_M]$ and coefficients $c_j\in\mathbb B$. Its mean is \begin{equation}\label{eq:LB-mean} f_\vartheta=\eta_f\sum_j w_j(F_{c_j}\circ\chi_j),\qquad \eta_f=c_fh^s,\quad 00$ and $p_-\le p_\vartheta/m_\vartheta\le p_+$; if $M\ge2$, $m_\vartheta\ge1/2$. Consequently, if $C_wc_f\le L_s$, $M\ge2$, and the hypothesis of \textup{(c)} holds, $\theta(\vartheta,V)\in\Theta$. \end{lemma} \begin{proof} The periodic extension is $\tilde f_\vartheta(x)=\eta_f\sum_{z\in\mathbb Z^d}G_{c_{[z]}}(x/h-z)$. At each point, at most $2^d$ terms contribute, and for a difference at two points, at most $2^{d+1}$ terms contribute. Scaling the profile bounds, we obtain \[ \|\partial^\gamma\tilde f_\vartheta\|_\infty\le2^dc_f \quad(|\gamma|\le\ell),\qquad [D^\ell\tilde f_\vartheta]_\alpha\le2^{d+1}c_f. \] Adding the $\binom{d+\ell}d$ sup norms proves (a). These bounds also keep the parameters of the response law in the ternary neighborhood, which implies (b). For (c), we have $m_\vartheta\ge1-c_m/M>0$ and \[ \frac{p_-+1/M}{1+c_m/M}\ge p_-, \qquad \frac{p_+-1/M}{1-c_m/M}\le p_+. \] The first inequality uses that $p_-c_m\le1$, and the second uses that $p_+c_m\le1$. When $M\ge2$, we have $m_\vartheta\ge1-1/M\ge1/2$. Finally, the normalized density is Borel, the periodic extension restricts to $\mathcal U$, and the displayed error kernel is measurable because $f_\vartheta$ is continuous. These facts verify all the model conditions. \end{proof} \medskip\noindent\emph{Parameter regime.} We set $q_{\rm rsp}=\max\{2,\lfloor d/(4s)\rfloor\}$ and $C_D=(d+8)/4$. For $K\ge1$, let $\mathsf{Reg}(K)$ consist of the parameters with \begin{equation}\label{eq:LB-regime} \begin{gathered} M\ge M_1\ \text{integer},\quad D=\lceil C_DM\rceil,\quad N\ge1,\quad 0<\mu<1,\quad \omega=\log(1/\mu),\\ M/K\le\omega\le KM,\quad M^2/K\le\log N\le KM^2,\\ 0<\eta_f\le\eta_0,\quad \eta_f^{4q_{\rm rsp}}N^d\le1. \end{gathered} \end{equation} The last inequality bounds the Taylor error of the fixed response rule. Define $C_\sharp=1/(p_-^2c_q)$ and $M_1'=\max\{M_1,\lceil2/(1-p_-)\rceil,\lceil2/(p_+-1)\rceil\}$. For $M\ge M_1'$, the density interval contains the value one with a fixed margin. The construction enters the statistical argument through its activity $B_{\rm loc}$ and its local score energy $\mathcal E$. For the admissible, annihilating and centered updates constructed below, \cref{lem:score-risk} implies that \begin{equation}\label{eq:LB-cost-risk} \mathfrak r_n^2\ge c_{\rm tr}\frac{\eta_f^4}{1+B_{\rm loc}^2+h^{-d}\mathcal E} -(v_+-v_-)^2\epsilon_n. \end{equation} Here $B_{\rm loc}$ bounds the signed mark weights, $\mathcal E$ is defined in \eqref{eq:LB-score-energy}, and $\epsilon_n$ is the mass-concentration bound \eqref{eq:LB-bad-mass}. \section{Interpolation on one patch}\label{sec:interp} In this section, we construct interpolation matrices whose evaluated kernels are diagonal. Their defect and band estimates control the local score energy and the aliases. We work with the frame \eqref{eq:frame} and an integer $d>4$, and $\|A\|_{\ell^1}$ denotes the sum of the absolute values of the entries of a matrix. A \emph{separated representation} of a permutation-invariant $E:Q^k\to\mathbb R^{I_D\times I_D}$ is \[ E(U)=\int_{\mathcal Z} \Sym\!\left[\prod_{l=1}^k b_l(U_l;\zeta)\right]A(\zeta)\,\sigma(d\zeta), \] where $\mathcal Z$ is standard Borel, $\sigma$ is finite and positive, the functions are measurable, $|b_l|\le1$, and $\int\|A\|_{\ell^1}d\sigma<\infty$. Here $\Sym$ averages over the $k!$ permutations of the points. We call the last integral the \emph{cost} of the representation. If $E$ is symmetric, then symmetrizing $A$ does not increase the cost. \begin{lemma}\label{lem:interp} There are $C_A,C_I\ge1$, depending only on $d$, such that for integers $D\ge2$, $2\le k\le D$ and real $T\ge1$ there are continuous, permutation-invariant symmetric matrices $E_{k,T}(U)$ and continuous weights $\widehat w_i(T;U)\in[0,1]$ satisfying \begin{equation}\label{eq:interp-diag} \phi(U_i)^TE_{k,T}(U)\phi(U_l) =\widehat w_i(T;U)\1_{\{i=l\}}. \end{equation} The weights increase with $T$. Set $E_{k,T}=\widehat w_i(T;\cdot)=0$ for $00). \end{equation} \end{lemma} \begin{proof} Let $\sigma$ be the uniform probability measure on $S^{d-1}$, and let $\gamma_d=\int|v_1|\,d\sigma(v)$. We draw Poisson hyperplanes \cite[Section~12.4]{LastPenrose2018} with intensity $(T/\gamma_d)\sigma(dv)\,db$ on $S^{d-1}\times[-\sqrt d,\sqrt d]$, and then give their cells independent fair signs. There are finitely many hyperplanes. A mark consists of their ordered list and a sign array indexed by the halfspace patterns. With a fixed boundary convention, the evaluation is Borel on a countable union of finite-dimensional spaces. Only the law of the Poisson count depends on $T$. Conditionally on the hyperplanes, averaging over the signs leaves one precisely when every occupied cell contains an even number of the indexed points. Two points share a cell when no hyperplane separates them, and the intensity of the hyperplanes that separate them is $(T/\gamma_d)\int|v\cdot(u-u')|\,d\sigma(v)=T|u-u'|$. This implies the covariance and odd-moment statements. A related representation of the exponential kernel, through the probability that two points lie in a common cell, is derived for STIT tessellations by O'Reilly and Tran \cite[Section~4.1, Theorem~4.1]{OReillyTran2022}. Such an occupancy partition of $j\ge4$ points contains two disjoint pairs of indices, with the points of each pair lying in the same cell. For any two chosen pairs, the union of their intervals of separating offsets has length at least half the sum of their lengths. The probability that each of the two pairs lies in a common cell is therefore at most $e^{-T(|u_a-u_b|+|u_c-u_e|)/2}$, and two independent spatial integrations show that the $L^2(Q^j)$ norm of this bound is at most $C^jT^{-d}$. A union bound over the at most $j^4$ choices of pairs proves \eqref{eq:hp-moment}. \end{proof} Let $T_0=Ne^{-c_0D}$, where $c_0=\tau+1$, and suppose that $T_0>1$. For $D\ge3$, we choose fixed symmetric matrices $A_0,A_1$ with kernels \[ \phi(u)^TA_0\phi(v)=1,\qquad \phi(u)^TA_1\phi(v)=-|u-v|^2. \] Indeed, $A_0=e_0e_0^\top$, and $-|u-v|^2=2u\cdot v-|u|^2-|v|^2$. Their coefficient norms depend only on $d$. Let $\mathrm{Rsp}(A)=\mathsf C_*\otimes\nu_A$, let $\mathcal L_T$ be the law of the marked field, and define \begin{equation}\label{eq:new-pair-packet} \begin{gathered} X=\tfrac12(\delta_{-1}+\delta_1)-\delta_0,\qquad \Gamma_T^{\rm fi} =\frac{b_M^2}{\vartheta(0)^2} \Lambda_{M+2}(dz)X(d\epsilon)\mathcal L_T(d\psi),\\ p_i^e=z+\frac{\vartheta(z)}{b_M}p_i\epsilon\psi(u_i). \end{gathered} \end{equation} Note that $\|X\|=2$ and that its moments vanish at degree zero and at every odd degree and equal one at every positive even degree. The fine update stays in $[a_M,b_M]$ and has slope less than $1/4$, by the same margin bound as for \eqref{eq:LB-loc-update}. We use three separate row tags: \begin{equation}\label{eq:new-pair-rows} \begin{gathered} \Xi_{\rm unit}=\Gamma_{2,D}\otimes\mathrm{Rsp}(A_0),\qquad \Xi_{\rm co}=\int_0^{T_0}T\, \Gamma_{2,D}\otimes\mathcal L_T\otimes\mathrm{Rsp}(A_1)\,dT, \\ \Xi_{\rm fi}=\int_{T_0}^NT\, \Gamma_T^{\rm fi}\otimes\mathrm{Rsp}(A_1)\,dT. \end{gathered} \end{equation} For the first two rows, we use the ordinary update \eqref{eq:LB-loc-update} with $b_1=b_2=1$ and $b_1=b_2=\psi$, respectively. Their extraction and admissibility follow from \cref{lem:local-operator}. We include $T$ in the mark, so that the displayed integrals are finite signed measures on standard Borel spaces. We denote their direct sum by $\Xi_{\rm pair}$. \begin{lemma}[Pair activity and cancellation]\label{lem:new-pair} In $\mathsf{Reg}(K)$, for sufficiently large $M$, the sum of these rows has variation at most $CN^2e^{\tau_MM}$ and annihilates counts zero and one. At count two, after subtracting the diagonal variance derivative, its numerator is \begin{equation}\label{eq:new-pair-defect} \eta_f^2w_1w_2p_1p_2(1+Nr)e^{-Nr} \partial_{f_1}\partial_{f_2}\mathsf P, \qquad r=|u_1-u_2|. \end{equation} This expression has squared norm at most $C\eta_f^4N^{-d}$ for every $N\ge1$. At $3\le k\le D$ its contribution is $\mathsf U_k+\mathsf V_k$, where $\mathsf U_k=0$ for $k\le M$ and \begin{align} \sum_{k=M+1}^D\frac{(C_\sharp\mu)^k}{k!}\|\mathsf U_k\|_{(k)}^2 &\le C\eta_f^4e^{2\tau_MM}N^{4-d} \frac{(C\mu)^{M+1}}{(M+1)!},\label{eq:new-pair-alias}\\ \sum_{k=3}^D\frac{(C_\sharp\mu)^k}{k!}\|\mathsf V_k\|_{(k)}^2 &\le C\eta_f^4e^{2\tau_MM}\mu^4T_0^{4-2d} \le C\eta_f^4\mu^2N^{-d}.\label{eq:new-pair-field} \end{align} \end{lemma} \begin{proof} The variations of the ordinary and fine packets give \[ C\{D^2e^{\tau_MD}(1+T_0^2)+N^2e^{\tau_MM}\} \le CN^2e^{\tau_MM}: \] eventually we have $\tau_M\le c_0$, so that the coarse term is at most $CD^2N^2e^{-c_0D}$, and the unit term is absorbed because $\log N\asymp M^2$. For the ordinary pair selector, we write $a_{k,m}=\mathfrak A_k^{2,m}$. This coefficient equals one at $k=2$, vanishes for $3\le k\le m$, and is bounded by $C^ke^{\tau_Mm}$. Expanding the fine density product, we obtain the same selected pair term and, for even $4\le j\le k$, the coefficients \[ c_{k,j}=b_M^{2-j}\vartheta(0)^{-2} \int z^{k-j}\vartheta(z)^j\,d\Lambda_{M+2}(z), \qquad |c_{k,j}|\le C^ke^{\tau_MM}. \] For the fine row, the vanishing zeroth and first moments of $X$ prove annihilation at counts zero and one. The ordinary rows annihilate these counts by count selection. At count two, the response operator is $p_1p_2(\mathcal B_{A_0}+W_N(r)\mathcal B_{A_1})\mathsf P$, where \[ W_N(r)=\int_0^NT e^{-Tr}\,dT =\frac{1-(1+Nr)e^{-Nr}}{r^2},\qquad W_N(0)=N^2/2. \] At count two, $q_{\rm rsp}\ge2$ and \eqref{eq:LB-response-error} give exact extraction of the Hessian. The diagonal contribution is the heat derivative, and the cross entry of the evaluated kernel is $1-r^2W_N(r)=(1+Nr)e^{-Nr}$, which proves \eqref{eq:new-pair-defect}. The norm bound follows from $\int_{\mathbb R^d}(1+N|x|)^2e^{-2N|x|}\,dx=CN^{-d}$, and this radial calculation itself holds for every $N\ge1$. For higher counts, we define $W^{\rm fi}(r)=\int_{T_0}^NT e^{-Tr}\,dT$ and \[ \mathsf U_k=a_{k,M}\sum_{i4$, we have $\|W^{\rm fi}(|\cdot-\cdot|)\|_2\le CT_0^{2-d/2}$. Using the full response bound and the bounded densities, and summing over pairs and responses, we find that $\|\mathsf U_k\|_{(k)}\le C^k\eta_f^2e^{\tau_MM}T_0^{2-d/2}$. The factorial tail starts at $M+1$. Replacing $T_0^{4-d}$ by $N^{4-d}$ costs a factor $e^{c_0(d-4)D}$, which is absorbed into the fixed $C^{M+1}$ because $D=O(M)$. This proves \eqref{eq:new-pair-alias}. For even $j\ge4$, let $H_j=\int_{T_0}^NT M_{T,j}\,dT$. By Minkowski's inequality and \eqref{eq:hp-moment}, we have $\|H_j\|_2\le C^jj^4T_0^{2-d}$. The remaining numerator is \[ \mathsf V_k=\sum_{\substack{4\le j\le k\\j\text{ even}}} c_{k,j}\sum_{|S|=j}\prod_{i\in S}p_i\, H_j(\mathbf u_S)\mathcal B_{A_1}\mathsf P. \] It follows that $\|\mathsf V_k\|_{(k)}\le C^k\eta_f^2e^{\tau_MM}T_0^{2-d}$, and this term vanishes for $k<4$. Summing these bounds gives the first inequality in \eqref{eq:new-pair-field}. Up to a fixed multiplicative constant, the ratio of its right side to $\eta_f^4\mu^2N^{-d}$ is $e^{2\tau_MM}\mu^2N^{4-d}e^{c_0(2d-4)D}=o(1)$, since $\log N\asymp M^2$, $D=O(M)$ and $d>4$. \end{proof} \medskip\noindent\emph{One cutoff for all target counts.} Recall that $c_0=\tau+1$. For $N\ge1$ and $\mu\in(0,1)$, let $H=1+\log N$, and define \begin{equation}\label{def:rows} \begin{gathered} \ell_\circ=Ne^{-c_0D},\qquad \nu_r=N(\mu H)^{(r-2)/d}\quad(2\le r\le D),\\ E_r^{\rm co}=E_{r,\min\{\nu_r,\ell_\circ\}},\qquad E_r^{\rm fi}=E_{r,\nu_r}-E_r^{\rm co},\qquad m_r^{\rm co}=D,\quad m_r^{\rm fi}=M. \end{gathered} \end{equation} As before, $E_{r,T}=0$ for $T<1$. We omit the cases $\nu_r<1$ and the zero bands. For sufficiently large $M$ in $\mathsf{Reg}(K)$, we have $\mu H\le1$ and $\ell_\circ>1$, and every nonzero fine band has $r\le M$. The scale calculation in \cref{sec:LB-alias} proves these guard conditions. The coarse guards equal $D$, so all aliases start at count $M+1$. For each selected band with $r\ge3$ and $b\in\{{\rm co},{\rm fi}\}$, we fix a separated representation from \cref{lem:interp} with symmetric $A(\zeta)$. For $r=1$, we take $b={\rm co}$, the constant matrix $E_1=e_0e_0^T$ and the guard $D$. Its representation has one point, with $b_1\equiv1$ and $A=e_0e_0^T$. For each such representation, we form the operator of \cref{lem:local-operator} with its prescribed guard: \begin{equation}\label{eq:LB-row-measure} \Xi_{r,b}=\Xi_{r,m_r^b,E_r^b}, \end{equation} where $E_1^{\rm co}=E_1$ and $m_1^{\rm co}=D$. At target count two, the row is $\Xi_{\rm pair}$ from \eqref{eq:new-pair-rows}. The singleton, the targets $r\ge3$ and the three pair components have distinct tags. Let $E_0$ be the finite disjoint union of their mark spaces, and define \begin{equation}\label{eq:LB-signed-measure} \Xi=\Xi_{\rm pair}\oplus\bigoplus_{r\ne2,b}\Xi_{r,b},\qquad B_{\rm base}=\|\Xi\|. \end{equation} We also denote the component at target $r$ and band $b$ by $\mathrm{Row}_r^b$. All mark spaces are standard Borel, all updates are measurable and admissible, and the singleton measure is nonzero. Hence $0M$; in that case, their targets satisfy $r\le M1$, $\mu H\le1$, and $\omega>2\log H$ eventually. A nonzero fine band requires $\nu_r>\ell_\circ$, which implies that \[ r<2+\frac{dc_0D}{\omega-\log H}\le R, \] for some fixed integer $R$ depending on $K$ and the model constants. Eventually $R\le M$, so every fine guard is legal. We choose a sufficiently large fixed $C_{\rm act}$ and define \begin{equation}\label{eq:LB-alias-constants} q_a=C_{\rm act}DH(\mu H)^{a/d},\qquad a=1,2. \end{equation} Since $\log q_a\le O(\log M)-aM/(Kd)$, we may choose $M_0(K)$ so that \begin{equation}\label{eq:LB-row-small} \begin{gathered} M\ge\max\{M_1',R\},\qquad \tau_M\le c_0,\qquad \mu H\le1,\\ \log H\le M,\qquad q_2\le q_1\le\tfrac12,\qquad D^2q_1\le1 \end{gathered} \end{equation} and all pair bounds hold. Further, any fixed multiple of $De^{c_0(D+1)}$ is at most $N$ eventually. This proves that the rows are well defined. \begin{lemma}\label{lem:activity} The centered update measure has $B_{\rm loc}\le C_BN^2e^{\tau_MM}$ for a fixed $C_B$, as asserted in \cref{prop:local}(ii). \end{lemma} \begin{proof} An ordinary target-$r$ band with guard $m$ and representation cost $\mathcal C$ has variation at most $C^rD^re^{\tau_Mm}\mathcal C$. For $r\ge3$, using the interpolation cost and $\min\{\nu_r,\ell_\circ\}^2\le\ell_\circ\nu_r$, we have \[ B_r^{\rm co}\le CD^2e^{\tau_MD}N\ell_\circ q_1^{r-2},\qquad B_r^{\rm fi}\le CD^2e^{\tau_MM}N^2q_2^{r-2}. \] Since $e^{\tau_MD}\ell_\circ/N\le1$, the geometric sums of these bounds are at most $CD^2N^2(q_1+e^{\tau_MM}q_2)\le CN^2e^{\tau_MM}$. The singleton costs at most $CDe^{c_0(D+1)}\le N$. It remains to add the pair bound from \cref{lem:new-pair} and to include the fixed factor from the reference centering. No factor depending exponentially on $M$ is added to the activity. \end{proof} \begin{lemma}\label{lem:alias} The aliases vanish through count $M$ and, for a fixed $C_6$, \[ \sum_{k=M+1}^D\frac{(C_\sharp\mu)^k}{k!} \|\mathsf R_k^{\rm al}\|_{(k)}^2 \le C_6\eta_f^4e^{2\tau_MM}N^{4-d} \frac{(C_6\mu)^{M+1}}{(M+1)!}. \] This is \cref{prop:local}(vi). \end{lemma} \begin{proof} The pair alias is bounded by \eqref{eq:new-pair-alias}. All other aliases come from nonzero fine bands with $3\le r\le R\le M$, so none of them occurs for $k\le M$. Let $E$ be any symmetric matrix kernel. By the uniform response and selector bounds, we have \begin{equation}\label{eq:LB-subset-lift} \left\|\mathfrak A_k^{r,M}\sum_{|S|=r}\prod_{i\in S}p_i\, \mathcal B_{E(\mathbf u_S)}\mathsf P\right\|_{(k)} \le C^k\eta_f^2e^{\tau_MM}\|E\|_{L^2(\ell^1)}. \end{equation} Because the response and selector bounds hold uniformly in the nuisance values at each tuple, they bound the supremum in \eqref{eq:LB-norm}. Summing over subsets and responses and integrating over the unused coordinates, we obtain the factor $\binom kr k^2 3^{k/2}|Q|^{(k-r)/2}\max\{1,p_+\}^k$, which is absorbed in $C^k$. For a nonzero fine band, interpolation gives \[ \|E_r^{\rm fi}\|_{L^2(\ell^1)}^2 \le C^r\ell_\circ^{4-d}(1+\log\ell_\circ)^{r-2} \le C^rN^{4-d}e^{C M}, \] where we used $r\le R$, $D=O(M)$ and $\log H\le M$. With at most $R$ such targets, their sum satisfies \[ \sum_{k=M+1}^D\frac{(C_\sharp\mu)^k}{k!} \|\mathsf R_k^{\rm al}-\mathsf U_k\|_{(k)}^2 \le C\eta_f^4e^{2\tau_MM}N^{4-d}e^{C M} \frac{(C\mu)^{M+1}}{(M+1)!}. \] Here we used $\sum_{k\ge M+1}x^k/k!\le e^xx^{M+1}/(M+1)!$ and $\mu<1$. We absorb $e^{C M}$ into a fixed constant raised to the power $M+1$, and we combine the result with the pair alias by the squared triangle inequality. This changes only the constant in the factorial tail, not the activity exponent. \end{proof} \section{A positive path of priors}\label{sec:LB-heaps} In this section, we use the centered marks $(E,\pi,\mathsf a^\circ)$ and the updates of \cref{sec:LB-local}, with $M\ge M_1'$, $n_h\ge4$ and vacuum \(\vartheta_{\rm vac}=\vartheta^1_{\rm vac}\). Recall that updates at disjoint patches commute, that their weights are state independent, and that $|\mathsf a^\circ|\le B_{\rm loc}$. We fix the constants \begin{equation}\label{eq:LB-prior-constants} \Delta=3^d,\qquad \varrho=2^{-(\Delta+1)},\qquad c_{\rm pos}=\frac1{2+8\Delta^2},\qquad C_{\rm mgf}=\frac{2^{2d}(p_+-p_-)^2}{8}. \end{equation} We assign polynomial weights to finite marked histories of local updates, and we identify histories that differ only in the order of disjoint updates. The next lemma proves positivity and a differentiation identity. Conditional annihilation then keeps the law of the raw density fixed, and bounded differences control its mass. The interval on which positivity holds depends on $B_{\rm loc}$, but not on the number of patches. \begin{lemma}[Positive realization]\label{lem:positive-realization} If $\delta B_{\rm loc}/\varrho\le c_{\rm pos}$, there are a standard Borel history space $\mathfrak H$, a jointly measurable raw-state map $\vartheta(\mathfrak h)$ and probability measures $\nu_t$, $0\le t\le\delta$, such that, for every bounded state observable $G$ with measurable history and mark evaluations, \begin{equation}\label{eq:LB-weak-transport} \frac d{dt}\int G\,d\nu_t=\int\mathcal AG\,d\nu_t. \end{equation} The product rule holds if $G$ is also continuously differentiable in $t$ in supremum norm. The law of $\mathfrak m=\int p_{\vartheta(\mathfrak h)}$ is independent of $t$. Moreover \begin{equation}\label{eq:LB-mass-mgf} \int\mathfrak m\,d\nu_0=1,\qquad \int e^{u(\mathfrak m-1)}\,d\nu_0 \le e^{C_{\rm mgf}u^2h^d}\quad(u\in\mathbb R). \end{equation} For fixed $n\ge1$, define \begin{equation}\label{eq:LB-tilted-prior} G_n=\int\mathfrak m^n\,d\nu_0,\qquad d\sigma_t=G_n^{-1}\mathfrak m^n\,d\nu_t. \end{equation} These are probability measures with time-independent mass law, and whenever $\varepsilon\ge8C_{\rm mgf}nh^d$, \begin{equation}\label{eq:LB-mass-tail} \sigma_t(|\mathfrak m-1|>\varepsilon) \le2\exp\left\{-\frac{\varepsilon^2} {8C_{\rm mgf}h^d}\right\}. \end{equation} \end{lemma} \begin{proof} We call two patch labels dependent if they are equal or adjacent, and independent otherwise. We identify finite words that differ by exchanges of consecutive independent labels, and we call an equivalence class a \emph{shape}. The pieces of a shape are indexed by label and occurrence. We order dependent pieces as in a representative word and take the transitive closure. Shapes are the elements of the free partially commutative monoid of Cartier and Foata \cite{CartierFoata1969}, and with this order they are heaps of pieces in the sense of Viennot \cite[Section~3]{Viennot1986}. Incomparable pieces have independent labels, and any compatible ordering represents the same shape. To see this, we move the first piece of one ordering to the front of another, passing only incomparable pieces, and repeat. Applying the marked updates therefore gives a well-defined raw state. As the shapes are countable, the disjoint union of their finite products of $E$ is standard Borel. We choose a canonical word for each shape. The finite update recursion along this word makes the density, field, mass and likelihoods jointly measurable, also after a mark is appended. In the cover graph, two pieces are joined when one covers the other in this partial order. Each piece has at most $2\Delta$ neighbors. Indeed, above and below it there is at most one covering piece of each dependent label, since equal labels are comparable. Maximal pieces have distinct labels. Removing a maximal piece is the inverse of appending its label and mark, and both operations are measurable finite deletions and insertions of coordinates. We give each shape $\mathfrak g$ weight $\varrho^{|\mathfrak g|}$ and, conditionally on the shape, we give its pieces independent $\pi$ marks. This defines a finite reference measure. Indeed, for fixed multiplicities $(n_j)$, the interleavings of every adjacent pair of label chains determine the shape, and their number is at most \[ \prod_{\{j,k\}:j\sim k}2^{n_j+n_k} \le2^{(\Delta-1)\sum_jn_j}. \] Summing over the multiplicities, we see that the normalizer is bounded by $\prod_j\sum_{r\ge0}(\varrho2^{\Delta-1})^r<\infty$, where $\varrho2^{\Delta-1}=1/4$. Let $\tau_0$ be the normalized reference law, with the empty shape included. For an upper set $U$ of a history $\mathfrak h$, let \[ v_t(U)=\operatorname{Leb}\{(s_x)_{x\in U}\in[0,t]^U: s_x0$ and $1/2\le F(S)/F(S\setminus\{x\})\le3/2$. We adapt the deletion induction of Dobrushin \cite{Dobrushin1996} for hard-core partition functions, as credited and presented by Scott and Sokal \cite[Section~5.1, Theorem~5.1]{ScottSokal2005}, to the connected upper sets of a history. The induction hypothesis on proper subsets gives $F(S\setminus\max P)/F(S\setminus\{x\})\le2^{|\max P|-1}$. A connected set of $k$ vertices containing $x$ is the vertex set of a depth-first walk of length $2(k-1)$ along a spanning tree, so there are at most $(2\Delta)^{2(k-1)}$ choices. Combining these bounds, we conclude that \[ \left|\frac{F(S)}{F(S\setminus\{x\})}-1\right| \le\sum_{k\ge1}(8\Delta^2)^{k-1}\epsilon^k =\frac\epsilon{1-8\Delta^2\epsilon}\le\frac12. \] At the stated threshold, the denominator is positive. Multiplying the ratios, we find $2^{-|R|}\le\varphi_t(\mathfrak h)\le(3/2)^{|R|}$. Next, we differentiate the order volumes. A coordinate can reach the upper face $s_x=t$ only when $x$ is maximal, and removing it leaves the order volume on the remaining pieces. Because the maximal pieces of an upper set are maximal in the whole history, we have \[ \partial_t\varphi_t(\mathfrak h)=\varrho^{-1} \sum_{x\in\max(\mathfrak h)}\mathsf a^\circ(e_x) \varphi_t(\mathfrak h\setminus x). \] Appending a specified label adds one independent $\pi$ mark and multiplies the reference weight of the shape by $\varrho$. Changing variables between a history and the history obtained by deleting a maximal piece, we obtain \[ \frac d{dt}\int G\varphi_t\,d\tau_0 =\int(\mathcal AG)\varphi_t\,d\tau_0. \] Here, on histories, $\mathcal A$ appends a marked update, and on state observables it is given by \eqref{eq:LB-generator}. Since $|R|\le|\mathcal J|$, the bounds on $\varphi_t$ and its derivative are uniform for fixed $n$. We then use dominated convergence to justify this identity, its product rule and continuous differentiability in total variation. Let $H$ be a bounded observable of the raw density. For a fixed incoming state and patch, its updated value depends on the mark only through $e_{\rm d}$, and \eqref{eq:LB-annihilation} then implies that $\mathcal AH=0$. Taking $H=1$, we find that $d\nu_t=\varphi_t\,d\tau_0$ is a probability measure, and general $H$ shows that the law of the raw density is invariant. As $a_M\le\mathfrak m\le b_M$, \eqref{eq:LB-tilted-prior} then defines probability measures with a constant normalizer and an invariant mass law. It remains to prove the mass bound under $\tau_0=\nu_0$. We condition on a shape. Its marks are independent, and by the centering identity \eqref{eq:LB-row-centered}, every fresh density update has expected output one at every point of its patch, for any incoming value. Starting from the vacuum, we find that the conditional mean of the mass is one. We group all marks of each label into one block, so that the blocks are independent. Changing one block changes the final density only on its patch, because all later density updates act pointwise. The resulting change in mass is at most $c=2^dh^d(p_+-p_-)$. Revealing these blocks in turn gives centered martingale differences whose conditional ranges have length at most $c$. The second derivative of the conditional log moment-generating function of each difference is its variance under exponential tilting, which is at most $c^2/4$. Integrating twice, we bound each log moment-generating function by $u^2c^2/8$. There are $h^{-d}$ blocks, and averaging over the shape proves \eqref{eq:LB-mass-mgf}. This is the method of bounded differences of McDiarmid \cite[Lemma~5.8 and Theorem~6.7]{McDiarmid1989}, applied conditionally on the shape. We abbreviate $Z=\mathfrak m-1$ and $a=C_{\rm mgf}h^d$. By Jensen's inequality, we have $G_n\ge1$. Further, $\mathfrak m^n\le e^{nZ}$. For $\varepsilon\ge8na$, the exponential Markov inequality with $u=\varepsilon/(2a)-n\ge0$ gives \[ \sigma_0(Z>\varepsilon) \le e^{-u\varepsilon}\int e^{(n+u)Z}\,d\nu_0 \le e^{n\varepsilon-\varepsilon^2/(4a)} \le e^{-\varepsilon^2/(8a)}. \] On the event $Z<-\varepsilon$, we have $\mathfrak m^n\le1$, and the same exponential bound gives $\sigma_0(Z<-\varepsilon)\le e^{-\varepsilon^2/(4a)}$. Adding the two tails and using the invariance of the mass law, we obtain \eqref{eq:LB-mass-tail} for every $t$. \end{proof} \section{From local updates to risk}\label{sec:LB-fisher} Throughout this section, we work with the constructed rows and assume that $M\ge M_1'$, $n_h\ge4$, $0<\eta_f=c_fh^s\le\eta_0$ and $C_wc_f\le L_s$. We set \begin{equation}\label{eq:LB-priors-constants} c_m=\frac1{2p_+},\qquad \mathcal A_{\mathfrak m} =\{\mathfrak h:|\mathfrak m(\mathfrak h)-1|\le c_m/M\}, \end{equation} and define the full local score energy \begin{equation}\label{eq:LB-score-energy} \mathcal E=\sum_{k\ge0}\frac{(C_\sharp\mu)^k}{k!} \|\mathsf R_k\|_{(k)}^2, \qquad C_\sharp=\frac1{p_-^2c_q},\qquad \mu=nh^d. \end{equation} Here the numerator and the norm are given by \eqref{eq:LB-numerator} and \eqref{eq:LB-norm}, and the series is finite by the full-response bound \eqref{eq:LB-full-response}. \begin{proposition}[A lower bound from local cost]\label{lem:score-risk} Suppose $c_m/M\ge8C_{\rm mgf}\mu$, and put \begin{equation}\label{eq:LB-bad-mass} \epsilon_n=2\exp\left\{-\frac{c_m^2}{8C_{\rm mgf}M^2h^d}\right\}. \end{equation} Then \eqref{eq:LB-cost-risk} holds with a constant $c_{\rm tr}>0$ depending only on the model constants. \end{proposition} \begin{proof} Set \[ I=\sqrt{3^dh^{-d}\mathcal E},\qquad \mathcal C=1+B_{\rm loc}+I,\qquad c_* =\min\{1,\varrho c_{\rm pos},\rho/\eta_0^2\},\qquad \delta=c_*/\mathcal C. \] Then $\delta B_{\rm loc}/\varrho\le c_{\rm pos}$ and $\eta_f^2\delta\le\rho$. \Cref{lem:positive-realization} applies and gives the probability priors \eqref{eq:LB-tilted-prior} on $[0,\delta]$. Let $V_t=v-\eta_f^2t$, and let $\mathbb E_t$ denote prior expectation followed by expectation under the normalized observation law at $V_t$. The choice of the ternary neighborhood and the bound on the variance decrease ensure that $V_t\in J_V$. By \cref{lem:legality}, every state in $\mathcal A_{\mathfrak m}$ is legal, and the mass tail bound implies that $\sup_{t\le\delta}\sigma_t(\mathcal A_{\mathfrak m}^c)\le\epsilon_n$. For bounded observables $T$ of the data, the mass tilt cancels the likelihood normalization exactly: \[ \mathbb E_tT=G_n^{-1}\int\nu_t(d\mathfrak h) \int T(\xi)\mathsf U(\vartheta(\mathfrak h),V_t;\xi) \,d\Lambda_n(\xi). \] The inner integral is continuously differentiable in supremum norm for fixed $n$, since the response factors are affine in $V_t$ and $\Lambda_n(\mathfrak X)=3^n$. Using the product rule in \eqref{eq:LB-weak-transport}, we obtain \begin{equation}\label{eq:LB-observable-derivative} \frac d{dt}\mathbb E_tT=\mathbb E_t[TR],\qquad R=\frac{\mathcal A\mathsf U-\eta_f^2\partial_V\mathsf U}{\mathsf U} =\sum_j\mathsf r_j,\qquad \mathbb E_tR=0. \end{equation} Here the score identity is \eqref{eq:LB-score-sum}, and the centering follows from \cref{lem:score}. At a fixed raw state, we group the observations according to their count $k$ in $Q_j$. Using the chart ratio, $dx=h^ddu$, and the bound of one on the probability that the remaining observations lie outside the patch, we have \[ \E_{\vartheta,V}\mathsf r_j^2 \le\sum_{k=0}^n\binom nk\Bigl(\frac{h^d}{m_\vartheta}\Bigr)^k \int_{Q^k}\sum_{\mathbf y}\frac{\mathsf R_k^2}{\mathsf L_k}\,d\mathbf u \le\mathcal E. \] In the second inequality, we used $m_\vartheta\ge p_-$, $\mathsf L_k\ge(p_-c_q)^k$, $\binom nk\le n^k/k!$ and the uniform bound in \eqref{eq:LB-norm}, whose supremum is taken before integration. By \cref{lem:score}, disjoint patches have orthogonal full scores. Each label has $3^d$ dependent labels, including itself, and applying $2ab\le a^2+b^2$ to the pairs of dependent labels yields \begin{equation}\label{eq:LB-residual-second-moment} \mathbb E_tR^2\le3^d\sum_j\mathbb E_t\mathsf r_j^2 \le3^dh^{-d}\mathcal E=I^2. \end{equation} The integrated score is at most $\delta I\le c_*$, and the separation is $\eta_f^2\delta$. Since $D_V\le v_+-v_-$, \cref{lem:path-risk} implies that \[ \mathfrak r_n^2\ge \frac{c_*^2\eta_f^4}{(2+c_*)^2\mathcal C^2}-(v_+-v_-)^2\epsilon_n. \] Since $\mathcal C^2\le3^{d+1}(1+B_{\rm loc}^2+h^{-d}\mathcal E)$, this proves \eqref{eq:LB-cost-risk} with $c_{\rm tr}=c_*^2/[3^{d+1}(2+c_*)^2]$. \end{proof} \begin{lemma}[Cost of the local construction]\label{lem:raw-energy} Let $(M,D,N,\mu,\eta_f)\in\mathsf{Reg}(K)$ with $M\ge M_0(K)$, $\mu=nh^d$ and $\eta_f=c_fh^s$. Set $\mathcal B=N^2e^{\tau_MM}$. There is a fixed $C_G\ge1$, depending only on the model and design constants and $K$, such that \begin{equation}\label{eq:LB-raw-energy} 1+B_{\rm loc}^2+h^{-d}\mathcal E \le C_G\{\mathcal B^2+\mathfrak E_1+\mathfrak E_2+\mathfrak E_3\}, \end{equation} where \begin{equation}\label{eq:LB-terms} \begin{aligned} \mathfrak E_1&=n^2h^{d+4s}N^{-d},\\ \mathfrak E_2&=n h^{4s}e^{2\tau_MM}N^{4-d} \frac{(C_G\mu)^M}{(M+1)!},\\ \mathfrak E_3&=n h^{4s}\mathcal B^2 \frac{(C_G\mu)^D}{(D+1)!}. \end{aligned} \end{equation} \end{lemma} \begin{proof} The score numerators at counts zero and one vanish. For $2\le k\le D$, we square the decomposition of \cref{prop:local} into defect, alias and field terms, at the cost of a factor three. For $k>D$, using \eqref{eq:LB-full-response}, the volume $2^{dk}$ and the $3^k$ response vectors, we have $\|\mathsf R_k\|_{(k)}^2\le(3\cdot2^dC_{\rm hi}^2)^k \eta_f^4(B_{\rm loc}+1)^2$. Let $c_{\rm h}=3\cdot2^dC_{\rm hi}^2C_\sharp$. Using $\sum_{k\ge a}x^k/k!\le e^xx^a/a!$, we find \[ \begin{aligned} \mathcal E\le{}&3(C_5+C_7)\eta_f^4\mu^2N^{-d}\\ &+3C_6\eta_f^4e^{2\tau_MM}N^{4-d} \frac{(C_6\mu)^{M+1}}{(M+1)!}\\ &+e^{c_{\rm h}}\eta_f^4(B_{\rm loc}+1)^2 \frac{(c_{\rm h}\mu)^{D+1}}{(D+1)!}. \end{aligned} \] We multiply by $h^{-d}$ and use $h^{-d}\mu=n$, $h^{-d}\mu^2=n^2h^d$ and $\eta_f=c_fh^s$. As $B_{\rm loc}\le C_B\mathcal B$ and $\mathcal B\ge1$, the claimed bound holds if we take $C_G$ to be the maximum of \[ 1,\ C_6,\ c_{\rm h},\ 1+C_B^2,\ 3(C_5+C_7)c_f^4,\quad 3C_6^2c_f^4,\ e^{c_{\rm h}}c_{\rm h}c_f^4(C_B+1)^2. \] All these constants are fixed before $n$. \end{proof} \section{Parameters and proof of \texorpdfstring{\cref{thm:lower}}{Theorem 6.1}}\label{sec:LB-proof} We choose the parameters by balancing the squared activity with the interpolation contribution $\mathfrak E_1$, which leads to $N^{d+4}=n^2h^{d+4s}e^{-2\tau_MM}$. The remaining contributions are factorial tails. The choice $M\asymp\sqrt{\log n}$ makes these tails smaller than the principal cost, while the correction $-\log M$ in $\omega$ produces the logarithmic power in the lower bound. We define \[ \begin{gathered} G=\frac{d-4s}{d},\qquad \gamma=\frac{G}{d+4},\qquad \theta=\frac{d\tau}{2(s-1)},\\ m_*^2=\frac{8\gamma(s-1)}{d\tau},\qquad C_D=\frac{d+8}{4}, \qquad c_f=\min\{1,L_s/C_w\}. \end{gathered} \] Fix $K$ such that $\theta$ and $\gamma/m_*^2$ lie strictly inside $[1/K,K]$. At this fixed $K$, we take the constants of \cref{prop:local} and then $C_G$ from \cref{lem:raw-energy}. Finally, we fix a constant $C_\omega>A+2$, where \[ A=1+\log C_G+\frac{d^2-16s}{d(d+4)}\theta+\frac{2d\tau}{d+4}. \] All these constants are fixed before $n$. With $L=\log n$, we set \begin{equation}\label{eq:LB-parameters} \begin{gathered} M=\lceil m_*\sqrt L\rceil,\qquad D=\lceil C_DM\rceil,\qquad \omega_0=\theta M-\log M+C_\omega,\qquad n_h=\lceil(ne^{\omega_0})^{1/d}\rceil,\\ h=1/n_h,\quad\mu=nh^d,\quad\omega=\log(1/\mu),\quad \eta_f=c_fh^s,\quad N=\left(n^2h^{d+4s}e^{-2\tau_MM}\right)^{1/(d+4)},\\ \mathcal B=N^2e^{\tau_MM},\qquad R_{\rm low}=\eta_f^2/\mathcal B. \end{gathered} \end{equation} \begin{lemma}\label{lem:saddle} For all sufficiently large $n$, the tuple $(M,D,N,\mu,\eta_f)$ belongs to $\mathsf{Reg}(K)$, and the local construction satisfies every hypothesis of \cref{lem:score-risk}. Moreover, \[ \begin{gathered} 1+B_{\rm loc}^2+h^{-d}\mathcal E\le3C_G\mathcal B^2,\\ R_{\rm low}\asymp n^{-\lambda}e^{-\kappa\sqrt L}L^{(s-1)/(d+4)}, \qquad \epsilon_n=o(R_{\rm low}^2). \end{gathered} \] \end{lemma} \begin{proof} Rounding upward gives \begin{equation}\label{eq:LB-omega} \omega_0\le\omega\le\omega_0+d\log2,\qquad \omega=\theta M-\log M+C_\omega+O(1), \qquad\tau_MM=\tau M+O(1). \end{equation} Since $h^d=e^{-\omega}/n$, we find from the definition of $N$ that \begin{equation}\label{eq:LB-N} (d+4)\log N=GL-(1+4s/d)\omega-2\tau_MM =GL+O(\sqrt L). \end{equation} We see that $\omega/M\to\theta$ and $\log N/M^2\to\gamma/m_*^2$, which verify the regime bounds at the fixed $K$. Also, $h\le n^{-1/d}$ eventually, and since $q_{\rm rsp}\ge d/(4s)-1$, we have \[ \eta_f^{4q_{\rm rsp}}N^d \le Cn^{(d-4s)/(d+4)-4q_{\rm rsp}s/d} \le Cn^{-4\gamma}\longrightarrow0. \] All fixed thresholds on $M,n_h,\eta_f$ hold eventually, and the mass condition $c_m/M\ge8C_{\rm mgf}\mu$ follows from $M\mu\to0$. The model bound $C_wc_f\le L_s$ holds by construction, and \cref{prop:local} supplies annihilation and reference centering. For the cost, we have $\mathfrak E_1=\mathcal B^2$ exactly. With $c_t=(d^2-16s)/\{d(d+4)\}$, direct substitution gives \begin{equation}\label{eq:LB-T2} \begin{aligned} \log\frac{\mathfrak E_2}{\mathcal B^2} &=4\gamma L+(c_t-M)\omega+ \left(\frac{2d\tau_M}{d+4}+\log C_G\right)M-\log(M+1)!\\ &=4\gamma L-M(\omega+\log M)+AM+O(\log M) \le(-C_\omega+A+o(1))M. \end{aligned} \end{equation} Here $A$ is independent of $C_\omega$, while the remainder may depend on that subsequently fixed constant. In the final inequality, we used $\theta M^2\ge4\gamma L$ and $\omega\ge\omega_0$. This shows that the alias tail tends to zero relative to $\mathcal B^2$. For the high-count tail, we have \[ \log\frac{\mathfrak E_3}{\mathcal B^2} \le GL-D\omega+O(D) =(G-4\gamma C_D)L+o(L)=-4\gamma L+o(L). \] The cost bound follows from \cref{lem:raw-energy}. For the risk scale, we compute \begin{equation}\label{eq:LB-saddle} \begin{aligned} \log R_{\rm low} &=-\lambda L-\frac{2(s-1)\omega+d\tau_MM}{d+4}+O(1)\\ &=-\lambda L-\frac{2d\tau}{d+4}M +\frac{2(s-1)}{d+4}\log M+O(1)\\ &=-\lambda L-\kappa\sqrt L+\frac{s-1}{d+4}\log L+O(1). \end{aligned} \end{equation} Finally, $M^2h^d\le CL/n$ implies that $\epsilon_n\le2e^{-cn/L}=o(R_{\rm low}^2)$. \end{proof} \begin{proof}[Proof of \cref{thm:lower}] We take the parameters \eqref{eq:LB-parameters} for all sufficiently large $n$. Combining \cref{lem:saddle} with \cref{lem:score-risk}, we obtain \[ \mathfrak r_n^2\ge\frac{c_{\rm tr}}{3C_G}R_{\rm low}^2 -(v_+-v_-)^2\epsilon_n\ge\frac{c_{\rm tr}}{6C_G}R_{\rm low}^2 \] eventually. Taking square roots and using \eqref{eq:LB-saddle} proves the stated lower bound. \end{proof}