% ---------- 01-intro.tex \chapter{Introduction}\label{sec:intro} \section{Background}\label{sec:intro-background} Missing-data means, average treatment effects, and expected conditional covariances under random design are averages over the covariate distribution of functions of regressions given the covariates. Their estimation rates depend on what is known about both the regressions and the covariate distribution. In missing-response problems, Robins and Ritov~\cite{robinsritov} showed that attainable rates depend on what is known about the nuisance functions. We study \emph{rough design}, in which the covariate density is unknown and subject only to fixed upper and lower bounds, while the regressions are assumed smooth. We show that below the root-$n$ threshold, rough design reduces the minimax polynomial exponent from $4s/(4s+d)$ for a known or sufficiently smooth density to $2s/d$, where $s$ is the average smoothness of the two nuisance functions and $d$ is the covariate dimension. The threshold $s=d/4$ itself is unchanged. For a missing-data mean, let $\alpha$ and $\beta$ be the H\"older smoothness indices of the inverse observation probability and of the outcome regression. For the average product of two conditional means, which underlies the expected conditional covariance, let them be the indices of the two means. Let $s=(\alpha+\beta)/2$. Estimators of these parameters impose assumptions on a density weight $g$, which is the covariate density or its product with the observation probability, depending on the parameter. When $g$ is known or belongs to a H\"older class of smoothness $\gamma>0$ such that $\gamma/(2\gamma+d)>\{\max(\alpha,\beta)/d\}(d-4s)/(d+4s)$, higher-order influence function (HOIF) estimators attain root-mean-square error of order \begin{equation}\label{eq:known-weight-rate} n^{-1/2}\vee n^{-4s/(4s+d)}, \end{equation} where $x\vee y=\max(x,y)$. This condition and the upper bound are from~\cite[Section~3.1]{robins}, the corrected version of~\cite{robins2017}, which treats the missing-data mean there and indicates the extension to other doubly robust functionals in~\cite[Section~11.3]{robins}. Below the threshold, the rate \eqref{eq:known-weight-rate} cannot be improved even when $g$ is known, as shown in~\cite{robins2009}; see~\cite[Section~3.1]{robins}. Rates of this form are classical for quadratic functionals of a density~\cite{bickelritov,birgemassart}. For the average product of two conditional means, McGrath and Mukherjee~\cite[Corollaries~1 and~16]{mcgrath} obtain this rate for a uniform or sufficiently smooth known density, and below $s=d/4$ their results also cover an unknown density that is sufficiently smooth. The rough-design question arises in the HOIF literature~\cite{hoif2008,robins2009,li2011,robins}. Immediately after their equation~(1.3), Robins et al.~\cite[p.~338]{hoif2008} conjecture that the rate given there is minimax, up to logarithmic factors, when the nuisance functions are equally smooth and the density smoothness $\gamma$ is too small for their condition~(1.2), which for equal indices is the condition stated before \eqref{eq:known-weight-rate}. They note that, as $\gamma\to0$, this rate tends to $n^{-2s/d}\log n$, which agrees up to the logarithmic factor with the minimax rate $n^{-2s/d}$ for a constant conditional variance under an equally spaced fixed design~\cite{wangvariance,caivariance}. This limit suggests that the exponent under rough design is $2s/d$. Robins et al.~\cite[Section~3.1]{robins} also state that heuristic arguments indicate a rate slower than \eqref{eq:known-weight-rate} for smaller $\gamma$. Kennedy, Balakrishnan, and Wasserman~\cite[Section~4]{kennedy2020discussion} describe the minimax problem with an unknown density as open. They suggest that the rate with an unknown covariate density might interpolate between the known-density rate and the fixed-design rate $n^{-2s/d}$, which would suggest agreement between rough random design and fixed design at the nonsmooth extreme. For a counterfactual mean, Liu, Mukherjee, and Robins~\cite[Section~4]{liumukherjeerobins2020} describe an empirical HOIF estimator, based on~\cite{liu}, that attains $n^{-2s/d}$ up to logarithmic factors without complexity assumptions on the covariate density. McClean et al.~\cite[Section~4.2 and Theorem~1]{mcclean} construct an estimator of the expected conditional covariance with a convergence rate in probability of $n^{-2s/d}$, up to a polylogarithmic factor, when $sd/4$, with only bounds on the covariate density. Our estimator has root-$n$ risk for $s\ge d/4$ in every dimension, for all smoothness indices, and for every functional in the class of \cref{thm:generic}. In his setting, Park~\cite[Theorem~1]{park} also obtains root-$n$ risk for $s\ge d/4$, including $s=d/4$. Stabilized higher-order estimators of bilinear functionals without sample splitting were recently developed in~\cite{stabilizedhoif}, and higher-order estimators of treatment parameters defined implicitly by estimating equations in~\cite{generalhoif}. \emph{Structure-agnostic models.} In the structure-agnostic models of Balakrishnan et~al.~\cite{structureagnostic}, in which only convergence rates are imposed on black-box nuisance estimators, first-order estimators are minimax optimal for the expected conditional covariance and, when the squared pilot-error radius is at least of order $n^{-1}$, for the two quadratic functionals studied there. Jin and Syrgkanis~\cite{jinsyrgkanis2025} prove the corresponding optimality for the average treatment effect, weighted average treatment effects, and the average treatment effect on the treated. Their subsequent generalization treats linear functionals of generalized regressions and includes a structure-agnostic lower bound for the expected conditional covariance even under a partially linear outcome restriction~\cite[Theorems~6.1--6.2 and~7.5]{jinsyrgkanis2026general}. Bonvini et~al.~\cite{bonvini2024smoothsa} study classes that add smoothness constraints to structure-agnostic ones. The H\"older smoothness that we assume is what allows higher-order estimators to improve on first-order ones in the minimax rate. Minimax-rate optimality does not imply admissibility: in the structure-agnostic model, Liu, Mukherjee, and Robins~\cite[Theorems~0--2]{liumukherjeerobins2026} show, under their other stated conditions and when the squared pilot-error radius is at least a constant multiple of $n^{-1/2}$, that second-order estimators asymptotically dominate first-order ones for two quadratic functionals. For the expected conditional covariance, neither estimator asymptotically dominates the other. \emph{Moment matching and polynomial approximation.} Recall that, as in minimax estimation of nonsmooth functionals~\cite{lepski,cailow,jiaovenkathanweissman,wuyangentropy,wuyang}, we pair moment-matching priors with polynomial approximation, with the reciprocal in place of the nonsmooth function. To estimate the support size of a discrete distribution, Wu and Yang~\cite{wuyang} use best polynomial approximation of the reciprocal on an interval bounded away from zero. Our priors use the framework of the lower bounds of Robins et al.~\cite[Theorem~2.1 and Sections~3--4]{robins2009} for these functionals: independent, randomly signed smooth bumps on the cells of a partition, with cell masses that do not depend on the prior parameters, and finite response spaces. Our lower bound follows the strategy of the lower bound of Aronow and Lopatto~\cite{aronowlopatto} for a constant conditional variance under rough design. There, the covariate density is random under both priors, approximate moment matching makes the local laws of the data under the two priors close simultaneously for every number $k\le M$ of observations in a local region, the sample is Poissonized, and $M\asymp\sqrt{\log n}$ gives a factor $e^{-C\sqrt{\log n}}$ in the lower bound, for some constant $C$. Relative to~\cite{aronowlopatto}, what is new here is the random phase of the density, under which the two priors differ only at frequency $M$, and the separation through the $M$th Fourier coefficient of the reciprocal density. This identifies the constant $\kappa$ in \eqref{eq:intro-refinement}, which the lower bound of~\cite{park} also identifies for the expected conditional covariance, and the lattice phases of \cref{sec:lattice} gain a further factor $\log n$. We compare the two constructions in more detail at the start of \cref{sec:lower-ideas}. \emph{Variance estimation.} In the semiparametric model $\Var(Y\mid X)=\sigma^2$, the variance can be estimated faster under a random design without density smoothness than under an equally spaced fixed design~\cite{wangvariance,caivariance} when $01$ and $d>4s$, Aronow and Lopatto~\cite{aronowlopatto} identify the minimax exponent $2(s+1)/(d+4)$, and Dobriban et al.~\cite{dobriban2026twoscale} give a two-scale estimator that attains it and whose analysis uses only upper and lower bounds on the design density. We discuss these results in \cref{sec:product-class}; they do not apply to the nonparametric models of this paper. \emph{Discrete covariates.} In a discrete analogue, Zeng et al.~\cite{zeng} give nearly matching minimax bounds for the average treatment effect with high-dimensional discrete covariates. Their lower bound~\cite[Section~3.2]{zeng} uses priors on the covariate probabilities, propensities, and regressions of the categories that match moments up to order $\log n$, through the duality between moment matching and best polynomial approximation. They also show that the expected conditional covariance and a variance-weighted average treatment effect can be estimated faster when the covariate distribution is known, but that this knowledge may not help for the average treatment effect~\cite[Section~5]{zeng}. \emph{Concurrent work.} Park~\cite{park} studies the expected conditional covariance $\E[\Cov(A,Y\mid X)]$ for responses bounded by one in absolute value, covariates in $[0,1]^d$ with a density between known bounds $00$, the ratio $\Var_P(f(X))/\Var_P(Y)$ is the population $R^2$. In a randomized trial with $P(A=1\mid X)=1/2$ and outcome $Y_{\mathrm{obs}}\in[0,1]$, we take $Y=(2A-1)(2Y_{\mathrm{obs}}-1)$. Then $f$ is the conditional average treatment effect (CATE), so that $\Var_P(f(X))$ measures the heterogeneity of treatment effects~\cite{levy}. \end{example} The response in the trial of \cref{ex:conditional} is obtained by centering the outcome at $1/2$ in the modified-outcome construction discussed by Tian et al.~\cite[Section~2.1]{tian2014}, who attribute it to Signorovitch's thesis. Its conditional mean under balanced randomization is the CATE. Functionals of the form \eqref{eq:generic-functional} also include some instrumental-variable parameters, and we obtain further treatment-effect parameters from them through ratios and differences (\cref{cor:applications}). \section{Model and nondegeneracy}\label{sec:exactclass} For $t>0$, we write $k=\lceil t\rceil-1$ and $\vartheta=t-k\in(0,1]$. For any $f:[0,1]^d\to\mathbb R$ with continuous partial derivatives of order at most $k$, we define the H\"older norm \[ \|f\|_{C^t([0,1]^d)} =\max_{|\mathbf j|\le k}\|D^{\mathbf j} f\|_\infty +\max_{|\mathbf j|=k}\sup_{x\ne y\in[0,1]^d} \frac{|D^{\mathbf j} f(x)-D^{\mathbf j} f(y)|}{\|x-y\|_2^\vartheta}. \] Here $D^{\mathbf j}$ is the partial derivative with multi-index $\mathbf j$, of total order $|\mathbf j|$. We set $\|f\|_{C^t([0,1]^d)}=+\infty$ if these derivatives do not exist or are not continuous, and we define the H\"older ball by $\cH^t(H)=\{f:\|f\|_{C^t([0,1]^d)}\le H\}$. In particular, for integer $t$ the functions in this ball have Lipschitz partial derivatives of order $t-1$. We fix $d\ge1$, $\alpha,\beta,H,\delta>0$, and $00. \end{aligned} \end{equation} The first line states that $P_0$ satisfies \eqref{eq:generic-class} strictly, since a constant has H\"older norm equal to its absolute value. The residuals $U-a_0D$ and $V-b_0D$ have mean zero under $\pi_0$; a positive definite covariance allows us to shift their means, and hence $a$ and $b$, independently. The second alternative covers the square of one conditional mean, for which this covariance is singular. For $U=D$, the first residual vanishes and neither alternative holds. For $r>0$, we define the neighborhood \begin{equation}\label{eq:local-class} \begin{aligned} \cP_r(\pi_0)=\bigl\{P\in\cP(\alpha,\beta):{}& |w-w_0|\le r,\ |a-a_0|\le r,\ |b-b_0|\le r\\ &\text{almost everywhere on }[0,1]^d\bigr\}. \end{aligned} \end{equation} Under \eqref{eq:generic-nondegeneracy}, this class contains $P_0$, and a lower bound on it holds on every larger class. For example, suppose that \eqref{eq:generic-class} is supplemented by bounds $w_-\le w\le w_+$, $a_-\le a\le a_+$, and $b_-\le b\le b_+$ with $w_-0),\qquad \nu=q_\alpha+q_\beta. \end{equation} The dimension $q_t$ counts the coefficients of a polynomial of degree $\lceil t\rceil-1$ in $d$ variables. We have $q_t=1$ if and only if $00$, we define \begin{equation}\label{eq:subcritical-scale} \kappa=2\sqrt{\theta(1-2\theta)\tau},\qquad a_n=a_n(\theta,\tau)=n^{-\theta}(\log n)^{\theta/2}e^{-\kappa\sqrt{\log n}}. \end{equation} We call $\theta=1/2$, that is, $s=d/4$, the threshold. \begin{theorem}[Minimax bounds] \label{thm:generic} Consider the parameter \eqref{eq:generic-functional} on the class \eqref{eq:generic-class}. \begin{enumerate}[label=(\alph*)] \item There are $C>0$ and an integer $n_0\ge3$, depending only on $d,\alpha,\beta,H,\delta,g_-,g_+,M_0,|\lambda|$, and $\|W\|_\infty$, such that, if $\cP(\alpha,\beta)$ is nonempty, then for every integer $n\ge n_0$ \begin{equation}\label{eq:upper-rates} \risk\le Cn^{-1/2}\quad\text{if }\theta\ge1/2,\qquad \risk\le C\,a_n(\log n)^{\nu/2+1/4}\quad\text{if }0<\theta<1/2. \end{equation} \item Suppose that a law $\pi_0$ for $Z$ with finite support satisfies \eqref{eq:generic-nondegeneracy}, and fix $r>0$. There are $c>0$ and an integer $n_0\ge3$, depending only on $d,\alpha,\beta,H,\delta,g_-,g_+,M_0$, the response functions $D,U,V,W$, $\lambda$, $\pi_0$, and $r$, such that the following holds for every class $\cP'$ with $\cP_r(\pi_0)\subseteq\cP'\subseteq\cP(\alpha,\beta)$ and every integer $n\ge n_0$. If $\theta\ge1/2$, then $\risk(\psi,\cP')\ge cn^{-1/2}$. If $0<\theta<1/2$, then $\risk(\psi,\cP')\ge c\,a_n\log n$ and \begin{equation}\label{eq:sharp-probability} \inf_{\widehat\psi}\sup_{P\in\cP'} \Pp_P\!\left\{ |\widehat\psi-\psi(P)|\ge 2c\,a_n\log n \right\}\ge\frac14. \end{equation} \end{enumerate} In particular, if such a $\pi_0$ exists, then for all $n\ge n_0$, with $n_0$ the larger of the two thresholds, \begin{equation}\label{eq:critical-rate} c n^{-1/2}\le\risk\le C n^{-1/2}\qquad(\theta\ge1/2), \end{equation} \begin{equation}\label{eq:sharp-bracket} c\,a_n\log n\le\risk\le C\,a_n(\log n)^{\nu/2+1/4} \qquad(0<\theta<1/2), \end{equation} and the same bounds hold for $\risk(\psi,\cP')$ for every class $\cP'$ as in~(b). \end{theorem} \begin{remark}[Logarithmic form and the remaining gap]\label{rem:gap} Taking logarithms in \eqref{eq:sharp-bracket}, we obtain \begin{equation}\label{eq:log-risk} \log\risk=-\theta\log n -2\sqrt{\theta(1-2\theta)\tau\log n}+O(\log\log n), \end{equation} which is \eqref{eq:intro-refinement}. The priors of \cref{lem:lattice-priors} with $J=0$, whose phases are constant on each block, give $\risk\ge c\,a_n$ by the proof of \cref{prop:local-lower}, and with the upper bound this implies \eqref{eq:log-risk}. The lattice phases, with $J\asymp\log\log n$, gain the factor $\log n$, so that $\risk/a_n\to\infty$, and $a_n$ is not the minimax rate up to fixed constants. The bounds leave the factor $(\log n)^{\nu/2-3/4}$ unresolved, which is $(\log n)^{1/4}$ when $0<\alpha,\beta\le1$. The constants may grow as $\theta\uparrow1/2$ or as the residual covariance in \eqref{eq:generic-nondegeneracy} approaches singularity. \end{remark} \begin{remark}[Smooth and nearly uniform densities] \label{rem:smooth-design}\label{rem:smooth-density} Every density in the lower-bound constructions of this paper is $C^\infty$, but its derivatives are not bounded uniformly in $n$. All lower bounds therefore persist if we require that $p$ be $C^\infty$, provided that no uniform bound on its derivatives is imposed. When $D=1$, as in \cref{ex:conditional}, we have $w=1$ and $g=p$. Applying \cref{thm:generic} with $g_\pm=1\pm\omega$ and $\delta<1$, we find that, under the remaining conditions of \eqref{eq:generic-nondegeneracy}, even the bounds $1-\omega\le p\le1+\omega$, for any fixed $0<\omega<1$, give the exponent $2s/d$ below the threshold. The same holds, with $p_\pm=1\pm\omega$, for the rows of \cref{tab:applications} with the interval $[p_-,p_+]$. The width of the interval changes only the correction in \eqref{eq:log-risk}. \end{remark} % ---------- 03-applications.tex \chapter{Applications}\label{sec:applications} \Cref{cor:applications} and \cref{tab:applications} collect the applications of \cref{thm:generic}, with proofs in \crefrange{sec:application-reductions}{sec:additional-effects}. For an interval $I=[i_-,i_+]$ with $00$ and $\nu\ge2$, we define \begin{equation}\label{eq:app-rates} \underline a_n(\theta,\tau)=a_n(\theta,\tau)\log n,\qquad \overline a_n(\theta,\tau,\nu)=a_n(\theta,\tau)(\log n)^{\nu/2+1/4} \end{equation} if $\theta<1/2$, with $a_n(\theta,\tau)$ as in \eqref{eq:subcritical-scale}, and $\underline a_n(\theta,\tau)=\overline a_n(\theta,\tau,\nu)=n^{-1/2}$ if $\theta\ge1/2$. \begin{definition}\label{def:bracket} Let $T$ be a real parameter on a class $\cP$ of observation laws, let $\theta>0$, $\nu\ge2$, and let $I$ be an interval as above. We say that $(T,\cP)$ satisfies the \emph{lower bracket} $(\theta,I)$ if there are $c>0$ and an integer $n_0\ge3$ such that, for every integer $n\ge n_0$, we have $\risk(T,\cP)\ge c\,\underline a_n(\theta,\tau_I)$ and, if $\theta<1/2$, \begin{equation}\label{eq:app-probability} \inf_{\widehat T}\sup_{P\in\cP} \Pp_P\bigl\{|\widehat T-T(P)|\ge2c\,\underline a_n(\theta,\tau_I)\bigr\} \ge\frac14 . \end{equation} We say that $(T,\cP)$ satisfies the \emph{upper bracket} $(\theta,I,\nu)$ if there are $C>0$ and $n_0\ge3$ such that $\risk(T,\cP)\le C\,\overline a_n(\theta,\tau_I,\nu)$ for every integer $n\ge n_0$. It satisfies the \emph{bracket} $(\theta,I,\nu)$ if it satisfies both. \end{definition} The targets and classes in \cref{tab:applications} are defined below. Most of these classes bound the covariate density by \begin{equation}\label{eq:design} 0{\raggedright\hangindent=1em\hangafter=1\arraybackslash}p{#1}} \begin{tabular*}{\linewidth}{@{\extracolsep{\fill}}L{0.215\linewidth}lllll@{}} \toprule Target & Class & $s$ & $I$ & $\nu$ & Lower bound from\\ \midrule MAR mean & $\mathcal M_\eta$ & $\frac{\alpha+\beta}2$ & $[\eta,\eta^{-1}]$ & $q_\alpha+q_\beta$ & \cref{thm:generic}(b)\\ ATE, ATT, ATU & $\mathcal T_\eta$ & $\frac{\alpha+\beta_*}2$ & $[\eta,\eta^{-1}]$ & $q_\alpha+q_{\beta_*}$ & MAR family, kernel\\ MAR mean & $\mathcal M^{\rm sep}$ & $\frac{\alpha+\beta}2$ & $[p_-,p_+]$ & $q_\alpha+q_\beta$ & MAR family\\ ATE, ATT, ATU & $\mathcal T^{\rm sep}$ & $\frac{\alpha+\beta_*}2$ & $[p_-,p_+]$ & $q_\alpha+q_{\beta_*}$ & MAR family, kernel\\ \addlinespace $\E[\Cov(U,V\mid X)]$, $\E[m_Um_V]$ & $\cP_{\rm bil}(\alpha,\beta)$ & $\frac{\alpha+\beta}2$ & $[p_-,p_+]$ & $q_\alpha+q_\beta$ & \cref{thm:generic}(b)\\ $\E[\Var(Y\mid X)]$, $\E[f^2]$ & $\cP_{\rm quad}(s)$ & $s$ & $[p_-,p_+]$ & $2q_s$ & \cref{thm:generic}(b)\\ $\Var(f(X))$ & $\cP_{\rm quad}(s)$ & $s$ & $[p_-,p_+]$ & $2q_s$ & Rademacher family\\ $R^2$ & $\cP_{\rm quad}(s;v_{\min})$ & $s$ & $[p_-,p_+]$ & $2q_s$ & Rademacher family\\ $\E[f^2]$, $\Var(f(X))$ in a trial & $\cP_{\rm trial}(s)$ & $s$ & $[p_-,p_+]$ & $2q_s$ & kernel from $\cP_{\rm quad}(s)$\\ \addlinespace Wald ratio & $\mathcal I_{\eta,c_W}$ & $\frac{\alpha+\beta}2$ & $[\eta,\eta^{-1}]$ & $q_\alpha+q_\beta$ & ATE, kernel\\ Average conditional Wald ratio & $\mathcal J_\eta$ & $\frac{\alpha+\beta}2$ & $[\eta,\eta^{-1}]$ & $q_\alpha+q_\beta$ & MAR mean, kernel\\ Overlap-weighted effect & $\cP_{\rm OW}(\alpha,\beta)$ & $\min\{\frac{\alpha+\beta}2,\alpha\}$ & $[p_-,p_+]$ & $q_\alpha+q_{\min(\alpha,\beta)}$ & two families\\ \bottomrule \end{tabular*} \end{table} In every row, $\risk(T,\cP)=n^{-\min\{1/2,2s/d\}+o(1)}$, and $\nu=2$ when all smoothness indices relevant to the row are at most one. When $s2$, and we let $\mathcal M_\eta(\alpha,\beta,H)$ consist of all laws of these observations such that $Y\in[0,1]$ and \[ 1/w\in\cH^\alpha(H),\qquad b\in\cH^\beta(H),\qquad \eta\le w\le1-\eta,\qquad \eta\le wp\le\eta^{-1}. \] This is the class $\cP(\alpha,\beta)$ with $M_0=1$, $\delta=\eta$, and $[g_-,g_+]=[\eta,\eta^{-1}]$, so that $\rho=(1-\eta)/(1+\eta)$, together with the overlap condition $w\le1-\eta$. For treatment effects, we observe $O=(X,A,Y)$ with $A\in\{0,1\}$ and $Y\in[0,1]$, and we write \[ \begin{gathered} w(x)=P(A=1\mid X=x),\qquad w_1=w,\qquad w_0=1-w,\\ a_j=1/w_j,\qquad g_j=w_jp,\qquad b_j(x)=\E_P[Y\mid A=j,X=x] \qquad(j=0,1). \end{gathered} \] Here the subscripts index the treatment arms; $w_0$ and $a_0$ are not the baseline constants of \cref{sec:models}. We let $\mathcal T_\eta(\alpha,\beta_0,\beta_1,H)$ consist of all laws of $O$ such that \begin{equation}\label{eq:two-arm-class} \eta\le w\le1-\eta,\qquad a_0,a_1\in\cH^\alpha(H),\qquad \eta\le g_0,g_1\le\eta^{-1},\qquad b_j\in\cH^{\beta_j}(H)\quad(j=0,1). \end{equation} These restrictions imply $2\eta\le p\le2\eta^{-1}$. We write $\mu_j=\E_P[b_j(X)]$ and consider \begin{equation}\label{eq:treatment-functionals} \begin{gathered} T_{\rm ATE}=\mu_1-\mu_0,\\ T_{\rm ATT}=\E_P[b_1(X)-b_0(X)\mid A=1],\qquad T_{\rm ATU}=\E_P[b_1(X)-b_0(X)\mid A=0]. \end{gathered} \end{equation} Under consistency, conditional exchangeability, and positivity~\cite{hernanrobins}, these are the average effects in the population, among treated units, and among untreated units. Their effective smoothness indices are \begin{equation}\label{eq:treatment-smoothness} s_{\rm ATE}=\frac{\alpha+\min(\beta_0,\beta_1)}2,\qquad s_{\rm ATT}=\frac{\alpha+\beta_0}2,\qquad s_{\rm ATU}=\frac{\alpha+\beta_1}2. \end{equation} To bound the difference from below, we vary one arm regression and keep the other constant (\cref{sec:causal-extension}). If we bound the covariate density and the propensity separately, the correction is determined by the density interval rather than by the range of $wp$. We fix $H_0>1$ and $0<\epsilon<1/2$. We let $\mathcal M^{\rm sep}$ consist of the laws of the missing-data observations such that $Y\in[0,1]$, $p$ satisfies \eqref{eq:design}, $w\in\cH^\alpha(H_0)$, $\epsilon\le w\le1-\epsilon$, and $b\in\cH^\beta(H_0)$. We let $\mathcal T^{\rm sep}$ consist of the laws of $(X,A,Y)$ such that $Y\in[0,1]$, $p$ satisfies \eqref{eq:design}, $w=P(A=1\mid X)$ satisfies the same conditions, and $b_j\in\cH^{\beta_j}(H_0)$ for $j=0,1$. Here $\alpha$ is the smoothness of $w$ itself, equivalently of $1/w$ up to the H\"older radius, since $\epsilon\le w\le1-\epsilon$. In all four classes, the conditional law of the response is otherwise unrestricted, and no smoothness of $p$ is assumed. For the upper bounds on $\mathcal M^{\rm sep}$ and $\mathcal T^{\rm sep}$, we divide $R$ and $RY$ by a smooth pilot estimate $\widehat w$ of $w$ computed from half of the sample. This preserves the target and replaces the weight $wp$ by $wp/\widehat w$, which lies within a factor $1\pm\zeta$ of $p$ except on an event of exponentially small probability (\cref{sec:broad-inclusions}). For $\mu_j$, we apply this construction with $R=\mathbf 1_{\{A=j\}}$ and observation probability $w_j$ as in \cref{ex:mar}; the arm-$0$ pilot estimates $1-w$. The correction computed from $I_\zeta$ tends to $\kappa_p(2s/d)$ as $\zeta\to0$, which gives \eqref{eq:sep-log-risk}. \section{Conditional covariances and variances}\label{sec:product-class} We take $D=1$ as in \cref{ex:conditional}, so that $g=p$, and we fix $H_0>1$ and the density bounds \eqref{eq:design}. Let $\cP_{\rm bil}(\alpha,\beta)$ consist of all laws of $(X,U,V)$ that satisfy \eqref{eq:design}, $|U|,|V|\le1$, $m_U\in\cH^\alpha(H_0)$, and $m_V\in\cH^\beta(H_0)$, where $m_U=\E_P[U\mid X]$ and $m_V=\E_P[V\mid X]$. For $U=V=Y$, let $\cP_{\rm quad}(s)$ consist of all laws of $(X,Y)$ that satisfy \eqref{eq:design}, $|Y|\le1$, and $f\in\cH^s(H_0)$, where $f=\E_P[Y\mid X]$. We write \[ Q_2=\E_P[f(X)^2],\qquad \bar\sigma^2=\E_P[\Var_P(Y\mid X)],\qquad E_*=\Var_P(f(X))=Q_2-(\E_PY)^2 . \] For fixed $v_{\min}\in(0,1)$, we let $\cP_{\rm quad}(s;v_{\min})$ be the subclass of $\cP_{\rm quad}(s)$ on which $\Var_P(Y)\ge v_{\min}$, and on it we define the population coefficient of determination $R_*^2=E_*/\Var_P(Y)$. In a randomized trial we observe $(X,A,Y_{\rm obs})$, and we let $\cP_{\rm trial}(s)$ consist of the laws such that $A\in\{0,1\}$, $Y_{\rm obs}\in[0,1]$, $P(A=1\mid X)=1/2$, $p$ satisfies \eqref{eq:design}, and the conditional average treatment effect (CATE) $f=\E_P[Y_{\rm obs}\mid A=1,X]-\E_P[Y_{\rm obs}\mid A=0,X]$ lies in $\cH^s(H_0)$. All these classes leave the conditional law of the response otherwise unrestricted. For $E_*$ and $R_*^2$, the upper bounds combine the estimator of $Q_2$ with sample means of $Y$ and $Y^2$, and the lower bounds use hard laws on which $\Var_P(Y)\to1$. In Gaussian linear regression, the same relations through observable moments imply that the signal strength, the noise variance, and the proportion of explained variance are equally hard to estimate, up to a parametric loss~\cite[Section~1.1]{verzelengassiat2018}. In the trial, the ATE $\E_P[f(X)]$ is an observable mean, whereas $\Var_P(f(X))$ is as hard to estimate as $E_*$; only $f$ needs to be smooth. Confidence intervals for this variance in experiments, which remain valid when it is zero, are constructed in~\cite{sanchezbecerra2023}. In \cref{cor:applications}(b), the loss $\E_P[(Y-f_0(X))^2]$ is an observable mean, but the excess loss is as hard to estimate as $\bar\sigma^2$. For the expected conditional covariance, the probability bound \eqref{eq:app-probability} on $\cP_{\rm bil}(\alpha,\beta)$ shows that the polynomial exponent in McClean et al.'s~\cite{mcclean} upper bound in probability is minimax optimal. Evans and Jones~\cite[Section~1(a) and Section~2]{evansjones} studied nearest-neighbor estimators of residual moments and covariance under Lipschitz regression, regularity conditions on the design density, and additive errors independent of the covariates; Williamson et al.~\cite[Section~2.3, Example~1, and Section~3, Theorems~1--2]{williamson} studied $R^2$ inference under convergence conditions on estimated regressions. Variance estimation has also been studied in the semiparametric model $\Var(Y\mid X)=\sigma^2$~\cite{hoif2008,shen}. There, when $01$, Aronow and Lopatto~\cite{aronowlopatto} resolved the uniform minimax formulation of an open question of Robins, recorded by Richardson and Rotnitzky~\cite[Section~6]{richardsonrotnitzky}. For $s>1$ and $d>4s$, the two-scale estimator of Dobriban et al.~\cite[Section~3.3, Proposition~3.4]{dobriban2026twoscale} has root-mean-square error $O(n^{-2(s+1)/(d+4)})$, the rate of the upper bound in~\cite[Theorem~1.1]{aronowlopatto}. The proof for this estimator uses only upper and lower bounds on the covariate density~\cite[Appendix~A.5]{dobriban2026twoscale}. \Cref{cor:applications} concerns the fully nonparametric model, to which these faster rates do not apply. \section{Wald ratios and the overlap-weighted effect}\label{sec:further-applications} We observe $(X,Z,A,Y)$ with a binary instrument $Z$, a binary treatment $A$, and $Y\in[0,1]$. We write $w_z(x)=P(Z=z\mid X=x)$, $m_{Y,z}(x)=\E_P[Y\mid Z=z,X=x]$, $m_{A,z}(x)=\E_P[A\mid Z=z,X=x]$, $\Delta_Y=m_{Y,1}-m_{Y,0}$, and $\Delta_A=m_{A,1}-m_{A,0}$. When $\E_P[\Delta_A(X)]>0$, the \emph{Wald ratio} is $T_{\rm W}=\E_P[\Delta_Y(X)]/\E_P[\Delta_A(X)]$. If $P(Z=1\mid X)=1/2$, $A\le Z$, and $q=P(A=1\mid Z=1,X)$ is bounded away from zero, then $\Delta_A=q$, and the \emph{average conditional Wald ratio} is $T_{\rm CW}=\E_P[b(X)]$ with $b=\Delta_Y/q$. We fix $\alpha,\beta>0$, $H>2$, $0<\eta<1/4$, and $01$ and $0<\epsilon<1/2$. In all three classes, the conditional law of the response is otherwise unrestricted. \Cref{sec:additional-effects} recalls the causal interpretations of the three targets. The numerator of $\gamma_{\rm OW}$ has effective smoothness $(\alpha+\beta)/2$; its denominator depends only on the propensity and has effective smoothness $\alpha$. The ratio has $s_{\rm OW}=\min\{(\alpha+\beta)/2,\alpha\}$, as in \cref{tab:applications}. % ---------- 04-upper.tex \chapter{Upper bounds}\label{sec:upper-ideas} We prove \cref{thm:generic}(a) by approximating $\int abg$ by polynomials in expectations of known functions of one observation. These have unbiased estimators, so we never estimate $g$. Our estimator is a sum of U-statistics of growing order, like the HOIF estimators of~\cite{hoif2008,robins}, which correct an initial estimate by U-statistics of increasing order. Our estimator uses no initial estimate of $g$ (\cref{sec:related}). For linear models, Kong and Valiant~\cite{kongvaliant2018} developed a method that approximates a quadratic form in the inverse of an unknown covariance matrix by polynomials, whose terms are estimated without bias by averages over distinct observations. Chen, Liu, and Mukherjee~\cite[Section~3.2]{chenliumukherjee2025} use Chebyshev polynomials of the reciprocal in this way. We apply this idea cell by cell, to the local Gram matrices of $g$. We use reciprocal series to approximate projection increments over nested partitions, bound the variance of their unbiased estimators, and choose the resolution and polynomial degrees. We first reduce to $W=0$ and $\lambda=1$, so that $\psi=\int abg$. For an estimator $\widehat T$ of $T(P)=\int abg$, by the triangle inequality, we have \[ \left\|\lambda\widehat T+\frac1n\sum_{i=1}^nW(O_i)-\psi(P)\right\|_{L^2(P^n)} \le |\lambda|\,\|\widehat T-T(P)\|_{L^2(P^n)} +\|W\|_\infty n^{-1/2}, \] and the last term is at most a fixed multiple of both upper rates. We may also assume that $\alpha\le\beta$, since exchanging $(a,U,\alpha)$ with $(b,V,\beta)$ changes neither $\int abg$ nor the class $\cP(\alpha,\beta)$, and the rates are symmetric in $(\alpha,\beta)$. We set $\ell=\lceil\beta\rceil-1$, the cellwise polynomial degree of the estimator, and $s_0=\min(\alpha,\beta)$. For $h\le1$, $h^{s_0}$ bounds both $h^\alpha$ and $h^\beta$. Recall from \eqref{eq:sharp-parameters} that $\theta=(\alpha+\beta)/d$ and $\tau=-\log\rho$, and that $\nu$ is given by \eqref{eq:poly-dimension}. In this section, constants depend only on $d,\alpha,\beta,H,\delta,g_-,g_+,M_0$, and, for the general target, on $|\lambda|$ and $\|W\|_\infty$. The estimator uses these known parameters and the values of $D,U,V,W$ at the observations. The arguments use only the bounds $0\le D\le M_0$, $|U|,|V|\le M_0$, and $\|W\|_\infty$, with no property of the response space or baseline law $\pi_0$. \section{The reciprocal series}\label{sec:proof-tools} Reciprocals enter both bounds. In the upper bound, we approximate products of reciprocals of cell moments of $g$ by polynomials in those moments. In the lower bound, the separation between the target means of two priors is a Fourier coefficient of the reciprocal of an oscillating density. In \cref{lem:reciprocal}, we record the Chebyshev series of $1/x$ on a positive interval $I$, with geometric rate $\rho_I$; the parameter $\rho_I^{-1}$ is that of the Bernstein ellipse through the pole at $x=0$, as in the classical theory of Chebyshev approximation of analytic functions~\cite[Chapter~8]{trefethen}. The Chebyshev coefficients of $1/(z-x)$ are recorded explicitly in~\cite[Lemma~1, equation~(10), p.~598]{mathar2006}. For $I=[g_-,g_+]$, or any positive multiple of it, $\rho_I$ equals the $\rho$ of \eqref{eq:sharp-parameters}. \begin{lemma}[Reciprocal expansion]\label{lem:reciprocal} Fix an interval $I=[x_-,x_+]$ with $00$, and suppose that $|f(\mathsf t)|\le A$ whenever $\|\mathsf t-\mathsf t_0\|_\infty<\varepsilon$. Then, for all $r\ge1$ and $v_1,\ldots,v_r\in\mathbb R^p$, \[ |D^rf(\mathsf t_0)[v_1,\ldots,v_r]|\le A\,r!\,(e/\varepsilon)^r \prod_{i=1}^r\|v_i\|_\infty . \] \end{enumerate} \end{lemma} \begin{proof} (a) The coefficient of $z^k$ in $\mathcal K(z;M)$ has degree at most $k$ in the entries of $M$. Since a determinant is a sum of products of entries, the coefficient of $z^k$ in $\det\mathcal K(z;M_i)$ also has total degree at most $k$. The same holds for $G_k$, and $S_m$ has degree at most $m$. For spectra in $[g_-,g_+]$, diagonalizing $M_i$, we have $\det\mathcal K(z;M_i)=\prod_{l=1}^{p_i}\mathcal K(z;\lambda_{il})$ for $|z|<\rho^{-1}$, where the $\lambda_{il}$ are the eigenvalues of $M_i$. The left side of \eqref{eq:matrix-polynomial} is therefore a product of $\nu'+1$ absolutely convergent series, and in each of them the coefficient of $z^k$ has norm at most $2\rho^k$. There are $\binom{k+\nu'}{\nu'}\le(k+1)^{\nu'}$ ways to distribute an index $k$ among $\nu'+1$ factors, so that $\|G_k\|_{\rm op}\le2^{\nu'+1}(k+1)^{\nu'}\rho^k$. In particular, $\sum_kG_k$ converges absolutely, and since $\mathcal K(1;M)=\bar gM^{-1}$ and $\det(\bar gM_i^{-1})=\bar g^{p_i}/\det M_i$, its sum is $\bar g^{\nu'+1}M_0^{-1}/(\det M_1\cdots\det M_N)$. Writing $k=m+j$ with $j\ge1$ and using $m+j+1\le(m+1)(j+1)$, we bound the left side of \eqref{eq:matrix-approximation-error} by \[ \bar g^{-(\nu'+1)}\sum_{k>m}\|G_k\|_{\rm op} \le2^{\nu'+1}\bar g^{-(\nu'+1)}(m+1)^{\nu'}\rho^m \sum_{j\ge1}(j+1)^{\nu'}\rho^j, \] which is at most the right side of \eqref{eq:matrix-approximation-error}. (b) Let $r_*=\rho^{-1/2}$ and $b=(1-\sqrt\rho)^2$. If $M$ is real symmetric with spectrum in $[g_-,g_+]$, then $Q(z;M)$ is diagonal in an orthonormal eigenbasis of $M$, with diagonal entries $(1+\rho ze^{i\phi})(1+\rho ze^{-i\phi})$ for real $\phi$, and these have absolute value at least $(1-\rho r_*)^2=b$ for $|z|\le r_*$. Thus $\|Q(z;M)^{-1}\|_{\rm op}\le b^{-1}$ for $|z|\le r_*$. Let $\Lambda\ge1$ be such that $\|M_i(\mathsf t)-M_i(\mathsf t_0)\|_{\rm op}\le\Lambda\|\mathsf t-\mathsf t_0\|_\infty$ for all $i$ and all complex $\mathsf t,\mathsf t_0$, and let $\varepsilon=\min\{1,d_gb/(4\Lambda)\}$. If $\mathsf t_0\in\mathcal V$ and $\|\mathsf t-\mathsf t_0\|_\infty<\varepsilon$, then \[ \|Q(z;M_i(\mathsf t))-Q(z;M_i(\mathsf t_0))\|_{\rm op} \le2\rho r_*\Lambda\varepsilon/d_g\le b/2,\qquad |z|\le r_*, \] and the Neumann series shows that $\|Q(z;M_i(\mathsf t))^{-1}\|_{\rm op}\le2/b$ and $\|\mathcal K(z;M_i(\mathsf t))\|_{\rm op}\le2(1+\rho)/b$ there. Since the inequality $\|\mathsf t-\mathsf t_0\|_\infty<\varepsilon$ is strict, the kernels are holomorphic on a neighborhood of $|z|\le r_*$. The fixed polynomials are bounded on the bounded set of such $\mathsf t$, and we conclude that $\mathcal S(\cdot;\mathsf t)$ is holomorphic and bounded by a uniform $A$ on $|z|\le r_*$. Its Taylor coefficients are $\mathcal S_k(\mathsf t)$ by \eqref{eq:matrix-chebyshev-series}; by Cauchy's inequality, we have $|\mathcal S_k(\mathsf t)|\le Ar_*^{-k}$. Summing over $k\le m$, we obtain the bound $A/(1-\sqrt\rho)$. (c) By multilinearity we may assume that $\|v_i\|_\infty=1$ for each $i$. The polynomial $\varphi(w_1,\ldots,w_r)=f(\mathsf t_0+\sum_iw_iv_i)$ on $\mathbb C^r$ is bounded by $A$ when $\max_i|w_i|<\varepsilon/r$, and \[ D^rf(\mathsf t_0)[v_1,\ldots,v_r]=\partial_{w_1}\cdots\partial_{w_r}\varphi(0). \] By Cauchy's formula on the torus $|w_1|=\cdots=|w_r|=\varepsilon'$ with $\varepsilon'<\varepsilon/r$, this derivative is at most $A\varepsilon'^{-r}$ in absolute value. Letting $\varepsilon'\uparrow\varepsilon/r$ and using $r^r\le e^rr!$, we obtain the claim. \end{proof} \section{Projection increments and their polynomial approximations} \label{sec:increments} We bisect $[0,1]^d$ successively in coordinates $1,\ldots,d$, and we repeat this cycle. At level $j$, the partition $\mathcal C_j$ has $2^j$ equal-volume rectangles, each of aspect ratio at most two, and each cell has two children at the next level. Let $\mathbb V_j$ be the space of functions whose restriction to each cell of $\mathcal C_j$ is a polynomial of total degree at most $\ell$, and let $\Pi_j$ be the orthogonal projection onto $\mathbb V_j$ in $L^2(g)$. These spaces are nested. In dimension one, they are the spaces of piecewise polynomials on dyadic intervals that underlie the multiwavelet bases of Alpert~\cite[Section~1.1]{alpert1993}. Haar projections of this kind were combined with U-statistics to estimate integral functionals of a density in~\cite{kerkyacharianpicard1996}. We define $\psi_j=\langle\Pi_j a,\Pi_j b\rangle_g$, where $\langle f_1,f_2\rangle_g=\int f_1f_2g$. By orthogonality, \begin{equation}\label{eq:main-projection} \psi-\psi_j=\langle a-\Pi_j a,b-\Pi_j b\rangle_g, \qquad |\psi-\psi_j|\le C2^{-j\theta}. \end{equation} Indeed, the cellwise Taylor polynomials of degrees $\lceil\alpha\rceil-1$ and $\lceil\beta\rceil-1$ belong to $\mathbb V_j$ and have supremum errors $C2^{-j\alpha/d}$ and $C2^{-j\beta/d}$. The orthogonal projection minimizes the $L^2(g)$ error, and the Cauchy--Schwarz inequality bounds the inner product by $g_+$ times the product of the two supremum errors. We note that the identity in \eqref{eq:main-projection} has the form of the truncation bias of the HOIF estimators, which is the expectation of a product of two projection residuals~\cite[Theorems~3.7--3.8]{hoif2008}. The decomposition $\psi_J=\psi_0+\sum_{j=1}^J(\psi_j-\psi_{j-1})$ allows us to use different approximation orders at different resolutions. Each increment is a sum over parent cells of terms in integrals of $g$, $ag$, and $bg$ against polynomials, since the projections act cell by cell. \begin{lemma}[Polynomials for projection increments] \label{lem:projection-polynomials} Fix a level $j\ge1$, and set $K=2^{j-1}$ and $h=K^{-1/d}$. For each parent cell $C\in\mathcal C_{j-1}$ there are a known observable vector $Z_C(O)$, its moment vector $\mathsf m_C=\E_PZ_C(O)$, and a function $F_C$ such that \[ \psi_j-\psi_{j-1}=K^{-1}\sum_{C\in\mathcal C_{j-1}}F_C(\mathsf m_C), \qquad |F_C(\mathsf m_C)|\le C_0h^{\alpha+\beta}, \] and \begin{equation}\label{eq:cell-observable-bounds} \|Z_C(O)\|_2\le C_0K1_C(X),\qquad P(X\in C)\le C_0/K. \end{equation} For every integer $m\ge0$, there is a polynomial $F_{C,m}$ in these moments of total degree at most $m+\nu+2$ such that \begin{equation}\label{eq:increment-approximation} |F_C(\mathsf m_C)-F_{C,m}(\mathsf m_C)| \le C_0h^{\alpha+\beta}(m+1)^\nu\rho^m. \end{equation} At every population moment vector, that is, at $\mathsf m_C=\mathsf m_C(P)$ for every $P\in\cP(\alpha,\beta)$, it satisfies \begin{align} |F_{C,m}(\mathsf t)|&\le C_0 &&\text{for complex }\mathsf t\text{ with } \|\mathsf t-\mathsf m_C\|_\infty<1/C_0,\label{eq:complex-bound}\\ \|\nabla F_{C,m}(\mathsf m_C)\|_2&\le C_0h^{s_0}. &&\label{eq:polynomial-gradient} \end{align} The functions $Z_C$ and the coefficients of $F_{C,m}$ depend only on the class parameters, the cell, and $m$. The dimension of $Z_C$ and the constant $C_0\ge1$ do not depend on $j$, $m$, $C$, or $P$. For the base term there are a known observable vector $Z_0(O)$ and, for every $m\ge0$, a polynomial $F_{0,m}$ of degree at most $m+2$ in $\mathsf m_0=\E_PZ_0(O)$ with $|\psi_0-F_{0,m}(\mathsf m_0)|\le C_0\rho^m$, for which \eqref{eq:cell-observable-bounds}, \eqref{eq:complex-bound}, and \eqref{eq:polynomial-gradient} hold with $K=h=1$. \end{lemma} \begin{proof} \emph{Bases.} For an increment, let $C\in\mathcal C_{j-1}$ be a parent cell, of volume $1/K$ and diameter at most $2\sqrt d\,h$. Its \emph{child space} consists of the functions on $C$ whose restriction to each child is a polynomial of total degree at most $\ell$. For $\eta=a,b$, its \emph{$\eta$-parent space} consists of the polynomials on $C$ of total degree at most $\lceil\alpha\rceil-1$ for $\eta=a$ and $\ell$ for $\eta=b$, and has dimension $p_a=q_\alpha$ and $p_b=q_\beta$, respectively. Since $\alpha\le\beta$, the $a$-parent space is contained in the $b$-parent space, which is contained in the child space. For the base term, we let $C=[0,1]^d$ and $K=h=1$, we let the child space consist of the polynomials of total degree at most $\ell$, and we let both parent spaces be $\{0\}$, so that $p_a=p_b=0$. Let $z_C$ be a basis of the child space that is orthonormal for $K\dd x$ on $C$, and for $\eta=a,b$, let $L_\eta^\top z_C$ be an orthonormal basis of the $\eta$-parent space for $K\dd x$ on $C$. Then $L_\eta$ has $p_\eta$ columns, $L_\eta^\top L_\eta=I$, and the column space of $L_a$ is contained in that of $L_b$. Each parent cell is the image of one of the $d$ reference rectangles $[0,\frac12]^r\times[0,1]^{d-r}$, $0\le r0$, and every cell, \[ \|Z_C(O)\|_2\le C_0K1_C(X),\qquad P(X\in C)\le C_0/K, \] that $|F_C(\mathsf t)|\le C_0$ for complex $\mathsf t$ with $\|\mathsf t-\mathsf m_C\|_\infty<1/C_0$, and that $\|\nabla F_C(\mathsf m_C)\|_2\le C_0h^{s_0}$. Then, with $C_v=2e^2C_0^7$, \begin{equation}\label{eq:variance-bound} \Var\widehat T\le\frac{C_vh^{2s_0}}n +\frac1n\sum_{r=2}^RC_v^r\,r!\left(\frac Kn\right)^{r-1}. \end{equation} If moreover $C_vRK/n\le1/2$, then \begin{equation}\label{eq:small-resolution-variance} \Var\widehat T\le\frac{C_vh^{2s_0}}n+\frac{4C_v^2K}{n^2}. \end{equation} \end{lemma} \begin{proof} We write $\mathsf m=(\mathsf m_C)_C$ and $H_r=D^rT(\mathsf m)[Z(O_1'),\ldots,Z(O_r')]$. By \cref{lem:lift} and \eqref{up:centering-contraction}, the term of order $r$ in the variance of $\widehat T$ is at most $\E|H_r|^2/\{r!(n)_r\}$, and the terms with $r>R$ vanish. For observations $o_l=(x_l,\zeta_l)$, we have \[ D^rT(\mathsf m)[Z(o_1),\ldots,Z(o_r)] =K^{-1}\sum_CD^rF_C(\mathsf m_C)[Z_C(o_1),\ldots,Z_C(o_r)]. \] Since the cells are disjoint and $Z_C$ is supported on $C$, the summand for $C$ vanishes unless $x_1,\ldots,x_r\in C$, and at most one summand is nonzero. Thus \[ |D^rT(\mathsf m)[Z(o_1),\ldots,Z(o_r)]|^2 =K^{-2}\sum_C|D^rF_C(\mathsf m_C)[Z_C(o_1),\ldots,Z_C(o_r)]|^2. \] For $r\ge2$, by \cref{lem:reciprocal-bounds}(c) with $\varepsilon=1/C_0$ and $A=C_0$, together with $\|Z_C\|_\infty\le\|Z_C\|_2\le C_0K$, the summand for $C$ is at most $C_0^2(r!)^2(eC_0)^{2r}(C_0K)^{2r}$ on the event that $X_1',\ldots,X_r'\in C$, which has probability at most $(C_0/K)^r$. Summing over the $K$ cells, we obtain \[ \E|H_r|^2 \le C_0^2(r!)^2(e^2C_0^5)^rK^{r-1} \le(r!)^2(e^2C_0^7)^rK^{r-1}. \] For $r=1$, by the gradient bound, we have $|\nabla F_C(\mathsf m_C)\cdot Z_C(o)|^2\le C_0^4h^{2s_0}K^21_C(x)$, and since the cells partition $[0,1]^d$, we obtain $\E|H_1|^2\le C_0^4h^{2s_0}\le C_vh^{2s_0}$. Since $2R\le n$, we have $(n)_r\ge(n/2)^r$ for $r\le R$, and these bounds imply \eqref{eq:variance-bound}. In \eqref{eq:variance-bound}, the ratio of the $(r+1)$st to the $r$th term is $C_v(r+1)K/n\le C_vRK/n$ for $r+1\le R$. If this is at most $1/2$, the sum over $r\ge2$ is at most twice its first term $2C_v^2K/n^2$, which proves \eqref{eq:small-resolution-variance}. \end{proof} For a fixed order $r$, variances of order $n^{-1}\max\{1,(K/n)^{r-1}\}$, with $K$ the number of basis functions, were obtained for the terms of the HOIF estimators in~\cite[Theorems~3.20--3.21]{hoif2008}. In \eqref{eq:variance-bound}, we keep the dependence on the order explicit, including the factorial, since our orders grow with $n$. \section{Choice of the degrees}\label{sec:tuning} For a level $j\ge1$ and an order $m\ge0$, we set $K_j=2^{j-1}$, the number of parent cells, and $h_j=K_j^{-1/d}$, and we let $T_{j,m}(P)=K_j^{-1}\sum_{C\in\mathcal C_{j-1}}F_{C,m}(\mathsf m_C)$ be the order-$m$ polynomial approximation of $\psi_j-\psi_{j-1}$. For the base term, we let $K_0=h_0=1$ and $T_{0,m}(P)=F_{0,m}(\mathsf m_0)$. In particular, $K_0=K_1=1$. We write $\widehat T_{j,m}$ for the lift of $T_{j,m}$ and $C_v=2e^2C_0^7$, with $C_0$ from \cref{lem:projection-polynomials}. By that lemma, if $R=m+\nu+2$ satisfies $2R\le n$, then the hypotheses of \cref{lem:variance} hold for $T_{j,m}$ with $K=K_j$, $h=h_j$, and this $R$, including the base term. Since $r!\le R^{r-1}$ for $2\le r\le R$, the bound \eqref{eq:variance-bound} then implies \begin{equation}\label{up:variance-levels} \Var\widehat T_{j,m}\le\frac{C_v}n \Bigl(h_j^{2s_0}+\sum_{r=1}^{R-1}y^r\Bigr),\qquad y=\frac{C_vRK_j}n. \end{equation} For a terminal level $J$ and orders $m_0,\ldots,m_J$ with $m_j+\nu+2\le n/2$, the estimator of $T(P)=\int abg$ is $\widehat T=\sum_{j=0}^J\widehat T_{j,m_j}$, the lift of $\sum_{j=0}^JT_{j,m_j}$. Its mean is $\sum_{j=0}^JT_{j,m_j}(P)$ by \cref{lem:lift}. By the telescoping decomposition of $\psi_J$, \eqref{eq:main-projection}, \eqref{eq:increment-approximation}, the base-term bound, and $h_j^{\alpha+\beta}=K_j^{-\theta}$, we obtain \begin{equation}\label{eq:bias-levels} |\E\widehat T-T(P)|\le C2^{-J\theta} +C_0\sum_{j=0}^Jf_j(m_j),\qquad f_j(m)=K_j^{-\theta}(m+1)^\nu\rho^m. \end{equation} By the $L^2$ triangle inequality, the standard deviation of $\widehat T$ is at most the sum of the standard deviations of the $\widehat T_{j,m_j}$. We choose the order on each level as the smaller of two values. The first, a bias rule, makes the approximation errors decrease geometrically from the terminal level toward the coarse levels. The second, a variance rule, is needed only below the threshold. We keep high orders only at coarse resolutions, and in this respect our estimator resembles that of Robins et al.~\cite[Section~9, (9.6)]{robins}. Their kernel of each order three or higher is a sum over the contiguous pairs of projection kernels in its product; in each summand, that pair is truncated below a hyperbola in the indices of the basis functions, and all other kernels are truncated at $n$ basis functions. On a level with $K$ parent cells and degree $R$, the bound \eqref{up:variance-levels} is at most about $n^{-2\theta}$ if $y^R\le n^{1-2\theta}$. If $x=\log y>0$, this limits the degree to about $(1-2\theta)\log n/x$. Since $K=ne^x/(C_vR)$, the level error $K^{-\theta}\rho^m$ at this degree is about $n^{-\theta}(C_vR)^\theta\exp\{-\theta x-\tau(1-2\theta)\log n/x\}$. The minimum of $\theta x+\tau(1-2\theta)\log n/x$ over $x>0$ is $\kappa\sqrt{\log n}$, and the relevant levels have $R$ of order $\sqrt{\log n}$. The three factors of $(\log n)^{\theta/2+\nu/2+1/4}$ in $a_n(\log n)^{\nu/2+1/4}/(n^{-\theta}e^{-\kappa\sqrt{\log n}})$ come from $(C_vR)^\theta$, which reflects the factorials in the variance, from the factor $(m+1)^\nu$ in \eqref{eq:increment-approximation}, and from the sum over levels. The next lemma bounds this sum. It is the only source of the factor $(\log n)^{1/4}$ in the upper bound. \begin{lemma}[Window sum]\label{lem:window-sum} Let $\theta,\tau,B,\bar X>0$, and let $S\subset(0,\bar X]$ be a finite set whose points are pairwise at distance at least $\log2$. Then \[ \sum_{x\in S}\exp\Bigl(-\theta x-\frac{\tau B}x\Bigr) \le\Bigl(2+\frac{(\pi\bar X/\theta)^{1/2}}{\log2}\Bigr) e^{-2\sqrt{\theta\tau B}}. \] \end{lemma} \begin{proof} Let $\xi=\sqrt{\tau B/\theta}$. Completing the square, we obtain \[ \theta x+\frac{\tau B}x=2\sqrt{\theta\tau B}+\frac{\theta(x-\xi)^2}x \ge2\sqrt{\theta\tau B}+\frac{\theta(x-\xi)^2}{\bar X}, \qquad00$ such that $\tau A\ge\theta\log2+1$. We write $K_*=2^J$ for the terminal resolution and $l_j=J-j+1$ for $0\le j\le J$. Then $K_j=K_*2^{-l_j}$ for $j\ge1$ and $K_0=1=2K_*2^{-l_0}$, so that $K_*2^{-l_j}\le K_j\le2K_*2^{-l_j}$ for every $j$. The \emph{bias rule} is the order $\lceil Al_j\rceil$. Since $(\lceil Al_j\rceil+1)^\nu\le(Al_j+2)^\nu$ and $\rho^{\lceil Al_j\rceil}\le e^{-\tau Al_j}$, we have \begin{equation}\label{eq:bias-rule-sum} \sum_{j=0}^Jf_j(\lceil Al_j\rceil) \le K_*^{-\theta}\sum_{l\ge1}(Al+2)^\nu e^{-l}=C_AK_*^{-\theta}, \end{equation} because the $l_j$ are distinct positive integers and $\tau A-\theta\log2\ge1$. In each case below, the order at level $j$ is $m_j=\min\{\lceil Al_j\rceil,M_j\}$ for a \emph{variance rule} $M_j\in\{0,1,2,\ldots\}\cup\{\infty\}$, and we call level $j$ \emph{capped} if $M_j<\lceil Al_j\rceil$. Then \eqref{eq:bias-levels}, \eqref{eq:bias-rule-sum}, and $2^{-J\theta}=K_*^{-\theta}$ imply that \begin{equation}\label{eq:bias-two-rules} |\E\widehat T-T(P)|\le CK_*^{-\theta} +C\sum_{j\text{ capped}}f_j(M_j). \end{equation} The degree $R_j=m_j+\nu+2$ satisfies $R_j\le Al_j+\nu+3$. In both cases below $J\le CL$, so that $R_j\le CL\le n/2$, and \eqref{up:variance-levels} applies to level $j$ with $y=y_j=C_vR_jK_j/n$. \emph{The case $\theta\ge1/2$.} Let $c_K\in(0,1]$ satisfy $c_K\le\{4C_v\sup_{l\ge1}(Al+\nu+3)2^{-l}\}^{-1}$, let $J=\lfloor\log_2(c_Kn)\rfloor\ge0$, and let $M_j=\infty$ for all $j$. Since $K_*\le c_Kn$, we have $C_vR_jK_j/n\le2C_vc_K(Al_j+\nu+3)2^{-l_j}\le1/2$, and \eqref{eq:small-resolution-variance} gives $\Var\widehat T_{j,m_j}\le C_vh_j^{2s_0}/n+4C_v^2K_j/n^2$. Since $\sum_{j\ge0}h_j^{s_0}\le1+\sum_{i\ge0}2^{-is_0/d}$ and $\sum_{j=0}^JK_j^{1/2}\le CK_*^{1/2}\le Cn^{1/2}$, the standard deviation of $\widehat T$ is at most $Cn^{-1/2}$. Since no level is capped, \eqref{eq:bias-two-rules} implies that the bias is at most $CK_*^{-\theta}\le C(c_Kn/2)^{-\theta}\le Cn^{-1/2}$. \emph{The case $0<\theta<1/2$.} Let $\Delta=1-2\theta$, so that $\kappa=2\sqrt{\theta\Delta\tau}$. We take \[ J=\Bigl\lceil\log_2n+\frac{\kappa\sqrt L}{\theta\log2}\Bigr\rceil, \qquad\text{so that}\qquad K_*^{-\theta}\le n^{-\theta}e^{-\kappa\sqrt L}\le\bar r_n, \quad\log\frac{K_*}n\le\frac\kappa\theta\sqrt L+\log2. \] We call level $j$ \emph{high} if $K_j>n/L^2$, and \emph{low} otherwise. The levels $0$ and $1$ are low. On a high level, $2^{l_j}=K_*/K_j0,\qquad M_j=\infty\quad\text{if }x_j\le0. \] If $x_j>0$, then $K_j>n/(C_v\bar R)\ge n/L^2$, so that level $j$ is high, and $00$, and we have $M_j+1\le\lceil Al_j\rceil\le\bar R$, so that $(M_j+1)^\nu\le CL^{\nu/2}$. Moreover $\rho^{M_j}\le e^{\tau(\nu+3)}e^{-\tau B/x_j}$ and $K_j^{-\theta}=n^{-\theta}e^{\theta c_L}e^{-\theta x_j} \le Cn^{-\theta}L^{\theta/2}e^{-\theta x_j}$, since $e^{\theta c_L}=(C_v\bar R)^\theta\le CL^{\theta/2}$. Combining these bounds, we find \[ f_j(M_j)\le Cn^{-\theta}L^{\theta/2+\nu/2} \exp\Bigl(-\theta x_j-\frac{\tau B}{x_j}\Bigr). \] The points $x_j$ of distinct capped levels lie in $(0,\bar X]$ and differ by nonzero multiples of $\log2$. By \cref{lem:window-sum} and $\bar X\le C\sqrt L$, the sum of these bounds is at most $Cn^{-\theta}L^{\theta/2+\nu/2+1/4}e^{-2\sqrt{\theta\tau B}}$. Finally, \[ \kappa\sqrt L-2\sqrt{\theta\tau B} =\frac{2\sqrt{\theta\tau}\,(\Delta L-B)}{\sqrt{\Delta L}+\sqrt B} \le\frac{4\kappa\sqrt{\theta\tau}}{\sqrt\Delta}=8\theta\tau, \] so that the bias is at most $C\bar r_n$. \emph{Variance.} Since $h_j\le1$, \eqref{up:variance-levels} implies that $\Var\widehat T_{j,m_j}\le C_vR_jn^{-1}\max\{1,y_j\}^{R_j-1}$. On a low level, we have $R_j\le CL$ and $K_j\le n/L^2$, so that $y_j\le CL^{-1}\le1$, and the variance is at most $CL/n$. There are at most $J+1\le CL$ low levels, and their standard deviations sum to at most $CL^{3/2}n^{-1/2}$. On a high level, we have $R_j\le\bar R$, so that $y_j\le C_v\bar RK_j/n=e^{x_j}$. If $x_j\le0$, then $y_j\le1$. If $x_j>0$, then $R_j\le M_j+\nu+2\le B/x_j$, and therefore $y_j^{R_j-1}\le e^{R_jx_j}\le e^B$. Since $e^B\ge1$, in both cases the variance is at most $C_v\bar Rn^{-1}e^B=C_v\bar Rn^{-2\theta}e^{-2\kappa\sqrt L}$, and the standard deviations of the at most $C\sqrt L$ high levels sum to at most $CL^{3/4}n^{-\theta}e^{-\kappa\sqrt L}$. Since $\nu\ge2$, this is at most $C\bar r_n$, and so is $CL^{3/2}n^{-1/2}$, since $L^{3/2}e^{-\Delta L/2+\kappa\sqrt L}$ is bounded. With $\lambda$ and the observable mean restored as at the start of the section, this proves \cref{thm:generic}(a), and \eqref{eq:basic-upper} follows because $L^{\theta/2+\nu/2+1/4}e^{-\kappa\sqrt L}$ is bounded below the threshold. All constants depend only on the parameters listed in \cref{thm:generic}(a). The estimator is an explicit function of the covariates $X_i$ and the values of $D,U,V,W$ at the observations, and it involves no baseline law. It remains to prove the uniformity assertions. For compact $\mathcal G$, $g_-$ and $d_g$ are bounded below, $g_+$ is bounded above, and $\rho$ lies in a compact subinterval of $(0,1)$. The approximation constants, the radii of the complex neighborhoods, and the polynomial bounds in \cref{lem:matrix-approximation,lem:projection-polynomials} can therefore be chosen uniformly, which gives common $C_0$ and $C_v$. We take $A$ with $A\inf_{\mathcal G}\tau\ge\theta\log2+1$, and then we choose $c_K$ uniformly. Below the threshold, $\Delta>0$ is fixed, and $\tau$ and $\kappa$ have positive lower and finite upper bounds on $\mathcal G$, so that $\bar R\le C\sqrt L$ and $\bar X\le C\sqrt L$ uniformly. Then every large-sample condition, rate comparison, and finite-sample enlargement above holds beyond one common $n_1$ and with one constant $C$. The reduction at the start of the section costs at most $M_Wn^{-1/2}$. Dependence of the observables on $X$ changes none of the moment identities or observable bounds; equivalently, we may regard $(X,Z)$ as the response. \end{proof} % ---------- 05-lower.tex \chapter{Lower bounds} \label{sec:lower-ideas} \label{sec:generic-lower} We prove \cref{thm:generic}(b) in this section and the next two. The root-$n$ bound follows from a two-point argument (\cref{sec:two-point}). Below the threshold, we use two fuzzy hypotheses~\cite[Section~2.7.4]{tsybakov}: the prior mixtures are close in total variation, while their target means differ by more than the prior fluctuations. In a finite-response model, we bound the Hellinger distance between Poissonized mixtures (\cref{lem:poisson-comparison}) and reduce the lower bound to coefficient and separation estimates (\cref{prop:reduction}). \Cref{sec:lattice} constructs priors giving $c\,a_n$ with block-constant phases and $c\,a_n\log n$ with lattice phases. \Cref{sec:generic-embedding} embeds them in the generic model; \cref{sec:hard-laws} establishes localization and completes the proof. For the construction below the threshold, we follow the strategy of the lower bound of Aronow and Lopatto~\cite[Section~1.2]{aronowlopatto} for estimating a constant conditional variance under rough design. There, the covariate density is random under both priors, approximate moment matching makes the local laws of the data under the two priors close simultaneously for every number $k\le M$ of observations in the same local region, and the sample is Poissonized, with a reweighting of the priors that cancels the exponential likelihood factor arising from the random total mass of the perturbed density. A matching order $M\asymp\sqrt{\log n}$, balanced against regions that hold about $e^{-C(\log n)/M}$ observations on average, gives a factor $e^{-C\sqrt{\log n}}$. Here, too, a random density hides the difference between the priors from every configuration with fewer than about $M$ observations in a pair of blocks, and we Poissonize the sample and take $M\asymp\sqrt{\log n}$ to balance the two losses. Relative to~\cite{aronowlopatto}, the hiding and the separation take a new form. Our priors share the law of a random phase $\Theta$ of the density and differ only through sign correlations at frequency $M$ in this phase. The two blocks of each pair oscillate with opposite signs, so that the mass of each pair is fixed and no reweighting is needed. The target sees the difference between the priors through the $M$th Fourier coefficient of the reciprocal density, which identifies the constant $\kappa$. The lattice phases of \cref{sec:lattice} gain a further factor $\log n$. \section{Divergences and a two-point bound}\label{sec:two-point} For probability laws $\mathbb P,\mathbb Q$, we write \[ \TV(\mathbb P,\mathbb Q)=\sup_E|\mathbb P(E)-\mathbb Q(E)|,\qquad d_{\rm H}^2(\mathbb P,\mathbb Q) =\int(\sqrt{\dd\mathbb P}-\sqrt{\dd\mathbb Q})^2 . \] Any test between $\mathbb P$ and $\mathbb Q$ has error probabilities that sum to at least $1-\TV(\mathbb P,\mathbb Q)$. We use the inequality $\TV\le d_{\rm H}$ and the fact that the affinity $1-d_{\rm H}^2/2$ is multiplicative over product laws~\cite[Section~2.4 and Lemma~2.3]{tsybakov}. Since $\prod_i(1-x_i)\ge1-\sum_ix_i$ for $x_i\in[0,1]$, the squared Hellinger distance between two product laws is at most the sum of the squared Hellinger distances between their factors. The next lemma applies the two-point method~\cite[Sections~2.3 and~2.4.2]{tsybakov} along a one-parameter path. \begin{lemma}[A one-dimensional testing bound]\label{lem:parametric-path} Suppose a class $\mathcal P$ contains probability laws $P_t$ for $t$ in an open interval $J$, with \[ \frac{\dd P_t}{\dd P_*}=f_0+t f_1,\qquad f_0+t f_1\ge c_f>0,\qquad |f_1|\le C_f \quad P_*\text{-almost surely}, \] where $P_*$ is a probability law and $c_f,C_f$ are independent of $t\in J$. If $t\mapsto T(P_t)$ is differentiable at an interior point $t_0$ and its derivative there is nonzero, then $\risk(T,\mathcal P)\ge c n^{-1/2}$ for all sufficiently large $n$. The constant $c>0$ is independent of $n$. \end{lemma} \begin{proof} Let $f_t=f_0+tf_1$. Since $(\sqrt{f_t}-\sqrt{f_s})^2=(t-s)^2f_1^2/(\sqrt{f_t}+\sqrt{f_s})^2$ and $f_t,f_s\ge c_f$, we have $d_{\rm H}^2(P_t,P_s)\le C_f^2(t-s)^2/(4c_f)$. Let $t_\pm=t_0\pm a/\sqrt n$, where $a>0$ is fixed and small enough that $C_fa/\sqrt{c_f}\le1/2$; for all large $n$, these points lie in $J$. By the subadditivity of $d_{\rm H}^2$ over products, we have $d_{\rm H}^2(P_{t_+}^{\otimes n},P_{t_-}^{\otimes n})\le C_f^2a^2/c_f\le1/4$, so that the total variation distance between the product laws is at most $1/2$. Since the derivative of the target at $t_0$ is nonzero, the separation $D_n=|T(P_{t_+})-T(P_{t_-})|$ is at least $c'n^{-1/2}$ for a constant $c'>0$ and all sufficiently large $n$. We threshold any estimator at the midpoint of the two target values. The two error probabilities of the resulting test sum to at least $1/2$. At one of the laws, the estimator therefore has probability at least $1/4$ of an estimation error of at least $D_n/2$, and its RMSE is at least $D_n/4$. \end{proof} \section{Phase priors and the Poisson comparison} \label{sec:lower-preparation} Throughout the lower-bound constructions, we fix $d\ge1$, $\alpha,\beta>0$ and $00,\quad \sum_zq_{z0}=1,\quad \sum_zq_{zu}=\sum_zq_{zv}=0. \end{equation} Here $u=u(x)$ and $v=v(x)$ are the perturbations. We call $\mathsf s_u(z)=q_{zu}/q_{z0}$ and $\mathsf s_v(z)=q_{zv}/q_{z0}$ the score functions of $u$ and $v$; they may coincide. Let $F$ be $C^4$ near $(0,0)$ with $F_{uv}(0,0)\ne0$, and assume that the target is a parameter that equals $T(P)=\int pF(u,v)$ on the laws that we construct. We state the lower bounds in terms of $a_n=a_n(\theta,\tau)$ from \eqref{eq:subcritical-scale}, with these $\theta$ and $\tau$. We next describe the priors. We fix a cube $\mathcal Q$ in the interior of $[0,1]^d$, of volume $v_0$, and we partition it into $B$ congruent subcubes, the \emph{blocks}, where $B$ is the $d$th power of an even integer. We group the blocks into $B/2$ \emph{pairs} of adjacent blocks, and we label the two blocks of each pair by $\ell\in\{1,2\}$. In each block we use rescaled coordinates in $\mathcal U=[0,1]^d$, and we identify a pair with $\mathcal U\times\{1,2\}$, which carries the product of Lebesgue measure and counting measure. Let $M\ge2$ be an integer, $C_0>0$, and $A_u,A_v>0$. We call two laws, the $+$ and $-$ \emph{priors}, of random functions $(p,u,v)$ on $[0,1]^d$ \emph{admissible with frequency $M$} if the following hold under both of them. \begin{enumerate}[label=(A\arabic*)] \item The parameters of distinct pairs are independent. \item The function $p$ is a smooth probability density with values in $[r_-,r_+]$, it equals a fixed constant $p_0$ outside $\mathcal Q$, and its mass on a pair is the same for every pair and in every realization. \item We have $u=v=0$ outside $\mathcal Q$, $|u|\le C_0A_u$, and $|v|\le C_0A_v$. \item The parameters of a pair are $(\Theta,H,\sigma_u,\sigma_v)$, where $(\Theta,H)$ has the same law under both priors, $\Theta$ is uniform on $[0,2\pi)$ and independent of $H$, and $\sigma_u,\sigma_v\in\{-1,1\}$. There are functions $p_c,p_1,p_2,h_u,h_v$ on the pair, which depend on $H$ but neither on $\Theta$ nor on the prior, such that, in block coordinates, \[ p=p_c+p_1\cos\Theta+p_2\sin\Theta,\qquad pu=A_u\sigma_uh_u,\qquad pv=A_v\sigma_vh_v. \] \item Conditional on $(\Theta,H)$, the signs satisfy \[ P_\pm\{\sigma_u=s,\sigma_v=t\mid\Theta,H\} =\tfrac14\{1\pm st\cos(M\Theta)\}\qquad(s,t\in\{-1,1\}). \] \end{enumerate} We write $E_\pm$ for expectation under the $\pm$ prior, and $E$ for expectation over the parameters $(\Theta,H)$ of one pair. Given a realization, the law of one observation is that of $O=(X,Z)$, where $X$ has density $p$ and $P(Z=z\mid X=x)=Q_z(u(x),v(x))$. In~\cite[Theorem~2.1 and Sections~3--4]{robins2009}, Robins et al.\ used priors with independent parameters on disjoint cells whose masses do not depend on the parameters, with a finite response space, to prove lower bounds for the missing-data mean and the expected conditional covariance. There the perturbations are bumps with independent random signs, and in the missing-data case the covariate density changes with the parameters without changing the mass of any cell. By (A4), the unsigned products $pu$ and $pv$ do not depend on $\Theta$; by (A5), only the sign correlation differs between the priors. Summing over the four sign values, for integers $k_u,k_v\ge0$ and every bounded function $f$ of $(\Theta,H)$, we obtain \begin{equation}\label{lo:sign-difference} (E_+-E_-)\bigl[\sigma_u^{k_u}\sigma_v^{k_v}f(\Theta,H)\bigr] =2\,\mathbf 1_{\{k_u,k_v\text{ odd}\}}\,E\bigl[\cos(M\Theta)f(\Theta,H)\bigr]. \end{equation} In particular, each prior is invariant under $(\sigma_u,\sigma_v)\mapsto(-\sigma_u,-\sigma_v)$, which maps $(p,u,v)$ to $(p,-u,-v)$, and the $-$ prior is the image of the $+$ prior under $\sigma_v\mapsto-\sigma_v$, which maps $(p,u,v)$ to $(p,u,-v)$. As in~\cite{jiaovenkathanweissman,wuyangentropy,aronowlopatto}, we Poissonize the sample: we compare the priors through the marked Poisson process on $[0,1]^d\times\mathcal E$ whose intensity, with respect to $\dd x$ times counting measure, is $2np(x)Q_z(u(x),v(x))$. Given the parameters, its restrictions to distinct pairs are independent. By (A2) and since $\int p=1$, the mean value of $p$ over a pair in block coordinates is $\bar p=\{1-p_0(1-v_0)\}/v_0$ in every realization. Let $\Lambda=2nv_0/B$. As reference law on a pair we use the marked Poisson process with intensity measure $\mu(\dd x,z)=\Lambda\bar p\,q_{z0}\dd x$ on $\mathcal Y=(\mathcal U\times\{1,2\})\times\mathcal E$. In block coordinates, the process on the pair has intensity $\Lambda p(x)Q_z(u(x),v(x))$, whose density with respect to $\mu$ is $1+\varphi$, where \begin{equation}\label{eq:poisson-point-factor} \varphi(x,z)=e_0(x)+e_u(x)\mathsf s_u(z)+e_v(x)\mathsf s_v(z),\qquad e_0=\frac p{\bar p}-1,\quad e_u=\frac{pu}{\bar p},\quad e_v=\frac{pv}{\bar p}. \end{equation} Since $\sum_zq_{zu}=\sum_zq_{zv}=0$ and the pair has mass $2\bar p$ in block coordinates, we have $\int\varphi\dd\mu=0$ in every realization. The next lemma is in the spirit of the affinity bound of Robins et al.~\cite[Theorem~2.1]{robins2009} for two mixtures of product laws under a product prior over the cells of a partition. We use the Poisson process in place of the sample, and we expand in the factors that the points carry. \begin{lemma}[Comparison of the Poisson mixtures] \label{lem:poisson-comparison} Let two priors be admissible with frequency $M$. Assume $B\ge n$, $A_u,A_v\le1$, and $\sup_{x,z}(|u(x)q_{zu}|+|v(x)q_{zv}|)/q_{z0}\le1/2$. Let $\mathbb P^{\rm Poi}_{B,\pm}$ be the mixtures, under the $\pm$ prior, of the laws of the marked Poisson process restricted to $\mathcal Q$. For $j\ge0$, define on $(\mathcal U\times\{1,2\})^{j+2}$, in the block coordinates of a given pair, \[ \Gamma_j(x,y,\zeta_1,\ldots,\zeta_j)=(E_+-E_-) \Bigl[(pu)(x)(pv)(y)\prod_{i=1}^j\{p(\zeta_i)-\bar p\}\Bigr], \] and let $\bar\Gamma_j^2$ be the maximum over the pairs of $\|\Gamma_j\|_2^2=\|\Gamma_j\|_{L^2((\mathcal U\times\{1,2\})^{j+2})}^2$. Then $\Gamma_j=0$ for $j0$, and fix $\epsilon_u,\epsilon_v\in(0,1]$. Suppose that there are constants $C_0,C_\Gamma\ge1$, $C_E\ge0$, and $c_2>0$ such that the following holds for all large $n$ and every $B\ge n$ that is the $d$th power of an even integer. There are two priors that are admissible with frequency $M$, with the constant $C_0$ and the amplitudes \[ A_u=\epsilon_u(BR)^{-a_*}m,\qquad A_v=\epsilon_v(BR)^{-b_*}m, \] such that, for every pair and every $j\ge0$, \begin{equation}\label{eq:coefficient-hypothesis} \|\Gamma_j\|_2^2\le C_\Gamma^{j+1}A_u^2A_v^2e^{C_EM} R^{1-M+(j-M)\sqrt M}\mathbf 1_{j\ge M}, \end{equation} and \begin{equation}\label{eq:separation-hypothesis} \Bigl|(E_+-E_-)\int puv\Bigr|\ge c_2A_uA_v\rho_n^M. \end{equation} Then, for all large $n$, there is such a $B$ with $B\ge n$, for which $A_u\le\epsilon_uB^{-a_*}$ and $A_v\le\epsilon_vB^{-b_*}$, with the following property. Let $\mathcal P$ be any class that contains the law of one observation for every realization of either prior with this $B$, and let $T$ be a parameter on $\mathcal P$ with $T(P)=\int pF(u,v)$ on these laws. Then there is $c_1>0$ such that, for all large $n$, \begin{equation}\label{eq:lower-testing-transfer} \inf_{\widehat T}\sup_{P\in\mathcal P} P\{|\widehat T-T(P)|\ge c_1m^2a_n\}\ge\frac38 . \end{equation} The infimum includes randomized estimators. Writing $B_n$ for this choice of $B$, we also have $B_n/\{n(\log n)^q\}\to\infty$ for every fixed $q>0$. The constant $c_1$ depends only on the fixed quantities, on $\epsilon_u\epsilon_v$, $F$, the response law, and the constants of the hypotheses. \end{proposition} \begin{proof} \emph{Scales.} Let $C\ge1$ be the constant of \cref{lem:poisson-comparison} for the constant $C_0$, let $C_1=2v_0CC_\Gamma$, and fix $c_0>1+\log C_1+C_E$. We define \[ t_n=\frac{\Delta_\theta L}M-\log M+c_0, \] we let $B$ be the $d$th power of the smallest even integer $k$ with $k^dR\ge ne^{t_n}$, and we set $x=BR/n$. Since $M$ is within one of $M_*=\sqrt{\theta\Delta_\theta L/\tau}$ and $\Delta_\theta L=(\tau/\theta)M_*^2$, we have $\Delta_\theta L/M=(\tau/\theta)M+O(1)$, so that $t_n=(\tau/\theta)M+O(\log M)$. As $\log R=O(\log M)$, the number $k_*=(ne^{t_n}/R)^{1/d}$ tends to infinity, and $k_*\le k0$. Since $m\le R^{a_*}$, we have $A_u\le\epsilon_uB^{-a_*}\le n^{-a_*}$, and similarly $A_v\le\epsilon_vB^{-b_*}\le n^{-b_*}$. Since $|u|\le C_0A_u$ and $|v|\le C_0A_v$, the amplitude assumptions of \cref{lem:poisson-comparison} hold for all large $n$. \emph{Indistinguishability.} Let $y=CC_\Gamma\Lambda=C_1n/B$. Since $\log y=-(\tau/\theta)M+O(\log M)$ by \eqref{eq:scales-basic}, and $\sqrt M\log R=O(\sqrt M\log M)=o(M)$, we have $yR^{\sqrt M}\le1/2$ for all large $n$. By \eqref{eq:coefficient-hypothesis}, the first term of \eqref{eq:poisson-comparison-bound} is at most \[ CC_\Gamma B\Lambda^2A_u^2A_v^2e^{C_EM}R^{1-M} \sum_{j\ge M}\frac{y^jR^{(j-M)\sqrt M}}{j!} \le2CC_\Gamma\mathcal I_n,\qquad \mathcal I_n=B\Lambda^2A_u^2A_v^2R^{1-M}e^{C_EM}\frac{y^M}{M!}, \] where we wrote $j=M+i$ and used $M!/(M+i)!\le1$ and $yR^{\sqrt M}\le1/2$. Since $B\Lambda^2=4v_0^2nR/x$, $(BR)^{-2\theta}=(nx)^{-2\theta}$, and $R^{1-M}y^M=R(C_1/x)^M$, we have \[ \mathcal I_n=4v_0^2(\epsilon_u\epsilon_v)^2n^{\Delta_\theta}R^2m^4 e^{C_EM}\frac{C_1^M}{M!\,x^{M+1+2\theta}}. \] Using $x\ge e^{t_n}$, $Mt_n=\Delta_\theta L-M\log M+c_0M$, and $\log M!\ge M\log M-M$, we obtain \[ \log\mathcal I_n\le\log(4v_0^2)+2\log R+4\log m-(1+2\theta)t_n -(c_0-1-\log C_1-C_E)M. \] The terms $2\log R$ and $4\log m$ are $O(\log M)$, since $m\le c_*M$, and $t_n>0$ for all large $n$. By the choice of $c_0$, we conclude that $\mathcal I_n\to0$. Let $s_*=\min(a_*,b_*)$. The second term of \eqref{eq:poisson-comparison-bound}, divided by $\mathcal I_n$, equals \[ C\Lambda^2(A_u^2+A_v^2)^2R^{M-1}e^{-C_EM}C_\Gamma^{-M} \le Cn^{-4s_*}e^{O(M\log M)}, \] since $\Lambda\le2v_0$, $A_u^2+A_v^2\le2n^{-2s_*}$, and $\log R=O(\log M)$. This tends to zero, since $M\log M=o(L)$. It follows that $d_{\rm H}(\mathbb P^{\rm Poi}_{B,+},\mathbb P^{\rm Poi}_{B,-})\to0$. \emph{Separation.} Let $D_n=|E_+T-E_-T|$, and let $F^{\rm oo}(u,v)=\tfrac14\{F(u,v)-F(-u,v)-F(u,-v)+F(-u,-v)\}$ be the part of $F$ that is odd in each argument. By the two symmetries that follow from (A5), we have \[ E_+T-E_-T=E_+\int p\{F(u,v)-F(u,-v)\}=2E_+\int pF^{\rm oo}(u,v), \] and in the same way $(E_+-E_-)\int puv=2E_+\int puv$. Moreover, $4F^{\rm oo}(u,v)=\int_{-u}^u\int_{-v}^vF_{uv}(s,t)\dd t\dd s$ for small $u,v$. The linear Taylor terms of $F_{uv}$ at $(0,0)$ integrate to zero over this symmetric rectangle, and since $F\in C^4$ the remainder is at most $C(s^2+t^2)$. This gives \[ F^{\rm oo}(u,v)=F_{uv}(0,0)uv+O\{|uv|(u^2+v^2)\}. \] Since $|u|\le C_0A_u$, $|v|\le C_0A_v$, and $\int p=1$, it follows from \eqref{eq:separation-hypothesis} that \[ D_n\ge|F_{uv}(0,0)|c_2A_uA_v\rho_n^M-CA_uA_v(A_u^2+A_v^2). \] Since $A_u^2+A_v^2\le2n^{-2s_*}$, while $\rho_n^{-M}=e^{O(\sqrt L)}$, the second term is at most half the first for all large $n$. Since $\log\rho_n-\log\rho=O(\delta_n)$ and $M\delta_n\to0$, we also have $\rho_n^M\ge\rho^M/2$ for all large $n$. With $f(z)=\theta\Delta_\theta L/z+\tau z$, we have \[ f(M)-f(M_*)=\frac{\tau(M-M_*)^2}M=O(M^{-1}),\qquad f(M_*)=2\sqrt{\theta\Delta_\theta\tau L}=\kappa\sqrt L, \] with $\kappa$ as in \eqref{eq:subcritical-scale}. By \eqref{eq:scales-basic}, we have $\log(BR)=L+t_n+o(1)$, and $\theta t_n+\tau M=f(M)-\theta\log M+\theta c_0$. Then \[ \log(A_uA_v\rho^M)=\log(\epsilon_u\epsilon_vm^2)-\theta L-f(M) +\theta\log M-\theta c_0+o(1)=\log(m^2a_n)+O(1), \] where we used $\theta\log M=(\theta/2)\log L+O(1)$. Thus $D_n\ge c_3m^2a_n$ for a constant $c_3>0$ and all large $n$. \emph{Testing.} The target $T$ is a constant plus a sum of independent contributions of the pairs, each bounded by $C/B$, since $F$ is bounded near the origin. Then $\Var_\pm T\le C/B$, and since $B\ge n$, $m\ge1$, and $\theta<1/2$, we have $BD_n^2\ge c_3^2na_n^2=c_3^2n^{1-2\theta}L^\theta e^{-2\kappa\sqrt L}\to\infty$. Outside $\mathcal Q$ the Poisson process has the same law in every realization, and it is independent of the process on $\mathcal Q$. Given the parameters, the number $N$ of points has law $\operatorname{Poisson}(2n)$, and given $N$ the points are independent observations from the law $P$. We order the points randomly, keep the first $n$ when $N\ge n$, and return a fixed sample otherwise. This defines a kernel that does not depend on the parameters. Its output mixtures are $(1-\pi_n)\mathbb Q_\pm^{(n)}+\pi_n\delta_*$, where $\mathbb Q_\pm^{(n)}$ are the mixtures of the laws of $n$ observations, $\delta_*$ is the law of the fixed sample, and $\pi_n=P\{\operatorname{Poisson}(2n)0$ and $\lambda\ge1$, let $M\ge2$ be an integer, and let $h_M=2\pi/M$, $\pi_\eta(x)=\eta^{-1}\Pi(x/\eta)$, and $K=6Q/\eta$. We define \[ G(x)=\Bigl(\frac{\sin(x/(\lambda\eta))}{x/(\lambda\eta)}\Bigr)^{2Q} \qquad(x\in\mathbb R),\qquad G(0)=1. \] If $4K\le M$, then the numbers $h_M\pi_\eta(x)$, $x\in h_M\mathbb Z$, sum to one. If $\xi$ has the law $P\{\xi=x\}=h_M\pi_\eta(x)$ on $h_M\mathbb Z$, then \begin{equation}\label{lat:gate-bounds} 0\le G\le1,\qquad 1-EG(\xi)\le\frac{Q\mu_2}{3\lambda^2},\qquad G(x)|x|^r\le(\lambda\eta)^r \quad(x\in\mathbb R,\ 0\le r\le2Q-2). \end{equation} Moreover, the function $\widehat\Xi(w)=E[G(\xi)^2e^{iw\xi}]$ of $w\in\mathbb R$ satisfies $|\widehat\Xi|\le1$, and $\widehat\Xi(w)=0$ unless $|w-lM|\le K$ for some $l\in\mathbb Z$. \end{lemma} \begin{proof} The identity $\sin y/y=\frac12\int_{-1}^1e^{i\omega y}\dd\omega$ shows that $(\sin y/y)^{2Q}$ is the Fourier transform of the $2Q$-fold convolution of $\frac12\mathbf 1_{[-1,1]}$, which is supported in $[-2Q,2Q]$. By Fourier inversion, $\widehat\pi_\eta$ and $\widehat G$ are supported in $[-2Q/\eta,2Q/\eta]$ and $[-2Q/(\lambda\eta),2Q/(\lambda\eta)]$, respectively. Since multiplication convolves Fourier transforms and adds their supports, and since $\lambda\ge1$, $\widehat{G^k\pi_\eta}$ is supported in $[-K,K]$ for $k=0,1,2$. Let $f=G^k\pi_\eta$ with $k\in\{0,1,2\}$, and let $w\in\mathbb R$. Since $0\le f(x)\le C_\eta(1+|x|)^{-2Q}$, the $h_M$-periodic function $y\mapsto\sum_{x\in h_M\mathbb Z}f(y+x)e^{iw(y+x)}$ is continuous, and its Fourier coefficients are $h_M^{-1}\widehat f(w-lM)$, $l\in\mathbb Z$. Since only finitely many of them are nonzero, the function is a trigonometric polynomial and equals its Fourier series everywhere. At $y=0$ we obtain an instance of the Poisson summation formula~\cite[Chapter~5, Section~3]{steinshakarchi}, \begin{equation}\label{lat:sampling} h_M\sum_{x\in h_M\mathbb Z}f(x)e^{iwx} =\sum_{l\in\mathbb Z}\widehat f(w+lM). \end{equation} At $w=0$ only the term $l=0$ remains, since $K0$. We set \[ r_0=\frac{r_-+r_+}2,\qquad h_n=\frac{r_+-r_-}2-\delta_n,\qquad p_0=\frac{1-v_0\bar\iota r_0}{1-v_0\bar\iota}, \] and we choose $v_0$ so small that $p_0$ lies in a fixed interior subinterval of $(r_-,r_+)$. Throughout this section, $n$ is large enough that \[ 0<\delta_n\le\tfrac12\min\{r_0-r_-,\ p_0-r_-,\ r_+-p_0\}. \] We write $s_0=\min(\alpha,\beta)$ and $\alpha_0=\min(s_0,1)/2$; we use only $0<\alpha_0<1$ and $\alpha_00$; in both cases $\varpi_\gamma(y)\ge\Upsilon(0)=1/2$. In the same way $\varpi_\gamma(y)\le1/2$ if $\lfloor y\rfloor$ is even. Therefore \begin{equation}\label{lat:soft-digit} |\varpi_\gamma(y)-e|\ge\tfrac12\qquad \text{if }e\in\{0,1\}\text{ and }e\ne\lfloor y\rfloor\bmod2 . \end{equation} \emph{Scales.} For a level $1\le a\le J$, we write $b=J-a$ and define \[ \gamma_a=\frac{\gamma_*}{(b+1)^2},\qquad \lambda_a=\lambda_*(b+1),\qquad \eta_a=\frac{\gamma_a2^{b\alpha_0}}{\lambda_am},\qquad K_a=\frac{6Q}{\eta_a}, \] where $\gamma_*=\min\{1/4,I_{\rm in}/(16d)\}$ and $\lambda_*=\max\{1,\sqrt{8Qd\mu_2/3}\}$. Thus $\lambda_a\eta_a=\gamma_a2^{b\alpha_0}/m$, so that coarse levels carry larger coefficients, and $K_a=K(b)m$ with $K(b)=6Q\lambda_*(b+1)^32^{-b\alpha_0}/\gamma_*$. We set $\overline K=\sup_{b\ge0}K(b)$ and $c_*=1/(4\overline K)$, and we require that $m\le c_*M$, so that \begin{equation}\label{lat:band} 4K_a\le M\qquad(1\le a\le J). \end{equation} All these constants depend only on $d,\alpha,\beta$ and $\iota_{\rm in}$, and they are fixed before $M$ is chosen. \emph{The phase and the gate.} Let $M\ge2$ be even with $m\le c_*M$, and let $h_M=2\pi/M$. In each pair, let $\xi_{a,q}$, for $1\le a\le J$ and $1\le q\le d$, be independent, with the law of $\xi$ in \cref{lat:gate} for $\eta=\eta_a$ and $\lambda=\lambda_a$; by \eqref{lat:band}, the lemma applies. We write $G_a$ and $\widehat\Xi_a$ for the corresponding functions $G$ and $\widehat\Xi$, and we define \[ S=\prod_{a,q}G_a(\xi_{a,q}),\qquad s_{a,q}(x)=\varpi_{\gamma_a}(2^ax_q),\qquad \phi(x)=\sum_{a=1}^J\sum_{q=1}^d\xi_{a,q}s_{a,q}(x) \] for $x\in\mathbb R^d$. By independence, \eqref{lat:gate-bounds}, and the inequality $\prod_i(1-t_i)\ge1-\sum_it_i$ for $t_i\in[0,1]$, we have \[ ES\ge1-\frac{Qd\mu_2}{3\lambda_*^2}\sum_{b\ge0}(b+1)^{-2}\ge\frac34, \qquad ES^2\ge(ES)^2\ge\frac9{16}, \] where we used $\sum_{b\ge0}(b+1)^{-2}<2$ and the choice of $\lambda_*$. Since $S\le G_a(\xi_{a,q})$ for every $(a,q)$ and $0\le G_a\le1$, the last bound in \eqref{lat:gate-bounds} implies that, for distinct pairs $(a_i,q_i)$ and integers $r_i\ge1$ with $\sum_ir_i\le2Q-2$, \begin{equation}\label{lat:gate-product} S\prod_i|\xi_{a_i,q_i}|^{r_i} \le\prod_i(\lambda_{a_i}\eta_{a_i})^{r_i}. \end{equation} Let $\mathcal U^\circ$ be the set of $x\in\mathcal U$ such that $2^ax_q$ has distance at least $\gamma_a$ from the integers for all $a$ and $q$. On $\mathcal U^\circ$ we have $s_{a,q}(x)=e_a(x_q)\in\{0,1\}$, so that $\phi(x)$ is a sum of elements of $h_M\mathbb Z$, and therefore \begin{equation}\label{lat:coherence} \cos(M\phi(x))=1\qquad(x\in\mathcal U^\circ). \end{equation} For each $a$ and $q$, the set of $x\in\mathcal U$ for which $2^ax_q$ has distance less than $\gamma_a$ from the integers has measure $2\gamma_a$. Since $\sum_a\gamma_a<2\gamma_*$, a union bound gives \begin{equation}\label{lat:collar} |\mathcal U\setminus\mathcal U^\circ|<4d\gamma_*\le I_{\rm in}/4 . \end{equation} \emph{The priors.} In each pair, let $\Theta$ be uniform on $[0,2\pi)$ and independent of $\xi=(\xi_{a,q})_{a,q}$. We call $(\Theta,\xi)$ the density parameters of the pair; both blocks of the pair use them, and their law is the same under both priors. On the block of the pair with label $\ell$, in block coordinates, we define \[ r(x)=r_0+(-1)^{\ell-1}h_n\cos(\Theta+\phi(x)),\qquad p=p_0+\iota_{\rm out}(r-p_0), \] and $p=p_0$ outside $\mathcal Q$. Every realization of $\phi$ is $C^\infty$, since it is a finite sum with finite coefficients. Since both blocks of a pair use the same function $\phi$ in block coordinates, the oscillations of the two blocks cancel in $\int p$, and the mass of $p$ on each pair is $2(v_0/B)\bar p$ in every realization, where \[ \bar p=p_0+\bar\iota(r_0-p_0) \] is the mean value of $p$ over a pair. The total mass is therefore $p_0(1-v_0)+v_0\bar p=p_0+v_0\bar\iota(r_0-p_0)=1$. Since $\iota_{\rm out}$ vanishes near the block boundaries, $p$ is a $C^\infty$ probability density. On the support of $\iota_{\rm out}$ it is a convex combination of $p_0$ and $r\in[r_-+\delta_n,r_+-\delta_n]$, so that $r_-+\delta_n\le p\le r_+-\delta_n$. Conditional on the density parameters, the signs $\sigma_u,\sigma_v\in\{-1,1\}$ of a pair have the law \[ P_\pm\{\sigma_u=s,\sigma_v=t\mid\Theta,\xi\} =\tfrac14\{1\pm st\cos(M\Theta)\}\qquad(s,t\in\{-1,1\}) \] under the $\pm$ prior, as in (A5) with $H=\xi$. As in \cref{sec:lower-preparation}, $E$ averages the density parameters of one pair. Let $\epsilon_u,\epsilon_v\in(0,1]$. With the amplitudes \begin{equation}\label{eq:lattice-amplitudes} A_u=\epsilon_u(BR)^{-a_*}m,\qquad A_v=\epsilon_v(BR)^{-b_*}m, \end{equation} we define \begin{equation}\label{lat:profiles} u=A_u\sigma_uS\iota_{\rm in}/p,\qquad v=A_v\sigma_vS\iota_{\rm in}/p \end{equation} on both blocks of each pair, and $u=v=0$ outside $\mathcal Q$. The parameters of distinct pairs are independent. Without the gate $S$ and the division by $p$, these are smooth bumps on the blocks with random signs. Classical lower-bound constructions use such localized bumps with binary coefficients~\cite[Section~2, p.~1045]{stone} or random signs~\cite[Section~3.5.2]{ingster},~\cite[Sections~3--4]{robins2009}; here the signs are correlated with the phase of the density. \begin{lemma}[Lattice priors]\label{lem:lattice-priors} \label{lat:smoothness}\label{lat:coefficients}\label{lat:separation} Let $J\ge0$, let $M\ge2$ be even with $m\le c_*M$, and let $B$ be the $d$th power of an even integer. There are constants $C,C_\Gamma\ge1$ and $C_E\ge0$, which depend only on $d,\alpha,\beta,r_\pm$ and the fixed choices of the construction, and not on $n,B,J,M,\epsilon_u,\epsilon_v$, such that the two priors above have the following properties. \begin{enumerate}[label=(\alph*)] \item They are admissible with frequency $M$, with $C_0=1/r_-$ and the amplitudes \eqref{eq:lattice-amplitudes}. Every realization of $p$ is a $C^\infty$ probability density with \begin{equation}\label{eq:lower-density-margin} r_-+\delta_n\le p\le r_+-\delta_n, \end{equation} and the functions $u$ and $v$ vanish outside $\mathcal Q$. \item In every realization, $u$ and $v$ are $C^\infty$, and \begin{equation}\label{eq:lower-norms} \|u\|_{C^\alpha}\le C\epsilon_u,\qquad \|v\|_{C^\beta}\le C\epsilon_v,\qquad \|u\|_\infty\le A_u/r_-,\qquad \|v\|_\infty\le A_v/r_-. \end{equation} Let $K$ be smooth on a neighborhood $N$ of $(0,0)$. If $A_u+A_v$ is below a threshold that depends on $K$, then \begin{equation}\label{eq:lower-smooth-map} \begin{aligned} K(0,v)=0\text{ on }N&\ \Longrightarrow\ \|K(u,v)\|_{C^\alpha}\le C_K\epsilon_u,\\ K(u,0)=0\text{ on }N&\ \Longrightarrow\ \|K(u,v)\|_{C^\beta}\le C_K\epsilon_v, \end{aligned} \end{equation} where $C_K$ depends only on $K$ and on the quantities on which $C$ depends. \item For every pair and every $j\ge0$, the coefficients $\Gamma_j$ of \cref{lem:poisson-comparison} satisfy \eqref{eq:coefficient-hypothesis}, that is, \[ \|\Gamma_j\|_2^2\le C_\Gamma^{j+1}A_u^2A_v^2 e^{C_EM}R^{1-M+(j-M)\sqrt M}\mathbf 1_{j\ge M}, \] where the norm is that of $L^2((\mathcal U\times\{1,2\})^{j+2})$. \item With $\rho_n$ as in \cref{prop:reduction}, we have \[ (E_+-E_-)\int_{[0,1]^d}puv\ge c_2A_uA_v\rho_n^M, \qquad c_2=\frac{9v_0I_{\rm in}}{16r_0}. \] \end{enumerate} \end{lemma} \emph{Block-constant phases.} For $J=0$, we have $\phi=0$, $S=1$, $R=m=1$, and $\mathcal U^\circ=\mathcal U$, with $\Theta$ the only density parameter. On the bump supports, $p=r_0\pm h_n\cos\Theta$, $pu=A_u\sigma_u\iota_{\rm in}$, and $pv=A_v\sigma_v\iota_{\rm in}$. After removing their signs, these products do not depend on $\Theta$. The $M$th Fourier coefficient of $1/p$ is of order $\rho_n^M$ by \eqref{eq:cosine-reciprocal}. The scale conditions of \cref{prop:reduction} reduce to $1\le c_*M$, which holds for all large $n$. That proposition, applied with \cref{lem:lattice-priors}(a), (c), and~(d), yields, for every class $\mathcal P$ and target $T$ as in the proposition, a constant $c_1>0$ such that, for all large $n$, \begin{equation}\label{eq:lattice-testing} \inf_{\widehat T}\sup_{P\in\mathcal P} P\{|\widehat T-T(P)|\ge c_1m^2a_n\}\ge\frac38 \end{equation} with $m=1$. The proof of \cref{prop:local-lower} uses the largest $J$ with $m\le c_*M$, for which $m^2\asymp\log n$. \section{Proof of \texorpdfstring{\cref{lem:lattice-priors}}{the lemma on lattice priors}} \label{lat:sec-properties}\label{lat:sec-proof} We use the H\"older norm $\|\cdot\|_{C^t}$ of the model class on $[0,1]^d$, and we write $\|\cdot\|_{C^t(\mathcal U)}$ for the norm defined in the same way on $\mathcal U$, in block coordinates. \begin{proof}[Proof of \cref{lem:lattice-priors}(a)] We showed above that $p$ is a $C^\infty$ probability density that satisfies \eqref{eq:lower-density-margin}, equals $p_0$ outside $\mathcal Q$, and has the same mass on every pair in every realization, which is (A2). By construction, the parameters of distinct pairs are independent, $u$ and $v$ vanish outside $\mathcal Q$, and the signs have the law in (A5). Since $p\ge r_-$ and $0\le S,\iota_{\rm in}\le1$, we have $|u|\le A_u/r_-$ and $|v|\le A_v/r_-$, which gives (A3) with $C_0=1/r_-$. In (A4), we have $H=\xi$, $pu=A_u\sigma_uS\iota_{\rm in}$, $pv=A_v\sigma_vS\iota_{\rm in}$, and, on the block with label $\ell$, \[ p_{\rm c}=p_0+\iota_{\rm out}(r_0-p_0),\qquad p_1=(-1)^{\ell-1}h_n\iota_{\rm out}\cos\phi,\qquad p_2=-(-1)^{\ell-1}h_n\iota_{\rm out}\sin\phi, \] and $h_u=h_v=S\iota_{\rm in}$. These functions depend on $\xi$ but neither on $\Theta$ nor on the prior. \end{proof} \begin{proof}[Proof of \cref{lem:lattice-priors}(b)] We work in block coordinates on a block with label $\ell$, and we set $k_{\max}=\lceil\max(\alpha,\beta)\rceil=Q-2$ and \[ f(y)=\frac1{r_0+(-1)^{\ell-1}h_n\cos(\Theta+y)},\qquad \mathsf H=S\iota_{\rm in}f(\phi). \] Since $p=r$ on the support of $\iota_{\rm in}$, we have $\mathsf H=S\iota_{\rm in}/p$ and $|\mathsf H|\le1/r_-$. The derivatives of $f$ through order $k_{\max}+1$ are bounded by constants that depend only on $r_\pm$ and $k_{\max}$, since its denominator lies in $[r_-,r_+]$. Below, constants may depend on $d,\alpha,\beta,r_\pm,\iota_{\rm in},\Upsilon$ and the fixed constants of the construction, but not on $n,B,J,M,\epsilon_u,\epsilon_v$, or the realization. \emph{Compositions.} Let $\kappa$ be a $C^\infty$ function on $[-1/r_-,1/r_-]$ with $\kappa(0)=0$, and let $s_0\le t\le\max(\alpha,\beta)$. We claim that \begin{equation}\label{lat:composition} \|\kappa(\mathsf H)\|_{C^t(\mathcal U)}\le C_\kappa(1+2^{Jt}/m), \end{equation} where $C_\kappa$ depends on $\kappa$ only through bounds on its derivatives of order at most $\lceil t\rceil+1$. We define \[ w_a=\lambda_a\eta_a=\frac{\gamma_a2^{(J-a)\alpha_0}}m,\qquad L_a=\frac{2^a}{\gamma_a},\qquad \phi_a=\sum_{c=1}^a\sum_{q=1}^d\xi_{c,q}s_{c,q}\quad(0\le a\le J). \] Since $\gamma_a\le1/4$ and $2^{(J-a)\alpha_0}\le2^{Js_0}=m$, we have $w_a\le1$ and $L_a\ge1$. Since $s_{a,q}$ depends only on $x_q$, we also have $|D^{\mathbf j}s_{a,q}|\le CL_a^{|\mathbf j|}$ for every multi-index $\mathbf j$, including $\mathbf j=0$. For an integer $1\le k\le\lceil t\rceil$ and $b=J-a$, we have $w_{a-i}L_{a-i}^k/(w_aL_a^k)=\{(b+i+1)/(b+1)\}^{2(k-1)}2^{-i(k-\alpha_0)}$ for $0\le i\alpha_0$. Let $F_a=\kappa(S\iota_{\rm in}f(\phi_a))$, $V_a=F_a-F_{a-1}$, and $W_a=\phi_a-\phi_{a-1}$ for $1\le a\le J$. By the fundamental theorem of calculus, \[ V_a=S\iota_{\rm in}W_a\int_0^1 f'(z_{a,u})\kappa'\bigl(S\iota_{\rm in}f(z_{a,u})\bigr)\dd u, \qquad z_{a,u}=\phi_{a-1}+uW_a, \] and the arguments of $\kappa$ lie in $[0,1/r_-]$. Let $0\le|\mathbf j|\le\lceil t\rceil$. By the chain and product rules, $D^{\mathbf j}V_a$ is a sum, with a number of terms that depends only on $|\mathbf j|$ and $d$, of products of a power $S^\varrho$ with $\varrho\ge1$, of derivatives of $\iota_{\rm in}$, $f$, and $\kappa$ of order at most $|\mathbf j|+1$, of $D^{\mathbf j_0}W_a$, and of $D^{\mathbf j_l}z_{a,u}$ for $1\le l\le r$, where $|\mathbf j_l|\ge1$ for $l\ge1$ and $|\mathbf j_0|+\sum_l|\mathbf j_l|\le|\mathbf j|$. In particular $r\le|\mathbf j|$. We expand $W_a$ and $z_{a,u}$ over levels and coordinates. The coefficients then enter through products \[ \xi_{a,q_0}D^{\mathbf j_0}s_{a,q_0} \prod_{l=1}^{r}\xi_{c_l,q_l}D^{\mathbf j_l}s_{c_l,q_l} \qquad(c_l\le a), \] with at most $|\mathbf j|+1\le k_{\max}+1=Q-1\le2Q-2$ coefficients, counted with multiplicity. Since $S^\varrho\le S$, we group repeated coefficients and apply \eqref{lat:gate-product}, which bounds $S^\varrho$ times such a product by $Cw_aL_a^{|\mathbf j_0|}\prod_lw_{c_l}L_{c_l}^{|\mathbf j_l|}$. Summing over $q_0$, and over the levels $c_l\le a$ and the coordinates $q_l$ by \eqref{lat:earlier-levels}, and using $w_a\le1$ and $L_a\ge1$, we obtain \[ \|D^{\mathbf j}V_a\|_\infty\le C_\kappa w_aL_a^{|\mathbf j|} \qquad(0\le|\mathbf j|\le\lceil t\rceil). \] With $N=\lceil t\rceil-1$ and $\vartheta=t-N\in(0,1]$, we find from the mean value theorem and this bound that, for $|\mathbf j|=N$, \[ |D^{\mathbf j}V_a(x)-D^{\mathbf j}V_a(y)| \le C_\kappa w_aL_a^N\min\{1,L_a|x-y|\} \le C_\kappa w_aL_a^t|x-y|^\vartheta, \] where we used $\min(1,z)\le z^\vartheta$ for $z\ge0$. Since $L_a\ge1$, the derivative suprema in the norm are also at most $C_\kappa w_aL_a^t$, and therefore $\|V_a\|_{C^t(\mathcal U)}\le C_\kappa w_aL_a^t$. The function $F_0=\kappa(S\iota_{\rm in}f(0))$ has norm at most $C_\kappa$, since $Sf(0)\in[0,1/r_-]$ is a number. Since $\kappa(\mathsf H)=F_J=F_0+\sum_{a=1}^JV_a$, we obtain \[ \|\kappa(\mathsf H)\|_{C^t(\mathcal U)} \le C_\kappa\Bigl\{1+\frac{2^{Jt}}m\gamma_*^{1-t} \sum_{b\ge0}(b+1)^{2(t-1)}2^{-b(t-\alpha_0)}\Bigr\} \le C_\kappa(1+2^{Jt}/m), \] where the series is bounded uniformly in $s_0\le t\le\max(\alpha,\beta)$ because $s_0>\alpha_0$. For $J=0$ the sum over $a$ is empty. This proves \eqref{lat:composition}. \emph{Conclusion.} On a block, we have $u=A_u\sigma_u\mathsf H$ and $v=A_v\sigma_v\mathsf H$ in block coordinates, which gives the sup-norm bounds. Let $g$ be a $C^\infty$ function on $[0,1]^d$ that vanishes outside $\mathcal Q$ and, on each block, equals $g_k$ in block coordinates, where every $g_k$ vanishes near $\partial\mathcal U$. The blocks have side $\ell_B=(v_0/B)^{1/d}\le1$. Rescaling multiplies a derivative of order $|\mathbf j|\le t$ by $\ell_B^{-|\mathbf j|}\le\ell_B^{-t}$ and a seminorm of order $\vartheta$ of a derivative of order $\lceil t\rceil-1$ by $\ell_B^{-t}$. If $x$ and $y$ do not lie in the same block, the segment from $x$ to $y$ meets the boundary of the block of $x$ at a point $x'$ and the boundary of the block of $y$ at a point $y'$, where all derivatives of $g$ vanish; the same holds if one of the points lies outside $\mathcal Q$. It follows that a difference quotient of order $\vartheta\le1$ between $x$ and $y$ is at most twice the largest such quotient within one block. Since the norm adds the derivative suprema and the H\"older seminorm, whose maxima may occur in different blocks, we obtain \begin{equation}\label{lat:rescale} \|g\|_{C^t([0,1]^d)}\le3(B/v_0)^{t/d}\max_k\|g_k\|_{C^t(\mathcal U)}. \end{equation} We apply \eqref{lat:composition} with $\kappa(z)=z$ and $t=\alpha$, and \eqref{lat:rescale}. Since $2^{J\alpha}=R^{a_*}$, $A_uB^{a_*}=\epsilon_uR^{-a_*}m$, and $m=2^{Js_0}\le R^{a_*}$, we obtain \[ \|u\|_{C^\alpha}\le CA_uB^{a_*}(1+R^{a_*}/m) =C\epsilon_u(mR^{-a_*}+1)\le2C\epsilon_u . \] The bound for $v$ is the same with $\beta$, $b_*$, and $\epsilon_v$. Let $K$ be smooth on a neighborhood of $(0,0)$ with $K(0,v)=0$ there, and fix a closed ball $N'$ of radius at most one around $(0,0)$ inside this neighborhood. On $N'$ we write $K(u,v)=u\widetilde K(u,v)$ with $\widetilde K(u,v)=\int_0^1\partial_uK(su,v)\dd s$. If $(A_u+A_v)/r_-$ is at most the radius of $N'$, then on each block $K(u,v)=A_u\sigma_u\kappa(\mathsf H)$ with $\kappa(z)=z\widetilde K(A_u\sigma_uz,A_v\sigma_vz)$, and $K(u,v)=0$ outside the blocks. The derivatives of $\kappa$ of order at most $k_{\max}+1$ on $[-1/r_-,1/r_-]$ are bounded by a constant that depends only on $K$ and $N'$, because $A_u\le\epsilon_uB^{-a_*}\le1$ and $A_v\le\epsilon_vB^{-b_*}\le1$. The argument above, with this $\kappa$ and $t=\alpha$, yields $\|K(u,v)\|_{C^\alpha}\le C_K\epsilon_u$. The case $K(u,0)=0$ is symmetric, with $\beta$ and $\epsilon_v$. \end{proof} \begin{proof}[Proof of \cref{lem:lattice-priors}(c)] We fix a pair. Since $pu=A_u\sigma_uS\iota_{\rm in}$ and $pv=A_v\sigma_vS\iota_{\rm in}$, \eqref{lo:sign-difference} gives \[ \Gamma_j(x,y,\zeta_1,\ldots,\zeta_j) =2A_uA_v\iota_{\rm in}(x)\iota_{\rm in}(y) E\Bigl[S^2\cos(M\Theta)\prod_{i=1}^j\{p(\zeta_i)-\bar p\}\Bigr]. \] Since $p$ and $\bar p$ lie in $[r_-,r_+]$, we have $|p-\bar p|\le r_+$, so that $\|\Gamma_j\|_\infty\le2A_uA_vr_+^j$. Since $\|\Gamma_j\|_2^2\le\|\Gamma_j\|_\infty\|\Gamma_j\|_1$, it suffices to prove that \begin{equation}\label{lat:L1-bound} \|\Gamma_j\|_1\le8A_uA_v(6\mathrm er_+)^je^{C_EM} R^{1-M+(j-M)\sqrt M}\mathbf 1_{j\ge M}, \end{equation} where $\mathrm e=\exp(1)$; then (c) holds with $C_\Gamma=16(1+6\mathrm er_+^2)$. As $(\mathcal U\times\{1,2\})^{j+2}$ has measure $2^{j+2}$, we have \begin{equation}\label{lat:crude} \|\Gamma_j\|_1\le8A_uA_v(2r_+)^j . \end{equation} If $j\ge M+\sqrt M$, then $(j-M)\sqrt M\ge M$, so that \eqref{lat:crude} implies \eqref{lat:L1-bound}. We expand each factor in Fourier modes in $\Theta$. On the block with label $\ell$, \[ \begin{gathered} p(\zeta)-\bar p=c(\zeta)+w(\zeta)\sum_{o=\pm1}e^{io(\Theta+\phi(\zeta))},\\ c=(\iota_{\rm out}-\bar\iota)(r_0-p_0),\quad w=(-1)^{\ell-1}\iota_{\rm out}h_n/2, \end{gathered} \] and $|c|\le r_+$ and $|w|\le r_+/2$. We write $\cos(M\Theta)=\frac12\sum_{\varepsilon=\pm1}e^{i\varepsilon M\Theta}$, and we fix $\varepsilon$ and a mode $o_i\in\{-1,0,1\}$ for each $i\in\{1,\ldots,j\}$. We denote the corresponding term of $\Gamma_j$ by $F_{\varepsilon,o}$; it is $A_uA_v\iota_{\rm in}(x)\iota_{\rm in}(y) \prod_{o_i=0}c(\zeta_i)\prod_{o_i\ne0}w(\zeta_i)$ times the expectation of $S^2e^{i\varepsilon M\Theta}\prod_{o_i\ne0}e^{io_i(\Theta+\phi(\zeta_i))}$. Since $S^2e^{i\sum_io_i\phi(\zeta_i)}=\prod_{a,q}G_a(\xi_{a,q})^2e^{i\xi_{a,q}\Phi_{a,q}}$, integrating over the uniform angle and the independent coefficients, we find that this expectation equals \begin{equation}\label{lat:alias-factor} \mathbf 1\Bigl\{\sum_io_i=-\varepsilon M\Bigr\} \prod_{a,q}\widehat\Xi_a(\Phi_{a,q}),\qquad \Phi_{a,q}=\sum_{i:o_i\ne0}o_is_{a,q}(\zeta_i). \end{equation} Since $|\sum_io_i|\le j$, it follows that $\Gamma_j=0$ for $j1$, we have $c_*<1$ and $m\le M$. With $2k<\sqrt M$, we obtain $\Omega k+d\sum_a\Omega_aK_a\le C_EM$, where $C_E=C_{\rm gate}+D$. For $i\in I_-\cup I_+$, by the lower bound on $\chi_{i,a,q}$, the independence of the digits under Lebesgue measure, and $de^{-\Omega_a/2}=(b+2)^{-2}$, we have \[ \int_{\mathcal U\times\{1,2\}}|w(\zeta_i)| e^{-\sum_{a,q}\Omega_a\chi_{i,a,q}}\dd\zeta_i \le r_+\prod_{a,q}\frac{1+e^{-\Omega_a/2}}2 \le\frac{r_+}R\exp\Bigl\{\sum_{b\ge0}\frac1{(b+2)^2}\Bigr\} \le\frac{\mathrm er_+}R . \] Each $i$ with $o_i=0$ contributes at most $2r_+$, and each of $x$ and $y$ at most $2$. Summing over the $R$ strings $e^*$, we obtain \[ \|F_{\varepsilon,o}\|_1\le4A_uA_vR\,e^{C_EM} \Bigl(\frac{\mathrm er_+}R\Bigr)^{M+2k}(2r_+)^{j-M-2k} \le4A_uA_v(2\mathrm er_+)^jR^{1-M}e^{C_EM}, \] which proves \eqref{lat:pattern-bound}. \end{proof} \begin{proof}[Proof of \cref{lem:lattice-priors}(d)] Let $I=[r_-+\delta_n,r_+-\delta_n]$, so that $\rho_I=\rho_n$ in the notation of \cref{lem:reciprocal}, with $c_I=r_0$, $d_I=h_n$, and $\bar g_I=\sqrt{r_0^2-h_n^2}\le r_0$. On the block with label $\ell$, we have $puv=A_uA_v\sigma_u\sigma_vS^2\iota_{\rm in}^2/r$ and $r=c_I+d_I\cos(\Theta+\phi+(\ell-1)\pi)$. Since the series \eqref{eq:cosine-reciprocal} converges absolutely, multiplying it by $\cos(M\Theta)$ and averaging over $\Theta$ with $\xi$ fixed keeps only its term of order $M$. We write $E_\Theta$ for this average. Since $M$ is even, we obtain on both blocks \[ E_\Theta\frac{\cos(M\Theta)}{r}=\frac{\rho_n^M\cos(M\phi)}{\bar g_I}. \] By \eqref{lo:sign-difference}, and since $\Theta$ is independent of $\xi$, each of the $B$ blocks has volume $v_0/B$, and the density parameters of all pairs have the same law, we obtain \[ (E_+-E_-)\int_{[0,1]^d}puv =\frac{2v_0A_uA_v\rho_n^M}{\bar g_I} E\Bigl[S^2\int_{\mathcal U}\iota_{\rm in}^2\cos(M\phi)\Bigr]. \] By \eqref{lat:coherence}, $\cos(M\phi)=1$ on $\mathcal U^\circ$. Since $0\le\iota_{\rm in}\le1$, \eqref{lat:collar} implies that the integral is at least $I_{\rm in}-2|\mathcal U\setminus\mathcal U^\circ|\ge I_{\rm in}/2$ in every realization. The expectation is therefore at least $(I_{\rm in}/2)ES^2\ge9I_{\rm in}/32$, although $S$ and $\phi$ are dependent. With $\bar g_I\le r_0$, this gives (d) with $c_2=9v_0I_{\rm in}/(16r_0)$. \end{proof} % ---------- 07-embedding.tex \chapter{Embedding and proof of the lower bounds} \label{sec:lower-proof}\label{emb:lower-proof} \section{Embedding the priors in the generic model} \label{sec:generic-embedding} We perturb $\pi_0$ along scores that shift $a,b$ by $u/w,v/w$ (both by $(u+v)/w$ in the diagonal case). We apply \cref{prop:reduction} to the resulting finite affine law. The applications use \cref{prop:local-lower} with arbitrary density intervals and scores. As in the lower bounds of~\cite[Sections~3--4]{robins2009}, where the responses are binary, we use a finite response space whose conditional laws are perturbed through the nuisance functions. We write a finite-support law for $Z$ as $\pi_0=\sum_{k=1}^{m_0}\pi_{0,k}\delta_{z_k}$, with every $\pi_{0,k}>0$, we let $w_0,a_0,b_0$ be as in \eqref{eq:generic-baseline-moments}, with $w_0>0$, and we define $\mathsf r=(U-a_0D,\,V-b_0D)^\top$, which has mean zero under $\pi_0$. We call bounded functions $\mathsf s_u,\mathsf s_v$ on the support of $\pi_0$ \emph{scores} for $\pi_0$ if $\E_{\pi_0}\mathsf s_u=\E_{\pi_0}\mathsf s_v=0$ and either \begin{equation}\label{eq:generic-scores} \E_{\pi_0}[\mathsf r(\mathsf s_u,\mathsf s_v)]=\mathrm{Id}, \quad\text{or}\quad U=V,\ \ \mathsf s_u=\mathsf s_v,\ \ \E_{\pi_0}[(U-a_0D)\mathsf s_u]=1, \end{equation} where $\mathrm{Id}$ is the $2\times2$ identity matrix. We call the second alternative the \emph{diagonal case}; in it we set $\mathsf r_{\rm d}=U-a_0D$, and we call the common score $\mathsf s_{\rm d}$ a \emph{diagonal score}. If the residual covariance $\Sigma=\E_{\pi_0}[\mathsf r\mathsf r^\top]$ is positive definite, then $(\mathsf s_u,\mathsf s_v)^\top=\Sigma^{-1}\mathsf r$ are scores. If $U=V$ and $\E_{\pi_0}\mathsf r_{\rm d}^2>0$, then $\mathsf s_{\rm d}=\mathsf r_{\rm d}/\E_{\pi_0}\mathsf r_{\rm d}^2$ is a diagonal score. Given scores, we define \begin{equation}\label{eq:generic-affine-response} \pi_{u,v}(\{z_k\}) =\pi_{0,k}\{1+u\mathsf s_u(z_k)+v\mathsf s_v(z_k)\}. \end{equation} These probabilities sum to one because the scores have mean zero, and, since the support is finite, they are at least $\pi_{0,k}/2$ when $|u|+|v|$ is below a fixed threshold. In the notation of \eqref{eq:finite-affine-law}, we have $q_{z_k0}=\pi_{0,k}$, $q_{z_ku}=\pi_{0,k}\mathsf s_u(z_k)$, and $q_{z_kv}=\pi_{0,k}\mathsf s_v(z_k)$. Let $(\ell_a,\ell_b)=(u,v)$, and $(\ell_a,\ell_b)=(u+v,u+v)$ in the diagonal case, and let $d_u=\E_{\pi_0}[D\mathsf s_u]$ and $d_v=\E_{\pi_0}[D\mathsf s_v]$. For every $f$, we have $\E_{\pi_{u,v}}f=\E_{\pi_0}f+u\E_{\pi_0}[f\mathsf s_u]+v\E_{\pi_0}[f\mathsf s_v]$. By $\E_{\pi_0}\mathsf r=0$ and \eqref{eq:generic-scores}, we obtain $\E_{\pi_{u,v}}\mathsf r=(\ell_a,\ell_b)^\top$; in the diagonal case, $a_0=b_0$ and both entries of $\mathsf r$ equal $\mathsf r_{\rm d}$. As $U=\mathsf r_1+a_0D$ and $V=\mathsf r_2+b_0D$, we find \begin{equation}\label{eq:generic-affine-moments} w(u,v)=w_0+d_uu+d_vv,\qquad \E_{\pi_{u,v}}U=a_0w(u,v)+\ell_a,\qquad \E_{\pi_{u,v}}V=b_0w(u,v)+\ell_b, \end{equation} where $w(u,v)=\E_{\pi_{u,v}}D$. The regression ratios are therefore \begin{equation}\label{eq:generic-lower-regressions} a(u,v)=a_0+\frac{\ell_a}{w(u,v)},\qquad b(u,v)=b_0+\frac{\ell_b}{w(u,v)}. \end{equation} Suppose that $X$ has density $p$ and the response law at $x$ is $\pi_{u(x),v(x)}$. With $m_U=\E_{\pi_{u,v}}U$, $m_V=\E_{\pi_{u,v}}V$, and $w=w(u,v)$, we have the identity \[ \frac{m_Um_V}w=a_0m_V+b_0m_U-a_0b_0w +\frac{(m_U-a_0w)(m_V-b_0w)}w . \] By \eqref{eq:generic-functional} and \eqref{eq:generic-affine-moments}, this identity implies that $\psi(P)=\int pF(u,v)$, where \begin{equation}\label{eq:generic-lower-integrand} F(u,v)=\E_{\pi_{u,v}}\widetilde W+\lambda\frac{\ell_a\ell_b}{w(u,v)}, \qquad \widetilde W=W+\lambda(a_0V+b_0U-a_0b_0D). \end{equation} The first term is affine in $(u,v)$, and $F$ is smooth on a fixed neighborhood of the origin on which $w(u,v)\ge w_0/2$. Since $\ell_a\ell_b$ vanishes to second order at the origin, we have $F_{uv}(0,0)=\lambda\,\partial_u\partial_v(\ell_a\ell_b)/w_0$, which is $\lambda/w_0$ or $2\lambda/w_0$ in the diagonal case, and is nonzero. \begin{proposition}[Local lower bounds]\label{prop:local-lower} Fix $d,\alpha,\beta$ with $\theta=(\alpha+\beta)/d$, a finite-support law $\pi_0$ for $Z$ with $w_0>0$, the functions $D,U,V,W$, and $\lambda\ne0$. Suppose that $\mathsf s_u,\mathsf s_v$ are scores for $\pi_0$, with $\alpha=\beta$ in the diagonal case, and let $\pi_{u,v}$ and $F$ be as above. Fix $00$, and let $\tau$ be computed from $[r_-,r_+]$ as in \cref{sec:lower-preparation}. \begin{enumerate} \item[(a)] Suppose that $\theta<1/2$. For all sufficiently large $n$ there are two priors on laws $P$ of $O=(X,Z)$, under which $X$ has a $C^\infty$ density $p$ and $Z$ given $X=x$ has law $\pi_{u(x),v(x)}$, such that every realization satisfies \eqref{eq:lower-density-margin}, $\|u\|_\infty\le Cn^{-\alpha/d}$, $\|v\|_\infty\le Cn^{-\beta/d}$ (so that $\|u\|_\infty+\|v\|_\infty=o(\delta_n)$), and, for every fixed smooth function $K$ on a neighborhood $N$ of $(0,0)$, for all sufficiently large $n$ depending on $K$, \begin{equation}\label{eq:generic-angular-ratio-split} \begin{gathered} K(0,v)=0\text{ on }N\ \Longrightarrow\ \|K(u,v)\|_{C^\alpha}\le C_K\epsilon,\\ K(u,0)=0\text{ on }N\ \Longrightarrow\ \|K(u,v)\|_{C^\beta}\le C_K\epsilon, \end{gathered} \end{equation} where $C_K$ does not depend on $\epsilon$ or $n$. We call the laws under these priors the \emph{hard laws}, and we write $\cP^{\rm h}_n$ for their set. On these laws $\psi(P)=\int pF(u,v)$. If $\mathcal P$ is a class that contains $\cP^{\rm h}_n$ and $T$ is a parameter with $T(P)=\psi(P)+t_*$ on $\cP^{\rm h}_n$ for a constant $t_*$, then there is $c>0$ such that, for all sufficiently large $n$, \[ \inf_{\widehat T}\sup_{P\in\mathcal P} P\{|\widehat T-T(P)|\ge c\,a_n\log n\}\ge\frac38, \qquad \risk(T,\mathcal P)\ge\sqrt{3/8}\,c\,a_n\log n, \] with $a_n=a_n(\theta,\tau)$. The infimum includes randomized estimators. \item[(b)] For every $u_*$ in a neighborhood of zero, the laws $P_t=\operatorname{Unif}([0,1]^d)\otimes\pi_{u_*,t}$, for $t$ near zero, have $\psi(P_t)=F(u_*,t)$. For every sufficiently small $u_*\ne0$ there is $c>0$, depending also on $u_*$, such that, on every class $\mathcal P$ that contains $P_t$ for all $t$ in an interval about zero, a parameter $T$ with $T(P_t)=\psi(P_t)+t_*$ for a constant $t_*$ satisfies $\risk(T,\mathcal P)\ge cn^{-1/2}$ for all sufficiently large $n$. \end{enumerate} The constants may depend on the fixed quantities above and on $\epsilon$. \end{proposition} \begin{proof} For (a), we apply \cref{prop:reduction} to the response law \eqref{eq:generic-affine-response}, which has the form \eqref{eq:finite-affine-law}, with the integrand $F$ and the interval $[r_-,r_+]$. Let $M$ be as in that proposition, let $c_*$ be the constant of \cref{sec:block-phases}, and let $s_0=\min(\alpha,\beta)$. For all large $n$ we have $c_*M\ge1$, and we set \[ J=\Bigl\lfloor\frac{\log_2(c_*M)}{s_0}\Bigr\rfloor\ge0,\qquad m=2^{Js_0},\qquad R=2^{Jd}, \] so that $2^{-s_0}c_*M2^{-s_0}c_*M$ and $M\ge\sqrt{\theta(1-2\theta)\log n/\tau}-1$, we have \[ m^2\ge2^{-2s_0-2}c_*^2\,\frac{\theta(1-2\theta)}{\tau}\log n \] for all large $n$, which gives the testing bound with the threshold $c\,a_n\log n$ for $c=2^{-2s_0-2}c_*^2\theta(1-2\theta)c_1/\tau$. The mean squared error is at least the squared threshold times the error probability, which gives the RMSE bound. With $J=0$ in place of this choice, so that $m=R=1$, these bounds hold with $a_n$ in place of $a_n\log n$, by the same argument. For (b), the constant perturbations $u=u_*$ and $v=t$, with $p=1$, give the law $P_t$, so that $\psi(P_t)=F(u_*,t)$. Let $\varphi(u)=F_v(u,0)$. Since $\varphi'(0)=F_{uv}(0,0)\ne0$, we have $\varphi(u_*)\ne0$ for every sufficiently small $u_*\ne0$, by continuity if $\varphi(0)\ne0$, and since $\varphi(u)=u\{\varphi'(0)+o(1)\}$ otherwise. Relative to $P_*=\operatorname{Unif}([0,1]^d)\otimes\pi_0$, the law $P_t$ has the density $1+u_*\mathsf s_u(Z)+t\mathsf s_v(Z)$, which is at least $1/2$ when $|u_*|+|t|$ is below a fixed threshold and has the bounded slope $\mathsf s_v(Z)$ in $t$. On an open interval about zero on which $P_t\in\mathcal P$, the map $t\mapsto T(P_t)=F(u_*,t)+t_*$ has the derivative $\varphi(u_*)\ne0$ at $t=0$. The bound now follows from \cref{lem:parametric-path}. \end{proof} \section{Localization}\label{sec:hard-laws} Every conditional moment of a hard law is affine in $(u,v)$: $\E_P[f(Z)\mid X]=\E_{\pi_0}f+u\E_{\pi_0}[f\mathsf s_u]+v\E_{\pi_0}[f\mathsf s_v]$. Every conditional-response nuisance used below, such as a conditional mean, a ratio of conditional means whose denominator is positive at the baseline, a propensity, or its reciprocal, is therefore of the form $K_P(x)=K(u(x),v(x))$ for a $C^\infty$ function $K$ near the origin. \begin{lemma}[Localization and membership] \label{lem:local}\label{cor:membership} Adopt the setting of \cref{prop:local-lower}. Let $K$ be a $C^\infty$ function on a neighborhood of $(0,0)$, let $k_0=K(0,0)$, and for a law $P$ with perturbations $u,v$ let $K_P(x)=K(u(x),v(x))$. Let $\gamma_K=\alpha$ if $K(0,v)=k_0$ for all small $v$, let $\gamma_K=\beta$ if $K(u,0)=k_0$ for all small $u$, and let $\gamma_K=\min(\alpha,\beta)$ otherwise. If both identities hold, either value may be used. There is $C_K$, independent of $n$ and $\epsilon$, such that, for all sufficiently large $n$, every hard law $P$ satisfies \[ \|K_P-k_0\|_\infty \le C_K(\|u\|_\infty+\|v\|_\infty)=o(\delta_n),\qquad \|K_P-k_0\|_{C^t}\le C_K\epsilon\quad(00$, also $r_-k_0\le pK_P\le r_+k_0$. For the laws $P_t$ of \cref{prop:local-lower}(b), $K_{P_t}$ is the constant $K(u_*,t)$, which tends to $k_0$ as $(u_*,t)\to0$. Consider finitely many conditions on $(p,u,v)$ of the following three types, each with its own $C^\infty$ function $K$ and $k_0=K(0,0)$: \begin{enumerate} \item[(H)] $\|K_P\|_{C^t}\le H'$, with $00$, $g_-'\le r_-k_0$, and $r_+k_0\le g_+'$. \end{enumerate} If $\epsilon$ is small enough, depending only on these conditions, then for all large $n$ every hard law satisfies all of them. Moreover, for every sufficiently small $u_*$, the laws $P_t$, which have $p=1$, satisfy all of them for $t$ in an interval about zero; this assertion does not use $\theta<1/2$. \end{lemma} \begin{proof} Shrinking the domain of $K$, we may assume that it is an open square $N$ centered at the origin, so that $(u,0)$ and $(0,v)$ lie in $N$ whenever $(u,v)$ does, and that $K$ is Lipschitz on $N$ with a constant $C_K$. By \cref{prop:local-lower}(a), we have $\|u\|_\infty+\|v\|_\infty\le Cn^{-\min(\alpha,\beta)/d}=o(\delta_n)$, uniformly over the hard laws, which gives the bound in the supremum norm. For the H\"older bounds, we write \[ K(u,v)-k_0=\{K(u,v)-K(0,v)\}+\{K(0,v)-k_0\} \] on $N$. The first brace vanishes when $u=0$, and the second when $v=0$. By \eqref{eq:generic-angular-ratio-split}, for all large $n$, their compositions with $(u,v)$ have norms at most $C\epsilon$ in $C^\alpha$ and in $C^\beta$, respectively, where $C$ does not depend on $\epsilon$. If $K(0,v)=k_0$, only the first brace remains, which gives the bound in $C^\alpha$. If $K(u,0)=k_0$, then $K-k_0$ itself vanishes when $v=0$, and the bound in $C^\beta$ follows from \eqref{eq:generic-angular-ratio-split}. In general, both braces are bounded in $C^{\min(\alpha,\beta)}$, since $\|f\|_{C^t}\le(1+d)\|f\|_{C^{t'}}$ for $00$ and $\eta=\|K_P-k_0\|_\infty$. Since $\eta=o(\delta_n)$, for all large $n$ we have $K_P>0$ and $r_+\eta\le k_0\delta_n$, and \eqref{eq:lower-density-margin} gives \[ pK_P\ge(r_-+\delta_n)(k_0-\eta) \ge r_-k_0+k_0\delta_n-r_+\eta\ge r_-k_0 , \] and in the same way $pK_P\le(r_+-\delta_n)(k_0+\eta)\le r_+k_0-k_0\delta_n+r_+\eta\le r_+k_0$. Since a constant has H\"older norm equal to its absolute value, we obtain (H) from $\|K_P\|_{C^t}\le|k_0|+C_K\epsilon\le H'$, once $C_K\epsilon\le H'-|k_0|$. The bound in the supremum norm implies (I), and the bounds on $pK_P$ imply (W). For the laws $P_t$, we have $p=1$ and $K_{P_t}=K(u_*,t)$, a constant. The three conditions hold strictly at $(u_*,t)=(0,0)$, for (W) because $g_-'\le r_-k_00$ such that, for all large $n$, the minimax probability of an error of at least $c_0a_n\log n$ is at least $3/8$, and $\risk(\psi,\cP_r(\pi_0))\ge\sqrt{3/8}\,c_0a_n\log n$. With $c=c_0/2$, this gives \eqref{eq:sharp-probability} and the lower bound in \eqref{eq:sharp-bracket}, since $3/8\ge1/4$ and $\sqrt{3/8}\ge1/2$. For every $\theta>0$, \cref{prop:local-lower}(b), with $\mathcal P=\cP_r(\pi_0)$ and $T=\psi$, implies that there exists $c'>0$ such that $\risk(\psi,\cP_r(\pi_0))\ge c'n^{-1/2}$ for all large $n$, which is the lower bound in \eqref{eq:critical-rate} when $\theta\ge1/2$. For $\theta\ge1/2$, we take $c=c'$ and the threshold of the root-$n$ argument. For $0<\theta<1/2$, we take $c=\min\{c_0/2,c'\}$ and the larger of the two thresholds for $n$. This gives all the assertions. \end{proof} % ---------- 08-conclusion.tex \chapter{Conclusion}\label{sec:conclusion} First, when $s0$. We say that $T$ is \emph{hard at scale $t_n$ on $(\cP_n)$} if \[ \liminf_{n\to\infty}\ \inf_{\widehat T}\ \sup_{P\in\cP_n} \Pp_P\{|\widehat T-T(P)|\ge t_n\}\ge\frac38, \] where, for each $n$, the infimum is over all estimators, including randomized ones, based on $n$ independent observations. \end{definition} Hardness at scale $t_n$ implies hardness at scale $ct_n$ for every fixed $c\in(0,1]$, and on every sequence of classes $\cP_n'\supseteq\cP_n$. \begin{lemma}[From hardness to risk] \label{lem:hard-risk}\label{lem:hardness-bracket} Let $T$ be hard at scale $t_n$ on $(\cP_n)$ and a real parameter on a class $\cP$. Suppose that $\cP_n\subseteq\cP$ for all large $n$. Then, for all large $n$, \begin{equation}\label{eq:hard-consequence} \inf_{\widehat T}\sup_{P\in\cP}\Pp_P\{|\widehat T-T(P)|\ge t_n\}\ge\frac14, \qquad \risk(T,\cP)\ge\frac{t_n}2 . \end{equation} The infimum includes randomized estimators. In particular, if $0<\theta<1/2$ and $t_n=c_1\underline a_n(\theta,\tau_I)$ for a fixed $c_1>0$ and an interval $I$ as in \cref{def:bracket}, then $(T,\cP)$ satisfies the lower bracket $(\theta,I)$ with $c=c_1/2$. \end{lemma} \begin{proof} By the $\liminf$ bound, the minimax probability is at least $1/4$ over $\cP_n$, and hence over $\cP\supseteq\cP_n$, for all large $n$. For every estimator, we have $\sup_{P\in\cP}\E_P(\widehat T-T(P))^2\ge t_n^2\sup_{P\in\cP}\Pp_P\{|\widehat T-T(P)|\ge t_n\}$, which gives the RMSE bound and, with $t_n=c_1\underline a_n=2c\,\underline a_n$, the lower bracket. \end{proof} We fix the inputs of \cref{prop:local-lower}: a finite-support law $\pi_0$ for the response, the functions $D,U,V,W$ and $\lambda$, scores for $\pi_0$ (in the diagonal case, one function $\mathsf s_{\rm d}$ with $\E_{\pi_0}\mathsf s_{\rm d}=0$ and $\E_{\pi_0}[(U-a_0D)\mathsf s_{\rm d}]=1$, used for both perturbations), smoothness indices $\alpha,\beta$ with $\theta=(\alpha+\beta)/d<1/2$, an interval $[r_-,r_+]$ with $00$. By \cref{prop:local-lower}(a), there exists a fixed $c_1>0$ such that, for every constant $t_*$, the parameter $\psi+t_*$ is hard at scale $c_1a_n(\theta,\tau)\log n$ on the hard laws $(\cP^{\rm h}_n)$, where $\tau$ is computed from $[r_-,r_+]$ and $\psi$ is the target of \cref{sec:models} for $(D,U,V,W,\lambda)$. On a hard law, the response law at $x$ is the perturbation of $\pi_0$ by $(u(x),v(x))$, and $\|u\|_\infty\le Cn^{-\alpha/d}$, $\|v\|_\infty\le Cn^{-\beta/d}$. For a fixed finite family of restrictions of the types (H), (I), and (W) of \cref{lem:local}, we take $\epsilon$ sufficiently small. Then, by that lemma, every hard law satisfies them for all large $n$, and the laws $P_t$ of \cref{prop:local-lower}(b) satisfy them for every sufficiently small $u_*$ and all $t$ in an interval about zero. The reductions preserve hardness, and the lower brackets on the classes of \cref{sec:applications} follow from \cref{lem:hard-risk}. \section{Three reductions}\label{sec:transfers} Let $h$ be a bounded measurable map from the observation space into $\mathbb R^k$, with $k\ge0$, and write $\mu(P)=\E_Ph(O)$. Let $\mathcal R$ be a compact rectangle that contains $\mu(P)$ for every law $P$ that we consider, and let $b_h^2=\sup_P\E_P\|h(O)-\mu(P)\|_2^2$. When $k=0$, the terms with $h$ and $\mu$ are absent and $b_h=0$. Let $I_1,\ldots,I_m$ be compact intervals. A map $\Phi$ on $I_1\times\cdots\times I_m\times\mathcal R$ is \emph{$L$-Lipschitz} if it is Lipschitz with constant $L\ge1$ for the sum of the absolute differences of the real arguments plus the Euclidean distance in $\mathcal R$. \begin{lemma}[Reductions]\label{lem:reductions} \leavevmode \begin{enumerate}[label=(\alph*)] \item (Combination.) Suppose that $T_1,\ldots,T_m$ are real parameters on a class $\cP$ with $T_i(P)\in I_i$, and that $T_0(P)=\Phi(T_1(P),\ldots,T_m(P),\mu(P))$ for all $P\in\cP$, where $\Phi$ is $L$-Lipschitz. Then, for every $n\ge1$, \[ \risk(T_0,\cP)\le L\sum_{i=1}^m\risk(T_i,\cP)+L\,b_h\,n^{-1/2}. \] \item (Reverse relations.) Let $m=1$, let $\cP'$ be a class, and suppose that $T_0(P)\in I_1$ and \begin{equation}\label{eq:transfer-relation} |T_1(P)-\Phi(T_0(P),\mu(P))|\le e\qquad(P\in\cP'), \end{equation} where $\Phi$ is $L$-Lipschitz and $e\ge0$. Then, for every $n\ge1$ and every $t>0$ with $t\ge2e$, \begin{gather} \risk(T_1,\cP')\le L\risk(T_0,\cP')+L\,b_h\,n^{-1/2}+e, \label{eq:transfer-rmse}\\ \inf_{\widehat T_0}\sup_{P\in\cP'} \Pp_P\Bigl\{|\widehat T_0-T_0(P)|\ge\frac t{4L}\Bigr\} \ge\inf_{\widehat T_1}\sup_{P\in\cP'} \Pp_P\{|\widehat T_1-T_1(P)|\ge t\}-\frac{16L^2b_h^2}{nt^2}. \label{eq:moment-probability-transfer} \end{gather} Consequently, suppose that for all large $n$ these hypotheses hold with $\cP'=\cP'_n$ and $e=e_n\le t_n/2$, where $h$, $\mathcal R$, $I_1$, and $\Phi$ do not depend on $n$. If $T_1$ is hard at scale $t_n$ on $(\cP'_n)$, and if $k=0$ or $nt_n^2\to\infty$, then $T_0$ is hard at scale $t_n/(4L)$ on $(\cP'_n)$. \item (Kernels.) Let $K$ be a Markov kernel from the observation space of a class $\cP_0$ to that of a class $\cP_1$ that does not depend on the unknown law, and let $KP$ be the law of the output when the input has law $P$. Suppose that $KP\in\cP_1$ and $T_1(KP)=T_0(P)$ for every $P\in\cP_0$. Then, for every $n\ge1$ and $t>0$, \[ \inf_{\widehat T_0}\sup_{P\in\cP_0}\Pp_P\{|\widehat T_0-T_0(P)|\ge t\} \le\inf_{\widehat T_1}\sup_{Q\in\cP_1}\Pp_Q\{|\widehat T_1-T_1(Q)|\ge t\}, \] and $\risk(T_0,\cP_0)\le\risk(T_1,\cP_1)$. In particular, upper bounds for $T_1$ on $\cP_1$ hold for $T_0$ on $\cP_0$, lower bounds and probability bounds for $T_0$ on $\cP_0$ hold for $T_1$ on $\cP_1$, and if $T_0$ is hard at scale $t_n$ on $(\cP_{0,n})$, then $T_1$ is hard at scale $t_n$ on $(K\cP_{0,n})$. \end{enumerate} \end{lemma} \begin{proof} Let $\widehat\mu$ be the coordinatewise projection onto $\mathcal R$ of the empirical mean of $h(O_1),\ldots,h(O_n)$, and let $\Pi_i$ be the projection onto $I_i$. Since these projections do not increase distance to points of $\mathcal R$ and $I_i$, we have $\E_P\|\widehat\mu-\mu(P)\|_2^2\le b_h^2/n$. Suppose that real parameters $S,S_1,\ldots,S_m$ satisfy $S_i(P)\in I_i$ and $|S(P)-\Phi(S_1(P),\ldots,S_m(P),\mu(P))|\le e$. Given estimators $\widehat S_i$, the estimator $\widehat S=\Phi(\Pi_1\widehat S_1,\ldots,\Pi_m\widehat S_m,\widehat\mu)$ satisfies \[ |\widehat S-S(P)| \le L\Bigl\{\sum_{i=1}^m|\widehat S_i-S_i(P)| +\|\widehat\mu-\mu(P)\|_2\Bigr\}+e . \] With $S=T_0$, $S_i=T_i$, and $e=0$, (a) follows from Minkowski's inequality in $L^2(P^n)$ by taking suprema and infima, without independence between estimators. With $m=1$, $S=T_1$, and $S_1=T_0$, the same argument on $\cP'$ proves \eqref{eq:transfer-rmse}. If moreover $|\widehat S-T_1(P)|\ge t\ge2e$, the display implies that $|\widehat T_0-T_0(P)|\ge t/(4L)$ or $\|\widehat\mu-\mu(P)\|_2\ge t/(4L)$, and by Chebyshev's inequality the latter event has probability at most $16L^2b_h^2/(nt^2)$. Since $\widehat S$ is an estimator of $T_1$, \eqref{eq:moment-probability-transfer} follows by taking the supremum over $P\in\cP'$ and then the infimum over $\widehat T_0$. Applying it with $t=t_n$ and taking the $\liminf$ gives the hardness assertion, since the error term vanishes when $k=0$ and otherwise tends to zero by assumption. For (c), we apply $K$ independently to the $n$ observations, with randomness independent of the sample, and evaluate $\widehat T_1$ on the outputs. Since, under $P$, the outputs have law $(KP)^n$ and $T_1(KP)=T_0(P)$, the resulting error has the law of the error of $\widehat T_1$ under $KP\in\cP_1$. The suprema and infima give both inequalities and the remaining assertions. \end{proof} \Cref{lem:reductions}(c) is the elementary direction of Le Cam's comparison of experiments by randomization~\cite[Chapter~2]{lecam}: an experiment is at least as informative as its image under a Markov kernel that does not depend on the unknown law. Composing an estimator for the image with the kernel, we obtain an estimator whose error has the same law. This gives both the inequality for error probabilities and the inequality for risks. For a ratio we use (a) with $\Phi(x,y)=x/y$ on $[-C_0,C_0]\times[\delta_{\mathcal D},C_0]$, which is $L$-Lipschitz with $L=\max\{1,\delta_{\mathcal D}^{-1},C_0\delta_{\mathcal D}^{-2}\}$, since $|x/y-x'/y'|\le|x-x'|/\delta_{\mathcal D}+C_0|y-y'|/\delta_{\mathcal D}^2$. A target that differs from another by a sign and an additive constant has the same estimation errors, so that all bounds transfer between the two. \begin{remark}[Combining rates]\label{rem:combining-rates} For fixed $\tau$ and $\nu$, and $0<\theta<1/2$, we have $\overline a_n(\theta,\tau,\nu)=n^{-\theta}e^{O(\sqrt{\log n})}$, since $(\log n)^c=e^{o(\sqrt{\log n})}$ for every fixed $c$. Hence, if $\theta_1<\min(\theta_2,1/2)$, then $\overline a_n(\theta_2,\tau,\nu_2)+n^{-1/2}=o(\overline a_n(\theta_1,\tau,\nu_1))$, and if $\theta_1=\theta_2$ and $\nu_1\le\nu_2$, then $\overline a_n(\theta_1,\tau,\nu_1)\le\overline a_n(\theta_2,\tau,\nu_2)$. For $\theta\ge1/2$ all these rates are $n^{-1/2}$. It follows from \cref{lem:reductions}(a) that if each $(T_i,\cP)$ satisfies the upper bracket $(\theta_i,I,\nu_i)$ with a common interval $I$, then $(T_0,\cP)$ satisfies the upper bracket $(\theta,I,\nu)$ with $\theta=\min_i\theta_i$ and $\nu=\max\{\nu_i:\theta_i=\theta\}$. \end{remark} % ---------- B-missing.tex \chapter{Missing-data means and treatment effects}\label{sec:causal-extension} The treatment effects inherit lower bounds below the threshold from one MAR family through a kernel with the observed regression in one arm and the constant $1/2$ in the other. Their arm means inherit upper bounds for the MAR mean by keeping one arm. \section{The MAR family and two maps}\label{sec:binary-embedding} We take $Z=(R,RY)$ with $R,Y\in\{0,1\}$ and $(D,U,V,W,\lambda)=(R,1,RY,0,1)$, and we let $\pi_0$ give $(R,RY)=(0,0),(1,0),(1,1)$ the probabilities $1/2,1/4,1/4$. Then $(w_0,a_0,b_0)=(1/2,2,1/2)$, and the residual $\mathsf r=(1-2R,\,RY-R/2)^\top$ takes the values $(1,0)$, $(-1,-1/2)$, and $(-1,1/2)$ at the three support points, so that $\Sigma=\Cov_{\pi_0}(\mathsf r)=\operatorname{diag}(1,1/8)$. Since $\Sigma$ is positive definite, the functions $(\mathsf s_u,\mathsf s_v)^\top=\Sigma^{-1}\mathsf r=(1-2R,\,8RY-4R)^\top$ are scores for $\pi_0$, as in \cref{sec:generic-embedding}. The perturbed response law $\pi_{u,v}(\{z\})=\pi_0(\{z\})\{1+u\mathsf s_u(z)+v\mathsf s_v(z)\}$ is \begin{equation}\label{eq:binary-embedding} \begin{gathered} P(R=0\mid X)=\frac{1+u}2,\\ P(R=1,Y=0\mid X)=\frac{1-u}4-v,\quad P(R=1,Y=1\mid X)=\frac{1-u}4+v, \end{gathered} \end{equation} so that \begin{equation}\label{eq:binary-moments} w=\frac{1-u}2,\qquad \frac1w=\frac2{1-u},\qquad \frac1{1-w}=\frac2{1+u},\qquad b=\frac12+\frac{2v}{1-u}, \end{equation} and $\psi(P)=\int abg=\E_P[b(X)]$ is the MAR mean. These observed laws are those of full-data laws in which $Y$ is Bernoulli with mean $b(X)$ and independent of $R$ given $X$. The lower bound for the missing-data mean in~\cite[Section~3]{robins2009} uses binary models of this kind. For indices $(\alpha,\beta')$ with $(\alpha+\beta')/d<1/2$ and an interval $[r_-,r_+]$ with $00$, and $\zeta_0\in(0,1)$. There are $H_1<\infty$, $c>0$, and, for each $n_1\ge1$, a map $\widehat w$ from $n_1$ observations of $(X,R)$ to $C^\infty$ functions on $[0,1]^d$, such that $(x,\text{data})\mapsto\widehat w(x)$ is measurable, $\epsilon\le\widehat w\le1-\epsilon$, $\|\widehat w\|_{C^\alpha}\le H_1$, and \[ \Pp_P\{\|\widehat w-w\|_\infty>\zeta_0\}\le c^{-1}e^{-cn_1} \] for every law of $(X,R)$ such that $R\in\{0,1\}$, the covariate density is at least $p_-$, and $w=\E[R\mid X]\in\mathcal W$. \end{lemma} \begin{proof} Let $\gamma=\min(\alpha,1)$ and $C_0=\sqrt d\,H_0$. Every $f\in\mathcal W$ satisfies $|f(x)-f(y)|\le C_0\|x-y\|_2^\gamma$, by its H\"older bound if $\alpha\le1$ and its gradient bound if $\alpha>1$. We fix an integer $k$ with $C_0(2\sqrt d/k)^\gamma\le\zeta_0/4$, we partition $[0,1]^d$ into $k^d$ congruent cubes $Q$, and we let $Q^+$ be the set of points of $[0,1]^d$ at $\ell^\infty$-distance at most $1/(2k)$ from $Q$, so that $\operatorname{diam}Q^+\le2\sqrt d/k$. We fix a $C^\infty$ partition of unity $(\varphi_Q)$ of $[0,1]^d$ with $\varphi_Q\ge0$ and $\supp\varphi_Q\subseteq Q^+$; for instance, products of one-dimensional ones. Let $N_Q$ be the number of observations with $X_i\in Q$, let $\widetilde w_Q$ be the average of their $R_i$, with $\widetilde w_Q=1/2$ if $N_Q=0$, let $\widehat w_Q$ be $\widetilde w_Q$ clipped to $[\epsilon,1-\epsilon]$, and let $\widehat w=\sum_Q\widehat w_Q\varphi_Q$. Then $\widehat w$ is $C^\infty$, depends measurably on the data, takes values in $[\epsilon,1-\epsilon]$ as a convex combination, and satisfies $\|\widehat w\|_{C^\alpha}\le H_1=\sum_Q\|\varphi_Q\|_{C^\alpha}$, which depends only on $k$, $d$, and $\alpha$. Let $\pi_Q=\Pp(X\in Q)\ge p_-k^{-d}$ and $m_Q=\E[R\mid X\in Q]$. Since $m_Q$ averages $w$ over $Q$, we have $m_Q\in[\epsilon,1-\epsilon]$ and $|m_Q-w(x)|\le\zeta_0/4$ for $x\in Q^+$. Since clipping does not increase $|\widetilde w_Q-m_Q|$, and $\varphi_Q(x)=0$ unless $x\in Q^+$, we have \[ \|\widehat w-w\|_\infty\le\max_Q|\widetilde w_Q-m_Q|+\zeta_0/4 . \] Conditionally on $N_Q=s\ge1$, the sum of the $R_i$ with $X_i\in Q$ is $\operatorname{Binomial}(s,m_Q)$, and Hoeffding's inequality~\cite[Theorem~1]{hoeffding1963} implies that $\Pp\{|\widetilde w_Q-m_Q|>\zeta_0/2\mid N_Q=s\}\le2e^{-s\zeta_0^2/2}$, a bound that also holds for $s=0$. Since $N_Q\sim\operatorname{Binomial}(n_1,\pi_Q)$, we obtain \[ \Pp\{|\widetilde w_Q-m_Q|>\zeta_0/2\} \le2\bigl(1-\pi_Q+\pi_Qe^{-\zeta_0^2/2}\bigr)^{n_1} \le2e^{-n_1\pi_Q(1-e^{-\zeta_0^2/2})}. \] A union bound over the $k^d$ cells gives $\Pp\{\|\widehat w-w\|_\infty>\zeta_0\} \le2k^de^{-n_1p_-k^{-d}(1-e^{-\zeta_0^2/2})}$, which is at most $c^{-1}e^{-cn_1}$ for a suitable $c>0$. \end{proof} \begin{proof}[Proof of the upper bracket on $\mathcal M^{\rm sep}$] We fix $\zeta\in(0,1)$, and we let $\theta=(\alpha+\beta)/d$ and $\nu=q_\alpha+q_\beta$. We apply \cref{lem:smooth-propensity-pilot} with $\zeta_0=\zeta\epsilon$ to the first $n_1=\lfloor n/2\rfloor$ observations. Conditionally on these, the remaining $N=n-n_1$ observations are independent with law $P$, and we use them with the observables \[ D'=\frac R{\widehat w(X)},\qquad U'=1,\qquad V'=\frac{RY}{\widehat w(X)},\qquad W'=0,\qquad\lambda'=1, \] which depend on $X$, as the assertions stated before the proof of \cref{thm:generic}(a) in \cref{sec:tuning} allow. These are bounded by $M_0'=\epsilon^{-1}$, and \[ w'=\frac w{\widehat w}\ge\frac\epsilon{1-\epsilon},\qquad a'=\frac{\widehat w}w,\qquad b'=b,\qquad g'=\frac{pw}{\widehat w},\qquad \int a'b'g'=\int bp . \] Since $\widehat w\in\cH^\alpha(H_1)$ and $w\in\mathcal W$, the function $a'=\widehat w\cdot(1/w)$ lies in a ball $\cH^\alpha(H_2)$, with $H_2$ depending only on $d,\alpha,H_0,H_1,\epsilon$. Indeed, by the Leibniz and chain rules, each derivative of $a'$ of order at most $\lceil\alpha\rceil-1$ is a polynomial in the derivatives of $\widehat w$ and $w$ of these orders and in $1/w\le\epsilon^{-1}$, and each factor is bounded and H\"older continuous with exponent $\alpha-\lceil\alpha\rceil+1$. On the event $\mathcal A=\{\|\widehat w-w\|_\infty\le\zeta\epsilon\}$, we have $|w/\widehat w-1|\le\zeta$, hence $g'\in I_\zeta$. Conditionally on a pilot outcome in $\mathcal A$, the remaining observations have law in $\cP(\alpha,\beta)$ for $(D',U',V',W',\lambda')$ with the constants $M_0'$, $\delta'=\epsilon/(1-\epsilon)$, $H'=\max(H_0,H_2)$, and $[g_-,g_+]=I_\zeta$, and its target is the MAR mean. We let $\widehat\psi$ be the estimator of \cref{thm:generic}(a) from \cref{sec:tuning}, applied with these constants and $\widehat w$ to the last $N$ observations and clipped to $[0,1]$. Its tuning depends only on these constants, so that $\widehat\psi$ is measurable, and on $\mathcal A$ its conditional mean squared error is at most $C\,\overline a_N(\theta,\tau_{I_\zeta},\nu)^2$, with $C$ independent of the pilot. Since its squared error is at most one off $\mathcal A$, its mean squared error is at most $C\,\overline a_N(\theta,\tau_{I_\zeta},\nu)^2+c^{-1}e^{-cn_1}$. Since $n/2\le N\le n$, the factors $N^{-\theta}$, the powers of $\log N$, and $e^{-\kappa\sqrt{\log N}}$ differ from their values at $n$ by bounded factors, the last because $\sqrt{\log n}-\sqrt{\log N}\le\log2/\sqrt{\log N}$. We conclude that $\overline a_N\le C'\overline a_n$. Since $e^{-cn_1}=o(n^{-1})$ and $n^{-1/2}\le\overline a_n$ for large $n$, we obtain $\risk(\psi,\mathcal M^{\rm sep})\le C_\zeta\overline a_n(\theta,\tau_{I_\zeta},\nu)$. \end{proof} \section{Proof of \texorpdfstring{\cref{cor:applications}}{the corollary} for these targets}\label{sec:broad-lower} \emph{Upper bounds.} We take the choices of \cref{sec:mar-implications}, with the response space $\{(0,0)\}\cup(\{1\}\times[0,1])$, $M_0=1$, $\delta=\eta$, and $[g_-,g_+]=[\eta,\eta^{-1}]$. Then $\mathcal M_\eta(\alpha,\beta,H)\subseteq\cP(\alpha,\beta)$, and \cref{thm:generic}(a) implies the upper bracket for the MAR mean on $\mathcal M_\eta$. The upper bracket on $\mathcal M^{\rm sep}$ with $I_\zeta$, for every $H_0>1$, was proved in \cref{sec:broad-inclusions}. The map that keeps arm $j$ sends $\mathcal T_\eta(\alpha,\beta_0,\beta_1,H)$ into $\mathcal M_\eta(\alpha,\beta_j,H)$, since its output has observation probability $w_j$ with $\eta\le w_j\le1-\eta$, $1/w_j=a_j\in\cH^\alpha(H)$, and $\eta\le w_jp=g_j\le\eta^{-1}$, and regression $b_j\in\cH^{\beta_j}(H)$. It sends $\mathcal T^{\rm sep}$ into the class $\mathcal M^{\rm sep}$ with indices $(\alpha,\beta_j)$ and radius $H_0+1$, since $\epsilon\le1-w\le1-\epsilon$ and $\|1-w\|_{C^\alpha}\le1+\|w\|_{C^\alpha}$. Since the MAR mean of the output is $\mu_j$, \cref{lem:reductions}(c) shows that $\mu_j$ satisfies the upper bracket $((\alpha+\beta_j)/d,I,q_\alpha+q_{\beta_j})$ on $\mathcal T_\eta$, with $I=[\eta,\eta^{-1}]$, and on $\mathcal T^{\rm sep}$ with $I_\zeta$. \Cref{lem:reductions}(a), applied to $T_{\rm ATE}=\mu_1-\mu_0$ and to the relations on the left of \eqref{eq:att-atu-identities}, together with \cref{rem:combining-rates}, gives the upper brackets of the three effects with the indices \eqref{eq:treatment-smoothness} and $\nu=q_\alpha+q_{\beta_*}$. \emph{Lower bounds.} We treat two cases at once. In the first, $\mathcal M=\mathcal M_\eta$, $\mathcal T=\mathcal T_\eta$, $I=[\eta,\eta^{-1}]$, and $[r_-,r_+]=[2\eta,2\eta^{-1}]$. In the second, $\mathcal M=\mathcal M^{\rm sep}$, $\mathcal T=\mathcal T^{\rm sep}$, and $I=[r_-,r_+]=[p_-,p_+]$. Then $r_-<12$ give $\max\{\delta,g_-\}=\eta1$. A Rademacher variable $Y$ satisfies the diagonal alternative with residual variance one. This gives the brackets of these four targets. \emph{Upper brackets.} We write $m_j=\E_P[Y^j]$ for $j=1,2$. For a fixed known measurable $f_0$ with $B_0=\|f_0\|_\infty<\infty$, we set $\ell_0=\E_P[f_0(X)^2-2Yf_0(X)]$ and $E_{f_0}=\E_P[(f(X)-f_0(X))^2]$. Since $\E_P[Yf_0(X)]=\E_P[f(X)f_0(X)]$, we have \[ E_*=Q_2-m_1^2,\qquad R_*^2=\frac{Q_2-m_1^2}{\max\{m_2-m_1^2,v_{\min}\}},\qquad E_{f_0}=Q_2+\ell_0, \] where the middle identity holds on $\cP_{\rm quad}(s;v_{\min})$. In \cref{lem:reductions}, we use $h(O)=(Y,Y^2)$, with $\mathcal R=[-1,1]\times[0,1]$, for $E_*$ and $R_*^2$, and $h(O)=f_0(X)^2-2Yf_0(X)$, with $\mathcal R=[-B_0^2-2B_0,B_0^2+2B_0]$, for $E_{f_0}$. Since the denominator is at least $v_{\min}$, the three maps are Lipschitz on $[0,1]\times\mathcal R$. By the bracket of $Q_2$ and \cref{lem:reductions}(a), the three targets satisfy the upper bracket $(2s/d,[p_-,p_+],2q_s)$ on their classes. \emph{Lower brackets below the threshold.} Let $s0$ depends only on $s,H_0$ and satisfies $\|\varphi\|_{C^s}\le H_0$ and $\|\varphi\|_\infty\le1/2$. Let $P_t$ have $p=1$ and $P_t(Y=y\mid X)=\{1+yt\varphi(X)\}/2$ for $|t|<1$. These laws lie in $\cP_{\rm quad}(s;v_{\min})$, with $\E_{P_t}Y=0$ and $\Var_{P_t}(Y)=1$. We set $A_\varphi=\int\varphi^2>0$ and $B_\varphi=\int\varphi f_0$. Then $E_*(P_t)=R_*^2(P_t)=t^2A_\varphi$ and $E_{f_0}(P_t)=t^2A_\varphi-2tB_\varphi+\int f_0^2$. Since the excess-loss derivatives at $t=1/4$ and $t=1/2$ differ by $A_\varphi/2$, at one of these points their absolute value is at least $A_\varphi/4$; the derivatives of $E_*$ and $R_*^2$ there are at least $A_\varphi/2$. The densities relative to the law with uniform design and Rademacher response are $1+tY\varphi(X)\ge1/2$, with bounded slope. \Cref{lem:parametric-path} yields the root-$n$ lower bounds for all three targets. Only the constants for $E_{f_0}$ may also depend on $B_0$. \emph{The trial.} The map $(X,A,Y_{\rm obs})\mapsto(X,(2A-1)(2Y_{\rm obs}-1))$ sends the trial class into $\cP_{\rm quad}(s)$, since its response lies in $[-1,1]$ and its conditional mean is $f$. Conversely, we draw an independent $A\sim\operatorname{Bernoulli}(1/2)$ and map $(X,Y)$ to $(X,A,\{1+(2A-1)Y\}/2)$. Since the outcome lies in $[0,1]$ and its arm means are $(1\pm f)/2$, this kernel sends $\cP_{\rm quad}(s)$ into the trial class. Both kernels preserve $p$, $f$, and the two targets. \Cref{lem:reductions}(c) transfers the upper bounds along the first kernel and the lower and probability bounds along the second. % ---------- D-further.tex \chapter{Wald ratios and the overlap-weighted effect}\label{sec:additional-effects} We recall the causal interpretations of the targets of \cref{sec:further-applications} and prove \cref{cor:applications}. The interpretations use identification assumptions; the classes restrict observed laws. \section{Wald ratios}\label{sec:wald} \begin{example}[Wald ratio]\label{ex:wald} Let $A(z)$ and $Y(a)$ denote potential treatments and outcomes, and suppose that $A=A(Z)$ and $Y=Y(A)$. If $Z\perp(A(0),A(1),Y(0),Y(1))\mid X$ and $A(1)\ge A(0)$, then $T_{\rm W}$ is the complier average treatment effect $\E[Y(1)-Y(0)\mid A(1)>A(0)]$, the local average treatment effect of~\cite{imbensangrist,angristimbensrubin}, in the form with covariates given in~\cite{frolich,dml}. Its numerator and denominator are average effects of the instrument, that is, differences of two MAR means. \end{example} \begin{example}[Average conditional Wald ratio]\label{ex:conditional-wald} Suppose that $P(Z=1\mid X)=1/2$ and $A\le Z$, so that the instrument is assigned with equal probability and treatment is unavailable when $Z=0$. Under the assumptions of \cref{ex:wald}, $b(X)$ is the conditional complier effect. If also $\E[Y(1)-Y(0)\mid X,A(1)]=\E[Y(1)-Y(0)\mid X]$, which says that the covariates account for differences in mean effects between compliers and never-takers, then $T_{\rm CW}$ is the ATE; see~\cite[Theorem~1 and the following discussion]{wangtchetgen}. The two Wald targets differ in general, since $T_{\rm W}$ averages $b(X)$ with weights proportional to $q(X)$. \end{example} \begin{proof}[Proof of \cref{cor:applications} for the Wald ratios] \emph{Wald ratio.} We write $\mathcal N=\E_P[\Delta_Y(X)]\in[-1,1]$ and $\mathcal D=\E_P[\Delta_A(X)]\in[c_W,1]$. With the instrument as the treatment, the deterministic maps $(X,Z,A,Y)\mapsto(X,Z,Y)$ and $(X,Z,A,Y)\mapsto(X,Z,A)$ send $\mathcal I_{\eta,c_W}(\alpha,\beta,H)$ into $\mathcal T_\eta(\alpha,\beta,\beta,H)$, and the ATEs of the images are $\mathcal N$ and $\mathcal D$. By \cref{cor:applications} for the ATE on $\mathcal T_\eta(\alpha,\beta,\beta,H)$ and \cref{lem:reductions}(c), both satisfy the upper bracket $((\alpha+\beta)/d,[\eta,\eta^{-1}],q_\alpha+q_\beta)$ on $\mathcal I_{\eta,c_W}(\alpha,\beta,H)$, and applying \cref{lem:reductions}(a) to the ratio, with $C_0=1$ and $\delta_{\mathcal D}=c_W$, together with \cref{rem:combining-rates}, we obtain the upper bracket for $T_{\rm W}$. For the lower bracket, the kernel $(X,Z,Y)\mapsto(X,Z,Z,Y)$, which sets $A=Z$, maps $\mathcal T_\eta(\alpha,\beta,\beta,H)$ into $\mathcal I_{\eta,c_W}(\alpha,\beta,H)$: its output has $m_{A,z}=z$, whose H\"older norm is $z\le1\alpha$. Without smoothness assumptions on the covariate density, and without the assumption that the conditional treatment effect is constant, Robins et al.~\cite[Remark~4.2]{hoif2008} refer to $n^{-2s/d}$, for $s\alpha$, we use indices $(\alpha,\alpha)$, $c=1/4$, $(\ell_1,\ell_2)=(u+v,0)$, and $(U,V,W,\lambda)=(A,A,A,-1)$, with the diagonal score $\mathsf s_1$. In both cases $D=1$, $[r_-,r_+]=[p_-,p_+]$, and the indices give the exponent $\theta$. For membership in $\cP_{\rm OW}(\alpha,\beta)$, we require $p_-\le p\le p_+$, $w\in\cH^\alpha(H_0)$, $m\in\cH^\beta(H_0)$, and $\epsilon\le w\le1-\epsilon$. The density and overlap conditions are of types (W) and (I). In the family with indices $(\alpha,\beta)$, both smoothness conditions are of type (H), with $k_0=1/2