% \documentclass[11pt]{article} \usepackage[letterpaper, margin=1.00in]{geometry} \usepackage[T1]{fontenc} \usepackage{lmodern} \usepackage{microtype} \usepackage{amsmath,amssymb,amsthm,mathtools,bm,mathrsfs} \usepackage{enumitem} \usepackage{xcolor} \usepackage{hyperref} \usepackage[nameinlink,noabbrev]{cleveref} \hypersetup{ colorlinks=true, linkcolor=blue!55!black, citecolor=blue!55!black, urlcolor=blue!55!black, pdftitle={Free energy of spherical and Ising perceptrons with arbitrary Lipschitz channels} } \newtheorem{theorem}{Theorem}[section] \newtheorem{proposition}[theorem]{Proposition} \newtheorem{lemma}[theorem]{Lemma} \newtheorem{corollary}[theorem]{Corollary} \theoremstyle{remark} \newtheorem{remark}[theorem]{Remark} \newcommand{\E}{\mathbb E} \newcommand{\Pp}{\mathbb P} \newcommand{\R}{\mathbb R} \newcommand{\1}{\mathbf 1} \newcommand{\dd}{\,\mathrm d} \newcommand{\Q}{\mathscr Q} \newcommand{\Pb}{\mathscr P_{\mathrm b}} \DeclareMathOperator{\Lip}{Lip} \DeclareMathOperator{\Var}{Var} \title{Parisi formula for Ising and Spherical Perceptrons } \author{} \date{October 7, 2026} \begin{document} \maketitle \begin{abstract} We prove a Parisi formula for Ising and spherical perceptrons with Gaussian disorder and i.i.d. random potentials having a common Lipschitz bound and integrable values at zero. We assume $M/N\to\alpha\in(0,\infty)$, where $M$ is the number of Gaussian vectors and $N$ is the number of spins. The proof combines cavity in $M$ and $N$ with the Aizenman–Sims–Starr scheme for the lower bound and Guerra’s interpolation for the upper bound. As consequences, we answer several problems posed by Talagrand (2011) and establish a quenched large deviation principle for the empirical distribution of the Gaussian projections, resolving a conjecture of Bolthausen and Kistler (2010). We also resolve the sharp-threshold conjecture of Aubin, Perkins, and Zdeborová (2019) for the $u$-function binary perceptron at every positive width, and establish the predicted variational formulas for the entropy and capacity of the negative spherical perceptron. Finally, the zero-temperature limit yields the scalar projection-pursuit formulas conjectured by Montanari and Zhou (2025, 2026), in both the supervised and unlabelled settings. \end{abstract} \setcounter{tocdepth}{1} \tableofcontents \subsection*{Disclaimer} This document and its arguments were generated through discussions with ChatGPT 5.6 Pro. An initial draft for the capacity of the negative spherical perceptron was generated on July 15, 2026, and we extended the argument to i.i.d. lipschitz potentials, both Ising and spherical, on September 17, 2026. While distilling the main ideas, carefully checking the proof, and polishing the draft, we learned on October 6, 2026, that OpenAI had released overlapping results across three manuscripts~\cite{OpenAIIsingPerceptron2026,OpenAISphericalPerceptron2026,OpenAIJamming2026}. Given these circumstances, we decided to post this working draft on Hexagon. \section{Introduction}\label{sec:introduction} We study Ising and spherical perceptrons with Gaussian disorder and i.i.d. random Lipschitz potentials. Let $\nu$ be a Borel probability measure on $\mathbb R$, and let $(\theta,x)\mapsto u_\theta(x)$ be Borel measurable on $\mathbb R^2$. We assume that $\Lip(u_\theta)\le L<\infty$ for $\nu$-almost every $\theta$ and that $\int_{\mathbb R}|u_\theta(0)|\nu(\dd\theta)<\infty$. Let $\Theta_1,\ldots,\Theta_M$ be i.i.d. with law $\nu$, and let $g_1,\ldots,g_M\in\mathbb R^N$ be independent standard Gaussian vectors, independent of $(\Theta_k)_{k\leq M}$. The Hamiltonian of the perceptron model is defined by \[ H_{N,M}^{\nu}(\sigma) =\sum_{k=1}^M u_{\Theta_k}(h_k(\sigma)),\quad\textnormal{where}\quad h_k(\sigma)=\frac{g_k\cdot\sigma}{\sqrt N} \] is the Gaussian projection. We consider the spherical and Ising configuration spaces \[ \Sigma_N^{\mathrm S} =\{\sigma\in\mathbb R^N:\|\sigma\|_2^2=N\}, \qquad \Sigma_N^{\mathrm I}=\{-1,1\}^N. \] Let \(\mu_N^{\mathrm S}\) and \(\mu_N^{\mathrm I}\) be the uniform probability measures on \(\Sigma_N^{\mathrm S}\) and \(\Sigma_N^{\mathrm I}\), respectively. For $\kappa\in\{\mathrm S,\mathrm I\}$, the partition function is \[ Z_{N,M}^{\kappa,\nu} =\int_{\Sigma_N^\kappa} e^{H_{N,M}^{\nu}(\sigma)}\,\mu_N^\kappa(\dd\sigma). \] Our main object is the normalized free energy $N^{-1}\log Z_{N,M}^{\kappa,\nu}$. Throughout, we take $M=M_N$ with $M_N/N\to\alpha\in(0,\infty)$. The study of perceptron model in statistical physics was intiated by Gardner and Gardner--Derrida \cite{Gardner1988,GardnerDerrida1988}. For constraints \(h_k(\sigma)\ge b\), they studied the volume of configurations satisfying all \(M\) inequalities and the density \(M/N\) at which solutions disappear. The entropy records the logarithmic volume or number of solutions; the storage capacity is the critical density for their existence. Krauth and M\'ezard subsequently predicted the capacity for Ising spins \cite{KrauthMezard1989}. For spherical spins and nonnegative margin \(b\), Shcherbina and Tirozzi proved Gardner's replica-symmetric entropy formula \cite{ShcherbinaTirozziGardner}. For Ising spins, Talagrand developed cavity arguments, comparing systems with slightly different numbers of spins or Gaussian vectors, and proved replica-symmetric formulas under restrictions on the potential and density \cite{TalagrandBookI}. Bolthausen, Nakajima, Sun, and Xu later obtained small-density formulas for a broad class of activations \cite{BolthausenNakajimaSunXu}. Nakajima and Sun proved concentration and universality results, together with sharp-threshold sequences for broad classes of Ising constraints \cite{NakajimaSun}. Their results control random fluctuations and changes in the disorder distribution, and locate the satisfiability transition around a sequence of densities. Proving that this sequence converges requires an additional argument. Beyond replica symmetry, Gy\"orgyi and Reimann developed formulas involving a full overlap distribution for the spherical perceptron at positive temperature \cite{GyorgyiReimann2000}. Franz and Parisi studied the constraints \(h_k(\sigma)\ge b\) with \(b<0\) and their jamming transition, where feasible configurations disappear as \(M/N\) increases \cite{FranzParisiJamming}. Their predictions concern both the entropy and the geometry near this transition. For Ising spins, Aubin, Perkins, and Zdeborov\'a studied the symmetric \(U\)-function constraints \(|h_k(\sigma)|\ge K\). They identified a large-\(K\) regime in which the first-moment estimate, based on the expected number of solutions, no longer gives the predicted capacity \cite{AubinPerkinsZdeborova}. Our main theorem determines the limiting free energy under the assumptions above. We also prove a quenched large-deviation principle for the empirical distribution of the Gaussian projections \(h_k(\sigma)\). For a typical fixed realization of the Gaussian vectors, this describes the exponential probabilities of atypical projection distributions when \(\sigma\) is sampled from \(\mu_N^\kappa\). Finally, the zero-temperature limit determines optimal empirical averages over one-dimensional projections, in both supervised and unlabelled settings. Define the overlap-path space by \[ \Q=\left\{q:[0,1]\to[0,1]: q\text{ is right-continuous and nondecreasing},~~~~ q(1)=1\right\}. \] We identify paths that agree almost everywhere on \([0,1)\) and equip \(\Q\) with the \(L^1\) metric. The spin cost \(\mathcal A_\kappa\) is defined in \eqref{eq:spherical-entropy} for spherical spins and \eqref{eq:intro-ising-cost} for Ising spins. The channel functional \(\Phi_u\) is defined in \eqref{eq:intro-channel-pde} through the scalar Parisi PDE \eqref{eq:intro-parisi-pde}. These definitions follow \cref{thm:iid-channels}. \begin{theorem}[Free energy for i.i.d.\ Lipschitz channels] \label{thm:iid-channels} For either prior, \begin{equation*} \frac1N\log Z_{N,M_N}^{\kappa,\nu} \longrightarrow \inf_{q\in\Q}\left\{\mathcal A_\kappa(q) +\alpha\int_{\mathbb R}\Phi_{u_\theta}(q)\,\nu(\dd\theta)\right\} \qquad\text{in }L^1. \end{equation*} Convergence holds in \(L^2\) if \(\int_{\mathbb R}|u_\theta(0)|^2\,\nu(\dd\theta)<\infty\), and almost surely and in every finite \(L^p\) if there is \(B<\infty\) such that \(|u_\theta(x)|\le B\) for every \(x\) and \(\nu\)-almost every \(\theta\). % For spherical spins, one may again restrict to \(q(1-)<1\). \end{theorem} Note that a deterministic potential \(u\) corresponds to a point-mass channel law. In this case, write \(H_{N,M}^{u}\), \(Z_{N,M}^{\kappa,u}\), and \(F_{N,\kappa}^{u}=N^{-1}\log Z_{N,M_N}^{\kappa,u}\), and put \begin{equation*} \mathcal P_{\kappa,u}(\alpha) =\inf_{q\in\Q}\{\mathcal A_\kappa(q)+\alpha\Phi_u(q)\}. \end{equation*} The theorem gives convergence in \(L^2\), since \(u(0)\) is finite. We next define the functionals \(\Phi_u\) and \(\mathcal A_\kappa\) on the overlap paths \(q\in\Q\). Recall that these paths are right-continuous and nondecreasing, take values in \([0,1]\), and satisfy \(q(1)=1\). Let \[ \Pb=\left\{a:[0,1)\to[0,\infty): a\text{ is bounded, right-continuous and nondecreasing}\right\}. \] These paths will enter the variational definition of the Ising spin cost. For \(c\in\Pb\), write \(c_*=c(1-)\) and \(\gamma_c(t)=\operatorname{Leb}\{r\in[0,1):c(r)\le t\}\). Thus \(\gamma_c\) is the distribution function of \(c(U)\) for \(U\) uniform on \([0,1)\). We use the same notation for \(q\in\Q\), restricting \(q\) to \([0,1)\). For \(T>0\), a nonnegative Lebesgue-integrable function \(\gamma:[0,T]\to[0,\infty)\), and a globally Lipschitz function \(h:\mathbb R\to\mathbb R\), let \(f_{\gamma,h}\) solve the backward Parisi PDE \begin{equation}\label{eq:intro-parisi-pde} \partial_t f+\frac12\bigl(f_{xx}+\gamma(t)f_x^2\bigr)=0, \qquad f(T,x)=h(x). \end{equation} We use the stochastic-control solution; its existence and stability are proved in \cref{lem:mz-control}. The channel functional is \begin{equation}\label{eq:intro-channel-pde} \Phi_u(q)=f_{\gamma_q,u}(0,0),\qquad T=1. \end{equation} The two spin costs are as follows. For spherical spins, \begin{equation}\label{eq:spherical-entropy} \mathcal A_{\mathrm S}(q) =\frac12\int_0^1\left\{ \frac1{\int_t^1\gamma_q(s)\,\dd s}-\frac1{1-t}\right\}\dd t. \end{equation} The outer integrand is nonnegative, and the value may be infinite. For Ising spins and \(a\in\Pb\), use the same PDE with terminal function \(\log\cosh\) and terminal time \(a_*\): \begin{equation}\label{eq:intro-ising-transform} \mathcal S_{\mathrm I}(a) =f_{\gamma_a,\log\cosh}(0,0)-\frac{a_*}{2},\qquad T=a_*. \end{equation} At \(a=0\), define $\mathcal S_{\mathrm I}(0)=0$. The Ising cost is \begin{equation}\label{eq:intro-ising-cost} \mathcal A_{\mathrm I}(q) =\sup_{a\in\Pb}\left\{\mathcal S_{\mathrm I}(a) +\frac12\int_0^1a(r)q(r)\,\dd r\right\}. \end{equation} Below, we write \(\langle a,q\rangle=\int_0^1a(r)q(r)\,\dd r\). % the diagonal value does not enter this pairing. \subsection{Direct consequences of the free energy formula} \label{sec:field-bipartite-applications} In this section, we prove direct consequences of Theorem~\ref{thm:iid-channels}, which may be of independent interest. \paragraph{Gaussian external fields.}\phantomsection \label{par:gaussian-field-extension} For a standard Gaussian vector \(g^0 \in \R^N\), independent of all other randomness, let \begin{equation}\label{eq:intro-gaussian-field-partition} Z_{N,M}^{\kappa,u,b} =\int e^{\sum_{k\le M}u(h_k(\sigma))+b\,g^0\cdot\sigma} \mu_N^\kappa(\dd\sigma),\qquad F_{N,\kappa}^{u,b}=\frac1N\log Z_{N,M_N}^{\kappa,u,b}, \end{equation} and denote the marked version by \(Z_{N,M}^{\kappa,\nu,b}\). Write \begin{equation}\label{eq:intro-field-value} \mathsf L(q)=\int_0^1(1-q(r))\,\dd r,\qquad \mathcal P_{\kappa,u}^{\mathrm{fld}}(\alpha,b) =\inf_{q\in\Q}\left\{\mathcal A_\kappa(q)+\alpha\Phi_u(q) +\frac{b^2}{2}\mathsf L(q)\right\}. \end{equation} For a deterministic potential, the Gaussian-field extension of \cref{thm:iid-channels} gives \(F_{N,\kappa}^{u,b}\to\mathcal P_{\kappa,u}^{\mathrm{fld}}(\alpha,b)\) in \(L^2\). For i.i.d.\ channels, replace \(\Phi_u\) by \(\E_\nu\Phi_{u_\Theta}\); the same \(L^1\) and \(L^2\) conclusions as in \cref{thm:iid-channels} hold. Mixing in linear channels proves this extension; see \cref{sec:gaussian-field-proof}. \paragraph{Spherical i.i.d.\ fields.} Let \(\pi\) be a Borel probability measure on \(\mathbb R\) satisfying \[ v_\pi:=\int_{\mathbb R}h^2\,\pi(\dd h)<\infty. \] Let \(h_1,h_2,\ldots\) be i.i.d.\ with law \(\pi\), independent of the Gaussian constraint vectors \((g_k)_{k\le M_N}\). Define \[ \begin{aligned} Z_{N,\mathrm S}^{u,\pi} &:= \int_{S_N} \exp\left\{ \sum_{k=1}^{M_N}u\!\left(\frac{g_k\cdot\sigma}{\sqrt N}\right) +\sum_{i=1}^N h_i\sigma_i \right\} \mu_N^{\mathrm S}(\dd\sigma),\\ F_{N,\mathrm S}^{u,\pi} &:=\frac1N\log Z_{N,\mathrm S}^{u,\pi}. \end{aligned} \] Here \(\E\) and \(\Pp\) denote expectation and probability with respect to the joint law of the constraint vectors \((g_k)_{k\le M_N}\) and external fields \((h_i)_{i\le N}\). \begin{corollary}[Spherical i.i.d.\ external fields] \label{cor:spherical-iid-field} Let $u:\R\to\R$ be real-valued and globally Lipschitz, let $M_N/N\to\alpha\in(0,\infty)$, and let $\pi$ satisfy $\int h^2\pi(\dd h)<\infty$ as above. Then \begin{equation*} F_{N,\mathrm S}^{u,\pi} \longrightarrow\mathcal P_{\mathrm S,u}^{\mathrm{fld}}(\alpha,\sqrt{v_\pi}) \quad\text{in }L^2. \end{equation*} % Thus convergence also holds in probability and in mean. \end{corollary} \begin{proof} For constraint vectors \(\mathbf g=(g_k)_{k=1}^{M_N}\) and a deterministic field \(\mathbf h\in\mathbb R^N\), define \[ \widehat f_N(\mathbf g,\mathbf h) := \frac1N\log\int_{S_N} \exp\left\{ \sum_{k=1}^{M_N} u\left(\frac{g_k\cdot\sigma}{\sqrt N}\right) +\mathbf h\cdot\sigma \right\} \mu_N^{\mathrm S}(\dd\sigma). \] For \(\mathbf h,\widetilde{\mathbf h}\in\mathbb R^N\), Cauchy--Schwarz gives $ |(\mathbf h-\widetilde{\mathbf h})\cdot\sigma| \le \sqrt N\,\|\mathbf h-\widetilde{\mathbf h}\|_2$. Comparing the two exponential weights therefore yields \begin{equation}\label{eq:spherical-field-vector-modulus} \left| \widehat f_N(\mathbf g,\mathbf h) -\widehat f_N(\mathbf g,\widetilde{\mathbf h}) \right| \le \frac{\|\mathbf h-\widetilde{\mathbf h}\|_2}{\sqrt N}. \end{equation} For an orthogonal matrix \(O\in\mathbb R^{N\times N}\), write \(O\mathbf g=(Og_k)_{k=1}^{M_N}\). The change of variables \(\sigma=O^{\mathsf T}\tau\), together with rotational invariance of \(\mu_N^{\mathrm S}\), gives $ \widehat f_N(\mathbf g,\mathbf h) = \widehat f_N(O\mathbf g,O\mathbf h).$ Under the Gaussian row law, \(O\mathbf g\) has the same distribution as \(\mathbf g\). Writing \(e_1=(1,0,\ldots,0)\in\mathbb R^N\) and choosing \(O\) such that \(O\mathbf h=\|\mathbf h\|_2e_1\), we obtain \[ \widehat f_N(\mathbf g,\mathbf h) \overset{\mathrm d}= \widehat f_N(\mathbf g,\|\mathbf h\|_2e_1) \qquad(\mathbf h\in\mathbb R^N), \] where the equality in distribution is with respect to the Gaussian constraint vectors. Let \(z_1,z_2,\ldots\) be independent standard Gaussian random variables, independent of both the external fields and the constraint vectors, and put $ \mathbf h^{(N)}=(h_1,\ldots,h_N)$ and $ z^{(N)}=(z_1,\ldots,z_N)$. In the remainder of the proof, \(\E\) also averages over this auxiliary Gaussian sequence. Define, using the same constraint vectors in both expressions, \[ X_N := \widehat f_N\bigl(\mathbf g,\|\mathbf h^{(N)}\|_2e_1\bigr), \qquad Y_N := \widehat f_N\bigl( \mathbf g,\sqrt{v_\pi}\,\|z^{(N)}\|_2e_1 \bigr). \] Conditioning on \(\mathbf h^{(N)}\) or \(z^{(N)}\) and applying the preceding rotational identity gives \[ \begin{aligned} X_N &\overset{\mathrm d}= \widehat f_N(\mathbf g,\mathbf h^{(N)}) =F_{N,\mathrm S}^{u,\pi},\\ Y_N &\overset{\mathrm d}= \widehat f_N(\mathbf g,\sqrt{v_\pi}\,z^{(N)}) \overset{\mathrm d}= F_{N,\mathrm S}^{u,\sqrt{v_\pi}}, \end{aligned} \] where \(F_{N,\mathrm S}^{u,\sqrt{v_\pi}}\) is the Gaussian-field free energy from \eqref{eq:intro-gaussian-field-partition} with \(b=\sqrt{v_\pi}\). By \eqref{eq:spherical-field-vector-modulus}, \[ |X_N-Y_N| \le \left| \frac{\|\mathbf h^{(N)}\|_2}{\sqrt N} -\sqrt{v_\pi}\frac{\|z^{(N)}\|_2}{\sqrt N} \right|. \] Using \((\sqrt x-\sqrt y)^2\le |x-y|\) for \(x,y\ge0\), we conclude that \[ \begin{aligned} \E|X_N-Y_N|^2 &\le \E\left| \frac1N\sum_{i=1}^N h_i^2 -\frac{v_\pi}{N}\sum_{i=1}^N z_i^2 \right|\\ &\le \E\left| \frac1N\sum_{i=1}^N h_i^2-v_\pi \right| +v_\pi\,\E\left| \frac1N\sum_{i=1}^N z_i^2-1 \right| \longrightarrow0. \end{aligned} \] Indeed, both terms tend to zero by the law of large numbers, since \(\E h_1^2=v_\pi<\infty\) and \(\E z_1^2=1\). Set \[ P_\pi := \mathcal P_{\mathrm S,u}^{\mathrm{fld}} (\alpha,\sqrt{v_\pi}). \] The Gaussian-field extension of Theorem~\ref{thm:iid-channels} and the distributional identity for \(Y_N\) imply that \(\E|Y_N-P_\pi|^2\to0\). Consequently, \[ \E\left|F_{N,\mathrm S}^{u,\pi}-P_\pi\right|^2 = \E|X_N-P_\pi|^2 \le 2\E|X_N-Y_N|^2+2\E|Y_N-P_\pi|^2 \longrightarrow0, \] which proves the asserted \(L^2\) convergence. \end{proof} % \begin{proof} % For fixed rows \(\mathbf g=(g_k)_{k \le M_N}\) and deterministic % \(\mathbf h \in\mathbb R^N\), set % \[ % \widehat f_N(\mathbf g,\mathbf h) % =\frac1N\log\int_{S_N}\exp\left\{ % \sum_{k \le M_N}u\!\left(\frac{g_k\cdot\sigma}{\sqrt N}\right) % +\mathbf h\cdot\sigma\right\}\,\mu_N^{\rm S}(\dd\sigma). % \] % Direct comparison on the sphere gives the estimate % \begin{equation}\label{eq:spherical-field-vector-modulus} % \left|\widehat f_N(\mathbf g,\mathbf h)- % \widehat f_N(\mathbf g,\widetilde{\mathbf h})\right| % \le \frac{\|\mathbf h-\widetilde{\mathbf h}\|_2}{\sqrt N}. % \end{equation} % Rotate \(\mathbf h\) to \(\|\mathbf h\|_2e_1\) and rotate the rows by the same % orthogonal map. Coupling \(\mathbf h^{(N)}=(h_i)_{i\le N}\) with an independent % standard Gaussian vector \(z^{(N)}\), the i.i.d.-field and Gaussian-field % pressures can therefore be coupled within error % \[ % \left|\frac{\|\mathbf h^{(N)}\|_2}{\sqrt N} % -\sqrt{v_\pi}\frac{\|z^{(N)}\|_2}{\sqrt N}\right|. % \] % This tends to zero in \(L^2\) by the \(L^1\) law of large numbers for % \(h_1^2\) and \((\sqrt x-\sqrt y)^2\le|x-y|\). Apply % the \hyperref[par:gaussian-field-extension]{Gaussian-field extension} with \(b=\sqrt{v_\pi}\). % \end{proof} \paragraph{Bipartite transposition.} Summing over one party converts a bipartite Ising model with a homogeneous deterministic field on only that party into a perceptron. For $\beta,b\in\R$, define \begin{equation*} Z_{N,M}^{\mathrm{bip}}(\beta,b) =2^{-(N+M)} \sum_{\substack{\sigma\in\{-1,+1\}^N\\ \tau\in\{-1,+1\}^M}} \exp\left\{ \frac\beta{\sqrt N}\sum_{k=1}^M\sum_{i=1}^N g_{ki}\tau_k\sigma_i +b\sum_{i=1}^N\sigma_i \right\}, \end{equation*} and put \begin{equation*} u_{\alpha,\beta,b}^{\rm bip}(x) =\log\cosh\bigl(b+\beta\sqrt\alpha\,x\bigr). \end{equation*} \begin{corollary}[Bipartite Ising formula with a field on one party] \label{cor:bipartite-one-sided-field} For every $\beta,b\in\R$, if $M_N/N\to\alpha\in(0,\infty)$, then \begin{equation*} \frac1N\log Z_{N,M_N}^{\mathrm{bip}}(\beta,b) \longrightarrow \alpha\,\mathcal P_{\mathrm I,u_{\alpha,\beta,b}^{\rm bip}}(1/\alpha) \end{equation*} in $L^2$ and hence in $L^1$ and probability. Moreover, \begin{equation}\label{eq:bipartite-one-sided-field-variance} \Var\left(\frac1N\log Z_{N,M_N}^{\mathrm{bip}}(\beta,b)\right) \le\frac{\beta^2M_N}{N^2}. \end{equation} \end{corollary} \begin{proof} Put \(\alpha_N=M_N/N\) and \(u_N^{\rm bip}(x)=\log\cosh(b+\beta\sqrt{\alpha_N}\,x)\). Summing first over \(\sigma\) gives the exact identity \begin{equation*} Z_{N,M_N}^{\mathrm{bip}}(\beta,b) =Z_{M_N,N}^{\mathrm I,u_N^{\rm bip}}. \end{equation*} Indeed, the columns of the \(M_N\)-by-\(N\) disorder matrix are independent standard Gaussian rows in \(\mathbb R^{M_N}\). Write \(G=(g_{ki})_{k \le M_N,\,i\le N}\) for this matrix. The one-Lipschitz property of \(\log\cosh\), Cauchy--Schwarz, and the standard Gaussian operator-norm estimate give \begin{equation*} \left|\frac1{M_N}\log Z_{M_N,N}^{\mathrm I,u_N^{\rm bip}} -\frac1{M_N}\log Z_{M_N,N}^{\mathrm I,u_{\alpha,\beta,b}^{\rm bip}}\right| \le |\beta|\,|\sqrt{\alpha_N}-\sqrt\alpha|\, \frac{\sqrt N\,\|G\|_{\rm op}}{M_N} =o_{L^2}(1). \end{equation*} Apply \cref{thm:iid-channels} with spin dimension \(M_N\), row count \(N\), and limiting density \(1/\alpha\), and multiply by \(M_N/N\to\alpha\). Finally, \(\partial_{g_{ki}}[N^{-1}\log Z^{\mathrm{bip}}] =\beta(N\sqrt N)^{-1}\langle\tau_k\sigma_i\rangle\), so Gaussian Poincar\'e gives \eqref{eq:bipartite-one-sided-field-variance}. \end{proof} \paragraph{Gaussian spins.} The spherical field formula gives the Euclidean Shcherbina--Tirozzi model \cite{ShcherbinaTirozziGardner}. Let \(g^0,g_1,g_2,\ldots\) be independent standard Gaussian vectors in \(\mathbb R^N\), and define, for \(c>0\), \begin{equation*} \mathcal Z_{N,M}^{\rm ST}(u,h,c) =\int_{\mathbb R^N}\exp\left\{ \sum_{k=1}^M u\left(\frac{g_k\cdot x}{\sqrt N}\right) +h\,g^0\cdot x-c\|x\|_2^2 \right\}\,\dd x. \end{equation*} For \(\rho>0\), put \begin{equation*} u_\rho(z)=u(\sqrt\rho\,z) \end{equation*} and \begin{align} \mathcal P_{\rm ST}(\alpha;u,h,c) :=\frac12\log(2\pi e) +\sup_{\rho>0}\Bigg\{ \frac12\log\rho-c\rho+\inf_{q\in\Q}\left[ \mathcal A_{\rm S}(q)+\alpha\Phi_{u_\rho}(q) +\frac{h^2\rho}{2}\mathsf L(q) \right]\Bigg\}. \label{eq:st-variational-value} \end{align} \begin{theorem}[Gaussian-spin free energy] \label{thm:st-free energy} Let \(u:\mathbb R\to\mathbb R\) be real-valued and globally Lipschitz, let \(c>0\), \(h\in\mathbb R\), and \(M_N/N\to\alpha\in(0,\infty)\). Then \begin{equation*} \frac1N\log\mathcal Z_{N,M_N}^{\rm ST}(u,h,c) \longrightarrow \mathcal P_{\rm ST}(\alpha;u,h,c) \qquad\text{in }L^1. \end{equation*} % In particular, convergence holds in probability and the expectations % converge to the same limit; all expectations are over the joint row and % external-field disorder. \end{theorem} In particular, under the hypotheses of Talagrand's Research Problem~3.1.2, \cref{thm:st-free energy} computes the limit \cite[Research Problem~3.1.2, p.~193]{TalagrandBookI}. It does not resolve the separate Research Problem~3.2.13 on uniqueness of the scalar replica-symmetric equations \cite[Research Problem~3.2.13, p.~222]{TalagrandBookI}; the optimizer in \eqref{eq:st-variational-value} is not assumed to be a constant path. % \begin{proof} % Polar disintegration with \(x=\sqrt\rho\,\sigma\) gives % \begin{equation}\label{eq:st-polar-disintegration} % \mathcal Z_{N,M}^{\rm ST}(u,h,c) % =C_N\int_0^\infty % \rho^{N/2-1}e^{-cN\rho} % Z_{N,M}^{\mathrm S,u_\rho,h\sqrt\rho}\,\dd\rho, % \qquad C_N=\frac{(\pi N)^{N/2}}{\Gamma(N/2)}. % \end{equation} % Stirling's formula yields % \begin{equation}\label{eq:st-polar-constant} % \frac1N\log C_N\longrightarrow\frac12\log(2\pi e). % \end{equation} % After the change of variables \(\rho=s^2\), % \cref{lem:st-radial-laplace} gives \(L^1\) convergence of the % normalized log radial integral to % \(\sup_{s>0}\{\mathcal P_{\mathrm S,u_{s^2}}^{\mathrm{fld}}(\alpha,hs) % +\log s-cs^2\}\). % Together with \eqref{eq:st-polar-constant}, this is % \eqref{eq:st-variational-value}. The lemma proves the compact-uniform % shell convergence, radial localization, and uniform integrability % needed for this passage. % \end{proof} \begin{proof} Polar disintegration with \(x=\sqrt\rho\,\sigma\), \(\sigma\in S_N\), gives \begin{equation}\label{eq:st-polar-disintegration} \mathcal Z_{N,M_N}^{\rm ST}(u,h,c) =C_N\int_0^\infty \rho^{N/2-1}e^{-cN\rho} Z_{N,M_N}^{\mathrm S,u_\rho,h\sqrt\rho}\,\dd\rho, \qquad C_N=\frac{(\pi N)^{N/2}}{\Gamma(N/2)}. \end{equation} Stirling's formula yields \begin{equation}\label{eq:st-polar-constant} \frac1N\log C_N \longrightarrow\frac12\log(2\pi e). \end{equation} For \(s>0\), define, using the same disorder for every \(s\), \[ Y_N(s):=\frac1N\log Z_{N,M_N}^{\mathrm S,u_{s^2},hs}, \qquad Y(s):=\mathcal P_{\mathrm S,u_{s^2}}^{\mathrm{fld}}(\alpha,hs). \] Since \(\rho=s^2\) gives \(\rho^{N/2-1}\,\dd\rho=2s^{N-1}\,\dd s\), we obtain \[ \frac1N\log\mathcal Z_{N,M_N}^{\rm ST}(u,h,c) = \frac1N\log(2C_N) +\frac1N\log\int_0^\infty s^{N-1}e^{N[Y_N(s)-cs^2]}\,\dd s. \] For these functions \(Y_N\) and \(Y\), \cref{lem:st-radial-laplace} gives \[ \frac1N\log\int_0^\infty s^{N-1}e^{N[Y_N(s)-cs^2]}\,\dd s \xrightarrow{L^1} \sup_{s>0}\{Y(s)+\log s-cs^2\}. \] Adding \eqref{eq:st-polar-constant}, with \(N^{-1}\log2\to0\), proves \(L^1\) convergence of the full free energy to \[ \frac12\log(2\pi e) +\sup_{\rho>0} \left\{ \frac12\log\rho-c\rho +\mathcal P_{\mathrm S,u_\rho}^{\mathrm{fld}} (\alpha,h\sqrt\rho) \right\}, \] where we have returned to \(\rho=s^2\). Substituting the definition of \(\mathcal P_{\mathrm S,u_\rho}^{\mathrm{fld}}\) from \eqref{eq:intro-field-value} identifies this expression with \(\mathcal P_{\rm ST}(\alpha;u,h,c)\) in \eqref{eq:st-variational-value}. \end{proof} \subsection{Quenched large deviations} \label{sec:quenched-ldp-statement} The free energy formula also determines nonlinear weights depending on the empirical distribution of the constraint fields. With the Gaussian rows \((g_k)_k\) held fixed, define \begin{equation*} L_{N,\sigma}=\frac1{M_N}\sum_{k\le M_N}\delta_{h_k(\sigma)}, \qquad \rho_N^\kappa=\mu_N^\kappa\circ L_{N,\cdot}^{-1}. \end{equation*} On the space \(\mathcal P(\mathbb R)\) of probability measures, equipped with its weak topology, put \begin{equation}\label{eq:quenched-ldp-rate} \mathcal J_{\kappa,\alpha}(\mu) =\sup_{u\in\mathrm{BL}(\mathbb R)} \left\{\int u\,\dd\mu-\frac{\mathcal P_{\kappa,u}(\alpha)}\alpha\right\}, \end{equation} where \(\mathrm{BL}(\mathbb R)\) denotes bounded Lipschitz functions. \begin{theorem}[Quenched empirical measure large deviations] \label{thm:quenched-empirical-ldp} For either prior, on one event of probability one, \(\rho_N^\kappa\) satisfies a large-deviation principle with speed \(M_N\) and good convex rate function \(\mathcal J_{\kappa,\alpha}\). On the same event, for every bounded weakly continuous \(\Psi:\mathcal P(\mathbb R)\to\mathbb R\), \begin{equation}\label{eq:nonlinear-perceptron-free energy} \frac1N\log\int e^{M_N\Psi(L_{N,\sigma})}\mu_N^\kappa(\dd\sigma) \longrightarrow \alpha\sup_\mu\{\Psi(\mu)-\mathcal J_{\kappa,\alpha}(\mu)\}. \end{equation} \end{theorem} Note that no convexity of \(\Psi\) is assumed in Theorem~\ref{thm:quenched-empirical-ldp}. For unnormalized Ising counting, add \(\log2\) to the limit in Eq.~\eqref{eq:nonlinear-perceptron-free energy}. At \(M_N=N\), this proves the quenched Laplace principle in \cite[Definition~2.1 and Conjecture~2.2]{BolthausenKistler} for the Gaussian Ising perceptron. The result concerns constraint fields under the spin prior; it does not assert replica symmetry for a conditioned spin measure. The proof in \cref{sec:quenched-ldp-proof} uses differentiability of the pressure in bounded channel directions. \subsection{Scalar projection pursuit and supervised learning} \label{sec:mz-statement} We next consider the problem of maximizing an empirical objective over a single projection direction. Let \((X_i)_i\) be independent standard Gaussian vectors in \(\R^d\), where \(n/d\to\alpha\in(0,\infty)\). Fix a teacher dimension \(k\ge1\), deterministic matrices \(W_{*,d}\in\R^{d\times k}\) with orthonormal columns, and labels \[ Y_i=\varphi(W_{*,d}^{\mathsf T}X_i,\varepsilon_i), \] where \(\varphi\) is measurable and real-valued and the \(\varepsilon_i\) are i.i.d.\ with a fixed law, independently of the covariates. The objective is \(n^{-1}\sum_i h(Y_i,w\cdot X_i)\), with \(\|w\|_2=1\) and bounded continuous \(h\). The limiting value uses the following scalar functional. Write \(C_b\) for bounded continuous functions and let \(\mathscr U\) be the nonnegative, nondecreasing integrable functions on \([0,1)\). For \(H\in C_b(\R)\) and \(c>0\), set \[ \mathsf H_cH(x)=\sup_z\{H(x+z)-z^2/(2c)\}. \] Let \(f_{\mu,c}^{H}\) solve \eqref{eq:intro-parisi-pde} on \([0,1]\) with coefficient \(\mu\) and terminal function \(h=\mathsf H_cH\). The function $h$ is bounded and Lipschitz, so the definition applies to every \(\mu\in\mathscr U\). \begin{theorem}[Scalar Montanari--Zhou formula] \label{thm:montanari-zhou} Under the preceding assumptions, for every fixed \(h\in C_b(\R^2)\), \begin{equation*} \max_{\|w\|_2=1}\frac1n\sum_{i=1}^n h(Y_i,w\cdot X_i) \longrightarrow \sup_{r\in B_k(1)}\inf_{\mu\in\mathscr U,\ c>0} \mathsf F_{\alpha,h}(\mu,c,r) \end{equation*} almost surely and in every \(L^p\) with $p<\infty$, where \(B_k(1)\) is the closed unit ball and \begin{equation}\label{eq:mz-supervised-functional} \mathsf F_{\alpha,h}(\mu,c,r) =\E\left[f_{\mu,c}^{h(Y,\cdot)}(|r|^2,r\cdot G)\right] +\frac1{2\alpha}\int_{|r|^2}^1 \frac{\dd t}{c+\int_t^1\mu(s)\,\dd s}. \end{equation} Here \(G\sim N(0,I_k)\) is independent of \(\varepsilon\), and \(Y=\varphi(G,\varepsilon)\). At \(|r|=1\), the inner infimum equals \(\E h(Y,r\cdot G)\). Moreover, for every fixed \(H\in C_b(\R)\), \begin{equation*} \max_{\|w\|_2=1}\frac1n\sum_{i=1}^n H(w\cdot X_i) \longrightarrow \inf_{\mu\in\mathscr U,\ c>0}\left\{ f_{\mu,c}^{H}(0,0)+\frac1{2\alpha}\int_0^1 \frac{\dd t}{c+\int_t^1\mu(s)\,\dd s}\right\} \end{equation*} almost surely and in every \(L^p\) with $p<\infty$. \end{theorem} These formulas establish the supervised scalar conjecture \cite[Conjecture~2.2, equations~(14)--(15)]{MZMultiIndex} and the \(m=1\) projection-pursuit conjecture \cite[Conjecture~2.1 and Remark~1]{MZProjection}. The first applies, in particular, under their sub-Gaussian label assumption. The limits also identify their feasible-distribution duals, which equal the corresponding probability limit inferiors \cite[equation~(9)]{MZMultiIndex}\cite[equation~(3)]{MZProjection}. \subsection{Hard constraints and capacities} We use the free-energy formula for Lipschitz potentials to derive formulas for the hard-constraint free energy at margin \(\theta\in\R\). Let \(C\subsetneq\mathbb R\) be a nonempty closed set containing either \([\theta,\infty)\) or \((-\infty,\theta]\). By reflection, we may assume that \begin{equation*} [\theta_C,\infty)\subset C \end{equation*} for some finite \(\theta_C\); we call such a set \emph{tail-accessible}. For \(Z\sim N(0,1)\), put \begin{equation*} d_C(x)=\operatorname{dist}(x,C),\qquad u_{\beta,C}(x)=-\beta d_C(x),\qquad \pi_C=\Pp(Z\in C)\in(0,1). \end{equation*} The feasible set and its normalized entropy are \begin{align*} \mathcal F_{M,N}^{\kappa}(C) &=\{\sigma\in\Sigma_N^\kappa:h_k(\sigma)\in C\text{ for all } k\le M\}, \\ \operatorname{Ent}_{N,\kappa}(C) &=\frac1N\log\mu_N^\kappa(\mathcal F_{M_N,N}^{\kappa}(C)), \qquad\log0=-\infty. \end{align*} For Ising spins, the unnormalized counting entropy is \(\log2+\operatorname{Ent}_{N,\mathrm I}(C)\). Define the hard channel and variational value by \begin{align*} \mathcal P_{\kappa,\infty,C}(\alpha) =\inf_{\substack{q\in\Q:\ \mathcal A_\kappa(q)<\infty}} \{\mathcal A_\kappa(q)+\alpha\Phi_{\infty,C}(q)\},\qquad \Phi_{\infty,C}(q)=\lim_{\beta\to\infty}\Phi_{u_{\beta,C}}(q), \end{align*} both taking values in \([-\infty,0]\), with \(\mathcal P_{\kappa,\infty,C}(0)=0\). \begin{theorem}[Closed-set hardening] \label{thm:closed-set-hardening} Let \(C\) be tail-accessible and \(M_N/N\to\alpha>0\). For either prior, \begin{equation}\label{eq:hard-variational-convergence} \mathcal P_{\kappa,u_{\beta,C}}(\alpha) \downarrow\mathcal P_{\kappa,\infty,C}(\alpha) \qquad(\beta\to\infty). \end{equation} For spherical spins, \begin{equation}\label{eq:hard-spherical-volume-limit} \operatorname{Ent}_{N,\mathrm S}(C) \xrightarrow{\Pp}\mathcal P_{\mathrm S,\infty,C}(\alpha) \end{equation} in the extended real line. For Ising spins, if \(\mathcal P_{\mathrm I,\infty,C}(\alpha)>-\log2\), then \begin{align*} \operatorname{Ent}_{N,\mathrm I}(C) &\xrightarrow{\Pp}\mathcal P_{\mathrm I,\infty,C}(\alpha), \\ \frac1N\log\#\mathcal F_{M_N,N}^{\mathrm I}(C) &\xrightarrow{\Pp}\log2+\mathcal P_{\mathrm I,\infty,C}(\alpha)>0. \end{align*} If \(\mathcal P_{\mathrm I,\infty,C}(\alpha)<-\log2\), then \begin{equation}\label{eq:hard-ising-exponential-unsat} \limsup_{N\to\infty}\frac1N\log\Pp \bigl(\mathcal F_{M_N,N}^{\mathrm I}(C)\ne\varnothing\bigr)<0. \end{equation} \end{theorem} Define the candidate Ising capacity by the variational ratio \begin{equation*} \alpha_c^{\mathrm I}(C) =\inf_{\substack{q\in\Q:\ \mathcal A_{\mathrm I}(q)<\infty}} \frac{\log2+\mathcal A_{\mathrm I}(q)}{-\Phi_{\infty,C}(q)}, \end{equation*} where the ratio is zero if \(\Phi_{\infty,C}(q)=-\infty\). The denominator is positive by \eqref{eq:ising-hard-annealed-bound}. Under a common infinite sequence of rows, define for either prior \begin{equation*} \mathfrak M_N^\kappa(C) =\sup\{M\in\mathbb N_0:\mathcal F_{M,N}^{\kappa}(C)\ne\varnothing\}. \end{equation*} \begin{theorem}[Ising capacity] \label{thm:closed-set-ising-capacity} For every tail-accessible \(C\), the critical constraint density satisfies \begin{equation}\label{eq:closed-set-capacity-upper-bound} 0<\alpha_c^{\mathrm I}(C) \le\frac{\log2}{-\log\pi_C}<\infty. \end{equation} It is characterized by positivity of the variational counting entropy: \begin{equation}\label{eq:closed-set-ising-capacity-boundary} \alpha_c^{\mathrm I}(C) = \sup\left\{ \alpha\ge0: \log2+\mathcal P_{\mathrm I,\infty,C}(\alpha)>0 \right\}. \end{equation} If \(\alpha\ne\alpha_c^{\mathrm I}(C)\), then the probability that all \(M_N\) constraints admit a common feasible Ising configuration satisfies \begin{equation}\label{eq:closed-set-ising-sat-unsat} \Pp\bigl( \mathcal F_{M_N,N}^{\mathrm I}(C)\ne\varnothing \bigr) \longrightarrow \begin{cases} 1,&0<\alpha<\alpha_c^{\mathrm I}(C),\\ 0,&\alpha>\alpha_c^{\mathrm I}(C). \end{cases} \end{equation} If \(0<\alpha<\alpha_c^{\mathrm I}(C)\), then \[ \frac1N\log\#\mathcal F_{M_N,N}^{\mathrm I}(C) \xrightarrow{\Pp} \log2+\mathcal P_{\mathrm I,\infty,C}(\alpha)>0, \] with the convention \(\log0=-\infty\). If \(\alpha>\alpha_c^{\mathrm I}(C)\), then \[ \limsup_{N\to\infty}\frac1N \log\Pp\bigl( \mathcal F_{M_N,N}^{\mathrm I}(C)\ne\varnothing \bigr)<0. \] Moreover, \(\mathfrak M_N^{\mathrm I}(C)<\infty\) almost surely for each \(N\), and \begin{equation}\label{eq:closed-set-random-capacity-limit} \frac{\mathfrak M_N^{\mathrm I}(C)}N \xrightarrow{\Pp}\alpha_c^{\mathrm I}(C). \end{equation} \end{theorem} % \begin{theorem}[Ising capacity] % \label{thm:closed-set-ising-capacity} % For every tail-accessible \(C\), % \begin{equation}\label{eq:closed-set-capacity-upper-bound} % 0<\alpha_c^{\mathrm I}(C)\le\frac{\log2}{-\log\pi_C}<\infty. % \end{equation} % The variational ratio is the boundary of positive raw entropy: % \begin{equation}\label{eq:closed-set-ising-capacity-boundary} % \alpha_c^{\mathrm I}(C) % =\sup\{\alpha\ge0:\mathcal P_{\mathrm I,\infty,C}(\alpha)>-\log2\}. % \end{equation} % For every \(M_N/N\to\alpha>0\) away from this threshold, % \begin{equation}\label{eq:closed-set-ising-sat-unsat} % \Pp\bigl(\mathcal F_{M_N,N}^{\mathrm I}(C) % \ne\varnothing\bigr) % \longrightarrow\mathbf1_{\{\alpha<\alpha_c^{\mathrm I}(C)\}}. % \end{equation} % Above threshold the decay is exponential as in % \eqref{eq:hard-ising-exponential-unsat}; at every positive subcritical load, % the raw entropy converges as in \eqref{eq:hard-ising-raw-limit}. % Moreover, \(\mathfrak M_N^{\mathrm I}(C)<\infty\) almost surely and % \begin{equation}\label{eq:closed-set-random-capacity-limit} % \frac{\mathfrak M_N^{\mathrm I}(C)}N % \xrightarrow{\Pp}\alpha_c^{\mathrm I}(C). % \end{equation} % No assertion is made at the critical load. % \end{theorem} \paragraph{Nonmonotone constraints.} The \(U\)-function constraint corresponds to \begin{equation*} C_K^{\rm U}=(-\infty,-K]\cup[K,\infty),\qquad K>0. \end{equation*} Its soft channel is \(-\beta(K-|x|)_+\). The two theorems give its spherical hard volume and Ising entropy and capacity, with \(\pi_{C_K^{\rm U}}=2\Pp(Z\ge K)\). For \(K>K_*\simeq0.817\), this gives the full-path variational characterization called for by the wide-regime analysis of \cite{AubinPerkinsZdeborova}; it does not numerically evaluate the capacity or identify an optimizing path. More examples and the limits of the tail-accessibility assumption are discussed in \cref{sec:ising-capacities}. For general \(C\), a spherical volume formula does not imply an emptiness threshold. \paragraph{Spherical half-spaces.} For \(C_\theta=[\theta,\infty)\), geometry turns the hard-volume formula into a satisfiability threshold. Define \begin{equation*} \alpha_c^{\rm S}(\theta) =\sup\{\alpha\ge0: \mathcal P_{\mathrm S,\infty,C_\theta}(\alpha)>-\infty\}. \end{equation*} \begin{theorem}[Spherical half-space capacity] \label{thm:spherical-halfspace-capacity} For every \(\theta\in\R\) and positive \(\alpha\ne\alpha_c^{\rm S}(\theta)\), \begin{equation*} \Pp\bigl(\mathcal F_{\lfloor\alpha N\rfloor,N}^{\rm S}(C_\theta) \ne\varnothing\bigr) \longrightarrow\mathbf1_{\{\alpha<\alpha_c^{\rm S}(\theta)\}}. \end{equation*} For every $\alpha \in (0, \alpha_c^{\mathrm S})$, the hard-volume limit in \eqref{eq:hard-spherical-volume-limit} is finite. Moreover, \(\mathfrak M_N^{\rm S}(C_\theta)<\infty\) almost surely and \begin{equation}\label{eq:random-capacity-limit} \frac{\mathfrak M_N^{\rm S}(C_\theta)}N \xrightarrow{\Pp}\alpha_c^{\rm S}(\theta). \end{equation} No assertion is made when \(M_N/N\to\alpha_c^{\mathrm S}(C_\theta)\). \end{theorem} For negative margins, the entropy formula and the capacity characterization above establish the variational predictions of Franz and Parisi \cite{FranzParisiJamming}. Their further predictions concerning the structure of the minimizing overlap distribution and the jamming critical exponents are not addressed here. \paragraph{Relation to three problems of Talagrand.} At zero Ising margin, \cref{thm:closed-set-hardening,thm:closed-set-ising-capacity} give a positive deterministic capacity, convergence of the normalized random capacity, the full subcritical variational entropy, and exponentially small supercritical satisfiability probability. These provide variational formulas for the threshold and entropy requested in Research Problem~2.1.1 \cite[Research Problem~2.1.1, p.~152]{TalagrandBookI}. The formulas optimize over all overlap paths; we do not identify a constant-path optimizer or evaluate the capacity numerically. The Gaussian-spin formula in \cref{thm:st-free energy} answers Research Problem~3.1.2. For every negative margin, with Talagrand's \(\tau\) equal to our \(\theta\), \cref{thm:closed-set-hardening,thm:spherical-halfspace-capacity} give the full-path hard-volume limit at every density and the SAT--UNSAT transition away from the critical density, answering the free energy and capacity aspects of Research Problem~8.4.3 \cite[Research Problem~8.4.3, p.~25]{TalagrandBookII}. We make no claim at the critical load and do not determine optimizer structure or critical exponents. % \paragraph{Proof strategy and further applications.} % Restoring one and two independent constraints reconstructs the derivative % correlations. A one-coordinate calculation for Ising spins and tangent % integration by parts on the sphere identify the overlap equations. % Inserting a fixed coordinate block gives the lower bound; the supporting % equalities supplied by the overlap equations fix the sign of the % interpolation for the upper bound. The finite-volume interpolation % uses finite-depth RPCs; general paths enter through the representation % of limiting overlap arrays. Sections~2--7 prove these statements % and the free energy formula; Section~8 proves the large-deviation % principle. Section~9 proves the remaining application results and % the radial passage used above, including the zero-temperature and % quadratic-penalty limits, hard constraints, and capacities. \section{Proof of the free energy formula} \label{sec:proof-overview} \label{sec:free energy-proof} In this section, we outline the proof of the Parisi formula for the soft free energy, \cref{thm:iid-channels}, and state the main ingredients. We establish matching lower and upper bounds for the limiting free energy. The lower bound follows the Aizenman--Sims--Starr scheme, using increments obtained by adding finitely many constraints and coordinates. The upper bound uses a Guerra interpolation. For both bounds, we add a Gaussian perturbation, independent of the constraint vectors, that changes the expected normalized log partition function by \(o(1)\). This perturbation enforces the joint Ghirlanda--Guerra identities for subsequential limits of the disorder-averaged replica laws of the spin overlaps and cascade ancestry labels. Together with synchronization, these identities yield the limiting-array representation used in the cavity calculations. Cavity in \(M\) determines the asymptotic effect of adding a fixed number of constraints. % The overlap equation follows from a one-coordinate cavity % calculation for Ising spins and from tangent integration by % parts for spherical spins. Cavity in \(N\) bounds the increment produced by adding a fixed block of coordinates. The overlap equation supplies the supporting equality needed to turn these calculations into the two variational bounds. We state these results below and deduce \cref{thm:iid-channels} from them at the end of this section. Sections~\ref{sec:cavity-M}--\ref{sec:rpc-synchronization} provide the proofs of these results and verify their limiting-array hypotheses. Let us first work with centered smooth channels; this assumption will be removed later. For the cavity arguments, assume \begin{equation}\label{eq:iid-smooth-core-assumption} u_\theta(0)=0,\qquad u_\theta\in C^3(\mathbb R),\qquad L_j:=\sup_{\theta\in\mathbb R}\|u_\theta^{(j)}\|_\infty<\infty \quad(1\le j\le3). \end{equation} Put \(L=L_1\), fix \(\kappa\in\{\mathrm I,\mathrm S\}\), and write \(u_k=u_{\Theta_k}\), where the marks \(\Theta_k\sim\nu\) are i.i.d.\ and independent of the Gaussian vectors and all further randomness introduced below. % A mark is held fixed during normalization and % Gaussian integration, and averaged afterward in \(\E\). We work at zero external field (meaning $b=0$) and abbreviate \(\mathcal A=\mathcal A_\kappa\) and \(\mathcal S=\mathcal S_\kappa\). \subsection{Conditional constraint insertion and reconstruction} \label{sec:overview-M-cavity} We first isolate the effect of adding a fixed number of independent constraints $u(h_k)$ to the Hamiltonian. % We first specify the finite-volume measure and the limiting overlap % array. We then define the scalar quantities that describe the effect % of inserting independent constraints. Throughout, the channels and marks satisfy the centered smooth assumptions stated above, and \(L=\sup_\theta\|u_\theta'\|_\infty\). The key result in this section is proposition~\ref{prop:0916-rows}, which has two conclusions. First, adding a fixed number of independent constraints preserves the limiting joint distribution of the spin overlaps and ancestry labels \((R_{\ell m},r_{\ell m})_{\ell,m\ge1}\), after averaging over the disorder. Second, under the reconstruction hypotheses (stated in the proposition), the proposition identifies the limiting distribution of the score-overlap array $Q$~\eqref{eq:overview-composite-score} jointly with \((R,r)\). \paragraph{The finite-volume measure and its overlap array.} Fix \(d\ge2\). Let \(\mathfrak R_d\) be the RPC with masses \(j/d\), \(1\le j0), \qquad T_{0,s}h(x)=\E h(x+\sqrt s Z),\quad Z\sim N(0,1). \end{equation} Write \(c=c_j\) on \([m_j,m_{j+1})\), where \(0=m_0<\cdots0\): under \(t=Ts\), its diffusion coefficient is \(T\) and its nonlinear coefficient is \(T\gamma(Ts)\). Additive constants give \(\Phi_u=u(0)+\Phi_{u-u(0)}\), while centering leaves normalized laws unchanged. \begin{corollary}[Independent constraint insertions] \label{lem:rpc-constraint-insertion} For a step path \(c\) on \(\pi\), every fixed \(k\ge1\), and every bounded Borel test \(F\) of finitely many replica ancestries, \begin{equation}\label{eq:rpc-insertion-log} \E\langle F\rangle_{\pi,c,+k}=\E_\pi F, \qquad \E\log Z_{\pi,c,k}=k\Phi(c). \end{equation} If these replicas include \(1,2\), then \begin{equation}\label{eq:rpc-insertion-score} \E\left\langle F\prod_{b=1}^k u_b'(X_b^{c,1})u_b'(X_b^{c,2})\right\rangle_{\pi,c,+k} =\E_\pi[F p_c(r^\pi_{12})^k]. \end{equation} For bounded Borel measurable \(\chi_b(\theta,x)\), continuous in \(x\), \begin{equation}\label{eq:rpc-insertion-one-replica} \E\left\langle\prod_{b=1}^k\chi_b(\Theta_b,X_b^{c,1})\right\rangle_{\pi,c,+k} =\prod_{b=1}^k\E\langle\chi_b(\Theta,X^{c,1})\rangle_{\pi,c}. \end{equation} For arbitrary \(c\in\Q\), these identities hold for the limiting replica laws, with \(r^\pi,\E_\pi\) replaced by \(U,\E_U\). The limiting expected log normalizer is \(k\Phi(c)\). The response \(p_c\) is nondecreasing, takes values in \([0,L^2]\), and is constant on plateaus of \(c\). Both \(\Phi(c)\) and \(p_c\) are continuous for \(L^1\) convergence of \(c\). The derivative formula \eqref{eq:0916-channel-derivative} also holds. \end{corollary} \begin{proof}[Finite-step proof] Condition on the fresh marks. Represent each row by independent ancestral Gaussian increments of variances \(c_0,c_1-c_0,\ldots,c_J-c_{J-1}\), integrate its terminal Gaussian, and apply \cref{rpc:prop-quasistationarity} conditional on the root Gaussians. The bound \(|u_b(x)|\le L|x|\) makes the recursion finite and its root log value integrable. Independent rows have factorized exponential moments, so the recursion adds and the tilted transitions factor even for different terminals \cite[Theorem~2.9 and (2.59)--(2.60)]{PanchenkoBook}. The expected log normalizer is therefore \(\sum_b\Phi_{u_b}(c)\). Conditional on any sampled genealogy, the entire tilted mark trees are independent across rows; their one-replica marginals do not depend on the genealogy. Averaging the marks proves \eqref{eq:rpc-insertion-log} and \eqref{eq:rpc-insertion-one-replica}. For one terminal \(u\), start with \(g_0\sim N(0,c_0)\) and tilt each transition \(N(g_j,c_{j+1}-c_j)\) by \(e^{m_{j+1}(X_{j+1}(g_{j+1})-X_j(g_j))}\), including the terminal step. Differentiating the recursion shows that \(M_j=X_j'(g_j)\) is a bounded martingale. Replicas branching after level \(j\) have conditionally independent continuations and terminal scores of conditional mean \(M_j\). Hence \begin{equation*} p_{c,u}(r)=\E M_j^2\quad\text{on }[m_j,m_{j+1}). \end{equation*} Other branches integrate to one, so this identity holds conditional on the full genealogy. Independent marks give \[ \E_{\Theta_1,\ldots,\Theta_k}\prod_{b=1}^k p_{c,u_{\Theta_b}}(r) =p_c(r)^k, \] proving \eqref{eq:rpc-insertion-score}. Conditional Jensen gives monotonicity and \(0\le p_c(r)\le p_{\rm diag}(c)\le L^2\); the inequality passes to general paths by their \(L^1\) limits. Zero-variance transitions give the plateau property. Gaussian covariance differentiation, with completed diagonal variance fixed at one, leaves only the two-replica score term and gives \eqref{eq:0916-channel-derivative}. The recursion and its derivatives are measurable in the marks, and the uniform derivative bounds justify mark averaging. Integrating the derivative gives \eqref{rpc:Phi-L1}. \end{proof} To pass through the random normalizer, we need positive and inverse weight moments. Under any base \(\mu\) independent of fresh row noises, the completed inputs at one sampled state are independent standard normals. Thus, for \(W=e^{\sum_{b\le k}u_b(X_b)}\) and every \(p>0\), \begin{equation}\label{eq:constraint-weight-moments} \E\mu(W^p+W^{-p})\le2(\E e^{pL|Z|})^k<\infty. \end{equation} For one fresh row under an independent probability measure \(\Gamma\), Jensen also gives, with \(A=\log\Gamma(e^{L|X|})\ge0\), \begin{equation}\label{eq:fresh-marked-row-bound} |\log\Gamma(e^{u_\Theta(X)})|\le A,\qquad \E A^p\le\E e^{pL|Z|}\quad(p\ge1). \end{equation} These bounds also control mark fluctuations and row-count changes. \begin{lemma}[Convergence after changing density] \label{lem:change-density-limit} Let \(\mu_n\) be random probability measures, \(W_n>0\) measurable weights, and \(A_n(s,s')\) maps into a fixed Polish space; single-sample marks may be included. For conditional independent samples, put \begin{equation*} Z_n=\mu_n(W_n),\qquad \mu_n^+(\dd s)=Z_n^{-1}W_n(s)\mu_n(\dd s). \end{equation*} Suppose that every finite sampled array, including its diagonal, converges: \begin{equation*} \left((A_n(s^\ell,s^m))_{\ell,m\le k}, (W_n(s^\ell))_{\ell\le k}\right) \Longrightarrow ((A_{\ell m})_{\ell,m\le k},(W^\ell)_{\ell\le k}), \qquad W^\ell>0, \end{equation*} and, for every \(p\ge1\), \begin{equation*} \sup_n\E\mu_n(W_n^p+W_n^{-p})<\infty. \end{equation*} Then the sampled arrays under \(\mu_n^+\) converge to a unique consistent replica law determined by the array on the right. There is also a positive random variable \(Z\), determined by that array, such that \begin{equation*} Z_n\Longrightarrow Z,\qquad \E\log Z_n\longrightarrow\E\log Z. \end{equation*} No random-measure representation of the limiting array is required. If one is given, the limits agree with its ordinary change of density. \end{lemma} \begin{proof} Fix a bounded continuous test \(F\) of a \(k\)-replica array, with \(|F|\le1\), and take \(P\) conditional independent samples \(s^1,\ldots,s^P\) from \(\mu_n\). Set \[ \begin{split} H_n(s^1,\ldots,s^k) &=F\bigl((A_n(s^\ell,s^m))_{\ell,m\le k}\bigr) \prod_{\ell=1}^kW_n(s^\ell),\\ B_n&=\mu_n^{\otimes k}(H_n),\qquad B_{n,P}=P^{-k}\sum_{i_1,\ldots,i_k\le P} H_n(s^{i_1},\ldots,s^{i_k}),\\ Z_{n,P}&=P^{-1}\sum_{i\le P}W_n(s^i),\qquad M_p=1+\sup_n\E\mu_n(W_n^p+W_n^{-p}). \end{split} \] Single-sample marks are included in the argument of \(F\) when present. Repeated sample indices are allowed in \(B_{n,P}\); this is why the array-convergence hypothesis includes the diagonal. Conditional independence gives \(\E|Z_{n,P}-Z_n|^2\le M_2/P\). To estimate the numerator, expand \(\E|B_{n,P}-B_n|^2\) as a sum over two ordered \(k\)-tuples of sample indices. If all \(2k\) indices are different, the corresponding centered product has conditional expectation zero. There are at most \(\binom{2k}{2}P^{2k-1}\) remaining pairs of tuples. For a tuple with distinct indices of multiplicities \(m_1,\ldots,m_j\), where \(\sum_i m_i=k\), H\"older with exponents \(k/m_i\) and \(k/(k-m_i)\) (or equality when \(m_i=k\)) gives \[ \E\bigl[|H_n(s^{i_1},\ldots,s^{i_k})|^2\mid\mu_n,W_n,A_n\bigr] \le\prod_{i=1}^j\mu_n(W_n^{2m_i}) \le\mu_n(W_n^{2k}). \] Also \(|B_n|^2\le\mu_n(W_n^{2k})\). Cauchy--Schwarz therefore bounds each remaining centered product in absolute expectation by \(4M_{2k}\), and hence \[ \E|B_{n,P}-B_n|^2 \le4\binom{2k}{2}M_{2k}/P. \] Jensen bounds all positive and inverse moments of \(Z_n,Z_{n,P}\). Since \(|B_n|\le Z_n^k\) and \(|B_{n,P}|\le Z_{n,P}^k\), expanding the difference using the larger denominator gives \[ \left|\frac{B_{n,P}}{Z_{n,P}^k}-\frac{B_n}{Z_n^k}\right| \le\frac{|B_{n,P}-B_n|}{\max(Z_n,Z_{n,P})^k} +k\frac{|Z_{n,P}-Z_n|}{\max(Z_n,Z_{n,P})}. \] Cauchy--Schwarz, with exponents \(2,2\), now yields \[ \sup_n\E\left|B_{n,P}/Z_{n,P}^k-B_n/Z_n^k\right| \le C_kP^{-1/2}, \] uniformly over \(|F|\le1\). Place the consistent limiting input arrays on one probability space. Given their first \(P\) samples, choose indices \(I_1,\ldots,I_k\) independently with probabilities \(W^i/\sum_{j\le P}W^j\), and let \(\Lambda_{P,k}\) be the averaged law of \((A_{I_\ell I_m})_{\ell,m\le k}\), together with any sampled marks. For fixed \(P\), its expectation of \(F\) is the limit of \(\E B_{n,P}/Z_{n,P}^k\): the normalized empirical integral is a bounded continuous function of the finite array and its positive weights. The preceding estimate shows, for all bounded continuous \(|F|\le1\), \[ |\Lambda_{P,k}(F)-\Lambda_{Q,k}(F)| \le C_k(P^{-1/2}+Q^{-1/2}). \] For probability measures on a Polish space, the supremum over these tests equals the supremum over bounded measurable tests. Thus \(\Lambda_{P,k}\) is Cauchy in total variation and has a probability limit \(\Lambda_k\). Consistency in \(k\) holds for each \(P\) and passes to the limit. Comparing the exact tilted law with its empirical law, then taking \(n\to\infty\) and \(P\to\infty\), proves convergence to \(\Lambda_k\) and its uniqueness. On the limiting input array put \(Z_P=P^{-1}\sum_{i\le P}W^i\). Finite-array convergence and the Portmanteau theorem transfer the bound \[ \E|Z_{n,P}-Z_{n,Q}|^2\le2M_2(P^{-1}+Q^{-1}) \] to \(Z_P,Z_Q\). Consequently \(Z_P\to Z\) in \(L^2\). Jensen and Fatou give \(\E(Z^p+Z^{-p})\le M_p\), so \(Z>0\). For any bounded Lipschitz function, comparison successively through \(Z_{n,P}\) and \(Z_P\), first sending \(n\to\infty\) and then \(P\to\infty\), proves \(Z_n\Rightarrow Z\). Finally \(|\log z|^2\le C(z+z^{-1})\) makes \(\log Z_n\) uniformly integrable; continuous mapping gives convergence of its expectation. If a random-measure representation of the limiting array is supplied, the same empirical estimates identify these limits with its ordinary change of density and normalizer. \end{proof} \begin{proof}[General paths in \cref{lem:rpc-constraint-insertion}] \emph{Covariance convergence.} Choose step paths \(c_n\to c\) in \(L^1\) on grids \(\pi_n\) of mesh \(h_n\to0\), adding zero-variance levels if necessary. Couple their finite genealogy arrays by \(r^n_{\ell m}=\kappa_{\pi_n}(U_{\ell m})\) off the diagonal, where \(\kappa_\pi\) rounds down to the left grid endpoint. This exact finite-RPC law follows from monotone rounding \cite[(2.118)--(2.122)]{PanchenkoBook} and the GG characterization \cite[Theorems~2.10 and~2.13]{PanchenkoBook}. Since \(U_{12}\) is uniform and \(c_n\) is constant on grid cells, \[ \E|c_n(r^n_{12})-c(U_{12})|=\|c_n-c\|_1. \] Every completed Gaussian covariance matrix therefore converges in probability, with diagonal fixed at one, regardless of \(c_n(1-)\). When true diagonals differ from same-leaf pair marks, adjoin an independent uniform label to each sampled state and assign diagonal values only to identical complete states. Distinct replicas then have distinct complete states almost surely; all weights and Gaussian fields ignore the extra labels. \emph{Normalized limits.} Condition on the fresh marks and couple finite Gaussian vectors by their matrix square roots, using independent copies across rows. This gives joint convergence of genealogy, completed inputs, and weights. The bound \eqref{eq:constraint-weight-moments} and \cref{lem:change-density-limit} define their unique normalized replica limits and expected log normalizers. They depend only on \(c\), not its approximation. Finite-RPC invariance passes to continuous tests and then to Borel tests by equality of finite measures, so the limiting genealogy still has law \(U\). Uniform moments permit mark averaging. \emph{Response continuity.} For continuous \(f\), define \[ \nu_c(f)=\E\langle f(U_{12})u_\Theta'(X^{c,1})u_\Theta'(X^{c,2})\rangle_c. \] Helly selection and \(0\le p_{c_n}\le L^2\) give an \(L^1\)-convergent further subsequence of every subsequence. The finite-step score identity identifies every such limit \(p\) by \[ \int_0^1 f(u)p(u)\,\dd u=\nu_c(f), \] since replacing \(f(r^n_{12})\) by \(f(U_{12})\) costs at most \(L^2\omega_f(h_n)\). Continuous tests determine finite measures, so \(p\) is unique almost everywhere. This defines \(p_c\) and proves full-sequence \(L^1\) convergence. Boundedness passes \(p_{c_n}^k\) to \(p_c^k\), giving all three insertion identities. The logarithmic bound \eqref{rpc:Phi-L1} passes to the limit as well. For an arbitrary sequence \(c_n\to c\), choose successively finer step approximations to each \(c_n\). Their completed arrays converge to the same array with covariance \(c(U)\); the uniform empirical bounds in \cref{lem:change-density-limit} give \(\nu_{c_n}(f)\to\nu_c(f)\). The same Helly argument proves \(p_{c_n}\to p_c\) in \(L^1\). To preserve a plateau of \(c\), use \(2^{-n}\lfloor2^nc\rfloor\) off the diagonal and refine its grid. These step approximations show that \(p_c\) is constant there. One-replica convergence gives \(p_{\rm diag}(c_n)\to p_{\rm diag}(c)\). Passing the finite-grid inequality \(p_{c_n}\le p_{\rm diag}(c_n)\) almost everywhere, then taking the essential supremum, gives \begin{equation*} p_{\rm diag}(c)\ge p_c(1-). \end{equation*} Convergence of \(p_{c_n}(1-)\) is neither needed nor asserted. \end{proof} \subsection{Array limits and reconstruction} \label{sec:constraint-proof} We transfer the insertion identities to the finite-volume measures of Section~2. Their limiting observed array is \((q(U_{\ell m}),\rho_d(U_{\ell m}))\) off the diagonal, with diagonal \((1,1)\). \begin{proof}[Proof of \cref{prop:0916-rows}] \emph{Insertion.} For every fixed \(k\), apply \cref{lem:change-density-limit} with \[ \mu_N=G_N\otimes\mathsf N_k,\qquad W_N=\exp\Bigl\{\sum_{b=1}^ku_b(X_b)\Bigr\}. \] Conditional on the old environment, sampled old states, and fresh marks, the completed Gaussian input vectors are centered and independent across rows, with off-diagonal covariance \(\mathsf C_N\) and true diagonal \(1+\varepsilon_N\to1\), including \(\varepsilon_N=0\). Thus \eqref{eq:overview-fresh-covariance-limit} gives joint convergence of the old array, completed inputs, and weights. The bound \eqref{eq:constraint-weight-moments}, with its Gaussian variance replaced by the bounded \(1+\varepsilon_N\), supplies all positive and inverse weight moments, uniformly in the marks. The lemma gives convergence of the normalized laws and expected logarithms; these conclusions pass through the mark average. Explicitly, \begin{align*} \E\log Z_{N,k}^+&\longrightarrow k\Phi(c), \\ \E\langle F\Psi\rangle_{+k} &\longrightarrow\E\langle F_\infty\Psi\rangle_{c,+k}, \end{align*} for bounded continuous old-array tests \(F\) and bounded \(\Psi\) measurable in the fresh marks and continuous in the inserted inputs. Here \(F_\infty\) evaluates \(F\) on the observed limiting array. Corollary~\ref{lem:rpc-constraint-insertion} preserves that array's law and, with \(\Psi=\prod_bu_b'(X_b^1)u_b'(X_b^2)\), gives \eqref{eq:overview-insertion-mixed-moments}. Restoring deleted rows gives the full measure exactly and preserves the limiting observed array. Its spin marginal, and hence its nondecreasing quantile \(q\), agree across the full and deleted systems. The common limiting rule \(c=K(q,\rho_d)\) identifies \(c\); a uniform error \(\|K_N-K\|_\infty\to0\) does not change the covariance limit. The same argument applies to any fixed collection of row counts linked by finitely many insertions. \emph{Reconstruction.} Put \(S_k=u_k'(X_k^1)u_k'(X_k^2)\), so \(|S_k|\le L^2\) and \(Q_{12}=M^{-1}\sum_kS_k\). For a bounded continuous old-array test \(F\), exchangeability gives \begin{equation*} \E\langle FQ_{12}\rangle=\E\langle FS_1\rangle,\qquad \E\langle FQ_{12}^2\rangle =\frac{M-1}{M}\E\langle FS_1S_2\rangle+O(M^{-1}). \end{equation*} The repeated-index error is at most \(\|F\|_\infty L^4/M\). One- and two-row restoration therefore give \begin{align*} \E\langle FQ_{12}\rangle &\longrightarrow\E_U[F_\infty p_c(U_{12})], \\ \E\langle FQ_{12}^2\rangle &\longrightarrow\E_U[F_\infty p_c(U_{12})^2]. \end{align*} The common covariance rule \(c=K(q,\rho_d)\) and the plateau property give \(p_c(u)=\bar p_c(q(u),\rho_d(u))\) almost everywhere, with \(\bar p_c\) bounded Borel; see \cref{lem:finite-rpc-representation}. For any joint subsequential limit \((A,Q)\), where \(A\) retains the finite observed overlap--ancestry array containing \((R_{12},r_{12})\), the first- and second-moment identities against continuous tests identify finite signed Borel measures. Hence \[ \E[Q\mid A]=\bar p_c(R_{12},r_{12}),\qquad \E[Q^2\mid A]=\bar p_c(R_{12},r_{12})^2. \] The conditional variance vanishes, so \(Q=\bar p_c(R_{12},r_{12})\) almost surely; in the canonical representation this equals \(p_c(U_{12})\). Boundedness gives compactness, and every joint limit has this form. For \(T_N=M^{-1}\sum_k\chi(\Theta_k,X_k^1)\), restore one row to compute \(\mathbb E\langle T_N\rangle\) and two rows to compute \(\mathbb E\langle T_N^2\rangle\). The product identity \eqref{eq:rpc-insertion-one-replica} gives the limits \(m_c(\chi)\) and \(m_c(\chi)^2\); repeated indices contribute \(O(M^{-1})\). This proves \eqref{eq:empirical-one-row-limit}. The scalar limits exist by the same covariance and weight-moment passage, so no separate approximation argument is needed. For the family of Section~\ref{sec:overview-N-cavity}, deletion leaves the row noises independent, and their covariance is \eqref{eq:overview-inserted-constraint-covariance}. The lookup \(q^{0,\sharp}\) is a function on the fixed finite ancestry alphabet, so convergence of the observed \((R,r)\)-array passes its covariance to the limit. Since \(q^{0,\sharp}(\rho_d(u))=q^0(u)\), this verifies \eqref{eq:overview-fresh-covariance-limit} with \(c=(1-t)q^0+tq\). \end{proof} The following comparison will control changes in the insertion field during coordinate lifting. \begin{lemma}[Comparison of tilted averages] \label{lem:tilted-comparison} Let \(\mu\) be a random probability measure on a measurable space. For \(Z_H=\mu(e^H)\in(0,\infty)\), write \(\langle F\rangle_H=(e^H\mu/Z_H)^{\otimes n}(F)\). Let \(F\) be a bounded \(n\)-replica test, with deterministic bound \(\|F\|_\infty\). \begin{enumerate}[label=(\roman*),leftmargin=*] \item \emph{General comparison.} For real measurable fields \(X,Y\), put \[ d^2=\E\mu((Y-X)^2),\qquad B=\max_{\substack{H\in\{X,Y\}\\\epsilon=\pm1}} \E\mu(e^{4\epsilon H}). \] If \(d,B<\infty\), then \begin{equation*} \begin{aligned} |\E\log Z_Y-\E\log Z_X|&\le\sqrt B\,d,\\ |\E\langle F\rangle_Y-\E\langle F\rangle_X| &\le2n\|F\|_\infty\sqrt B\,d. \end{aligned} \end{equation*} No independence among \(\mu,X,Y,F\) is required. If \(|Y-X|\le\varepsilon\), both bounds hold before averaging with \(\sqrt B\,d\) replaced by \(\varepsilon\), requiring only finite positive normalizers; the averaged log bounds then hold whenever the logarithms are integrable. \item \emph{Gaussian comparison.} Condition on an environment determining \(\mu\) and \(F\). Let \(X,Y\) be conditionally centered Gaussian fields with covariance kernels \(C_X,C_Y\), whose diagonal values are bounded by a finite deterministic constant. If \(\|C_Y-C_X\|_\infty\le\delta\), then \begin{equation*} \begin{aligned} |\E\log Z_Y-\E\log Z_X|&\le\delta,\\ |\E\langle F\rangle_Y-\E\langle F\rangle_X| &\le n(2n+1)\|F\|_\infty\delta. \end{aligned} \end{equation*} \end{enumerate} \end{lemma} \begin{proof} For (i), set \(D=Y-X\), \(H_t=X+tD\), and \(Z_t=\mu(e^{H_t})\). Holding \(F\) fixed, differentiation gives \begin{equation*} \frac{\dd}{\dd t}\log Z_t=\langle D\rangle_t,\qquad \frac{\dd}{\dd t}\langle F\rangle_t =\left\langle F\left(\sum_{i=1}^nD(s^i)-nD(s^{n+1})\right)\right\rangle_t. \end{equation*} Convexity bounds \(\E\mu(e^{\pm4H_t})\) by \(B\). Jensen gives \(Z_t^{-4}\le\mu(e^{-4H_t})\), so H\"older yields \[ \E\langle|D|\rangle_t \le d\,(\E\mu(e^{4H_t}))^{1/4}(\E\mu(e^{-4H_t}))^{1/4} \le\sqrt B\,d. \] Integration proves (i); truncation and this bound justify differentiation. For \(|D|\le\varepsilon\), use \(\langle|D|\rangle_t\le\varepsilon\) directly, before averaging. For (ii), take conditionally independent endpoint copies and set \(V_t=\sqrt{1-t}X+\sqrt tY\), \(Z_t=\mu(e^{V_t})\), and \(\Delta_{ij}=(C_Y-C_X)(s^i,s^j)\). Gaussian integration by parts gives \begin{equation}\label{eq:gaussian-log-interpolation} \frac{\dd}{\dd t}\E\log Z_t =\frac12\E\langle\Delta_{11}-\Delta_{12}\rangle_t \end{equation} and \begin{equation*} \begin{split} \frac{\dd}{\dd t}\E\langle F\rangle_t =\frac12\E\Big\langle F\Big[&\sum_{i,j\le n}\Delta_{ij} -2n\sum_{i\le n}\Delta_{i,n+1}\\ &-n\Delta_{n+1,n+1} +n(n+1)\Delta_{n+1,n+2}\Big]\Big\rangle_t. \end{split} \end{equation*} These identities hold first for finite state spaces. If \(B_0\) bounds the field variances, then \(\E\mu(e^{pV_t}+e^{-pV_t})\le2e^{p^2B_0/2}\). The empirical measure argument of \cref{lem:change-density-limit}, using conditional samples independent of the Gaussian noises, passes the integrated identities to general \(\mu\). Bounding each covariance difference by \(\delta\) and summing the absolute coefficients proves (ii). \end{proof} \subsection{The quadratic coefficient for block insertion} \label{sec:simple-empirical-row-recovery} The block calculation at \(t=1\) uses the coefficient \begin{equation*} B_N=\frac1M\sum_{k=1}^M\{u_k''(h_k)-h_k u_k'(h_k)\}. \end{equation*} The diagonal correlation \(Q_{11}\) is already covered by \eqref{eq:empirical-one-row-limit}. We recover \(B_N\) by Gaussian integration by parts and the limiting GG identity. The spectral bound below controls its unbounded factor \(h_k\). \begin{lemma}[Uniform constraint bounds] \label{lem:spectral-moments} Let \(G\) be the \(M\times N\) Gaussian constraint matrix and put \(\Lambda_N=\|G\|_{\rm op}^2/M\). If \(M/N\) stays in a compact subset of \((0,\infty)\), then, for every \(r,c\ge0\), \[ \sup_N\E\Lambda_N^r<\infty,\qquad \sup_N\E e^{c\sqrt{\Lambda_N}}<\infty. \] On either spin space, \begin{equation}\label{eq:simple-spectral-empirical} \frac1M\sum_k h_k(\sigma)^2\le\Lambda_N,\qquad \frac1M\sum_k|h_k(\sigma)|\le\sqrt{\Lambda_N}. \end{equation} \end{lemma} \begin{proof} Use \(\|\sigma\|^2=N\), Cauchy--Schwarz, and the Gaussian matrix tail bound \(\mathbb P\{\|G\|_{\rm op}\ge\sqrt M+\sqrt N+s\}\le e^{-s^2/2}\). \end{proof} In particular, the common bounds on \(u_k',u_k''\) give the uniform envelope \begin{equation}\label{eq:quadratic-coefficient-exponential-envelope} |B_N|\le C_u(1+\sqrt{\Lambda_N}). \end{equation} \begin{lemma}[The quadratic coefficient] \label{lem:empirical-row-coefficients} At \(t=1\), assume the full measure and its one- and two-constraint deletions satisfy the array-limit hypotheses of \cref{prop:0916-rows}, with common spin path \(q\). Put \begin{equation*} \overline B=\langle q,p_q\rangle-p_{\rm diag}(q). \end{equation*} Then \begin{align*} \E\langle|Q_{11}-p_{\rm diag}(q)|^2\rangle&\longrightarrow0, \\ \E\langle|B_N-\overline B|^2\rangle&\longrightarrow0. \end{align*} \end{lemma} \begin{proof} The first limit is \eqref{eq:empirical-one-row-limit} with \(\chi(\theta,x)=u_\theta'(x)^2\). To recover \(B_N\), first add back its diagonal derivative term. Put \[ C_N=B_N+Q_{11},\qquad K_{\ell m}=R_{\ell m}Q_{\ell m},\qquad \overline K_N=\E\langle K_{12}\rangle. \] For an \(n\)-replica test \(F\), let \[ D_n(F)=\frac1{M\sqrt N}\sum_{k,i} \E\langle\sigma_i^1u_k'(h_k^1)\partial_{g_{ki}}F\rangle, \] where the derivative holds the replica states and marks fixed. Gaussian integration by parts conditional on the marks, using independence of the perturbation from the constraint vectors, gives \begin{equation}\label{eq:reaction-replica-ibp} \E\langle F C_N^1\rangle =\E\left\langle F\left(nK_{1,n+1} -\sum_{\ell=2}^nK_{1\ell}\right)\right\rangle-D_n(F). \end{equation} The derivatives of \(u_k'(h_k^1)\) and of the first replica's density produce \(u_k''(h_k^1)\) and \(u_k'(h_k^1)^2\); the other density derivatives produce the displayed overlaps. With \(n=1,F=1\), this yields \(\E\langle C_N\rangle=\overline K_N\). The uniform derivative bounds and \(\|\sigma^\ell\|=\sqrt N\) give \[ \|\nabla_{g_k}C_N^1\|\le\frac{C_u}{M}(1+|h_k^1|),\qquad \|\nabla_{g_k}K_{12}\|\le\frac{C_u}{M}. \] Contracting with \(\sigma^1\) and using \eqref{eq:simple-spectral-empirical} yields \begin{equation*} |D_1(C_N^1)|+|D_2(K_{12})| \le\frac{C_u}{M}\bigl(1+\E\sqrt{\Lambda_N}\bigr)=O(M^{-1}). \end{equation*} The same bound and \eqref{eq:quadratic-coefficient-exponential-envelope} justify these integrations by parts by truncation. Applying \eqref{eq:reaction-replica-ibp} to these two tests gives \[ \E\langle(C_N-\overline K_N)^2\rangle =\{2\E\langle K_{12}K_{13}\rangle -\E\langle K_{12}^2\rangle-\overline K_N^2\} -D_2(K_{12})-D_1(C_N^1). \] Row reconstruction identifies the limiting \(K_{12}=q(U_{12})p_q(U_{12})\). The GG identity for \(U\), applied to the bounded Borel function \(u\mapsto q(u)p_q(u)\), makes the expression in braces vanish. This is the identity of \cite[Theorem~2.10]{PanchenkoBook}, extended by \cite[Theorem~2.17]{PanchenkoBook} and equality of the resulting finite dimensional laws to bounded Borel tests. Thus \(C_N\to\langle q,p_q\rangle\) in averaged \(L^2\). Subtract the established limit of \(Q_{11}\) to obtain the result. \end{proof} \section{Scalar transforms and supporting inequalities} \label{sec:spin-laws} The Ising response is a two-spin correlation. The spherical response is the explicit function obtained from a Gaussian quadratic integral. We prove that both responses give supporting pairs for the spin cost. The same Gaussian integral supplies the spherical block increment. Write \(\mathcal S_{\rm I},\mathcal S_{\rm S}\), \(\mathcal A_{\rm I},\mathcal A_{\rm S}\), and \(\widehat q_a^{\rm I},\widehat q_a^{\rm S}\) for the two priors. For a partition \(\pi=(x_0,\ldots,x_n)\), write a step source as \begin{equation}\label{eq:spin-step-source} 0=x_00\), then \begin{equation*} K_{a_n}(\beta_n)\to K_a(\beta),\qquad \|q^{\rm co}_{a_n,\beta_n}-q^{\rm co}_{a,\beta}\|_{L^1}\to0, \qquad V_{a_n}(\beta_n)\to V_a(\beta). \end{equation*} For each fixed \(R>0\), the quantities \(K_{a_n,R}(\beta_n)\) also converge to a limit \(K_{a,R}(\beta)\) independent of the approximation, and \begin{equation}\label{eq:scalar-log-cutoff} K_{a,R}(\beta)\uparrow K_a(\beta)\qquad(R\to\infty). \end{equation} \end{lemma} \begin{proof} \emph{Gaussian recursion.} For a step source, set \begin{equation*} d_n=\beta,\qquad d_j=d_{j+1}-x_j(a_{j+1}-a_j),\qquad 1\le j\beta_-(a)\). The leaf integral is \(\beta^{-1/2}e^{w^2/(2\beta)}\). Square completion at level \(j\), with \(\Delta_j=a_{j+1}-a_j\), gives \[ \frac1{x_j}\log\E_Z e^{x_j[C+(w+\sqrt{\Delta_j}Z)^2/(2d_{j+1})]} =C+\frac1{2x_j}\log\frac{d_{j+1}}{d_j} +\frac{w^2}{2d_j}. \] This is the Gaussian recursion of \cite[Lemma~3.5]{TalagrandSpherical}; positivity of \(d_j\) checks exactly the fractional moments in \cref{rpc:prop-quasistationarity}. Averaging the unchanged root Gaussian of variance \(a_1\) gives \begin{equation}\label{eq:scalar-log-normalizer} K_a(\beta)=-\frac12\log\beta+\frac{a_1}{2d_1} +\frac12\sum_{j=1}^{n-1}\frac1{x_j}\log\frac{d_{j+1}}{d_j}. \end{equation} We need only the one-coordinate marginal of the normalized scalar integral, to control the logarithmic cutoff below. Under the tilted transition \eqref{eq:rpc-mark-transition}, the next field, given its current value \(w\), is Gaussian with mean \(d_{j+1}w/d_j\) and variance \(\Delta_jd_{j+1}/d_j\). At the leaf, the coordinate has conditional law \(\mathcal N(w/\beta,\beta^{-1})\). Consequently its averaged law is \begin{equation}\label{eq:scalar-gaussian-representation} \tau\stackrel{\rm d}= \frac{Z_n}{\sqrt\beta}+\frac{\sqrt{a_1}}{d_1}Z_0 +\sum_{j=1}^{n-1}\sqrt{\frac{a_{j+1}-a_j}{d_jd_{j+1}}}\,Z_j, \end{equation} with independent standard Gaussians. Its variance is \begin{equation*} V_a(\beta)=\frac1\beta+\frac{a_1}{d_1^2} +\sum_{j=1}^{n-1}\frac{a_{j+1}-a_j}{d_jd_{j+1}}. \end{equation*} On \((0,a_1]\), \(\pi_{a,\beta}=d_1\); on \((a_j,a_{j+1}]\), \(\pi_{a,\beta}(t)=d_j+x_j(t-a_j)\). Integrating \(\pi^{-1}\) and \(\pi^{-2}\) gives \eqref{eq:general-scalar-normalizer} and \eqref{eq:general-response-and-variance}. In particular, \begin{equation}\label{eq:scalar-coordinate-response} q^{\rm co}_{a,\beta}(r)=\frac{a_1}{d_1^2} +\sum_{i=1}^{j-1}\frac{a_{i+1}-a_i}{d_id_{i+1}}, \qquad x_{j-1}\le r0\), apply \cite[Proposition~3.1 and (3.34)]{TalagrandSpherical} with \[ \xi(s)=a_*s^2/2,\qquad h=0,\qquad b=\beta,\qquad x(s)=\operatorname{Leb}\{r\in[0,1):a(r)\le a_*s\}. \] The overlap levels are \(a_j/a_*\), with masses \(x_j\); coinciding terminal levels are allowed in \cite[(3.2)]{TalagrandSpherical}. Both surface measures are normalized, so his spin pressure is \(\varphi_K(0)=\mathcal S_{{\rm S},K}(a)+a_*/2\). His \(d(s)=a_*\int_s^1x(u)\dd u\) satisfies \(d(0)=\beta_-(a)\) and \(\beta-d(s)=\pi_{a,\beta}(a_*s)\). Changing variables in (3.34) gives \[ W(x,\beta)-a_*/2 =K_a(\beta)+(\beta-1-a_*)/2=F_a(\beta). \] Proposition~3.1 therefore proves \eqref{eq:spherical-large-block}. Its restriction \(b>1\) loses no minimizer, since \(V_a(\bar\beta(a))=1\) implies \(\bar\beta(a)>1\) for \(a_*>0\). For \(a=0\), both sides are zero and \(\bar\beta(0)=1\). For general \(a\), choose endpoint-preserving step approximations \(a_m\to a\). Their minimizing precisions lie in \([1,1+a_*]\) and satisfy \(\pi\ge1\). The stability in \cref{prop:rpc-spherical-insertion} shows that every subsequential limit solves \(V_a(\beta)=1\), hence \(\bar\beta(a_m)\to\bar\beta(a)\) and \(F_{a_m}(\bar\beta(a_m))\to F_a(\bar\beta(a))\). The uniform bound \eqref{eq:spin-transform-lipschitz} then passes both \eqref{eq:general-scalar-spherical-variational} and \eqref{eq:spherical-large-block} to \(a\). For \(q_*=q(1-)<1\), \(\gamma_q=1\) above \(q_*\), so \eqref{eq:spherical-entropy} becomes \begin{equation}\label{eq:spherical-entropy-classical} \mathcal A_{\rm S}(q) =\frac12\int_0^{q_*}\frac{\dd t}{\int_t^1\gamma_q(s)\,\dd s} +\frac12\log(1-q_*). \end{equation} For its finite-grid expression, put \(\mathscr Q_{<1}=\{q\in\mathscr Q:q(1-)<1\}\) and \(\lambda_{\gamma_q}(t)=\int_t^1\gamma_q(s)\dd s\). If \(q=q_j\) on the grid intervals and \(q_n<1\), its susceptibilities are \begin{equation*} \lambda_n=1-q_n,\qquad \lambda_j=\lambda_{\gamma_q}(q_j) =\lambda_{j+1}+x_j(q_{j+1}-q_j),\quad j\max\{1,\|a\|_\infty\}\), and put \[ u_A(x)=\log\cosh(\sqrt A\,x)-A/2,\qquad c_a(r)=a(r)/A\ (r<1),\quad c_a(1)=1. \] For \(X=(w^a+\sqrt{A-a_*}\,z)/\sqrt A\), terminal Gaussian integration gives \begin{align*} \int e^{u_A(X)}\,\gamma(\dd z) &=e^{-a_*/2}\cosh w^a,\\ \frac{\int u_A'(X)e^{u_A(X)}\,\gamma(\dd z)} {\int e^{u_A(X)}\,\gamma(\dd z)} &=\sqrt A\,\tanh w^a. \end{align*} The first identity identifies \(K\) spin coordinates with \(K\) independent constraint insertions; the second identifies their response. The recursion for this spin insertion is the step-coefficient solution of \eqref{eq:intro-parisi-pde} with terminal \(\log\cosh\) and horizon \(a_*\), followed by subtraction of \(a_*/2\). Thus it is exactly \eqref{eq:intro-ising-transform}. For general sources, use endpoint-preserving step approximations \(a_n\) on partitions \(\pi_n\) of vanishing mesh. The finite bases \(\mathfrak R_{\pi_n}\otimes(\delta_{-1}+\delta_1)/2\) and weights \(e^{\tau w^{a_n}-a_*/2}\) have uniformly bounded positive and inverse moments of every order. Their sampled Gaussian arrays converge, with off-diagonal path \(a\) and fixed diagonal \(a_*\). Lemma~\ref{lem:change-density-limit} therefore defines the limiting averaged Ising replica laws \(\E\langle\cdot\rangle_a^{\rm I}\). Control stability and the quantile identity pass the transform identification to the limit. The constraint identities in \cref{lem:rpc-constraint-insertion} give ancestry preservation and \begin{equation}\label{eq:ising-tensorization} \mathcal S_{{\rm I},K}(a)=\mathcal S_{{\rm I},1}(a) =\mathcal S_{\rm I}(a),\qquad K\ge1, \end{equation} \begin{equation*} \mathcal S_{\rm I}(a)=\Phi_{u_A}(c_a),\qquad \widehat q_a^{\rm I}=A^{-1}p_{c_a,u_A}. \end{equation*} In particular, \(\widehat q_a^{\rm I}\) is nonnegative, nondecreasing, bounded by one, and constant on the plateaus of \(a\). Choose one \(A\) for both endpoints of a segment \(a_s=(1-s)a+s\widetilde a\). The constraint derivative and response continuity give \begin{equation}\label{eq:ising-derivative-from-constraint} \left.\frac{\dd}{\dd s}\right|_{s=0+}\mathcal S_{\rm I}(a_s) =-\tfrac12\langle\widehat q_a^{\rm I},\widetilde a-a\rangle. \end{equation} For general sources, apply the integrated finite-grid identity and pass to the limit using the \(L^1\)-continuity of the response and \(0\le\widehat q^{\rm I}\le1\). The fixed \(A\) keeps every completed constraint variance equal to one. \paragraph{The Chen--Issa--Mourrat concavity input.} \label{par:ising-cim-normalization} \label{prop:centered-Ising-response-input} Concavity follows by identifying the spin transform with their symmetric-Ising functional: \begin{equation*} \mathcal S_{\rm I}(a)=-\psi^\circ(a/2). \end{equation*} Indeed, \(\sqrt2\,w^{a/2}\) has covariance \(a\), the diagonal subtraction is \(a_*/2\), and the Ising prior is normalized as in \eqref{eq:spin-finite-block}. The identity extends from finite grids by \eqref{eq:spin-transform-lipschitz}. Convexity of \(\psi^\circ\) in \cite[Proposition~2.2]{CIM} gives concavity of \(\mathcal S_{\rm I}\). % This is the only Chen--Issa--Mourrat input and supplies the supporting % inequality; constraint insertion gives the response and derivative. \begin{corollary}[Responses and supporting equalities] \label{prop:scalar-coordinate-laws} For either prior, \(\widehat q_a^\kappa\), with true diagonal one, belongs to \(\mathscr Q\) and is constant on every plateau of \(a\). The finite and limiting Ising scalar laws preserve genealogy: \begin{equation*} \E\langle F\rangle_{\pi;a}^{\rm I}=\E_\pi F, \qquad \E\langle F\rangle_a^{\rm I}=\E_U F \end{equation*} for every bounded Borel test of finitely many replica ancestries. Their supporting equality is \begin{equation}\label{eq:ising-exposure-identity} \mathcal A_{\rm I}(\widehat q_a^{\rm I}) =\mathcal S_{\rm I}(a)+\tfrac12\langle a,\widehat q_a^{\rm I}\rangle <\infty. \end{equation} For the spherical response, \(\widehat q_a^{\rm S}(1-)=1-\bar\beta(a)^{-1}\le a_*/(1+a_*)<1\), and \begin{align*} K_a(\bar\beta(a)) &=\mathcal S_{\rm S}(a)+\tfrac12\{1+a_*-\bar\beta(a)\}, \\ \mathcal A_{\rm S}(\widehat q_a^{\rm S}) &=\mathcal S_{\rm S}(a)+\tfrac12\langle a,\widehat q_a^{\rm S}\rangle. \end{align*} \end{corollary} \begin{proof}[Proof of the Ising assertions] The constraint specialization gives genealogy preservation and the response properties. Concavity and \eqref{eq:ising-derivative-from-constraint} give \[ \mathcal S_{\rm I}(\widetilde a) \le\mathcal S_{\rm I}(a) -\tfrac12\langle\widehat q_a^{\rm I},\widetilde a-a\rangle. \] The supremum over \(\widetilde a\in\Pb\), after adding \(\langle\widehat q_a^{\rm I},\widetilde a\rangle/2\), is attained at \(a\), proving \eqref{eq:ising-exposure-identity}. \end{proof} \subsection{The spherical supporting equalities} \label{sec:spherical-supporting-equalities} \begin{proof}[Proof of the spherical assertions of \cref{prop:scalar-coordinate-laws}] The definition \eqref{eq:general-scalar-spherical-variational} gives the normalizer identity. Also \(\widehat q_a^{\rm S}(1-)=1-\bar\beta(a)^{-1}\le a_*/(1+a_*)\); monotonicity and the plateau property follow from \eqref{eq:general-response-and-variance}. One finite-grid calculation proves the supporting equality and the Fenchel inequality needed below. Given \(q,a\) on a common grid with \(q_n<1\), set \(d_1=\lambda_1^{-1}\) and \(d_{j+1}=d_j+x_j(a_{j+1}-a_j)\). Define \(h(z)=z-1-\log z\ge0\), and recall \(x_n=1\). Substituting \eqref{eq:scalar-log-normalizer} and \eqref{eq:spherical-step-entropy}, then collecting terms, gives \begin{equation}\label{eq:spherical-fenchel-gap} 2F_a(d_n)+\langle q,a\rangle =2\mathcal A_{\rm S}(q) -\sum_{j=2}^n\left(\frac1{x_{j-1}}-\frac1{x_j}\right)h(d_j\lambda_j). \end{equation} Here \(\lambda_1=1-\int_0^1q\) cancels the root term. For the remaining terms, \(\lambda_{j-1}=\lambda_j+x_{j-1}(q_j-q_{j-1})\) makes the coefficient of \(d_j\) equal to \(-\lambda_j(1/x_{j-1}-1/x_j)\), including \(j=n\). Since \(\mathcal S_{\rm S}(a)\le F_a(d_n)\), this yields \begin{equation}\label{eq:spherical-fenchel} \mathcal S_{\rm S}(a)+\tfrac12\langle q,a\rangle \le\mathcal A_{\rm S}(q). \end{equation} For \(q=\widehat q_a^{\rm S}\), the explicit response formula \eqref{eq:scalar-coordinate-response} and \(V_a(\bar\beta)=1\) give, with the minimizing precisions, \begin{equation*} q_1=a_1d_1^{-2},\qquad q_n=1-d_n^{-1},\qquad q_{j+1}-q_j=\frac{d_j^{-1}-d_{j+1}^{-1}}{x_j}. \end{equation*} Thus \(\lambda_j=d_j^{-1}\), the matching relation in \cite[Lemma~4.1, (4.8)--(4.11)]{TalagrandSpherical}; all gaps vanish and \(d_n\) is the minimizer. Consequently \begin{equation*} \mathcal S_{\rm S}(a) =\mathcal A_{\rm S}(q)-\tfrac12\langle q,a\rangle. \end{equation*} Endpoint-preserving step approximations of \(a\), the stability in \cref{prop:rpc-spherical-insertion}, and the common response bound \(q(1-)\le a_*/(1+a_*)\) pass this equality to general sources by \eqref{eq:spherical-entropy-continuity}. The same continuity and \eqref{eq:spin-transform-lipschitz} extend \eqref{eq:spherical-fenchel} to all \(q\in\mathscr Q_{<1}\), \(a\in\Pb\). \end{proof} For later use, let \(\overline q_x\) be the constant off-diagonal path: \begin{equation*} \overline q_x(r)=x\quad(r<1),\qquad \overline q_x(1)=1, \qquad 0\le x<1. \end{equation*} \begin{theorem}[Spherical conjugate formula] \label{thm:spherical-entropy-identification} For every \(q\in\mathscr Q\), the cost defined in \eqref{eq:spherical-entropy} satisfies \begin{equation}\label{eq:spherical-conjugate-identification} \mathcal A_{\rm S}(q) =\sup_{a\in\Pb}\{\mathcal S_{\rm S}(a)+\tfrac12\langle a,q\rangle\}. \end{equation} For every \(L^1\)-continuous \(\Psi:\mathscr Q\to\mathbb R\), \begin{equation}\label{eq:spherical-endpoint-density} \inf_{q\in\mathscr Q}\{\mathcal A_{\rm S}(q)+\Psi(q)\} =\inf_{q\in\mathscr Q_{<1}}\{\mathcal A_{\rm S}(q)+\Psi(q)\}. \end{equation} \end{theorem} \begin{proof} Denote the supremum in \eqref{eq:spherical-conjugate-identification} by \(C(q)\). The Fenchel inequality gives \(C(q)\le\mathcal A_{\rm S}(q)\) when \(q(1-)<1\). For step \(q\), construct its inverse source by \[ a_1=\frac{q_1}{\lambda_1^2},\qquad a_{j+1}-a_j=\frac1{x_j} \left(\frac1{\lambda_{j+1}}-\frac1{\lambda_j}\right). \] This is a bounded nonnegative nondecreasing source. With \(d_j=\lambda_j^{-1}\), its scalar response is \(q\) and \(V_a(d_n)=q_n+d_n^{-1}=1\), so the trial precision is the minimizer. All gaps in \eqref{eq:spherical-fenchel-gap} vanish, giving equality, also when \(n=1\) and the sum is empty. For general \(q(1-)\le\bar q<1\), take step approximations \(q_m\) with this same endpoint bound. Their inverse sources obey \(\|a_m\|_\infty\le\bar q/(1-\bar q)^2\), and hence \[ C(q)\ge\mathcal S_{\rm S}(a_m)+\tfrac12\langle q,a_m\rangle =\mathcal A_{\rm S}(q_m)+\tfrac12\langle q-q_m,a_m\rangle \longrightarrow\mathcal A_{\rm S}(q). \] For any \(q\in\mathscr Q\), put \(q_\delta=(1-\delta)q\) off the diagonal and \(q_\delta(1)=1\). Nonnegativity of the sources and interchange of the two suprema give \[ C(q_\delta) =\sup_{a\in\Pb}\{\mathcal S_{\rm S}(a) +\tfrac12(1-\delta)\langle a,q\rangle\} \uparrow C(q). \] Since \(\gamma_{q_\delta}\downarrow\gamma_q\), monotone convergence in \eqref{eq:spherical-entropy} gives \(\mathcal A_{\rm S}(q_\delta)\uparrow\mathcal A_{\rm S}(q)\). The identity \(C(q_\delta)=\mathcal A_{\rm S}(q_\delta)\) therefore proves \(C(q)=\mathcal A_{\rm S}(q)\), including infinite values. In particular, the spherical supporting equality is \begin{equation*} \mathcal A_{\rm S}(\widehat q_a^{\rm S}) =\mathcal S_{\rm S}(a)+\tfrac12\langle a,\widehat q_a^{\rm S}\rangle. \end{equation*} Finally, \(q_\delta\to q\) in \(L^1\) and \(\mathcal A_{\rm S}(q_\delta)\to\mathcal A_{\rm S}(q)\). Applying this approximation to every finite-cost competitor proves \eqref{eq:spherical-endpoint-density}; the reverse inequality follows from inclusion of domains. \end{proof} \begin{proof}[Proof of \cref{prop:scalar-proof-input}] The Ising conjugate formula is \eqref{eq:intro-ising-cost}; the spherical formula is \cref{thm:spherical-entropy-identification}. Corollary~\ref{prop:scalar-coordinate-laws} gives the supporting equalities. The left-hand side of \eqref{eq:overview-spin-endpoint} is \(\mathcal S_{\kappa,N}(a)\) from \eqref{eq:spin-finite-block}. Its asserted value and limit follow from \eqref{eq:ising-tensorization} and \eqref{eq:spherical-large-block}, respectively. \end{proof} \section{Cavity in $N$} \label{sec:cavity-N} We first prove the overlap equation in \cref{prop:0916-coordinate}. For Ising spins, separating one coordinate determines its normalized law. On the sphere, integration by parts gives the equation directly. We then compute the logarithmic increment from inserting a fixed block of coordinates at \(t=1\). This last calculation is the only one that uses the quadratic coefficient \(B_N\). \subsection{The Ising overlap equation} \label{sec:simple-coordinate-equation} \begin{proof}[Proof of the Ising assertion of \cref{prop:0916-coordinate}] Put \(D=N+1\), write \(\sigma^D=(\sigma,\tau)\), and separate the last entries \(z_k\) of the constraint vectors. In the measure with the last coordinate deleted, retain the denominator \(\sqrt D\): \[ X_k^- =\sqrt{t/D}\,g_k\cdot\sigma+\sqrt{1-t}\,Y_k, \qquad X_k^D=X_k^-+\sqrt{t/D}\,z_k\tau. \] Use \(V_N\) in this deleted measure. Its fresh constraint variance is \(1-t/D\), and its off-diagonal covariance is \(t(N/D)R+(1-t)q^{0,\sharp}(r)\). The variable-diagonal version of \cref{prop:0916-rows} therefore applies. Select a common refinement for this measure and its one- and two-constraint deletions, and write \(q^-\) for their common limiting spin path. Set \[ c^-=(1-t)q^0+tq^-,\qquad a^-=(1-t)a^0+t\alpha p_{c^-}. \] No comparison with the usual \(N\)-dimensional denominator is made. The lifted and full overlaps differ by at most \(2/D\), while their ancestries agree. The covariance of \(V_N\) is \(N^{3/4}\) times a uniformly bounded Lipschitz kernel. Consequently replacing \(V_D(\sigma,\tau,\eta)\) by \(V_N(\sigma,\eta)\) changes its covariance by \(O(N^{-1/4})\), uniformly in the states and amplitudes. Apply \cref{lem:tilted-comparison}(ii) with the perturbations removed from the common density. The replacement changes bounded replicated tests by \(o(1)\). Since \(\tau^2=1\), Taylor expansion gives \begin{equation*} \begin{gathered} \sum_{k\le M}\{u_k(X_k^D)-u_k(X_k^-)\} =\tau Z_N+\frac{tM}{2D}C_N+\varepsilon_N,\\ Z_N=\sqrt{t/D}\sum_{k\le M}z_k u_k'(X_k^-),\quad C_N=M^{-1}\sum_{k\le M}u_k''(X_k^-). \end{gathered} \end{equation*} Before insertion, the centered quadratic term has conditional second moment at most \(2t^2ML_2^2/D^2\). The cubic remainder has averaged second moment \(O(N^{-1})\), by the uniform third-derivative bound and Gaussian moments. Thus \(\E\mu_N(\varepsilon_N^2)=O(N^{-1})\), where \(\mu_N\) is the deleted measure times the symmetric Ising prior of \(\tau\). The exact increment is uniformly Lipschitz in \((z_k)\), with squared constant at most \(tML^2/D\), and its conditional Gaussian mean has absolute value at most \(tML_2/(2D)\). The approximate increment has the same positive and inverse exponential moment bounds, since \(C_N\) is bounded and \(Z_N\) has bounded conditional variance. The fresh source term \(\tau W_{t,D}\) also has bounded Gaussian variance. Lemma~\ref{lem:tilted-comparison}(i) therefore transfers the Taylor estimate through the normalizer. Apply \eqref{eq:empirical-one-row-limit} with \(\chi(\theta,x)=u_\theta''(x)\) and with \(\chi(\theta,x)=u_\theta'(x)^2\). It gives \[ C_N\longrightarrow c'',\qquad Q_{11}^-\longrightarrow p_{\rm diag}(c^-) \quad\text{in averaged }L^2, \] where \(c''\) is deterministic on the selected refinement. Row reconstruction also gives \(Q_{12}^-\Rightarrow p_{c^-}(U_{12})\), jointly with the retained array. Conditional on the old replicas, \(Z_N+W_{t,D}\) is Gaussian, with off-diagonal covariance converging to \(a^-(U_{12})\) and diagonal converging to \((1-t)a_*^0+t\alpha p_{\rm diag}(c^-)\). The covariance and coefficient limits, together with the preceding moment bounds, verify \cref{lem:change-density-limit}. Thus the normalized insertion law passes to its scalar limit. Represent the excess diagonal variance by an independent terminal Gaussian of variance \(\delta=t\alpha\{p_{\rm diag}(c^-)-p_{c^-}(1-)\}\). Its integral is \(e^{\delta\tau^2/2}=e^{\delta/2}\). The factor from \(C_N\) tends to \(e^{t\alpha c''/2}\). Both factors are constants and cancel from normalization. The remaining law is exactly the scalar Ising law with weight \(e^{\tau w^{a^-}-a_*^-/2}\). This argument uses concentration of \(C_N\); the equality \(\tau^2=1\) alone would not remove a coefficient depending on the old configuration. The scalar Ising law preserves the genealogy, by \cref{prop:scalar-coordinate-laws}. The old array after insertion therefore still has law \((q^-(U),\rho_d(U))\). Since full and old overlaps differ by \(O(N^{-1})\), the limiting path \(q\) of the full system equals \(q^-\). Coordinate exchangeability gives \begin{equation*} \E\langle\tau^1\tau^2f(R_{12}^D,r_{12})\rangle_D =\E\langle R_{12}^Df(R_{12}^D,r_{12})\rangle_D. \end{equation*} Passing this bounded identity to the scalar law yields \[ \int_0^1 f(q(r),\rho_d(r)) \{\widehat q_a^{\rm I}(r)-q(r)\}\,\dd r=0, \qquad a=(1-t)a^0+t\alpha p_{(1-t)q^0+tq}. \] Both paths factor through \((q,\rho_d)\), by their plateau properties. Continuous tests determine the resulting signed measure, so \(q=\widehat q_a^{\rm I}\) almost everywhere. The supporting equality follows from \cref{prop:scalar-coordinate-laws}. \end{proof} \subsection{The spherical overlap equation} \begin{proof}[Proof of the spherical assertion of \cref{prop:0916-coordinate}] Let \(F=f(R_{12},r_{12})\), with \(f\) continuously differentiable in its first variable; the second variable lies in the finite ancestry alphabet. First retain only the first \(J\) terms of \(V_N\), and write \(H^{[J]}\) and \(\langle\cdot\rangle_J\) for this Hamiltonian and its replica averages. Put \[ v=\sigma^2-R_{12}\sigma^1,\qquad d_\ell=N^{-1}v\cdot\sigma^\ell; \qquad d_1=0,\quad d_2=1-R_{12}^2,\quad d_3=R_{23}-R_{12}R_{13}. \] The vector \(v\) is tangent to the sphere at \(\sigma^1\). Its spherical divergence and its action on the test are \[ \operatorname{div}_{\mathbb S}(v/N)=-(1-N^{-1})R_{12},\qquad (v/N)\cdot\nabla F=N^{-1}d_2\partial_Rf. \] Integrating the divergence of \((v/N)F e^{H^{[J]}(\sigma^1)}\) in the first replica therefore gives \begin{equation*} (1-N^{-1})\E\langle R_{12}F\rangle_J =N^{-1}\E\langle d_2\partial_Rf\rangle_J +N^{-1}\E\langle Fv\cdot\nabla H^{[J]}(\sigma^1)\rangle_J. \end{equation*} For the Gaussian integrations, condition on the underlying cascade weights and channel marks, and hold each replica state fixed inside its integral. Differentiating a two-replica average with respect to a Gaussian coordinate \(g\) gives \[ \partial_g\langle A\rangle_J =\left\langle\partial_gA+ A\{\partial_gH^{[J],1}+\partial_gH^{[J],2} -2\partial_gH^{[J],3}\}\right\rangle_J. \] The two numerator densities give the coefficients \(1,1\), and the square of the normalizer gives \(-2\). Neither \(F\) nor \(v\) has an explicit Gaussian derivative; their dependence through the sampling measure is included in these density terms. Write \(u_k^{\prime\ell}=u_k'(X_k^\ell)\) and \(H_{\rm con}=\sum_k u_k(X_k)\). Since \(\partial_{g_{ki}}X_k^\ell=\sqrt{t/N}\,\sigma_i^\ell\), Gaussian integration by parts gives the full constraint contribution: \[ \begin{aligned} &N^{-1}\E\langle Fv\cdot\nabla H_{\rm con}(\sigma^1)\rangle_J\\ &\quad=\frac tN\sum_k\E\left\langle F\left[ d_1u_k''(X_k^1)+u_k^{\prime1} (d_1u_k^{\prime1}+d_2u_k^{\prime2}-2d_3u_k^{\prime3}) \right]\right\rangle_J\\ &\quad=\frac{tM}{N}\E\langle F(d_2Q_{12}-2d_3Q_{13})\rangle_J. \end{aligned} \] The \(u''\) term and the diagonal \((u')^2\) term both multiply \(d_1=0\), which is the required cancellation. For the source, set \(b_{11}=(1-t)a_*^0\) and \(b_{1\ell}=(1-t)a^{0,\sharp}(r_{1\ell})\) for \(\ell=2,3\). Its covariance and the same density derivative give \[ N^{-1}\E\langle Fv\cdot W_t(\eta^1)\rangle_J =\E\langle F(d_1b_{11}+d_2b_{12}-2d_3b_{13})\rangle_J. \] Again the diagonal term vanishes. Thus the constraints and source together contribute \(\E\langle F(d_2A_{12}^N-2d_3A_{13}^N)\rangle_J\), where \[ A_{\ell m}^N=t(M/N)Q_{\ell m}+(1-t)a^{0,\sharp}(r_{\ell m}). \] The truncated perturbation has covariance \(N^{3/4}k_J(R,r)\), where \(k_J(R,r)=\sum_{j\le J}4^{-j}x_j^2K_j((1+R)/2,r)\). Differentiating this covariance in its first spin argument shows that its contribution is exactly \[ N^{-1/4}\E\left\langle F\left[ d_2\partial_Rk_J(R_{12},r_{12}) -2d_3\partial_Rk_J(R_{13},r_{13})\right]\right\rangle_J. \] Here the self term also vanishes because \(d_1=0\). Since \(|d_2|,|d_3|\le1\) and \(\|\partial_Rk_J\|_\infty\le\frac12\sum_j4^{-j}x_j^2\le2/3\), this is bounded by \(2\|F\|_\infty N^{-1/4}\), uniformly in \(J\). Now keep \(N\) fixed and let \(J\to\infty\). The covariance tail is \(O(N^{3/4}4^{-J})\), so \cref{lem:tilted-comparison}(ii) passes all the bounded replica averages in the combined identity to the full perturbation. The covariance-derivative tail is \(O(4^{-J})\), which also passes the displayed perturbation term to the limit. No convergence of sample gradients is needed. Dropping the truncation subscript, we obtain \begin{equation}\label{eq:spherical-replicated-equation} \begin{split} (1-N^{-1})\E\langle R_{12}F\rangle ={}&N^{-1}\E\langle d_2\partial_Rf\rangle +\E\langle F(d_2A_{12}^N-2d_3A_{13}^N)\rangle+\varepsilon_N,\\ &|\varepsilon_N|\le2\|F\|_\infty N^{-1/4}. \end{split} \end{equation} Row reconstruction gives \(A_{\ell m}^N\Rightarrow a(U_{\ell m})\), jointly with the overlap--ancestry array. We may now let \(N\to\infty\) in \eqref{eq:spherical-replicated-equation}: the test-derivative term and \(\varepsilon_N\) vanish. To evaluate the limit, recall the conditional triple law of the uniform RPC: for bounded \(H\), \begin{equation}\label{eq:uniform-rpc-triple-law} \begin{split} 2\E[H(U_{13},U_{23})\mid U_{12}=r] ={}&\int_0^r H(s,s)\,\dd s\\ &+\int_r^1\{H(s,r)+H(r,s)\}\,\dd s+rH(r,r). \end{split} \end{equation} The GG identity gives the joint law of \((U_{12},U_{13})\) as one half product measure plus one half diagonal measure; ultrametricity and replica symmetry then give this formula. Set \[ \begin{gathered} B=\int_0^1qa,\qquad C(r)=1-B+\int_r^1a,\qquad I(r)=\int_0^rqa,\\ D(r)=1-\int_r^1q,\qquad \Delta(r)=D(r)-rq(r). \end{gathered} \] Apply \eqref{eq:uniform-rpc-triple-law} to \(H(s,v)=\{q(v)-q(r)q(s)\}a(s)\). Collecting terms yields \[ \int_0^1 f(q(r),\rho_d(r)) \{q(r)C(r)+I(r)-a(r)\Delta(r)\}\,\dd r=0. \] On every common fiber of \(q\) and \(\rho_d\), the path \(a\) is constant. The expression in braces is also constant there: the derivatives of its integral terms cancel. Thus it factors through \((q,\rho_d)\), by the factorization argument in \cref{lem:finite-rpc-representation}. Smooth tests in the first variable and arbitrary tests on the finite ancestry alphabet are dense in continuous tests. Equality of the resulting finite measures therefore implies \begin{equation}\label{eq:spherical-path-equation} q(r)C(r)+I(r)=a(r)\Delta(r)\quad\text{for a.e. }r. \end{equation} This equation can be solved without assuming differentiability of the paths. The function \(CD+rI\) is absolutely continuous, and \eqref{eq:spherical-path-equation} gives \[ (CD+rI)'=-aD+Cq+I+rqa=0\quad\text{a.e.} \] Its value at \(r=1\) is one. Expanding the product and using \eqref{eq:spherical-path-equation} proves \begin{equation}\label{eq:spherical-precision-matching} P(r)\Delta(r)=1,\qquad P(r):=C(r)+ra(r). \end{equation} Right-continuous versions give this identity at every \(r<1\). At the terminal point it yields \[ (1+a_*-B)(1-q_*)=1. \] Thus \(q_*<1\), and \(\beta:=1+a_*-B=(1-q_*)^{-1}\). Layer cake gives \(P(r)=\pi_{a,\beta}(a(r))\), using the precision function defined in Section~2. Moreover \(P(0)=1/\Delta(0)>0\), so \(\beta\in\mathfrak D_a\). In the sense of Stieltjes measures, \(\dd P=r\,\dd a\) and \(\dd\Delta=-r\,\dd q\). The continuous parts of \eqref{eq:spherical-precision-matching} give \(\dd q=\dd a/P^2\) away from zero. At a jump at \(r>0\), the same identity gives exactly \[ q(r)-q(r-)=\frac{a(r)-a(r-)}{P(r-)P(r)} =\int_{a(r-)}^{a(r)} \frac{\dd v}{[P(r-)+r\{v-a(r-)\}]^2}. \] At zero, \eqref{eq:spherical-path-equation} gives \(q(0)=a(0)/P(0)^2\). Together these identities prove \[ q(r)=\int_0^{a(r)}\pi_{a,\beta}(v)^{-2}\,\dd v =q^{\rm co}_{a,\beta}(r). \] The terminal identity says \(V_a(\beta)=q_*+\beta^{-1}=1\). Uniqueness of this precision gives \(\beta=\bar\beta(a)\), hence \(q=\widehat q_a^{\rm S}\). The supporting equality follows from \cref{prop:scalar-coordinate-laws}. \end{proof} \subsection{The logarithmic increment of a fixed block} \label{sec:simple-small-coordinate} \begin{proof}[Proof of \cref{prop:log-block-insertion}] Fix \(K\ge1\), put \(D=N+K\), and use \(t=1\). The old measure has \(M\) constraints in dimension \(N\). For \(\tau\in\mathbb R^K\), use \[ \iota_\tau^{\rm I}\sigma=(\sigma,\tau),\qquad \iota_\tau^{\rm S}\sigma= \left(\sqrt{\frac{D-\|\tau\|^2}{N}}\,\sigma,\tau\right). \] In the Ising case \(\tau\in\{-1,1\}^K\). On every fixed cube \([-R,R]^K\), the old and full overlaps satisfy \begin{equation*} \max_{\ell,m}|R_{\ell m}^D-R_{\ell m}^N| \le C_{K,R}/N, \end{equation*} and their ancestry labels agree. The covariance comparison used above therefore replaces \(V_D\circ\iota\) by \(V_N\) with error \(O_{K,R}(N^{-1/4})\) in expected logarithms. In this comparison the common density contains neither perturbation; any old partition function in the insertion ratio cancels from the difference. For spherical spins, the density of the last \(K\) coordinates is \begin{equation*} f_{D,K}(\tau)= \frac{\Gamma(D/2)}{(\pi D)^{K/2}\Gamma((D-K)/2)} \left(1-\frac{\|\tau\|^2}{D}\right)_+^{(D-K-2)/2}. \end{equation*} Conditionally on \(\tau\), the first coordinates have the spherical law specified by \(\iota_\tau^{\rm S}\). Stirling's formula gives \begin{equation*} f_{D,K}(\tau)=(2\pi)^{-K/2}e^{-\|\tau\|^2/2} \{1+O_{K,R}(N^{-1})\} \quad\text{uniformly on }[-R,R]^K. \end{equation*} Let \(\nu_{N,R}\) be this prior conditioned on the cube, and \(\zeta_{N,R}\) its mass. For Ising spins take \(R\ge1\), \(\nu_{N,R}\) the product Ising prior and \(\zeta_{N,R}=1\). Write \(\mu_N=G_{N,M,1}\otimes\nu_{N,R}\). Separate \(g_k^D=(g_k,z_k)\), where \(z_k\in\mathbb R^K\), and put \[ c_D(\tau)=\sqrt{1-\|\tau\|^2/D},\qquad \xi_k=D^{-1/2}z_k\cdot\tau,\qquad Z_i=D^{-1/2}\sum_k z_{ki}u_k'(h_k). \] The exact new constraint input is \(c_D(\tau)h_k+\xi_k\), for both priors. Taylor expansion gives \begin{equation*} \begin{split} U_N&:=\sum_k\{u_k(c_Dh_k+\xi_k)-u_k(h_k)\} =E_N+\tfrac12T_N+\mathcal R_N,\\ E_N&=\tau\cdot Z+\frac{M\|\tau\|^2}{2D}B_N,\qquad T_N=\sum_k u_k''(h_k) \{\xi_k^2-\|\tau\|^2/D\}. \end{split} \end{equation*} Here \(B_N=M^{-1}\sum_k\{u_k''(h_k)-h_ku_k'(h_k)\}\). Before insertion the \(z_k\)'s are independent of \(\mu_N\), so \[ \E_zT_N^2=\frac{2\|\tau\|^4}{D^2}\sum_k u_k''(h_k)^2 \le C_{K,R}/N. \] Taylor's theorem also gives \[ |\mathcal R_N|\le C_{K,R}\left\{ D^{-2}\sum_k(|h_k|+h_k^2) +D^{-1}\sum_k|h_k\xi_k|+\sum_k|\xi_k|^3 +D^{-3}\sum_k|h_k|^3\right\}. \] Using \(\sum_k h_k^2\le M\Lambda_N\), \(\sum_k|h_k|^3\le(M\Lambda_N)^{3/2}\), and Gaussian moments, \cref{lem:spectral-moments} gives \begin{equation*} \E\mu_N(T_N^2+|\mathcal R_N|^2)\le C_{K,R}/N. \end{equation*} For fixed old state and \(\tau\), the squared Lipschitz constant of \(U_N\) in \((z_k)\) is at most \(ML^2\|\tau\|^2/D\), and \[ |\E_zU_N|\le L|c_D-1|\sum_k|h_k| +\frac{ML_2\|\tau\|^2}{2D} \le C_{K,R}(1+\sqrt{\Lambda_N}). \] Gaussian concentration bounds its positive and negative exponential moments. The same bound holds for \(E_N\), since its conditional Gaussian variance is bounded and \(|B_N|\le C_u(1+\sqrt{\Lambda_N})\). H\"older and the spectral exponential bound therefore give, for each fixed \(p<\infty\), \begin{equation}\label{eq:coordinate-exponent-moment-envelope} \sup_{0\le s\le1}\E\mu_N \left(e^{p\{(1-s)E_N+sU_N\}}+ e^{-p\{(1-s)E_N+sU_N\}}\right)\le C_{p,K,R}. \end{equation} Only the exact and quadratic endpoint exponents are used here; no exponential moment of the cubic Taylor bound is required. Lemma~\ref{lem:tilted-comparison}(i) now changes the expected log normalizer by \(O_{K,R}(N^{-1/2})\) when \(U_N\) is replaced by \(E_N\). On the selected refinement, row reconstruction and \cref{lem:empirical-row-coefficients} give \[ Q_{12}\Rightarrow p_q(U_{12}),\qquad Q_{11}\longrightarrow p_{\rm diag}(q),\qquad B_N\longrightarrow\langle q,p_q\rangle-p_{\rm diag}(q), \] jointly with the old array; the last two limits hold in averaged \(L^2\). Conditional on the old measure and sampled states, the inserted Gaussian vector has covariance \[ \E_z Z_i^\ell Z_j^m=\delta_{ij}(M/D)Q_{\ell m}. \] Put \(a=\alpha p_q\). Its limit consists of \(K\) independent Gaussian trees of covariance path \(a\), each completed by independent terminal variance \(\alpha\{p_{\rm diag}(q)-p_q(1-)\}\). The covariance limits and \eqref{eq:coordinate-exponent-moment-envelope} verify \cref{lem:change-density-limit} at fixed \(K,R\). Equivalently, approximate the paths by finite cascades, preserving \(a_*\); sampled covariances, coefficients, and logarithmic normalizers have the same limit. The prior mass \(\zeta_{N,R}\) is restored after using the probability base \(\mu_N\). Integrating the terminal Gaussians cancels the \(p_{\rm diag}(q)\) term. For Ising spins the resulting block weight is \[ e^{K\langle q,a\rangle/2} \prod_{i=1}^K e^{\tau_iw_i^a-a_*/2} \left(\frac{\delta_{-1}+\delta_1}{2}\right)^{\!\otimes K} (\dd\tau). \] The independent Gaussian trees have product normalized transitions; their logarithmic RPC recursion is additive. Its expected logarithm is \(K\{\mathcal S_{\rm I}(a)+\langle q,a\rangle/2\}\). The Ising overlap equation already proved identifies this as \(K\mathcal A_{\rm I}(q)\). For spherical spins the limiting cutoff weight is \[ \mathbf1_{[-R,R]^K}(\tau) \exp\left\{\sum_{i=1}^K\tau_iw_i^a -\frac\beta2\|\tau\|^2\right\} \frac{\dd\tau}{(2\pi)^{K/2}},\qquad \beta=1+a_*-\langle q,a\rangle. \] The same product recursion gives expected logarithm \(K K_{a,R}(\beta)\). Restriction to the cube can only decrease the finite dimensional normalizer. The spherical overlap equation gives \(\beta=\bar\beta(a)\), and the logarithmic cutoff limit \eqref{eq:scalar-log-cutoff} consequently yields \[ \liminf_N\{A_{N+K}(M)-A_N(M)\} \ge K K_a(\bar\beta(a))=K\mathcal A_{\rm S}(q). \] The last equality uses the supporting identity and \(1+a_*-\bar\beta(a)=\langle q,a\rangle\). This proves the claimed increment for both priors. Only the old array was used; no limiting array in dimension \(N+K\) is needed. Finally, take \(R=1\) to obtain the uniform lower bound, without any array-limit hypothesis. The preceding lifting and Taylor comparisons are uniform for \(M/N\) in a compact subset of \((0,\infty)\) and for all amplitudes. The Gaussian part of \(E_N\) has mean zero under \(\mu_N\). Gaussian integration by parts gives \[ \E\langle B_N\rangle =\E\langle R_{12}Q_{12}-Q_{11}\rangle\ge-2L^2. \] The coordinate prior is independent of the old measure and \(\nu_{N,1}(\|\tau\|^2)\le K\). Jensen therefore gives \[ A_{N+K}(M)-A_N(M) \ge\log\zeta_{N,1}+\E\log\mu_N(e^{E_N})+o(1) \ge\log\zeta_{N,1}-MKL^2/D+o(1)\ge-C_K. \] Here \(\zeta_{N,1}=1\) for Ising spins and \(\zeta_{N,1}\to\gamma([-1,1])^K>0\) on the sphere. \end{proof} \section{Perturbation, positivity, and synchronization} \label{sec:rpc-synchronization} The free energy bounds require the conditional cavity hypotheses to hold simultaneously for the interpolating measures and their fixed row and dimension shifts. We construct a perturbation that enforces the joint GG identities at vanishing pressure cost. Pressure concentration permits a common subsequence selection; synchronization then represents the limiting joint array through a limiting genealogy array. Throughout, assume \eqref{eq:iid-smooth-core-assumption} and use the common bound \(L=L_1\). \subsection{The perturbation and its covariance} \label{sec:overview-enriched-free energy} \label{sec:overview-RJ-perturbation} Fix \(\kappa\in\{\mathrm I,\mathrm S\}\), \(d\ge2\), and the trial paths \(q^0\in\Q_d\) and \(a^0\) used in Section~\ref{sec:overview-N-cavity}. Use the finite RPC \(\mathfrak R_d\) with masses \(j/d\), \(1\le j0\), and let \(\mathcal Z(g)=\sum_\eta v_\eta L(g,Y_\eta)\) satisfy the conditional moment hypotheses of \cref{rpc:prop-quasistationarity}. Here \(g\) is a standard Gaussian vector independent of the Poisson and branch-mark trees, whose laws do not depend on \(g\). If \(\log L(g,y)\) is \(D\)-Lipschitz in \(g\), uniformly in the branch marks \(y\), then \begin{equation}\label{eq:finite-rpc-pressure-concentration} \operatorname{Var}(\log\mathcal Z)\le D^2+Cm^{-2}, \end{equation} where \(C\) is universal, independent of the later masses and the weight. \end{lemma} \begin{proof} We separate the Gaussian fluctuations of the conditional mean from the cascade fluctuations at fixed \(g\). Put \(b(g)=\log S_0(g)\), with \(S_0\) from the marked recursion. For unnormalized cascade weights \(w_\eta\), set \(T=\sum_\eta w_\eta\) and \(T'=\sum_\eta w_\eta L(g,Y_\eta)/S_0(g)\). Conditional only on \(g\), \cref{rpc:prop-quasistationarity} gives \begin{equation}\label{eq:rpc-pressure-centered-totals} T'\stackrel{\rm d}=T,\qquad \log\mathcal Z-b(g)=\log T'-\log T. \end{equation} The total-mass calculation in \cite[Lemma~2.4]{PanchenkoBook} identifies \(T\) as a deterministic scale times a positive \(m\)-stable variable. Later levels enter only through finite fractional moments, since their masses are strictly increasing. For any \(Q>0\) with \(\E e^{-sQ}=e^{-cs^m}\), take \(E\sim\operatorname{Exp}(1)\) independent of \(Q\). Then \[ \mathbb P(E/Q>s)=e^{-cs^m},\qquad \log(E/Q)\stackrel{\rm d}=m^{-1}(\log E'-\log c), \quad E'\sim\operatorname{Exp}(1). \] Since \(\log E\) and \(\log(E/Q)\) have finite second moments, so does \(\log Q\). Independence of \(E\) and \(Q\) now yields \[ \operatorname{Var}(\log Q) =(m^{-2}-1)\operatorname{Var}(\log E)\le Cm^{-2}. \] Centering the two identically distributed logarithms in \eqref{eq:rpc-pressure-centered-totals} therefore gives \[ \E[\log\mathcal Z\mid g]=b(g),\qquad \operatorname{Var}(\log\mathcal Z\mid g)\le Cm^{-2}. \] No independence of \(T,T'\) is needed. Finally, the bound \(f(g,y)\le f(\widetilde g,y)+D\|g-\widetilde g\|\) passes through every logarithmic moment transform in the backward recursion by order preservation and translation equivariance. Thus \(b\) is \(D\)-Lipschitz. Gaussian Poincar\'e and total variance prove \eqref{eq:finite-rpc-pressure-concentration}. \end{proof} Condition on all channel marks; the following constants are uniform in those marks. The common Gaussian vector consists of the constraint vectors, common trial-row and spin-source increments, and the root noises of \(V_N\). All higher ancestral increments are independent branch marks; terminal Gaussians are integrated inside \(L_N\). Using \(\|\sigma\|^2=N\), \(|u_k'|\le L\), and \eqref{eq:RJ-mixture-summability}, the squared Lipschitz constant of the Hamiltonian in the common Gaussian vector is at most \[ D_N^2\le tML^2+(1-t)q^0(0)ML^2 +(1-t)Na^0(0)+CNa_N^2\le C'N. \] Logarithmic spin and terminal integration preserves this bound. The first cascade mass is \(m=1/d\), so \cref{lem:finite-rpc-pressure-concentration} gives \[ \operatorname{Var}(\log Z_N^{\rm pert}\mid\text{channel marks}) \le CN+Cd^2. \] The bound is uniform over finite perturbation sums and bounded amplitudes. To pass to the full sum, let \(V_N^{(J)}\) retain its first \(J\) fields and let \(\langle\cdot\rangle_J\) be the associated Gibbs average. Independence and the vanishing diagonal variance of the Gaussian tail give, at fixed \(N,M\), \[ \mathbb E\left|\log Z_N^{\rm pert}-\log Z_N^{(J)}\right|^2 \le\mathbb E\left\langle e^{2|V_N-V_N^{(J)}|}-1\right\rangle_J\longrightarrow0. \] Here \(|\log\langle e^V\rangle|\le\log\langle e^{|V|}\rangle\) and Jensen's inequality give the displayed bound. Thus it also holds for the full perturbation, after fixed constraint or dimension shifts. To include mark fluctuations, resample one mark and delete its complete row. Under the deleted law \(\Gamma\), the change in \(\log Z\) is at most \(2A\), where \(A=\log\Gamma(e^{L|X|})\) and \(X\) is the fresh completed row input. By \eqref{eq:fresh-marked-row-bound}, \(\E A^2\le\E e^{2L|Z|}\) for a standard Gaussian \(Z\). Conditional Jensen and Efron--Stein therefore bound the variance of the conditional mean given the marks by \(CM\). Returning to the full disorder law, total variance gives \begin{equation}\label{eq:iid-variance} \Var(\log Z_{N,M,t})\le CN+Cd^2+CM=O_d(N), \end{equation} which proves \eqref{eq:rj-pressure-variance} and supplies concentration for the joint GG identities below. For the original centered model, omit the cascade, source fields, and perturbation. The same Gaussian gradient bound and mark-resampling argument give \(\Var(\log Z_{N,M}^{\kappa,\nu})\le ML^2+CM\). Dividing by \(N^2\) proves \eqref{eq:unperturbed-iid-variance}. \subsection{Positivity, Ghirlanda--Guerra identities, and synchronization} \label{sec:simple-rj-perturbation} The following representation identifies the array limits once the joint GG identities hold. \begin{lemma}[Uniform refinement of a finite ancestry array] \label{lem:finite-rpc-representation} Let \((R,r)\) be jointly weakly exchangeable positive semidefinite arrays with diagonal \((1,1)\), satisfying the joint GG identities, and suppose that \(r\) has the ancestry law of \(\mathfrak R_d\). Then \(R_{12}\ge0\) almost surely and the joint array has the representation \eqref{eq:overview-RJ-synchronization} for a deterministic \(q\in\Q\). This path is uniquely determined almost everywhere by the law of \(R_{12}\). Every bounded Borel function \(h\) that is constant almost everywhere on each common fiber of \(q\) and \(\rho_d\) admits a Borel factorization \(h(u)=\bar h(q(u),\rho_d(u))\) almost everywhere, with \(\bar h:[0,1]^2\to\mathbb R\). If a nondecreasing covariance path \(c\) factors through \((q,\rho_d)\), then so does \(p_c\). In particular, for the trial paths of Section~\ref{sec:overview-N-cavity}, \(t\in[0,1]\), and \(\alpha>0\), \[ c=(1-t)q^0+tq,\qquad p_c,\qquad a=(1-t)a^0+t\alpha p_c,\qquad \widehat q_a \] all have such factorizations. Their evaluations at \(U_{12}\) are therefore measurable functions of the retained pair \((R_{12},r_{12})\), although the uniform refinement \(U\) itself need not be observable. \end{lemma} \begin{proof} The scalar GG identities for \(R\) and the positivity principle \cite[Theorem~2.16]{PanchenkoBook} give \begin{equation*} \mathbb P(R_{12}<0)=0. \end{equation*} Put \(S=(R+r)/2\). Panchenko's synchronization theorem \cite[Theorem~4]{PanchenkoMultiSpecies}, applied with weights \(1/2,1/2\), gives deterministic nondecreasing \(2\)-Lipschitz maps \(L_1,L_2\) such that \[ R_{\ell m}=L_1(S_{\ell m}),\qquad r_{\ell m}=L_2(S_{\ell m}). \] The scalar array \(S\) also satisfies GG. The characterization and existence results \cite[Theorems~2.13, 2.17 and Remark~2.1]{PanchenkoBook} represent its off-diagonal law as \(s_0(U_{\ell m})\), where \(s_0\) is the nondecreasing quantile of \(S_{12}\) and \(U\) is the uniform limiting genealogy array from finite RPCs. These results apply to the off-diagonal Gram--de Finetti array via the Dovbysh--Sudakov representation, as in the proof of Theorem~2.17; the specified true diagonal is kept separately. Thus \(q=L_1\circ s_0\) and \(\rho=L_2\circ s_0\) give the joint representation. Since \(U_{12}\) is uniform, the nondecreasing \(\rho\) is the quantile of \(r_{12}\), hence equals \(\rho_d\) almost everywhere. Likewise \(q\) is the quantile of \(R_{12}\). Choose right-continuous versions and set \(q(1)=1\); then \begin{equation*} ((R_{\ell m},r_{\ell m}))_{\ell0\), \(\kappa\in\{\mathrm S,\mathrm I\}\), and \(f_1,\ldots,f_k\in\mathrm{BL}(\R)\). Define \begin{equation*} \Lambda_f(\lambda) =\frac1\alpha\mathcal P_{\kappa,u_\lambda}(\alpha), \qquad u_\lambda=\sum_{i=1}^k\lambda_i f_i, \qquad C_f=\frac14\sum_{i=1}^k\operatorname{osc}(f_i)^2. \end{equation*} Then \(\Lambda_f\) is convex and belongs to \(C^{1,1}(\R^k)\), with \begin{equation*} \|\nabla\Lambda_f(\lambda)-\nabla\Lambda_f(\lambda')\|_2 \le C_f\|\lambda-\lambda'\|_2. \end{equation*} \end{proposition} \begin{proof} First let \(q\in\Q\) be a step path, and choose a finite grid \(\pi\) containing its jump points. On the finite cascade \(\mathfrak R_\pi\), use the completed Gaussian field \(X^q\) from \eqref{eq:overview-completed-row-field}, whose construction does not depend on \(\lambda\), and define \[ G_{q,\lambda}^{\pi}(\dd\eta,\dd z) =\frac{e^{u_\lambda(X^q(\eta,z))}} {Z_{q,\lambda}^{\pi}}\,\mathfrak R_\pi(\dd\eta)\gamma(\dd z), \qquad Z_{q,\lambda}^{\pi} =\int e^{u_\lambda(X^q(\eta,z))} \mathfrak R_\pi(\dd\eta)\gamma(\dd z). \] The finite-step one-constraint identity \eqref{eq:rpc-insertion-log} gives \(\Phi_{u_\lambda}(q)=\E\log Z_{q,\lambda}^{\pi}\); for bounded Lipschitz channels the identity follows also by uniform mollification. Differentiation in \(\lambda\) is justified by boundedness of the \(f_i\). For \(h\in\R^k\), writing \(h\cdot f=\sum_i h_i f_i\), it gives \begin{equation}\label{eq:rpc-channel-curvature} \begin{aligned} D_\lambda^2\Phi_{u_\lambda}(q)[h,h]=\E\operatorname{Var}_{G_{q,\lambda}^{\pi}} \bigl(h\cdot f(X^q)\bigr)\le\tfrac14\operatorname{osc}(h\cdot f)^2 \le C_f\|h\|_2^2. \end{aligned} \end{equation} The first inequality is the variance bound for a random variable taking values in an interval of length \(\operatorname{osc}(h\cdot f)\). Thus \(\lambda\mapsto\Phi_{u_\lambda}(q)-C_f\|\lambda\|_2^2/2\) is concave for every step path \(q\). For an arbitrary \(q\in\Q\), choose step paths \(q_n\to q\) in \(L^1\). On each compact set \(K\subset\R^k\), \(L_K:=\sup_{\lambda\in K}\Lip(u_\lambda)<\infty\), and scalar path continuity \eqref{rpc:Phi-L1} gives \[ \sup_{\lambda\in K} |\Phi_{u_\lambda}(q_n)-\Phi_{u_\lambda}(q)| \le\tfrac12 L_K^2\|q_n-q\|_{L^1}\longrightarrow0. \] Concavity passes to this limit. In particular, for every \(q\in\Q\), \begin{equation*} \Phi_{u_{\lambda+h}}(q)+\Phi_{u_{\lambda-h}}(q) -2\Phi_{u_\lambda}(q)\le C_f\|h\|_2^2. \end{equation*} Only this semiconcavity bound is needed for general paths. The main variational formula now shows that \begin{equation}\label{eq:channel-pressure-semiconcavity} \Lambda_f(\lambda)-\tfrac{C_f}{2}\|\lambda\|_2^2 =\inf_{\substack{q\in\Q\\\mathcal A_\kappa(q)<\infty}} \left\{\frac{\mathcal A_\kappa(q)}\alpha +\Phi_{u_\lambda}(q)-\tfrac{C_f}{2}\|\lambda\|_2^2\right\} \end{equation} is concave: each function inside the infimum is concave by the finite-step bound \eqref{eq:rpc-channel-curvature} and the preceding limit argument. The set of finite-cost paths is nonempty and independent of \(\lambda\). On the other hand, \(\Lambda_f\) is finite and convex, since it is the pointwise limit of the convex functions \(M_N^{-1}\E\log Z_{N,M_N}^{\kappa,u_\lambda}\). For completeness, if a finite convex function \(g\) satisfies that \(g-C\|\cdot\|_2^2/2\) is concave, then its central differences are at most \(C\|h\|_2^2\). For any two subgradients \(p,p'\) at \(x\), the subgradient inequalities therefore give \(t(p-p')\cdot h\le Ct^2\|h\|_2^2\). Letting \(t\downarrow0\) shows that the subgradient is unique, so \(g\) is differentiable. Convolution with a compactly supported smooth probability density preserves convexity and the same upper curvature bound; hence \(0\le D^2g_\varepsilon\le C I_k\). The gradients converge locally to \(\nabla g\), proving that \(\nabla g\) is \(C\)-Lipschitz. Apply this fact to \eqref{eq:channel-pressure-semiconcavity}. No selection or uniqueness of a minimizing path is used. \end{proof} \subsection{finite dimensional large deviations} Fix \(\kappa,\alpha\), and abbreviate \[ \Lambda(u)=\frac{\mathcal P_{\kappa,u}(\alpha)}\alpha, \qquad \Lambda_N(u)=\frac1{M_N}\log Z_{N,M_N}^{\kappa,u}. \] All finite dimensional large-deviation principles must hold on one probability-one event. We construct this event for a countable class of channels, then extend the pressure limits to every bounded Lipschitz channel using the spectral bound. Let \(\mathbf G_N\) be the \(M_N\times N\) matrix with rows \(g_k^{\mathsf T}\), and fix \(C>(1+\alpha^{-1/2})^2\). The Gaussian matrix tail estimate in \cref{lem:spectral-moments} and Borel--Cantelli imply, almost surely, \begin{equation}\label{eq:ldp-eventual-compact-support} \sup_{\sigma\in\Sigma_N^\kappa}\int x^2\,L_{N,\sigma}(\dd x) \le\frac{\|\mathbf G_N\|_{\rm op}^2}{M_N}\le C \quad\text{for all sufficiently large }N. \end{equation} Thus \(\rho_N^\kappa\) is eventually supported on \[ K_C=\left\{\mu\in\mathcal P(\R):\int x^2\,\mu(\dd x)\le C\right\}. \] This set is weakly compact, by tightness and lower semicontinuity of the second moment. For \(u,v\in\mathrm{BL}(\R)\), set \(d_C(u,v)=\sup_{\mu\in K_C}\int|u-v|\,\dd\mu\). When \(\rho_N^\kappa(K_C)=1\), direct comparison of the integrands gives \begin{equation}\label{eq:ldp-channel-modulus} |\Lambda_N(u)-\Lambda_N(v)|\le d_C(u,v), \qquad |\Lambda(u)-\Lambda(v)|\le d_C(u,v). \end{equation} To justify the deterministic second inequality, first fix \(u,v\). The first inequality holds with probability tending to one, while \(|\Lambda_N(u)-\Lambda_N(v)|\le\|u-v\|_\infty\) always. Take expectations and use \cref{thm:iid-channels}. Let \(\mathscr D\) be the countable rational vector space generated by the constant function one and compactly supported piecewise linear functions with rational breakpoints and values. It is dense for \(d_C\) in \(\mathrm{BL}(\R)\): cutting off \(u\) outside \([-R-1,R+1]\), while leaving it unchanged on \([-R,R]\), costs at most \(C\|u\|_\infty/R^2\); the resulting compactly supported function can be approximated uniformly by rational piecewise linear functions. Bounded-channel almost-sure convergence in \cref{thm:iid-channels}, applied to a deterministic channel, gives one event of probability one on which \(\Lambda_N(v)\to\Lambda(v)\) for all \(v\in\mathscr D\), and \eqref{eq:ldp-eventual-compact-support} holds. On this event, \eqref{eq:ldp-channel-modulus} and approximation give \begin{equation}\label{eq:ldp-simultaneous-pressure-limits} \Lambda_N(u)\longrightarrow\Lambda(u) \qquad\text{for every }u\in\mathrm{BL}(\R). \end{equation} Indeed, the limit superior of the absolute difference is at most \(2d_C(u,v)\) for each \(v\in\mathscr D\). Fix an environment in this event. For any finite family \(f=(f_1,\ldots,f_k)\) of bounded Lipschitz functions, define \[ T_f(\mu)=\left(\int f_1\,\dd\mu,\ldots,\int f_k\,\dd\mu\right), \qquad Y_N(\sigma)=T_f(L_{N,\sigma}). \] Its logarithmic moment generating function satisfies the exact identity \begin{equation*} \frac1{M_N}\log\int e^{M_N\lambda\cdot Y_N(\sigma)} \mu_N^\kappa(\dd\sigma) =\Lambda_N(\lambda\cdot f) \longrightarrow\Lambda_f(\lambda). \end{equation*} The limit is finite and differentiable on all of \(\R^k\), by \cref{prop:channel-pressure-differentiability}; there is no boundary of its effective domain at which a steepness condition is needed. The vectors \(Y_N\) lie in a fixed compact rectangle. The G\"artner--Ellis theorem \cite[Theorem~2.3.6]{DemboZeitouni} therefore gives a full large-deviation principle at speed \(M_N\), with good rate \begin{equation}\label{eq:ldp-finite-moment-rate} J_f(z)=\sup_{\lambda\in\R^k} \{\lambda\cdot z-\Lambda_f(\lambda)\}. \end{equation} This conclusion holds for every such finite family on the same event. \subsection{The empirical measure principle and nonlinear free energies} We continue on this event and write \(\mathcal J=\mathcal J_{\kappa,\alpha}\). Its definition \eqref{eq:quenched-ldp-rate} is a supremum of continuous affine functions of \(\mu\), so \(\mathcal J\) is convex and lower semicontinuous. It is nonnegative because \(\Lambda(0)=0\). It is infinite outside \(K_C\): if \(\mu(x^2)>C\), choose \(R<\infty\) with \(\mu(f_R)>C\), where \(f_R(x)=\min\{x^2,R\}\) is bounded and Lipschitz. Since \(\nu(f_R)\le C\) for \(\nu\in K_C\), \eqref{eq:ldp-simultaneous-pressure-limits} gives \(\Lambda(t f_R)\le tC\) for every \(t>0\). Consequently \[ \mathcal J(\mu)\ge t\{\mu(f_R)-C\}\longrightarrow\infty. \] Thus all sublevel sets of \(\mathcal J\) are compact. \emph{Passage to empirical measures.} Enumerate \(\mathscr D=\{f_j:j\ge1\}\). The continuous coordinate map \[ T:\mathcal P(\R)\longrightarrow\R^{\mathbb N},\qquad T(\mu)=(\mu(f_j))_{j\ge1}, \] restricts to a homeomorphism of \(K_C\) onto its compact image in the product topology, since the functions in \(\mathscr D\) separate measures. Every finite projection of \(T_\#\rho_N^\kappa\) satisfies the good LDP \eqref{eq:ldp-finite-moment-rate} on the same quenched event. The Dawson--G\"artner theorem \cite[Theorem~4.6.1]{DemboZeitouni} gives a good LDP in the product space, with rate \(z\mapsto\sup_kJ_{(f_1,\ldots,f_k)}(z_1,\ldots,z_k)\). The laws are eventually supported on the closed set \(T(K_C)\). Restricting to that set and transporting by \(T^{-1}\) gives the LDP on \(K_C\), with rate \[ \sup_{u\in\operatorname{span}_{\R}\mathscr D} \{\mu(u)-\Lambda(u)\}=\mathcal J(\mu). \] The equality follows from \(d_C\)-density and \eqref{eq:ldp-channel-modulus}. Eventual support in \(K_C\) and \(\mathcal J=\infty\) outside it extend this to the weak topology on \(\mathcal P(\R)\), proving the asserted large-deviation principle. \emph{Nonlinear free energies.} For bounded weakly continuous \(\Psi\), Varadhan's lemma \cite[Theorem~4.3.1]{DemboZeitouni} now gives \[ \frac1{M_N}\log\int e^{M_N\Psi(\mu)}\,\rho_N^\kappa(\dd\mu) \longrightarrow\sup_\mu\{\Psi(\mu)-\mathcal J(\mu)\}. \] The integral on the left equals the spin integral in \eqref{eq:nonlinear-perceptron-free energy}. Multiplication by \(M_N/N\to\alpha\) proves that formula; replacing normalized Ising counting measure by the sum adds \(\log2\). This completes the proof of \cref{thm:quenched-empirical-ldp}. \section{Applications and singular limits} \label{sec:applications-proofs} We first prove the Gaussian-field extension, then pass to zero temperature for scalar projection pursuit and justify the radial passage used in \cref{thm:st-free energy}. Removing quadratic truncations and hardening distance penalties give the ground-state and capacity results. Throughout this section, \(\mathbf G_N\) denotes the \(M_N\times N\) matrix with rows \(g_k^{\mathsf T}\). \subsection{Gaussian external fields} \label{sec:gaussian-field-proof} We now prove the \hyperref[par:gaussian-field-extension]{Gaussian-field extension}, using the notation of \eqref{eq:intro-gaussian-field-partition}--\eqref{eq:intro-field-value}. \begin{proof}[Proof of the {\hyperref[par:gaussian-field-extension]{Gaussian-field extension}}] Put \(\ell_b(x)=bx\). On a step path, a Gaussian recursion step of mass \(m\) and variance \(\Delta\) contributes \(b^2m\Delta/2\). Summing and then using path continuity gives \[ \Phi_{\ell_b}(q)=\frac{b^2}{2}\int_0^1\gamma_q(s)\,\dd s =\frac{b^2}{2}\mathsf L(q). \] Apply \cref{thm:iid-channels} at zero field to \(T_N=\lfloor(\alpha+1)N\rfloor\) rows, each having channel \(u_\Theta\) with probability \(\alpha/(\alpha+1)\) and \(\ell_b\) otherwise. This is a Borel family indexed by \(\mathbb R\): assign label \(e^\theta>0\) to \(u_\theta\), label \(0\) to \(\ell_b\), and the zero channel to negative labels. This mixture has a common Lipschitz bound and the required origin moment. Its variational value is \[ \inf_{q\in\Q}\left\{\mathcal A_\kappa(q) +\alpha\E_\nu\Phi_{u_\Theta}(q)+\frac{b^2}{2}\mathsf L(q)\right\}. \] Conditional on the counts \(M'_N,K'_N=T_N-M'_N\), the linear rows sum to \(b\sqrt{K'_N/N}\,g^0\cdot\sigma\). Thus the mixture partition \(Z_N^{\rm mix}\) can be coupled to the desired partition \(Z_N(b)\) using one original-row stream and one \(g^0\), both independent of the counts and of each other; \(g^0\) is standard Gaussian even when \(K'_N=0\), when its coefficient vanishes. For a fresh original row, centering the channel in \eqref{eq:fresh-marked-row-bound} bounds its logarithmic increment in \(L^p\) by \(C_p+\|u_\Theta(0)\|_p\), \(p\in\{1,2\}\), uniformly over the independent deleted law. Conditional-count Minkowski and \(N^{-1}|(b'-b)g^0\cdot\sigma| \le|b'-b|\|g^0\|_2/\sqrt N\) give \[ \left\|\frac{\log Z_N^{\rm mix}-\log Z_N(b)}N\right\|_p \le\frac{C_p}{N}\|M'_N-M_N\|_p +|b|\left\|\sqrt{K'_N/N}-1\right\|_p=o(1). \] The last step uses the binomial count moments and \(M_N/N\to\alpha\). This proves the asserted \(L^1\) and \(L^2\) limits; the spherical entropy representation is inherited from the mixture theorem. For a deterministic channel, the squared gradient norms of the normalized pressure in the rows and in \(g^0\) are bounded by \(M_NL^2/N^2\) and \(b^2/N\), respectively. Gaussian Poincar\'e gives \begin{equation}\label{eq:gaussian-field-variance} \Var(F_{N,\kappa}^{u,b}) \le\frac{M_N\Lip(u)^2}{N^2}+\frac{b^2}{N}. \end{equation} \end{proof} \subsection{Zero temperature and the Montanari--Zhou formula} \label{sec:mz-proof} The zero-temperature passage uses control values for integrable coefficients and measurable marked terminals. We prove their stability and integrability by extending the standard step-coefficient Parisi control argument; see \cite[Lemma~18]{JagannathTobasco}. \begin{lemma}[Scalar control with integrable coefficients] \label{lem:mz-control} Let \(\lambda\in L^1_+([0,1])\), \(a\ge0\), and let \(K:\mathbb R\to\mathbb R\) be \(L\)-Lipschitz. Define \begin{equation*} \mathcal T_\lambda^a K(t,x) =\sup_v\E\left[ K\left(x+\sqrt a(B_1-B_t)+\int_t^1\lambda(s)v_s\,\dd s\right) -\frac12\int_t^1\lambda(s)v_s^2\,\dd s\right]. \end{equation*} Here \(B\) is standard Brownian motion. The controls are progressively measurable for its increments after \(t\), also when \(a=0\), and satisfy \(\E\int_t^1\lambda(s)v_s^2\,\dd s<\infty\). The value is finite and \(L\)-Lipschitz in \(x\), with \begin{equation}\label{eq:mz-control-growth} |\mathcal T_\lambda^aK(t,x)-K(x)| \le L\sqrt{\frac{2a(1-t)}\pi} +\frac{L^2}{2}\int_t^1\lambda(s)\,\dd s. \end{equation} For bounded \(K\), it is also bounded in absolute value by \(\|K\|_\infty\). The supremum is unchanged by restricting to \(|v|\le L\). The terminal map preserves order and is a contraction in the sup norm. For another such triple \((\widetilde\lambda,\widetilde a,\widetilde K)\) with the same Lipschitz bound \(L\), \begin{equation}\label{eq:mz-control-stability} \begin{aligned} \|\mathcal T_\lambda^aK(t,\cdot) -\mathcal T_{\widetilde\lambda}^{\widetilde a} \widetilde K(t,\cdot)\|_\infty \le{}&\|K-\widetilde K\|_\infty +\frac32L^2\|\lambda-\widetilde\lambda\|_{L^1(t,1)}\\ &+L\sqrt{\frac{2(1-t)}{\pi}}\, |\sqrt a-\sqrt{\widetilde a}|. \end{aligned} \end{equation} For step \(\lambda\) and \(a>0\), the value is the backward Cole--Hopf/heat recursion solving, between the step endpoints, \[ \partial_t f+\frac a2f_{xx}+\frac{\lambda(t)}2f_x^2=0, \qquad f(1,\cdot)=K. \] For general \(\lambda\), it is the uniform limit on \([0,1]\times\mathbb R\) of these values for any nonnegative step approximation in \(L^1\). This approximation-independent extension is the \emph{stochastic-control solution}. It satisfies dynamic programming at every deterministic intermediate time. At zero diffusion, \begin{equation}\label{eq:mz-zero-diffusion} \mathcal T_\lambda^0K(t,x) =\mathsf H_{\int_t^1\lambda(s)\,\dd s}K(x), \qquad \mathsf H_0K=K. \end{equation} If \((\theta,x)\mapsto K_\theta(x)\) is jointly Borel measurable on \(\mathbb R^2\), with common Lipschitz bound \(L\), then \((\theta,t,x)\mapsto\mathcal T_\lambda^aK_\theta(t,x)\) is jointly Borel measurable. For any Borel probability measure \(\nu\) on \(\mathbb R\) satisfying \(\int_{\mathbb R}|K_\theta(0)|\,\nu(\dd\theta)<\infty\), these values are integrable at every \((t,x)\), and the preceding comparison bounds may be integrated over the mark. \end{lemma} \begin{proof} \emph{Clipping and comparison.} Weighted Cauchy--Schwarz makes the drift of every admissible control integrable. Since \(|K(x)|\le|K(0)|+L|x|\), its terminal payoff and cost are integrable. For \(w=(-L)\vee(v\wedge L)\), the terminal payoff lost by clipping is at most \(L\int_t^1\lambda|v-w|\), whereas \[ \frac12\int_t^1\lambda(v^2-w^2) \ge L\int_t^1\lambda|v-w|. \] Thus clipping improves the payoff pathwise. The zero control gives the lower growth bound. The upper bound follows from \(L|v|-v^2/2\le L^2/2\), proving \eqref{eq:mz-control-growth}. For bounded terminals the zero control and nonnegative cost give the sup-norm bound. Comparing translations and terminals gives the Lipschitz bound, order preservation, and contraction. For the same clipped control, changing \(\lambda\) changes the terminal expectation by at most \(L^2\|\lambda-\widetilde\lambda\|_1\), and the cost by at most half this amount. Changing \(a\) changes only the Brownian displacement, with expected magnitude \(\sqrt{2(1-t)/\pi}\,|\sqrt a-\sqrt{\widetilde a}|\). Taking suprema proves \eqref{eq:mz-control-stability}. All these arguments also hold on any subinterval. \emph{Step coefficients.} For \(a>0\), on an interval of length \(\Delta\) where \(\lambda=m>0\), the backward operator is \[ F(x)\longmapsto \frac a m\log\E\exp\left\{\frac m a F(x+\sqrt{a\Delta}\,G)\right\}, \qquad G\sim N(0,1). \] For \(m=0\), use \(\E F(x+\sqrt{a\Delta}\,G)\). Mollify \(K\) uniformly, preserving its Lipschitz bound. The resulting recursion solves the PDE, has \(|f_x|\le L\), and has bounded \(f_{xx}\) for each fixed mollification. For \(X_s=x+\sqrt a(B_s-B_t)+\int_t^s\lambda(u)v_u\,\dd u\), apply It\^o's formula on the successive intervals, stopping first when \(|X_s|\) exceeds a finite level. For clipped controls, linear growth of \(f\), bounded \(f_x\), and the integrable Brownian maximum justify removing the stopping. This gives \[ \E\left[K(X_1)-\frac12\int_t^1\lambda(s)v_s^2\,\dd s\right] =f(t,x)-\frac12\E\int_t^1 \lambda(s)\bigl(v_s-f_x(s,X_s)\bigr)^2\,\dd s. \] Here \(K\) denotes the mollified terminal. The feedback \(v_s=f_x(s,X_s)\) has a bounded Lipschitz drift and gives equality. Thus the recursion equals the control value. Applying the same verification on subintervals gives dynamic programming. Terminal contraction removes the mollification. At \(a=0\), put \(A=\int_t^1\lambda\). If \(A>0\), the displacement \(z=\int_t^1\lambda v\) satisfies \(\int_t^1\lambda v^2\ge z^2/A\) pathwise; constant controls \(v=z/A\) give the reverse inequality after taking the supremum over deterministic \(z\). This proves \eqref{eq:mz-zero-diffusion}. If \(A=0\), the value is \(K\). Splitting the quadratic cost between two intervals gives the Hopf--Lax semigroup identity and hence dynamic programming. \emph{Integrable coefficients and marks.} Take nonnegative step functions \(\lambda_n\to\lambda\) in \(L^1\); interval averages preserve monotonicity when \(\lambda\) is nondecreasing. Estimate \eqref{eq:mz-control-stability} gives uniform convergence to the control value, independently of the approximation. The same estimate on subintervals and terminal contraction pass dynamic programming to the limit. For marked terminals, the finite-step Gaussian recursions are jointly Borel measurable in \((\theta,t,x)\). At zero diffusion the Hopf--Lax supremum can be taken over rational displacements. The uniform limits therefore establish measurability for integrable coefficients. Clipped payoffs are bounded in absolute value by \[ |K_\theta(0)|+L|x|+L\sqrt a\,|B_1-B_t| +\tfrac32L^2\|\lambda\|_1. \] The assumed first moment makes this bound integrable over marks and Brownian paths. Fubini and dominated convergence therefore apply; no measurable choice of an optimizing control is needed. \end{proof} Write \(\mathcal T_\lambda=\mathcal T_\lambda^1\). For \(H\in C_b(\mathbb R)\) with \(\|H\|_\infty\le B\), maximizing displacements in \(\mathsf H_cH\) have magnitude at most \(2\sqrt{Bc}\). Comparison using an optimizer gives a spatial increment bound \(2\sqrt{B/c}|h|+h^2/(2c)\); subdivision removes the quadratic term. Thus \(\mathsf H_cH\) is bounded by \(B\) and \(2\sqrt{B/c}\)-Lipschitz. More generally, if \(H\) is \(L\)-Lipschitz, its Hopf--Lax transform is finite, \(L\)-Lipschitz, and \(H\le\mathsf H_cH\le H+cL^2/2\). If \((\theta,x)\mapsto H_\theta(x)\) is jointly Borel measurable on \(\mathbb R^2\), with each \(H_\theta\) continuous, the transform is jointly Borel measurable by its rational-displacement supremum. Consequently \cref{lem:mz-control} defines \(f_{\mu,c}^{H_\theta}=\mathcal T_\mu(\mathsf H_cH_\theta)\), with the mark expectation taken after this scalar evolution. \begin{lemma}[Zero-temperature variational comparison] \label{lem:mz-variational-comparison} Fix \(\alpha>0\), a nonempty parameter set \(\Lambda\), \(b:\Lambda\to\mathbb R\), and a Borel probability measure \(\nu\) on \(\mathbb R\). For each \(\lambda\in\Lambda\), let \((\theta,z)\mapsto H_{\theta,\lambda}(z)\) be jointly Borel measurable on \(\mathbb R^2\), and assume \[ \sup_{\theta\in\mathbb R,\,\lambda\in\Lambda}\Lip(H_{\theta,\lambda})\le L, \qquad \int_{\mathbb R}|H_{\theta,\lambda}(0)|\,\nu(\dd\theta)<\infty \quad\text{for every }\lambda. \] Let \(\mathscr C\) be the nondecreasing right-continuous functions \(\gamma:[0,1]\to[0,1]\) equal to one on some terminal interval, and write \(D_\gamma(t)=\int_t^1\gamma(s)\,\dd s\). Define \begin{equation}\label{eq:mz-variational-objectives} \begin{aligned} \mathcal J_\beta(\lambda,\gamma) &=\E\mathcal T_{\beta\gamma}H_{\Theta,\lambda}(0,0) +\frac1{2\alpha\beta}\int_0^1 \left\{\frac1{D_\gamma(t)}-\frac1{1-t}\right\}\dd t+b(\lambda),\\ \mathcal F(\lambda,\mu,c) &=\E\mathcal T_\mu(\mathsf H_cH_{\Theta,\lambda})(0,0) +\frac1{2\alpha}\int_0^1\frac{\dd t}{c+\int_t^1\mu(s)\,\dd s} +b(\lambda). \end{aligned} \end{equation} For every fixed \((\lambda,\mu,c)\in\Lambda\times\mathscr U\times(0,\infty)\), there are \(\gamma_\beta\in\mathscr C\) such that \(\mathcal J_\beta(\lambda,\gamma_\beta)\to\mathcal F(\lambda,\mu,c)\). Conversely, for every \(\beta\ge2\) and \(\gamma\in\mathscr C\), \(\mathcal F(\lambda,\beta\gamma,\beta^{-1}) \le\mathcal J_\beta(\lambda,\gamma) +C_{\alpha,L}(1+\log\beta)/\beta\), uniformly in \(\lambda\) and the mark law. Consequently, in \(\mathbb R\cup\{-\infty\}\), \begin{equation*} \lim_{\beta\to\infty}\inf_{\lambda\in\Lambda,\,\gamma\in\mathscr C} \mathcal J_\beta(\lambda,\gamma) =\inf_{\lambda\in\Lambda,\,\mu\in\mathscr U,\,c>0} \mathcal F(\lambda,\mu,c). \end{equation*} \end{lemma} \begin{proof} \emph{Uniform reverse comparison.} For \(c>0\), use the Hopf--Lax bound \(H\le\mathsf H_cH\le H+cL^2/2\) and the elementary inequality \[ \frac1{c+x}-\frac1x+\frac1y\le\frac1{c+y} \qquad(0c\), set \(T=1-c/\beta\), \(\mu_\beta=\mu\wedge\beta\), and \(\gamma_\beta=\mu_\beta/\beta\) below \(T\), and one above. This is admissible, and its entropy is \begin{equation*} \frac1{2\alpha\beta}\int_0^1 \left\{\frac1{D_{\gamma_\beta}(t)}-\frac1{1-t}\right\}\dd t =\frac1{2\alpha}\int_0^T \frac{\dd t}{c+\int_t^T\mu_\beta(s)\dd s} +\frac{\log(c/\beta)}{2\alpha\beta}. \end{equation*} It converges to the entropy in \(\mathcal F\), since \(\mu_\beta\to\mu\) in \(L^1\) and the denominators are at least \(c\). For the channel, split at \(T\) and apply \cref{lem:mz-control}: remove diffusion from the terminal interval, whose drift mass is \(c\); replace \(\mu_\beta\) by \(\mu\) below \(T\); and restore the omitted \(\mu\)-evolution on \([T,1]\). Terminal contraction bounds the total error by \[ 2L\E|G|\sqrt{c/\beta} +\frac32L^2\|\mu_\beta-\mu\|_1 +\frac{L^2}{2}\int_T^1\mu(s)\dd s\longrightarrow0. \] This proves recovery uniformly in the mark. Taking the infimum over fixed candidates proves the upper limit, also when it is \(-\infty\). \end{proof} \begin{proposition}[Marked spherical ground states] \label{prop:marked-ground-state} Let \(\nu\) be a Borel probability measure on \(\mathbb R\), and let \((\theta,x)\mapsto H_\theta(x)\) be jointly Borel measurable on \(\mathbb R^2\), with \(\sup_{\theta\in\mathbb R}\|H_\theta\|_\infty\le B\) and \(\sup_{\theta\in\mathbb R}\Lip(H_\theta)\le L\). For i.i.d.\ marks \(\Theta_k\sim\nu\), independent of the Gaussian rows and \(M_N/N\to\alpha\in(0,\infty)\), almost surely, \begin{equation*} \begin{aligned} E_N&:=\max_{\|v\|_2=1}\frac1{M_N}\sum_{k\le M_N}H_{\Theta_k}(g_k\cdot v) \longrightarrow \inf_{\mu\in\mathscr U,\ c>0}\mathcal F_\nu(\mu,c),\\ \mathcal F_\nu(\mu,c)&:=\E_\nu f_{\mu,c}^{H_\Theta}(0,0) +\frac1{2\alpha}\int_0^1\frac{\dd t}{c+\int_t^1\mu(s)\,\dd s}. \end{aligned} \end{equation*} Convergence also holds in every finite \(L^p\). \end{proposition} \begin{proof} \emph{Finite temperature.} The class \(\mathscr C\) in \cref{lem:mz-variational-comparison} consists precisely of the distribution functions \(\gamma_q\) with \(q(1-)<1\), up to endpoint conventions. This restriction does not change the finite-temperature infimum: replacing \(q\) by \(q\wedge(1-\varepsilon)\) on \([0,1)\), retaining \(q(1)=1\), decreases the spherical entropy, while \eqref{rpc:Phi-L1}, applied to \(\beta H_\Theta\), changes its \(\beta^{-1}\)-scaled channel term by at most \(\beta L^2\varepsilon/2\). Let \(\varepsilon\downarrow0\). The recursion \eqref{rpc:channel-semigroup} and \cref{lem:mz-control} give \(\beta^{-1}\Phi_{\beta H}(q)=\mathcal T_{\beta\gamma_q}H(0,0)\), first on finite grids and then by approximation. Consequently \cref{thm:iid-channels} and \eqref{eq:spherical-entropy-classical} give, for the partition function with channels \(\beta H_{\Theta_k}\), \begin{equation}\label{eq:mz-finite-temperature} \lim_N\frac{\log Z_N(\beta)}{\beta M_N} =\inf_{\gamma\in\mathscr C}\mathcal J_\beta(\gamma), \end{equation} where \(\mathcal J_\beta\) is \eqref{eq:mz-variational-objectives} with a singleton parameter set and zero penalty. The limit holds almost surely, simultaneously for integer \(\beta\). By \cref{lem:mz-variational-comparison}, \(\lim_{\beta\to\infty}\inf_\gamma\mathcal J_\beta(\gamma) =\inf_{\mu,c}\mathcal F_\nu(\mu,c)\). \emph{From pressure to the maximum.} The objective defining \(E_N\) has Lipschitz constant at most \(L\|\mathbf G_N\|_{\rm op}/\sqrt{M_N}\). The cap estimate used in \eqref{eq:spherical-zero-temperature-comparison} therefore gives, for \(0<\eta\le1\), \[ 0\le E_N-\frac{\log Z_N(\beta)}{\beta M_N} \le\frac{L\|\mathbf G_N\|_{\rm op}}{\sqrt{M_N}}\eta +\frac{N}{\beta M_N}\log\frac8\eta. \] Gaussian operator-norm tails imply that the first coefficient is eventually bounded almost surely. Take \(\eta=\beta^{-1}\), let \(N\to\infty\) on the common event in \eqref{eq:mz-finite-temperature}, then let integer \(\beta\to\infty\). This proves the almost-sure limit. Boundedness gives every finite \(L^p\) convergence. \end{proof} \begin{proof}[Proof of \cref{thm:montanari-zhou}] First suppose \(h\) is bounded, jointly Borel measurable, and uniformly \(L\)-Lipschitz in its prediction coordinate. Rotate the teacher coordinates to write \(X_i=(G_i,Z_i)\), where \(G_i\sim N(0,I_k)\) and \(Z_i\sim N(0,I_{d-k})\) are independent, and \(Y_i=\varphi(G_i,\varepsilon_i)\). For \(r\in B_k(1)\), set \(s_r=\sqrt{1-|r|^2}\) and \[ H_{r,(G,Y)}(z)=h(Y,r\cdot G+s_rz),\qquad M_{n,d}(r)=\max_{\|v\|_2=1}\frac1n\sum_i H_{r,(G_i,Y_i)}(Z_i\cdot v). \] These channels are uniformly bounded and \(L\)-Lipschitz, and their marks are independent of the residual rows. Identify \((G,Y)\in\mathbb R^{k+1}\) with a real label through a fixed Borel isomorphism, the same for every \(r\); the resulting channels are jointly Borel measurable with the same bounds. Since \(n/(d-k)\to\alpha\), \cref{prop:marked-ground-state} gives \(M_{n,d}(r)\to V(r)\) almost surely at each fixed \(r\), with \(V(r)\) given by its marked functional. To maximize over the teacher overlap, we need uniform convergence. Comparing the objectives on the same disorder gives \begin{equation*} |M_{n,d}(r)-M_{n,d}(r')| \le L|r-r'|\left(\frac1n\sum_i|G_i|^2\right)^{1/2} +L|s_r-s_{r'}|\frac{\|Z\|_{\rm op}}{\sqrt n}. \end{equation*} The two random coefficients are eventually bounded almost surely, and \(|s_r-s_{r'}|\le\sqrt{2|r-r'|}\). The deterministic limit inherits this modulus. Intersect this probability-one event with the convergence events for a countable dense set of overlaps. Finite nets then give uniform convergence on \(B_k(1)\). The original maximum is exactly \(\max_r M_{n,d}(r)\), so its limit is \(\sup_r V(r)\). For \(a=|r|^2<1\), change variables in the marked functional by \[ t_{\rm new}=a+(1-a)t,\qquad x_{\rm new}=r\cdot G+\sqrt{1-a}\,x, \qquad c_{\rm new}=(1-a)c, \qquad \mu_{\rm new}(t)=\mu\left(\frac{t-a}{1-a}\right). \] The last formula applies on \([a,1)\), with zero extension below \(a\). For step coefficients it transforms both the scalar evolution and its reciprocal-integral term into \eqref{eq:mz-supervised-functional}; \eqref{eq:mz-control-stability} extends the identity to integrable coefficients. Conversely every admissible coefficient restricted to \([a,1)\) has an admissible inverse under this change, so the infima agree. At \(|r|=1\) the channel is constant in its residual argument and its value is \(\E h(Y,r\cdot G)\). In \eqref{eq:mz-supervised-functional} the entropy interval is then empty; letting \(c\downarrow0\) gives the same value. Applying \cref{prop:marked-ground-state} directly to a deterministic channel gives the unlabelled assertion. Finally let \(h\in C_b(\R^2)\), \(\|h\|_\infty\le B\), and regularize only the prediction coordinate: \[ h_L^-(y,z)=\inf_w\{h(y,w)+L|z-w|\},\qquad h_L^+(y,z)=\sup_w\{h(y,w)-L|z-w|\}. \] These functions satisfy \(-B\le h_L^-\le h\le h_L^+\le B\) and are uniformly \(L\)-Lipschitz in \(z\). Taking the extrema over rational \(w\) proves joint measurability. For each fixed \(R,y\), continuity of \(h(y,\cdot)\) gives \[ \eta_L(R,y):=\sup_{|z|\le R}(h_L^+-h_L^-)(y,z)\longrightarrow0, \qquad 0\le\eta_L(R,y)\le2B. \] The supremum can again be taken over rational points, so \(\eta_L\) is measurable. Writing \(X\) for the full covariate matrix, \begin{equation*} \sup_{\|w\|_2=1}\frac1n\sum_i (h_L^+-h_L^-)(Y_i,w\cdot X_i) \le\frac1n\sum_i\eta_L(R,Y_i) +\frac{2B\|X\|_{\rm op}^2}{nR^2}. \end{equation*} For each integer \(L,R\), bounded-variable concentration and Borel--Cantelli give convergence of the empirical first term to \(\E\eta_L(R,Y)\), also for triangular arrays. Dominated convergence makes this expectation tend to zero as \(L\to\infty\); Gaussian operator-norm control bounds the other coefficient almost surely. Thus the difference of the two envelope optimization limits tends to zero, first as \(L\to\infty\), then \(R\to\infty\). The functional for \(h\) is well defined by \cref{lem:mz-control} and the Hopf--Lax terminal bounds. Its terminal transform, scalar evolution, infimum, and outer supremum preserve order, so its value lies between the two envelope values. The same squeeze proves convergence of the empirical maximum to this value. The one-variable envelopes give the unlabelled case. Boundedness gives convergence in every finite \(L^p\). \end{proof} \subsection{Radial passage for Gaussian spins} \label{sec:gaussian-spins} The following lemma justifies the radial optimization used in \cref{thm:st-free energy}. \begin{lemma}[Radial Laplace passage] \label{lem:st-radial-laplace} Let \(u,h,c\) and \(M_N\) satisfy the assumptions of \cref{thm:st-free energy}. On the same disorder, define \begin{align*} Y_N(s)&:=\frac1N\log Z_{N,M_N}^{\mathrm S,u_{s^2},hs}, \\ Y(s)&:=\inf_{q\in\Q}\left\{ \mathcal A_{\rm S}(q)+\alpha\Phi_{u_{s^2}}(q) +\frac{h^2s^2}{2}\mathsf L(q)\right\}. \end{align*} Then \begin{equation}\label{eq:st-quenched-laplace-principle} \frac1N\log\int_0^\infty s^{N-1}e^{N[Y_N(s)-cs^2]}\,\dd s \xrightarrow{L^1}\sup_{s>0}\{Y(s)+\log s-cs^2\}. \end{equation} The supremum is finite. \end{lemma} \begin{proof} \emph{Shell convergence.} For each \(s>0\), the \hyperref[par:gaussian-field-extension]{Gaussian-field extension} and \eqref{eq:gaussian-field-variance} give \(Y_N(s)\to Y(s)\) in \(L^2\). Writing \(L=\Lip(u)\), \(u_0=u(0)\), and \(\alpha_N=M_N/N\), comparison of the Gibbs weights gives \begin{align} |Y_N(s)-Y_N(t)|&\le D_N|s-t|, &D_N&=L\sqrt{\alpha_N}\frac{\|\mathbf G_N\|_{\rm op}}{\sqrt N} +|h|\frac{\|g^0\|_2}{\sqrt N}, \notag\\ |Y_N(s)-\alpha_Nu_0|&\le sD_N, &&s,t\ge0. \label{eq:st-shell-linear-envelope} \end{align} Here \(Y_N(0)=\alpha_Nu_0\), and \cref{lem:spectral-moments} and Gaussian norm moments give \(\sup_N\E D_N^p<\infty\) for every finite \(p\). The expected first bound makes \(Y\) locally Lipschitz; finite nets then give \begin{equation}\label{eq:st-local-uniform-shell-limit} \sup_{s\in[a,A]}|Y_N(s)-Y(s)|\longrightarrow0 \quad\text{in probability and in }L^1, \qquad 00\), and put \begin{equation*} \ell_\theta(x)=(\theta-x)_+, \qquad u_{\beta,\theta}^{\mathrm{har}}(x) =-\frac{\beta}{2}\ell_\theta(x)^2. \end{equation*} For \(R>0\), its capped Lipschitz approximation is \begin{equation*} u_{\beta,\theta,R}(x) =-\frac{\beta}{2}\bigl(\ell_\theta(x)\wedge R\bigr)^2. \end{equation*} Define \begin{align*} \Phi_{\beta,\theta}^{\mathrm{har}}(q) &=\lim_{R\to\infty}\Phi_{u_{\beta,\theta,R}}(q), \\ \mathcal P_{\kappa,\beta,\theta}^{\mathrm{har}}(\alpha) &=\inf_{q\in\Q}\left\{ \mathcal A_\kappa(q)+\alpha\Phi_{\beta,\theta}^{\mathrm{har}}(q) \right\}, \\ F_{N,\kappa}^{\mathrm{har}}(\beta,\theta) &=\frac1N\log\int_{\Sigma_N^\kappa} \exp\left\{-\frac\beta2\sum_{k=1}^{M_N} (\theta-h_k(\sigma))_+^2\right\}\mu_N^\kappa(\dd\sigma), \\ e_{N,\kappa}(\theta) &=\frac1{2N}\min_{\sigma\in\Sigma_N^\kappa} \sum_{k=1}^{M_N}(\theta-h_k(\sigma))_+^2. \end{align*} \begin{theorem}[Quadratic free energy and ground state] \label{thm:harmonic-free energy-ground-state} Let \(M_N/N\to\alpha\in(0,\infty)\), \(\theta\in\mathbb R\), and \(\kappa\in\{\mathrm S,\mathrm I\}\). For every \(\beta>0\), \begin{equation} F_{N,\kappa}^{\mathrm{har}}(\beta,\theta) \longrightarrow \mathcal P_{\kappa,\beta,\theta}^{\mathrm{har}}(\alpha) \qquad\text{in }L^1, \label{eq:harmonic-free energy-limit} \end{equation} and \begin{equation} \mathcal P_{\kappa,\beta,\theta}^{\mathrm{har}}(\alpha) =\lim_{R\to\infty}\mathcal P_{\kappa,u_{\beta,\theta,R}}(\alpha). \label{eq:harmonic-variational-truncation} \end{equation} There is a deterministic \(e_{\mathrm{gs},\kappa}(\alpha,\theta)\ge0\) such that \begin{equation} e_{N,\kappa}(\theta)\longrightarrow e_{\mathrm{gs},\kappa}(\alpha,\theta) \qquad\text{in }L^1, \label{eq:quadratic-ground-state-limit} \end{equation} and \begin{equation} e_{\mathrm{gs},\kappa}(\alpha,\theta) =\lim_{\beta\to\infty} -\frac1\beta\mathcal P_{\kappa,\beta,\theta}^{\mathrm{har}}(\alpha). \label{eq:quadratic-zero-temperature-identity} \end{equation} The proof first sends \(N\to\infty\) at fixed \((\beta,R)\), then \(R\to\infty\) at fixed \(\beta\), and finally \(\beta\to\infty\). \end{theorem} \begin{proof}[Proof of Theorem~\ref{thm:harmonic-free energy-ground-state}] \emph{Finite temperature.} Put \[ d_R(x)=u_{\beta,\theta,R}(x)-u_{\beta,\theta}^{\mathrm{har}}(x) =\frac\beta2\bigl(\ell_\theta(x)^2-R^2\bigr)_+, \qquad \varepsilon_R=\E d_R(Z)\longrightarrow0, \quad Z\sim N(0,1). \] We first compare one row. Let \(\Gamma\) be a possibly random probability measure under which \(X\) has annealed standard-Gaussian marginal. Write \(\mathcal N_R=\Gamma(e^{u_{\beta,\theta,R}(X)})\), \(\mathcal N_\infty=\Gamma(e^{u_{\beta,\theta}^{\rm har}(X)})\), and \(\Gamma_R=\mathcal N_R^{-1}e^{u_{\beta,\theta,R}(X)}\Gamma\). Jensen gives \[ 0\le\log\mathcal N_R-\log\mathcal N_\infty =-\log\Gamma_R(e^{-d_R(X)}) \le\Gamma_R(d_R(X))\le\Gamma(d_R(X)). \] The last inequality follows because \(d_R\) increases and \(e^{u_{\beta,\theta,R}}\) decreases as functions of \(\ell_\theta^2\). Consequently \begin{equation}\label{eq:harmonic-one-row-capped-comparison} 0\le\E(\log\mathcal N_R-\log\mathcal N_\infty) \le\varepsilon_R. \end{equation} For the finite-volume capped pressure \(F_{N,\kappa}^R\), replace the rows successively. The normalized measure with the current row deleted is independent of that row, whose marginal at each spin is standard normal. Thus \eqref{eq:harmonic-one-row-capped-comparison} gives \begin{equation}\label{eq:harmonic-finite-volume-telescope} 0\le\E\bigl[F_{N,\kappa}^R(\beta,\theta) -F_{N,\kappa}^{\rm har}(\beta,\theta)\bigr] \le\frac{M_N}{N}\varepsilon_R, \end{equation} with a pointwise nonnegative difference. For a step path \(q\), the completed field on its finite cascade has the same standard-Gaussian marginal under the averaged unweighted measure. Jensen therefore gives \[ -\frac\beta2\E(\theta-Z)_+^2 \le\Phi_{u_{\beta,\theta,R}}(q)\le0. \] At fixed \(\beta,R\), the capped channel is Lipschitz, so \eqref{rpc:Phi-L1} extends these bounds to every \(q\in\Q\). They then persist in the decreasing limit \(\Phi_{\beta,\theta}^{\rm har}(q)\). Commuting \(\inf_R\) and \(\inf_q\) on \(\{q:\mathcal A_\kappa(q)<\infty\}\) therefore gives \[ \mathcal P_{\kappa,u_{\beta,\theta,R}}(\alpha) \downarrow\inf_{q:\mathcal A_\kappa(q)<\infty} \{\mathcal A_\kappa(q)+\alpha\Phi_{\beta,\theta}^{\rm har}(q)\} =\mathcal P_{\kappa,\beta,\theta}^{\rm har}(\alpha). \] The values are finite by the Jensen bound, \(\mathcal A_\kappa\ge0\), and the trial path \(\overline q_0\). This proves \eqref{eq:harmonic-variational-truncation}. At fixed \(R\), \cref{thm:iid-channels} gives convergence of \(F_{N,\kappa}^R\) to its variational value in \(L^2\). The triangle inequality and \eqref{eq:harmonic-finite-volume-telescope}, followed by \(R\to\infty\), prove \eqref{eq:harmonic-free energy-limit}. \emph{Zero temperature.} Set \(X_{N,\kappa}(\beta)=-F_{N,\kappa}^{\rm har}(\beta,\theta)/\beta\). Since the prior is normalized, H\"older's inequality makes this nonincreasing in \(\beta\), with \(X_{N,\kappa}(\beta)\ge e_{N,\kappa}(\theta)\ge0\). A minimizing Ising atom has mass \(2^{-N}\), so \begin{equation*} 0\le X_{N,\mathrm I}(\beta)-e_{N,\mathrm I}(\theta) \le\frac{\log2}{\beta}. \end{equation*} For the spherical prior write \(x=\sigma/\sqrt N\), put \(\theta_+=\max\{\theta,0\}\), and set \[ \mathscr L_{N,\theta}(x)=\frac1{2N} \|(\theta\mathbf1-\mathbf G_Nx)_+\|_2^2, \qquad L_{N,\theta}=\frac{\|\mathbf G_N\|_{\rm op}}{\sqrt N} \left(\theta_+\sqrt{\frac{M_N}{N}} +\frac{\|\mathbf G_N\|_{\rm op}}{\sqrt N}\right). \] Contraction of the positive-part map and \(\|(\theta\mathbf1-\mathbf G_Nx)_+\|_2 \le\theta_+\sqrt{M_N}+\|\mathbf G_N\|_{\rm op}\) show that \(\mathscr L_{N,\theta}\) is \(L_{N,\theta}\)-Lipschitz on the unit sphere. For \(N\ge2\) and \(00$, \(M_N\le\bar\alpha N\) for all large $N$. Assume that, for some $c_0>0$ with $c_0<\log2$ when $\kappa=\mathrm I$ and for every fixed $\beta>0$, \begin{equation}\label{eq:hard-soft-lower-input} \Pp\bigl(F_{N,\kappa}^{\beta,C}<-c_0\bigr)\longrightarrow0. \end{equation} Then, for every $\varepsilon>0$, \begin{equation*} \lim_{\beta\to\infty}\limsup_{N\to\infty} \Pp\left(F_{N,\kappa}^{\beta,C} -\operatorname{Ent}_{N,\kappa}(C)>\varepsilon\right)=0. \end{equation*} \end{proposition} \begin{proof} \emph{Single-constraint claim.} Fix $\rho,\zeta>0$ and a tail-accessible closed set $C$, and choose $\theta_C\in\R$ such that $[\theta_C,\infty)\subset C$. There are constants $b_{\rm row},c,A,w_0>0$ and $N_0<\infty$, depending only on $(\rho,\zeta,C,\theta_C)$, such that the following holds for every $N\ge N_0$. Let $\nu$ be any deterministic probability measure on a subset $\Omega_N\subset S_N$ satisfying \begin{equation}\label{eq:hard-small-ball-hypothesis} \sup_{\sigma_0\in\Omega_N} \nu\left\{\sigma:\|\sigma-\sigma_0\|_2\le\rho\sqrt N\right\} \le e^{-\zeta N}. \end{equation} For $g\sim\mathsf N_N$, set \begin{align*} X_g(\sigma)&=N^{-1/2}g\cdot\sigma,\\ p_{\nu,C}(g)&=\nu\{X_g\in C\},\\ s_{\nu,\beta,C}(g)&=\int e^{-\beta d_C(X_g(\sigma))}\nu(\dd\sigma),\\ D_{\nu,\beta,C}(g)&=\log\frac{s_{\nu,\beta,C}(g)}{p_{\nu,C}(g)}, \end{align*} We suppress the common argument \(g\) in the remainder of this claim and its proof. Set $D_{\nu,\beta,C}=+\infty$ when $p_{\nu,C}=0$. Uniformly for $w_0\le w\le b_{\rm row}N$, \begin{equation}\label{eq:hard-one-row-tail} \Pp_g\bigl(p_{\nu,C}(g)\le e^{-w}\bigr) \le Ae^{-cw}+Ae^{-cN}. \end{equation} Moreover, there is a deterministic function $\omega_{\rho,\zeta,C}(\beta)\downarrow0$ as $\beta\to\infty$ such that, for each fixed $\beta$ and all sufficiently large $N$, \begin{equation}\label{eq:hard-one-row-mean} \E_g(D_{\nu,\beta,C}\wedge b_{\rm row}N) \le\omega_{\rho,\zeta,C}(\beta)+ANe^{-cN}. \end{equation} All constants are uniform over $\nu$ satisfying \eqref{eq:hard-small-ball-hypothesis}. The same conclusions hold conditionally if $\nu$ is random, independent of $g$, and \eqref{eq:hard-small-ball-hypothesis} holds almost surely. \emph{Proof of the claim.} Decrease $b_{\rm row}$ so that $2b_{\rm row}<\zeta$. Given $w\in[w_0,b_{\rm row}N]$, let $K=\lfloor e^w/4\rfloor$ and sample $\sigma^1,\ldots,\sigma^K$ independently from $\nu$. A union bound and \eqref{eq:hard-small-ball-hypothesis} show that the $\nu^{\otimes K}$-probability of a pair within distance $\rho\sqrt N$ is at most $K^2e^{-\zeta N}\le e^{-cN}$. Conditional on pairwise separation, the centered Gaussian variables $X_g(\sigma^i)$ have variance one and canonical distances at least $\rho$. Sudakov minoration gives $\E_g\max_{i\le K}X_g(\sigma^i)\ge c_\rho\sqrt{\log K}$, and the Gaussian concentration inequality for the maximum then gives, after increasing $w_0$, \begin{equation}\label{eq:hard-separated-sample} \Pp_g\left(\left.\max_{i\le K}X_g(\sigma^i)<\theta_C \,\right|\,\sigma^1,\ldots,\sigma^K\right)\le e^{-cw} \end{equation} for every pairwise separated realization $(\sigma^1,\ldots,\sigma^K)$. On the other hand, on $\{p_{\nu,C}(g)\le e^{-w}\}$, conditional independence of the samples gives \[ \Pp\left(X_g(\sigma^i)\notin C\text{ for every }i\le K\mid g\right) =(1-p_{\nu,C}(g))^K\ge\frac34 \] once $w_0$ is large. Averaging first in the samples and then in $g$, and noting that $x\notin C$ implies $x<\theta_C$, combining this inequality with \eqref{eq:hard-separated-sample} proves \eqref{eq:hard-one-row-tail}. Write $\Delta_{\nu,\beta,C}(g)=s_{\nu,\beta,C}(g)-p_{\nu,C}(g)$ and $Y_{\nu,C}(g)=-\log p_{\nu,C}(g)$, with $Y_{\nu,C}=+\infty$ when $p_{\nu,C}=0$. Since $X_g(\sigma)$ is standard Gaussian for every $\sigma\in S_N$, Fubini gives the uniform estimate \begin{equation}\label{eq:hard-soft-excess-mean} \E_g \Delta_{\nu,\beta,C} =\varepsilon_C(\beta) :=\E\left[e^{-\beta d_C(Z)}\1_{\{Z\notin C\}}\right] \longrightarrow0. \end{equation} Here dominated convergence applies because closedness gives $d_C(Z)>0$ on $\{Z\notin C\}$. For $w\in[w_0,b_{\rm row}N)$, \begin{equation}\label{eq:hard-one-row-split} D_{\nu,\beta,C}\wedge b_{\rm row}N \le e^w \Delta_{\nu,\beta,C} +(Y_{\nu,C}\wedge b_{\rm row}N)\1_{\{Y_{\nu,C}>w\}}. \end{equation} Indeed, on $\{Y_{\nu,C}\le w\}$ use $\log(1+\Delta_{\nu,\beta,C}/p_{\nu,C}) \le \Delta_{\nu,\beta,C}/p_{\nu,C} \le e^w\Delta_{\nu,\beta,C}$, while on $\{Y_{\nu,C}>w\}$ use $s_{\nu,\beta,C}\le1$, hence $D_{\nu,\beta,C}\le Y_{\nu,C}$. Integrating \eqref{eq:hard-one-row-tail} yields \[ \E_g\bigl[(Y_{\nu,C}\wedge b_{\rm row}N) \1_{\{Y_{\nu,C}>w\}}\bigr] \le A(1+w)e^{-cw}+ANe^{-cN}. \] For large $\beta$, take \[ w_\beta=w_0\vee\frac12\log\frac1{\varepsilon_C(\beta)}; \] for fixed $\beta$ this lies below $b_{\rm row}N$ when $N$ is large. Then $e^{w_\beta}\varepsilon_C(\beta)\to0$ and $(1+w_\beta)e^{-cw_\beta}\to0$. Equations \eqref{eq:hard-soft-excess-mean} and \eqref{eq:hard-one-row-split} prove \eqref{eq:hard-one-row-mean} with some $\omega_{\rho,\zeta,C}(\beta)\downarrow0$. \emph{Sequential hardening.} Choose $c_1>c_0$, with $c_1<\log2$ when $\kappa=\mathrm I$. We first verify the spread hypothesis in the single-constraint claim. If a probability measure $\nu$ on $S_N$ satisfies $\dd\nu/\dd\mu_N^{\rm S}\le e^{c_1N}$, then a spherical cap estimate gives, for fixed $0<\rho<1$, \[ \sup_{\sigma_0\in S_N}\nu\{\|\sigma-\sigma_0\|_2\le\rho\sqrt N\} \le e^{c_1N}C_\rho\sqrt N\,\rho^{N-3}. \] In the spherical case, choose $\rho$ sufficiently small and then $\zeta>0$ so that this bound is at most $e^{-\zeta N}$ for all large $N$. On the cube, if $\max_{\sigma\in\Sigma_N^{\rm I}}\nu(\sigma)\le e^{-sN}$ for some $s>0$, a ball of radius $\rho\sqrt N$ contains at most \[ \exp\{N\mathsf h(\rho^2/4)+o(N)\} \] vertices, where \(\mathsf h(x)=-x\log x-(1-x)\log(1-x)\) is binary entropy. Choosing $\rho$ so that $\mathsf h(\rho^2/4)0$ in the Ising case. Fix the resulting pair $(\rho,\zeta)$ for the chosen prior. Write \[ U_k(\sigma)=e^{-\beta d_C(h_k(\sigma))}, \qquad I_k(\sigma)=\1_{\{h_k(\sigma)\in C\}}. \] For $0\le m\le M_N$, introduce the soft--hard hybrids \begin{equation*} Z_m=\int \left(\prod_{k\le m}I_k(\sigma)\right) \left(\prod_{m0$; otherwise set $\nu_m=\mu_N^\kappa$. Conditional on all rows except $g_m$, $\nu_m$ is independent of $g_m$. When $\widehat Z_m>0$, \begin{equation*} Z_{m-1}=\widehat Z_m s_{\nu_m,\beta,C}(g_m), \qquad Z_m=\widehat Z_m p_{\nu_m,C}(g_m). \end{equation*} Let $D_m=D_{\nu_m,\beta,C}(g_m)$. Whenever $Z_{m-1}>0$, $D_m=\log Z_{m-1}-\log Z_m\in[0,\infty]$. On $\{\widehat Z_m\ge e^{-c_1N}\}$ the unnormalized density defining $\nu_m$ is at most one. Hence \begin{align*} \frac{\dd\nu_m}{\dd\mu_N^{\rm S}}&\le e^{c_1N} &&(\kappa=\mathrm S),\\ \max_\sigma\nu_m(\sigma)&\le2^{-N}e^{c_1N} =e^{-(\log2-c_1)N}&&(\kappa=\mathrm I). \end{align*} The preceding spread estimates and the single-constraint claim, conditionally on all constraints except $g_m$, give constants $b_{\rm row},c,A>0$ and $\omega_{\rho,\zeta,C}(\beta)\downarrow0$, uniform in $m$, such that \begin{equation}\label{eq:hard-active-cost} \E\left[(D_m\wedge b_{\rm row}N) \1_{\{\widehat Z_m\ge e^{-c_1N}\}}\right] \le\omega_{\rho,\zeta,C}(\beta)+ANe^{-cN}. \end{equation} To telescope even when a hard constraint makes a partition function zero, truncate logarithms at a common floor. For \(z\ge0\), put \(\ell_N(z)=\log(z\vee e^{-c_1N})\) and \(L_m=\ell_N(Z_{m-1})-\ell_N(Z_m)\ge0\). Both hybrids lie below the floor when \(\widehat Z_m-c_1N\); hence both endpoint logs are unclipped and their difference is at most \(\eta N\). Markov's inequality gives \[ \Pp\left(F_{N,\kappa}^{\beta,C} -\operatorname{Ent}_{N,\kappa}(C)>\varepsilon\right) \le\Pp(Z_0 -\infty$ for spherical spins, or that $P_\infty> -\log2$ for Ising spins. Choose $c_0>-P_\infty$, with $c_0<\log2$ in the Ising case. Because $P_\beta\ge P_\infty$, the main free energy theorem at the fixed finite Lipschitz channel $u_{\beta,C}$ implies \eqref{eq:hard-soft-lower-input}. Apply \cref{prop:hard-entropy-preserving-hybrid}. Choose first $\beta$ with $P_\beta-P_\infty$ small and then send $N\to\infty$. The hybrid controls the lower tail of $\operatorname{Ent}_{N,\kappa}(C)$, while \eqref{eq:hard-below-soft} controls the upper tail. Thus $\operatorname{Ent}_{N,\kappa}(C)\to P_\infty$ in probability, proving the finite spherical case and the Ising entropy assertions. If $P_\infty=-\infty$ for spherical spins, fix $L>0$ and choose a finite $\beta$ with $P_\beta<-L-1$. The fixed-$\beta$ soft formula and \eqref{eq:hard-below-soft} give $\Pp(\operatorname{Ent}_{N,\mathrm S}(C)>-L)\to0$; then let $L\to\infty$. Finally suppose $P_\infty<-\log2$ for Ising spins. Choose $\eta>0$ and then a finite $\beta$ so that $P_\beta<-\log2-3\eta$. For all large $N$, \[ \E F_{N,\mathrm I}^{\beta,C} <-\log2-2\eta. \] On the satisfiable event, one configuration has zero penalty and hence $F_{N,\mathrm I}^{\beta,C}\ge-\log2$. Since $\Lip(u_{\beta,C})=\beta$, Gaussian concentration with row-matrix Lipschitz constant $\beta\sqrt{M_N}/N$ gives \[ \Pp\bigl(\mathcal F_{M_N,N}^{\mathrm I}(C)\ne\varnothing\bigr) \le2\exp\left\{-\frac{\eta^2N^2}{2M_N\beta^2}\right\}. \] As $M_N/N\to\alpha$, this proves \eqref{eq:hard-ising-exponential-unsat}. \end{proof} \subsection{Capacity consequences} \label{sec:ising-capacities} This subsection proves the general Ising capacity theorem, the spherical half-space threshold, and the explicit spherical trial bound summarized in \cref{sec:introduction}. \subsubsection{Ising capacities for tail-accessible sets} \begin{proof}[Proof of \cref{thm:closed-set-ising-capacity}] Fix a tail-accessible closed set \(C\) and \(\beta>0\). For a step path \(q\), choose a finite grid \(\pi\) containing its jump points. Jensen's inequality on the finite cascade gives \[ \Phi_{u_{\beta,C}}(q) \le\log\E\int e^{-\beta d_C(X^q(\eta,z))} \mathfrak R_\pi(\dd\eta)\mathsf N_1(\dd z) =\log\E e^{-\beta d_C(Z)}. \] The equality uses the standard-Gaussian marginal of the completed field under the averaged unweighted measure. At this fixed \(\beta\), \(\Lip(u_{\beta,C})=\beta\), so \eqref{rpc:Phi-L1} extends the inequality to every \(q\in\Q\) by step-path approximation. Only then let \(\beta\to\infty\); dominated convergence yields \begin{equation}\label{eq:ising-hard-annealed-bound} \Phi_{\infty,C}(q)\le\log\pi_C<0. \end{equation} Put \(B_C(q)=-\Phi_{\infty,C}(q)\), \(r_C=\alpha_c^{\rm I}(C)\), and \[ s_C(\alpha)=\log2+\mathcal P_{\mathrm I,\infty,C}(\alpha) =\inf_{\substack{q\in\Q:\ \mathcal A_{\rm I}(q)<\infty}} \{\log2+\mathcal A_{\rm I}(q)-\alpha B_C(q)\}, \qquad \alpha>0. \] If \(0<\alpha0. \] The last bound is uniform in \(q\), so \(s_C(\alpha)>0\). If \(\alpha>r_C\), some trial ratio is smaller than \(\alpha\), giving \(s_C(\alpha)<0\); if its denominator is infinite, the trial value is \(-\infty\). Thus \[ s_C(\alpha)>0\quad(0<\alphar_C). \] Together with \(\mathcal P_{\mathrm I,\infty,C}(0)=0\), these inequalities prove \eqref{eq:closed-set-ising-capacity-boundary}, including the possibility \(r_C=0\) at this stage. The path $\overline q_0$ has $\mathcal A_{\rm I}(\overline q_0)=0$. Its scalar recursion has no common Gaussian increment, and hence \[ \Phi_{u_{\beta,C}}(\overline q_0) =\log\E e^{-\beta d_C(Z)}\longrightarrow\log\pi_C. \] Therefore $B_C(\overline q_0)=-\log\pi_C$, giving the upper bound in \eqref{eq:closed-set-capacity-upper-bound}. Apply \cref{thm:closed-set-hardening} to any sequence \(M_N/N\to\alpha>0\). For \(\alpha0\) in probability, hence satisfiability with high probability. For \(\alpha>r_C\), it gives exponential decay of the satisfiability probability. This proves \eqref{eq:closed-set-ising-sat-unsat}. For strict positivity, choose \(\theta_C\) with \([\theta_C,\infty)\subseteq C\). The small-density theorem of Bolthausen--Nakajima--Sun--Xu applies to the half-space indicator \(U(x)=\1_{\{x\ge\theta_C\}}\) and gives satisfiability with high probability at some density \(\alpha_0>0\) \cite[Theorem~1.1 and Proposition~1.3]{BolthausenNakajimaSunXu}. Every such solution also satisfies the constraints in \(C\). The supercritical conclusion just proved therefore forces \(r_C\ge\alpha_0>0\), completing \eqref{eq:closed-set-capacity-upper-bound}. For fixed $N$, each of the $2^N$ configurations survives the first $m$ independent rows with probability $\pi_C^m$, and hence \[ \Pp\bigl(\mathfrak M_N^{\rm I}(C)\ge m\bigr) \le 2^N\pi_C^m. \] Letting $m\to\infty$ shows $\mathfrak M_N^{\rm I}(C)<\infty$ almost surely. For fixed \(00\). Here \(\pi_C=1/2\), so \(\alpha_c^{\rm I}(C)\le1\); the full-path formula does not by itself establish the optimized value one predicted in \cite{BexSerneelsVandenBroeck}. For the square-wave set \[ C=\bigcup_{j\in\mathbb Z}[(2j+1)\delta,(2j+2)\delta],\qquad\delta>0, \] \cref{thm:iid-channels} gives the soft formula for \(-\beta\operatorname{dist}(\,\cdot\,,C)\), but the hardening theorem does not apply because \(C\) contains no half-line. The present results also do not cover the correlated-label teacher--student model of \cite{BenedettiEtAl}. \subsubsection{Spherical half-spaces} \label{sec:halfspace-spherical-capacity} In this subsection we abbreviate the half-space quantities by \begin{gather*} u_{\beta,\theta}=u_{\beta,C_\theta},\qquad \Phi_{\infty,\theta}=\Phi_{\infty,C_\theta},\qquad \mathcal P_{\kappa,\infty,\theta}=\mathcal P_{\kappa,\infty,C_\theta},\\ \alpha_c^{\rm I}(\theta)=\alpha_c^{\rm I}(C_\theta),\qquad \mathcal F_{M,N}^{\kappa}(\theta)=\mathcal F_{M,N}^{\kappa}(C_\theta),\qquad \mathfrak M_N^\kappa(\theta)=\mathfrak M_N^\kappa(C_\theta). \end{gather*} Write \(\varphi_{\rm G}(x)=(2\pi)^{-1/2}e^{-x^2/2}\) and \(\overline\Phi_{\rm G}(x)=\Pp(Z\ge x)\), with \(Z\sim N(0,1)\). \begin{equation*} \mathcal B_\theta(q)=-\Phi_{\infty,\theta}(q), \end{equation*} \begin{equation*} D_\theta=\E(\theta-Z)_+^2. \end{equation*} The hard-volume formula does not itself imply emptiness on the sphere. The next three results supply the threshold geometry, margin transfer, and robust thickening needed to prove \cref{thm:spherical-halfspace-capacity}. \begin{lemma}[Constant-path bound and the spherical finiteness threshold] \label{lem:hard-capacity-geometry} Fix \(\theta\in\mathbb R\). For $0\le x<1$, the constant path satisfies \begin{align} \mathcal A_{\rm S}(\overline q_x) &=\frac12\left(\frac{x}{1-x}+\log(1-x)\right), \label{eq:spherical-rs-spin-cost}\\ \Phi_{\infty,\theta}(\overline q_x) &=\E\log\overline\Phi_{\rm G}\left( \frac{\theta-\sqrt{x}Z}{\sqrt{1-x}}\right). \label{eq:constant-hard-channel} \end{align} The set of $\alpha$ for which $\mathcal P_{\mathrm S,\infty,\theta}(\alpha)>-\infty$ is downward closed. Hence the spherical hard value is finite for $\alpha<\alpha_c^{\rm S}(\theta)$ and equals $-\infty$ for $\alpha>\alpha_c^{\rm S}(\theta)$. Moreover, \begin{equation*} \alpha_c^{\rm S}(\theta)\le D_\theta^{-1}<\infty. \end{equation*} \end{lemma} \begin{proof} The spherical identity in \eqref{eq:spherical-rs-spin-cost} follows from \eqref{eq:spherical-step-entropy}. For the channel, conditional on the common field $\sqrt xZ$, the soft normalizer decreases to \[ \overline\Phi_{\rm G}\left( \frac{\theta-\sqrt xZ}{\sqrt{1-x}}\right). \] Its logarithm has integrable negative part by Mills' inequalities, which proves \eqref{eq:constant-hard-channel}. The same inequalities and dominated convergence give, as $x\uparrow1$, \[ (1-x)\mathcal A_{\rm S}(\overline q_x)\longrightarrow\frac12, \qquad (1-x)\mathcal B_\theta(\overline q_x) \longrightarrow\frac12D_\theta. \] Thus constant paths make the spherical objective tend to $-\infty$ when $\alpha>D_\theta^{-1}$. Since \(\Phi_{\infty,\theta}(q)\le0\), the spherical hard value is nonincreasing in \(\alpha\). Finiteness at \(\alpha_2\) therefore gives a finite lower bound at every \(0<\alpha_1<\alpha_2\), while the trial path \(\overline q_0\) gives a finite upper bound. This proves downward closure. \end{proof} \begin{proposition}[Supercriticality persists after lowering the margin] \label{prop:spherical-margin-relaxation} For every \(\theta\in\mathbb R\) and every $\alpha>\alpha_c^{\rm S}(\theta)$, there exists $\delta>0$ such that \begin{equation}\label{eq:spherical-relaxed-supercritical} \mathcal P_{\mathrm S,\infty,\theta-\delta}(\alpha)=-\infty. \end{equation} \end{proposition} \begin{proof} We bound the cost of shifting the margin by the spherical spin cost. First, every path of finite spherical spin cost satisfies \begin{align} \mathcal A_{\rm S}(q) &\ge\frac12\left(\frac1{\mathsf L(q)}-1+\log\mathsf L(q)\right), \notag\\ \frac1{\mathsf L(q)}&\le4\bigl(1+\mathcal A_{\rm S}(q)\bigr). \label{eq:spherical-susceptibility-bound} \end{align} In particular, \(\mathsf L(q)>0\). To prove these, first suppose \(q(1-)<1\), put \(\ell=\mathsf L(q)=\int_0^1\gamma_q\), and write \(D(t)=\int_t^1\gamma_q(s)\,\dd s\). Since \(D(t)\le\min\{\ell,1-t\}\), \eqref{eq:spherical-entropy-classical} gives \[ 2\mathcal A_{\rm S}(q) =\int_0^1\left\{\frac1{D(t)}-\frac1{1-t}\right\}\dd t \ge\int_0^{1-\ell}\left\{\frac1\ell-\frac1{1-t}\right\}\dd t =\ell^{-1}-1+\log\ell. \] For general \(q\), cap it at \(1-\eta\) on \([0,1)\), retaining the diagonal value \(q_\eta(1)=1\). The conjugate formula \eqref{eq:spherical-conjugate-identification} gives \(\mathcal A_{\rm S}(q_\eta)\le\mathcal A_{\rm S}(q)\), while \(\mathsf L(q_\eta)\to\mathsf L(q)\). The preceding bound rules out \(\mathsf L(q)=0\) and passes to the limit. Applying \(x\le4+2(x-1-\log x)\), \(x\ge1\), proves \eqref{eq:spherical-susceptibility-bound}. The second estimate, for every \(\delta,\varepsilon>0\), is \begin{equation}\label{eq:spherical-margin-shift} \mathcal B_{\theta+\delta}(q) \le(1+\varepsilon)\mathcal B_\theta(q) +\left(1+\frac1\varepsilon\right)\frac{\delta^2}{2\mathsf L(q)}. \end{equation} The scalar recursion and \cref{lem:mz-control}, as in the finite-temperature passage above, identify \(\Phi_{u_{\beta,\vartheta}}(q) =\mathcal T_{\gamma_q}u_{\beta,\vartheta}(0,0)\) directly for the Lipschitz terminal \(u_{\beta,\vartheta}(x)=-\beta(\vartheta-x)_+\). Put \(\ell=\mathsf L(q)\). For any admissible control \(v\), set \(X^v=B_1+\int_0^1\gamma_qv\) and \(w=v+\delta/\ell\). The shifted control has finite weighted quadratic cost and is therefore admissible; the optional clipping bound need not be imposed. For \(p>1\), \[ u_{\beta,\theta+\delta}(X^w) =u_{\beta,\theta}(X^v)\ge p\,u_{\beta,\theta}(X^v),\qquad \int_0^1\gamma_qw^2 \le p\int_0^1\gamma_qv^2+\frac{p\delta^2}{(p-1)\ell}. \] The first inequality uses nonpositivity of the terminal and the second is Young's inequality. Taking control suprema, then letting \(\beta\to\infty\) and setting \(p=1+\varepsilon\), proves \eqref{eq:spherical-margin-shift}. Choose \(\alpha_c^{\rm S}(\theta)<\alpha_0<\alpha\), and then \(\delta>0\) so small that \[ \alpha_0\left\{\frac{1+\delta}{\alpha} +2\delta(1+\delta)\right\}<1. \] If the hard value at \((\alpha,\theta-\delta)\) were finite, some \(C_0\ge0\) would give, for every finite-cost path, \[ \mathcal B_{\theta-\delta}(q) \le\frac{\mathcal A_{\rm S}(q)+C_0}{\alpha}. \] Apply \eqref{eq:spherical-margin-shift} from \(\theta-\delta\) to \(\theta\), with \(\varepsilon=\delta\), and use \eqref{eq:spherical-susceptibility-bound}. The result is \[ \mathcal B_\theta(q) \le\left\{\frac{1+\delta}{\alpha}+2\delta(1+\delta)\right\} \mathcal A_{\rm S}(q)+C_\delta. \] Thus \(\mathcal A_{\rm S}(q)-\alpha_0\mathcal B_\theta(q)\) is uniformly bounded below, contradicting \cref{lem:hard-capacity-geometry}. This proves \eqref{eq:spherical-relaxed-supercritical}. \end{proof} \begin{proposition}[Feasibility at margin \(\theta\) gives exponential volume at margin \(\theta-\delta\)] \label{prop:spherical-robust-thickening} Suppose $M_N/N\to\alpha\in(0,\infty)$. For every $\theta\in\R$ and $\delta>0$, there is a finite $C=C(\alpha,\theta,\delta)$ such that \begin{equation}\label{eq:spherical-thickening} \Pp\left( \mathcal F_{M_N,N}^{\rm S}(\theta)\ne\varnothing,\quad \frac1N\log\mu_N^{\rm S} \bigl(\mathcal F_{M_N,N}^{\rm S}(\theta-\delta)\bigr)<-C \right)\longrightarrow0. \end{equation} \end{proposition} \begin{proof} Fix $\bar\alpha>\alpha$. Gaussian norm tails show that, with probability tending to one, $M_N\le\bar\alpha N$ and $\max_{k\le M_N}\|g_k\|\le2\sqrt N$. The remaining argument is deterministic in any rows satisfying these bounds. Identify $S_N$ with the unit sphere by $s=\sigma/\sqrt N$, so constraints become $g_k\cdot s\ge\theta$ and $\mu_N^{\rm S}$ remains normalized surface measure. For an independent $W\sim N(0,I_N/N)$, \v{S}id\'ak's inequality \cite[Corollary~1, p.~628]{Sidak1967} gives \[ \Pp_W\left(\max_k|g_k\cdot W|\le L\right) \ge p_{L/2}^{M_N}\ge e^{-\kappa_LN},\qquad p_x=\Pp(|Z|\le x),\quad \kappa_L=-\bar\alpha\log p_{L/2}. \] Here each $g_k\cdot W$ has variance at most four; arbitrary correlations are allowed. Also $\Pp_W(\|W\|>2)\le e^{-c_0N}$ for a universal $c_0>0$. Fix $L$ large enough that $\kappa_L\alpha_c^{\rm S}(\theta)$, choose $\delta>0$ from \cref{prop:spherical-margin-relaxation}. The hard-wall theorem at margin $\theta-\delta$ and \cref{prop:spherical-robust-thickening} imply $\Pp(\mathcal F_{M_N,N}^{\rm S}(\theta)\ne\varnothing)\to0$. For fixed $N$, cover the unit sphere by finitely many Euclidean caps $\{s:\|s-v_j\|_2