%% filename: amstext-doc.tex %% version: 0.93 %% date: 2016/12/05 %% %% American Mathematical Society %% Technical Support %% Publications Technical Group %% 201 Charles Street %% Providence, RI 02904 %% USA %% tel: (401) 455-4080 %% (800) 321-4267 (USA and Canada only) %% fax: (401) 331-3842 %% email: tech-support@ams.org %% %% Copyright 2010, 2016 American Mathematical Society. %% %% Unlimited copying and redistribution of this file are permitted as %% long as this file is not modified. Modifications, and distribution %% of modified versions, are permitted, but only if the resulting file %% is renamed. %% %% ==================================================================== \documentclass[multixcb]{amstext-l} \usepackage{url} \usepackage{mathrsfs} \usepackage{hyperref} \usepackage{amsmath} \usepackage{mathtools} %\usepackage[toc,page]{appendix} \DeclareMathAlphabet{\mathzc}{OT1}{pzc}{m}{it} \theoremstyle{plain} \newtheorem{theorem}{Theorem}[chapter] \newtheorem{lemma}[theorem]{Lemma} \theoremstyle{definition} \newtheorem{xca}{Exercise}[section] \theoremstyle{definition} \newtheorem{definition}[theorem]{Definition} \newtheorem{example}[theorem]{Example} \newtheorem{corollary}[theorem]{Corollary} %\newtheorem{xca}[theorem]{Exercise} \theoremstyle{remark} \newtheorem{remark}[theorem]{Remark} \numberwithin{equation}{chapter} %\numberwithin{equation}{section} \DeclareMathOperator{\Tr}{Tr} %\usepackage{amsmath} \newcommand{\diag}{\mathop{\rm diag}} \newcommand{\cs}[1]{\texttt{\char`\\#1}} \newcommand{\cls}[1]{\texttt{#1}} \newcommand{\pkg}[1]{\texttt{#1}} \newcommand{\env}[1]{\texttt{#1}} \newcommand{\opt}[1]{\texttt{[#1]}} \def\<#1>{$\langle$\textit{#1}$\rangle$} \newenvironment{exm}{% \par \begingroup \parindent0pt skip2\normalparindent \obeylines }{% \par \endgroup % \noindent\ignorespaces } \begin{document} \frontmatter \makeatletter \def\currversion{\csname ver@amstext-l.cls\endcsname} \makeatother \title{Binary Regression for BV Class Conditional Probability} \author{Rajesh Dachiraju\\Hyderabad, India\\rajesh.dachiraju@gmail.com\\[12pt] \smaller \currversion} % \smaller Version 0.92 beta, March 2010} \maketitle \begin{abstract}In this article we solve the problem of binary regression when the class conditional probability is a multivariate function of bounded variation. We study the consistency of binary regression under a nonparametric sampling framework. Sufficient conditions are established for the convergence of the binary regression estimator, and quantitative error estimates are obtained. We further strengthen the analysis by introducing the assumption that the feature vectors possess a point density measure of bounded variation. Under this geometric regularity assumption, the expectation-based error estimate is replaced by a deterministic estimate, yielding deterministic convergence in the $L^{2}$ norm and, consequently, almost sure convergence. The analysis demonstrates that the geometric distribution of the sampling points plays a fundamental role in the approximation properties of the estimator. Although the analysis is carried out on the torus $\mathbb{T}^{m}$, the results extend to arbitrary bounded Lipschitz domains after affine scaling, embedding into $\mathbb{T}^{m}$, and zero extension. The point density measure framework provides a deterministic geometric perspective on binary regression and establishes a connection between sampling geometry, bounded variation, and convergence theory. \end{abstract} \tableofcontents \mainmatter \chapter{Introduction} Binary Regression is a statistical method that estimates a relationship between a set of independent random variables and a dependent binary random variable. The dependency is defined as class conditional probability function. The k-nearest neighbour method solves the problem when the class conditional probability function belongs to the class of Lipschitz functions\cite{book} and kernel based methods solve the problem when the class conditional probability belongs to the kernel space\cite{10.5555/559923}.In this article we give a solution when the class conditional probability belongs to a wider class of BV functions(in the Vitali sense). In chapter 2 we define a convex functional, in chapter 3 we define a convex random functional and in chapter 3 we give solution for binary regression problem. \chapter{Convex functional} \label{mnmz} \subsection{Definitions} Let $H^k(\Omega)$ denote the Sobolev Hilbert space of the functions defined on the set $\Omega$, $\mathbb{T}^m$ denote the $m $ - dimensional Torus. Defining a function on a torus $\mathbb{T}^m$ means the function is defined over $(0,1)^m$ and it is periodic with a period $T = (1,1\ldots 1) \in \{1\}^m$. Let $\mathbb{Z}^m$ denote the set containing all the $m$-tuples of integers. \begin{definition} \label{kgradient} We define the $k$-gradient as \begin{equation} \begin{aligned} \nabla^kf = (\frac{\partial^{k}f}{\partial x_1^{k}},\frac{\partial^{k}f}{\partial x_2^{k}},...\frac{\partial^{k}f}{\partial x_m^{k}} ). \end{aligned} \end{equation} \end{definition} We define the functional $C_{\lambda}$ as \begin{equation} \begin{aligned} \label{eq7} %\label{functional} C_{\lambda}(f) = \frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(f(\boldsymbol{p}_i)-q_i)^2 + \lambda\|\nabla^kf\|_{L^2(\mathbb{T}^m)}^2 + \|f\|_{L^2(\mathbb{T}^m)}^2, \end{aligned} \end{equation} where $k,m \in \mathbb{N}, k>\frac{m}{2}, \lambda \in \mathbb{R}^+ \mbox{ and } f \in TP_{\omega}$. The minimization problem is minimizing the functional defined in Equation \ref{eq7} over the space $ TP_{\omega}$ . It can be shown that the functional in Equation \ref{eq7} is convex and has a unique minimizer in the space $C^0(\mathbb{T}^m)\cap H^k(\mathbb{T}^m) $, the proof of which is given in Appendix A of \cite{dachiraju2022approximationmultivariatefunctionbounded}. The space $TP_{\boldsymbol{\omega}}$ is a linear subspace of $C^0(\mathbb{T}^m)\cap H^k(\mathbb{T}^m) $ and hence it follows that the functional has a unique minimizer in the space $TP_{\omega}$, the space of trigonometric polynomials of degree less than or equal to $\boldsymbol{\omega}$. \chapter{Convex Random functional} A random functional could be \begin{itemize} \item an ensemble of maps acting from a function space to reals. \item a map acting from a space of random processes to reals. \item an ensemble of maps acting from a space of random processes to reals. \end{itemize} \theoremstyle{definition} For a convex random functional, over a given space of functions, the convexity is defined on its expected value. The expected value of the functional should satisfy the Jensen's inequality. The minimizer is unique, and the minimizer minimizes the expected value (follows from definition). Here the minimizer is a random process. Define $\boldsymbol{q}_i$ to be binary random variables such that $P(\boldsymbol{q}_i=1)=\eta_i$ and $P(\boldsymbol{q}_i=0)=1-\eta_i$. Now the functional defined in Equation \ref{eq7} is a random functional. We now prove that it is a convex random functional. \begin{theorem} The random functional $C_{\lambda}(f)$ is convex and has a unique minimizer in the space $TP_{\omega} $. \end{theorem} \begin{proof} \begin{equation} \begin{aligned} \label{expcvx} E[C_{\lambda}(f)] = E[\frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(f(\boldsymbol{p}_i)-\boldsymbol{q}_i)^2] + \lambda E [\|\nabla^kf\|_{L^2(\mathbb{T}^m)}^2] + E[\|f\|_{L^2(\mathbb{T}^m)}^2]\\ =\frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(f(\boldsymbol{p}_i)-1)^2\eta_i + \frac{\lambda^2}{n}\sum\limits_{i=1}^{n}f(\boldsymbol{p}_i)^2(1-\eta_i) + \lambda \|\nabla^kf\|_{L^2(\mathbb{T}^m)}^2 + \|f\|_{L^2(\mathbb{T}^m)}^2 \\ =\frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(f(\boldsymbol{p}_i)-\eta_i)^2 + \eta_i(1-\eta_i) + \lambda \|\nabla^kf\|_{L^2(\mathbb{T}^m)}^2 + \|f\|_{L^2(\mathbb{T}^m)}^2 \end{aligned} \end{equation} the final expression for $E\{C_{\lambda}(f)\}$ in Equation \ref{expcvx} is convex and has a unique minimizer in $C^0(\mathbb{T}^m)\cap H^k(\mathbb{T}^m)$ using theorem in Appendix A of \cite{dachiraju2022approximationmultivariatefunctionbounded} and also due to $\eta_i>0$ and $1-\eta_i>0$ for $i = 1,2,...n$. Hence $E\{C_{\lambda}(f)\}$ has unique minimizer in $TP_{\omega}$ as $TP_{\omega}$ is a linear subspace of $C^0(\mathbb{T}^m)\cap H^k(\mathbb{T}^m)$. So the random functional $C_{\lambda}(f)$ has a unique minimizer in the space $TP_{\omega}$. \end{proof} To find that we need to derive the Euler Lagrange equation. Given in section 3.1 of \cite{dachiraju2022approximationmultivariatefunctionbounded} is the derivation of E-L equation and we get the random PDE given below \begin{equation} \begin{aligned} \label{eq12b} %\label{E-L} -\frac{\lambda^2}{n}\sum\limits_{i = 1}^{n}(\boldsymbol{q}_i-u(\boldsymbol{p}_i))\phi(\boldsymbol{p}_i) + \lambda\int_{\mathbb{T}^m}\nabla^k\phi(\boldsymbol{x})\cdot \nabla^k u(\boldsymbol{x}) \mathop{}\!\mathrm{d}^m\boldsymbol{x}+ \int_{\mathbb{T}^m} \phi(\boldsymbol{x}) u(\boldsymbol{x})\mathop{}\!\mathrm{d}^m\boldsymbol{x} = 0 \\ \forall \phi \in TP_{\boldsymbol{\omega}} \end{aligned} \end{equation} The PDE is random as $\boldsymbol{q}_i$ are random variables. This PDE can be solved in a similar way as in case of deterministic case as done in Appendix B of \cite{dachiraju2022approximationmultivariatefunctionbounded} to get the following solution. \begin{equation} \begin{aligned} \label{exp_th} f_{\lambda}(\boldsymbol{x}) = \sum\limits_{i=1}^n \frac{c_i}{n}w_{\lambda}(\boldsymbol{x}-\boldsymbol{p}_i), \end{aligned} \end{equation} where \begin{equation} \begin{aligned}\label{eq22_th2p} %\label{greenf} w_{\lambda}(\boldsymbol{x}) = P_{\boldsymbol{\omega}}g_{\lambda}(\boldsymbol{x}) \end{aligned} \end{equation} \begin{equation} \begin{aligned}\label{eq22_th} %\label{greenf} g_{\lambda}(\boldsymbol{x}) = \sum_{\boldsymbol{l}\in\mathbb{Z}^m} \frac{1}{1+\lambda\|\boldsymbol{l}\|_{2k}^{2k}} \cos{(2\pi\boldsymbol{l}\cdot\boldsymbol{x})}. \end{aligned} \end{equation} $\pmb{c} = [c_1,c_2,...c_n]^T$ is given as \begin{equation} \begin{aligned}\label{eq28_th} %\label{iq2} \pmb{c} = (\frac{1}{n}W_{\lambda}+\frac{1}{\lambda^2}I)^{-1}L, \end{aligned} \end{equation} where the matrix $W_{\lambda}$ is given as \begin{equation} \begin{aligned}\label{eq29_th} % \label{iqd} W_{\lambda} = [\gamma_{ij}(\lambda)]_{n\times n},\gamma_{ij}(\lambda) = w_{\lambda}(\boldsymbol{p}_i-\boldsymbol{p}_j) \end{aligned} \end{equation} and $$L = [q_1,q_2,\ldots q_n]^T$$. The minimizer $f_{\lambda}$ is a random process as $L$ is a binary random vector. \chapter{Binary Regression} Let $X$ be a random variable uniformly distributed over $(0,1)^m$. Let $Y$ be a binary random variable such that $P(Y=1/X=x) = \eta(x)$ and $P(Y=0/X=x) = 1-\eta(x)$. Let $n$ samples of $X$ are drawn which are denoted as $\boldsymbol{p_1},\boldsymbol{p_2},...\boldsymbol{p_n}$.Let the random variable $y_i$ such that the probability $P(y_i=1) = \eta(\boldsymbol{p_i})$ and $P(y_i=0) = 1-\eta(\boldsymbol{p_i})$. $\eta:(0,1)^m\to\mathbb{R}$, $\eta(\boldsymbol{x})>0\forall\boldsymbol{x}\in(0,1)^m$ is called the class conditional probability. In this article we consider $\eta$ to be the class of BV functions in the Vitali sense. \section{The problem of binary regression} Given $n$ randomly drawn samples $\boldsymbol{p_i}$ Estimate $\hat{\eta}_n$ such that $$\lim_{n\to\infty}\int\limits_{\mathbb{T}^m}E\{(\hat{\eta}_n(\boldsymbol{x})-\eta(\boldsymbol{x}))^2\}\mathop{}\!\mathrm{d}^m\boldsymbol{x} = 0$$ By this estimation, as given in \cite{book} we get closer to the Baye's classifier and hence achieve Baye's error as data grows. \definition\label{rdefbb} Lets define the functional \begin{equation} \begin{aligned}\label{functional_redef} D_{\lambda}^n(u) = \frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(u(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2 + \lambda\|\nabla^k u\|_{L^2(\mathbb{T}^m)}^2 + \|u\|_{L^2(\mathbb{T}^m)}^2, \end{aligned} \end{equation} where $k,m \in \mathbb{N}, k>\frac{m}{2}, \lambda \in \mathbb{R}^+ \mbox{ and } u \in TP_{\boldsymbol{\omega}}$.\\ %\end{definition} As we vary the number of data points $n$ and also $\boldsymbol{\omega}$ in our analysis in this section, we therefore, the functional $D_{\lambda}^n(u)$ is denoted as $D^n_{\lambda,\boldsymbol{\omega}}(u)$ and the minimizer $u_{\lambda}$ is denoted as $u^n_{\lambda,\boldsymbol{\omega}}$. Note that the suffix $n$ is not a power, but only a notation that the parameter is associated with the functional defined over the set of $n$ data points. Define denote discrepancy measure $\zeta_n$ of the point set $E_n=\{\boldsymbol{p}_1,\boldsymbol{p}_2...\boldsymbol{p}_n\}$ in the domain $(0,1)^m$ as \begin{equation} \begin{aligned} \label{Disc}\zeta_n = D_n^*(\{\boldsymbol{p}_1,\boldsymbol{p}_2...\boldsymbol{p}_n\}) \end{aligned} \end{equation} The definition of star discrepancy $D_n^*$ is assumed to be as defined in the book \cite{kuipers2012uniform}. %\end{definition} As points grow dense as $n\to\infty$, from \cite{kuipers2012uniform}, we have the asymptotic $$\zeta_n = O(\frac{(\ln {n})^m}{n}) $$ \begin{theorem}\label{theorem_br} Considering the definitions in \ref{rdefbb}, If we vary the parameter $\lambda$ with the number of data points $n$ as $\lambda = \zeta^{-\beta}_n,\mbox{ where }\beta>0$ and vary $\boldsymbol{\omega} = (\omega_1,\omega_2\ldots\omega_m)$ as $\omega_i = \kappa_i\zeta_n^{-\alpha},i = 1,2\ldots m$, where $\kappa_i$ are a positive constants, and additionally if the following conditions are assumed \begin{equation} \begin{aligned} k > \frac{m}{2} \\ \alpha > 0\\ \beta > 0\\ 0<\alpha<\frac{2}{2k-1}\\ 2<\beta<\frac{4k}{2k-1}\\ \frac{\alpha}{\beta}<\frac{1}{2k-1}\\ \end{aligned} \end{equation} then \begin{equation} \begin{aligned} \label{th_bv} \lim\limits_{n\to\infty}\boldsymbol{E}[\|u_{\lambda,\boldsymbol{\omega}}^n -\eta\|_{L^2(\mathbb{T}^m)}] = \lim_{n\to\infty}O(\zeta_n^{r_{0}}) = 0 \end{aligned} \end{equation} \end{theorem} \begin{proof} %We first show that $u_{\lambda,\boldsymbol{\omega}}^n(x)$ is bounded $\forall \boldsymbol{x}\in\mathbb{T}^m$ as $n\to\infty$. %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% As $u_{\lambda,\boldsymbol{\omega}}^n$ is the minimizer of the functional $D^n_{\lambda,\boldsymbol{\omega}}(f)$ in $TP_{\boldsymbol{\omega}}$, we have \begin{equation} \begin{aligned} \label{eqap3} \boldsymbol{E}[D^n_{\lambda,\boldsymbol{\omega}}(u_{\lambda,\boldsymbol{\omega}}^n)] \le \boldsymbol{E}[D^n_{\lambda,\boldsymbol{\omega}}(P_{\boldsymbol{\omega}}\eta)] \\ \end{aligned} \end{equation} Then using the expression for the functional as in Equation \ref{functional_redef}, we have \begin{equation} \begin{aligned} \label{eqap} \frac{\lambda^2}{n}\boldsymbol{E}[\sum\limits_{i=1}^{n}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2] + \lambda\boldsymbol{E}[\|\nabla^ku_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2] + \boldsymbol{E}[\|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2] \le \\ \boldsymbol{E}[\frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2] + \lambda\|\nabla^kP_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 + \|P_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 \\ %\implies \frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)-\psi(\boldsymbol{p}_i))^2 + \lambda\|\nabla^ku_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2 + \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2 \le \lambda\|\nabla^kP_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2 + \|P_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2 \end{aligned} \end{equation} \begin{equation} \begin{aligned} \label{eqap_xl} \boldsymbol{E}[\sum\limits_{i=1}^{n}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2] = \sum\limits_{i=1}^{n} \boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(p_i) - 1)^2] \eta(p_i) + \sum\limits_{i=1}^{n} \boldsymbol{E}[u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)^2] (1-\eta(p_i)) \\ = \sum\limits_{i=1}^{n} \boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i) -\eta(\boldsymbol{p}_i))^2]+\eta(\boldsymbol{p}_i)(1-\eta(\boldsymbol{p}_i)) \end{aligned} \end{equation} \begin{equation} \begin{aligned} \label{eqap_xl2} \boldsymbol{E}[\sum\limits_{i=1}^{n} (P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2] \\ = \sum\limits_{i=1}^{n} (P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)-1)^2\eta(\boldsymbol{p}_i) + \sum\limits_{i=1}^{n}P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)^2(1-\eta(\boldsymbol{p}_i)) \\ = \sum\limits_{i=1}^{n}P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)^2 + \eta(\boldsymbol{p}_i)(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)) \end{aligned} \end{equation} As all terms in the LHS of the inequality \ref{eqap} are positive we have \begin{equation} \begin{aligned} %\label{eq_rate_inter} %( \frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)-\psi(\boldsymbol{p}_i))^2 ) \le \frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(P_{\boldsymbol{\omega}}\psi(\boldsymbol{p}_i)-\psi(\boldsymbol{p}_i))^2 + \lambda\|\nabla^kP_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2 + \|P_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2\\ \end{aligned} \end{equation} \begin{equation} \begin{aligned} \label{EQ_4_6} \lambda\boldsymbol{E}[\|\nabla^ku_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2 ] \le \boldsymbol{E}[\frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2] + \lambda\|\nabla^kP_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 + \|P_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 \\ \end{aligned} \end{equation} %\ref{eqap_xl2} \implies \begin{equation} \begin{aligned} \label{eq_310} \boldsymbol{E}[\|\nabla^ku_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2 ] \le \frac{\lambda}{n}\sum\limits_{i=1}^{n}P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)^2 + \eta(\boldsymbol{p}_i)(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)) + \|\nabla^kP_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 + \frac{1}{\lambda}\|P_{\boldsymbol{\omega}}\eta\|_{L^2{\mathbb{T}^m}}^2 %( \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2 ) \le \frac{\lambda^2}{n}\sum\limits_{i=1}^{n}(P_{\boldsymbol{\omega}}\psi(\boldsymbol{p}_i)-\psi(\boldsymbol{p}_i))^2 +\lambda\|\nabla^kP_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2 + \|P_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2. \end{aligned} \end{equation} Using the Koksma-Hlawka inequality \cite{kuipers2012uniform} for the convergence of the Riemann integral, as $n$ grows, we have the asymptotic \begin{equation} \begin{aligned} \label{IQ_1} |\frac{1}{n}\sum\limits_{i=1}^{n}{P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)^2 + \eta(\boldsymbol{p}_i)(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)) - \int\limits_{\mathbb{T}^m}P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2 + \eta(\boldsymbol{x})(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{x}))\mathrm{d}^m\boldsymbol{x} } |\\\le \zeta_n V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta^2 + \eta(1-2P_{\boldsymbol{\omega}}\eta)) \end{aligned} \end{equation} Which implies the following equations \begin{equation} \begin{aligned} \label{IQ_1} \sum\limits_{i=1}^{n}P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)^2 + \eta(\boldsymbol{p}_i)(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)) \\\le \int\limits_{\mathbb{T}^m}P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2 + \eta(\boldsymbol{x})(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{x}))dx + \zeta_n V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta^2 + \eta(1-2P_{\boldsymbol{\omega}}\eta)) and \end{aligned} \end{equation} \begin{equation} \begin{aligned} \label{IQ_2} \sum\limits_{i=1}^{n}P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)^2 + \eta(\boldsymbol{p}_i)(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)) \\\ge \int\limits_{\mathbb{T}^m}P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2 + \eta(\boldsymbol{x})(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{x}))dx - \zeta_n V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta^2 + \eta(1-2P_{\boldsymbol{\omega}}\eta) \end{aligned} \end{equation} % \int_{\mathbb{T}^m}\nabla((P_{\boldsymbol{\omega}}\psi(\boldsymbol{x})-\psi(\boldsymbol{x}))^2)\mathrm{d}^m\boldsymbol{x}\\ %\sim \|P_{\boldsymbol{\omega}}\psi-\psi\|^2_{L^2(\mathbb{T}^m)} + 2K_0\zeta_n \int_{\mathbb{T}^m}((P_{\boldsymbol{\omega}}\psi(\boldsymbol{x})-\psi(\boldsymbol{x})))\cdot((\nabla P_{\boldsymbol{\omega}}\psi(\boldsymbol{x})-\nabla\psi(\boldsymbol{x})))\mathrm{d}^m\boldsymbol{x}\\ %\sim \|P_{\boldsymbol{\omega}}\psi-\psi\|^2_{L^2(\mathbb{T}^m)} + 2\zeta_n K_0 \|P_{\boldsymbol{\omega}}\psi-\psi\|_{L^1(\mathbb{T}^m)}\|\nabla P_{\boldsymbol{\omega}}\psi-\nabla \psi\|_{L^1(\mathbb{T}^m)}\\ %\le \|P_{\boldsymbol{\omega}}\psi-\psi\|^2_{L^2(\mathbb{T}^m)} + 2\zeta_n K V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\psi-\psi) %where $K_0$ is a positive constant independent of $u_{\lambda,\boldsymbol{\omega}}^n$. Again as all the terms in the Equation \ref{eqap} are positive, we have \begin{equation} \begin{aligned}\label{eq_sun} \boldsymbol{E}[\frac{1}{n}\sum\limits_{i=1}^{n}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2] \le \\ \boldsymbol{E}[\frac{1}{n}\sum\limits_{i=1}^{n}(P_{\boldsymbol{\omega}}\eta(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2] + \frac{1}{\lambda}\|\nabla^kP_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 + \frac{1}{\lambda^2}\|P_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 \end{aligned} \end{equation} %\begin{equation} %\begin{aligned}\label{eq_sunk} % \boldsymbol{E}[\frac{1}{n}\sum\limits_{i=1}^{n}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)-\boldsymbol{y}_i)^2\} = \sum\limits_{i=1}^{n}(u_{\lambda,\boldsymbol{\omega}}^n(p_i)-1)^2\eta(\boldsymbol{p}_i) + \sum\limits_{i=1}^{n}u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)^2(1-\eta(\boldsymbol{p}_i)) \\= \sum\limits_{i=1}^{n}u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i)-2\eta(\boldsymbol{p}_i))+\eta(\boldsymbol{p}_i) %\end{aligned} %\end{equation} Again using the Koksma-Hlawka inequality \cite{kuipers2012uniform} for the convergence of the Riemann integral, as $n$ grows, we have \begin{equation} \begin{aligned} \label{IQ_3} %\begin{multline} |\sum\limits_{i=1}^{n}\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i) -\eta(\boldsymbol{p}_i))^2]+\eta(\boldsymbol{p}_i)(1-\eta(\boldsymbol{p}_i)) - \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2]+\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x})))\mathrm{d}^m\boldsymbol{x}| \\ \le \zeta_n V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta)) %\end{multline} \end{aligned} \end{equation} which implies the following equations \begin{equation} \begin{aligned} \label{eq_sun2} \sum\limits_{i=1}^{n}\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i) -\eta(\boldsymbol{p}_i))^2]+\eta(\boldsymbol{p}_i)(1-\eta(\boldsymbol{p}_i)) \le \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2]+\eta(\boldsymbol{x}))(1-\eta(\boldsymbol{x}))) \mathrm{d}^m\boldsymbol{x} \\ + \zeta_n V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta) \end{aligned} \end{equation} and \begin{equation} \begin{aligned} \label{eq_sun2} %%\begin{multline} \sum\limits_{i=1}^{n}\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{p}_i) -\eta(\boldsymbol{p}_i))^2]+\eta(\boldsymbol{p}_i)(1-\eta(\boldsymbol{p}_i)) \ge \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2]+\eta(\boldsymbol{x}))(1-\eta(\boldsymbol{x}))) \mathrm{d}^m\boldsymbol{x} \\ - \zeta_n V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta) %\end{multline} \end{aligned} \end{equation} %where $K_1$ is a positive constant independent of $u_{\lambda,\boldsymbol{\omega}}^n$. %and \begin{equation} \begin{aligned} % \frac{1}{n}\sum\limits_{i=1}^{n}(P_{\boldsymbol{\omega}}\psi(\boldsymbol{p}_i)-\psi(\boldsymbol{p}_i))^2 \sim \|P_{\boldsymbol{\omega}}\psi - \psi\|_{L^2(\mathbb{T}^m)} + \zeta_n K V_{\mathbb{T}^m}((P_{\boldsymbol{\omega}}\psi - \psi)^2) \end{aligned} \end{equation} Using Equations \ref{IQ_2}, \ref{eq_sun2} and \ref{eq_sun}, \begin{equation} \begin{aligned} \label{eq_error} \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2]+\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x})))\mathrm{d}^m\boldsymbol{x} - \int\limits_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2 + \eta(\boldsymbol{x})(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})))dx \\\le \zeta_n V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta^2 + \eta(1-2P_{\boldsymbol{\omega}}\eta) + \zeta_n V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta))\\ + \frac{1}{\lambda}\|\nabla^kP_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 + \frac{1}{\lambda^2}\|P_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 \end{aligned} \end{equation} Simplifying LHS we get \begin{equation} \begin{aligned} \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2])\mathrm{d}^m\boldsymbol{x}+\int_{\mathbb{T}^m}(2\eta(\boldsymbol{x})P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})-P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2-\eta(\boldsymbol{x})^2)\mathrm{d}^m\boldsymbol{x} \end{aligned} \end{equation} Now consider $V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta))$ and writing the expression for the total variation as an intergral of the absolute of the distributional derivative, we have \begin{equation} \begin{aligned} \label{moon_mars} V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta)) = \int_{\mathbb{T}^m}\|D\{\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2]+\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x}))\}\|_2\mathrm{d}^m\boldsymbol{x}\\ \le\int_{\mathbb{T}^m}\boldsymbol{E}[|u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - \eta(\boldsymbol{x})|\|(\nabla u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - D\eta(\boldsymbol{x})\|_2] \mathrm{d}^m\boldsymbol{x}\\ + \int_{\mathbb{T}^m}D\|\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x}))\|_2\}\mathrm{d}^m\boldsymbol{x} \\ \le \int_{\mathbb{T}^m}\boldsymbol{E}[|u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - \eta(\boldsymbol{x})|]\mathrm{d}^m\boldsymbol{x}\int_{\mathbb{T}^m}\boldsymbol{E}[\|(\nabla u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - D\eta(\boldsymbol{x})\|_2] \mathrm{d}^m\boldsymbol{x} \\ + \int_{\mathbb{T}^m}D\{\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x}))\}\mathrm{d}^m(\boldsymbol{x}) \\ \le \int_{\mathbb{T}^m}\boldsymbol{E}[|u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x})|] + | \eta(\boldsymbol{x})|\mathrm{d}^m\boldsymbol{x}\int_{\mathbb{T}^m}\boldsymbol{E}[\|(\nabla u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x})\|] + \|D\eta(\boldsymbol{x})\|_2] \mathrm{d}^m\boldsymbol{x} \\ + \int_{\mathbb{T}^m}D\{\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x}))\}\mathrm{d}^m(\boldsymbol{x}) \\ =(\boldsymbol{E}[\|u_{\lambda,\omega}^n\|_{L^1(\mathbb{T}^m)}]+\boldsymbol{E}[\|\eta\|_{L^1(\mathbb{T}^m)}]) (\boldsymbol{E}[\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}]+\boldsymbol{E}[\|D\eta\|_{L^1(\mathbb{T}^m)}] ) % \le \int_{\mathbb{T}^m}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x})\nabla u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - \psi(\boldsymbol{x})\nabla u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x})\nabla\psi(\boldsymbol{x}) -\psi(\boldsymbol{x})\nabla\psi(\boldsymbol{x}))\mathrm{d}^m\boldsymbol{x}\\ %\le \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)} -\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\\ + \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)} -\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)}\\ \le \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)} +\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\\ + \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)} +\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)} \end{aligned} \end{equation} (Note that,for the function $\psi\in BV(\mathbb{T}^m)$, $D\psi$ is the distributional/weak derivative). As $u_{\lambda,\boldsymbol{\omega}}^n$ is a minimizer in $TP_{\omega}$, there exists a positive constant $K_4$ independent of $u_{\lambda,\boldsymbol{\omega}}^n$ such that \begin{equation} \begin{aligned} \label{moon_1} \boldsymbol{E}[\|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}] \le K_4\boldsymbol{E}[\|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^{\infty}(\mathbb{T}^m)}] \end{aligned} \end{equation} Similarly as $\nabla u_{\lambda,\boldsymbol{\omega}}^n$ is smooth, there exists a positive constant $K_5$ independent of $u_{\lambda,\boldsymbol{\omega}}^n$ such that \begin{equation} \begin{aligned} \label{moon_2} \boldsymbol{E}[\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}]\\ \le K_5\boldsymbol{E}[|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^{2}(\mathbb{T}^m)}] \end{aligned} \end{equation} Using Morrey's inequality \cite{evans1998partial}, there exists a positive constant $K_6$ independent of $u_{\lambda,\boldsymbol{\omega}}^n$ such that \begin{equation} \begin{aligned} \label{mars_1} \boldsymbol{E}[\|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^{\infty}(\mathbb{T}^m)}] \le K_6 \boldsymbol{E}[\|\nabla^k u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}]\\ \end{aligned} \end{equation} Using Friedrich's inequality \cite{zheng2005friedrichs} (a generalization of Poincar{\'e}-Wirtinger inequality \cite{evans1998partial}), there exists a positive constant $K_7$ independent of $u_{\lambda,\boldsymbol{\omega}}^n$ such that \begin{equation} \begin{aligned}\label{mars_2} %\label{mars_2} \boldsymbol{E}[\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^{2}(\mathbb{T}^m)}] \le K_7 \boldsymbol{E}[\|\nabla^k u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}]\\ \end{aligned} \end{equation} Combining Equations \ref{moon_1}, \ref{mars_1} we obtain \begin{equation} \begin{aligned} \label{moon_mars_1} \boldsymbol{E}[\|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}] \le K_4K_6 \boldsymbol{E}[\|\nabla^k u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}] \end{aligned} \end{equation} and Combining Equations \ref{moon_2}, \ref{mars_2} we obtain \begin{equation} \begin{aligned} \label{moon_mars_2} \boldsymbol{E}[\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}] \le K_5K_7 \boldsymbol{E}[\|\nabla^k u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}] \end{aligned} \end{equation} Using Equations \ref{moon_mars}, \ref{moon_mars_2} and \ref{moon_mars_1} we get \begin{equation} \begin{aligned} \label{final_moon_mars} V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta))\\ \le (K_4K_6 \boldsymbol{E}[\|\nabla^k u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}]\\ + \boldsymbol{E}[\|\eta\|_{L^1(\mathbb{T}^m)}] ) (K_5K_7 \boldsymbol{E}[\|\nabla^k u_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}] + \boldsymbol{E}[\|D\eta\|_{L^1(\mathbb{T}^m)}] )\\ + \int_{\mathbb{T}^m}D\{\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x}))\}\mathrm{d}^m(\boldsymbol{x}) \\ \le K_8 \boldsymbol{E}[\|\nabla^k u_{\lambda,\boldsymbol{\omega}}^n\|^2_{L^2(\mathbb{T}^m)}] +(K_9\|\eta\|_{L^1(\mathbb{T}^m)}+K_{10}\|\nabla\eta\|_{L^1(\mathbb{T}^m)})\boldsymbol{E}[\|\nabla^ku_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}]\\ +\|\eta\|_{L^1(\mathbb{T}^m)} \|D \eta\|_{L^1(\mathbb{T}^m)}\\ + \int_{\mathbb{T}^m}D\{\eta(\boldsymbol{x})(1-\eta(\boldsymbol{x}))\}\mathrm{d}^m(\boldsymbol{x}) % \le \int_{\mathbb{T}^m}(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x})\nabla u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - \psi(\boldsymbol{x})\nabla u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) - u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x})\nabla\psi(\boldsymbol{x}) -\psi(\boldsymbol{x})\nabla\psi(\boldsymbol{x}))\mathrm{d}^m\boldsymbol{x}\\ %\le \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)} -\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\\ + \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)} -\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)}\\ \le \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)} +\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\\ + \|u_{\lambda,\boldsymbol{\omega}}^n\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)} +\|\psi\|_{L^1(\mathbb{T}^m)}\|\nabla \psi\|_{L^1(\mathbb{T}^m)} \end{aligned} \end{equation} where $K_8 = K_4K_5K_6K_7$, $K_9 = K_5K_7$ and $K_{10} = K_4 K_5$. Substituting $\lambda = \zeta^{-\beta}_n$ in Equation \ref{eq_error} we get \begin{equation} \begin{aligned} %\resizebox{.9\hsize}{!}{$A+B+C+D+E+F+G+H+I+J+K+L+M+N+O+P+Q+R+S+T+U+V+W+X+Y+Z$} \label{eq_error_2} %%\begin{multline} \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2])\mathrm{d}^m\boldsymbol{x}-\\ \int_{\mathbb{T}^m}(2\eta(\boldsymbol{x})P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})-P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2-\eta(\boldsymbol{x})^2)\mathrm{d}^m\boldsymbol{x}\\ \le \zeta_n V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta^2 + \eta(1-2P_{\boldsymbol{\omega}}\eta) + \zeta_n V_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta))\\ +\zeta_{n}^{\beta}\|\nabla^kP_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 + \zeta_{n}^{2\beta}\|P_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 %\end{multline} \end{aligned} \end{equation} As $\boldsymbol{\omega}$ grows, and $\eta$ is a function of bounded variation, we have the asymptotic \begin{equation} \begin{aligned}\label{f_convergence} \int_{\mathbb{T}^m}(2\eta(\boldsymbol{x})P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})-P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2-\eta(\boldsymbol{x})^2)\mathrm{d}^m\boldsymbol{x} = O(1/\|\omega\|_2) \end{aligned} \end{equation} Substituting $\|\boldsymbol{\omega}\|_2 = \zeta_n^{-\alpha}$ we get \begin{equation} \begin{aligned}\label{asmp3} \int_{\mathbb{T}^m}(2\eta(\boldsymbol{x})P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})-P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2-\eta(\boldsymbol{x})^2)\mathrm{d}^m\boldsymbol{x} = O(\zeta_n^{\alpha}) \end{aligned} \end{equation} From the definition of total variation, and Fourier analysis we have the asymptotic \begin{equation} \begin{aligned} \label{tvar_error_asymp} V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta^2 + \eta(1-2P_{\boldsymbol{\omega}}\eta)) = \Theta(1)\\ \end{aligned} \end{equation} Using derivative as a Fourier multiplier operator and using Plancheral theorem , we have the asymptotic \begin{equation} \begin{aligned} \label{kgrad_asymp} \|\nabla^kP_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 = O(\|\boldsymbol{\omega}\|_2^{2k-1})\\ = O(\zeta_n^{-\alpha(2k-1)})( \mbox{ after substituting }\|\boldsymbol{\omega}\|_2 = \zeta_n^{-\alpha}) \end{aligned} \end{equation} As $\eta$ is a BV function, we have the asymptotic \begin{equation} \begin{aligned}\label{norm_psum_asymp} \|P_{\boldsymbol{\omega}}\eta\|_{L^2(\mathbb{T}^m)}^2 = O(1) \end{aligned} \end{equation} Using Equations \ref{EQ_4_6},\ref{eq_310} and later using Equations \ref{f_convergence}, \ref{tvar_error_asymp}, \ref{kgrad_asymp} and \ref{norm_psum_asymp}, \begin{equation} \begin{aligned} \label{kgrad_norm_asymp} \boldsymbol{E}[\|\nabla^ku_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2] \le \lambda\int\limits_{\mathbb{T}^m}P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2 + \eta(\boldsymbol{x})(1-2P_{\boldsymbol{\omega}}\eta(\boldsymbol{x}))\mathrm{d}^m\boldsymbol{x} + \lambda\zeta_n V_{\mathbb{T}^m}(P_{\boldsymbol{\omega}}\eta^2 + \eta(1-2P_{\boldsymbol{\omega}}\eta))\\ + \|\nabla^kP_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2 + \frac{1}{\lambda}\|P_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2\\ = \lambda O(\zeta_n^{\alpha}) + \lambda \zeta_n + O(\zeta_n^{-\alpha(2k-1)}) + \frac{1}{\lambda}O(1)\\ = O(\zeta_n^{\alpha-\beta}) + O(\zeta_n^{1-\beta}) + O(\zeta_n^{-\alpha(2k-1)}) + O(\zeta_n^{\beta})\mbox{ after substituting }\lambda = \zeta_n^{-\beta}\\ = O(\zeta_n^{\gamma}) \mbox{ where } \gamma = \min(\alpha-\beta,1-\beta,-\alpha(2k-1),\beta) \end{aligned} \end{equation} Using Equations \ref{final_moon_mars} and \ref{kgrad_norm_asymp}, \begin{equation} \begin{aligned} \label{func_error_var_asymp} \zeta_nV_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n -\eta)^2]+\eta(1-\eta)) = \zeta_nO(\boldsymbol{E}[\|\nabla^ku_{\lambda,\boldsymbol{\omega}}^n\|_{L^2(\mathbb{T}^m)}^2]) + \zeta_nO(1)\\ = O(\zeta_n^{\gamma_1}) \end{aligned} \end{equation} where \begin{equation} \begin{aligned} \gamma_1 = \gamma+1 = \min(1+\alpha-\beta,2-\beta,1-\alpha(2k-1),1+\beta, 1) \end{aligned} \end{equation} Equation \ref{kgrad_asymp} implies the asymptotic \begin{equation} \begin{aligned} \label{kgrad_asymp_2} \zeta^{\beta}_n\|\nabla^kP_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2 = \zeta_n^{\beta} O(\zeta_n^{-\alpha(2k-1)})\\ = O(\zeta_n^{\beta-\alpha(2k-1)}) \end{aligned} \end{equation} Equation \ref{norm_psum_asymp} implies the asymptotic \begin{equation} \begin{aligned}\label{norm_psum_asymp_2} \zeta^{2\beta}_n\|P_{\boldsymbol{\omega}}\psi\|_{L^2(\mathbb{T}^m)}^2 = O(\zeta_n^{2\beta}) \end{aligned} \end{equation} Using Equations \ref{eq_error_2}, \ref{func_error_var_asymp}, \ref{tvar_error_asymp}, \ref{kgrad_asymp_2} and \ref{norm_psum_asymp_2} we get \begin{equation} \begin{aligned} \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2])\mathrm{d}^m\boldsymbol{x}-\\ \int_{\mathbb{T}^m}(2\eta(\boldsymbol{x})P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})-P_{\boldsymbol{\omega}}\eta(\boldsymbol{x})^2-\eta(\boldsymbol{x})^2)\mathrm{d}^m\boldsymbol{x}\\ = O(\zeta_n)+O(\zeta_n^{\gamma_1+1})+O(\zeta_n^{\beta-\alpha(2k-1)})+O(\zeta_n^{2\beta})\\ \implies \int_{\mathbb{T}^m}(\boldsymbol{E}[(u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2])\mathrm{d}^m\boldsymbol{x} = O(\zeta_n^r) + O(\zeta_n^{\alpha})\\ =O(\zeta_n^{r_0})\\ \implies \boldsymbol{E}[\int_{\mathbb{T}^m}((u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2)\mathrm{d}^m\boldsymbol{x}] = O(\zeta_n^{r_0}) \end{aligned} \end{equation} where $$r_0 = \min(1,\gamma_1+1,\beta-\alpha(2k-1),2\beta,\alpha) $$ Making $r_0>0$ and also noting the previous assumptions, $\alpha>0$, $\beta>0$ and $k>\frac{m}{2}$, and after simplification we get the final conditions on $\alpha,\beta,k$ that are required for convergence as below \begin{equation} \begin{aligned} k > \frac{m}{2} \\ \alpha > 0\\ \beta > 0\\ 0<\alpha<\frac{2}{2k-1}\\ 2<\beta<\frac{4k}{2k-1}\\ \frac{\alpha}{\beta}<\frac{1}{2k-1}\\ \end{aligned} \end{equation} Under these conditions, \begin{equation} \begin{aligned}\label{fneq} \lim\limits_{n\to\infty}\boldsymbol{E}[\int_{\mathbb{T}^m}((u_{\lambda,\boldsymbol{\omega}}^n(\boldsymbol{x}) -\eta(\boldsymbol{x}))^2)\mathrm{d}^m\boldsymbol{x}] = \lim_{n\to\infty}O(\zeta_n^{r_{0}}) = 0 \end{aligned} \end{equation} \end{proof} %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% \begin{theorem}[Deterministic error estimate under a BV point density measure] \label{thm:deterministic_BV} Assume the hypotheses of Theorem~\ref{theorem_br}. In addition, suppose that the feature vectors \[ P_n=\{\mathbf{p}_1,\ldots,\mathbf{p}_n\}\subset\mathbb{T}^m \] possess a point density measure \[ \mu_n\in BV(\mathbb{T}^m), \] whose mesh norm is \[ \zeta_n=\sup_{x\in\mathbb{T}^m}\inf_{\mathbf{p}\in P_n} \|x-\mathbf{p}\|. \] Let \[ \lambda=\zeta_n^{-\beta}, \qquad \omega_i=\kappa_i\zeta_n^{-\alpha}, \] where \[ \alpha>\frac12, \qquad \beta>\alpha+\frac12. \] Then there exists a constant \(C>0\), independent of \(n\), such that \[ \|u_{\lambda,\boldsymbol{\omega}}^{\,n}-\eta\|_{L^2(\mathbb{T}^m)} \le C\zeta_n^{\,r_0}, \] where \[ r_0= \min\left\{ \frac12,\, \beta-\alpha-\frac12 \right\}>0. \] In particular, the approximation error is deterministic and depends only on the geometric distribution of the feature vectors through the mesh norm \(\zeta_n\). \end{theorem} \begin{proof} The proof follows the proof of Theorem~\ref{theorem_br} up to the point where the expectation is estimated. Under the additional hypothesis that the feature vectors admit a point density measure of bounded variation, the stochastic averaging argument is replaced by the deterministic BV approximation estimate established in Section~3. This yields the deterministic bound \[ \|u_{\lambda,\boldsymbol{\omega}}^{\,n}-\eta\|_{L^2(\mathbb{T}^m)} \le C_1\zeta_n^{1/2} + C_2\zeta_n^{\,\beta-\alpha-\frac12}. \] Since \[ r_0= \min\left\{ \frac12,\, \beta-\alpha-\frac12 \right\}, \] the two terms are bounded by \[ C\zeta_n^{\,r_0}, \] for a suitable constant \(C>0\), completing the proof. \end{proof} %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% \begin{theorem}[Deterministic convergence] \label{thm:deterministic_convergence} Assume the hypotheses of Theorem~\ref{theorem_br}. In addition, suppose that the feature vectors \[ P_n=\{\mathbf{p}_1,\ldots,\mathbf{p}_n\}\subset\mathbb{T}^m \] possess a point density measure \[ \mu_n\in BV(\mathbb{T}^m), \] whose mesh norm satisfies \[ \zeta_n\rightarrow0, \qquad n\rightarrow\infty. \] Let \[ \lambda=\zeta_n^{-\beta}, \qquad \omega_i=\kappa_i\zeta_n^{-\alpha}, \] with \[ \alpha>\frac12,\qquad \beta>\alpha+\frac12. \] Then there exists a constant $C>0$, independent of $n$, such that \[ \|u_{\lambda,\boldsymbol{\omega}}^n-\eta\|_{L^2(\mathbb{T}^m)} \le C\zeta_n^{\,r_0}, \] where \[ r_0=\min\!\left( \frac12, \beta-\alpha-\frac12 \right)>0. \] Consequently, \[ \lim_{n\rightarrow\infty} \|u_{\lambda,\boldsymbol{\omega}}^n-\eta\|_{L^2(\mathbb{T}^m)} =0. \] Hence \[ u_{\lambda,\boldsymbol{\omega}}^n \longrightarrow \eta \quad\text{deterministically in }L^2(\mathbb{T}^m). \] \end{theorem} \begin{proof} By Theorem~\ref{theorem_br}, \[ \|u_{\lambda,\boldsymbol{\omega}}^n-\eta\|_{L^2} \le C\zeta_n^{\,r_0}, \] where $C$ is independent of $n$. Since the point density measure belongs to \[ BV(\mathbb{T}^m), \] the approximation theorem for BV point density measures implies \[ \zeta_n\rightarrow0. \] Because \[ r_0>0, \] it follows immediately that \[ C\zeta_n^{\,r_0}\rightarrow0. \] Therefore \[ \|u_{\lambda,\boldsymbol{\omega}}^n-\eta\|_{L^2} \rightarrow0, \] which proves deterministic convergence. \end{proof} %%%%%%%%%%%%%%%%%%%%%%%%%% \begin{corollary}[Almost sure convergence] \label{cor:as_convergence} Assume the hypotheses of Theorem~\ref{thm:deterministic_convergence}. Then \[ u_{\lambda,\boldsymbol{\omega}}^{\,n} \longrightarrow \eta \qquad\text{almost surely in }L^2(\mathbb{T}^m), \] that is, \[ \mathbb{P}\!\left( \lim_{n\to\infty} \|u_{\lambda,\boldsymbol{\omega}}^{\,n}-\eta\|_{L^2(\mathbb{T}^m)} =0 \right)=1. \] \end{corollary} \begin{proof} By Theorem~\ref{thm:deterministic_convergence}, \[ \lim_{n\to\infty} \|u_{\lambda,\boldsymbol{\omega}}^{\,n}-\eta\|_{L^2(\mathbb{T}^m)} =0. \] This convergence is deterministic, i.e., it holds for every realization of the data satisfying the assumptions. Hence it holds on the entire underlying probability space, and therefore also on a set of probability one. Consequently, \[ \mathbb{P}\!\left( \lim_{n\to\infty} \|u_{\lambda,\boldsymbol{\omega}}^{\,n}-\eta\|_{L^2(\mathbb{T}^m)} =0 \right)=1, \] which is precisely almost sure convergence. \end{proof} %%%%%%%%%%%%%%%%%%%%%%%%%%%% %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% \begin{remark} Although the analysis has been presented on the torus $\mathbb{T}^{m}$, the results extend immediately to any bounded Lipschitz domain $\Omega\subset\mathbb{R}^{m}$. Indeed, after an affine scaling, $\Omega$ may be regarded as a Lipschitz subdomain of the unit cube $(0,1)^m$, which is identified with the torus $\mathbb{T}^{m}$. The functions and the point density measure are then extended by zero to the complement $\mathbb{T}^{m}\setminus\Omega$. Since the extension preserves the required regularity on $\Omega$, all estimates established on $\mathbb{T}^{m}$ apply to the zero-extended functions. Restricting the resulting estimates back to $\Omega$ yields the same approximation, stability, and convergence results. Consequently, every theorem in this paper remains valid for arbitrary bounded Lipschitz domains after this scaling and zero-extension procedure. \end{remark} \section{Computational Fill Distance} Let $X$ be a random variable with values in a closed compact $\Omega \subset (0,1)^m$. Assume $\Omega$ has a Lipschitz boundary. Let $p(x)$ be the probability density function of $X$ and assume $p(x)>0\ \forall x\in\Omega$. Now we draw two independent sets of $n$ iid copies of $X$. We denote them $A_n = \{p_1,p_2\ldots p_n\}$ and $B_n = \{q_1,q_2\ldots q_n\}$. \definition Define $\xi_n$ as the computational fill distance of the data $(A_n,B_n)$, which is given as $$\xi_n = \max\limits_{x\in B_n}(d(x,A_n)) = \max\limits_{x\in B_n}\min\limits_{y\in A_n} \|x-y\|_2$$ The rate of convergence of the computational fill distance as the data become dense in a domain is derived here \cite{391340} and is given as \begin{equation} \begin{aligned}\label{eqil} \xi_n=O_P(n^{-1/m}\ln^{1/m}n) \end{aligned} \end{equation} In Theorem \ref{theorem_br} inequalities \ref{IQ_1} and \ref{IQ_3} are valid if we replace star discrepancy $\zeta$ with computational fill distance $\xi$, as the rate of convergence of computational fill distance is slower than that of star discrepancy. Hence Theorem \ref{theorem_br} holds if we replace star discrepancy with computational fill distance. %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% \section{Conclusion} In this paper, we studied the binary regression problem and established convergence results for the proposed estimator under suitable assumptions on the sampling points and the regression function. Error estimates and convergence results were obtained, providing a theoretical justification for the method and clarifying the role of the sampling density in the approximation process. We further showed that when the feature vectors possess a point density measure of bounded variation \cite{Dachiraju2026PointDensityMeasure}, the convergence theory can be strengthened from an expectation-based formulation to deterministic error estimates, yielding deterministic convergence and, consequently, almost sure convergence. This demonstrates that geometric regularity of the sampling points provides a natural framework for obtaining stronger convergence guarantees. Although the analysis was presented on the torus $\mathbb{T}^{m}$, the results extend to arbitrary bounded Lipschitz domains after an affine scaling, embedding into $\mathbb{T}^{m}$, and zero extension. Thus, the theory is applicable to a broad class of practical data domains. The results presented here provide a rigorous mathematical foundation for binary regression based on geometric sampling principles. Future work includes extending the analysis to multi-class classification, adaptive sampling strategies, and more general nonparametric learning problems. \section{Data Availability Statement} We declare that no data has been used in this document and none has been made available. %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% \section{Compliance with Ethical Standards} \subsection{Funding} This study was not funded by any grant or organization. \subsection{Conflict of Interest} The author Rajesh Dachiraju declares that he has no conflict of interest. \subsection{Ethical approval} This article does not contain any studies with human participants or animals performed by any of the authors. %\\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ \\ %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%% \bibliographystyle{amsplain} \bibliography{refs} \end{document}