Nearly Minimax Variance Estimation Under Rough Random Design

Primarily AI-generated textHuman understanding: some partsmath.ST — Statistics Theory

Contributed by P. M. Aronow ↗, Patrick Lopatto ↗

Version 1 / Oct 08, 2026 / CC BY 4.0

Abstract

We determine, up to a power of log⁡n\log n, the minimax risk for constant conditional variance estimation under rough random design. The unknown design density is bounded above and away from zero, with no smoothness assumption, and the conditional error laws may depend on the covariates and have uniformly bounded fourth moments. For an ss-Hölder regression function with s>1s>1 in dimension d>4sd>4s, the minimax root-mean-square risk lies, for all sufficiently large nn, between cΨnc\Psi_n and CΨn(log⁡n)ΓC\Psi_n(\log n)^{\Gamma}, where Ψn=n−2(s+1)/(d+4)e−κlog⁡n(log⁡n)(s−1)/(d+4)\Psi_n=n^{-2(s+1)/(d+4)}e^{-\kappa\sqrt{\log n}}(\log n)^{(s-1)/(d+4)} and the constants κ>0\kappa>0 and Γ>0\Gamma>0 are explicit. In particular, the minimax exponent is 2(s+1)/(d+4)2(s+1)/(d+4), the minimax risk is smaller than n−2(s+1)/(d+4)n^{-2(s+1)/(d+4)} by a stretched-exponential factor whose constant κ\kappa is identified, and the rate proposed by Robins is not uniformly attainable over this model class. For 0<s≤10<s\le1, we show that the exact minimax rate is n−1/2∨n−4s/(d+4s)n^{-1/2}\vee n^{-4s/(d+4s)}; for s>1s>1 and d≤4sd\le4s, it is n−1/2n^{-1/2}.

Provenance statement

Proof discovery and verification were done with a combination of GPT 5.4, GPT 5.5, GPT 5.6 Sol, GPT 6 Sol, GPT 6 Astra, GPT 6.1 Sol, Claude Fable 5, Claude Fable 5.1, and Claude Opus 5.5. Lean 4.34.1 formalizations were produced by GPT 6.1 Sol as part of verification. Prior versions of the paper involved substantial human drafting and editing, but the final version, involving stronger results, was substantially redrafted by Opus 5.5, GPT 6.1 Sol, and GPT 6 Astra.

Tools used

Lean
LeanVersion 4.34.1
Anthropic
Claude FableVersion 5.1
Anthropic
Claude FableVersion 5
OpenAI
ChatGPTVersion 6.1
OpenAI
ChatGPTVersion 6
OpenAI
ChatGPTVersion 5.6

References

  1. Aronow, P. M.; Lopatto, P. On rates attainable under random design: A negative answer to a problem of Robins. 2026. \hrefhttps://arxiv.org/abs/2607.13170v1arXiv:2607.13170v1
  2. Aronow, P. M.; Lopatto, P. On rates attainable under random design: A negative answer to a problem of Robins. 2026. \hrefhttps://arxiv.org/abs/2607.13170v2arXiv:2607.13170v2
  3. Aronow, P. M.; Lopatto, P. Nearly-minimax variance estimation under rough random design. 2026. \hrefhttps://arxiv.org/abs/2607.13170v3arXiv:2607.13170v3
  4. Aronow, P. M.; Kallus, N.; Lopatto, P. Nearly minimax rates for functional estimation under rough random design. 2026. \hrefhttps://hexagonmath.org/2610.00157?version=1hexagon:2610.00157v1
  5. Brown, L. D.; Levine, M. Variance estimation in nonparametric regression via the difference sequence method. The Annals of Statistics, vol. 35, no. 5, pp. 2219–2232. 2007DOI
  6. Brown, L. D.; Low, M. G. A constrained risk inequality with applications to nonparametric functional estimation. The Annals of Statistics, vol. 24, no. 6, pp. 2524–2535. 1996DOI
  7. Cai, T. T.; Low, M. G. Testing composite hypotheses, Hermite polynomials and optimal estimation of a nonsmooth functional. The Annals of Statistics, vol. 39, no. 2, pp. 1012–1041. 2011DOI
  8. Cai, T. T.; Levine, M.; Wang, L. Variance function estimation in multivariate nonparametric regression with fixed design. Journal of Multivariate Analysis, vol. 100, no. 1, pp. 126–136. 2009DOI
  9. Cartier, P.; Foata, D. Problèmes Combinatoires de Commutation et Réarrangements. Springer, vol. 85. 1969DOI
  10. Chen, X.; Liu, L.; Mukherjee, R. Method-of-moments inference for GLMs and doubly robust functionals under proportional asymptotics. 2025. \hrefhttps://arxiv.org/abs/2408.06103v3arXiv:2408.06103v3
  11. Dette, H.; Munk, A.; Wagner, T. Estimating the variance in nonparametric regression–what is a reasonable choice?. Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 60, no. 4, pp. 751–764. 1998DOI
  12. NIST Digital Library of Mathematical Functions. 2026. F. W. J. Olver, A. B. Olde Daalhuis, D. W. Lozier, B. I. Schneider, R. F. Boisvert, C. W. Clark, B. R. Miller, B. V. Saunders, H. S. Cohl, and M. A. McClain, eds.
  13. Dobriban, E.; Mukherjee, R.; Robins, J. M.; Wang, Z. Improved variance estimation in homoskedastic nonparametric random-design regression via a two-scale approach. 2026. \hrefhttps://arxiv.org/abs/2609.08783v1arXiv:2609.08783v1
  14. Dobrushin, R. L. Estimates of semi-invariants for the Ising model at low temperatures. Topics in Statistical and Theoretical Physics, vol. 177, pp. 59–81. 1996DOI
  15. Donoho, D. L.; Nussbaum, M. Minimax quadratic estimation of a quadratic functional. Journal of Complexity, vol. 6, no. 3, pp. 290–323. 1990DOI
  16. Du, J.; Schick, A. A covariate-matched estimator of the error variance in nonparametric regression. Journal of Nonparametric Statistics, vol. 21, no. 3, pp. 263–285. 2009DOI
  17. Fan, J.; Yao, Q. Efficient estimation of conditional variance functions in stochastic regression. Biometrika, vol. 85, no. 3, pp. 645–660. 1998DOI
  18. Gasser, T.; Sroka, L.; Jennen-Steinmetz, C. Residual variance and residual pattern in nonlinear regression. Biometrika, vol. 73, no. 3, pp. 625–633. 1986DOI
  19. Giné, E.; Nickl, R. A simple adaptive estimator of the integrated square of a density. Bernoulli, vol. 14, no. 1, pp. 47–61. 2008DOI
  20. Hall, P.; Marron, J. S. On variance estimation in nonparametric regression. Biometrika, vol. 77, no. 2, pp. 415–419. 1990DOI
  21. Hall, P.; Kay, J. W.; Titterington, D. M. Asymptotically optimal difference-based estimation of variance in nonparametric regression. Biometrika, vol. 77, no. 3, pp. 521–528. 1990DOI
  22. Hoeffding, W. A class of statistics with asymptotically normal distribution. The Annals of Mathematical Statistics, vol. 19, no. 3, pp. 293–325. 1948DOI
  23. Huang, L.-S.; Fan, J. Nonparametric estimation of quadratic regression functionals. Bernoulli, vol. 5, no. 5, pp. 927–949. 1999DOI
  24. Jiao, J.; Venkat, K.; Han, Y.; Weissman, T. Minimax estimation of functionals of discrete distributions. 2015. \hrefhttps://arxiv.org/abs/1406.6956v5arXiv:1406.6956v5
  25. Kong, W.; Valiant, G. Estimating learnability in the sublinear data regime. Advances in Neural Information Processing Systems, vol. 31, pp. 5455–5464. 2018. Locators refer to \hrefhttps://arxiv.org/abs/1805.01626v3arXiv:1805.01626v3link
  26. Kraus, J. K.; Vassilevski, P. S.; Zikatanov, L. T. Polynomial of best uniform approximation to x−1x^-1 and smoothing in two-level methods. 2012. \hrefhttps://arxiv.org/abs/1002.1859v3arXiv:1002.1859v3
  27. Last, G.; Penrose, M. Lectures on the Poisson Process. Cambridge University Press, vol. 7. 2018DOI
  28. Laurent, B. Efficient estimation of integral functionals of a density. The Annals of Statistics, vol. 24, no. 2, pp. 659–681. 1996DOI
  29. Lehmann, E. L.; Casella, G. Theory of Point Estimation. Springer. 1998DOI
  30. Lepski, O.; Nemirovski, A.; Spokoiny, V. On estimation of the LrL_r norm of a regression function. Probability Theory and Related Fields, vol. 113, no. 2, pp. 221–253. 1999DOI
  31. Liu, L.; Mukherjee, R.; Robins, J. M. Rejoinder: On nearly assumption-free tests of nominal confidence interval coverage for causal parameters estimated by machine learning. Statistical Science, vol. 35, no. 3, pp. 545–554. 2020. Section~4 is cited from \hrefhttps://arxiv.org/abs/2008.03288v1arXiv:2008.03288v1DOI
  32. Liu, L.; Mukherjee, R.; Robins, J. M.; Tchetgen Tchetgen, E. Adaptive estimation of nonparametric functionals. Journal of Machine Learning Research, vol. 22, no. 99, pp. 1–66. 2021link
  33. Mason, J. C.; Handscomb, D. C. Chebyshev Polynomials. Chapman & Hall/CRC. 2003
  34. Mathar, R. J. Chebyshev series expansion of inverse polynomials. Journal of Computational and Applied Mathematics, vol. 196, no. 2, pp. 596–607. 2006. Locators refer to \hrefhttps://arxiv.org/abs/math/0403344v4arXiv:math/0403344v4DOI
  35. McClean, A.; Balakrishnan, S.; Kennedy, E. H.; Wasserman, L. Double cross-fit doubly robust estimators: Beyond series regression. Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 88, no. 4, pp. 1469–1491. 2026. Locators refer to \hrefhttps://arxiv.org/abs/2403.15175v3arXiv:2403.15175v3DOI
  36. McDiarmid, C. On the method of bounded differences. Surveys in Combinatorics, 1989, pp. 148–188. 1989DOI
  37. Mukherjee, R.; Newey, W. K.; Robins, J. M. Semiparametric efficient empirical higher order influence function estimators. 2017. \hrefhttps://arxiv.org/abs/1705.07577v1arXiv:1705.07577v1
  38. Müller, U. U.; Schick, A.; Wefelmeyer, W. Estimating the error variance in nonparametric regression by a covariate-matched U-statistic. Statistics, vol. 37, no. 3, pp. 179–188. 2003DOI
  39. Munk, A.; Bissantz, N.; Wagner, T.; Freitag, G. On difference-based variance estimation in nonparametric regression when the covariate is high dimensional. Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 67, no. 1, pp. 19–41. 2005DOI
  40. Newey, W. K.; Robins, J. M. Cross-fitting and fast remainder rates for semiparametric estimation. 2018. \hrefhttps://arxiv.org/abs/1801.09138v1arXiv:1801.09138v1
  41. O'Reilly, E.; Tran, N. M. Stochastic geometry to generalize the Mondrian process. SIAM Journal on Mathematics of Data Science, vol. 4, no. 2, pp. 531–552. 2022. Locators refer to \hrefhttps://arxiv.org/abs/2002.00797v3arXiv:2002.00797v3DOI
  42. Park, S. Minimax estimation of the expected conditional covariance under bounds on the covariate density. 2026. \hrefhttps://arxiv.org/abs/2610.05006v1arXiv:2610.05006v1
  43. Rice, J. Bandwidth choice for nonparametric regression. The Annals of Statistics, vol. 12, no. 4, pp. 1215–1230. 1984DOI
  44. Richardson, T. S.; Rotnitzky, A. Causal etiology of the research of James M. Robins. Statistical Science, vol. 29, no. 4, pp. 459–484. 2014DOI
  45. Rivlin, T. J. Chebyshev Polynomials: From Approximation Theory to Algebra and Number Theory. Wiley. 1990
  46. Robins, J.; Li, L.; Tchetgen, E.; van der Vaart, A. Higher order influence functions and minimax estimation of nonlinear functionals. Probability and Statistics: Essays in Honor of David A. Freedman, vol. 2, pp. 335–421. 2008DOI
  47. Robins, J.; Tchetgen Tchetgen, E.; Li, L.; van der Vaart, A. Semiparametric minimax rates. Electronic Journal of Statistics, vol. 3, pp. 1305–1321. 2009DOI
  48. Robins, J. M.; Li, L.; Mukherjee, R.; Tchetgen Tchetgen, E.; van der Vaart, A. Minimax estimation of a functional on a structured high-dimensional model. The Annals of Statistics, vol. 45, no. 5, pp. 1951–1987. 2017DOI
  49. Scott, A. D.; Sokal, A. D. The repulsive lattice gas, the independent-set polynomial, and the Lovász local lemma. Journal of Statistical Physics, vol. 118, no. 5–6, pp. 1151–1261. 2005. Locators refer to \hrefhttps://arxiv.org/abs/cond-mat/0309352v2arXiv:cond-mat/0309352v2DOI
  50. Shen, Y.; Gao, C.; Witten, D.; Han, F. Optimal estimation of variance in nonparametric regression with random design. The Annals of Statistics, vol. 48, no. 6, pp. 3589–3618. 2020. Locators refer to \hrefhttps://arxiv.org/abs/1902.10822v2arXiv:1902.10822v2DOI
  51. Song, X. Batched and complete U-statistics for trace-polynomial estimation from classical shadows. 2026. \hrefhttps://arxiv.org/abs/2608.22962v1arXiv:2608.22962v1
  52. Spokoiny, V. Variance estimation for high-dimensional regression models. Journal of Multivariate Analysis, vol. 82, no. 1, pp. 111–133. 2002. Locators refer to the \hrefhttps://www.wias-berlin.de/people/spokoiny/publications/6_Spokoiny_a4_02/jmva99n.pdfauthor manuscriptDOI
  53. Stanley, R. P. Two poset polytopes. Discrete & Computational Geometry, vol. 1, no. 1, pp. 9–23. 1986DOI
  54. van der Vaart, A. Higher order tangent spaces and influence functions. Statistical Science, vol. 29, no. 4, pp. 679–686. 2014DOI
  55. Viennot, G. X. Heaps of pieces, I: Basic definitions and combinatorial lemmas. Combinatoire énumérative, vol. 1234, pp. 321–350. 1986DOI
  56. Wang, L.; Brown, L. D.; Cai, T. T.; Levine, M. Effect of mean on variance function estimation in nonparametric regression. The Annals of Statistics, vol. 36, no. 2, pp. 646–664. 2008DOI
  57. Wu, Y.; Yang, P. Minimax rates of entropy estimation on large alphabets via best polynomial approximation. IEEE Transactions on Information Theory, vol. 62, no. 6, pp. 3702–3720. 2016DOI
  58. Wu, Y.; Yang, P. Chebyshev polynomials, moment matching, and optimal estimation of the unseen. The Annals of Statistics, vol. 47, no. 2, pp. 857–883. 2019. Locators refer to \hrefhttps://arxiv.org/abs/1504.01227v2arXiv:1504.01227v2DOI

Version history

  1. v1Submitted by Patrick LopattoInitial depositCurrentOct 08, 2026