Nearly Minimax Rates for Functional Estimation Under Rough Random Design

Primarily AI-generated textHuman understanding: some partsmath.ST — Statistics Theory

Contributed by P. M. Aronow ↗, Nathan Kallus ↗, Patrick Lopatto ↗claimed

A claim means an ORCID-authenticated account linked itself to the matching contributor entry. It grants no control of the record and is not a Hexagon endorsement.

Version 1 / Oct 08, 2026 / CC BY 4.0

Abstract

We establish nearly minimax bounds for missing-at-random means, treatment effects, and expected conditional covariances under rough random design. For two nuisance functions with average H\"older smoothness ss in dimension dd, the minimax root-mean-square error is n−2s/d+o(1)n^{-2s/d+o(1)} when s<d/4s<d/4 and of order n−1/2n^{-1/2} when s≥d/4s\ge d/4. The first rate confirms the rough-design exponent suggested by higher-order influence function theory. In the generic model with an unknown bounded density weight, our upper and lower bounds differ by only polylogarithmic factors and identify a leading correction e−κlog⁡ne^{-\kappa\sqrt{\log n}} when s<d/4s<d/4, with κ\kappa explicit in terms of s/ds/d and the weight bounds. For the expected conditional covariance, the same exponent and the same constant κ\kappa were obtained independently and concurrently by S. Park (arXiv:2610.05006). For models with separately bounded density and propensity, we identify the same polynomial exponent and the explicit leading correction, with an o(log⁡n)o(\sqrt{\log n}) remainder in the logarithm of the risk.

Provenance statement

Proof discovery and verification were done with a combination of GPT 5.6 Sol, GPT 6 Sol, GPT 6 Astra, GPT 6.1 Sol, Claude Fable 5 and 5.1, and Claude Opus 5.5. Lean formalizations were produced by GPT 6.1 Sol as part of verification. An early version of the Abstract and Introduction was drafted by humans, and earlier version of the proofs had edits performed by humans for clarity. The final version, involving stronger results, was substantially redrafted by Opus 5.5 and GPT 6 Astra.

Tools used

Anthropic
Claude OpusVersion 5.5
Anthropic
Claude FableVersion 5.1
Anthropic
Claude FableVersion 5
OpenAI
ChatGPTVersion 6.1
OpenAI
ChatGPTVersion 6
OpenAI
ChatGPTVersion 5.6
Lean
LeanVersion 4.34.1

References

  1. B. K. Alpert. A class of bases in L2L^2 for the sparse representation of integral operators. SIAM Journal on Mathematical Analysis , 24 0 (1): 0 246--262, 1993. ọi 10.1137/0524016 .DOI
  2. J. D. Angrist, G. W. Imbens, and D. B. Rubin. Identification of causal effects using instrumental variables. Journal of the American Statistical Association , 91 0 (434): 0 444--455, 1996. ọi 10.1080/01621459.1996.10476902 .DOI
  3. P. M. Aronow and P. Lopatto. Nearly-minimax variance estimation under rough random design, 2026. https://arxiv.org/abs/2607.13170v3 arXiv:2607.13170v3 .arXiv
  4. S. Balakrishnan, E. H. Kennedy, and L. Wasserman. The fundamental limits of structure-agnostic functional estimation. Statistical Science , 41 0 (3): 0 659--670, 2026. ọi 10.1214/25-STS997 .DOI
  5. P. J. Bickel and Y. Ritov. Estimating integrated squared density derivatives: Sharp best order of convergence estimates. Sankhy ā : The Indian Journal of Statistics, Series A , 50 0 (3): 0 381--393, 1988.
  6. L. Birg é and P. Massart. Estimation of integral functionals of a density. The Annals of Statistics , 23 0 (1): 0 11--29, 1995. ọi 10.1214/aos/1176324452 .DOI
  7. M. Bonvini, E. H. Kennedy, O. Dukes, and S. Balakrishnan. Doubly-robust inference and optimality in structure-agnostic models with smoothness, 2024. https://arxiv.org/abs/2405.08525v2 arXiv:2405.08525v2 .arXiv
  8. T. T. Cai and M. G. Low. Testing composite hypotheses, Hermite polynomials and optimal estimation of a nonsmooth functional. The Annals of Statistics , 39 0 (2): 0 1012--1041, 2011. ọi 10.1214/10-AOS849 .DOI
  9. T. T. Cai, M. Levine, and L. Wang. Variance function estimation in multivariate nonparametric regression with fixed design. Journal of Multivariate Analysis , 100 0 (1): 0 126--136, 2009. ọi 10.1016/j.jmva.2008.03.007 .DOI
  10. X. Chen, L. Liu, and R. Mukherjee. Method-of-moments inference for GLMs and doubly robust functionals under proportional asymptotics, 2025. https://arxiv.org/abs/2408.06103v3 arXiv:2408.06103v3 .arXiv
  11. V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal , 21 0 (1): 0 C1--C68, 2018. ọi 10.1111/ectj.12097 .DOI
  12. V. Chernozhukov, C. Hansen, N. Kallus, M. Spindler, and V. Syrgkanis. Applied Causal Inference Powered by ML and AI . Online, 2026. URL r̆l https://causalml-book.org/ . Version 0.1.2, 3 May 2026.link
  13. R. K. Crump, V. J. Hotz, G. W. Imbens, and O. A. Mitnik. Moving the goalposts: Addressing limited overlap in estimation of average treatment effects by changing the estimand. IZA Discussion Paper 2347, Institute for the Study of Labor (IZA), Bonn, 2006. https://docs.iza.org/dp2347.pdf docs.iza.org/dp2347.pdf .link
  14. E. Dobriban, R. Mukherjee, J. M. Robins, and Z. Wang. Improved variance estimation in homoskedastic nonparametric random-design regression via a two-scale approach, 2026. https://arxiv.org/abs/2609.08783v1 arXiv:2609.08783v1 .arXiv
  15. D. Evans and A. J. Jones. Non-parametric estimation of residual moments and covariance. Proceedings of the Royal Society A , 464 0 (2099): 0 2831--2846, 2008. ọi 10.1098/rspa.2007.0195 .DOI
  16. M. Fr ö lich. Nonparametric IV estimation of local average treatment effects with covariates. Journal of Econometrics , 139 0 (1): 0 35--75, 2007. ọi 10.1016/j.jeconom.2006.06.004 .DOI
  17. P. R. Halmos. The theory of unbiased estimation. The Annals of Mathematical Statistics , 17 0 (1): 0 34--43, 1946. ọi 10.1214/aoms/1177731020 .DOI
  18. M. A. Hern á n and J. M. Robins. Causal Inference: What If . Chapman & Hall/CRC, Boca Raton, 2020.
  19. W. Hoeffding. A class of statistics with asymptotically normal distribution. The Annals of Mathematical Statistics , 19 0 (3): 0 293--325, 1948. ọi 10.1214/aoms/1177730196 .DOI
  20. W. Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association , 58 0 (301): 0 13--30, 1963. ọi 10.1080/01621459.1963.10500830 .DOI
  21. G. W. Imbens and J. D. Angrist. Identification and estimation of local average treatment effects. Econometrica , 62 0 (2): 0 467--475, 1994. ọi 10.2307/2951620 .DOI
  22. Yu. I. Ingster and I. A. Suslina. Nonparametric Goodness-of-Fit Testing Under Gaussian Models , volume 169 of Lecture Notes in Statistics . Springer, New York, 2003. ọi 10.1007/978-0-387-21580-8 .DOI
  23. J. Jiao, K. Venkat, Y. Han, and T. Weissman. Minimax estimation of functionals of discrete distributions. IEEE Transactions on Information Theory , 61 0 (5): 0 2835--2885, 2015. ọi 10.1109/TIT.2015.2412945 .DOI
  24. J. Jin and V. Syrgkanis. Structure-agnostic optimality of doubly robust learning for treatment effect estimation (extended abstract). In Proceedings of the Thirty Eighth Conference on Learning Theory , volume 291 of Proceedings of Machine Learning Research , pages 3159--3160, 2025 a . Full paper: https://arxiv.org/abs/2402.14264v4 arXiv:2402.14264v4 .arXiv
  25. J. Jin and V. Syrgkanis. Sharp structure-agnostic lower bounds for general linear functional estimation, 2025 b . https://arxiv.org/abs/2512.17341v2 arXiv:2512.17341v2 .arXiv
  26. E. H. Kennedy, S. Balakrishnan, and L. Wasserman. Discussion of `` On nearly assumption-free tests of nominal confidence interval coverage for causal parameters estimated by machine learning''. Statistical Science , 35 0 (3): 0 540--544, 2020. ọi 10.1214/20-STS796 .DOI
  27. G. Kerkyacharian and D. Picard. Estimating nonquadratic functionals of a density using Haar wavelets. The Annals of Statistics , 24 0 (2): 0 485--507, 1996. ọi 10.1214/aos/1032894450 .DOI
  28. W. Kong and G. Valiant. Estimating learnability in the sublinear data regime. In Advances in Neural Information Processing Systems , volume 31, pages 5455--5464, 2018. Extended version: https://arxiv.org/abs/1805.01626v3 arXiv:1805.01626v3 .arXiv
  29. J. K. Kraus, P. S. Vassilevski, and L. T. Zikatanov. Polynomial of best uniform approximation to 1/x1/x and smoothing in two-level methods. Computational Methods in Applied Mathematics , 12 0 (4): 0 448--468, 2012. ọi 10.2478/cmam-2012-0026 .DOI
  30. G. Last and M. Penrose. Lectures on the Poisson Process , volume 7 of Institute of Mathematical Statistics Textbooks . Cambridge University Press, Cambridge, 2017. ọi 10.1017/9781316104477 .DOI
  31. G. Last and M. D. Penrose. Poisson process Fock space representation, chaos expansion and covariance inequalities. Probability Theory and Related Fields , 150 0 (3--4): 0 663--690, 2011. ọi 10.1007/s00440-010-0288-5 .DOI
  32. L. Le Cam. Asymptotic Methods in Statistical Decision Theory . Springer Series in Statistics. Springer, New York, 1986. ọi 10.1007/978-1-4612-4946-7 .DOI
  33. O. Lepski, A. Nemirovski, and V. Spokoiny. On estimation of the LrL_r norm of a regression function. Probability Theory and Related Fields , 113 0 (2): 0 221--253, 1999. ọi 10.1007/s004409970006 .DOI
  34. J. Levy, M. van der Laan, A. Hubbard, and R. Pirracchio. A fundamental measure of treatment effect heterogeneity. Journal of Causal Inference , 9 0 (1): 0 83--108, 2021. ọi 10.1515/jci-2019-0003 .DOI
  35. F. Li, K. L. Morgan, and A. M. Zaslavsky. Balancing covariates via propensity score weighting. Journal of the American Statistical Association , 113 0 (521): 0 390--400, 2018. ọi 10.1080/01621459.2016.1260466 .DOI
  36. L. Li, E. Tchetgen Tchetgen, A. van der Vaart, and J. M. Robins. Higher order inference on a treatment effect under low regularity conditions. Statistics & Probability Letters , 81 0 (7): 0 821--828, 2011. ọi 10.1016/j.spl.2011.02.030 .DOI
  37. L. Liu, R. Mukherjee, W. K. Newey, and J. M. Robins. Semiparametric efficient empirical higher order influence function estimators, 2017. https://arxiv.org/abs/1705.07577v5 arXiv:1705.07577v5 .arXiv
  38. L. Liu, R. Mukherjee, and J. M. Robins. Rejoinder: On nearly assumption-free tests of nominal confidence interval coverage for causal parameters estimated by machine learning. Statistical Science , 35 0 (3): 0 545--554, 2020. ọi 10.1214/20-STS804 .DOI
  39. L. Liu, R. Mukherjee, J. M. Robins, and E. Tchetgen Tchetgen. Adaptive estimation of nonparametric functionals. Journal of Machine Learning Research , 22 0 (99): 0 1--66, 2021. URL r̆l https://jmlr.org/papers/v22/19-892.html .link
  40. L. Liu, R. Mukherjee, and J. M. Robins. On the asymptotic inadmissibility of double machine learning estimators under structure-agnostic models, 2026 a . https://arxiv.org/abs/2606.22391v2 arXiv:2606.22391v2 .arXiv
  41. N. Liu, C. Li, Y. Gu, and L. Liu. Stabilized higher-order influence functions: Statistical theory of a class of bilinear forms, 2026 b . https://arxiv.org/abs/2607.04743v3 arXiv:2607.04743v3 .arXiv
  42. R. J. Mathar. Chebyshev series expansion of inverse polynomials. Journal of Computational and Applied Mathematics , 196 0 (2): 0 596--607, 2006. ọi 10.1016/j.cam.2005.10.013 .DOI
  43. A. McClean, S. Balakrishnan, E. H. Kennedy, and L. Wasserman. Double cross-fit doubly robust estimators: Beyond series regression. Journal of the Royal Statistical Society Series B: Statistical Methodology , 88 0 (4): 0 1469--1491, 2026. ọi 10.1093/jrsssb/qkag057 . Theorem numbering refers to https://arxiv.org/abs/2403.15175v3 arXiv:2403.15175v3 .arXivDOI
  44. S. McGrath and R. Mukherjee. Nuisance function tuning and sample splitting for optimally estimating a doubly robust functional, 2022. To appear in The Annals of Statistics . https://arxiv.org/abs/2212.14857v5 arXiv:2212.14857v5 .arXiv
  45. W. K. Newey and J. M. Robins. Cross-fitting and fast remainder rates for semiparametric estimation, 2018. https://arxiv.org/abs/1801.09138v1 arXiv:1801.09138v1 .arXiv
  46. S. Park. Minimax estimation of the expected conditional covariance under bounds on the covariate density, 2026. https://arxiv.org/abs/2610.05006v1 arXiv:2610.05006v1 .arXiv
  47. T. S. Richardson and A. Rotnitzky. Causal etiology of the research of James M. Robins . Statistical Science , 29 0 (4): 0 459--484, 2014. ọi 10.1214/14-STS505 .DOI
  48. J. Robins, L. Li, E. Tchetgen, and A. van der Vaart. Higher order influence functions and minimax estimation of nonlinear functionals. In Probability and Statistics: Essays in Honor of David A. Freedman , volume 2 of IMS Collections , pages 335--421. Institute of Mathematical Statistics, 2008. ọi 10.1214/193940307000000527 .DOI
  49. J. Robins, E. Tchetgen Tchetgen, L. Li, and A. van der Vaart. Semiparametric minimax rates. Electronic Journal of Statistics , 3: 0 1305--1321, 2009. ọi 10.1214/09-EJS479 .DOI
  50. J. M. Robins and Y. Ritov. Toward a curse of dimensionality appropriate ( CODA ) asymptotic theory for semi-parametric models. Statistics in Medicine , 16 0 (3): 0 285--319, 1997. ọi 10.1002/(SICI)1097-0258(19970215)16:3<285::AID-SIM535>3.0.CO;2-# .DOI
  51. J. M. Robins, A. Rotnitzky, and L. P. Zhao. Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association , 89 0 (427): 0 846--866, 1994. ọi 10.1080/01621459.1994.10476818 .DOI
  52. J. M. Robins, L. Li, L. Liu, R. Mukherjee, E. Tchetgen Tchetgen, and A. van der Vaart. Minimax estimation of a functional on a structured high-dimensional model (corrected version), 2015. https://arxiv.org/abs/1512.02174v3 arXiv:1512.02174v3 .arXiv
  53. J. M. Robins, L. Li, R. Mukherjee, E. Tchetgen Tchetgen, and A. van der Vaart. Minimax estimation of a functional on a structured high-dimensional model. The Annals of Statistics , 45 0 (5): 0 1951--1987, 2017. ọi 10.1214/16-AOS1515 .DOI
  54. P. M. Robinson. Root- NN -consistent semiparametric regression. Econometrica , 56 0 (4): 0 931--954, 1988. ọi 10.2307/1912705 .DOI
  55. A. Rotnitzky, E. Smucler, and J. M. Robins. Characterization of parameters with a mixed bias property. Biometrika , 108 0 (1): 0 231--238, 2021. ọi 10.1093/biomet/asaa054 .DOI
  56. D. B. Rubin. Inference and missing data. Biometrika , 63 0 (3): 0 581--592, 1976. ọi 10.1093/biomet/63.3.581 .DOI
  57. A. S á nchez-Becerra. Robust inference for the treatment effect variance in experiments using machine learning, 2023. https://arxiv.org/abs/2306.03363v1 arXiv:2306.03363v1 .arXiv
  58. R. J. Serfling. Approximation Theorems of Mathematical Statistics . Wiley, New York, 1980. ọi 10.1002/9780470316481 .DOI
  59. Y. Shen, C. Gao, D. Witten, and F. Han. Optimal estimation of variance in nonparametric regression with random design. The Annals of Statistics , 48 0 (6): 0 3589--3618, 2020. ọi 10.1214/20-AOS1944 .DOI
  60. X. Song. Batched and complete U -statistics for trace-polynomial estimation from classical shadows, 2026. https://arxiv.org/abs/2608.22962v1 arXiv:2608.22962v1 .arXiv
  61. E. M. Stein and R. Shakarchi. Fourier Analysis: An Introduction , volume 1 of Princeton Lectures in Analysis . Princeton University Press, Princeton, 2003.
  62. C. J. Stone. Optimal global rates of convergence for nonparametric regression. The Annals of Statistics , 10 0 (4): 0 1040--1053, 1982. ọi 10.1214/aos/1176345969 .DOI
  63. L. Tian, A. A. Alizadeh, A. J. Gentles, and R. Tibshirani. A simple method for estimating interactions between a treatment and a large number of covariates. Journal of the American Statistical Association , 109 0 (508): 0 1517--1532, 2014. ọi 10.1080/01621459.2014.951443 .DOI
  64. L. N. Trefethen. Approximation Theory and Approximation Practice . SIAM, Philadelphia, extended edition, 2019. ọi 10.1137/1.9781611975949 .DOI
  65. L. N. Trefethen and D. Bau, III. Numerical Linear Algebra . SIAM, Philadelphia, 1997. ọi 10.1137/1.9780898719574 .DOI
  66. A. B. Tsybakov. Introduction to Nonparametric Estimation . Springer Series in Statistics. Springer, New York, 2009. ọi 10.1007/b13794 .DOI
  67. A. W. van der Vaart. Asymptotic Statistics . Cambridge University Press, Cambridge, 1998. ọi 10.1017/CBO9780511802256 .DOI
  68. S. Vatedka, N. Kashyap, and A. Thangaraj. Secure compute-and-forward in a bidirectional relay. IEEE Transactions on Information Theory , 61 0 (5): 0 2531--2556, 2015. ọi 10.1109/TIT.2015.2412114 .DOI
  69. N. Verzelen and E. Gassiat. Adaptive estimation of high-dimensional signal-to-noise ratios. Bernoulli , 24 0 (4B): 0 3683--3710, 2018. ọi 10.3150/17-BEJ975 .DOI
  70. L. Wang and E. Tchetgen Tchetgen. Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables. Journal of the Royal Statistical Society Series B: Statistical Methodology , 80 0 (3): 0 531--550, 2018. ọi 10.1111/rssb.12262 .DOI
  71. L. Wang, L. D. Brown, T. T. Cai, and M. Levine. Effect of mean on variance function estimation in nonparametric regression. The Annals of Statistics , 36 0 (2): 0 646--664, 2008. ọi 10.1214/009053607000000901 .DOI
  72. B. D. Williamson, P. B. Gilbert, N. R. Simon, and M. Carone. A general framework for inference on algorithm-agnostic variable importance. Journal of the American Statistical Association , 118 0 (543): 0 1645--1658, 2023. ọi 10.1080/01621459.2021.2003200 .DOI
  73. Y. Wu and P. Yang. Minimax rates of entropy estimation on large alphabets via best polynomial approximation. IEEE Transactions on Information Theory , 62 0 (6): 0 3702--3720, 2016. ọi 10.1109/TIT.2016.2548468 .DOI
  74. Y. Wu and P. Yang. Chebyshev polynomials, moment matching, and optimal estimation of the unseen. The Annals of Statistics , 47 0 (2): 0 857--883, 2019. ọi 10.1214/17-AOS1665 .DOI
  75. Z. Zeng, S. Balakrishnan, Y. Han, and E. H. Kennedy. Causal inference with high-dimensional discrete covariates, 2024. https://arxiv.org/abs/2405.00118v3 arXiv:2405.00118v3 .arXiv
  76. Y. Zhang, L. Liu, and Z. Zhang. Higher-order debiased estimators for general treatment models. Econometric Theory , pages 1--43, 2026. ọi 10.1017/S0266466626100516 . First View; remark numbering refers to https://arxiv.org/abs/2606.01706v2 arXiv:2606.01706v2 .arXivDOI

Version history

  1. v1Submitted by P. M. AronowInitial depositCurrentOct 08, 2026