Conditional independence testing for discrete distributions: Beyond χ2-and G-tests

Ilmun Kim, Matey Neykov, Sivaraman Balakrishnan, Larry Wasserman

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

This paper is concerned with the problem of conditional independence testing for discrete data. In recent years, researchers have shed new light on this fundamental problem, emphasizing finite-sample optimality. The non-asymptotic viewpoint adapted in these works has led to novel conditional independence tests that enjoy certain optimality under various regimes. Despite their attractive theoretical properties, the considered tests are not necessarily practical, relying on a Poissonization trick and unspecified constants in their critical values. In this work, we attempt to bridge the gap between theory and practice by reproving optimality without Poissonization and calibrating tests using Monte Carlo permutations. Along the way, we also prove that classical asymptotic χ2-and G-tests are notably sub-optimal in a high-dimensional regime, which justifies the demand for new tools. Our theoretical results are complemented by experiments on both simulated and real-world datasets. Accompanying this paper is an R package UCI that implements the proposed tests.

Original languageEnglish (US)
Pages (from-to)4767-4794
Number of pages28
JournalElectronic Journal of Statistics
Volume18
Issue number2
DOIs
StatePublished - 2024

Funding

We would like to thank the reviewers for their thoughtful comments that significantly improved our paper. This work was partially supported by funding from the NSF grants DMS-2113684 and DMS-2310632, as well as an Amazon AI and a Google Research Scholar Award to SB. MN acknowledges support from the NSF grant DMS-2113684. IK acknowledges support from the Yonsei University Research Fund of 2023-22-0419 as well as support from the Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (2022R1A4A1033384), and the Korea government (MSIT) RS-2023-00211073.

Keywords

  • Depoissonization
  • conditional independence
  • negative association
  • permutation tests
  • sample complexity

ASJC Scopus subject areas

  • Statistics and Probability
  • Statistics, Probability and Uncertainty

Fingerprint

Dive into the research topics of 'Conditional independence testing for discrete distributions: Beyond χ2-and G-tests'. Together they form a unique fingerprint.

Cite this