Embedding-Bias in Conditional Independence Testing
To test conditional independence of $X$ and $Y$ given a text or an image $Z$, one conditions on an embedding $ψ(Z)$ in place of $Z$.
Key points
- The embedded test is valid if $Z$ is independent of $X$ or of $Y$ given $ψ(Z)$, which cannot be confirmed from data, and when this fails, the rejection probability under the null hypothesis can tend to one.
- We study this failure, and show that focusing on a specific form of dependence relaxes what the embedding must retain.
- For a residual correlation test inspired by the Generalised Covariance Measure, validity only requires that the parts of $\mathbb{E}[X \mid Z]$ and $\mathbb{E}[Y \mid Z]$ missed by $\mathbb{E}[X \mid ψ(Z)]$ and $\mathbb{E}[Y \mid ψ(Z)]$ are uncorrelated.
- On synthetic data and text embeddings, the robust test holds its level approximately.
Sources (1)
- [1]Embedding-Bias in Conditional Independence TestingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:33 AM
To test conditional independence of $X$ and $Y$ given a text or an image $Z$, one conditions on an embedding $ψ(Z)$ in place of $Z$.
The embedded test is valid if $Z$ is independent of $X$ or of $Y$ given $ψ(Z)$, which cannot be confirmed from data, and when this fails, the rejection probability under the null hypothesis can tend to one.
Extractive summary: sentences quoted from the sources.