Executing Causal Structure Learning with Linear-Attention Transformers
We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.
Key points
- Transformers can execute algorithms on data given in their input.
- We explicitly construct a fixed-weight transformer whose forward pass exactly reproduces one update of this method, so repeated blocks reproduce its optimization trajectory.
- We show that retaining the multiplier is essential for exact execution, since different multiplier values can lead to different next updates.
- Experiments show that the constructed block agrees with a reference update to floating-point precision, while arithmetic replay on synthetic data and seven published benchmark network topologies inherits the reference solver's successes and failures.
Sources (1)
- [1]Executing Causal Structure Learning with Linear-Attention TransformersarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:48 PM
We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity.
Transformers can execute algorithms on data given in their input.
Extractive summary: sentences quoted from the sources.