AION
Research paperEfficiency & Inference · Training & Scaling1 source · Oct 6, 2026

Cleave: Scaling Tensor Program Optimization via Decoupled Algebraic Search and Operator Scheduling

We propose Cleave, an ML compiler built on symbolic decoupling: Cleave discovers transformations by performing superoptimization on a graph with symbolic shapes, and then schedules each resulting graph on concrete shapes.

Key points

  • Optimized kernels such as FlashAttention and FlashDecoding are crucial for accelerating today's large models.
  • Producing such kernels requires fusing computations with multiple reductions, which requires both algebraic transformation of the computation graph and operator scheduling of the transformed graph.
  • Evaluation on common LLM subgraphs shows that Cleave generates kernels up to 2.8x faster than the best baseline (1.6x on average) and reduces compilation time by 5.9x on average compared to Mirage.
  • For dynamic workloads captured from production serving traces, Cleave compiles each operator once and achieves geometric mean speedups of 1.4x and 1.7x over FlashInfer's handwritten FA2 and FA3 backends.

Sources (1)

  • [1]Cleave: Scaling Tensor Program Optimization via Decoupled Algebraic Search and Operator Scheduling
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:37 AM
    We propose Cleave, an ML compiler built on symbolic decoupling: Cleave discovers transformations by performing superoptimization on a graph with symbolic shapes, and then schedules each resulting graph on concrete shapes.
    Optimized kernels such as FlashAttention and FlashDecoding are crucial for accelerating today's large models.

Extractive summary: sentences quoted from the sources.