Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses
Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
Key points
- BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict.
- We show that this rule has a systematic blind spot: positions whose type is predictable but whose value must be computed, such as the number after "=" in a worked math solution.
- The entropy trigger of Scratchpad Patching is likewise indistinguishable from random scratchpads on final answers (5.6% vs 5.9%, 5 seeds), while answer-start scratchpads give 38.1%.
- Boundary dependence, the rise in the model's own loss when a patch start is removed, measured per two-byte context, finds these positions without labels: combined with entropy it beats the hand-written rule on computed results.
Sources (1)
- [1]Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching missesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:03 PM
Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict.
Extractive summary: sentences quoted from the sources.