Evaluating Trajectory Features for Routing Final-Layer Attention
Attention routing requires a signal that predicts the value of attention on the current prefix.
Key points
- We evaluate whether hidden-state extrapolation error, curvature and error change improve this prediction beyond uncertainty, one-step displacement, position and state projections.
- Paired executions of the final attention layer supply signed next-token loss differences in frozen SmolLM3-3B-Base and Qwen3.5-4B-Base checkpoints.
- In Qwen3.5, a parameter-matched fixed-projection control lowers NLL by 0.00356 nats/token relative to the trajectory router (95 percent interval 0.00218 to 0.00487).
- Actual selected-query execution yields small long-sequence latency reductions with increased NLL, while learned routers remain slower during cached continuation.
Sources (1)
- [1]Evaluating Trajectory Features for Routing Final-Layer AttentionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:12 AM
Attention routing requires a signal that predicts the value of attention on the current prefix.
We evaluate whether hidden-state extrapolation error, curvature and error change improve this prediction beyond uncertainty, one-step displacement, position and state projections.
Extractive summary: sentences quoted from the sources.