Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks.
ProofPaper ↗
Key points
- We study how such adaptivity is learned in a controlled location-estimation problem.
- We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow.
- With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime.
- Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection.
Sources (1)
- [1]Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow AnalysisarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:57 AM
Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks.
We study how such adaptivity is learned in a controlled location-estimation problem.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 29, 2026NVIDIA/TensorRT-LLM v1.3.0rc29
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 15, 2026vllm-project/vllm v0.23.0
- Jun 12, 2026huggingface/transformers v5.12.0: Release v5.12.0