ResearchResearch paperTraining & Scaling · Interpretability · Large Language Models1 source · Oct 6, 2026

Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks.

Key points

  • We study how such adaptivity is learned in a controlled location-estimation problem.
  • We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow.
  • With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime.
  • Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 29, 2026NVIDIA/TensorRT-LLM v1.3.0rc29
  2. Aug 10, 2026vllm-project/vllm v0.27.0
  3. Jul 11, 2026vllm-project/vllm v0.25.0
  4. Jun 29, 2026vllm-project/vllm v0.24.0
  5. Jun 15, 2026vllm-project/vllm v0.23.0
  6. Jun 12, 2026huggingface/transformers v5.12.0: Release v5.12.0

Related