A Systematic Study of Small Language Models on Abstract Reasoning Tasks
Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities.
ProofPaper ↗
Key points
- We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer.
- Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning.
- We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences.
- Executable-rule induction also yields correct solutions not observed under direct grid generation.
Sources (1)
- [1]A Systematic Study of Small Language Models on Abstract Reasoning TasksarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:01 PM
Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities.
We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
- Oct 6, 2026Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
- Oct 5, 2026Sharing AI progress in mathematics
- Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model