ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

A Systematic Study of Small Language Models on Abstract Reasoning Tasks

Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities.

Key points

  • We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer.
  • Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning.
  • We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences.
  • Executable-rule induction also yields correct solutions not observed under direct grid generation.

Sources (1)

  • [1]A Systematic Study of Small Language Models on Abstract Reasoning Tasks
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:01 PM
    Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities.
    We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
  2. Oct 6, 2026Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
  3. Oct 5, 2026Sharing AI progress in mathematics
  4. Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
  5. Jun 10, 2026DiffusionGemma: 4x faster text generation
  6. Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Related