ResearchResearch paperLarge Language Models · Reasoning & Planning · Efficiency & Inference1 source · Oct 8, 2026

Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog

We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog.

Key points

  • Current frontier LLMs can theoretically process long contexts with 1M tokens or more.
  • But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts?
  • We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length.
  • We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related