Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog
We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog.
ProofPaper ↗
Key points
- Current frontier LLMs can theoretically process long contexts with 1M tokens or more.
- But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts?
- We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length.
- We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.
Sources (1)
- [1]Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with PrologarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:42 AM
We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog.
Current frontier LLMs can theoretically process long contexts with 1M tokens or more.
Extractive summary: sentences quoted from the sources.