AION
Research paperReinforcement Learning · Large Language Models · Training & Scaling1 source · Oct 7, 2026

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.

Key points

  • Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation.
  • Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier.
  • We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements.
  • Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems.

Sources (1)

  • [1]From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:46 AM
    To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.
    Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation.

Extractive summary: sentences quoted from the sources.