ResearchResearch paperLarge Language Models · Training & Scaling · Efficiency & Inference1 source · Oct 6, 2026

Forecast Accuracy Is Not Trading Profit: Evolving Small Recurrent Networks for Stock Return Prediction

We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy.

Key points

  • Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for, and a lower forecast error does not imply a better decision downstream.
  • A parallel debate asks whether modern transformer architectures forecast better than recurrent and other lightweight models.
  • Across four mid-cap portfolios and three trading years, the evolved networks rank first on both forecast accuracy and net trading performance, while the second most accurate model loses money once positions are formed and costs are charged.
  • They are also the cheapest end to end: a CPU-only search of 16 minutes yields 66-weight networks that predict in 10.8 $μ$s on a Raspberry Pi Zero, against transformer baselines of up to 817,153 parameters that require GPU training.

Sources (1)

  • [1]Forecast Accuracy Is Not Trading Profit: Evolving Small Recurrent Networks for Stock Return Prediction
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:20 AM
    We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy.
    Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for, and a lower forecast error does not imply a better decision downstream.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 5, 2026LiquidAI/d1-omni-600M
  2. Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
  3. Sep 30, 2026Cloudflare/clef-flash
  4. Sep 29, 2026microsoft/AesCode-32B
  5. Sep 29, 2026Language Models for Text Classification: From Bag-of-Words to Jev
  6. Aug 26, 2026vllm-project/vllm v0.28.0

Related