AION
Model

SmolLM

4stories this week
4last 30 days
4all time

In the model registry

Timeline

  1. Oct 11, 2026 · Open-source release · 1 source
    Converting dense models into Mixture-of-Experts
    For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
  2. Oct 7, 2026 · Research paper · 1 source
    SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning
    We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B.
  3. Oct 7, 2026 · Research paper · 1 source
    BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
    We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
  4. Oct 7, 2026 · Research paper · 1 source
    Evaluating Trajectory Features for Routing Final-Layer Attention
    Attention routing requires a signal that predicts the value of attention on the current prefix.

Often appears with