AION
Technique

Sparse autoencoders

Also known as: SAE, SAEs, sparse autoencoder

6stories this week
6last 30 days
6all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
    Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.
  2. Oct 8, 2026 · Research paper · 1 source
    Regularized Small Area Estimation with Graph Laplacian Benchmarking priors
    The resulting family includes a Benchmarking Prior that incorporates the benchmarking restrictions without additional regularization across areas, and Single and Multi-View Laplacian Benchmarking Priors that introduce regularization through graph Laplacians constructed from area similarities based on external covariate information.
  3. Oct 7, 2026 · Research paper · 1 source
    Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
    Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
  4. Oct 7, 2026 · Research paper · 1 source
    Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models
    Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models.
  5. Oct 6, 2026 · Research paper · 1 source
    AnyBottle: A Recipe to Only Keep the Concepts You Really Need
    We propose AnyBottle, a single recipe for building compact, task-specific CBMs. AnyBottle assumes only a frozen backbone and an unsupervised concept pool, such as a sparse autoencoder.
  6. Oct 6, 2026 · Research paper · 1 source
    Illusory Pattern Perception Drives Spurious Inference in Large Language Models
    Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random.

Often appears with