Sparse autoencoders
Also known as: SAE, SAEs, sparse autoencoder
6stories this week
6last 30 days
6all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceRouterInterp: Understanding Superposed Specialisation in Mixture of Experts RoutingLeveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations.
- Oct 8, 2026 · Research paper · 1 sourceRegularized Small Area Estimation with Graph Laplacian Benchmarking priorsThe resulting family includes a Benchmarking Prior that incorporates the benchmarking restrictions without additional regularization across areas, and Single and Multi-View Laplacian Benchmarking Priors that introduce regularization through graph Laplacians constructed from area similarities based on external covariate information.
- Oct 7, 2026 · Research paper · 1 sourceDisentangling Linguistic and Paralinguistic Information with Routed Sparse AutoencodersSelf-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
- Oct 7, 2026 · Research paper · 1 sourceSparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action ModelsVision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models.
- Oct 6, 2026 · Research paper · 1 sourceAnyBottle: A Recipe to Only Keep the Concepts You Really NeedWe propose AnyBottle, a single recipe for building compact, task-specific CBMs. AnyBottle assumes only a frozen backbone and an unsupervised concept pool, such as a sparse autoencoder.
- Oct 6, 2026 · Research paper · 1 sourceIllusory Pattern Perception Drives Spurious Inference in Large Language ModelsIllusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random.