AION
Technique

Transformer

Also known as: transformer architecture, transformers

134stories this week
141last 30 days
182all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens
    Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€
  2. Oct 11, 2026 · Opinion / analysis · 1 source
    Has machine learning research gotten more "wordy"? [D]
    I feel like I am unable to digest most machine learning research these days.
  3. Oct 9, 2026 · Open-source release · 1 source
    Cloudflare/clef-omni
    Cloudflare published the model clef-omni on Hugging Face.
  4. Oct 8, 2026 · Research paper · 2 sources
    One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
    In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
  5. Oct 8, 2026 · Research paper · 2 sources
    LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
    Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
  6. Oct 8, 2026 · Research paper · 1 source
    LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC
    We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes.
  7. Oct 8, 2026 · Research paper · 1 source
    Marformer: A Transformer for Predicting Missing Data Distributions
    We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values.
  8. Oct 8, 2026 · Research paper · 1 source
    PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
    Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.
  9. Oct 8, 2026 · Research paper · 1 source
    AdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution Drift
    Accurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained models rapidly.
  10. Oct 8, 2026 · Research paper · 1 source
    Verification with Transfer: Exact Information Frontiers and Their Price in Calls
    A verifier that accepts or rejects whole answers reveals little: under a flat prior over $k$-bit answers, zero error needs $2^k-1$ verifications.
  11. Oct 8, 2026 · Research paper · 1 source
    SciTBERT: A family of chronologically consistent language models for scientific and technological language processing
    We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025.
  12. Oct 8, 2026 · Research paper · 1 source
    Unifying Policy Learning and State Prediction through Spatial Language Modeling
    We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens.
  13. Oct 8, 2026 · Research paper · 1 source
    Structure Tax: How Structured Output affects LLMs Performance
    Deploying large language models in production often requires constraining outputs to structured formats such as JSON or XML, and prior work treats the resulting accuracy loss as an inherent structure tax'.
  14. Oct 8, 2026 · Research paper · 1 source
    TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning
    To address this problem, we present TACROSS, a scalable system for learning from human touch and transferring it to robots that bridges this heterogeneity by aligning tactile streams at the level of contact events rather than raw sensor values.
  15. Oct 8, 2026 · Research paper · 1 source
    Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
    This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.
  16. Oct 8, 2026 · Research paper · 1 source
    Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses
    Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
  17. Oct 8, 2026 · Research paper · 1 source
    Correlational Training of Morphological Neural Networks
    In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme.
  18. Oct 8, 2026 · Research paper · 1 source
    Timer-M1: A Multivariate Time Series Foundation Model via Learning Primitives
    We introduce Timer-M1, a pretrained multivariate time series foundation model that learns with primitives for zero-shot forecasting.
  19. Oct 8, 2026 · Research paper · 1 source
    4-Tensor Attention Model for Semantic Physical Reality
    We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning.
  20. Oct 8, 2026 · Research paper · 1 source
    $σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
    Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.
  21. Oct 8, 2026 · Research paper · 1 source
    Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
    Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.
  22. Oct 8, 2026 · Research paper · 1 source
    Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection
    Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance.
  23. Oct 8, 2026 · Research paper · 1 source
    SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving
    We present SDPAD, a fully spike-driven end-to-end planning pipeline that closes this gap.
  24. Oct 8, 2026 · Research paper · 1 source
    SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection
    To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules.
  25. Oct 8, 2026 · Research paper · 2 sources
    Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
    In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update.
  26. Oct 8, 2026 · Research paper · 1 source
    From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers
    We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.
  27. Oct 8, 2026 · Research paper · 1 source
    Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains
    To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets.
  28. Oct 8, 2026 · Research paper · 1 source
    Causal-fate dynamics of unrealized influence
    Here we formulate causal-fate dynamics, in which generated influence may be realized, remain latent, or be transformed by subsequent dynamics, and give an exact finite-transport representation when the relevant maps are specified.
  29. Oct 8, 2026 · Research paper · 1 source
    Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
    Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).
  30. Oct 8, 2026 · Research paper · 1 source
    Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
    In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.

Often appears with