Transformer
Also known as: transformer architecture, transformers
134stories this week
141last 30 days
182all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceI trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokensHiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€
- Oct 11, 2026 · Opinion / analysis · 1 sourceHas machine learning research gotten more "wordy"? [D]I feel like I am unable to digest most machine learning research these days.
- Oct 9, 2026 · Open-source release · 1 sourceCloudflare/clef-omniCloudflare published the model clef-omni on Hugging Face.
- Oct 8, 2026 · Research paper · 2 sourcesOne Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed ExpertsIn this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
- Oct 8, 2026 · Research paper · 2 sourcesLEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video GenerationGenerating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
- Oct 8, 2026 · Research paper · 1 sourceLeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPCWe introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes.
- Oct 8, 2026 · Research paper · 1 sourceMarformer: A Transformer for Predicting Missing Data DistributionsWe present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values.
- Oct 8, 2026 · Research paper · 1 sourcePLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action PoliciesLearning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.
- Oct 8, 2026 · Research paper · 1 sourceAdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution DriftAccurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained models rapidly.
- Oct 8, 2026 · Research paper · 1 sourceVerification with Transfer: Exact Information Frontiers and Their Price in CallsA verifier that accepts or rejects whole answers reveals little: under a flat prior over $k$-bit answers, zero error needs $2^k-1$ verifications.
- Oct 8, 2026 · Research paper · 1 sourceSciTBERT: A family of chronologically consistent language models for scientific and technological language processingWe introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025.
- Oct 8, 2026 · Research paper · 1 sourceUnifying Policy Learning and State Prediction through Spatial Language ModelingWe introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens.
- Oct 8, 2026 · Research paper · 1 sourceStructure Tax: How Structured Output affects LLMs PerformanceDeploying large language models in production often requires constraining outputs to structured formats such as JSON or XML, and prior work treats the resulting accuracy loss as an inherent structure tax'.
- Oct 8, 2026 · Research paper · 1 sourceTACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot LearningTo address this problem, we present TACROSS, a scalable system for learning from human touch and transferring it to robots that bridges this heterogeneity by aligning tactile streams at the level of contact events rather than raw sensor values.
- Oct 8, 2026 · Research paper · 1 sourceSoftmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It MustThis work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies.
- Oct 8, 2026 · Research paper · 1 sourceEasy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching missesByte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
- Oct 8, 2026 · Research paper · 1 sourceCorrelational Training of Morphological Neural NetworksIn this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme.
- Oct 8, 2026 · Research paper · 1 sourceTimer-M1: A Multivariate Time Series Foundation Model via Learning PrimitivesWe introduce Timer-M1, a pretrained multivariate time series foundation model that learns with primitives for zero-shot forecasting.
- Oct 8, 2026 · Research paper · 1 source4-Tensor Attention Model for Semantic Physical RealityWe describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning.
- Oct 8, 2026 · Research paper · 1 source$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.
- Oct 8, 2026 · Research paper · 1 sourceWhere to Adapt Matters: Layer-Selective Fine-Tuning for Capability RetentionParameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.
- Oct 8, 2026 · Research paper · 1 sourceAdapting English Quality Classifiers for Multilingual LLM Pretraining Data SelectionRecent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance.
- Oct 8, 2026 · Research paper · 1 sourceSDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous DrivingWe present SDPAD, a fully spike-driven end-to-end planning pipeline that closes this gap.
- Oct 8, 2026 · Research paper · 1 sourceSV-TAD: Native Sparse Convs for Efficient Temporal Action DetectionTo adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules.
- Oct 8, 2026 · Research paper · 2 sourcesScaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention ResidualsIn this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update.
- Oct 8, 2026 · Research paper · 1 sourceFrom Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped TransformersWe introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace.
- Oct 8, 2026 · Research paper · 1 sourceZatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across DomainsTo this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets.
- Oct 8, 2026 · Research paper · 1 sourceCausal-fate dynamics of unrealized influenceHere we formulate causal-fate dynamics, in which generated influence may be realized, remain latent, or be transformed by subsequent dynamics, and give an exact finite-transport representation when the relevant maps are specified.
- Oct 8, 2026 · Research paper · 1 sourceRewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream TransformerVision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).
- Oct 8, 2026 · Research paper · 1 sourceDeflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion TransformersIn diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.