ImageNet
22stories this week
23last 30 days
23all time
Timeline
- Oct 8, 2026 · Research paper · 2 sourcesOne Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed ExpertsIn this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
- Oct 8, 2026 · Research paper · 1 sourceFew-Step Generation via Data-Space IterationFlow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations.
- Oct 8, 2026 · Research paper · 1 sourceRecovery Guarantees for Posterior Sampling of One-Bit Compressed SensingWe study the sample complexity of noisy one-bit compressed sensing for signals drawn from a prior distribution.
- Oct 8, 2026 · Research paper · 1 sourceDino Forcing Flow Models: Do not denoise what you can predictCo-denoising pretrained representations such as DINO can substantially improve the training speed and quality of flow matching models, but it introduces a second denoising trajectory and requires carefully designed schedules.
- Oct 8, 2026 · Research paper · 1 sourceHAND: A Biologically-Inspired Activation Function that Improves Generalisation and Sample Efficiency in Image ClassificationWe incorporate a biologically-inspired inductive bias into a new activation function, HAND (Homeostasis, Accelerating Nonlinearity, and Divisive-nomalisation), and show its effectiveness with CNNs trained on image classification.
- Oct 8, 2026 · Research paper · 1 sourceMCL: Meta Convolution LayerDynamic convolution enhances convolutional neural networks (CNNs) by adapting kernels to input content, but it expresses the effective kernel as a linear mixture of a small number of basis kernels, which limits expressivity and complicates optimization as the mixture size grows.
- Oct 7, 2026 · Research paper · 1 sourceVelocity Scaling in Flow MatchingScaling a learned flow-matching velocity field $vθ$ by a gain $γ(t)$ was recently shown to greatly improve generation quality.
- Oct 7, 2026 · Research paper · 1 sourceQuadTok: Quadtree Visual Tokenizer for Autoregressive Image GenerationWe introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation.
- Oct 7, 2026 · Research paper · 1 sourceKoopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow MeasurementsWe introduce an observation-corrected Koopman framework for accelerating frozen diffusion models.
- Oct 7, 2026 · Research paper · 1 sourceKinetic Langevin Meets Split Gibbs: Accelerated Posterior Sampling for Imaging Inverse Problems with Diffusion PriorsSplit Gibbs sampling (SGS) is a popular framework for posterior sampling in Bayesian imaging inverse problems.
- Oct 7, 2026 · Research paper · 1 sourceProgress and Prospect of AI in ARPES WorkflowArtificial intelligence (AI) is becoming an increasingly useful tool across the experimental sciences, including angle-resolved photoemission spectroscopy (ARPES), which routinely produces large, multidimensional datasets of electronic structure.
- Oct 7, 2026 · Research paper · 1 sourceDisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation ModelsWe introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone.
- Oct 7, 2026 · Research paper · 1 sourceBeyond Group Splits: Specimen-Level Cross-Validation and Visual Attribution for Remaining-Shelf-Life Regression in Climacteric FruitEstimating remaining shelf life (RSL) from images could provide affordable decision support for perishable produce, but evaluation protocols can substantially affect reported performance when repeated images are available from the same biological specimen.
- Oct 7, 2026 · Research paper · 1 sourceMRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific MaskingIn post-deployment time, inputs to deep learning models may or may not be adversarially patched.
- Oct 7, 2026 · Research paper · 1 sourcePooling Representation Autoencoders for Efficient DiffusionMotivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens.
- Oct 6, 2026 · Research paper · 1 sourceConsistent Distribution Matching for Data-Free Diffusion DistillationIn this work, we propose Consistent Distribution Matching, a simulation-free and data-free distillation method for accelerating diffusion and flow models while preserving strong generative capacity.
- Oct 6, 2026 · Research paper · 1 sourceCo-Evolving Paths and Flows via Path-Flow AlignmentWe study path-flow alignment as a unified training objective for flow matching.
- Oct 6, 2026 · Research paper · 1 sourceFrom the Drosophila Visual Connectome to General-Purpose Computer VisionWe develop ConnectomeX around FlyVision, a trainable architecture that preserves parallel ON/OFF processing, recurrent computation and population-level graph interaction while scaling model capacity across tasks.
- Oct 6, 2026 · Research paper · 1 sourceTest-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned RecalibrationWe propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters.
- Oct 6, 2026 · Research paper · 1 sourceCompact Robot Policies Need Fine-Grained Visual RepresentationsTo test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator.
- Oct 6, 2026 · Research paper · 1 sourceTwo Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image GenerationRecent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly.
- Oct 6, 2026 · Research paper · 1 sourceLater Is Better: Token Reduction for ViTs Under Distribution ShiftTraining-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute.
- Oct 1, 2026 · Model release · 1 sourcenvidia/PixelDiT2-ImageNetnvidia published the model PixelDiT2-ImageNet on Hugging Face.