Vision Transformer
Also known as: ViT
12stories this week
12last 30 days
22all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceHealthy Counterfactual Generation via Diffusion Inpainting for Mammography ClassificationTo address this issue, we propose a counterfactual data augmentation strategy that generates healthy mammograms by "erasing" lesions from anomalous images, thereby enriching the training distribution.
- Oct 8, 2026 · Research paper · 1 sourceSAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language DecodersBuilding on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3.
- Oct 8, 2026 · Research paper · 1 sourceRethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text EncoderBuilding on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.
- Oct 8, 2026 · Research paper · 1 sourceDissecting Representation Structure in Vision Transformers: A Rigorous Architectural StudyRepresentation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior.
- Oct 8, 2026 · Research paper · 1 sourceMCL: Meta Convolution LayerDynamic convolution enhances convolutional neural networks (CNNs) by adapting kernels to input content, but it expresses the effective kernel as a linear mixture of a small number of basis kernels, which limits expressivity and complicates optimization as the mixture size grows.
- Oct 7, 2026 · Research paper · 1 sourceOn the Necessity of Attention-FFN Split in Vision TransformersIn this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs).
- Oct 7, 2026 · Research paper · 1 sourceORCA: Hunting Compositional Failures in Text-to-Image DiffusionText-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.
- Oct 7, 2026 · Research paper · 1 sourceLightweight and Versatile Learned Optimization by Recombination of Gradient HistoryThis paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans.
- Oct 6, 2026 · Research paper · 1 sourceData Leakage in Patch-Based Hyperspectral Image Classification: Quantifying the Impact of Spatial OverlapPatch-based learning improves hyperspectral image (HSI) classification by exploiting local spectral-spatial information, but random train-test sampling from the same image can cause spatial patch overlap, leading to data leakage and optimistic performance estimates.
- Oct 6, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.19.0: Release v5.19.0EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
- Oct 6, 2026 · Research paper · 1 sourceTest-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned RecalibrationWe propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters.
- Oct 6, 2026 · Research paper · 1 sourceCompact Robot Policies Need Fine-Grained Visual RepresentationsTo test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator.
- Aug 26, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.16.0: Release: v5.16.0Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Aug 10, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.15.0: Release: v5.15.0Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases.
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.23.0DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
- Jun 3, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.10.1: Release v5.10.1Sorry everyone, this happens when we rush a release!!!
- May 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.22.0DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
- May 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.21.0Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
- Apr 3, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.19.0We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.