Technique

Vision Transformer

Also known as: ViT

12stories this week
12last 30 days
22all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    Healthy Counterfactual Generation via Diffusion Inpainting for Mammography Classification
    To address this issue, we propose a counterfactual data augmentation strategy that generates healthy mammograms by "erasing" lesions from anomalous images, thereby enriching the training distribution.
  2. Oct 8, 2026 · Research paper · 1 source
    SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
    Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3.
  3. Oct 8, 2026 · Research paper · 1 source
    Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
    Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.
  4. Oct 8, 2026 · Research paper · 1 source
    Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study
    Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior.
  5. Oct 8, 2026 · Research paper · 1 source
    MCL: Meta Convolution Layer
    Dynamic convolution enhances convolutional neural networks (CNNs) by adapting kernels to input content, but it expresses the effective kernel as a linear mixture of a small number of basis kernels, which limits expressivity and complicates optimization as the mixture size grows.
  6. Oct 7, 2026 · Research paper · 1 source
    On the Necessity of Attention-FFN Split in Vision Transformers
    In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs).
  7. Oct 7, 2026 · Research paper · 1 source
    ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
    Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count.
  8. Oct 7, 2026 · Research paper · 1 source
    Lightweight and Versatile Learned Optimization by Recombination of Gradient History
    This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans.
  9. Oct 6, 2026 · Research paper · 1 source
    Data Leakage in Patch-Based Hyperspectral Image Classification: Quantifying the Impact of Spatial Overlap
    Patch-based learning improves hyperspectral image (HSI) classification by exploiting local spectral-spatial information, but random train-test sampling from the same image can cause spatial patch overlap, leading to data leakage and optimistic performance estimates.
  10. Oct 6, 2026 · Open-source release · 1 source
    huggingface/transformers v5.19.0: Release v5.19.0
    EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
  11. Oct 6, 2026 · Research paper · 1 source
    Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration
    We propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters.
  12. Oct 6, 2026 · Research paper · 1 source
    Compact Robot Policies Need Fine-Grained Visual Representations
    To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator.
  13. Aug 26, 2026 · Open-source release · 1 source
    huggingface/transformers v5.16.0: Release: v5.16.0
    Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).
  14. Aug 10, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.0
    Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
  15. Aug 10, 2026 · Open-source release · 1 source
    huggingface/transformers v5.15.0: Release: v5.15.0
    Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases.
  16. Jul 27, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.26.0
    New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
  17. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
  18. Jun 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.23.0
    DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
  19. Jun 3, 2026 · Open-source release · 1 source
    huggingface/transformers v5.10.1: Release v5.10.1
    Sorry everyone, this happens when we rush a release!!!
  20. May 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.22.0
    DeepSeek V4 maturity: DeepSeek V4 received a major hardening pass this cycle — the model was reorganized into a dedicated vllm/models/deepseekv4/ package (#43004, #43039, #43073, #43077, #43149), gained NVFP4 fused MoE support (#42209), full + piecewise CUDA graph (#42604), and MTP speculative decoding (#43385).
  21. May 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.21.0
    Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
  22. Apr 3, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.19.0
    We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.

Often appears with