Long context
Also known as: context window, long-context
30stories this week
32last 30 days
43all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceQwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysisSo far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceOMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
- Oct 10, 2026 · Opinion / analysis · 1 sourceI Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)[P]ALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.
- Oct 10, 2026 · Opinion / analysis · 1 sourceIs anyone running Qwen3.8 Flash Next with a 1M context?So, while doing a search to read the model card again, and because I didn't memorize the huggingface URL, I saw an AI "answer" at the top of the search, stating that while it natively supports a ~244k context, it could go to 1M using YaRN.
- Oct 8, 2026 · Research paper · 1 sourceVFold: Symmetry-Aware Cross-Layer Value Cache CompressionIn this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.
- Oct 8, 2026 · Research paper · 1 sourceRevisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative FlowsTo explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media.
- Oct 8, 2026 · Research paper · 1 sourceInternalizer: Portable Context-to-Parameter Mapping for Very Large Language ModelsWe present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work.
- Oct 8, 2026 · Research paper · 1 sourceBioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical TextTo address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies.
- Oct 8, 2026 · Research paper · 2 sourcesREMORY: Learning Residual Memory for Context CompactionWe introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens.
- Oct 8, 2026 · Research paper · 1 sourceRead What Matters: Query-Adaptive Quantization for KV CachesKV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.
- Oct 8, 2026 · Research paper · 1 sourceQUILT: Rethinking Sparse-Attention Prefill through Shared Query ExecutionWe present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
- Oct 7, 2026 · Research paper · 1 sourceLarge Language Models for Machine Translation Quality Annotation: Humans and Models Are Both ChallengedWe present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains.
- Oct 7, 2026 · Research paper · 2 sourcesReal Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than RecomputeWe test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
- Oct 7, 2026 · Research paper · 1 sourceRunningTab: Direct Workspace Interaction with Environment-Side TabsTo address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent.
- Oct 7, 2026 · Research paper · 1 sourceHeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video UnderstandingTo close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery.
- Oct 7, 2026 · Research paper · 1 sourceMechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid PositionThe architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension.
- Oct 7, 2026 · Research paper · 1 sourceDual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV CachesWe introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict.
- Oct 7, 2026 · Research paper · 1 sourceDecoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context PollutionWe study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters.
- Oct 7, 2026 · Research paper · 1 sourceRelevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QADrawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set.
- Oct 6, 2026 · Research paper · 1 sourceBookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write ResolverMemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting".
- Oct 6, 2026 · Research paper · 1 sourceA Self-Pruning Transformer: Extreme KV-Cache Compression with Universal AttentionRecent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference.
- Oct 6, 2026 · Product / feature launch · 1 sourceEmbeddingGemma 2: an open, lightweight multimodal embedding modelEmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026 · Research paper · 1 sourceSPIN: Shadow Predictive Indexer for Sparse AttentionWe propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead.
- Oct 6, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.903-beta: New Browser + Voice CloningThis release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
- Oct 6, 2026 · Research paper · 1 sourcePHBA: Prefix-State Hybrid Block AttentionIn this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.
- Oct 6, 2026 · Research paper · 1 sourceUNREAL: Unifying Retrieval and Long-Context with a Single ModelWe introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference.
- Oct 6, 2026 · Research paper · 1 sourceMacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric VideosAudio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference.
- Oct 6, 2026 · Research paper · 1 sourceHybrid Latent Attention for Looped Language ModelsWe propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values.
- Oct 6, 2026 · Research paper · 1 sourceReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon AgentsTo overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context.
- Oct 6, 2026 · Research paper · 1 sourcePersistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can TellDecomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds.