AION
Concept

Long context

Also known as: context window, long-context

30stories this week
32last 30 days
43all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis
    So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
  2. Oct 11, 2026 · Tutorial / explainer · 1 source
    OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!
    I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
  3. Oct 10, 2026 · Opinion / analysis · 1 source
    I Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)[P]
    ALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.
  4. Oct 10, 2026 · Opinion / analysis · 1 source
    Is anyone running Qwen3.8 Flash Next with a 1M context?
    So, while doing a search to read the model card again, and because I didn't memorize the huggingface URL, I saw an AI "answer" at the top of the search, stating that while it natively supports a ~244k context, it could go to 1M using YaRN.
  5. Oct 8, 2026 · Research paper · 1 source
    VFold: Symmetry-Aware Cross-Layer Value Cache Compression
    In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.
  6. Oct 8, 2026 · Research paper · 1 source
    Revisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative Flows
    To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media.
  7. Oct 8, 2026 · Research paper · 1 source
    Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
    We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work.
  8. Oct 8, 2026 · Research paper · 1 source
    BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text
    To address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies.
  9. Oct 8, 2026 · Research paper · 2 sources
    REMORY: Learning Residual Memory for Context Compaction
    We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens.
  10. Oct 8, 2026 · Research paper · 1 source
    Read What Matters: Query-Adaptive Quantization for KV Caches
    KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.
  11. Oct 8, 2026 · Research paper · 1 source
    QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
    We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
  12. Oct 7, 2026 · Research paper · 1 source
    Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged
    We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains.
  13. Oct 7, 2026 · Research paper · 2 sources
    Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
    We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
  14. Oct 7, 2026 · Research paper · 1 source
    RunningTab: Direct Workspace Interaction with Environment-Side Tabs
    To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent.
  15. Oct 7, 2026 · Research paper · 1 source
    HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding
    To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery.
  16. Oct 7, 2026 · Research paper · 1 source
    Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
    The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension.
  17. Oct 7, 2026 · Research paper · 1 source
    Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches
    We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict.
  18. Oct 7, 2026 · Research paper · 1 source
    Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution
    We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters.
  19. Oct 7, 2026 · Research paper · 1 source
    Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA
    Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set.
  20. Oct 6, 2026 · Research paper · 1 source
    Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
    MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting".
  21. Oct 6, 2026 · Research paper · 1 source
    A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention
    Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference.
  22. Oct 6, 2026 · Product / feature launch · 1 source
    EmbeddingGemma 2: an open, lightweight multimodal embedding model
    EmbeddingGemma 2: an open, lightweight multimodal embedding model
  23. Oct 6, 2026 · Research paper · 1 source
    SPIN: Shadow Predictive Indexer for Sparse Attention
    We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead.
  24. Oct 6, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.903-beta: New Browser + Voice Cloning
    This release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
  25. Oct 6, 2026 · Research paper · 1 source
    PHBA: Prefix-State Hybrid Block Attention
    In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.
  26. Oct 6, 2026 · Research paper · 1 source
    UNREAL: Unifying Retrieval and Long-Context with a Single Model
    We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference.
  27. Oct 6, 2026 · Research paper · 1 source
    MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos
    Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference.
  28. Oct 6, 2026 · Research paper · 1 source
    Hybrid Latent Attention for Looped Language Models
    We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values.
  29. Oct 6, 2026 · Research paper · 1 source
    ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents
    To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context.
  30. Oct 6, 2026 · Research paper · 1 source
    Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell
    Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds.

Often appears with