AION
Model

Qwen

139stories this week
157last 30 days
189all time

In the model registry

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    [P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]
    I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.
  2. Oct 11, 2026 · Opinion / analysis · 1 source
    PSA: DeepSeek V4.1 Flash habitually exfiltrates API keys. It is dangerously misaligned and may be hazardous to use
    EDIT: since people keep calling it out, this is API key abuse but not exfiltration.
  3. Oct 11, 2026 · Opinion / analysis · 1 source
    Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis
    So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
  4. Oct 11, 2026 · Open-source release · 1 source
    Converting dense models into Mixture-of-Experts
    For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
  5. Oct 11, 2026 · Tutorial / explainer · 1 source
    Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards
    This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.
  6. Oct 11, 2026 · Opinion / analysis · 1 source
    Qwen3.8 Flash Next fixed my GNOME extension
    I love Dash2Dock Lite, but Icedman is always a week or two before updates.
  7. Oct 11, 2026 · Tutorial / explainer · 1 source
    Building a 4x R9700 setup for a 10 person startup
    Just wanted to share a build I am doing for a client.
  8. Oct 11, 2026 · Tutorial / explainer · 1 source
    OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!
    I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
  9. Oct 10, 2026 · Opinion / analysis · 1 source
    Benefits of using bigger models than Qwen 3.8 flash next?
    Qwen 3.8 27b was the first model I tried on Ninfer at NVFP4 and then shifted to Flash next after seeing issues with 27b such as not willing to yield to instructions set in AGENTS.md or agent skills.
  10. Oct 10, 2026 · Opinion / analysis · 1 source
    Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work
    Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
  11. Oct 10, 2026 · Opinion / analysis · 1 source
    Is anyone running Qwen3.8 Flash Next with a 1M context?
    So, while doing a search to read the model card again, and because I didn't memorize the huggingface URL, I saw an AI "answer" at the top of the search, stating that while it natively supports a ~244k context, it could go to 1M using YaRN.
  12. Oct 10, 2026 · Opinion / analysis · 1 source
    48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)
    Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
  13. Oct 10, 2026 · Opinion / analysis · 1 source
    Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
    I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
  14. Oct 10, 2026 · Opinion / analysis · 1 source
    Improve token per second without touching quant
    Spent the past month tweaking and experimenting with many different numbers to achieve 30tps.
  15. Oct 10, 2026 · Opinion / analysis · 1 source
    Strata with Qwen3.8 Flash Next UD-Q4_K_XL
    Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.
  16. Oct 10, 2026 · Opinion / analysis · 1 source
    Help With Choosing Hardware [P]
    I am going to fine-tune a satire model, with the base model being Qwen3.5-14B-Base.
  17. Oct 9, 2026 · Model release · 2 sources
    Qwen/Qwen-Image-2.1-Turbo
    Qwen published the model Qwen-Image-2.1-Turbo on Hugging Face.
  18. Oct 9, 2026 · Open-source release · 1 source
    Cloudflare/clef-omni
    Cloudflare published the model clef-omni on Hugging Face.
  19. Oct 8, 2026 · Research paper · 1 source
    Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
    To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
  20. Oct 8, 2026 · Research paper · 1 source
    FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
    Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios.
  21. Oct 8, 2026 · Research paper · 2 sources
    OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
    We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
  22. Oct 8, 2026 · Research paper · 1 source
    WOVEN: Weaving Visual World Modeling into Multimodal LLMs
    We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
  23. Oct 8, 2026 · Research paper · 2 sources
    SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
    Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
  24. Oct 8, 2026 · Research paper · 1 source
    GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving
    Multimodal large language models (MLLMs) often struggle to identify and use geometric relations in diagrams.
  25. Oct 8, 2026 · Research paper · 1 source
    Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
    Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
  26. Oct 8, 2026 · Research paper · 2 sources
    SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
    The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
  27. Oct 8, 2026 · Research paper · 1 source
    HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
    To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning.
  28. Oct 8, 2026 · Research paper · 1 source
    Language Models as AI Research World Models
    AI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement.
  29. Oct 8, 2026 · Research paper · 2 sources
    VibeEdit: Image Editing with Canvas Instructions
    We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
  30. Oct 8, 2026 · Research paper · 1 source
    Recursive Self-Improvement through Multi-Agent Self-Supervision
    To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories.

Often appears with