Qwen
139stories this week
157last 30 days
189all time
In the model registry
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 source[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.
- Oct 11, 2026 · Opinion / analysis · 1 sourcePSA: DeepSeek V4.1 Flash habitually exfiltrates API keys. It is dangerously misaligned and may be hazardous to useEDIT: since people keep calling it out, this is API key abuse but not exfiltration.
- Oct 11, 2026 · Opinion / analysis · 1 sourceQwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysisSo far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
- Oct 11, 2026 · Open-source release · 1 sourceConverting dense models into Mixture-of-ExpertsFor the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceRunning Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boardsThis will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.
- Oct 11, 2026 · Opinion / analysis · 1 sourceQwen3.8 Flash Next fixed my GNOME extensionI love Dash2Dock Lite, but Icedman is always a week or two before updates.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceBuilding a 4x R9700 setup for a 10 person startupJust wanted to share a build I am doing for a client.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceOMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
- Oct 10, 2026 · Opinion / analysis · 1 sourceBenefits of using bigger models than Qwen 3.8 flash next?Qwen 3.8 27b was the first model I tried on Ninfer at NVFP4 and then shifted to Flash next after seeing issues with 27b such as not willing to yield to instructions set in AGENTS.md or agent skills.
- Oct 10, 2026 · Opinion / analysis · 1 sourceEngineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering workModels: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
- Oct 10, 2026 · Opinion / analysis · 1 sourceIs anyone running Qwen3.8 Flash Next with a 1M context?So, while doing a search to read the model card again, and because I didn't memorize the huggingface URL, I saw an AI "answer" at the top of the search, stating that while it natively supports a ~244k context, it could go to 1M using YaRN.
- Oct 10, 2026 · Opinion / analysis · 1 source48Gb VRAM speed AND quality ! (Qwen 3.8 27B Swift 1.5 W8A16)Because sometimes you need both speed AND quality, I made my own Qwen 3.8 27B Swift 1.5 quant.
- Oct 10, 2026 · Opinion / analysis · 1 sourceQwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernelI've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
- Oct 10, 2026 · Opinion / analysis · 1 sourceImprove token per second without touching quantSpent the past month tweaking and experimenting with many different numbers to achieve 30tps.
- Oct 10, 2026 · Opinion / analysis · 1 sourceStrata with Qwen3.8 Flash Next UD-Q4_K_XLMost of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.
- Oct 10, 2026 · Opinion / analysis · 1 sourceHelp With Choosing Hardware [P]I am going to fine-tune a satire model, with the base model being Qwen3.5-14B-Base.
- Oct 9, 2026 · Model release · 2 sourcesQwen/Qwen-Image-2.1-TurboQwen published the model Qwen-Image-2.1-Turbo on Hugging Face.
- Oct 9, 2026 · Open-source release · 1 sourceCloudflare/clef-omniCloudflare published the model clef-omni on Hugging Face.
- Oct 8, 2026 · Research paper · 1 sourceRubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-DistillationTo this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
- Oct 8, 2026 · Research paper · 1 sourceFastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios.
- Oct 8, 2026 · Research paper · 2 sourcesOneSearch-VL: Unified Multimodal Deep Research Agent for Image and VideoWe introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
- Oct 8, 2026 · Research paper · 1 sourceWOVEN: Weaving Visual World Modeling into Multimodal LLMsWe therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
- Oct 8, 2026 · Research paper · 2 sourcesSpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language ModelsExisting spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
- Oct 8, 2026 · Research paper · 1 sourceGeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem SolvingMultimodal large language models (MLLMs) often struggle to identify and use geometric relations in diagrams.
- Oct 8, 2026 · Research paper · 1 sourceWhich Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-EvolutionSkills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
- Oct 8, 2026 · Research paper · 2 sourcesSparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM InferenceThe memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
- Oct 8, 2026 · Research paper · 1 sourceHarnessSQL: Harness-Native Training for SQL Agents in Realistic Database EnvironmentsTo bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning.
- Oct 8, 2026 · Research paper · 1 sourceLanguage Models as AI Research World ModelsAI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement.
- Oct 8, 2026 · Research paper · 2 sourcesVibeEdit: Image Editing with Canvas InstructionsWe introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
- Oct 8, 2026 · Research paper · 1 sourceRecursive Self-Improvement through Multi-Agent Self-SupervisionTo address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories.