Llama
Also known as: LLaMA
6stories this week
6last 30 days
6all time
Timeline
- Oct 11, 2026 · Tutorial / explainer · 1 sourceRunning Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boardsThis will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.
- Oct 8, 2026 · Research paper · 1 sourceSpectral Weight Decay: Inducing Low-Rank Structure in Neural Network WeightsWe introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.
- Oct 8, 2026 · Research paper · 1 sourceLadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMsLifelong editing of LLMs requires storing thousands of edits after acquisition.
- Oct 7, 2026 · Research paper · 1 sourceDocument-Level Text Simplification in Estonian Using Large Language ModelsDespite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored.
- Oct 6, 2026 · Research paper · 1 sourceFew Bits, One Law: Toward W2A4KV2We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation.
- Oct 6, 2026 · Research paper · 1 sourceHow Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault AnalysisAs small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern.