AION
Technique

LoRA

Also known as: QLoRA, low-rank adaptation

41stories this week
47last 30 days
61all time

Timeline

  1. Oct 10, 2026 · Opinion / analysis · 1 source
    Help With Choosing Hardware [P]
    I am going to fine-tune a satire model, with the base model being Qwen3.5-14B-Base.
  2. Oct 8, 2026 · Research paper · 1 source
    Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
    Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it.
  3. Oct 8, 2026 · Research paper · 1 source
    Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
    To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration.
  4. Oct 8, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.905-beta: Sandboxing is here!
    We're introducing Windows, Mac and Linux sandboxing in Unsloth!
  5. Oct 8, 2026 · Research paper · 1 source
    DPPM: Dual-Path Parametric Memory for Personalized Language Models
    Long-term personalization requires language models to use interaction history to track users' preferences across sessions.
  6. Oct 8, 2026 · Research paper · 1 source
    Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
    We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work.
  7. Oct 8, 2026 · Research paper · 1 source
    YOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous Teleoperation
    We present YOCO, a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses.
  8. Oct 8, 2026 · Research paper · 1 source
    Harness Evolution Hits a Ceiling: When Weight Training Should Begin
    Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights.
  9. Oct 8, 2026 · Research paper · 1 source
    Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
    Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.
  10. Oct 8, 2026 · Research paper · 1 source
    Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
    We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
  11. Oct 8, 2026 · Research paper · 1 source
    Parametric Trajectory Distillation for Few-Step Video Generation
    We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.
  12. Oct 8, 2026 · Research paper · 1 source
    Beyond Resolution: Object-to-Image Ratio Mismatch in Instance Retrieval
    We show that the dominant cause is usually not resolution loss but object-to-image (O2I) ratio mismatch: the object occupies different fractions of the two images.
  13. Oct 8, 2026 · Research paper · 1 source
    Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?
    We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness.
  14. Oct 8, 2026 · Research paper · 1 source
    LadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMs
    Lifelong editing of LLMs requires storing thousands of edits after acquisition.
  15. Oct 8, 2026 · Opinion / analysis · 1 source
    The model that didn't exist, so you made it yourself
    Last week, I wanted a small version of the prompt rewriter that ships with Qwen-Image 2.1.
  16. Oct 7, 2026 · Research paper · 1 source
    Beyond Owls: Subliminal Learning Can Transfer Learned Capabilities and Backdoors
    In subliminal learning (SL), a teacher model passes on a trait to a student model by distillation on data semantically unrelated to the trait.
  17. Oct 7, 2026 · Research paper · 1 source
    AutoAdapt: Automatic Domain Discovery Enables Low-Cost Extensibility
    We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters.
  18. Oct 7, 2026 · Research paper · 1 source
    SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
    Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing.
  19. Oct 7, 2026 · Research paper · 1 source
    Leakage-Controlled Multimodal Learning for Diagnosis and Progression Prediction in Alzheimer's Disease Research
    Alzheimer's disease prediction involves irregular visits, heterogeneous measurements and incomplete modalities.
  20. Oct 7, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.904-beta: Train your own Decision model
    Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.
  21. Oct 7, 2026 · Research paper · 1 source
    Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
    We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.
  22. Oct 7, 2026 · Research paper · 1 source
    BagDINO: Multi-View Baggage Re-Identification with DINOv3
    This paper investigates baggage re-identification as an instance-level retrieval problem in a multi-camera setting, leveraging DINOv3 foundation-model representations to match a query image against a gallery of registered baggage images.
  23. Oct 7, 2026 · Research paper · 1 source
    Marrying Pricing and Advertising with LLMs
    We study a sequential pricing problem in which a seller jointly posts a price and an advertisement generated by a large language model (LLM).
  24. Oct 7, 2026 · Research paper · 1 source
    Juno: Taming Predictive Latents for Vision-Language-Action Models
    Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models.
  25. Oct 7, 2026 · Research paper · 1 source
    Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR
    We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3).
  26. Oct 7, 2026 · Research paper · 1 source
    A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks
    Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data.
  27. Oct 7, 2026 · Research paper · 1 source
    Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution
    We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters.
  28. Oct 7, 2026 · Research paper · 1 source
    Shaer: Controlled Arabic Poetry Generation with Meter Subform and Semantic Conditioning
    We present Shaer, a controllable Classical Arabic poetry generation framework jointly conditioned on natural-language descriptions, meter subforms, and target hemistich counts.
  29. Oct 7, 2026 · Research paper · 1 source
    When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
    We introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model.
  30. Oct 6, 2026 · Research paper · 1 source
    Are Parameter-Efficient Fine-tuning Methods Really Different?
    Parameter-efficient fine-tuning (PEFT) offers many parameterizations, yet their methodological and functional differences remain unclear.

Often appears with