LoRA
Also known as: QLoRA, low-rank adaptation
41stories this week
47last 30 days
61all time
Timeline
- Oct 10, 2026 · Opinion / analysis · 1 sourceHelp With Choosing Hardware [P]I am going to fine-tune a satire model, with the base model being Qwen3.5-14B-Base.
- Oct 8, 2026 · Research paper · 1 sourceCited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought FaithfulnessLarge language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it.
- Oct 8, 2026 · Research paper · 1 sourceSyn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal EmbeddingsTo address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration.
- Oct 8, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.905-beta: Sandboxing is here!We're introducing Windows, Mac and Linux sandboxing in Unsloth!
- Oct 8, 2026 · Research paper · 1 sourceDPPM: Dual-Path Parametric Memory for Personalized Language ModelsLong-term personalization requires language models to use interaction history to track users' preferences across sessions.
- Oct 8, 2026 · Research paper · 1 sourceInternalizer: Portable Context-to-Parameter Mapping for Very Large Language ModelsWe present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work.
- Oct 8, 2026 · Research paper · 1 sourceYOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous TeleoperationWe present YOCO, a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses.
- Oct 8, 2026 · Research paper · 1 sourceHarness Evolution Hits a Ceiling: When Weight Training Should BeginImproving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights.
- Oct 8, 2026 · Research paper · 1 sourceWhere to Adapt Matters: Layer-Selective Fine-Tuning for Capability RetentionParameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.
- Oct 8, 2026 · Research paper · 1 sourceResidual Advantage: Student-Relative Teacher Guidance for RL with Verifiable RewardsWe propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
- Oct 8, 2026 · Research paper · 1 sourceParametric Trajectory Distillation for Few-Step Video GenerationWe introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.
- Oct 8, 2026 · Research paper · 1 sourceBeyond Resolution: Object-to-Image Ratio Mismatch in Instance RetrievalWe show that the dominant cause is usually not resolution loss but object-to-image (O2I) ratio mismatch: the object occupies different fractions of the two images.
- Oct 8, 2026 · Research paper · 1 sourceHarness Compilation: Which Decisions Should a Small Vision-Language Model Keep?We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness.
- Oct 8, 2026 · Research paper · 1 sourceLadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMsLifelong editing of LLMs requires storing thousands of edits after acquisition.
- Oct 8, 2026 · Opinion / analysis · 1 sourceThe model that didn't exist, so you made it yourselfLast week, I wanted a small version of the prompt rewriter that ships with Qwen-Image 2.1.
- Oct 7, 2026 · Research paper · 1 sourceBeyond Owls: Subliminal Learning Can Transfer Learned Capabilities and BackdoorsIn subliminal learning (SL), a teacher model passes on a trait to a student model by distillation on data semantically unrelated to the trait.
- Oct 7, 2026 · Research paper · 1 sourceAutoAdapt: Automatic Domain Discovery Enables Low-Cost ExtensibilityWe present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters.
- Oct 7, 2026 · Research paper · 1 sourceSLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic RoutingMotivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing.
- Oct 7, 2026 · Research paper · 1 sourceLeakage-Controlled Multimodal Learning for Diagnosis and Progression Prediction in Alzheimer's Disease ResearchAlzheimer's disease prediction involves irregular visits, heterogeneous measurements and incomplete modalities.
- Oct 7, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.904-beta: Train your own Decision modelTurn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.
- Oct 7, 2026 · Research paper · 1 sourceBeyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question AnsweringWe present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.
- Oct 7, 2026 · Research paper · 1 sourceBagDINO: Multi-View Baggage Re-Identification with DINOv3This paper investigates baggage re-identification as an instance-level retrieval problem in a multi-camera setting, leveraging DINOv3 foundation-model representations to match a query image against a gallery of registered baggage images.
- Oct 7, 2026 · Research paper · 1 sourceMarrying Pricing and Advertising with LLMsWe study a sequential pricing problem in which a seller jointly posts a price and an advertisement generated by a large language model (LLM).
- Oct 7, 2026 · Research paper · 1 sourceJuno: Taming Predictive Latents for Vision-Language-Action ModelsJoint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models.
- Oct 7, 2026 · Research paper · 1 sourceItgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASRWe describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3).
- Oct 7, 2026 · Research paper · 1 sourceA Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never SpeaksContinual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data.
- Oct 7, 2026 · Research paper · 1 sourceDecoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context PollutionWe study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters.
- Oct 7, 2026 · Research paper · 1 sourceShaer: Controlled Arabic Poetry Generation with Meter Subform and Semantic ConditioningWe present Shaer, a controllable Classical Arabic poetry generation framework jointly conditioned on natural-language descriptions, meter subforms, and target hemistich counts.
- Oct 7, 2026 · Research paper · 1 sourceWhen Routing Reveals Membership: Privacy Leakage from MoE Router TelemetryWe introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model.
- Oct 6, 2026 · Research paper · 1 sourceAre Parameter-Efficient Fine-tuning Methods Really Different?Parameter-efficient fine-tuning (PEFT) offers many parameterizations, yet their methodological and functional differences remain unclear.