Fine-tuning
Also known as: fine tuning, finetuning
152stories this week
158last 30 days
163all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceSpatialHarness: Test-Time Spatial Scaffolding for Fine Robotic ManipulationWe introduce SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic manipulation without policy fine-tuning or changes to the physical sensing setup.
- Oct 8, 2026 · Research paper · 1 sourceRounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State QuantizationQuantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates.
- Oct 8, 2026 · Research paper · 1 sourceVioLA: Learning Generalist Humanoid Control Policies from Human DataWe introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands.
- Oct 8, 2026 · Research paper · 1 sourcePredicting Alignment Generalization with Value RepresentationsIn this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values.
- Oct 8, 2026 · Research paper · 2 sourcesSpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language ModelsExisting spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input.
- Oct 8, 2026 · Research paper · 1 sourceLearning Kilometer-Scale Weather Prediction with Global-Regional AlignmentWe propose ScaleCast, a regional forecasting framework that addresses these challenges through Global-Regional Alignment.
- Oct 8, 2026 · Research paper · 1 sourceARC: A Reasoning Recipe for Robot Foundation ModelsWe show that there exists an effective and efficient complementary approach: the right reasoning recipe can substantially improve the zero-shot task performance of existing state-of-the-art RFMs. We refer to this recipe as ARC.
- Oct 8, 2026 · Research paper · 1 sourceOvercoming Prior Barriers: Supervised Fine-Tuning under Long-Tail DistributionSupervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support.
- Oct 8, 2026 · Research paper · 1 sourceAmbient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient LearningWe introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications.
- Oct 8, 2026 · Research paper · 1 sourcePrior or Feedback? What an LLM Uses When Adapting Neural OperatorsDo LLM scientific agents rely only on their initial task context, or do they adapt their decisions in response to experimental feedback?
- Oct 8, 2026 · Research paper · 1 sourceControllable Exaggeration for Generative Motion Models via Training-Time Adaptation and Inference-Time GuidanceRecent motion generative models have demonstrated strong capabilities in synthesizing physically plausible character motion, but often overlook established animation principles used by professional animators to ground and design their animation work.
- Oct 8, 2026 · Research paper · 1 sourceHarnessSQL: Harness-Native Training for SQL Agents in Realistic Database EnvironmentsTo bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning.
- Oct 8, 2026 · Research paper · 2 sourcesVibeEdit: Image Editing with Canvas InstructionsWe introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image.
- Oct 8, 2026 · Research paper · 1 sourceRecursive Self-Improvement through Multi-Agent Self-SupervisionTo address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories.
- Oct 8, 2026 · Research paper · 1 sourceIs Real-World Training Data Necessary for Generalist Graph Anomaly Detection?Generalist graph anomaly detection (GAD) aims to build a foundation model that detects anomalies on arbitrary unseen graphs without retraining or fine-tuning.
- Oct 8, 2026 · Research paper · 1 sourceAn Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under ResamplingA large body of research measures model coherence based on output variance without adequately considering competing causes.
- Oct 8, 2026 · Research paper · 2 sourcesSuperNav: An Agentic Navigation System for Any Task in Any SceneGeneral-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
- Oct 8, 2026 · Research paper · 1 sourceWhen Should Agents Think? Adaptive Reasoning via Cross-Turn EstimationBased on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning.
- Oct 8, 2026 · Research paper · 1 sourceDo Not Train Away Uncertainty: Early Uncertainty Anchored CalibrationEUA-Cal introduces early prediction regularization to preserve early predictive uncertainty and prototype structure regularization to exploit uncertainty reflected in the early feature space, jointly mitigating overconfidence.
- Oct 8, 2026 · Research paper · 1 sourceNatural Language to First-Order Logic LLM-based AutoformalizationThis paper addresses this gap: we first provide a principled definition for the FOL-autoformalization task by distinguishing Ontology Extraction from Logical Translation, showing how their conflation obscures (cross-study) evaluation; we review existing datasets, evaluation metrics, and LLM-based methods, including fine-tuning, prompting, and verification-based refinement; we identify open challenges in benchmarking, semantic evaluation, ontology-aware methods, and end-to-end applications.
- Oct 8, 2026 · Research paper · 1 sourcePulseBound: Future-Beat State Forecasting Under an Explicit Information BoundaryWe introduce PulseBound, a PPG representation learner combining physiologically structured future-beat prediction with an explicit stored-window information boundary.
- Oct 8, 2026 · Research paper · 1 sourceTest-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and LimitsWhich forms of test-time compute improve the predictions of strong pretrained tabular foundation models (TFMs)?
- Oct 8, 2026 · Research paper · 1 sourceProject Greenhouse: Progress Toward Fully Open and Sovereign Agentic SearchProject Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources.
- Oct 8, 2026 · Research paper · 1 sourceFrom Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON ExtractionWe present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections.
- Oct 8, 2026 · Research paper · 1 sourceSeek-and-View Reasoning for Multi-View Spatial UnderstandingTo realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model: a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning.
- Oct 8, 2026 · Research paper · 1 sourceEasy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching missesByte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
- Oct 8, 2026 · Research paper · 1 sourceMAP4CS: A Multi-dimensional Data Pruning Framework for Efficient Code Retriever Fine-tuningTo address these challenges, we propose MAP4CS (Multi-dimensional Awareness Pruning for Code Search), an adaptive data pruning framework.
- Oct 8, 2026 · Research paper · 1 sourceFrom Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object InteractionWe propose a hierarchical framework that converts a single-agent HOI policy into a reusable Object-oriented Motion Skill.
- Oct 8, 2026 · Research paper · 1 sourceAutoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal RetrievalWe introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval.
- Oct 8, 2026 · Research paper · 1 sourceYOCO: You Only Calibrate Once! Fast Mocap Calibration for Dexterous TeleoperationWe present YOCO, a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses.