Function calling
Also known as: tool calling, tool use
36stories this week
42last 30 days
52all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 source[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.
- Oct 10, 2026 · Open-source release · 1 sourceopen-webui/open-webui v0.12.0Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.
- Oct 8, 2026 · Research paper · 1 sourceDex-One2Many: Learning Dexterous Manipulation from a Single Human DemonstrationTo address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video.
- Oct 8, 2026 · Research paper · 1 sourceDataSense-Bench: The First Step Toward an AI ScientistWe introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.
- Oct 8, 2026 · Research paper · 1 sourceForms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent SystemLarge language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent.
- Oct 8, 2026 · Research paper · 1 sourceError-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent SystemsWe propose Error-Propagation Modeling for Failure Attribution (EMFA).
- Oct 8, 2026 · Research paper · 1 sourceReTeach: Building a Self-Teacher through Multi-Round Reflection and RetryWe introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
- Oct 8, 2026 · Research paper · 1 sourceEvaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility WorkflowsContext: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited.
- Oct 8, 2026 · Research paper · 1 sourceDuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent CollaborationVoice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation.
- Oct 8, 2026 · Research paper · 1 sourceVAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video UnderstandingWe introduce VAMR (Video Agent for Multi-Question Reasoning), which coordinates all questions about a video through one shared tool-use trajectory.
- Oct 8, 2026 · Research paper · 1 sourceWhen Interfaces Speak: Data-Aware Generative UI Harness for Active InteractionWe propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces.
- Oct 8, 2026 · Research paper · 1 sourceAgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use TasksTo this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
- Oct 7, 2026 · Research paper · 1 sourceRH-Detect: A Unified Benchmark for Reward Hacking DetectionWe present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
- Oct 7, 2026 · Product / feature launch · 2 sourcesClaude Haiku 5.5As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.
- Oct 7, 2026 · Research paper · 1 sourcePlan-and-Patch: Diffusion Language Models for Agentic PlanningWe introduce Plan-and-Patch, a plan-and-act framework in which a diffusion language model (dLLM) generates a structured, program-like plan through parallel unmasking and repairs it by filling in selected regions while keeping the surrounding steps fixed.
- Oct 7, 2026 · Research paper · 1 sourceRECAST: Learning to Compute the Right Context through Adaptive Evidence RoutingIn this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved.
- Oct 7, 2026 · Research paper · 1 sourceBefore They Can Solve: Predicting Post-Training Coding-Agent Performance from Base ModelsHow can we predict which base checkpoint is worth an expensive round of agentic post-training?
- Oct 7, 2026 · Research paper · 1 sourceA Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and TeachingReinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate.
- Oct 7, 2026 · Research paper · 1 sourceHarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image RestorationReal-world low-quality images suffer from complex mixed degradations, including but not limited to noise, blur, atmospheric effects, etc. Recent agentic methods usually model real-world image restoration (Real-IR) as a sequential tool calling problem over task-specific single-degradation restoration models.
- Oct 7, 2026 · Research paper · 1 sourceLoud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model AgentsPrior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents.
- Oct 7, 2026 · Research paper · 1 sourceLiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving MarketsWe introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents.
- Oct 7, 2026 · Research paper · 1 sourceHow Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by SuppressionTo obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., write) with an analysis-verb (e.g., discuss) reliably flips the decision, suggesting it is mediated by a compact internal state.
- Oct 7, 2026 · Research paper · 1 sourceFrom Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert TriageTool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert.
- Oct 6, 2026 · Research paper · 1 sourceToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and EvaluationWe present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent.
- Oct 6, 2026 · Research paper · 2 sourcesFrozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AILarge language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.
- Oct 6, 2026 · Research paper · 1 sourceBreaking the Space Barrier and its Application to Language Model InferenceLanguage models are more and more often asked for structured output: JSON that follows a schema, or a tool call with typed arguments.
- Oct 6, 2026 · Research paper · 2 sourcesCADFather Reconstructs Parametric CAD Programs Using Coordinated ToolsCADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.
- Oct 6, 2026 · Research paper · 1 sourceAgent-Controlled Forgetting for Tool-Using Agents: Reversible Context Curation in PracticeWe study agent-controlled forgetting: the acting model selects previously observed tool results, replaces each with a short note at its original position, and retains the exact original in a recoverable archive.
- Oct 6, 2026 · Research paper · 1 sourceEMHO: EMbodied Agent Harness Optimization via Experience TracesWe propose EMbodied Agent Harness Optimization (EMHO), a self-evolving framework that keeps the embodied model frozen and iteratively revises its harness by analyzing execution trajectories and prior harness history.
- Oct 6, 2026 · Research paper · 1 sourceTool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type GreekWe compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes.