AION
Technique

Function calling

Also known as: tool calling, tool use

36stories this week
42last 30 days
52all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    [P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]
    I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.
  2. Oct 10, 2026 · Open-source release · 1 source
    open-webui/open-webui v0.12.0
    Approvals and questions from tools still have to be answered in the chat, calls need the Allow Call permission and end after an hour, models can set their own Realtime Voice in the model editor, and the settings can also be given with the "AUDIOREALTIMEENABLED", "AUDIOREALTIMEOPENAIAPIBASEURL", "AUDIOREALTIMEOPENAIAPIKEY", "AUDIOREALTIMEMODEL", "AUDIOREALTIMEVOICE", "AUDIOREALTIMETRANSCRIPTIONMODEL" and "REALTIMECALLPROMPTTEMPLATE" environment variables.
  3. Oct 8, 2026 · Research paper · 1 source
    Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration
    To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video.
  4. Oct 8, 2026 · Research paper · 1 source
    DataSense-Bench: The First Step Toward an AI Scientist
    We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.
  5. Oct 8, 2026 · Research paper · 1 source
    Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System
    Large language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent.
  6. Oct 8, 2026 · Research paper · 1 source
    Error-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent Systems
    We propose Error-Propagation Modeling for Failure Attribution (EMFA).
  7. Oct 8, 2026 · Research paper · 1 source
    ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
    We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
  8. Oct 8, 2026 · Research paper · 1 source
    Evaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility Workflows
    Context: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited.
  9. Oct 8, 2026 · Research paper · 1 source
    DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
    Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation.
  10. Oct 8, 2026 · Research paper · 1 source
    VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding
    We introduce VAMR (Video Agent for Multi-Question Reasoning), which coordinates all questions about a video through one shared tool-use trajectory.
  11. Oct 8, 2026 · Research paper · 1 source
    When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
    We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces.
  12. Oct 8, 2026 · Research paper · 1 source
    AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
    To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
  13. Oct 7, 2026 · Research paper · 1 source
    RH-Detect: A Unified Benchmark for Reward Hacking Detection
    We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema.
  14. Oct 7, 2026 · Product / feature launch · 2 sources
    Claude Haiku 5.5
    As previously promised, here's Anthropic's new fast, low cost model: Introducing Claude Haiku 5.5.
  15. Oct 7, 2026 · Research paper · 1 source
    Plan-and-Patch: Diffusion Language Models for Agentic Planning
    We introduce Plan-and-Patch, a plan-and-act framework in which a diffusion language model (dLLM) generates a structured, program-like plan through parallel unmasking and repairs it by filling in selected regions while keeping the surrounding steps fixed.
  16. Oct 7, 2026 · Research paper · 1 source
    RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
    In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved.
  17. Oct 7, 2026 · Research paper · 1 source
    Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
    How can we predict which base checkpoint is worth an expensive round of agentic post-training?
  18. Oct 7, 2026 · Research paper · 1 source
    A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
    Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate.
  19. Oct 7, 2026 · Research paper · 1 source
    HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration
    Real-world low-quality images suffer from complex mixed degradations, including but not limited to noise, blur, atmospheric effects, etc. Recent agentic methods usually model real-world image restoration (Real-IR) as a sequential tool calling problem over task-specific single-degradation restoration models.
  20. Oct 7, 2026 · Research paper · 1 source
    Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents
    Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents.
  21. Oct 7, 2026 · Research paper · 1 source
    LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
    We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents.
  22. Oct 7, 2026 · Research paper · 1 source
    How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression
    To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., write) with an analysis-verb (e.g., discuss) reliably flips the decision, suggesting it is mediated by a compact internal state.
  23. Oct 7, 2026 · Research paper · 1 source
    From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
    Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert.
  24. Oct 6, 2026 · Research paper · 1 source
    ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
    We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent.
  25. Oct 6, 2026 · Research paper · 2 sources
    Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
    Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.
  26. Oct 6, 2026 · Research paper · 1 source
    Breaking the Space Barrier and its Application to Language Model Inference
    Language models are more and more often asked for structured output: JSON that follows a schema, or a tool call with typed arguments.
  27. Oct 6, 2026 · Research paper · 2 sources
    CADFather Reconstructs Parametric CAD Programs Using Coordinated Tools
    CADFather is an autonomous agentic system that coordinates complementary tools and a vision-language assistant to reconstruct parametric CAD models from 3D meshes without additional training.
  28. Oct 6, 2026 · Research paper · 1 source
    Agent-Controlled Forgetting for Tool-Using Agents: Reversible Context Curation in Practice
    We study agent-controlled forgetting: the acting model selects previously observed tool results, replaces each with a short note at its original position, and retains the exact original in a recoverable archive.
  29. Oct 6, 2026 · Research paper · 1 source
    EMHO: EMbodied Agent Harness Optimization via Experience Traces
    We propose EMbodied Agent Harness Optimization (EMHO), a self-evolving framework that keeps the embodied model frozen and iteratively revises its harness by analyzing execution trajectories and prior harness history.
  30. Oct 6, 2026 · Research paper · 1 source
    Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek
    We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes.

Often appears with