AION
Research paperRetrieval, RAG & Search · Large Language Models1 source · Oct 6, 2026

Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek

We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes.

Key points

  • Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG).
  • With Claude Haiku 4.5 as router and answer model, a reconstruction of the platform's tool agent answered 71.6% of canonical Greek questions correctly and vector RAG 95.3% (difference $-23.6$ percentage points, 95% CI $-33.1$ to $-15.1$).
  • The tool agent's losses arose in retrieval.
  • Tool interfaces for community knowledge bases need search that tolerates how users type.

Sources (1)

  • [1]Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:58 AM
    We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes.
    Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG).

Extractive summary: sentences quoted from the sources.