ResearchResearch paperEvaluation & Benchmarks · Large Language Models · Business, Funding & Industry1 source · Oct 6, 2026

Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces.

Key points

  • Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role.
  • A signed-number-at-pointer baseline explains most recovery over exact quotation checks.
  • Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence.
  • The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed.

Sources (1)

  • [1]Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:57 PM
    We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces.
    Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  2. Oct 2, 2026Chatham scales its capital markets expertise with OpenAI
  3. Oct 1, 2026Claude-shaped science
  4. Sep 29, 2026Introducing GPT-6.1 Sol
  5. Sep 28, 2026Holo4: powering generalist computer-use agents
  6. Sep 28, 2026Basis completes a tax workbook 2x faster with GPT-6 Astra

Related