Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge
We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces.
ProofPaper ↗
Key points
- Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role.
- A signed-number-at-pointer baseline explains most recovery over exact quotation checks.
- Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence.
- The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed.
Sources (1)
- [1]Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence JudgearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:57 PM
We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces.
Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Oct 2, 2026Chatham scales its capital markets expertise with OpenAI
- Oct 1, 2026Claude-shaped science
- Sep 29, 2026Introducing GPT-6.1 Sol
- Sep 28, 2026Holo4: powering generalist computer-use agents
- Sep 28, 2026Basis completes a tax workbook 2x faster with GPT-6 Astra