Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis
We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist.
Key points
- AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data.
- Such automation requires repeatability: when the underlying evidence is unchanged, the agent's categories, priorities, and counts should not shift materially between runs, even if each individual answer appears plausible.
- We evaluate three recurring customer-feedback tasks across eight frontier models, corpus sizes from 100 to 5,000 records, multiple prompts, and three execution designs: raw generation, taxonomy-free hierarchical decomposition, and a taxonomy-grounded agent (TGA) using persistent themes, subthemes, and record-level predictions.
- Overall, these results show that taxonomy grounding produces more consistent and repeatable outputs for recurring knowledge work.
Sources (1)
- [1]Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback AnalysisarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 09:29 AM
We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist.
AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data.
Extractive summary: sentences quoted from the sources.