ResearchResearch paperInterpretability · Multimodal Models · Large Language Models1 source · Oct 6, 2026

The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns.

Key points

  • EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions.
  • Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large).
  • We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category.
  • The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to $κw=0.254$-$0.423$.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding
  2. Oct 6, 2026Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning
  3. Oct 6, 2026SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
  4. Oct 4, 2026SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
  5. Oct 2, 2026Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience

Related