REVIEW 13 cited by
EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs). We assess the ability of LLMs to understand complex emotions and social interactions by asking them to predict the intensity of emotional states of characters in a dialogue. The benchmark is able to discriminate effectively between a wide range of models. We find that EQ-Bench correlates strongly with comprehensive multi-domain benchmarks like MMLU (Hendrycks et al., 2020) (r=0.97), indicating that we may be capturing similar aspects of broad intelligence. Our benchmark produces highly repeatable results using a set of 60 English-language questions. We also provide open-source code for an automated benchmarking pipeline at https://github.com/EQ-bench/EQ-Bench and a leaderboard at https://eqbench.com
Forward citations
Cited by 13 Pith papers
-
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
CompanionBench, a real-data-grounded bilingual benchmark with a hidden disclosure gate and IRT-corrected judging, ranks 28 AI companions and finds most fail to earn deeper disclosure, often substituting warmth for substance.
-
StoryScope: Investigating idiosyncrasies in AI fiction
StoryScope extracts narrative features showing AI stories favor tidy plots and over-explain themes while human stories show more moral ambiguity and temporal complexity, enabling strong detection and attribution.
-
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
HSS-Synth generates 230k instruction-tuning samples for 14 humanities/social-science fields and reports state-of-the-art fine-tuning results on 16 benchmarks.
-
BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
BridgeAlign's rubric-guided 'bridge' degradation creates near-boundary hard preference pairs, letting Qwen3-8B beat 11 baselines on average across 17 human-preference and knowledge benchmarks.
-
PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models
Personality-aware prompting plus contrastive retrieval of aligned scenarios measurably lifts LLM accuracy on EmoBench emotional-understanding tasks, especially for GPT models with CoT.
-
Information-Theoretic Limits of Reliability and Scaling in Language Models
A theoretical framework derives a reliability ceiling and a max-form Chinchilla-type scaling law for LLMs from task entropy and dependency spectra.
-
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.
-
Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation
A skeleton-first reasoning generation method reduces answer anchoring in reverse chain-of-thought traces, while semantic suppression increases latent anchoring.
-
Escaping the Verifier: Learning to Reason via Demonstrations
RARO trains reasoning LLMs from expert demonstrations alone using an adversarial game between a policy and a relativistic critic, outperforming verifier-free baselines and approaching verifier-based RL.
-
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
A new attack framework, TrojanClimb, shows that adversaries can place models with embedded backdoors or biases on public leaderboards while retaining competitive rankings, across text embeddings, text generation, spee...
-
PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
A post-training framework with persona-specific LoRA experts and a situation-aware router improves LLM emotional responses, but the evidence on preserving general ability is undercut by missing base-model comparisons.
-
DynamicBench: Evaluating Real-Time Report Generation in Large Language Models
DynamicBench is a proposed benchmark for real-time report generation, and the authors claim their retrieval-augmented system outperforms GPT-4o, but the evidence is inconsistent and the comparison is not apples-to-apples.
-
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
Fine-tuning a 1B medical chatbot on LLM-rewritten emotional dialogues improves its emotion scores with only small changes in n-gram overlap with the original medical responses.
Discussion (0). Continue with ORCID to comment.