Pith. sign in

REVIEW 13 cited by

EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06281 v2 pith:KDK42KR5 submitted 2023-12-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords eq-benchbenchmarkemotionalintelligencemodelsaspectshttpslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs). We assess the ability of LLMs to understand complex emotions and social interactions by asking them to predict the intensity of emotional states of characters in a dialogue. The benchmark is able to discriminate effectively between a wide range of models. We find that EQ-Bench correlates strongly with comprehensive multi-domain benchmarks like MMLU (Hendrycks et al., 2020) (r=0.97), indicating that we may be capturing similar aspects of broad intelligence. Our benchmark produces highly repeatable results using a set of 60 English-language questions. We also provide open-source code for an automated benchmarking pipeline at https://github.com/EQ-bench/EQ-Bench and a leaderboard at https://eqbench.com

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

    cs.CL 2026-08 conditional novelty 8.0 of 10

    CompanionBench, a real-data-grounded bilingual benchmark with a hidden disclosure gate and IRT-corrected judging, ranks 28 AI companions and finds most fail to earn deeper disclosure, often substituting warmth for substance.

  2. StoryScope: Investigating idiosyncrasies in AI fiction

    cs.CL 2026-04 conditional novelty 7.0 of 10

    StoryScope extracts narrative features showing AI stories favor tidy plots and over-explain themes while human stories show more moral ambiguity and temporal complexity, enabling strong detection and attribution.

  3. HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    HSS-Synth generates 230k instruction-tuning samples for 14 humanities/social-science fields and reports state-of-the-art fine-tuning results on 16 benchmarks.

  4. BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

    cs.CL 2026-07 conditional novelty 6.0 of 10

    BridgeAlign's rubric-guided 'bridge' degradation creates near-boundary hard preference pairs, letting Qwen3-8B beat 11 baselines on average across 17 human-preference and knowledge benchmarks.

  5. PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Personality-aware prompting plus contrastive retrieval of aligned scenarios measurably lifts LLM accuracy on EmoBench emotional-understanding tasks, especially for GPT models with CoT.

  6. Information-Theoretic Limits of Reliability and Scaling in Language Models

    cs.CL 2026-05 conditional novelty 6.0 of 10

    A theoretical framework derives a reliability ceiling and a max-form Chinchilla-type scaling law for LLMs from task entropy and dependency spectra.

  7. MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.

  8. Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

    cs.CL 2026-02 conditional novelty 6.0 of 10

    A skeleton-first reasoning generation method reduces answer anchoring in reverse chain-of-thought traces, while semantic suppression increases latent anchoring.

  9. Escaping the Verifier: Learning to Reason via Demonstrations

    cs.LG 2025-11 conditional novelty 6.0 of 10

    RARO trains reasoning LLMs from expert demonstrations alone using an adversarial game between a policy and a relativistic critic, outperforming verifier-free baselines and approaching verifier-based RL.

  10. Exploiting Leaderboards for Large-Scale Distribution of Malicious Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A new attack framework, TrojanClimb, shows that adversaries can place models with embedded backdoors or biases on public leaderboards while retaining competitive rankings, across text embeddings, text generation, spee...

  11. PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions

    cs.CL 2025-09 reject novelty 5.0 of 10

    A post-training framework with persona-specific LoRA experts and a situation-aware router improves LLM emotional responses, but the evidence on preserving general ability is undercut by missing base-model comparisons.

  12. DynamicBench: Evaluating Real-Time Report Generation in Large Language Models

    cs.LG 2025-06 reject novelty 4.0 of 10

    DynamicBench is a proposed benchmark for real-time report generation, and the authors claim their retrieval-augmented system outperforms GPT-4o, but the evidence is inconsistent and the comparison is not apples-to-apples.

  13. Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuning a 1B medical chatbot on LLM-rewritten emotional dialogues improves its emotion scores with only small changes in n-gram overlap with the original medical responses.

Pith tools