Pith. sign in

Paper Citation Record · LEDGER

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

As of 10 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 3 inbound Pith citation observations for arXiv:2502.04976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04976 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:49:51.683856Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:28:56.798473Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:18:53.193305Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact5
  • verified fuzzy10
  • unresolved62
  • parse uncertain4
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb87eba8-4fc5-4cc3-a2a9-697bfd994a21 · outbound

This paper cites Qwen Technical Report.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.402163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.402163Z digest=sha256:dad5a93759fb871f29daf31ca708cea9031a813bf781791f595b2337531245a9

Observation 31147a08-07ff-4534-9c32-ccbbf8074a64 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.406713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.406713Z digest=sha256:4b03f1f61e07bb50650dc5944c824040a9164fcd8bd721d5a5cfeb8e640d509f

Observation 056093af-70b6-445d-bb1a-871fb52b9cc5 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Gonzalez, Ion Stoica, and Eric P

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.410253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.410253Z digest=sha256:aca0d19ffe58a566e806e15bd80740f2f63962f11270c7966f71d9e6b6c0da25

Observation 66ebfe9e-e93b-4e7d-9fb7-86552c74a084 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.413339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.413339Z digest=sha256:3bb40cbb10b7e41beb72d6466187cada783edbeadc33fefdd6f2236ddde4a128

Observation 29821088-ca91-40cf-a029-9405535b50de · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.416963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.416963Z digest=sha256:9048d6bd8265154e04a734354cb3892ce5402ae7776c1586495d5ba8726e011a

Observation 5ed0215c-ad68-4e10-b656-50786cd0b997 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.420407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.420407Z digest=sha256:0af1d7d1d1d0dcc24a86d28f78933c9bdb5eb7e87376875b5e3507ed20c36d65

Observation d775e18f-52dd-4ee1-ae6b-23aa0f2522ac · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.420055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.424211Z digest=sha256:73735c119df5c264f9723373a8a3037ad66c546499711249943d81d704542fa7

Observation 0a18e02f-bb40-4054-a0cf-d9235503413d · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.410994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.427160Z digest=sha256:aa4a06b5409f68e5c9b4a789094f2318fca8745730a7962a6ad730383b243238

Observation bed9aafa-425c-4deb-a3d3-5c9862d2d37c · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.391618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.434111Z digest=sha256:f88a712edb3b3833d0f1a93fc82e7cb7bd625f5ffeaf7f3046ec56573510edf5

Observation f6d23236-04d6-4b56-a5d5-57ccb2b86412 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.382153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.438529Z digest=sha256:e48b9d93e71a6cf7bd74c9d345eee4f7eaaa0e9d3034f4aa27699119b6a80f5c

Observation da6d4218-8a5a-4b4e-bf2f-c80d7351c8f3 · outbound

This paper cites EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.934715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.444014Z digest=sha256:187c0e5d72b21e4cbb63ee6865ba27a5d148fe2d9a09940719f01a81a24cbac1

Observation 503c6581-b745-444c-ac62-337a5dc32756 · outbound

This paper cites In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): Tutorial Summaries.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): Tutorial Summaries

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.372704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.441406Z digest=sha256:688a2af3322c94fc2ca912fd5c97ffdab8bd5735a4061b7e35321ed977964c1c

Observation 26abf529-89f8-465c-82c9-a342d2febada · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.353222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.450379Z digest=sha256:8ca09c0eb796430073dc0cbadd7c5884db31bbe43b3223d1feb5964538343005

Observation 026f6c2b-820e-432e-97a8-e7b6408c99f8 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.362524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.447368Z digest=sha256:8a47b9f81ab6ff6e6e763842b2d0acc75858209f1cd8e66588abc7e0b44b2203

Observation e3d29065-5f7a-4c8f-8de0-e02348cc64a3 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.456352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.456352Z digest=sha256:25a166176610b3adfd5062b29b78de3cd2cadf98ebf122ffad41b54ee19ed150

Observation 4189bc9b-c508-4943-b6ba-5366ab3247f7 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.453494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.453494Z digest=sha256:9b538a101bc8e9614067c6f026b4460618d5b8523381956fde92fd4acdf82010

Observation 18f12589-cb3a-4c75-ad01-c5cd9541715d · outbound

This paper cites DiaASQ : A Benchmark of Conversational Aspect-based Sentiment Quadruple Analysis.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark DiaASQ : A Benchmark of Conversational Aspect-based Sentiment Quadruple Analysis

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.902059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.463088Z digest=sha256:5e998e2fad52ff8b666b7c17a1995e4c036b9cad00c513b3222195ad2a91a1e7

Observation 639262ca-b2a4-499c-90d3-901024897bd1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Auto-Encoding Variational Bayes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.459647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.459647Z digest=sha256:8fa0a60b237dd311a26a5a68ca3ffcc079aba229a4b30da3e5764295ba6df801

Observation 00378d85-45e0-4e90-a7a5-2c533c9f3cee · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.338392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.469739Z digest=sha256:7984b1a74dcdcabe36214d274ee0374a3163a6fa3b622c59a13980845275290a

Observation 6357f5ef-71d5-49d2-b059-d866fe875d4c · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark A Diversity-Promoting Objective Function for Neural Conversation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.466277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.466277Z digest=sha256:bddda9d504763dbb12d7948bce72ecb181df27dc6f5bc48a45b1744abc23eaff

Observation 0f81845f-abf1-49a0-bc00-421084672c58 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.328872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.476098Z digest=sha256:22ebfbd73e2d9f42cd9d52ca391435dbedf163fc343711f397d9d1f7b514faab

Observation dc6c9300-5af4-4b18-aca5-a9ab9bb1ea3d · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark A Survey on Benchmarks of Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.472790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.472790Z digest=sha256:f0a739d4e72f85b492e3b6d1f1f5cffcd68f5d95ab41c602df9b3fde7bcd87af

Observation c16bd71b-5db5-4cce-acc5-dd42d2412db8 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.482452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.482452Z digest=sha256:e1e4bc149e4f35e82c57fbaeab0a2c400ac25e52cbae1576816c98b6c2913ccf

Observation 7cc22715-9a48-4d91-9b05-62f5586f76c5 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.319703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.479145Z digest=sha256:c126f8c819187f0e6d18c633ce029be41d12cff87d09e552a7b1435c1448f8c4

Observation d7a011df-0e4a-4441-ad5c-8c795905297a · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.489205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.489205Z digest=sha256:aa48e18a8a83611d8ecaf610cdeb0fa29cb80650d15532a595792be76c25b880

Observation 774632b9-a894-4948-8628-4ed7e5677e85 · outbound

This paper cites MoEL: Mixture of Empathetic Listeners.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark MoEL: Mixture of Empathetic Listeners

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.861607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.485665Z digest=sha256:283bdf30e5866638fbb2e98a3cc971c82388fc7de4cd517ef3da0b9ced1dcb1f

Observation 8c2ad0f5-e375-4a3a-b729-bea345d8a0a8 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.304549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.495423Z digest=sha256:7cd3e160357bd0961f04a9a60ab47479d7f31ab094d239d4a487912721add84a

Observation 1e125c31-4ad3-4b76-9d4d-69cb9c4b4d23 · outbound

This paper cites The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.492309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.492309Z digest=sha256:250378872c76c6ab285c3859353817fce1a2fd107856736b4e0c8a90e2a9d811

Observation 519ba272-27db-40c5-ae17-409449357493 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.295005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.502866Z digest=sha256:7b1b33ac40cd030a19fd27b8748a83ed500770f7fe0a9789cf3a9bb312809a8e

Observation 3377d4c9-0abc-42e4-8b07-7bbe50a1a9ac · outbound

This paper cites PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.499157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.499157Z digest=sha256:f9a1360e80dd3d0b49f2967e51887f74a1c7fbeceb47d0070a749c3a5cc5af1a

Observation adda083f-3989-4b76-9f15-1470a0218d7d · outbound

This paper cites MIME: MIMicking Emotions for Empathetic Response Generation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark MIME: MIMicking Emotions for Empathetic Response Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.510158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.510158Z digest=sha256:e0812f8ceff27c33516378cbc531efbe9b7cca324fc7be5d497b1b89baf7c1a8

Observation ad011661-eb30-4f38-b868-0fa2e5831a66 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.506385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.506385Z digest=sha256:e83b6afc894fa1df9da9c81e7e9bdda91e1641902177cd7c5518da9b37c4d661

Observation 0b57493c-9938-4813-b92a-dd57ea777fc7 · outbound

This paper cites Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.518375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.518375Z digest=sha256:42bcdb5283516b3716c2b4c854ac93a261973f1f82b4e5b304a109e3ed7fbd86

Observation 87ec28d0-353a-4027-adbc-14929db4ddfe · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.285217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.514162Z digest=sha256:f94459c5ad1f836190e062c476e36342416e270220ad375aea4a6ae3ccf2e325

Observation ebd0be06-2ead-48bf-b3f9-bc832c413447 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.525875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.525875Z digest=sha256:6ddf9209f789067e92aa6baaa0b77d3b52266b002995628165aeac818619e2c8

Observation f3f87b12-e718-4dba-b923-859d5d12b45d · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.276019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.522263Z digest=sha256:c60b7067e813d0f01dceefe8a9302e5ef4ee463bbd0631f3de987cccdd216350

Observation 01c515dd-d82c-4aaf-a231-3b81af018c48 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.261092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.533397Z digest=sha256:8a78386821cb42c3ed09a22e86793a618d048ccd203f20329959cbcb7efc76e4

Observation 7e4d92b7-ad2c-4f5c-b3b6-f4b5fa189708 · outbound

This paper cites Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.529599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.529599Z digest=sha256:37f74d2044db8fef9667962cf90b6bf67ee3bbd30f3dece70be44bdc8c7096b7

Observation a940ed49-f0b6-4a23-9e7b-7e18e28f4774 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.540459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.540459Z digest=sha256:18e4f85eb32a5ff99ece83dffba16c08122aee66ad110578b163922c7d157633

Observation 80047f13-f03c-4e15-a98f-c317e2a2cb7f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.251869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.536999Z digest=sha256:5a2028aa60286df850c856a352e3cc0896e23cf3b2e976d3c7e792af7893586e

Observation 1f9e4a3e-c6ca-467a-ae58-0b6a4160d5bb · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.551860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.551860Z digest=sha256:b839fdf567bbecc5e8bc840ff4a77909b91dda473f74828605a8412a8bf0fb70

Observation 840e6f52-38dc-4251-88b6-e2d43a08c4e3 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.555321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.555321Z digest=sha256:ca5d9074c4f5e14aa6c70d07d652bd66f87c3fa91309f7115fc82a14ac987930

Observation cc5ae234-023d-442f-9b5c-527e3301931e · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.236565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.548221Z digest=sha256:d636f9b6933309f981e2df10eb110050e68439f47206456c18478300e6c2c124

Observation 1343a509-e0af-4f92-aea5-ceffad9ab69f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.205795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.562421Z digest=sha256:ae9ff48c8eeae377a98d4034db8945400776e0c945c6710abea99e50466b7302

Observation e03bcb67-6ab2-4b67-97b7-44db29d4d48d · outbound

This paper cites Towards Semantic Equivalence of Tokenization in Multimodal LLM.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Towards Semantic Equivalence of Tokenization in Multimodal LLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.565771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.565771Z digest=sha256:8f819858de31cbcf4add3560f1ca4a6b61728730b66580d818a5c05f7b654cb9

Observation 1ba78d90-f4a9-4817-b575-1e48d260b4e2 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.215408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.558870Z digest=sha256:e949610ce10b4a87c72f55a01adf0d5dc3d7db489a7db7cf02200863553bdb59

Observation c1ec96a0-5b61-47cc-b32d-ca99d38961b3 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.572967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.572967Z digest=sha256:1203576db61e93393f351cc11e5c1071a0e4dbc8ae72bef7a9a8bda6e6b207d5

Observation 32f64eee-8cdd-4936-a868-d7c5cc0d2cbb · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.180118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.580202Z digest=sha256:0b07393dbd64ad800a29b1905f7fd880d6e039f5f517dc913dfe86662f9b3080

Observation 6fa01968-e790-43c7-840f-f341027f7931 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.195317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.569609Z digest=sha256:8074e7f8a9088420193bf84a02d1022090e18cbe0d325124d26ec5867ed8cf0b

Observation 8ea6d1d5-c5a6-4848-bca1-e89bf56e8cd2 · outbound

This paper cites Enhancing Empathetic Response Generation by Augmenting LLMs with Small-scale Empathetic Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Enhancing Empathetic Response Generation by Augmenting LLMs with Small-scale Empathetic Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.586828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.586828Z digest=sha256:717d216ae0dfe778bd790f515682b1f61e25bd9836c21789e3e77b615e7186cf

Observation 139bc4c6-ce29-4d47-ac3c-e0c0ccb5b7eb · outbound

This paper cites Faithful Logical Reasoning via Symbolic Chain-of-Thought.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Faithful Logical Reasoning via Symbolic Chain-of-Thought

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.576651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.576651Z digest=sha256:bc0b653a999728e4f5685df017d0849d95daf9c82272a786e09a8b27c38ca20b

Observation 7e637f55-148f-4a4e-bfad-048a9e2e7e83 · outbound

This paper cites A Survey of Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark A Survey of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.593437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.593437Z digest=sha256:23efab3dcd20a971fdc81afc239707d45fdf334d14f12c220864ff4e85fd0631

Observation 7bece7b9-8aef-43ac-b64e-ace0af46964f · outbound

This paper cites Exploiting Emotion-Semantic Correlations for Empathetic Response Generation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Exploiting Emotion-Semantic Correlations for Empathetic Response Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.769322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.583466Z digest=sha256:7728a569ad5cde3d2a6ac6d2a3025c23ea21024a02a65643d5b7dfbb8a1b5ebc

Observation 2533db02-107c-4585-a428-c44e8307f18e · outbound

This paper cites CASE: Aligning Coarse-to-Fine Cognition and Affection for Empathetic Response Generation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark CASE: Aligning Coarse-to-Fine Cognition and Affection for Empathetic Response Generation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.728632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.600355Z digest=sha256:d2204edd42d82be7e1da30473fe0c0129ec813b5cb802b7e55970d3c0ff7a377

Observation 972c834a-7b94-4739-87de-b0325cc573ab · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.171186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.590277Z digest=sha256:466c6d2caa4ac8454b407328cca1be3a7a9d35c45d90ea495e25afd40688fe52

Observation 96151797-fc49-416e-9120-772d8c712263 · outbound

This paper cites ECQED: Emotion-Cause Quadruple Extraction in Dialogs.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark ECQED: Emotion-Cause Quadruple Extraction in Dialogs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.596829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.596829Z digest=sha256:e70d161c4f1968f4c8ab3f88ceab011f6ed4b0cce69b52b1539a853b7799e656

Observation 41e073f6-2b6d-4735-9a4a-4e2ab616a733 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.604550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.604550Z digest=sha256:1b9f3039d52faf02aba76803c8913bc8135f1f1630c072d2396e2cceab3dcea9

Observation c7931415-92b0-4347-b766-87de8e798a2f · outbound

This paper cites dia_id":.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark dia_id":

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.162028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.608745Z digest=sha256:74e726d8062e101c1164d2152d49c83e2e2aed074f01446830f93ab9dcc1a64e

Observation b1153dac-d97b-4c20-b916-4eba867c6306 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.152215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.612368Z digest=sha256:d551abc6c36feb07ca6edc7604527ee4fe0f00d1183fc0500a1ef37709c3de44

Observation 38d1a22a-a33d-44db-ba4c-2661f8c4c309 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.143087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.616140Z digest=sha256:0eb021d033a16422b7153ba281279c7bc3dc364b3541971e7d14ea5d5711e317

Observation b5e822e1-d806-478b-8b4b-35fbae1cd6b4 · outbound

This paper cites It’s like they have no empathy or think about what if it was them.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark It’s like they have no empathy or think about what if it was them

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.133737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.619386Z digest=sha256:4d76c7f82263a45f73fcfa1128bfe32a5799536d61b27f49c6d1dd871f579622

Observation a4fdbe06-2545-46be-ba4a-060b3e34e4e4 · outbound

This paper cites Achievements and Self-Realization.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Achievements and Self-Realization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.124516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.622729Z digest=sha256:b949246106b168f70ac4cea95a738f543b70ee795d56ce22d21c362b2a893551

Observation 40a52432-062e-4a44-a55c-1e18d064b5e9 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.115094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.627132Z digest=sha256:5831f8173d05168d2da2d52cdb7a7461b0b7eb32195116f53e8e36b97384e604

Observation 17733aba-a035-4fb3-bf77-9ad0bfc0205f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.105871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.630499Z digest=sha256:80a918f34a6d1e3cf459368e0f1487a0f3a07e47fb957d0c99d62e2f6094878c

Observation 2be31d4b-7a47-4c47-a6bc-762fee5f8602 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.096776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.634207Z digest=sha256:5765680841ca4d9c271ea36e0e98ecf37c28083f116c701147bd1ca87c76e475

Observation 074d344d-e9e9-4799-8589-d15636c8c968 · outbound

This paper cites Avg. Score.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Avg. Score

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.087871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.637772Z digest=sha256:2306ed8b0321a13eba665f8d8efeb94927fbbe8c7a2f0033549e6deeb7932c00

Observation 649fcae0-1cb1-4ad0-ae2c-4dfa1520ef2e · outbound

This paper cites As shown, the model’s performance peaks when the number of tokens reaches 16.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark As shown, the model’s performance peaks when the number of tokens reaches 16

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.078641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.641508Z digest=sha256:68d01dec9c8534430b9d5b3ac2c20700edf649815d5031303070d90c6ef60a5c

Observation b7132b85-21e9-4e4d-bc31-2fe924d2828f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.069295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.645210Z digest=sha256:d09c7e5b0bb1644780f3d12327ec673850f43cf284ad54e977535081b403fe7f

Observation 1daad0ad-245f-4243-a96b-e0ffdb572373 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 71

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:52.060370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.648662Z digest=sha256:20edd1ff430f68e5703055d85fbfa295243c1211bf506c9c91173b4aad5964bc

Observation 5b7a6db1-39ae-4259-b7e8-59f202305a64 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.050814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.652248Z digest=sha256:e8310b16863ceea16a45642adbd2a93c1032a4195ab8f9a9fab00a0bf1a1ef7a

Observation b7d9000f-65a5-4168-9fe4-edcb5fa5a0c0 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.041555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.655747Z digest=sha256:485cee3a5cdff1eb430c43acae5db1779cba8bdd11b5b0c1f14dadfba9e8761a

Observation f8f960a5-79f9-4072-b7d1-17889a6b1ee1 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 74

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:52.031828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.659041Z digest=sha256:1ccff7bd31ae95e9386129dde2b84b00dce234fb6c26f9651de3f7101adf6503

Observation 76328e9c-ba14-44f1-b375-d964621910fd · outbound

This paper cites 4.Goal to Response: Validate the speaker’s sense of relief and preparedness, acknowledging the stressful situation they avoided.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark 4.Goal to Response: Validate the speaker’s sense of relief and preparedness, acknowledging the stressful situation they avoided

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.022313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.662921Z digest=sha256:be4e66972e4025429f6579c9a3d088650b523a47a839581e5606bc03688e105b

Observation 2dc9975c-9905-41f8-84f1-950c27fe3330 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.011982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.666565Z digest=sha256:399f30bedd3fc474c8020f33b0f233f67ab9ff04d6b27b6ce7aae20e4052cc8b

Observation 6b757b37-4780-49a9-8028-102a74959d60 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 77

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:52.001767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.670365Z digest=sha256:261a246a7936949bce4a329d3f28f547ea4cde0cc5bbf1587973b8563b9939c7

Observation 2540466f-e338-4f96-8906-664f4489e4e6 · outbound

This paper cites 4.Goal to Response: To provide empathy and acknowledge the speaker’s excitement and enjoyment of the trip.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark 4.Goal to Response: To provide empathy and acknowledge the speaker’s excitement and enjoyment of the trip

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:51.992188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.673892Z digest=sha256:4081bd101ecf27ccd6ff64d1e671f6aff71bb059c7181cf174f8090c798e89e3

Observation a130b068-e74f-44ae-9ffe-0ac0e290fbd6 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:51.982285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.677229Z digest=sha256:9b2a9b33776f7c3f3125f9277bbf5ee3d4e6225aac3ea64dd8a5e1419ebb0486

Observation e92b5957-dfe8-4510-b368-284088896f31 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 80

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:51.972365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.680586Z digest=sha256:96121deb00faceed43554b4ed9c1b806c9e69dd34f69e87faa7a99ede110a03a

Observation 83ef674b-50de-4c29-aee8-73302ef10a5e · outbound

This paper cites 4.Goal to Response: Offer empathy and validation for the speaker’s feelings, acknowledging the challenge of parenting and the importance of communication.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark 4.Goal to Response: Offer empathy and validation for the speaker’s feelings, acknowledging the challenge of parenting and the importance of communication

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:51.962545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.683856Z digest=sha256:ac18e33708a563bd8b01357291ea5abea341e49a537b4ecda20a572cba6bc1d8

Observation b6542455-5c43-442a-88bb-038ad27c90d1 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark PandaGPT: One Model To Instruction-Follow Them All

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.543944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.543944Z digest=sha256:2b72e6ea5d31888e7a76efe4b556b27956c7cb9631adb10cbd728def711524d0

Observation 84b2134c-7597-4a83-a21e-9fefb77419fd · outbound

This paper cites Proceedings of the Advances in neural information processing systems.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Proceedings of the Advances in neural information processing systems

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.401199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:49:51.430540Z digest=sha256:225d86dbd0a0b854ccd85491749259eba121a1c30fcd945e622cd882e36512f4

Pith citing papers

Observation a5c5a808-e9d9-47ff-b2fc-35e12ed5f1a6 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.196060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:bc1e57ab12e843b8eb5848f927f30019160d9b8d6f01789a927b0158d8b85520

Observation d9dd7114-7e1d-4658-bdd7-c897424e51b0 · inbound

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation cites this paper.

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:04.092789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:31:05.546782Z digest=sha256:ca347410f49f5b5acd4ea3d22b64ca02ac5d4a79caf4f95e2305b90e77b2bdf6

Observation e42e6a58-8452-43d4-b7fd-a389285c8dce · inbound

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot cites this paper.

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:56.798473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:56.798473Z digest=sha256:10401db2a7bb487e5fb7e86b197e18a38ac8feabd4b2e940412ae36ecaf15383