Pith. sign in

Paper Citation Record · LEDGER

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

As of 9 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 3 inbound Pith citation observations for arXiv:2502.04976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04976 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:49:51.683856Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:28:56.798473Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:18:53.193305Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact5
  • verified fuzzy10
  • unresolved62
  • parse uncertain4
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb87eba8-4fc5-4cc3-a2a9-697bfd994a21 · outbound

This paper cites Qwen Technical Report.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.402163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.402163Z digest=sha256:c37b5699de59788bdcca8bd3af8c16e1532bdeb156d9c38525aff3fd660c58e9

Observation 31147a08-07ff-4534-9c32-ccbbf8074a64 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.406713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.406713Z digest=sha256:4b03f1f61e07bb50650dc5944c824040a9164fcd8bd721d5a5cfeb8e640d509f

Observation 056093af-70b6-445d-bb1a-871fb52b9cc5 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Gonzalez, Ion Stoica, and Eric P

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.410253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.410253Z digest=sha256:aca0d19ffe58a566e806e15bd80740f2f63962f11270c7966f71d9e6b6c0da25

Observation 66ebfe9e-e93b-4e7d-9fb7-86552c74a084 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.413339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.413339Z digest=sha256:3bb40cbb10b7e41beb72d6466187cada783edbeadc33fefdd6f2236ddde4a128

Observation 29821088-ca91-40cf-a029-9405535b50de · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.416963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.416963Z digest=sha256:9048d6bd8265154e04a734354cb3892ce5402ae7776c1586495d5ba8726e011a

Observation 5ed0215c-ad68-4e10-b656-50786cd0b997 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.420407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.420407Z digest=sha256:0af1d7d1d1d0dcc24a86d28f78933c9bdb5eb7e87376875b5e3507ed20c36d65

Observation d775e18f-52dd-4ee1-ae6b-23aa0f2522ac · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.420055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.424211Z digest=sha256:eba3eb58216ebd80dd20b44384e8a7dd68e05a60a2334bce4529a3a012fcf3f6

Observation 0a18e02f-bb40-4054-a0cf-d9235503413d · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.410994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.427160Z digest=sha256:929c4c254cc31615a13cfe4c4778801637ce496d0dbcd739178f20f34ac612f5

Observation bed9aafa-425c-4deb-a3d3-5c9862d2d37c · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.391618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.434111Z digest=sha256:eb86fdd84eeac962be729eb801eb3ba976369daeff95e21a7772438675c94060

Observation f6d23236-04d6-4b56-a5d5-57ccb2b86412 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.382153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.438529Z digest=sha256:33f1bee1e150082af3e414834dc2af73c0d47db8cae91379964f3ff74d04d34d

Observation da6d4218-8a5a-4b4e-bf2f-c80d7351c8f3 · outbound

This paper cites EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.934715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.444014Z digest=sha256:f01efd2db6145c37714be8b28fbd854992118f3ced25b392d51ef7c19cb08d75

Observation 503c6581-b745-444c-ac62-337a5dc32756 · outbound

This paper cites In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): Tutorial Summaries.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): Tutorial Summaries

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.372704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.441406Z digest=sha256:0bd03571b3d4270d6fa74ed6e0adc3bc1b94a895bff52e9e8897e3a48b9202f5

Observation 26abf529-89f8-465c-82c9-a342d2febada · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.353222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.450379Z digest=sha256:cb897167f51898dd88bf3c2d1b6d206aaf69b3b324cbd3582dd7b83838e4d9d9

Observation 026f6c2b-820e-432e-97a8-e7b6408c99f8 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.362524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.447368Z digest=sha256:77e3d166c54a11385eb5cdd8231fde93e98f6d4882044940e27ff1442610e2c6

Observation e3d29065-5f7a-4c8f-8de0-e02348cc64a3 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.456352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.456352Z digest=sha256:25a166176610b3adfd5062b29b78de3cd2cadf98ebf122ffad41b54ee19ed150

Observation 4189bc9b-c508-4943-b6ba-5366ab3247f7 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.453494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.453494Z digest=sha256:9b538a101bc8e9614067c6f026b4460618d5b8523381956fde92fd4acdf82010

Observation 18f12589-cb3a-4c75-ad01-c5cd9541715d · outbound

This paper cites DiaASQ : A Benchmark of Conversational Aspect-based Sentiment Quadruple Analysis.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark DiaASQ : A Benchmark of Conversational Aspect-based Sentiment Quadruple Analysis

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.902059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.463088Z digest=sha256:ee6169f4aaf2a203fbf9c405dfd2d973bbd7398caa424b3b3c60d88b62f240a7

Observation 639262ca-b2a4-499c-90d3-901024897bd1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Auto-Encoding Variational Bayes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.459647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.459647Z digest=sha256:8fa0a60b237dd311a26a5a68ca3ffcc079aba229a4b30da3e5764295ba6df801

Observation 00378d85-45e0-4e90-a7a5-2c533c9f3cee · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.338392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.469739Z digest=sha256:dfe7213c0692c43eba2d33311d1139e4545f6cd770c4c2e9ebf2fb0ca7dd5f0c

Observation 6357f5ef-71d5-49d2-b059-d866fe875d4c · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark A Diversity-Promoting Objective Function for Neural Conversation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.466277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.466277Z digest=sha256:bddda9d504763dbb12d7948bce72ecb181df27dc6f5bc48a45b1744abc23eaff

Observation 0f81845f-abf1-49a0-bc00-421084672c58 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.328872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.476098Z digest=sha256:9179a1fd994b4f326e4f8d93c5960052200ef856d28362a4d3c3e21cce74dff9

Observation dc6c9300-5af4-4b18-aca5-a9ab9bb1ea3d · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark A Survey on Benchmarks of Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.472790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.472790Z digest=sha256:f0a739d4e72f85b492e3b6d1f1f5cffcd68f5d95ab41c602df9b3fde7bcd87af

Observation c16bd71b-5db5-4cce-acc5-dd42d2412db8 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.482452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.482452Z digest=sha256:e1e4bc149e4f35e82c57fbaeab0a2c400ac25e52cbae1576816c98b6c2913ccf

Observation 7cc22715-9a48-4d91-9b05-62f5586f76c5 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.319703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.479145Z digest=sha256:74a00806071045d392c030ebbf1495dec75462e4e0ee746d8b1aed4b1b248fa3

Observation d7a011df-0e4a-4441-ad5c-8c795905297a · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.489205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.489205Z digest=sha256:aa48e18a8a83611d8ecaf610cdeb0fa29cb80650d15532a595792be76c25b880

Observation 774632b9-a894-4948-8628-4ed7e5677e85 · outbound

This paper cites MoEL: Mixture of Empathetic Listeners.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark MoEL: Mixture of Empathetic Listeners

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.861607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.485665Z digest=sha256:12fb716d85a52ec5207fdd0bfdfbceb03555cb9556d7991af2183328d69c7609

Observation 8c2ad0f5-e375-4a3a-b729-bea345d8a0a8 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.304549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.495423Z digest=sha256:db8e4ccc6b6dabcb870d037b334818cc6446e2fe8c99bfbe64fffea1501d63f8

Observation 1e125c31-4ad3-4b76-9d4d-69cb9c4b4d23 · outbound

This paper cites The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.492309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.492309Z digest=sha256:250378872c76c6ab285c3859353817fce1a2fd107856736b4e0c8a90e2a9d811

Observation 519ba272-27db-40c5-ae17-409449357493 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.295005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.502866Z digest=sha256:9e7a7f1805ee93ce713b64aab8f068867bb59e5d6f74d27282dfe1759affbf2b

Observation 3377d4c9-0abc-42e4-8b07-7bbe50a1a9ac · outbound

This paper cites PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.499157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.499157Z digest=sha256:f9a1360e80dd3d0b49f2967e51887f74a1c7fbeceb47d0070a749c3a5cc5af1a

Observation adda083f-3989-4b76-9f15-1470a0218d7d · outbound

This paper cites MIME: MIMicking Emotions for Empathetic Response Generation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark MIME: MIMicking Emotions for Empathetic Response Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.510158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.510158Z digest=sha256:e0812f8ceff27c33516378cbc531efbe9b7cca324fc7be5d497b1b89baf7c1a8

Observation ad011661-eb30-4f38-b868-0fa2e5831a66 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.506385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.506385Z digest=sha256:e83b6afc894fa1df9da9c81e7e9bdda91e1641902177cd7c5518da9b37c4d661

Observation 0b57493c-9938-4813-b92a-dd57ea777fc7 · outbound

This paper cites Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.518375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.518375Z digest=sha256:42bcdb5283516b3716c2b4c854ac93a261973f1f82b4e5b304a109e3ed7fbd86

Observation 87ec28d0-353a-4027-adbc-14929db4ddfe · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.285217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.514162Z digest=sha256:5ae70660d794f7c05996d7a8360910fbb4d23418706443c62b3a54ea239665e6

Observation ebd0be06-2ead-48bf-b3f9-bc832c413447 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.525875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.525875Z digest=sha256:6ddf9209f789067e92aa6baaa0b77d3b52266b002995628165aeac818619e2c8

Observation f3f87b12-e718-4dba-b923-859d5d12b45d · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.276019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.522263Z digest=sha256:10581e51ca557d0592c449e86ecefcfce65095dbd531a155378f75447b08c112

Observation 01c515dd-d82c-4aaf-a231-3b81af018c48 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.261092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.533397Z digest=sha256:8ea30fd2da682b725b0c35e24eab5bc97c757f557e9f4f8f4102fe07eca9798b

Observation 7e4d92b7-ad2c-4f5c-b3b6-f4b5fa189708 · outbound

This paper cites Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.529599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.529599Z digest=sha256:37f74d2044db8fef9667962cf90b6bf67ee3bbd30f3dece70be44bdc8c7096b7

Observation a940ed49-f0b6-4a23-9e7b-7e18e28f4774 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.540459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.540459Z digest=sha256:18e4f85eb32a5ff99ece83dffba16c08122aee66ad110578b163922c7d157633

Observation 80047f13-f03c-4e15-a98f-c317e2a2cb7f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.251869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.536999Z digest=sha256:eccee70300434b8911ad952cf789f76ceaa1ddc44a4297fe97a9c47d4ab569ff

Observation 1f9e4a3e-c6ca-467a-ae58-0b6a4160d5bb · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.551860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.551860Z digest=sha256:b839fdf567bbecc5e8bc840ff4a77909b91dda473f74828605a8412a8bf0fb70

Observation 840e6f52-38dc-4251-88b6-e2d43a08c4e3 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.555321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.555321Z digest=sha256:ca5d9074c4f5e14aa6c70d07d652bd66f87c3fa91309f7115fc82a14ac987930

Observation cc5ae234-023d-442f-9b5c-527e3301931e · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.236565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.548221Z digest=sha256:3518b8ba1b17e5353e417e853b65f56ef80a6a45b10b43886922f5d485fd8999

Observation 1343a509-e0af-4f92-aea5-ceffad9ab69f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.205795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.562421Z digest=sha256:3ae17413f1070950c5ca57eacce97d6a60831291248dd116f641cb2585a6c050

Observation e03bcb67-6ab2-4b67-97b7-44db29d4d48d · outbound

This paper cites Towards Semantic Equivalence of Tokenization in Multimodal LLM.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Towards Semantic Equivalence of Tokenization in Multimodal LLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.565771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.565771Z digest=sha256:8f819858de31cbcf4add3560f1ca4a6b61728730b66580d818a5c05f7b654cb9

Observation 1ba78d90-f4a9-4817-b575-1e48d260b4e2 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.215408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.558870Z digest=sha256:7f4dd6e62311bd5eabdd2daec2b03a4ea3b0a3dbbaf38e0b0fee076bacf44447

Observation c1ec96a0-5b61-47cc-b32d-ca99d38961b3 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.572967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.572967Z digest=sha256:1203576db61e93393f351cc11e5c1071a0e4dbc8ae72bef7a9a8bda6e6b207d5

Observation 32f64eee-8cdd-4936-a868-d7c5cc0d2cbb · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.180118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.580202Z digest=sha256:92e4ddfd41bbeb26a04166f8c25687c5bdc156e875f3aa07004c191e40d5cf0a

Observation 6fa01968-e790-43c7-840f-f341027f7931 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.195317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.569609Z digest=sha256:c9ebd7acaedd8e205eadbe75831432e0bcdaea1ae4a294c6a646378a2fdce1b3

Observation 8ea6d1d5-c5a6-4848-bca1-e89bf56e8cd2 · outbound

This paper cites Enhancing Empathetic Response Generation by Augmenting LLMs with Small-scale Empathetic Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Enhancing Empathetic Response Generation by Augmenting LLMs with Small-scale Empathetic Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.586828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.586828Z digest=sha256:717d216ae0dfe778bd790f515682b1f61e25bd9836c21789e3e77b615e7186cf

Observation 139bc4c6-ce29-4d47-ac3c-e0c0ccb5b7eb · outbound

This paper cites Faithful Logical Reasoning via Symbolic Chain-of-Thought.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Faithful Logical Reasoning via Symbolic Chain-of-Thought

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.576651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.576651Z digest=sha256:bc0b653a999728e4f5685df017d0849d95daf9c82272a786e09a8b27c38ca20b

Observation 7e637f55-148f-4a4e-bfad-048a9e2e7e83 · outbound

This paper cites A Survey of Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark A Survey of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.593437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.593437Z digest=sha256:23efab3dcd20a971fdc81afc239707d45fdf334d14f12c220864ff4e85fd0631

Observation 7bece7b9-8aef-43ac-b64e-ace0af46964f · outbound

This paper cites Exploiting Emotion-Semantic Correlations for Empathetic Response Generation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Exploiting Emotion-Semantic Correlations for Empathetic Response Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.769322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.583466Z digest=sha256:a0d68075222b87298eeba0c0b83b8253d40ce302dc3a0e808e98d6e97138aa0a

Observation 2533db02-107c-4585-a428-c44e8307f18e · outbound

This paper cites CASE: Aligning Coarse-to-Fine Cognition and Affection for Empathetic Response Generation.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark CASE: Aligning Coarse-to-Fine Cognition and Affection for Empathetic Response Generation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-08T20:49:51.728632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.600355Z digest=sha256:aa6812ebb44b42aa52e2c97db3cd0e1bd819ef862de2c0e1bd554f02f4bf59ef

Observation 972c834a-7b94-4739-87de-b0325cc573ab · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.171186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.590277Z digest=sha256:f6bfbff3ab44df8306947b3a65301d57b646c9985efb7d433df97cff3cfef668

Observation 96151797-fc49-416e-9120-772d8c712263 · outbound

This paper cites ECQED: Emotion-Cause Quadruple Extraction in Dialogs.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark ECQED: Emotion-Cause Quadruple Extraction in Dialogs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.596829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.596829Z digest=sha256:e70d161c4f1968f4c8ab3f88ceab011f6ed4b0cce69b52b1539a853b7799e656

Observation 41e073f6-2b6d-4735-9a4a-4e2ab616a733 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.604550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.604550Z digest=sha256:1b9f3039d52faf02aba76803c8913bc8135f1f1630c072d2396e2cceab3dcea9

Observation c7931415-92b0-4347-b766-87de8e798a2f · outbound

This paper cites dia_id":.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark dia_id":

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.162028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.608745Z digest=sha256:1661bd984b5156266a4f49e0a647d06e5f120c6dcf20e81d41f1c98ba0350297

Observation b1153dac-d97b-4c20-b916-4eba867c6306 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.152215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.612368Z digest=sha256:6d94ee17f7f4dd8d6fa10ffe2302acaa20ce6fda1cff8f033ab06e5e7815ec97

Observation 38d1a22a-a33d-44db-ba4c-2661f8c4c309 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.143087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.616140Z digest=sha256:ab819900e1984822cbf8dbf898712a9051866aa0394a84f0695d09fc79228202

Observation b5e822e1-d806-478b-8b4b-35fbae1cd6b4 · outbound

This paper cites It’s like they have no empathy or think about what if it was them.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark It’s like they have no empathy or think about what if it was them

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.133737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.619386Z digest=sha256:1caee903d912bab6e2ada74e1007bffe475dcab337a01a2ee5dda0e362f0f2e3

Observation a4fdbe06-2545-46be-ba4a-060b3e34e4e4 · outbound

This paper cites Achievements and Self-Realization.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Achievements and Self-Realization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.124516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.622729Z digest=sha256:b4d7694913397391b88d323376a838060d8ae51d16f29b7da3d3e51eb2340450

Observation 40a52432-062e-4a44-a55c-1e18d064b5e9 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.115094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.627132Z digest=sha256:4592087513fdecaa047021b84461a2ab60140714d574303df5a762d36737d93f

Observation 17733aba-a035-4fb3-bf77-9ad0bfc0205f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.105871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.630499Z digest=sha256:26f182b5e4834f3acb21d2bac1d2ad367764c80fb58f7b900c698f34d2410fd9

Observation 2be31d4b-7a47-4c47-a6bc-762fee5f8602 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.096776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.634207Z digest=sha256:599f24de76c72c2b08376ddcf5c95528117577e3db14fccd7988071b2744b59d

Observation 074d344d-e9e9-4799-8589-d15636c8c968 · outbound

This paper cites Avg. Score.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Avg. Score

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.087871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.637772Z digest=sha256:64ba5afc3af5c7b8efb5982e97536d1402bc243d44d719e922224416191a845b

Observation 649fcae0-1cb1-4ad0-ae2c-4dfa1520ef2e · outbound

This paper cites As shown, the model’s performance peaks when the number of tokens reaches 16.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark As shown, the model’s performance peaks when the number of tokens reaches 16

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.078641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.641508Z digest=sha256:165a161ae5667bb521ae91f59bb0a8e12b7d852b02dd58915959dd3305e9d0a0

Observation b7132b85-21e9-4e4d-bc31-2fe924d2828f · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.069295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.645210Z digest=sha256:621ad4b06ac3a64b2133d72526aa3ad91f67af2972d6131a1e321274f8c981d0

Observation 1daad0ad-245f-4243-a96b-e0ffdb572373 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 71

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:52.060370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.648662Z digest=sha256:ffaf486963f2a2fd678b4e15b50fc5c1367ad6b08fab4a61675415e5dc5cfb6f

Observation 5b7a6db1-39ae-4259-b7e8-59f202305a64 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.050814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.652248Z digest=sha256:cb0014ac6b172f64e361840126b1b6f29f44bd84283f48ebca0fc69b046f3744

Observation b7d9000f-65a5-4168-9fe4-edcb5fa5a0c0 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.041555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.655747Z digest=sha256:0bd06f7b7215c30934a4d4cec1900507ad1fb0534d44ab95e8361ac2005bfd9f

Observation f8f960a5-79f9-4072-b7d1-17889a6b1ee1 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 74

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:52.031828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.659041Z digest=sha256:7be56fa902f34863255a802dca970638372941be457efc4660d56560d3da6c38

Observation 76328e9c-ba14-44f1-b375-d964621910fd · outbound

This paper cites 4.Goal to Response: Validate the speaker’s sense of relief and preparedness, acknowledging the stressful situation they avoided.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark 4.Goal to Response: Validate the speaker’s sense of relief and preparedness, acknowledging the stressful situation they avoided

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.022313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.662921Z digest=sha256:2c1464f5855f9d1dd40134a7b0cb2d50813dae17fd03dd2baa8d6341106e2d5f

Observation 2dc9975c-9905-41f8-84f1-950c27fe3330 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:52.011982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.666565Z digest=sha256:5f723ba3480af13fa20f01f1091d876280a438bcdd7824f5490fb6ccff031b3a

Observation 6b757b37-4780-49a9-8028-102a74959d60 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 77

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:52.001767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.670365Z digest=sha256:c364727383d0f02daa97bfaaf22684bd68efe294c12e23972fed89036a8c7191

Observation 2540466f-e338-4f96-8906-664f4489e4e6 · outbound

This paper cites 4.Goal to Response: To provide empathy and acknowledge the speaker’s excitement and enjoyment of the trip.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark 4.Goal to Response: To provide empathy and acknowledge the speaker’s excitement and enjoyment of the trip

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:51.992188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.673892Z digest=sha256:58a5a512c10c37d6eb58ea9baca80a55b21531b805757bb22f290cf9922bb7eb

Observation a130b068-e74f-44ae-9ffe-0ac0e290fbd6 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-08T20:49:51.982285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.677229Z digest=sha256:86818fde022a25271c8cd0e06c0d345176c1c395b715fffe59b455f00b797a12

Observation e92b5957-dfe8-4510-b368-284088896f31 · outbound

This paper cites an unresolved cited work.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Unresolved cited work

Reference 80

Resolution
parse uncertain
raw_fallback, observed 2026-08-08T20:49:51.972365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.680586Z digest=sha256:ca9a05cf25bd3a52593f2f76e4150cd9355e0e4e22b8cea7324f7e2144c69daf

Observation 83ef674b-50de-4c29-aee8-73302ef10a5e · outbound

This paper cites 4.Goal to Response: Offer empathy and validation for the speaker’s feelings, acknowledging the challenge of parenting and the importance of communication.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark 4.Goal to Response: Offer empathy and validation for the speaker’s feelings, acknowledging the challenge of parenting and the importance of communication

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:51.962545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.683856Z digest=sha256:4a4e7220bc199f29d2f0b7042707f645fb529523066b8059ce97931d98053478

Observation b6542455-5c43-442a-88bb-038ad27c90d1 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark PandaGPT: One Model To Instruction-Follow Them All

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T20:49:51.543944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:49:51.543944Z digest=sha256:2b72e6ea5d31888e7a76efe4b556b27956c7cb9631adb10cbd728def711524d0

Observation 84b2134c-7597-4a83-a21e-9fefb77419fd · outbound

This paper cites Proceedings of the Advances in neural information processing systems.

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark Proceedings of the Advances in neural information processing systems

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:49:52.401199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T20:49:51.430540Z digest=sha256:579490011723e291e32d87448b2df7601e6b805bf975e338e4a5e409e4d60369

Pith citing papers

Observation a5c5a808-e9d9-47ff-b2fc-35e12ed5f1a6 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.196060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:6847daf66b67a461fa5ff194f12b3821586e2a1a4ddaeed7993abff899e5be17

Observation d9dd7114-7e1d-4658-bdd7-c897424e51b0 · inbound

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation cites this paper.

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:04.092789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:31:05.546782Z digest=sha256:ea5f1661c51a9b598dba0223ab0a63602f71ceaa5af06f4eca1e5e9f8708eb6e

Observation e42e6a58-8452-43d4-b7fd-a389285c8dce · inbound

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot cites this paper.

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:56.798473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:56.798473Z digest=sha256:10401db2a7bb487e5fb7e86b197e18a38ac8feabd4b2e940412ae36ecaf15383