Pith. sign in

Paper Citation Record · LEDGER

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues

As of 15 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2412.17292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17292 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:41:33.390250Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10a95dde-38d3-4839-87f4-e1ca93295a73 · outbound

This paper cites Empathy Through Multimodality in Conversational Interfaces.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Empathy Through Multimodality in Conversational Interfaces

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:32.836243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:32.836243Z digest=sha256:325f4074024a1b3138bbde3543a7ec3ee29c0118549f95e36ec81205cb476272

Observation 5caa6b3b-b711-4899-95df-ed6e1491726d · outbound

This paper cites Rahmani, and Ramesh Jain.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Rahmani, and Ramesh Jain

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.464912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:32.842235Z digest=sha256:9876ad6bb066d2a30c3e0a929aa42f88097f3e9557ddd1c36bbe769016b69df3

Observation 85c12ff7-8a16-40da-8665-2e0a3715138a · outbound

This paper cites GPT-4 Technical Report.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:32.847576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:32.847576Z digest=sha256:4ed740c307186e13676c766c9b9402007ae580792e25ba11ebd4d34adb12b70c

Observation 97c2493c-aef2-490d-89d0-d52181a4266e · outbound

This paper cites Facechat: An emotion-aware face-to-face dialogue framework.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Facechat: An emotion-aware face-to-face dialogue framework

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.447857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.069472Z digest=sha256:eb7bddfd795967430d97eb74e08be991fe641118f9be0aabc3f692bbae410b14

Observation bbe1a122-f4fc-4dbc-b32f-0f1a8bd7e0ac · outbound

This paper cites PaLM 2 Technical Report.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues PaLM 2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.074736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.074736Z digest=sha256:2a69dc87689028e0767d5f137d1e507fbf275d5e00774e8d69db13c700909fdc

Observation 91b55cb7-7ec2-419a-9915-154f9b0191c5 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Common Voice: A Massively-Multilingual Speech Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.080449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.080449Z digest=sha256:2eaa163fc97f1eb03a65392a0aa8fb4950fe3d9b3dde62049691a39c7ffd28f3

Observation 1165a748-2539-49b5-94f8-188866968abd · outbound

This paper cites Non-verbal communication in human social interaction.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Non-verbal communication in human social interaction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.430997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.086252Z digest=sha256:014465d6a6e801edf105e8356e8ec3ab56c2657d715c34f4771895e11d178757

Observation bd8c445e-b357-4cb2-986f-38a17e003b6b · outbound

This paper cites Qwen Technical Report.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.091119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.091119Z digest=sha256:47c11095304e830d54898025c8291d1192cf484f580c01f9aab0de8d326d3f1c

Observation 5e29f143-3003-4d08-976d-b6ee62fba3f2 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.414928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.097389Z digest=sha256:325aa70dcffb7c7191d54dfb0fa4cf87dd1f2a7ae58a735ef66554f25a70316b

Observation e1fcdb86-810a-40de-8830-5dbe02ddfaca · outbound

This paper cites A neural probabilistic language model.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A neural probabilistic language model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.397506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.101838Z digest=sha256:892d5af514cf49966e4b3f0df6bfb689acfa863bebb409f3d833752b278159db

Observation 9a2599c3-1b19-4067-bd61-18105ac3e172 · outbound

This paper cites Lan- guage models are few-shot learners.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Lan- guage models are few-shot learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.107326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.107326Z digest=sha256:f9a0a427cafad3dcaad986fdaf1ad7bcea29c09302cdb884ca8cbdd3974e4288

Observation aad6f30a-782c-4647-a182-b26077c14612 · outbound

This paper cites How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks).

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.112033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.112033Z digest=sha256:f721533d517aa21fe5a0c4dacc04f2cc1a1df49e9a8f3eea4e2165fb99316d56

Observation fa8681dc-d0c3-4886-8af0-bf68a0f6dc01 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Crema-d: Crowd-sourced emotional multimodal actors dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.357690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.117351Z digest=sha256:067f84eefae0708cc0c57b51e29ee68a48346342dd621ad5385c612a4f1a62fc

Observation 7cf50eda-3079-4e43-8608-02e13d757dec · outbound

This paper cites End-to- end object detection with transformers.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues End-to- end object detection with transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.122541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.122541Z digest=sha256:b1759c7b0ef9c27c52ebe4d6329c24f131d201a59c5cc99775dba1c7c8dd6d46

Observation 24beb63f-628e-4a50-ad59-36a4a5a335d0 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.127633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.127633Z digest=sha256:0bf90946b8acb8809ec13393e69d7aa0dbcfbeecc4334925beec555bd6abcf80

Observation a5f0462c-2ba7-41fb-b957-f79247347c4f · outbound

This paper cites Palm: Scaling language modeling with pathways.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Palm: Scaling language modeling with pathways

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.133408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.133408Z digest=sha256:3a4214ee5235749d3521b1009348c9c63148dfe823a8c38ed3b28867f32b8662

Observation bb900e1d-4d39-45d1-b8c2-4cc399c98be3 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.138391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.138391Z digest=sha256:2e161d92ceb88a5f4f0c572e4a3a2c43af112ff64849e999419c7687dffda85b

Observation 053d322a-ba49-424b-86ff-8dbfe45ddf01 · outbound

This paper cites Towards multimodal emotional support con- versation systems, 2024.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Towards multimodal emotional support con- versation systems, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.320642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.143449Z digest=sha256:d51dc68108a2cd15fedf58751d763b0963752929ba7eb0a1773e7c3c65101218

Observation a1aef6e3-e7d0-4f97-8552-90ed59f3653d · outbound

This paper cites GoEmotions: A Dataset of Fine-Grained Emotions.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GoEmotions: A Dataset of Fine-Grained Emotions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.148396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.148396Z digest=sha256:de4c6bc5b87782124b83d96704f9626fa3abc8eb8e8e5ef643b5835b26b34e1d

Observation 1b9dc277-c2cf-49ef-8775-d2a1dbd601ac · outbound

This paper cites Retinaface: Single-shot multi- level face localisation in the wild.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Retinaface: Single-shot multi- level face localisation in the wild

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.304532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.153793Z digest=sha256:c2418ac172f2266fda326870db96b4f8694c4a41c1185d52d42632582848fc68

Observation a801b367-aceb-47b3-9db8-eb8a763541c1 · outbound

This paper cites EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.159002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.159002Z digest=sha256:53343f52fc98b3a8ac72408af5e7addecb20544b136b857f47b6157bb9aece79

Observation df1fcccf-4332-4586-9d3c-7e4eceaa3ab8 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Imagebind: One embedding space to bind them all,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.287083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.164662Z digest=sha256:3a4bb6a3f4970748e90b56b66ae514ed32353694db6364b85daa87d8b65a9cac

Observation c9633150-aece-4a77-9ed0-9059406c1eb7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.171100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.171100Z digest=sha256:3ed558b48184aa715755fb9d75b681b65b68f86e407381f533bf22ef1d838b1f

Observation f4f11a56-8622-43b2-9f1f-fe98f6afb096 · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.176818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.176818Z digest=sha256:df0d975dddb20cb529dca517656e455f685e4eabc30a63ffe02332d0b143a024

Observation 32dce23d-ac17-4de5-8b3b-99ab082b7ad5 · outbound

This paper cites Perceiver: General perception with iterative attention.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Perceiver: General perception with iterative attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.182934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.182934Z digest=sha256:366dda9816aedd8a2b97a0f0898eedf7557783bfd58629f218cc4d792257eaab

Observation e4b3b818-95d7-4f05-ab8b-ab48d29d055f · outbound

This paper cites CoLLaVO: Crayon Large Language and Vision mOdel.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues CoLLaVO: Crayon Large Language and Vision mOdel

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.187905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.187905Z digest=sha256:3d3d2a5c5a7aaee77b76e94c50779efa81e75342309004b22343dcde79287208

Observation 49604cd8-84cc-4946-8f70-1c133ee6d42a · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A Diversity-Promoting Objective Function for Neural Conversation Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.193401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.193401Z digest=sha256:0e30b94819532ee5c04555fc8cd606283bc48d0c241871138924f7841d717172

Observation f49ec87b-b44e-4ce6-98d7-e511d6024f4e · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.198727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.198727Z digest=sha256:8ac1a303cdd88eacd51333c09c21c96eae10b8f7c5f10145bbbbc97b5e552e44

Observation ef1bc8fc-ff9a-49cc-9b51-e64d8ca4cce7 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Rouge: A package for automatic evaluation of summaries

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.248264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.203694Z digest=sha256:4362ff89263dbe0b9aa160d3d13fa59ee2110387e0369c86c4c66652bfca3da4

Observation 8fef244b-6d9e-4e28-ab4a-ecd1e1e3a98a · outbound

This paper cites Paralinguistics-enhanced large language modeling of spoken dialogue.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Paralinguistics-enhanced large language modeling of spoken dialogue

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.231184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.209055Z digest=sha256:27608da4dba698847894393c46664cecd4159687679c11ac6358de26a6bafc7a

Observation 141e1eaa-952c-45d7-b268-6b88002b2ca9 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.213822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.213822Z digest=sha256:39ee2575533a418c3e150527e8191b2f37dd86a68e8e6d80a0268505f8e67116

Observation a010c5e9-99e9-4651-aef4-227a4274d111 · outbound

This paper cites Visual instruction tuning.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.218859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.218859Z digest=sha256:a9526ae9ea5c4f928e352785467293905c0e3d229d4026ce283e1519a0a98dfe

Observation f538aa0b-225b-40f7-ad69-b948551e7fcb · outbound

This paper cites The ryer- son audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues The ryer- son audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.192672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.224336Z digest=sha256:3ac2edb2fa2a306cb0ed736bb0ff24614da77b88539562f0ad11abaebe86ecc2

Observation 78d20b3b-01ce-4009-b82c-e1629321318e · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.229353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.229353Z digest=sha256:286d314658c4427d05a1ab5433605bc3741d58c02892ab22b173ca2652c21bda

Observation b5ac3c96-bf86-46e3-aa66-1b9d6e0a142d · outbound

This paper cites Gen- erative spoken dialogue language modeling.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Gen- erative spoken dialogue language modeling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.174437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.235122Z digest=sha256:f57b831b50b30c8a0d3ec31c5a1a1c76a8f8613207f51d5a2f4b490d0e5c6727

Observation d28c59e3-69d6-4a40-8262-5099d4ed297e · outbound

This paper cites Librispeech: an asr corpus based on public do- main audio books.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Librispeech: an asr corpus based on public do- main audio books

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.157323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.240627Z digest=sha256:43948100707974d034a5e42865c83d6895b24c60450a4762b89ee57675cd4b6d

Observation cd831bb4-08ac-497c-8bc1-5c98edf90e94 · outbound

This paper cites Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.245811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.245811Z digest=sha256:42b71c8eba60c89713ab3dbf581fb80a883f0073e17a5489d864c011954cc789

Observation 888f3fa5-ca73-42f2-8aeb-0ca558f1a949 · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A Call for Clarity in Reporting BLEU Scores

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.253158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.253158Z digest=sha256:e43fec6a89efa8592fe2a71389d1e3d5ccb0988998262263897371aba31c1207

Observation e0bbf93c-7a3d-4862-954e-6e8abc9e4d89 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Learning transferable visual models from natural language supervi- sion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.259314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.259314Z digest=sha256:284af8a305b816a1ae6277f98365c713d2255499b602572e452080ce3e12d265

Observation 743ee57f-053e-45ce-b1d1-695af3a57bc1 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Robust speech recognition via large-scale weak supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.131058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.264135Z digest=sha256:948c1d00e35eed31619ccbecd6a2e01588395e80ed58641ec5ad39198ebad00d

Observation 56dcf946-7027-4f02-acaf-3cdb3ed3000d · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.268906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.268906Z digest=sha256:e726c7d6d0a93de1fde08fccd80a7af06dc5161973a30c79c43461ef35eb824e

Observation 603a6181-8037-4858-8586-39910ee2cfbb · outbound

This paper cites Deepface: Closing the gap to human-level per- formance in face verification.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Deepface: Closing the gap to human-level per- formance in face verification

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.101671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.273920Z digest=sha256:61aec49589efcf2595203289f302b0737cbcb0dfec634f13a46972971cbdfc27

Observation 0f5ec78b-f67e-4117-a13d-22e74e8b8cf3 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.279337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.279337Z digest=sha256:6dfb756c642e2ac86290a43d0accf199516f5c158d504863a14c1794e1260804

Observation a93e3875-9208-4627-a1ef-b2b7cd2e37cc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Gemini: A Family of Highly Capable Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.284458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.284458Z digest=sha256:405a0c8f2a70b6cbad08e73ca9a4cb27c93c170ab02a59dc81dc7d6a3de9f878

Observation 4c6cbd7c-a95e-417c-be6e-c94589a3539c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.289815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.289815Z digest=sha256:5c401285daa13432f40753f17ce20cf14daea0bffa1ae8993b6c0143607b27ee

Observation a66f0a91-7e53-451e-9c10-df680246853f · outbound

This paper cites Nonverbal cues in human–robot interaction: A communication studies perspec- tive.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Nonverbal cues in human–robot interaction: A communication studies perspec- tive

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.083757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.294830Z digest=sha256:e02643e65ae6282cca0bc3c186ac842d586e63a2e16252c64699f442f9100de1

Observation 197bb5ce-6c80-4c14-a1ec-71a2e2db2c3a · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.300088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.300088Z digest=sha256:bd29b20839db0c82d4881a4dcc5c8545b163c05d2a428eb45643720c025b0942

Observation ef9e7d27-876e-44ac-8e72-8338b337f9a1 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues NExT-GPT: Any-to-Any Multimodal LLM

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.305307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.305307Z digest=sha256:593dd8d22279a819715c0da1a895b367979f5aa6d6a0b0e616da63bceccf58d7

Observation 232acf97-afc6-4654-9994-9bd2ff31758a · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.310250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.310250Z digest=sha256:9fd00e71cf888521ed38971b7018ce2f9f2b08f5b435925fbbd7b42669b87da7

Observation 992ada81-9b8d-49a7-87fa-bc54fdeb9c55 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.316145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.316145Z digest=sha256:d5d7180c211af34d408655af20329d8c74ae4a41ea8dcc8cf4c6dba55307d4f5

Observation 8a4ad359-db96-4a07-a166-3948da8617d7 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.321332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.321332Z digest=sha256:45664cd686f222d8d9f9c82bb9f9c71ea9dcfefaa64ba4d5672491a3a03f4a28

Observation 9af723fe-c657-40b5-999e-7fabb25417e7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues OPT: Open Pre-trained Transformer Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.326680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.326680Z digest=sha256:41c6273f68463a45f1aed61857eba2f28c7dc84c4df22b0d7370288db1058d60

Observation dd18ea5b-f986-47ae-a955-b61271b4d3bf · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues BERTScore: Evaluating Text Generation with BERT

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.332586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.332586Z digest=sha256:005e82c360d9d27852f8f1b0e849aa40a993cecc282565d6121e25d47ed9da21

Observation 6ac453cb-0805-4ce5-9ad7-f2d2d0388e5e · outbound

This paper cites DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.337719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.337719Z digest=sha256:0c6c6e24fea2f7499b64343fed277d72878bd141c4fa31c53a6b66e77c20fb5a

Observation b02e6bb9-6af8-4011-87e8-2eba579bab59 · outbound

This paper cites A higher BLEU score indicates a more natural and engaging dialogue model.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A higher BLEU score indicates a more natural and engaging dialogue model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.065645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.343532Z digest=sha256:8355945782b967dbec8008cf06c6903928791e12b623de4945700f5becbed43d

Observation 5ea4968a-f7f4-4369-8911-42d70cb02f94 · outbound

This paper cites Fluency evaluates the grammatical correctness, smoothness, and natural flow of the response.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Fluency evaluates the grammatical correctness, smoothness, and natural flow of the response

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.049017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.348630Z digest=sha256:0b3afff589c3bb4485f723fd8bcbc4b2b7b54297393f3d9de1eb408f5d9628ff

Observation fae329b9-1064-4ffa-a20b-f112b6440eed · outbound

This paper cites an unresolved cited work.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:41:34.030674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.354052Z digest=sha256:2ab10906432d49a69b866c5297202790caed66f2d9a44c245711feb805b1aa15

Observation dbc4d716-2063-4973-b899-f2669da73c0b · outbound

This paper cites It con- tains 2,615 hours of English speech from 92,325 voices with diverse genders, ages, and accents.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues It con- tains 2,615 hours of English speech from 92,325 voices with diverse genders, ages, and accents

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.013229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.359058Z digest=sha256:fd6645a53eab1cfc1092712a0d153136b3003e34cf205b510fcdc5138ceacb36

Observation aa9cfbb6-a444-4f41-9f1c-b09b7e64ca78 · outbound

This paper cites The prompt given to the GPT is: These are the frames in a video.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues The prompt given to the GPT is: These are the frames in a video

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.994769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.364317Z digest=sha256:8928ba5f3d927addb7fc00a0d57a846106dc2094b0b51d089446e52b938a7156

Observation e7b34924-928c-47b7-a190-3fd516d0764e · outbound

This paper cites Yeah, I think it’s unfair how the FD burns 6 tons of books.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Yeah, I think it’s unfair how the FD burns 6 tons of books

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.976852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.369504Z digest=sha256:92207ef35f4b28dc283fcb41b19f5e561de0353e6848b66f8fb764076d35df89

Observation cbdd5d48-a78e-4d9c-8bd0-c8f1cc0766b8 · outbound

This paper cites Understanding the context is crucial for a fair evaluation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Understanding the context is crucial for a fair evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.960759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.374722Z digest=sha256:9d686693ee5add90a1042c240890f6cd5fe7be08c5f20eaf5446233fc57c19ac

Observation f59e432a-3a1e-44cd-9012-1442a6576f8b · outbound

This paper cites Consider if the emotion expressed is suitable for the situation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Consider if the emotion expressed is suitable for the situation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.943115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.379738Z digest=sha256:eafdd6aa4b2db6898a8a0ee1f8e406b40350608a8f1973edad5ba5bd12de96f7

Observation c9647bcc-93cc-485d-a050-452bec147607 · outbound

This paper cites an unresolved cited work.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:41:33.926595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.384727Z digest=sha256:127f64441145aa24e4b2fff72d99fb950a64b0a1befc2941ef007eee5efeaeb4

Observation 8a3b6a30-15fb-463b-aba8-f6a6c29d223a · outbound

This paper cites I've been feeling really down lately. Nothing seems to cheer me up.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues I've been feeling really down lately. Nothing seems to cheer me up

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.910063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:41:33.390250Z digest=sha256:981183bca375575eb330be526d901365114a6c5daa591fa8634b03631e25b4ed

Pith citing papers

No inbound Pith citation observations are available.