Pith. sign in

Paper Citation Record · LEDGER

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

As of 14 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 2 inbound Pith citation observations for arXiv:2508.02429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.02429 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:02:24.778729Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:19:34.140857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T13:55:53.243562Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97ea4dbf-5925-42e3-87e0-8cbc9aab1064 · outbound

This paper cites A review of affective computing: From unimodal analysis to multimodal fusion,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A review of affective computing: From unimodal analysis to multimodal fusion,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.340081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.340081Z digest=sha256:74c4ad3b4f7f7c196baddec66e70b4f0ce5b93d378feb8196f29456879f02058

Observation 11db27f3-e04a-4364-a9a2-a32fe4fb8e4a · outbound

This paper cites A systematic review on affective computing: Emotion models, databases, and recent advances,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A systematic review on affective computing: Emotion models, databases, and recent advances,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:31.520898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.345319Z digest=sha256:774655a816ac9e6422730ca10313bbae3ffb282b40b8c742d7c79551cfe9a59c

Observation de78fbb7-3a1d-4f76-84a8-a9aca737d6bd · outbound

This paper cites An effective data fusion methodology for multi-modal emotion recognition: A survey,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting An effective data fusion methodology for multi-modal emotion recognition: A survey,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:31.232058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.350453Z digest=sha256:9165ac9ce7fb1d419a1fd4651158466cec0a6a02702a6928bf68998d364f6a37

Observation b7525118-6794-41e5-a078-930cd45e178b · outbound

This paper cites Surveying the mllm landscape: A meta-review of current surveys,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Surveying the mllm landscape: A meta-review of current surveys,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.355118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.355118Z digest=sha256:77b1979f16aab2ffd45827a82fb5eb105295a1efb76732bf3234ef862b4d3c66

Observation f1b077f7-53f5-49e0-b917-d81568e64559 · outbound

This paper cites Learning by comparing: Boosting multi- modal affective computing through ordinal learning,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Learning by comparing: Boosting multi- modal affective computing through ordinal learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:31.087023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.359998Z digest=sha256:3d12161815a865fd83dbbaaddd90783d7c3c543ed5358b4fca2a2c2b180532d3

Observation e75fac92-5937-4d9f-99cf-2ab6992642da · outbound

This paper cites Llm-based nlg evaluation: Current status and challenges,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Llm-based nlg evaluation: Current status and challenges,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.871599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.365403Z digest=sha256:6b1af1c1ea3c484ca264e4485e8be2efad003204cf4d76590a68d11d341ff97e

Observation f3f1bbd2-9bca-4223-b3f0-c63677aba923 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.371526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.371526Z digest=sha256:2a9ad01594cb04553dbe507703188673741ab4d58c8873bf8dd71d5cdeecec12

Observation 81d76dc5-c191-4d76-9687-2d6e99d50007 · outbound

This paper cites Visual instruction tuning,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Visual instruction tuning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.718937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.378442Z digest=sha256:259f5591d45042040cd8ae8c44eed85442da7858523c3af8f60e43773dce239d

Observation 0fe48994-3171-4398-b5a2-43388ce28bca · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Gemini: A Family of Highly Capable Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.382903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.382903Z digest=sha256:46d2fdb2df4e20158378ab9d6661107a30d15420f5359e4525749982fb9997ff

Observation bd611aaf-b1b6-4e25-96de-8140eb598eae · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen-vl: A versatile vision-language model for understanding, localization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.529258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.388809Z digest=sha256:c2226c40f56e4acbfa1aeacd16f16e31af0c5e6440ead44cfcb63cf15fea06a6

Observation df5a41c4-2648-4435-a407-e74517b4513f · outbound

This paper cites A survey on multimodal large language models,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A survey on multimodal large language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.385999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.394458Z digest=sha256:d6885f1fa81cee3019baa80d239a8d22b1978c9621bed839fdf96eb9109bd222

Observation b0d994dc-4d41-410e-91a1-54be9d9d8e9a · outbound

This paper cites EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.399808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.399808Z digest=sha256:6ff2a9c1e4db97be0f359b8da7079bf31819c6c2ec1b851feddf184b55a54a37

Observation 54470714-89de-4455-97f4-f8f45267dceb · outbound

This paper cites Eemo-bench: A benchmark for multi-modal large lan- guage models on image evoked emotion assessment,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Eemo-bench: A benchmark for multi-modal large lan- guage models on image evoked emotion assessment,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.405487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.405487Z digest=sha256:6dc8440ae01dd2711fcdcfde321f780595955d64bfc8a221817527573f8730a9

Observation d1c5d6d0-bdc4-4b2a-8647-0b82c892956b · outbound

This paper cites Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.412048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.412048Z digest=sha256:b0a18c83f907fc8fab8afbaa026a8fe5955302f4b0d938f135115dbc7944cb3e

Observation 42d0d87b-261b-41e3-8ed8-a2c60f5f6119 · outbound

This paper cites The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.417545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.417545Z digest=sha256:f29103f2db13b6e0380b23e9c9e478f73095e1338925bf765aed30ff271d273c

Observation b1b62d41-cbfa-452e-a7a0-e9afabb2e9b6 · outbound

This paper cites MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.423638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.423638Z digest=sha256:02bd1df74c9b06866f48816acc4a16960891ab858bef873c27eaf2275c9cf84c

Observation 21db1851-7af0-41ee-ba55-5898ff420f84 · outbound

This paper cites Ch- sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Ch- sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.238274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.429162Z digest=sha256:6fca58068098551cba99093a2ade9d913ce9b5cccf00f23734967312ab6876aa

Observation 3309977f-1b35-4894-a074-94440234caa7 · outbound

This paper cites Make acoustic and visual cues matter: Ch-sims v2. 0 dataset and av-mixup consistent module,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Make acoustic and visual cues matter: Ch-sims v2. 0 dataset and av-mixup consistent module,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:30.092305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.434158Z digest=sha256:169f200217e53291cb467bfba4ff9caf14ac8f8888f3170dcaff191a9244f92d

Observation d915f76c-e091-4c4e-8496-012f69e9870b · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.438791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.438791Z digest=sha256:87c86d90db0ee13f0d0fb76604e5bb41b08dc9b01897bd3a3c605d041d256394

Observation 0855267a-b265-4c1c-838b-74924086cff2 · outbound

This paper cites UR-FUNNY: A Multimodal Language Dataset for Understanding Humor.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting UR-FUNNY: A Multimodal Language Dataset for Understanding Humor

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.443408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.443408Z digest=sha256:923c72eebaf25ce6aa1140fac9f5be1366bd487dde95e61c7c86f8516b2edf28

Observation fd73413d-c3a3-48cf-b9f0-4314da5146bc · outbound

This paper cites Generated Knowledge Prompting for Commonsense Reasoning.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Generated Knowledge Prompting for Commonsense Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.447997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.447997Z digest=sha256:a689c1c9874ea448f843a8d57a05006d953a743ca121538f35903523482af8bd

Observation 8fc8751e-6fb9-4930-8744-859f0a979492 · outbound

This paper cites Self-attentive feature-level fusion for multimodal emotion detection,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Self-attentive feature-level fusion for multimodal emotion detection,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.971592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.452615Z digest=sha256:bab68af65ff0da3282841d69cf188090679178098a9e079e057ea307fa201cbc

Observation adfaf75b-99c2-44bb-a848-01b9d6828b6c · outbound

This paper cites Emotion recognition using feature-level fusion of facial expressions and body gestures,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Emotion recognition using feature-level fusion of facial expressions and body gestures,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.832233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.457912Z digest=sha256:115fbee4d11ddeff7834c0a875796520e8787cbc0dcdd043b693f80425ee7681

Observation 090ae2fc-4055-4976-84ac-48158ff0fd74 · outbound

This paper cites Decision-level fusion method for emotion recognition using multimodal emotion recog- nition information,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Decision-level fusion method for emotion recognition using multimodal emotion recog- nition information,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.705955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.463005Z digest=sha256:3d6b44d1e0f4e33510774d141841d54d92cc4fe730383cacc458fe23e8bc1290

Observation 1e10bca3-bf67-4ee0-9c05-2e71ca3b791f · outbound

This paper cites Deep learning-based late fusion of multi- modal information for emotion classification of music video,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Deep learning-based late fusion of multi- modal information for emotion classification of music video,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.514991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.467810Z digest=sha256:dd2ca44fed0a5bb0ea961ea156eb95726fef8819b420cf8aeb7efc7442716550

Observation 368a5029-19d1-4bcd-9ea0-746739f59219 · outbound

This paper cites A joint cross-attention model for audio-visual fusion in dimensional emotion recognition,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A joint cross-attention model for audio-visual fusion in dimensional emotion recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.311160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.472652Z digest=sha256:0bb6c33d7c676845509f924c1532d302f355d63e1dd46e9a608a27eeda0ad252

Observation 404f1533-fea9-4bfe-a18d-5545045a8f36 · outbound

This paper cites Speech emotion recognition with co-attention based multi-level acoustic information,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Speech emotion recognition with co-attention based multi-level acoustic information,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:29.122837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.476988Z digest=sha256:1e0e23b967f686e4b34653456dd14c8e6ec257ef50912c7407c5a8f604aa28a4

Observation b284676b-38d0-44fd-978f-2acca5ce5b4a · outbound

This paper cites Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.481438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.481438Z digest=sha256:6c5292df67ffa211dacba17b56e01a238a21ab32f15b9351a2644087cc9cd0af

Observation 74964bcf-796f-4350-a856-84deb9c68674 · outbound

This paper cites OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.486614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.486614Z digest=sha256:c3008679527c7a3cd1e9cd12cd91f261bd7bb6f81a43fd0db7fd991f0d3dc789

Observation a0fec81d-4e6d-40dc-9b2f-8d553f64eefc · outbound

This paper cites Mellm: Exploring llm-powered micro-expression understanding enhanced by subtle motion perception,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Mellm: Exploring llm-powered micro-expression understanding enhanced by subtle motion perception,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.491839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.491839Z digest=sha256:99b6c3a6cb79d7b2707b9b4831dc29221f3bca15031aac94f2c4462ae70177ad

Observation a3f3bde7-2787-49c7-9fca-fcaec3ec39ca · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Learning transferable visual models from natural language supervision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.497289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.497289Z digest=sha256:d20d47d49308e97de0d6128765c5db9bffc8cd8691c6df372393bebcbe24632c

Observation c35e7560-2238-4b01-b812-b28bcbea21f0 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.930253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.504123Z digest=sha256:cc1edd0f1f886d84e9053ea84c096c1facbbaae554fb3181bccfa1c35d521e91

Observation d32b479c-3c7f-4f03-9026-72d988877b7b · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.509860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.509860Z digest=sha256:094baf0e5743a81b7334b70606ce32b64ce969a8de1ce5da0f8ca6fbbed5f28d

Observation c7731c3b-89bc-49ca-a089-6d1ee0720fdf · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.516633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.516633Z digest=sha256:b254fe3a8e1a33b8ca47a50233821c97d0172d5b4e69a3e91d8b8d53f763379b

Observation 20b4fd2d-8357-4fa8-905d-821aa6e42340 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.523176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.523176Z digest=sha256:2b96cae0dc18f4a8e7f7d86de583c5b5a98a540c302e80e9577b73a97fb0630f

Observation 5af11d9b-7bd0-4ea7-ac40-7c30cab0911d · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.529880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.529880Z digest=sha256:62f65f51b166b27cf1b635b68d3861c3f7622833c4283cabcd26d5bef289ed55

Observation 11f2de67-3088-49d6-8256-7b9f687d8eef · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.535612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.535612Z digest=sha256:74dd26590c620ed9d5753424d0e00f89a569bec6d39c3553bceac462686d0530

Observation fa56c0e6-2f13-4c61-8dc4-f259ad81b243 · outbound

This paper cites HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.540778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.540778Z digest=sha256:0c9648b97801fa4c8865e6ec29f3e05a19de85a1064c987e90573fbeeda6683a

Observation a779eae7-fa98-4831-9eb1-a1dd5b7829b9 · outbound

This paper cites Ola: Pushing the frontiers of omni-modal language model with progressive modality alignment,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Ola: Pushing the frontiers of omni-modal language model with progressive modality alignment,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.768680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.548636Z digest=sha256:f7572e366eb335c488f3998bbf4a7da88d588d2ee2eae619e7a2e5506a4511d7

Observation ce2dff1b-3623-49ca-8b58-96642e13d7f8 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2.5-Omni Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.555547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.555547Z digest=sha256:e80a9b734ae587b920a3d3bb0a8354a032608e47eface53e7e6b4cf3c6a1b750

Observation 23714ce4-ca41-49be-a199-fdb7fe05d0c0 · outbound

This paper cites Learning emotional prompt features with multiple views for visual emotion analysis,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Learning emotional prompt features with multiple views for visual emotion analysis,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.612232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.563409Z digest=sha256:9948662fccf6bd90aee6d8918b064581e6e599ce12f7a5603ec3358c4700d3e2

Observation fd4ab5fb-87cc-402c-988d-3457955a13c0 · outbound

This paper cites Visual and textual prompts in vllms for enhancing emotion recognition,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Visual and textual prompts in vllms for enhancing emotion recognition,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.478909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.570432Z digest=sha256:e3aa65a294ce2f08bab7a4a23803afcd905df9aba8156738b295ef11601b54ed

Observation 86ed16c1-d44c-4cd9-80cf-0b421f060ff1 · outbound

This paper cites Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.328838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.578864Z digest=sha256:44b6ff213efb4939e5f04c03d42da45a93c0cd6e3ec7d0c13908ae13519e1b48

Observation b4ec7c38-d45f-40db-9361-da55708317ab · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.586681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.586681Z digest=sha256:0ecab0012a21c4be2dfe22f2100059049520db95aa8f0eb7fbdaeaa9b3e4cb5c

Observation 0dd638df-6b46-4131-9a35-4098b317b9dd · outbound

This paper cites Emotion-llama: Multimodal emotion recognition and reasoning with instruction tuning,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Emotion-llama: Multimodal emotion recognition and reasoning with instruction tuning,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:28.130166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.593485Z digest=sha256:1b3df454524e196f9b5d9c2555fd221fd5edab84758748e51ad36f06c7dedc99

Observation a138cafe-fb39-4808-a7e6-ba7c95b95a7d · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting PandaGPT: One Model To Instruction-Follow Them All

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.597936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.597936Z digest=sha256:09e0110dd0180634b1effb5893db9ffef4c05f3918e7fcdfeba99ec28d25a8d1

Observation 42933506-5da1-4470-bf52-7a1a9093cb09 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Lora: Low-rank adaptation of large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.979157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.602861Z digest=sha256:b4776ccd7d0d9831bb903dd33ecc7219517eab01f00f4441cb40f54d6b165f66

Observation 23935ab7-5e3e-4cd0-bf5e-dfdfc236d63f · outbound

This paper cites Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.838128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.608419Z digest=sha256:2ac46c388805bcfa44dfe15714b2c7a24d45e8a057656d9fd962abe42984ce72

Observation f4ade031-da1b-491c-a62b-695a1d283a13 · outbound

This paper cites Injecting multimodal informa- tion into pre-trained language model for multimodal sentiment analysis,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Injecting multimodal informa- tion into pre-trained language model for multimodal sentiment analysis,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.684303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.614031Z digest=sha256:886a281b0ab7a3dc3fa2d88b61520d53d1e141481faae50808495e17be29857a

Observation 9364ce03-6343-4e38-8aba-e1f5a2243737 · outbound

This paper cites Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:02:25.031605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.619395Z digest=sha256:59a3d08233a3b5f75a5a22c423368c2bc693520aae2de0830ca7900c825b08c2

Observation 07d8b9bb-8567-4c82-b9d9-4716b492c995 · outbound

This paper cites Hgtfm: Hierarchical gating-driven transformer fusion model for robust multimodal sentiment analysis,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Hgtfm: Hierarchical gating-driven transformer fusion model for robust multimodal sentiment analysis,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.533321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.624420Z digest=sha256:df67c8d5a88d97fb777eeb51df609d4907337ab1b4f1720c064944a97952b09e

Observation 19298fd8-686c-4506-ab67-daa8ca63f8cf · outbound

This paper cites End-to-end Semantic-centric Video-based Multimodal Affective Computing.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting End-to-end Semantic-centric Video-based Multimodal Affective Computing

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:02:24.999298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.630238Z digest=sha256:4c4b9c79a76d07ee42f0b2aeccaf74b4a53a3a45e2942ab55b3dc7f6e4472b3f

Observation 721e03b3-f413-49f6-b0d5-41838fea01f9 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.635326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.635326Z digest=sha256:c616e97665995659df7633489bde4200538d289ec83c7ea66be262e29a4bcd53

Observation 50dd4fc4-8c65-4259-999c-7920a4cd450d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.641304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.641304Z digest=sha256:48d346421655d3007271d41758f36435ca5937c84067bfa7ac6887a562c79123

Observation ef9dcab4-4f63-40a7-a256-d0d59273805d · outbound

This paper cites Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.338026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.647280Z digest=sha256:f0fc265069a8adbe10977dd56d5efc5ba8a3b74c56bb9525414bb327e8a12b96

Observation e1fd8b1b-21fa-48b1-9fa3-0f4edd6ef6ae · outbound

This paper cites Qwen2.5-Coder Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2.5-Coder Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.652010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.652010Z digest=sha256:376bb26d75ade06f1f23018510f4c4b794bb2c11e9ecac11a971174794fd3f7d

Observation 85140fb3-e5f7-4992-910b-cca5bdd125be · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Robust speech recognition via large-scale weak supervi- sion,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.657476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.657476Z digest=sha256:7dc12e1d700d7a00369e200a42f4bc7dac422f8a9e759a4efea42564e70ad3ad

Observation 8c2e31f1-ad87-48ef-89f8-083ee16eb6b6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2.5-VL Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.671751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.671751Z digest=sha256:cb2e353ffa8424b72ad1e825a86ea39bad1f6668e78790a95bbc76327bd27840

Observation 4f549136-c179-4339-8972-b9a90ce551b8 · outbound

This paper cites A survey on vision transformer,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting A survey on vision transformer,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.174956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.677944Z digest=sha256:9e11c3241d25def4aadeeefd258893fbf9d0455a9d409684d4e642e13fbd8cf6

Observation 4e5f5b46-67cf-4e38-9266-19f98ad5916a · outbound

This paper cites Sigmoid loss for language image pre-training,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Sigmoid loss for language image pre-training,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:27.008775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.683813Z digest=sha256:25633d88b5ec345c39d62bfc82e382dc6127e663088d687f84f11bfc1da0c73a

Observation 8b1132f2-adf1-4797-9d46-f0ba22dcad40 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting LLaVA-OneVision: Easy Visual Task Transfer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.688796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.688796Z digest=sha256:5001a8c45971f07b9a94d07d66666e9892b7091d8997bef9acb11d4de39a4205

Observation b95fbbae-b9c3-43f3-8cf9-7f2c29168dbd · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.696608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.696608Z digest=sha256:4261debbe4c9b5ac161317af6015964b56da34033eb78f5bff28ea685c7f201d

Observation 39714fe6-3995-4275-b61f-95aee8211602 · outbound

This paper cites Qwen2 Technical Report.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.703321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.703321Z digest=sha256:4d1fc3b322725ded26ce2c6518fa75feafa3091de05f1275199fabb3a110e4ea

Observation 5548e265-db02-49f1-82f6-2416ffd59237 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Imagebind: One embedding space to bind them all,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:26.848544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.710416Z digest=sha256:758729144d4bf30c4d41f5c0c7553ca80ab70160cf6dd05f7785c268e03f69df

Observation 3fa0f8c0-2f9e-4206-a2a3-56bda65f93b1 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.717254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.717254Z digest=sha256:f51cacc1b1b0374c12adc1d3844e825630618c49dd4736f8b605f9dc9fd5be24

Observation 263337ce-aea8-41fb-9da2-b727ba36fc0d · outbound

This paper cites Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:26.617560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.723159Z digest=sha256:9f19a5b0aafd5500312d22535b25bde098015c8eee740d50ee525c3cf9404698

Observation 0b18373b-846d-452e-9c72-040a0914caa9 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale,.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Eva: Exploring the limits of masked visual representation learning at scale,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:26.349226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.730210Z digest=sha256:6f202f8d5461a1ab7677c77fa873eaf4436c37beeb7d8dcf7d9e266e3cf480ea

Observation 74da8b45-edc1-4c74-897f-a14243688431 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting LLaMA: Open and Efficient Foundation Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.737622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.737622Z digest=sha256:ee3bbf29b6eeef910cdee1d472ffeff56e9fd64863acc964fd43de5c0e5a8aff

Observation cde85087-202c-4242-87b8-d29ff4cbc18a · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:26.217538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.743337Z digest=sha256:34837f798758170a3a801efd08d1a00a5fa14f549a4c9cd2eb6a5a5fe5154f2a

Observation 917be28d-0b19-4dab-b624-cd745e496966 · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:26.031375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.755334Z digest=sha256:309ec2d37c186204f5bdc726fd84888c7eda9c6f170f29cef3fe12f08d53e045

Observation 311dbcf4-fc63-4d0f-8ad0-e80cfbc0cb9b · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:25.873698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.762492Z digest=sha256:3560b6dea934a467a846b68150d561989c20e19a733ae92eceb66ac44e8853ae

Observation abe15f6e-c07c-4049-80bb-ed134b901b3a · outbound

This paper cites an unresolved cited work.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:02:25.826168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.768191Z digest=sha256:d843705f4019c9f2771a98794dd28d17ebb8a9d037dede735a4b87bab087f786

Observation 99597bf0-f020-44c7-84c0-680b7b5a62ee · outbound

This paper cites Its core architecture follows the Thinker-Talker design.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Its core architecture follows the Thinker-Talker design

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:25.786279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.773095Z digest=sha256:1769abe454cb61d0530eda6d0f5d38dd9e8542b234d1e51c41b8bd103e02eb52

Observation 27a231c1-e83b-44f8-b2fd-748b9bd889d0 · outbound

This paper cites Its key innovation lies in the ability to simultaneously process visual and speech information in human-centric scenes.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Its key innovation lies in the ability to simultaneously process visual and speech information in human-centric scenes

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:02:25.746097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T05:02:24.778729Z digest=sha256:bdcd7d58b7212ce6232560cf0bfbea826f2d17cc04a307479b40d52c8632d604

Pith citing papers

Observation 8b65755c-ea84-4a75-8437-4557e50e4303 · inbound

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis cites this paper.

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T11:19:34.140857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:19:34.140857Z digest=sha256:96410a6e4b23fcff56b708468f2dd7a784452b95361535b6e20d469e84f5484b

Observation 73165972-423e-47ac-9fdc-f2fb77c25237 · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:55:53.245819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:e5ace3f55051621924ce6876f5b1013281b6673714e3c6cf781493e663bc7949