Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

As of 19 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 4 inbound Pith citation observations for arXiv:2504.16427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16427 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:08:40.817531Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:02:24.412048Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:46:23.764672Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3a14c33-b305-4cca-8569-e28103db6f68 · outbound

This paper cites Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Infinibench: A comprehensive benchmark for large multimodal models in very long video understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:37.871037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:37.871037Z digest=sha256:a7630c2eff8cd1e2449f5ae38d97d7587232ec8a420642d1dbfa25d786e5d338

Observation 19c4bf43-44e3-4fbd-8124-2d4d8622084c · outbound

This paper cites Language models are few-shot learners.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:37.915598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:37.915598Z digest=sha256:a6faa2f552d3553b41b295ddc62f71d18030904811b253954544380b5ea4e085

Observation d71a2953-d875-450c-9bcb-aede9d2a644f · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Iemocap: Interactive emotional dyadic motion capture database

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:37.975196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:37.975196Z digest=sha256:780a3692796e3e0b60f94cf6ec60292dd7a37d0ce0ee807322e8a52695e1fa3b

Observation 5e7792af-859f-4853-9d18-ebf1dc08ab71 · outbound

This paper cites InternLM2 Technical Report.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark InternLM2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:37.981020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:37.981020Z digest=sha256:0e355de4112abcefb7c39979223297c76f227544871aff0256dd400cd15be89b

Observation da28b228-c315-47cf-ac66-63f5daa2113a · outbound

This paper cites Towards multimodal sarcasm detection (an _obviously_ perfect paper).

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Towards multimodal sarcasm detection (an _obviously_ perfect paper)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:37.986334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:37.986334Z digest=sha256:39879d07c6eff1bc69a089f39866f676fde1ac079d3f97e981bf3dc754f86a1a

Observation 2e1a1f20-47d1-492b-8675-0b085c49b697 · outbound

This paper cites D2R: Dual-branch dynamic routing network for multimodal sentiment detection.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark D2R: Dual-branch dynamic routing network for multimodal sentiment detection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.061928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.061928Z digest=sha256:1b5056f1e684c234a3f795c6088fc2abeaeda479a580febf5cd8657a69bded6f

Observation cb5e928b-a854-4cbd-bff8-d033367bf2ba · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.067609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.067609Z digest=sha256:0b1c965f8d447f9882119d91b666816a3f8214e368c0e761a62fffe5960829d6

Observation f72e9338-a6f8-494d-842c-ed09abd8cbcb · outbound

This paper cites Schuller.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Schuller

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.072726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.072726Z digest=sha256:bb9ceab75d673cee953e0e15ee5610534e1a2b319fab5d0afbb7c7492e88ad2f

Observation c7d3d620-22fe-482f-b681-eb7b5937a43e · outbound

This paper cites User attention- guided multimodal dialog systems.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark User attention- guided multimodal dialog systems

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.587668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.077382Z digest=sha256:79bad8956363a4733a0e9eb40a411e6073263d6bcfc877c6c82d8a3415a8ceb6

Observation cd176bee-9f3e-46f7-b325-2b00b7cb63bc · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.436901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.167088Z digest=sha256:9541e114ba0e0db0d46d68fea27008d0318f3b1ac759c98650ca96c3e3f03dc5

Observation a40b1bb8-3ddc-4b07-870c-50e4a4e7ba34 · outbound

This paper cites Multimodal sentiment analysis: a survey of methods, trends, and challenges.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multimodal sentiment analysis: a survey of methods, trends, and challenges

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.421500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.230979Z digest=sha256:ace568af099846300201ce044e91e4ce1eb97484952d632ac2b01cb4d1b51d16

Observation f0c6d9ac-95cf-4362-be3e-11c3b62ed52e · outbound

This paper cites Simmmdg: A simple and effective framework for multi-modal domain generalization.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Simmmdg: A simple and effective framework for multi-modal domain generalization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.333453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.235799Z digest=sha256:1aa3944e3fdd8b58940b97d3feb4e99a0b5b771fd319e3abd399d67a9547d388

Observation a4076023-83c3-494b-a17d-6c8d5857fcfd · outbound

This paper cites The Llama 3 Herd of Models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.304925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.304925Z digest=sha256:c8b6a7897c1fa07b00ed54b45c4ddfa339da77b0ade0643d849db15d09b4cfca

Observation ee6e944a-e3da-4951-bb5a-b211ae727139 · outbound

This paper cites An argument for basic emotions.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark An argument for basic emotions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.310466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.310466Z digest=sha256:c4d5fe2fd839714c509e025edf040721902c07ff4f91603bb18dc0cca2c816bf

Observation 5c88186a-1981-4524-82b7-716d9018c872 · outbound

This paper cites Facial signs of emotional experience.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Facial signs of emotional experience

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.268639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.314962Z digest=sha256:942472a480afa161465a87b838346a6bfdab8b02e1ae83d9769d3ed77a425f89

Observation ade2840b-5cb6-4e3e-9f89-dc08e6d45abd · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.374553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.374553Z digest=sha256:6b82ee2d8a09ebf655dfed3cbbd98c91b24781313ead3b502f844977c1123ad1

Observation cc3502c1-e0f3-40a5-900a-0e306aa5a16f · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.380021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.380021Z digest=sha256:ec6886282878cb50d490d71783edda98a365cd2aff6db3f23a8b76420b671415

Observation fb7acf6a-1fb0-4e43-9e50-6158578d61c7 · outbound

This paper cites Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.252117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.384727Z digest=sha256:b2d0d935aa14d6e5e84921105511aba1353d8a2cb25a759a0bd38b3dcc655915

Observation 0d2ba2ae-1739-4f6b-b1bb-31d9ab244cca · outbound

This paper cites Switchboard: Telephone speech corpus for research and development.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Switchboard: Telephone speech corpus for research and development

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.186699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.469275Z digest=sha256:a3d1e3d6a1a3f288ac183de6cc66454b332947bec377873cce96906390fad6d7

Observation d2842141-74af-45cf-b2f4-07beecf28654 · outbound

This paper cites Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.098269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.474867Z digest=sha256:0f097cb66622fb219d937e092a2807c4863182b071ebf2dbaaa074c66850b054

Observation 3b583e2d-850d-4898-addc-ee3c7a929174 · outbound

This paper cites Ur-funny: A multimodal language dataset for understanding humor.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Ur-funny: A multimodal language dataset for understanding humor

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:45.079780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.479980Z digest=sha256:ba667784807dc69c3fbcf338636f8defa6046a6132a1939e1c688e040ef18c3b

Observation 90a6f9f0-8ab0-4483-ac05-14ec3d0ed785 · outbound

This paper cites Humor knowledge enriched transformer for understanding multimodal humor.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Humor knowledge enriched transformer for understanding multimodal humor

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.984181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.502754Z digest=sha256:17c21aa230477ab5ff8ba0204dc197d2606e4984ef51d160f332ee71ed282c51

Observation 5a9c67c5-d140-4787-8d53-c864e09bb09c · outbound

This paper cites Misa: Modality-invariant and-specific representations for multimodal sentiment analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Misa: Modality-invariant and-specific representations for multimodal sentiment analysis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.569431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.569431Z digest=sha256:46c1468c5a5f1e6af3a60a87e72423a345823f6368e106817165a7c342e2323d

Observation af3a8dec-cb73-4483-ac8d-f9822faf315a · outbound

This paper cites Mm-dfn: Multimodal dynamic fusion network for emotion recognition in conversations.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Mm-dfn: Multimodal dynamic fusion network for emotion recognition in conversations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.875311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.573896Z digest=sha256:38257e836d3f17c5fafde2be819d2cbb0be6a2afeffd2777541d987e907f799f

Observation 7b9f2641-17cf-41c3-9381-0feafde64202 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark LoRA: Low-rank adaptation of large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.803295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.579630Z digest=sha256:43340038f97535322248b44ba387f63c9e2343f0616fbc3112ef30381d316015

Observation 27654d8c-a8b6-49de-8d4e-63c34ee3d276 · outbound

This paper cites UniMSE: Towards unified multimodal sentiment analysis and emotion recognition.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark UniMSE: Towards unified multimodal sentiment analysis and emotion recognition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.762669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.660914Z digest=sha256:0f411ae197e3045eb6dc626157bf6dac8b736ef052ab728053097cf10e2edcc4

Observation 8d28c96e-8552-494d-b893-a00c36223d37 · outbound

This paper cites GPT-4o System Card.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark GPT-4o System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.681895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.681895Z digest=sha256:79a7fad89b18c13ec3b92b4fd8e2c3997d5a135767493a500cb2bff706327e8c

Observation 947d491a-78f2-462e-945a-22f3d5def60f · outbound

This paper cites Mistral 7B.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.701324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.701324Z digest=sha256:5f6c12d246ed4233daa0c23a061cc662727c459f80884ce7916479453b140cf1

Observation 2fc860b5-8636-4ed9-bc59-2eccf212bf26 · outbound

This paper cites Scaling Laws for Neural Language Models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Scaling Laws for Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.840249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.840249Z digest=sha256:a5f5eb3e18cb304a926d55892348572b27a5de2195ed3ac6996de069af950747

Observation 39e28ba5-f31a-435d-9280-533ff3fd60f1 · outbound

This paper cites Integrating text and image: Determining multimodal document intent in Instagram posts.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Integrating text and image: Determining multimodal document intent in Instagram posts

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.638892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:38.952110Z digest=sha256:b2006d8a2455288f5f1e37e27bc7c356d1cabef6ec3dc338a965da4894011143

Observation d47ec44b-c702-46d8-b5dc-3c08168203c7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:39.065101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:39.065101Z digest=sha256:8b2c11d03aefd684e038d74922f6b37b4bdaf4e477a3b510dd9609c7ab016222

Observation 1aa8c0e9-67b1-4c19-9fc5-0d34f9a1919f · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Seed-bench: Benchmarking multimodal large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:39.070651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:39.070651Z digest=sha256:3de42e5f316bbc170659dd9749028194e1b8749667b4ec9e4dfe759bc54fafa2

Observation 119f6cae-a9ee-48cb-ba02-7e8cea7b47f8 · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.546997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.112243Z digest=sha256:a3c4fc65ac726401e02ea595a4cf03ef51f7eed94e661368c5193a844bf3753e

Observation aeea295b-3595-4442-b04f-2ab106c40975 · outbound

This paper cites Multibench: Multiscale benchmarks for multimodal representation learning.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multibench: Multiscale benchmarks for multimodal representation learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.445770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.198496Z digest=sha256:418a5fe37db0e820730feaef000ab00d03342f81846ce2a626843d229fdad429

Observation 865cd7c5-ccd7-4bac-82a8-131b61b04391 · outbound

This paper cites Visual instruction tuning.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.328505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.203226Z digest=sha256:f9b7aaf1e6d3b0a059653c61f0f6c6c2d57e72103571fbc30a0c0fff73c3e4be

Observation 9c393b94-bb3f-4f76-9dd4-1b6faf78d448 · outbound

This paper cites Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:39.206914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:39.206914Z digest=sha256:0a2b595f22d49ab48174d8b4ff230b015382fe1d22e4b22011afe05a979cf3fb

Observation d66dccdc-d4da-4f65-ab8b-05ff1d6c8c2b · outbound

This paper cites Make Acoustic and Visual Cues Matter: CH-SIMS v2.0 Dataset and A V-Mixup Consistent Module.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Make Acoustic and Visual Cues Matter: CH-SIMS v2.0 Dataset and A V-Mixup Consistent Module

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.218914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.292548Z digest=sha256:82a210e4a187dc5cb13c6eb5954068272046a444138e755d6b300e08c9f7008e

Observation 5da65a07-2fff-490e-a7a4-4bdbb5e3029a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, pages 216–233, 2025.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, pages 216–233, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.204115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.297300Z digest=sha256:4a8ca88ee37db887c634bf46bc70363b29b8a8d30c843c73473477f7ec29a917

Observation a5b70296-aaff-48db-9a22-4464df8b72d6 · outbound

This paper cites Efficient low-rank multimodal fusion with modality-specific factors.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Efficient low-rank multimodal fusion with modality-specific factors

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.146054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.380216Z digest=sha256:1f2fdeb2545d5cfddb90e84a1501b8c9e3d1733b7a88f8e38ae12efad665275d

Observation 7a7813f3-33cc-4abd-93fd-10a4346d0b53 · outbound

This paper cites Towards determining perceived audience intent for multimodal social media posts using the theory of reasoned action.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Towards determining perceived audience intent for multimodal social media posts using the theory of reasoned action

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.088831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.399304Z digest=sha256:ab40674526830a0bd93efea6ae855c4dea9a9de3d11a801967804a77b7ed762e

Observation c64c6f4a-895c-4860-8d40-76d0a9a08a4f · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:39.404640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:39.404640Z digest=sha256:29d6465dd12a610a62ae95259bcc29a24221b16bc0244267c76305e9751b8826

Observation 4d1b4716-6586-4069-b2af-b81f8705b495 · outbound

This paper cites Patro, Mayank Lunayach, Deepankar Srivastava, Sarvesh, Hunar Singh, and Vinay P.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Patro, Mayank Lunayach, Deepankar Srivastava, Sarvesh, Hunar Singh, and Vinay P

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:44.024931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.409559Z digest=sha256:c8d95676bae38b4ba1c858c90dd6cf430545f8a6a8363c3bd7d36ddd9c4f0af3

Observation d61b2698-0210-4184-afec-e0ec84552b10 · outbound

This paper cites Psychological aspects of natural language use: Our words, our selves.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Psychological aspects of natural language use: Our words, our selves

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.957526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.458920Z digest=sha256:ae6d75ce651dfde0c7f0786a6b7920d75f1e31e64986d02a2d49744c4f8f52cb

Observation 471a97ad-02c0-4949-90c5-57cf372717f5 · outbound

This paper cites Utterance-level multimodal sentiment analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Utterance-level multimodal sentiment analysis

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.839056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.552314Z digest=sha256:32fae78f777e5a0c0eed95ea9786be6a4b2b985a7e15a985daf41d8cda8acd5c

Observation 8dde7a8a-103a-47b3-acfd-d2c0f2ce05d7 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversations.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark MELD: A multimodal multi-party dataset for emotion recognition in conversations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.669348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.557588Z digest=sha256:d22c6b4be9c1f09bc870bc8d6b12fd6be0983623f067caee3046592c3a821476

Observation e1c4bca9-55d1-475a-a117-b80058d1196c · outbound

This paper cites Patel Johns.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Patel Johns

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.586289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.563240Z digest=sha256:f2f466e2a817f180108323324b7c949b693aa6ce0f2ea1df48b0991b7ae9cb99

Observation e3a577f4-1fa3-49f9-96f2-0b5d2599c2fc · outbound

This paper cites A survey on spoken language understanding: Recent advances and new frontiers.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark A survey on spoken language understanding: Recent advances and new frontiers

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.535546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.567692Z digest=sha256:fecd422c0f2a727e6687c4390bad5ecacccbe360de247f5fc52cbbd2db893ec7

Observation 8b3d3dab-013f-49ab-8e5a-22acf26307ba · outbound

This paper cites Integrating multimodal information in large pretrained transformers.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Integrating multimodal information in large pretrained transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.442573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.572024Z digest=sha256:90d8de147a2d32bdf3fc64adb9633ade35b116f0d8c6b15e449a6d6cc9a8a2a4

Observation 53f3cfa4-3d40-40eb-8826-9105539b7f86 · outbound

This paper cites Towards emotion-aided multi-modal dialogue act classification.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Towards emotion-aided multi-modal dialogue act classification

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.389091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.577621Z digest=sha256:0469af94e8c65a718445288d8b9e0a19522cabd19e2cb2823da57fd26990c35b

Observation 0a399ca0-9dce-4b9c-ab3f-c50a2b826550 · outbound

This paper cites Intention, emotion, and action: A neural theory based on semantic pointers.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Intention, emotion, and action: A neural theory based on semantic pointers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.305364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.655300Z digest=sha256:ff3cbe8d29197db47cd0d2627630acf27cac840dd9e975ade1f439b75cd15c77

Observation 36512453-103a-49f1-907f-3b41840f257e · outbound

This paper cites Dialogue act modeling for automatic tagging and recognition of conversational speech.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Dialogue act modeling for automatic tagging and recognition of conversational speech

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.221849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.688200Z digest=sha256:276b4977f727ffada6c703433680bb8d982be9b8010811cccd68399281872423

Observation c8c4f1f2-f0cf-4393-becc-a2f3e982e5c5 · outbound

This paper cites Contextual augmented global contrast for multi- modal intent recognition.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Contextual augmented global contrast for multi- modal intent recognition

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.126297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.692841Z digest=sha256:a953a87b46b102e9d4832a9cf2606edcd0ad2a00e24500dfa38bc22846ce1590

Observation e18dc8fb-0e50-4793-af3a-8786857f0e7f · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multimodal transformer for unaligned multimodal language sequences

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:43.057434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.698130Z digest=sha256:2a76eae17ba273c07caa94be7592082773fa584b18a1d2e226b01588648f835a

Observation 47a8fc1f-c42a-4b15-9dd7-a8007c8c5401 · outbound

This paper cites Social signal processing: Survey of an emerging domain.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Social signal processing: Survey of an emerging domain

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:39.766476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:39.766476Z digest=sha256:d6d53be0f69fcc3bb3f75f694036a08f3ef22a0d4c83c66e9e50d9c7c9f182bf

Observation 15855b66-ef3b-4a10-b816-89cc3d2682a9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:39.848301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:39.848301Z digest=sha256:e9af8b2067cd8a45abb041750f0086c5645bdb638aafa4cee1ef679da6b0acca

Observation 74645432-16b9-43d7-bc8a-8014222e58a3 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Cogvlm: Visual expert for pretrained language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.944634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.854625Z digest=sha256:5028fe541d2f553a04f037b2576c0dc8be20137a5ddf1e3808f709064194bf3a

Observation fe9744bb-c876-4374-b105-e1e2b768ddb4 · outbound

This paper cites Words can shift: Dynamically adjusting word representations using nonverbal behaviors.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Words can shift: Dynamically adjusting word representations using nonverbal behaviors

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.804229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.969951Z digest=sha256:7c949033d55688eea35ce215e0db4c8b1cd0a5824ea43d971e98818238125ec0

Observation 4975890b-8ff7-49e9-9f94-a86c53b828b3 · outbound

This paper cites SIMMC-VR: A task-oriented multimodal dialog dataset with situated and immersive VR streams.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark SIMMC-VR: A task-oriented multimodal dialog dataset with situated and immersive VR streams

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.707828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:39.975509Z digest=sha256:ca94c880af3a54e746ad18f93fae27d0b8baecc2a9df8c9f223688b797518afe

Observation 9d5946ea-0cd3-4758-9b79-c198e98830d2 · outbound

This paper cites Multimodal multi-loss fusion network for sentiment analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multimodal multi-loss fusion network for sentiment analysis

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.620570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.022390Z digest=sha256:3dec8300298a8e06ce3aaacd7c6b89b011c493cef42b60f7322a8319c8da073c

Observation 444fefc8-97d3-4964-b0c0-ec38721e8162 · outbound

This paper cites Anno-mi: A dataset of expert-annotated counselling dialogues.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Anno-mi: A dataset of expert-annotated counselling dialogues

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.517477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.073385Z digest=sha256:b4058f845c0b92acac85cc6cc0ecf6052872897d2b5f0af35f61340a59e2dd79

Observation 7b53f925-57c5-4072-bd8e-234c899b29c1 · outbound

This paper cites Qwen2 Technical Report.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Qwen2 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.078964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.078964Z digest=sha256:b2c4538377dde0228311e06c5f0f8a75daaa28af902aa530df345ec0c19aae2c

Observation a2c0f49c-9a01-4393-9945-ce5ef6283e82 · outbound

This paper cites Disentangled representation learning for multimodal emotion recognition.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Disentangled representation learning for multimodal emotion recognition

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.452315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.083645Z digest=sha256:2cd47e78f3e9441731260becc80d6128420dd258a738a57556b7c98a72b95ad9

Observation f3231b74-9e57-4b71-bdd6-c7f295d0b554 · outbound

This paper cites EmoLLM: Multimodal Emotional Understanding Meets Large Language Models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark EmoLLM: Multimodal Emotional Understanding Meets Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.180088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.180088Z digest=sha256:8236dd17615d9eef2fd1df5487740195cdbff863c97cb0f9e030e348498c7cb2

Observation a730faf5-5f2e-4d2f-aa19-838a99afd977 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.185939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.185939Z digest=sha256:a32169679e2cd7734207f4737b2a075f2da07e074c796676b5b531a67cf9333c

Observation c8a3b504-dfda-40e6-9f92-d8065474c606 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark A Survey on Multimodal Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.190730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.190730Z digest=sha256:e5234a0170c20710123055d00e9c0aa7a691fc871f507b3bfb332d652faa2cec

Observation e063028a-a47a-4d86-b325-ec06d6a4a4fc · outbound

This paper cites CH-SIMS: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark CH-SIMS: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.410954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.305700Z digest=sha256:dd88b5a7652eb8e63ac9d83ede8e4f6976ced5f1d1e88244acbd914b44f1aa79

Observation a30ec358-0d92-4ae3-82cc-c27a3991314a · outbound

This paper cites Learning modality-specific representations with self- supervised multi-task learning for multimodal sentiment analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Learning modality-specific representations with self- supervised multi-task learning for multimodal sentiment analysis

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.336158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.311223Z digest=sha256:1d1f5e6fbac85ffd97381985058dd4d4a2a25d35acb0229fcd6843de231a8cc7

Observation 53344c47-db9b-4e9f-bc69-3d8761af55e1 · outbound

This paper cites MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.316302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.316302Z digest=sha256:b55778a0f06e36f631814285216a8a72269da87dcaef0f7c73384d36bd51d78f

Observation f820054d-e7ea-4929-bf7f-57d64a971283 · outbound

This paper cites Tensor fusion network for multimodal sentiment analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Tensor fusion network for multimodal sentiment analysis

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.280744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.321699Z digest=sha256:33faf9960797c4ec97c7498bc4066f82ebc706e305bc71d6954de66546739e92

Observation 87aadbad-d4cb-4676-b4ed-53029761aab5 · outbound

This paper cites Memory fusion network for multi-view sequential learning.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Memory fusion network for multi-view sequential learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.229896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.380625Z digest=sha256:db7c99a33750fcf3dc28f0b8a868f07b0397d529c3f66bf69ee5c926a8a15051

Observation 6afd95ff-6a93-408b-8e52-9540350f0b31 · outbound

This paper cites Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.173414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.385181Z digest=sha256:0bc6b66a22562a933b4451908858567790bc7bf99682ae3385416312f1744ff3

Observation 7fae6616-b871-4df7-81a0-a74f112328c0 · outbound

This paper cites Multimet: A multimodal dataset for metaphor understanding.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multimet: A multimodal dataset for metaphor understanding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.093030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.469761Z digest=sha256:2964fd85433c4da749cf5d04ac18cb4420e1f89cb5dcd8eb1c4a7260b037c622

Observation 434bee27-b950-4083-ac4e-94e29c5b7c4d · outbound

This paper cites Mintrec: A new dataset for multimodal intent recognition.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Mintrec: A new dataset for multimodal intent recognition

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:42.032039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.474310Z digest=sha256:442186444f86c75e96d170b33af873237b6e964c982fd643ec2b9bf046285c1c

Observation 735f697e-b445-44bf-a5fc-24c274cf09b5 · outbound

This paper cites Mintrec2.0: A large-scale benchmark dataset for multimodal intent recognition and out-of-scope detection in conversations.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Mintrec2.0: A large-scale benchmark dataset for multimodal intent recognition and out-of-scope detection in conversations

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:41.963455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.479575Z digest=sha256:a27b84e4bca0c1bf238f96cd1c25185995fcfec18ff7a72cf34adfd5b0dfdb18

Observation dd3dc599-cb9e-4264-96ca-2556a3e56b0c · outbound

This paper cites Unsupervised multimodal clustering for seman- tics discovery in multimodal utterances.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Unsupervised multimodal clustering for seman- tics discovery in multimodal utterances

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:41.886102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.485078Z digest=sha256:4cd4c64a3ac8bb8830fe411ed056decec9db7c902752fc22685b8c22556fc56f

Observation 99ee4d5f-0a8e-4a7f-854b-f3f40dbdf274 · outbound

This paper cites Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.489992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.489992Z digest=sha256:6bc1a4720a4850b87bfd2e5fabe10f44524c47b97aa2def2da12d6e08e4c9cfb

Observation 08287610-78e6-485a-84cc-a29ea89042c1 · outbound

This paper cites Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.557915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.557915Z digest=sha256:edc1d772ed20e27e4d193253d7f5cd1bdc1ef0d4f2a556bc1e87ed3a345350ae

Observation 1bcfecd4-2a30-47ef-8a61-93d782b773f2 · outbound

This paper cites Amanda: Adaptively modality-balanced domain adap- tation for multimodal emotion recognition.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Amanda: Adaptively modality-balanced domain adap- tation for multimodal emotion recognition

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:41.821189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.564358Z digest=sha256:8be74df05ba3e9ca1bf45a84c36877ed65dbdee648aa16a5ae945edb7fc9b242

Observation 76dc3447-b618-4e74-a32b-13beec9dc583 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.568945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.568945Z digest=sha256:280c6200ea8cd7858999cdb73c577ccddac353c5c8d692e9a06c7a8710ea7385

Observation a01149ec-22df-468d-a885-fba48df25f6b · outbound

This paper cites Multi-modal sarcasm generation: dataset and solution.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Multi-modal sarcasm generation: dataset and solution

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:41.725468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.648262Z digest=sha256:1f59d491436e3ea944c2fed2109dee0676541b859e67c68fdf32c5c05e819db1

Observation b439ce78-287e-4c90-a8fa-283939355fa9 · outbound

This paper cites Swift: a scalable lightweight infrastructure for fine-tuning.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Swift: a scalable lightweight infrastructure for fine-tuning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:40.696604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:40.696604Z digest=sha256:035865a41ab62d3a5af693a00d32f0947c4f83a9bc22d7f8f4b1e95948d4b48e

Observation c5e33433-966e-48d4-baeb-86c36a6e1bac · outbound

This paper cites LlamaFactory: Unified efficient fine-tuning of 100+ language models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark LlamaFactory: Unified efficient fine-tuning of 100+ language models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:41.632759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.802761Z digest=sha256:1ec804eef1524e04a3a18b933d21a0ee2520c039b845cd06e8efa4a72a06bbed

Observation 1d44e8d7-d645-4208-bbd8-967fccf1fe33 · outbound

This paper cites Token-level contrastive learning with modality-aware prompting for multimodal intent recognition.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Token-level contrastive learning with modality-aware prompting for multimodal intent recognition

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:41.579008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.807580Z digest=sha256:3b1ec618aa0eefc7b26e21bfd79998d05d7ef685bad4561ec85c6ff2997f7f8f

Observation c3d8a5fd-5d8f-48e9-81b7-c961293c21c7 · outbound

This paper cites Minigpt-4: Enhancing vision- language understanding with advanced large language models.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Minigpt-4: Enhancing vision- language understanding with advanced large language models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:08:41.462964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.811975Z digest=sha256:bbbdafae15481a023c5ebf8ad145f42558ced7ea394c93796153957c07c750d3

Observation cb92f005-81a2-4e57-bc89-9fd505493ac1 · outbound

This paper cites Inmu-net: advancing multi-modal intent detection via information bottleneck and multi- sensory processing.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark Inmu-net: advancing multi-modal intent detection via information bottleneck and multi- sensory processing

Reference 85

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T11:08:41.078665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:08:40.817531Z digest=sha256:0c07f1beb5ad5f030855ca64ba911cfeb7c25b20aa6833c387119ce7649a3a78

Pith citing papers

Observation d1c5d6d0-bdc4-4b2a-8647-0b82c892956b · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.412048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.412048Z digest=sha256:52101b0f88fd7f3f09439b8f43df7a5cfbbf13a064a824305f6c9550fbdbcc9a

Observation 18792731-eeab-4556-ab9c-2aeb3c878e2a · inbound

Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy cites this paper.

Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:46:23.768172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T17:43:32.593075Z digest=sha256:65009d72738425819bfbce66f25c4da832f2a62a2622b6ac507320629c10c569

Observation dbec60f5-baec-4a30-9d45-2ce614f53014 · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:55:53.277250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:5396de64bee31e30146883b4a6732fd4fd6853c67f70adb84c3c010616d49a16

Observation faf1b855-e48f-42da-8822-792aaf5a78a9 · inbound

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention cites this paper.

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:52:08.446084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:52:08.446084Z digest=sha256:aa6675af7be4b8e3d7c16397220d37aa08968c6e38f7944111226716105e47b1