Pith. sign in

Paper Citation Record · LEDGER

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2409.18042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.18042 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:13:57.080098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T00:50:13.938169Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9d0958cb-b41e-4f06-ad3b-cfb81e1d753f · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:13.940657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:95ddfa573a6365c908638ab65037c5f4e2708803f1dbf7fffcd90f151c139e6d

Observation c02b3be4-debd-4f67-8ee1-6d6aa4965463 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.080098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.080098Z digest=sha256:1ce8fd5441e1ec740ba53af26926c9755ffc6609a08f2681461402473f8cf1b0

Observation 57e20474-ba2e-4d9a-90ac-ef77d669ecd1 · inbound

Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners cites this paper.

Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:12:23.961929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:12:23.961929Z digest=sha256:1a9e15828b3de3ba4bec20c7f1c3d1a255c5f87f3c2d0bdb2ac250b810ef2d20

Observation 9f44a0a4-1d8f-4374-96d4-87bb95f2eb1b · inbound

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance cites this paper.

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:29:52.565561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:29:52.565561Z digest=sha256:5b4d3bf8c9b5280b2dd45d2a108eaa8cc90aa3f5d64f054cf68b7b59d34fdd62

Observation 780a2c6b-aba3-47fd-a748-1eb4aa5c67bf · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.020705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.020705Z digest=sha256:7e57a238f573564446caa348626db5ebdf76733d3fd9b8509db85ece025b4de3

Observation 52a13339-d729-401b-a598-487c74ecd778 · inbound

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios cites this paper.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.906306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.906306Z digest=sha256:41f02602ac78dd10e25fdc8dc037bc791d67f24af3a6fcdbb4e4d75238e1d0df

Observation e168c0e2-ddbc-4808-9a19-bdca34480153 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.079119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:12.079119Z digest=sha256:a0fc82c22eb5a046bfa8c724117511b810c94373d48264e5582ec2f954390ea7

Observation 3f5707ab-84f4-4a61-aa79-9f94835d864c · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.410942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.410942Z digest=sha256:8a2a9b460d7eefed877d2b7085fb6997140fc1083d1e9ce125ae5e9fb9440053

Observation 5ceabdae-6d93-4d0a-8776-9f84088b11d8 · inbound

ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving cites this paper.

ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:50.600560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:50.600560Z digest=sha256:a5ebf8a42e6a63b073c6df84355c5ac67cbf189becc8515aa113acf7c3fd9702

Observation dd75c2f4-27d3-494d-ab24-e907d7a05ddc · inbound

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding cites this paper.

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:55.202093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:55.202093Z digest=sha256:efece2d2623fb38460a67094bf7e5fb60e42b0a2f47950b95771e655f8eb2f74

Observation 047c6f91-bf8e-4dd3-b40d-9c1b001d6043 · inbound

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models cites this paper.

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T19:28:32.792370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:28:32.792370Z digest=sha256:b2342c7ed930ed97ef9a3fee849f1a045d9800b1268c4886c739a03471a8feda

Observation d65ae3e8-d898-4b20-bd84-c976d9391b04 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:12.988342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:12.988342Z digest=sha256:a744e43dd85fad9568a7c8633b280078cda84f4a32c50f6cbaff0108296f8f8b