Pith. sign in

Paper Citation Record · LEDGER

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

As of 17 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2507.04955.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04955 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:06.423278Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:02.091028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:40:07.168991Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e40e093-95c3-427f-b276-9e5220945e43 · outbound

This paper cites EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.236634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.091028Z digest=sha256:81e6fdd64c16bc924cfe13f301c9f3cf586d41b9eca6ef1413bfa7f8ad808259

Observation a5161bbe-ecc6-451f-9590-4f7a5a01a468 · outbound

This paper cites Visual and Motion-Based Control Early interactive sys- tems mapped facial or bodily features directly to sound.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual and Motion-Based Control Early interactive sys- tems mapped facial or bodily features directly to sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.582848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.126591Z digest=sha256:df2a630bae5ec79dee7660a3c5d4ff0ab93efd3b24490da1719fd0f1b2ae8d68

Observation f50d091c-f22a-40d2-8453-b09dcb9ae82a · outbound

This paper cites We froze the parame- ters of the vanilla Musicgen during training to preserve its text understanding ability.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We froze the parame- ters of the vanilla Musicgen during training to preserve its text understanding ability

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.439452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.174306Z digest=sha256:afcc47e1b6c95d3272fc8bae7b2fb3442b8f7492d76a00063c70c6051f41f720

Observation c1dde247-c637-4b46-943d-0d3213644c8a · outbound

This paper cites We recruited volunteers to record their facial expressions and upper body movements while listening to 30-second audio clips.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We recruited volunteers to record their facial expressions and upper body movements while listening to 30-second audio clips

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.334019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.274703Z digest=sha256:6cd46950e395111815114e55a0853a4b10099ca7fcb5943d4d09936c12493fc2

Observation 28781a49-92c9-4ce2-9d9a-24267fca8f1c · outbound

This paper cites an unresolved cited work.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:40:13.173796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.335175Z digest=sha256:3ac1f6d8e93cb73d0960acb51efa9a4d5076f22641476f59ab9df9213f4e5e19

Observation 33c5a219-3486-401f-9e06-230fdbce7c2e · outbound

This paper cites an unresolved cited work.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:40:13.017260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.420427Z digest=sha256:cef11b7159ac1f83b124366434751716b637ae94a59db9f6172a01eb5debf52b

Observation 5411faeb-ed22-49a3-b1a1-44ddb2381623 · outbound

This paper cites Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthe- sis,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthe- sis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.895724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.506180Z digest=sha256:f43da38c45a33f82a179a7f739615b7c2cb5f4bf0718379c36909ac490943f3a

Observation 9edde2b6-174f-4fb3-9ae6-bb3ffff98c6f · outbound

This paper cites Diff-Foley: Syn- chronized Video-to-Audio Synthesis with Latent Dif- fusion Models,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Diff-Foley: Syn- chronized Video-to-Audio Synthesis with Latent Dif- fusion Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.771095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.572574Z digest=sha256:2da66e8849f31f715bacce6d77e3f074dab5bb79029b6f5d15bf10744c88380a

Observation 280c36b4-c927-457d-bcaa-f567ebd5a954 · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Temporally Aligned Audio for Video with Autoregression,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.623336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.639668Z digest=sha256:49db1a9b39260b1e182eff84b650a13daa045e17261961ead60f0b5a0aab3eb2

Observation 6607a917-256c-4f2f-bd66-2a9b9f04c558 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:02.722459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:02.722459Z digest=sha256:a34dadc5f1771c9fd7ecb04a3ea32f6083f154563a3350fae3d214915d8967fd

Observation 7ca01989-1023-4bec-a53f-64ee01474ca4 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.493699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.784427Z digest=sha256:b09132b9122bb3b31784fcad7a1df55a86fe2601adc891621841dc0bdb9c72cd

Observation 6d1a5a91-9a00-4e55-92d7-7bf31a0b806f · outbound

This paper cites MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.370291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.871106Z digest=sha256:47344cb5e02796400836850633cdda6dc3ca1bb26cb34069feafd9b33940a293

Observation 3591a8b3-2211-42de-a461-5c3226339e1f · outbound

This paper cites Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.069029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.934968Z digest=sha256:7cf5e8f4023fc1150ba3430641b240ba9854b3b9d6ca038a2fa96336cd61ad9a

Observation 347c438d-e88c-4438-a8f0-87b421fa314a · outbound

This paper cites Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:02.986485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:02.986485Z digest=sha256:ffa8b622cd38f476c7d8f06fa0ebd6fc611fe8b54bb9109f99b021721d031d44

Observation 98613fea-9f29-40f2-a7ed-af9861c212ed · outbound

This paper cites InstaGen: Enhancing Object Detection by Training on Synthetic Dataset.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:40:06.861777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.033155Z digest=sha256:ec663c71beb5ceaac0fea53b4a4661f724cfacae967dd3b1127d0d3a3aa695e0

Observation 1914b229-406f-4b65-95f1-44bd64a2e926 · outbound

This paper cites Orbital angular momentum of Bloch electrons: equilibrium formulation, magneto-electric phenomena, and the orbital Hall effect.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Orbital angular momentum of Bloch electrons: equilibrium formulation, magneto-electric phenomena, and the orbital Hall effect

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:40:06.700771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.124928Z digest=sha256:8cda57667587aaccbb644e278a22197af751db7d74cdb3f5e9c490c5c25f299b

Observation 3268cad6-05c1-4629-991e-60b6452c29b3 · outbound

This paper cites Conditional Generation of Audio from Video via Fo- ley Analogies,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Conditional Generation of Audio from Video via Fo- ley Analogies,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.236091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.203009Z digest=sha256:13a1ef73b7124fa7f574a398f61b4af29256e1472adb5777a4ddc3030f0516f5

Observation 570ab3ab-09c6-4f44-a80b-eb1002082a44 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:03.310979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:03.310979Z digest=sha256:789131dd798a6b3f33314af9676ce9315247c09f232de97f5934b2ff1cc4f889

Observation 5c4325b2-3013-4bf5-94e1-01675403956d · outbound

This paper cites VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.123750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.401713Z digest=sha256:91af790c4ca69b79670a8788fd8d91c66fc89e2c17c8e3737e37164811ecb5a8

Observation ace31055-2ce3-4c57-9ada-33435a441f6d · outbound

This paper cites Video2Music: Suitable Music Generation from Videos using an Affec- tive Multimodal Transformer model,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2Music: Suitable Music Generation from Videos using an Affec- tive Multimodal Transformer model,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.993773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.480431Z digest=sha256:80220edc84591972b8f111f77b8335ea6e4a56dfcf0c52c4f5c1e1cabafed9c7

Observation 272d2f30-f92d-4de2-88de-12a9bd7183b7 · outbound

This paper cites V2Meow: Meowing to the Visual Beat via Video-to-Music Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation V2Meow: Meowing to the Visual Beat via Video-to-Music Generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.847780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.578377Z digest=sha256:339220be3d09affb247c7b95081fac91ae087aaead138913b8198604960e2430

Observation df50cb85-48d6-4390-8928-3e959bdb753e · outbound

This paper cites DanceCom- poser: Dance-to-Music Generation Using a Progressive Conditional Music Generator,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation DanceCom- poser: Dance-to-Music Generation Using a Progressive Conditional Music Generator,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.704582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.686152Z digest=sha256:143a28ee5df14ce369786d26d0603e40f8a684aa3a8dea923b593d6d117619de

Observation 7e4af507-c12d-48d4-ad6b-2f3f28224bbe · outbound

This paper cites Discrete Contrastive Diffusion for Cross- Modal Music and Image Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Discrete Contrastive Diffusion for Cross- Modal Music and Image Generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.528547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.792531Z digest=sha256:c7c9614c9e69fc982fbcf5ab65368eaf7e0d3aed22393c7a55b06b6779ead498

Observation 7c3a79c2-7f4f-4cd3-a3e4-95985327a9f3 · outbound

This paper cites Long- Term Rhythmic Video Soundtracker,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Long- Term Rhythmic Video Soundtracker,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.368600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.912137Z digest=sha256:e5e71e6657f799b33250dd5fb124acc399b5d174b221fcd53b3b7f9beaedc33e

Observation 20b4b726-aaae-4943-aaff-ce82af8fc7c4 · outbound

This paper cites Simple and control- lable music generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Simple and control- lable music generation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.195413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:03.992590Z digest=sha256:bff1091e462450e898ecb06aed3e218157218573f0da94380bc60f2865340b44

Observation 8f0bba4b-2550-4491-86ea-2112d40c1c49 · outbound

This paper cites Content-based Controls For Music Large Language Modeling.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Content-based Controls For Music Large Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.068660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.068660Z digest=sha256:93c13afb23c335905551102a2284e19e0aa0d071b05a147d26cf1c326ee4dfb1

Observation 867e4e4d-3533-4f57-9782-533fc52b2236 · outbound

This paper cites BiMediX: Bilingual Medical Mixture of Experts LLM.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation BiMediX: Bilingual Medical Mixture of Experts LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.143427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.143427Z digest=sha256:43ddfe065421030111993a4cf8f617f25d8cbf4e8de08c6c7e72d486579b926c

Observation 95569952-453b-4786-a063-b217798b55b4 · outbound

This paper cites Audioclip: Extending clip to image, text and audio,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audioclip: Extending clip to image, text and audio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.013911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.196472Z digest=sha256:b6ecb489da6d7f0ae59ef60fa879ff07ff918fdd9d36c8dbe429e65863046bd6

Observation 660763c1-7f85-49a8-abe1-d7d88733bbf8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.281404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.281404Z digest=sha256:0117249a41a2630ae3f481d1446b62cce591d76c7a9bed5e905434881500b4bd

Observation 1a77d445-28de-4259-85b5-7470c0e565af · outbound

This paper cites Vidmuse: A simple video-to- music generation framework with long-short-term mod- eling,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vidmuse: A simple video-to- music generation framework with long-short-term mod- eling,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.832849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.379020Z digest=sha256:a7790dcca5a85a2f112a91d942a7998c336546d96c3f472ef9e2f7f5f54a63b9

Observation b3655a34-319f-49d9-8e80-239d6fc3e059 · outbound

This paper cites Sonify Your Face: Facial Expressions for Sound Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Sonify Your Face: Facial Expressions for Sound Generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.626318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.446437Z digest=sha256:f75a139f3f87cb72a890ca327831599faa3a6db5b3f66cfd81d9a6f917354a28

Observation 213f40a7-211f-4da6-aeae-8ee1e726f045 · outbound

This paper cites Movement to emotions to music: using whole body emotional expression as an interaction for electronic music genera- tion,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Movement to emotions to music: using whole body emotional expression as an interaction for electronic music genera- tion,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.373609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.532491Z digest=sha256:dd316cf20818c7f0b2e1d7ade161d0e0e4a67f2d46359c825a3b0999c56350bc

Observation 5a30467f-308d-41b2-b8e5-4398e4ffedfb · outbound

This paper cites D2MNet for music generation jointly driven by facial expressions and dance movements,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation D2MNet for music generation jointly driven by facial expressions and dance movements,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.145199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.603651Z digest=sha256:29801fa3c8666936bbbf369a92ef3b7231ad0f828876e84c1defa355218fd61c

Observation 8beac812-bcff-45fc-9d2a-dfb7b735dc7c · outbound

This paper cites Deep- Tunes: Music Generation based on Facial Emotions using Deep Learning,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Deep- Tunes: Music Generation based on Facial Emotions using Deep Learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.946722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.683328Z digest=sha256:6278f5f742e77a4fd0f18fec94f085d06322975e16607ecdaf26028222d1b0f4

Observation 95e5e6ef-e7cd-4911-b0e4-dd7226c2d195 · outbound

This paper cites A Contin- uous Emotional Music Generation System Based on Facial Expressions,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation A Contin- uous Emotional Music Generation System Based on Facial Expressions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.797652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.816135Z digest=sha256:04c24422d87daddcce5ed54ba65b7c967cb49479bdc7c585fbf1297027b9740f

Observation 290f3bdf-e9f3-492b-8641-be59b2b62b56 · outbound

This paper cites Musiclm: Generating music from text,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Musiclm: Generating music from text,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.655800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:04.908406Z digest=sha256:5d450a921d41bc73d68fa34b0a9173fafe39b76f3727794af6e830ec59e02978

Observation 437cbddc-94b7-4e31-b553-70831a24277a · outbound

This paper cites Audio set: An on- tology and human-labeled dataset for audio events,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audio set: An on- tology and human-labeled dataset for audio events,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.502002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.007593Z digest=sha256:8531d93df837610ff9f1122791f14ae96649c2b7bf8a819a8847bca8c8bbe28e

Observation 8d4a48bf-9cba-4a3e-a001-b3488749bedb · outbound

This paper cites Vg- gsound: A large-scale audio-visual dataset,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vg- gsound: A large-scale audio-visual dataset,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.346892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.092253Z digest=sha256:59200d14fe8161316300592a17867fe38d098654c9583fd8e061ca92fe195ca0

Observation 22542eab-239f-45f9-a1d1-3db28816f211 · outbound

This paper cites Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.182866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.158914Z digest=sha256:a6cce1ca78e7f25464712d70ee8f3a688a1c29f35f7e94d7f0585ad1643341fd

Observation c29f1b56-bd19-4960-bb1b-e28fac85f60b · outbound

This paper cites Marlin: Masked autoencoder for facial video representation learning,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Marlin: Masked autoencoder for facial video representation learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.043876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.230502Z digest=sha256:10b5a7333251997f60384b8f831af128305459ff07db284df63c0930ede70d31

Observation 83659852-de37-4dcc-a2fc-0941b0e9eaea · outbound

This paper cites Synch- former: Efficient synchronization from sparse cues,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Synch- former: Efficient synchronization from sparse cues,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.910894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.317236Z digest=sha256:7325da7a903a6f7d25e149623335759adcd5b60984a49d02a56d85780b33e649

Observation 8ee1745e-fb97-48c2-8d7c-8b64bbcdc768 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Raft: Recurrent all-pairs field transforms for optical flow,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.790842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.387260Z digest=sha256:2ef0b8f26378949d3d5826516080c8aad143194b204b1120bd9bcafc3663c1b2

Observation 30df38f1-ba3e-4b42-9166-3b4bc2c8f2e6 · outbound

This paper cites Salmonn: Towards generic hear- ing abilities for large language models,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Salmonn: Towards generic hear- ing abilities for large language models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.638745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.475020Z digest=sha256:366415a0e9423897622e71af6b3f770a2dcd0bc6678ec7c966bf9b2af7cc4e6c

Observation a906a825-ceb7-4767-a58c-469be428b1c0 · outbound

This paper cites High Fidelity Neural Audio Compression.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation High Fidelity Neural Audio Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:05.562715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:05.562715Z digest=sha256:10ca9380564b157089f2b9f873fbffd816335dbd00645c38a3c5ca891c2760a6

Observation 3570c2dd-fc12-4839-94ba-a278eac465da · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.478892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.646574Z digest=sha256:5644a2862f479f361ce42b3f2970091d7912f28a0d44849f6d958c360ab94083

Observation 4d498afb-aabf-454d-b1f7-7cb6a779273d · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:06.423278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:06.423278Z digest=sha256:7155bddf7a6a1f46ad07d18080a07950d3d94da9a4319822c5ec05a5a11f5a31

Observation 2109cd34-ce37-4fb4-ab09-158242398ce3 · outbound

This paper cites Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.324059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.802192Z digest=sha256:67369ff80a340f5dbfcf98f41cf39a5b41dd7aab15c9e2a9a299b0d3e92a16f6

Observation b3b0825d-f5ef-4181-bdd0-f6dda41b6f6c · outbound

This paper cites Cnn architectures for large-scale audio classification,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Cnn architectures for large-scale audio classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.184873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:05.915584Z digest=sha256:61577468dc05b990f7f95936238b1127a8634e93d6ed66114b271fdf58428015

Observation 9bc0bd82-ac09-487d-ae91-6404e14bdf0a · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:06.008469Z digest=sha256:5938d7ed1fffefd2b971669dbe9b3d91e6dc20167f291fdd0fb023785e02d13c

Observation 4ae24de1-8228-4f7d-ad82-b95dbfabf1d2 · outbound

This paper cites Efficient training of audio transform- ers with patchout,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Efficient training of audio transform- ers with patchout,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.890135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:06.095672Z digest=sha256:eac8cd958f2f61b1db3daa961d1efe2065802010ff130e6257833e47ecab8c3a

Observation a7efb45e-843f-4d91-96e9-5a8b69baac36 · outbound

This paper cites Improved techniques for training gans,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Improved techniques for training gans,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.747301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:06.183184Z digest=sha256:dc760d84757e81b3abb382c4156a70057c9e20917fce56796537c681b01e922c

Observation ddce1c67-46ab-452f-9717-4bbfefbf307d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.534298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:06.270020Z digest=sha256:14a0e018e382aec000d6092b42b64d2ead6ee3f93e50ce3faab174c01153adec

Observation f793ccb1-7e6e-4408-82c8-6d38868b0d67 · outbound

This paper cites Available: https://arxiv.org/abs/2211.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: https://arxiv.org/abs/2211

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.382378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:06.341162Z digest=sha256:4fefb4b72890a19a01800cbb0c9ec1ddcae289e4440deaa548e5229f4198a618

Observation 7a790e88-2b6f-4552-95e8-4e401d7c3a91 · outbound

This paper cites Available: http://dx.doi.org/10.1016/j.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: http://dx.doi.org/10.1016/j

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:40:05.719311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:05.719311Z digest=sha256:ce65a3af4975358c907f8c754b71e5a0bb802d6d6c2f533c1a1fa56b29fb59cc

Pith citing papers

Observation 9e40e093-95c3-427f-b276-9e5220945e43 · inbound

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation cites this paper.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.236634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T19:40:02.091028Z digest=sha256:81e6fdd64c16bc924cfe13f301c9f3cf586d41b9eca6ef1413bfa7f8ad808259