Pith. sign in

Paper Citation Record · LEDGER

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.09529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09529 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.381968Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5499e55c-4442-485c-a0ce-bd274fcb657b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.106949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.106949Z digest=sha256:efd9ef748b5a11125b6e659b145806a33d3406ac4620e42533df565460003373

Observation 7e1016fa-d14f-4487-a833-e02b546b800a · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.117045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.117045Z digest=sha256:fd4baa5c870c4e68ef4cfa064e95f49207a02975c47de957ce2681b2cea8f8b1

Observation 54d3592b-0ec5-49dc-8047-bca52ce43250 · outbound

This paper cites Qwen3-VL Technical Report.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.121620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.121620Z digest=sha256:e9b229144c8353b57088a0ef4cabbc5da41f47d12c021b5d8ac8514fbf8ed84a

Observation 77c8da77-0574-4eaa-bfe3-c6507eb04917 · outbound

This paper cites SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.126404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.126404Z digest=sha256:442df1913b5924e3b2c72ce684e77539d28dbad7a717d00ec0fd370916d2816a

Observation 1f347c07-c602-4f2c-9057-c1e130551bbe · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.403358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.131345Z digest=sha256:390b46aa88d2c3fd8f0fd0e6416b39c5619fe7db50404b3da0d02af94b1b8462

Observation c7f64e7b-04f9-4c96-b155-f6ad8df9f481 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.391346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.136526Z digest=sha256:2395d0abd3d113a0e22e5c667e6458bfab513976c0412204d1c7a844d3388d90

Observation 7d258208-eb36-4d96-9d87-f42ae61d879b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.141840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.141840Z digest=sha256:5d042f270823bf68584c6a855bb8d6a14cb29de440680416cba553c40c3e7d11

Observation 9b954811-bcdf-4f68-98ca-badcb11de8cf · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.147229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.147229Z digest=sha256:243666639f527c6ca60fbaf19a04026d2bf7c99a767a1634f7c3839620406185

Observation 81f7bfc2-eef1-4023-80a4-976ee2762f86 · outbound

This paper cites FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.152062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.152062Z digest=sha256:e277a5fa9aa66f45a0817495ba3052ea826707c0473e8a21ee780fcf83ef505d

Observation 81131255-de76-44bc-adcd-a2201c8416e7 · outbound

This paper cites Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.156810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.156810Z digest=sha256:f1a1c6c410cf258e472a22bb2710d420a5ea4d5e374beb53fe907157fe08dcc9

Observation 407959c6-c6ae-48f5-8ce6-4c796e047570 · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Qwen2-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.161186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.161186Z digest=sha256:0e4b61f8e1bb0bf4acd8baae25d13cd0041219c6331dde34637fc28290a98d62

Observation 445908df-85c5-4b6e-8f1c-177512540198 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.165702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.165702Z digest=sha256:245a0686b7f8cd23c561528d3b28a23d995b8eb5238ae866be50df51c0d44da8

Observation 44a65de6-7316-4f2d-bf81-582b7ea5dc66 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.371299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.169825Z digest=sha256:96e4b3504bd053ca259bd612c54f0b778c0247f33464e374471ff60d777a281a

Observation 170ce7ee-8829-44a9-b2ec-0311c63bab55 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.174885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.174885Z digest=sha256:8163fd13ef99b9e352a50094724bb50fd99c74a7c9c2ee3d191a62ded22fd1aa

Observation 7ab50835-eb3a-444c-92f0-63f2e4bfdd6c · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.178845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.178845Z digest=sha256:cad37dd9de425b317d1b1275f2f194fab13028bf6cc16b61ee49865d256d8372

Observation fdfeec41-52cb-4c17-ab5e-f51637f8aae6 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.340192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.182233Z digest=sha256:2db085d67657d6f7de8065e3859c893a102fa7fcf22893710700a3b981e0443a

Observation 085e0dd4-a64a-4cbc-9805-7e14955564b2 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.185572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.185572Z digest=sha256:a06b3cf49fb81fb252e74ebe626fa8d09b175e10e6293249976877c041d94dd7

Observation b158ec18-97bb-4b2a-bbd7-8c095c677568 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.191958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.191958Z digest=sha256:d17492da170937925803553452bbe98a8adcc04f12b8575a7e353617b50111bb

Observation 51445caf-5a85-41d0-afb7-0efef22bf28f · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.195746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.195746Z digest=sha256:bf9fe170ee86c03e65671f286ba41abeec7b380223620cb047e342214f17a7b6

Observation 05a2c0ef-b839-4460-b0bb-414505746313 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.310410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.202252Z digest=sha256:fca7016ff780586ba88b7def2b9a4cc13564162f4b75a35ef10a807c36774349

Observation c0e26a85-064d-4ffb-bf78-a034169f0d44 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.296132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.223442Z digest=sha256:f17382e0c6306acb2eabcf3ee2c8badf1120a35455f2e98f7b3245d4dc68c9c5

Observation 47d65d10-c76a-4f9e-8db1-8db6158f11fc · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.283319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.227907Z digest=sha256:8be7e0d3611306bb2fe4661f7a40e46bee2991eeb483903d72d20e87f42d0ac6

Observation 0a176863-2ca2-4c67-8b38-1b3a937649ec · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.233176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.233176Z digest=sha256:3546ac7a49927f9d0d6d6f5fc6e33ae97beb42486f465385bba1905c72ff883e

Observation ed524262-779f-452a-aa52-7855b4c1be89 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.261879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.237280Z digest=sha256:0413323f4f94188a03706435bb77856155acd5eae598c134c9f69213dcc8005b

Observation 48d333c0-3f15-4fb8-93cb-f3b4b4fe15c9 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-11T15:41:34.422290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.241666Z digest=sha256:ca730dcbe8b73e7358891b3dc3006dd90b49ad742ce7ef4c5f3463707128785e

Observation 529399a4-06e1-420c-977a-77791d751c37 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.248327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.246374Z digest=sha256:28292d918e26494fa7204f2b1db8454a237ea843a4e51df43a5c0747bbe868fc

Observation 5881b4de-0b83-4946-9684-cdc6ca2e0973 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.251401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.251401Z digest=sha256:c9ef385f30771739db35b7b685927c6f23b7cfe84d052b5a28a309261619fa0b

Observation 0580e7e2-d925-4381-b93b-d9ef06abf2a6 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.256085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.256085Z digest=sha256:d26337eac1b7751dd7390e25ebffb68c1273753453fc9d316db00cf04b91b278

Observation 3114ed37-9df1-4fda-b656-cc91ef055d32 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.227361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.260375Z digest=sha256:480ef677da7ea584826dc1d2bb9399f2f932d3fc05fc256b8219d7e401a0bf1e

Observation c02742d9-3e0c-45fc-b9a8-24c9ffe28e52 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.264355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.264355Z digest=sha256:496f398d61a9ffae7ae8906bd0e7ee19d5779b71884fc3b90b54ac45e91f7388

Observation 996eb128-6663-429c-a08d-09ed058974a5 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.203076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.268402Z digest=sha256:47213aba2af4d31296d21c242422bd423c5ba0ba87ed76f62755cacbc170f964

Observation 285cf977-bf1b-4acb-b96e-aabff348d07e · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.272510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.272510Z digest=sha256:15db04950cef44516c686150b44d15a83a92e9d8d1b68bd5867bb985b6bb9550

Observation c544954d-47fa-487f-901c-17aba8c82d8d · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.276452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.276452Z digest=sha256:712a8791030f114ed350e144e7410f9099345f7b8094e492ae8659917df92ad0

Observation b5435a5d-54f9-42ea-a39f-2c32235422e8 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.280600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.280600Z digest=sha256:b41f5c5898368cbcb37d38c8ca97c26340c0f3bedf3093e35ec333c7485403a3

Observation 8b61f7b0-894b-45bb-817c-5084ad65406b · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.174848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.284566Z digest=sha256:c790a49eff2195720dd455781015cfcc9f86bef79fc230e6275782f496b10184

Observation b94af2fb-3928-47b3-abdf-9f2d7a1a054e · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.288493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.288493Z digest=sha256:51255ff3f60c4abe4a7afceccd51f83eba6f069b825a95c7bd854caa3d00831f

Observation 8eaed4ba-c82c-49fc-9570-f147d7200303 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.152737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.292361Z digest=sha256:4f1ac484632a5f6d734d34b85811dea5fdb76c3a2c4dac903378c208951ae6d9

Observation be5bf654-8cbe-4ad3-9153-98eccb2f164f · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.297516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.297516Z digest=sha256:8380410e836bea51f3d74d0f6f7a38265b9fd5c78698dafe133262902ae324df

Observation 49af4d98-5205-44f9-89f9-27b18638e561 · outbound

This paper cites InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.302159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.302159Z digest=sha256:5bcfe3ce9b9dc41ce0b4bae7bad1f34a900b5c09d4e014ae25916da690d34b61

Observation 72ffc7f8-406c-4ebc-8961-f11d36e33b6d · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.307150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.307150Z digest=sha256:5504360c447c718d0e1346b5b010b3d0ea7a00c17558d4f3db7085c37ff7b3cb

Observation b3ef91ca-ce95-4f76-b966-c5dd5866864d · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.137389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.311858Z digest=sha256:50af254318452f32d2230310c0692d660b5d451a3be1bcecad46396a94985497

Observation ac2da736-ac44-45da-bf7c-b87ffbb70170 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.102644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.321229Z digest=sha256:0ec726b75e598c1a8365f29071c90b8a1055e4d7eb126ef442be14e9ca9f7766

Observation 06941a77-2e5c-418b-8df4-ee7edef33f36 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.325382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.325382Z digest=sha256:c6e857dfe38c7ecabc3389ac6812f05cba1a023a665cab383c27ae318b83ba6c

Observation 445351ab-91db-4c6e-b5e2-ffdb4e3c8056 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.076228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.329856Z digest=sha256:fd38becffc17ee196d29de522ade26f676a2030b7b8b1b148c3f85ddf704594f

Observation a761130c-9057-4301-afa8-f83b212722c5 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.334245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.334245Z digest=sha256:7ef6de3a5fd9b82ef3c92029a96efcd19cd915d6e840a9a355346f00241fad30

Observation 8fe144ba-39ec-4501-b027-5122f5c9f762 · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.050662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.338582Z digest=sha256:ac3fff909eceecdc1f8759ef546556db170a0e6238ed1f97093ffb7d48b0eef8

Observation c0db9894-a99d-4c82-8c84-6694902529e0 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.343280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.343280Z digest=sha256:da526c797b2d71b3d4fb27eb3a0de0dcbc57618675b17bf30ce76827381ed485

Observation aef01029-75b3-4962-9c29-7563eb6c0a20 · outbound

This paper cites AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.347508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.347508Z digest=sha256:c7f812e8e68045beabf4cda8f67205085af58080e7d308f581fa394da3c5c06f

Observation e1a51c55-39ff-4bbf-bb94-bdbf3ca4cc5a · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.352108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.352108Z digest=sha256:76e724b932328293a430e493af3398eff2f77ec4c520b737fadd6ba0ac039ee2

Observation 1e3c2a70-445c-49fa-a886-e9efa78515e5 · outbound

This paper cites Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.355920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.355920Z digest=sha256:f779f0fbecbf2867d5164f62e9f451f4ee6180844168dfdde91c02add8b3086a

Observation 6173af98-3b2e-4a96-9062-41097d0cc30c · outbound

This paper cites FoleySpace: Vision-Aligned Binaural Spatial Audio Generation.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:41:34.554124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.359873Z digest=sha256:735c71c54131e60635405252563820ceb9f4aaf67bb092668d70f19cad95f739

Observation 33e7086b-aa82-4375-83bf-94946de8c5eb · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:41:35.033078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.363774Z digest=sha256:146c52bc9d5aa61febd2f4f97a8a1bdec1de16e79001be9e73d2de425285415d

Observation 8ceabb6b-f3d7-4e72-bc19-07aedac7bb3a · outbound

This paper cites an unresolved cited work.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Unresolved cited work

Reference 53

Resolution
verified exact
raw_fallback, observed 2026-08-11T15:41:34.532298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.367982Z digest=sha256:2dce0c6f1564e1090c3dac4fad0a3c7932382ff0d60d74266c01236e4ff032c3

Observation 8ac456d1-b78b-402b-b432-68c7559eafbe · outbound

This paper cites ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.372388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.372388Z digest=sha256:fa86a9584e00f1a0012ef4d0e179109c2333277c825a02bcb01b878e39e389e9

Observation d44e608b-7bd0-49d1-8b02-933e5ccad5ec · outbound

This paper cites silent frames.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework silent frames

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:35.015532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.377167Z digest=sha256:7fb2acab3d07abcb9c2707c25c03e795985276f11505a8614d0b04421228820b

Observation 9bd2df92-e91e-4ca8-aa53-5fa0a9f34f8e · outbound

This paper cites the sound of a vehicle driving.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework the sound of a vehicle driving

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:34.992918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.381968Z digest=sha256:581125fd867397d1e53f9b635f0945c01c5a28e5aed5a60d2fc0d34c3457ddfc

Observation f7eef15a-5a49-4d27-a3b0-31cd6b8af678 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:35.119557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:41:34.316499Z digest=sha256:15abfe27ff9a126d48a896362fee70b173b0eaf0f52458989ac0383377a35a94

Observation dbdff0f5-4c34-4ab9-b1ad-5ccd20273336 · outbound

This paper cites In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.112150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.112150Z digest=sha256:26a57a663cf24ffeb5123bec67c616b22a88588584ebf437ab61b01b26d3782d

Pith citing papers

No inbound Pith citation observations are available.