Pith. sign in

Paper Citation Record · LEDGER

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

As of 18 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.11913.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11913 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:27:56.984255Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e7819bb-46a5-4095-844b-9aa85c87df2c · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion AudioGen: Textually Guided Audio Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.712678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.712678Z digest=sha256:8ffecf4c880a21f5aa567c3b2b74b5a664ce1b50618c7a0bed723ac53b6df3a4

Observation 091b67d6-4056-4fb5-b0de-355b858815dc · outbound

This paper cites In: International Conference on Machine Learning, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: International Conference on Machine Learning, pp

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.913775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.718048Z digest=sha256:7fed1ccbabee2c30adbad1de2ceebfdd0ac4319ffd9d13e040480fa810659468

Observation 5f8384a3-22bf-42f0-bfa5-ca7b7925f6e5 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.722383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.722383Z digest=sha256:007b6475416f2e52bdc61885c8abac344261864b2a260dbe579251f8cdcf3f1e

Observation 732a2beb-af5e-4de3-ae60-8c4583503fb1 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.900236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.726676Z digest=sha256:79fa8ccfdd9ce3de962d3c89927a58728f3dd4e25a880abc67ad9ce77f8db9c6

Observation f2814413-27d7-4f0f-ad66-1b8ce175a5f4 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.887526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.731597Z digest=sha256:5d0a0db9b4baf826b6e01e9f43bc1a01a4a631fab4d1a73c42c2c0b1a8905f96

Observation 52ae80db-ecd6-4720-bd3e-77e755cb7060 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.735765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.735765Z digest=sha256:bbf1e127e38fb8ec5a4afc545db498b7637f513d6bd9cdb558627b994559001d

Observation 8252999e-aeb5-4a3c-bbc1-d58235d76c56 · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36(2024)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.873825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.740329Z digest=sha256:51454d60ec52b51666a76c5234f3388dac15cd5c2b3b57b982add40a185fe59c

Observation c7850795-4826-465d-92fa-e522197828a4 · outbound

This paper cites In: 25 Proceedings of the AAAI Conference on Artificial Intelligence, vol.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: 25 Proceedings of the AAAI Conference on Artificial Intelligence, vol

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.858771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.744066Z digest=sha256:5f045a974527c93aeb997c092114850fc6a5e1692a74e7f299e786ff6062bc18

Observation d56a2bbe-43c4-4577-8b00-6f5b14d534d2 · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.747595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.747595Z digest=sha256:d242434ec3becba082969a9f8db8d205ab63e467e673bf4778fc21f142c009ab

Observation 8c6298df-201d-4fae-a2c8-03a640fb03c0 · outbound

This paper cites Jukebox: A Generative Model for Music.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Jukebox: A Generative Model for Music

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.752114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.752114Z digest=sha256:c84506bddb18de52caba94de03a4a01395f199614988e30dc4eaf55870b2c37b

Observation a87b5197-e653-482e-b3a6-0de78d8b2111 · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36(2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.844657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.756554Z digest=sha256:c5573ab45355848c93f21c00c96af39e2594ead328abbbfce5f01929f28b8f60

Observation 375c3569-f328-43d4-824d-aace6b6e536e · outbound

This paper cites Advances in neural information processing systems36, 14005–14034 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems36, 14005–14034 (2023)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.830642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.760900Z digest=sha256:dc577cc7ee18a48c461231f84b8829b9dad2151e5b4af08fce477300b31f4db3

Observation 99b3b07d-b06a-4258-8fc5-763fce20dff8 · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.764846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.764846Z digest=sha256:b0b679f75ffb0dd5c8105d88d90359a8b1db27c20e8f8079b01a93afcdbc2057

Observation 6323bf76-1628-4056-9d74-21c204b9d7ef · outbound

This paper cites In: NAACL-HLT (Findings) (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: NAACL-HLT (Findings) (2024)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.816176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.769736Z digest=sha256:fb0a17e2775fd498705fed55ba3bda503a5cbb9761f43cbe733ab472c8f5dd42

Observation cd8c9d4c-7258-4381-913b-ba42fc6b457d · outbound

This paper cites V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.773970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.773970Z digest=sha256:f5e0efa537f7f855242c1252571f527a089024e779d333d14a16b67cac3ae912

Observation 81112b9e-3a70-40e2-b223-891ac4bcf10d · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing31, 1720–1733 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion IEEE/ACM Transactions on Audio, Speech, and Language Processing31, 1720–1733 (2023)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.801617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.778338Z digest=sha256:702f9cb8514e2b0a3f1406e075297f1dc270512b472423d5a99b7282d6563126

Observation 4a26f552-1407-462e-b59e-f41d73cb198d · outbound

This paper cites In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.786691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.782497Z digest=sha256:dd0b119014efe1b5c345492433ef94a084cdf120de9f07331217032529c19d84

Observation a5a43a36-4f0f-4c93-b81a-e3bc9e2bf2b8 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion IEEE/ACM Transactions on Audio, Speech, and Language Processing (2024)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.772134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.786388Z digest=sha256:39a21501671e62763f9c3a0922b4523a5a6d560f888753b2be8afc8d7f213b35

Observation b15cb50b-ba18-4504-adfd-2d9bd3e953b6 · outbound

This paper cites In: Forty-first International Conference on Machine Learning (2024) 26.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Forty-first International Conference on Machine Learning (2024) 26

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.757133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.790367Z digest=sha256:8f33fc75624f1cba3f99e60a8244ffc9a305bb7c4b40d782a1f003b722eb89d6

Observation d6e21cea-612a-485f-ab37-1bb8f76c625b · outbound

This paper cites In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.742362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.794702Z digest=sha256:7eb28ba4c993005d331dc9ba73635c2cd03bf13bdfa3ccfacd83982e2d4b3f79

Observation 65c58431-bd3e-4ba4-8909-31fb2fe49cbe · outbound

This paper cites In: British Machine Vision Conference (BMVC) (2021).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: British Machine Vision Conference (BMVC) (2021)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.727619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.799087Z digest=sha256:420609cd8f5dffeef51569687c2661f209aaf3a5874dd3a0168804f421f3cf47

Observation 80fbd0f6-4583-4055-90ef-a5eae7d5d9d6 · outbound

This paper cites In: European Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: European Conference on Computer Vision, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.712813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.803267Z digest=sha256:ae6eef07d18c0c97e436c880cca6748cead1836a9123e2fb131befc5e8933dad

Observation f0a71da5-dc07-4b08-8159-5abc0f61ece8 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.807382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.807382Z digest=sha256:a308d775fbbc9cf42b8e6346b4b41eb2c9a52a48f5080cff9a4f7511f3b7e524

Observation bee83027-0efa-4327-a425-230eaa6c2862 · outbound

This paper cites Advances in Neural Information Processing Systems37, 128118–128138 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems37, 128118–128138 (2024)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.812139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.812139Z digest=sha256:270f5b85101b028719b90174372fe89eb653ac1b31138f2cfe180793a58a1c4e

Observation 10418fc3-bb32-4f0a-9106-8eb9122011e9 · outbound

This paper cites In: European Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: European Conference on Computer Vision, pp

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.688460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.816312Z digest=sha256:d854e02fc0965fa8d72588329af933f8f86ce155d5cdee0ffb03d9fd96d62c96

Observation 463212b0-d45c-4c69-91d8-1a837e07dc4a · outbound

This paper cites In: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.674109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.820327Z digest=sha256:43709c4bf5e0d557bebd9be74fc32fdbdf1229cac4d6b40b735c156f42cde630

Observation 293e2e7d-31bd-492c-a703-3dc06404301a · outbound

This paper cites Expert Systems with Applications249, 123640 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Expert Systems with Applications249, 123640 (2024)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.660151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.824643Z digest=sha256:5ad88796d91e94435fe4a27ef0fb0e8f6287e50a31dcf1ac4aebd60d0a407967

Observation 57fd110b-6e5c-4d40-a358-8c2ba84d2b09 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the Computer Vision and Pattern Recognition Conference, pp

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.646119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.829173Z digest=sha256:b41b401e8b0ed495c1fcdbf88f7713b3eb0da744830c0e7d2bd33765b1c1987c

Observation 1cd0ce97-2a41-4f14-b4d2-d1623e23a208 · outbound

This paper cites In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.630402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.833475Z digest=sha256:4727199370128ea1acbe164be23a54699c6b47d92a27115400934edef786a469

Observation b7d643b2-c701-4277-adfb-3cbfe477e85a · outbound

This paper cites In: Proceedings of the 32nd ACM International Conference on Multimedia, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the 32nd ACM International Conference on Multimedia, pp

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.616375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.837480Z digest=sha256:42f620409743463aadb8ae6c83da6ee7f0bdb4d08a0cb5d267ee020c5c663a14

Observation f07b40b7-08f8-4344-aebf-3158027ea702 · outbound

This paper cites In: Pro- ceedings of the AAAI Conference on Artificial Intelligence, vol.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Pro- ceedings of the AAAI Conference on Artificial Intelligence, vol

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.601398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.841715Z digest=sha256:8b83fede414fe1f00f8b0ff91efc7f72ef4ffa513f50741da630d0ee89896fcb

Observation 57d58cc0-26c3-4354-937f-d7c441c7c267 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence, vol.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the AAAI Conference on Artificial Intelligence, vol

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.587214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.845812Z digest=sha256:c0a2795ecdfbacee72f613a47c7571633979405ff09f1e8dd7c7e2678b1af9b9

Observation a1f59108-06d7-4902-80fb-7dab094bf2af · outbound

This paper cites In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.573885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.850012Z digest=sha256:a561faf6331fe3b1012b8cef93fed84faffdc52d9a059ddc450821522c1ccd65

Observation 7dbec8da-ac2c-4f17-9efa-eb93f6857390 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.559378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.854403Z digest=sha256:eceb097a56aa41481f0162fad50a9110780cddcfd6c382ac13ba5bb3f323245a

Observation 06a8869e-afd8-497b-a31d-23ee28793e30 · outbound

This paper cites MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.858527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.858527Z digest=sha256:6fa784f2d823378816821e4cab089c76654cf29f91e2601d13d39c1c6f0740b3

Observation 6a83718f-d1e1-417e-9475-4ce5a1ed377c · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.863037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.863037Z digest=sha256:9bb01760108dc3439d1e3e907c95477b6261e15be79e0cf020cd48e409edd2a1

Observation 394b22ae-1db9-415b-90bc-547dc7b013a3 · outbound

This paper cites an unresolved cited work.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:27:57.542854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.867299Z digest=sha256:d8f4a36b25aadac194fe09fa9b09707e5c4ca0c38dbd15439695d8fd8ca5a085

Observation 7d27d25d-af74-4969-8c6c-92baafdd4d82 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.528185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.871072Z digest=sha256:eab4b29c661372cd8c7cd40c40854da4c1f628c878e1d137128fe2f9d9b335ec

Observation 509ca0ea-9b3f-4d54-bdc9-be1ffefaa7f0 · outbound

This paper cites Advances in neural information processing systems30(2017).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems30(2017)

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.874666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.874666Z digest=sha256:3d9f305a34fdbc009029e3422eef501fb607c8328ce0b028590f6308f3f439ea

Observation b497b9f1-69dc-41ee-8ec8-eb70a53a3c0a · outbound

This paper cites Advances in neural information processing systems33, 3008–3021 (2020).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems33, 3008–3021 (2020)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.504639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.878323Z digest=sha256:e4bb93abc7adfa52284588cb6f9e5e61adaef07678bb675ba167a700b5836810

Observation f88d134d-575b-42b4-ace0-cfe37d0431ee · outbound

This paper cites Advances in neural information processing systems35, 27730–27744 (2022).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems35, 27730–27744 (2022)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.881893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.881893Z digest=sha256:db2ba96cb0e2914741a7450a1c7c94029a1941348b6cfcd1cf03c6194e6b2363

Observation 627e4055-c6d2-452d-b298-755be48300a0 · outbound

This paper cites Advances in Neural Information Processing Systems36, 53728–53741 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36, 53728–53741 (2023)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.480657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.885400Z digest=sha256:6f1bc56e34c3f194cc8788ad6e7b0ffe10557a0dfb6e61e09d9a52c646161232

Observation 017b35da-32a1-4171-bc5a-c986da2b7c18 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.889364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.889364Z digest=sha256:e0fa5f5d656df2065b89e16c45d13cf793cdda34fdade3e514355c5673c64d96

Observation fa21302e-f917-4716-a010-b2d2a414ddef · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.893803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.893803Z digest=sha256:f81e99200216aed5738258586362040eb236686d59b8022b1f57b4dede5928d1

Observation c22c0a52-147b-49a1-94dd-c87ecf787fdb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Constitutional AI: Harmlessness from AI Feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.897942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.897942Z digest=sha256:680fcb657fc64fd2d269d7df1bb1b9658fd7a81ce2d03dfc2b305c8a160d1a33

Observation 1c5d24ae-298d-4fa4-9ae8-de5340b82c53 · outbound

This paper cites Advances in Neural Information Processing Systems36, 15903–15935 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36, 15903–15935 (2023)

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.902231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.902231Z digest=sha256:3e27127ae978e0bb650e5193d38a45b5b19e6e643714178dcbdbacf6e90a96ee

Observation c8a9da8e-d582-4f0f-8579-e218e17ecd70 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.453843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.906308Z digest=sha256:d615289844ba182296ed3b5dd733e8eeeb3da301d7f6a8d9becc425231f0492e

Observation 31449035-8014-4085-8416-fb6394ee92f0 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.910237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.910237Z digest=sha256:c52844d27a79378f54bcfd8960820cec3f651e8eb6f7d662865bd468f293478a

Observation b27ff933-b8de-4cb5-b04e-88ee0e7455fb · outbound

This paper cites Advances in neural information processing systems33, 6840–6851 (2020).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems33, 6840–6851 (2020)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.914138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.914138Z digest=sha256:017bb764b5f8c280f3031da8a42f6ed101b420370919f5fb036891c528e7b891

Observation c30dd17c-72f5-4c70-99ce-fb504e5ce0c1 · outbound

This paper cites In: Proceedings of the 32nd ACM International Conference on Multimedia, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the 32nd ACM International Conference on Multimedia, pp

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.421466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.918107Z digest=sha256:c2a62413c7f43d489ca26a95c556662c8140bd809f055a6b39db10683c30c2b2

Observation a3a0e9c7-dc9c-49e1-b335-482581e42b40 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.921927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.921927Z digest=sha256:27795222eb29e26edc20bc2ccb31d2e1bbb963b4d2734cd3b4dc1fed0962312b

Observation da1d74bb-9cb2-4016-be8b-872fd52240e2 · outbound

This paper cites Neurocomputing568, 127063 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Neurocomputing568, 127063 (2024)

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.927264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.927264Z digest=sha256:31dcc3bf4652c93606bb88eccfc5f8b3d46d2c8f8881f1564ec095e2ffc6085d

Observation 291410dc-31fd-44ac-8e07-8a66cb894816 · outbound

This paper cites Journal of Machine Learning Research25(70), 1–53 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Journal of Machine Learning Research25(70), 1–53 (2024)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.397921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.931337Z digest=sha256:8af10784e565de343fa1dfa15a70165f4e41ed891f838b01d0fc7268d42146a3

Observation 82135de5-b739-4df1-b12a-9d2478d93f9c · outbound

This paper cites Self-Taught Evaluators.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Self-Taught Evaluators

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.935166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.935166Z digest=sha256:eb450c66b7dc27b304e462ab276cdc8e8edbbe122ff6d35350e891b8f1acea8b

Observation 9a3c7be1-1c5c-46b4-a022-e8d7256a9cfa · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Contrastive Audio-Visual Masked Autoencoder

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.939307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.939307Z digest=sha256:96a7007bea394cdef144aa8e9385f71e54bfca7dbbbb7f3785094d0fa2ff59bb

Observation 39b4b44d-c61a-4f15-a3e5-7edb5a2953c5 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.943597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.943597Z digest=sha256:6b88ee46a89603228aa8da08bea4a8f0c3c6b8d496278ad2edae0c54e83ee514

Observation 3c9305fc-45c0-4fd9-915c-d21bf107b95d · outbound

This paper cites In: ICASSP 2022-2022 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2022-2022 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.383585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.947797Z digest=sha256:fc9710961021b07ca0b3b8ec92972ad510f49b5a794895efe1b84fd282f34159

Observation 2ef3e406-8604-47d4-8e3c-55c6c4367650 · outbound

This paper cites Audio-Synchronized Visual Animation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Audio-Synchronized Visual Animation

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:27:57.056509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.951652Z digest=sha256:556a7ea6de343ff82e6bd9a72345f9b3e86f95eaa70fdd00bac48d815c614aaa

Observation 1e167d94-415c-4193-8688-668b854a9410 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.955809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.955809Z digest=sha256:6711336f83443666a5f8e8d6b90fe1dd55e6516c93830bda0b4bb302c7b66556

Observation c67033e6-21a6-441d-a757-4f5aacc88296 · outbound

This paper cites Taming Visually Guided Sound Generation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Taming Visually Guided Sound Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.960058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.960058Z digest=sha256:146823bcf2c20368a7c32433b81cb8e878bc7cb1cf16d4661f33607d8a1a1e30

Observation 8e8120bc-b658-4ade-81b8-d6507c21881d · outbound

This paper cites Advances in neural information processing systems30(2017).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems30(2017)

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.964446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.964446Z digest=sha256:551da52b87c3b6d47bc254d4418fc29282ca8cc1eb41df1fd401f3562d3252f8

Observation 8e43efb9-e006-433e-a0df-087fe1b17fc8 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.968322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.968322Z digest=sha256:d8b3c3062812b31f3851da70ce992e99ff7666ab1d206fb30ce61c20da5a2849

Observation f85598bf-480d-45c5-bd31-b50260cd735e · outbound

This paper cites In: 2017 Ieee International Conference on Acoustics, Speech and Signal Processing (icassp), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: 2017 Ieee International Conference on Acoustics, Speech and Signal Processing (icassp), pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.352105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.972384Z digest=sha256:017e88fedc4758fa3ca678d0dc02e9dd9c95dc14fe25d36579bec3b336107f1f

Observation 75138c27-91a7-4047-8f9c-bab2aa7b1110 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.338428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.976488Z digest=sha256:74dbe79fa65a6d79c79b70fd273ce28b32a6211ac2b214d8c0ae889e01de1a5a

Observation 4d3be700-3182-49c0-9d76-244937e34b80 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.324496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.980300Z digest=sha256:627c86c2dbc6db9feb6285ae901aef0b3f78d05d131ebd14d7d37d76ed1b53ce

Observation 33e1ab7a-c8cb-4dfa-af12-52064fc4aeb7 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.310492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:27:56.984255Z digest=sha256:3ca89ab2239f465c9caf75bd9bc6df21f0702b3b249daafcb8528af83ddf3b79

Pith citing papers

No inbound Pith citation observations are available.