Pith. sign in

Paper Citation Record · LEDGER

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2509.04957.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04957 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:50:17.733574Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f842bb75-a603-4117-9da5-2b066087a606 · outbound

This paper cites Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:50:18.135702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.519723Z digest=sha256:9c39aea418ef7f7128ee184d48385c7289f9be0af057b1c1cc6d9a4105ff314e

Observation d92612f5-0f85-4dc3-8993-f96e18ac2a2f · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.606540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.524856Z digest=sha256:c699b34f3994dd2995ec22e5e942a1c255dc67a656b25b576e9b5fdbc96471d9

Observation 64b80587-6110-4140-a7e1-9e808d6a1dfe · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.591054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.529177Z digest=sha256:82482103e511f661574168763ccc94471a94a576a15b6e30c2870c6ba3de1b46

Observation b083d87c-44e8-4c1f-b798-83715b8af64d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.533640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.533640Z digest=sha256:329430bb770683dfb9d4f7b8c2b56e6d4ad76d4a6666501f35c67e3cf5f18598

Observation 0a4974b8-f164-4803-9450-e0320ed4036e · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.576690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.538704Z digest=sha256:7f077e85a3d50e40a79b739a620c3bf54dc453bd35c1d451fe3f6a818b341c46

Observation 8f7afa3b-c735-43ec-b80b-1d9279f0be44 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Scaling Instruction-Finetuned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.543238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.543238Z digest=sha256:d64d715bfbe829ff72ca6facf0302886c4ed916f0b8d01a60f77db3c985ca32b

Observation 920e750c-87cb-44a9-90a8-487cc2e9d6d9 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.548517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.548517Z digest=sha256:4832daa46df518addaa5d00b2b7f87a86211c5a5d46613935cdeef41ea6ef083

Observation 2477458d-f157-461f-a77c-3dec81e27925 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.562436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.553217Z digest=sha256:ede8ce0dcc70d36200369ff021c8f6ec3c37a685875acaff6b74e3f643428d98

Observation 8b65355f-770c-42a0-9601-7638765eca5f · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.547989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.557627Z digest=sha256:adbda467e43f5656447d82c6cfbd0c4496772de1cf9998b311a827574b8ff813

Observation 01ee2aa1-7ee7-4fd0-95d3-481ff5b6f6b9 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.533267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.562046Z digest=sha256:4c1ad2db7617270b5998407d0f71e46501c990d8d6240bb8f8866a664cafef23

Observation ccad6bb4-a840-49e3-8c12-ce157375b9ed · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.566407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.566407Z digest=sha256:199ad4b0f1364275f1552235bf5aab700515fc0afd8a8a55b13bdd8703ced42e

Observation f2c81e68-5173-487d-a076-ae32126e42b7 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.570895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.570895Z digest=sha256:92b5b56121bad6e03a1340876ce603ff96710aa3244d389da8c813fdf08fe1b4

Observation 2daa0361-24c5-4086-8f2b-9dbb3aa95cba · outbound

This paper cites Classifier-Free Diffusion Guidance.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Classifier-Free Diffusion Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.575495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.575495Z digest=sha256:caebd9fde11fe1e6c52620d8c37ad4557500846fae5b754257bfe567677abe33

Observation adad0761-b7c1-48a1-868f-c62c50dd4d3e · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.499605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.580136Z digest=sha256:fe61402eaf383ec3658b39b30232f727bd9a88e7bf4c599b30268defd2ffadb6

Observation 99e9fa52-c8de-4799-bf42-bdf00c1033ba · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.484811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.584412Z digest=sha256:8af2f74e88a55f0edee4f51cc2e7f6066f150ff0cb499c7f5ec10bef3bffa967

Observation 18893818-3c1a-4f82-b771-ea52e6e531cf · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.469860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.588724Z digest=sha256:6367000d33066259eaad789a5ab46e166a1908d3340dfe69e6b2104730764191

Observation b7ccc4d3-bb00-44a2-a790-a168969d09b1 · outbound

This paper cites Kingma and Jimmy Ba.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Kingma and Jimmy Ba

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:50:18.456063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.593020Z digest=sha256:d07e03fd6a6b66282e2f075de470769be3e55624bdc1801c65b7f31fc2fe73d2

Observation 17df4267-1f6d-436d-96a8-c81db972266b · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.441960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.597490Z digest=sha256:d31a7181ea75452b4490cd7fc0c6d62579c081116a94d5f7b0ba1cd803c96749

Observation 1108955e-afe4-412a-af83-d5f6f9f2b367 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.426643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.601835Z digest=sha256:565ebd25281c247962530e0e8a315739dc8d5f1b96439a59c67da25c476fac2e

Observation f37833bb-2ede-4a41-a0f5-d21a26a422e8 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.606300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.606300Z digest=sha256:d70f95955df1f41531ec105aceb35e797f689c4677790be53c402130fbc1f5f5

Observation 1bf7075f-6193-42a6-bd20-d612f08a4319 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.411781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.610788Z digest=sha256:0cd5c9e7852ed16d10d992cb4bb9db64c5f05b40fb46d6e5f9c17dbe0fa4f7cb

Observation 975cc498-caf8-4376-b9ce-d025e0bd1d04 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.614967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.614967Z digest=sha256:2f487732a6e603ada8d3342e60bbdf403e9903a47022cc0409902f61aebdab13

Observation 66111a53-d5f5-48bc-9bb2-83037c25e61c · outbound

This paper cites Flow Matching for Generative Modeling.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Flow Matching for Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.619399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.619399Z digest=sha256:682ca08ce016438f35df5c3f2d5ad83b4562d5ee7d1b70e7e2dfea9cf3107f1d

Observation 0d32f6cb-e52f-4189-9917-074c376419be · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.397292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.623593Z digest=sha256:7be819101b808167e0f513cea673ee5c51a41c2473f879f8802c16a329bbd839

Observation d382af3b-d5a5-497c-b424-bca310888960 · outbound

This paper cites Plumbley.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Plumbley

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T05:50:17.945212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.627972Z digest=sha256:c49ad02d815a1cd53aeed8a86b24f128ecfb181590076b000b60f4b4ac3230f4

Observation 67980384-74a7-4b10-a69f-40e82ae3e346 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.383059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.632277Z digest=sha256:0b92583db24dabd4f9f265a96b909a79e9f407e8cf279129158915e6ce48d3bc

Observation 2004328e-cec0-4701-ad35-823601a41df6 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.368955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.636526Z digest=sha256:851008c150086f9acba843f237e4fc37e13e794cceb58181971bf81126ca229d

Observation 52cdf0cd-88c2-47ff-8f27-7cf897e9e091 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.354042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.640702Z digest=sha256:bf7d519a3018c588fd26f51f0aeadc0ea54199e69173380829caa818aefb7cd3

Observation 31785116-109a-4ac6-a46d-f4f50b0ebf38 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.339848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.645138Z digest=sha256:fc54c5050c49febe58f52ccdfc865d149eb8beccfa91a5345341ce061294b5eb

Observation e9fa5e75-9146-4be3-b3ce-2a4d7c1b470c · outbound

This paper cites FoleyGen: Visually-Guided Audio Generation.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper FoleyGen: Visually-Guided Audio Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.649506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.649506Z digest=sha256:3dd2aea30a3fb91f685e353ad95e3e91eacd3736cbe3c2dc04813ab518efcc36

Observation 319ee91f-0d9e-4a29-a079-c133b1f3be5e · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.325583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.654064Z digest=sha256:a1ddd8825c2153aad174bdaeaef7bfc1d15a171911a9033adb12399d26667353

Observation 4b766f44-b95b-4d07-87df-f4552eaff303 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Movie Gen: A Cast of Media Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.658718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.658718Z digest=sha256:58950cf5f2c6a96b1a3fe6b233ca7b11552a53fb73caa57c50c0361a387b711b

Observation 17ed6f0b-e814-4182-b78b-2bb016ad8d92 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.311338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.664093Z digest=sha256:087290498718eb3cf3f04a8cd4941eb94d9e101669e23efa476849eba6fc9915

Observation 24abc524-30d2-4a93-b5fd-f162697875f6 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.297367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.668943Z digest=sha256:b6cb9e2f1be028a02b6849683a1e9ff56fc0967abbc82c13beb5a6e75872a527

Observation b7494c12-2a84-418d-92fb-3603cd230f03 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.283544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.673895Z digest=sha256:7ab13abcb98ae62d50bcdcaa26b89a9304316acda7bf3f5b0daf16ae22e8e4a6

Observation 18a912f6-1e23-49fe-8f26-677563723251 · outbound

This paper cites STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.678651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.678651Z digest=sha256:a7ae614eccdbe79676b120c027bde172a7899ef4d50602bc19ff5084915690ae

Observation 7da1f194-15b0-4b70-b537-2d88fe5436de · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.268582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.684187Z digest=sha256:1533403b3dab3eefda56a4b37b3d6b3d6d73beaa10510d530c5eec3e60f0409c

Observation 104718a5-30aa-4099-98b6-1dcbbc970e99 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.254648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.688905Z digest=sha256:e1c87c4090ffe02819ca21dbc8ce901d7c2fd1ea9d2b178982d306c64362b295

Observation 99e60a90-e50d-4f49-8c2b-09ca4e1d6b7e · outbound

This paper cites Stilwell.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Stilwell

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:50:18.240302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.693279Z digest=sha256:9a1f601b6e97517b0f94662f331ce73758fce901b6aa2320afc91efab90111b6

Observation 105322e9-47c3-4fe6-a387-86a7f29ce25e · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:50:18.226263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.697683Z digest=sha256:e9a75b6bfb67de7138ab16f44f1b06081f554c36e30442a6c82c9f0d46af80c7

Observation 0fdf94d3-d15b-4f2c-90d6-0d5ded89b42b · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Temporally Aligned Audio for Video with Autoregression

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:50:17.811973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.701833Z digest=sha256:504b6a75c916b69d49a11c2e8fcaa70d9f6091a5f0b02fccd0780e024c86c0b1

Observation a01090d7-0673-45c0-9e59-9b88d94a7785 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.210766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.706257Z digest=sha256:99797b7019cf0b9779b2ff7839d01bfd0654ca8e458dbda9bf06dcaacb25c0c4

Observation 91c118c5-e7c3-42f4-b2af-1d7d9026b95d · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.195399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.710751Z digest=sha256:402d84bfd62a30359923f22ef57f9ba9570d5021f45daa545985c2da6163d3af

Observation 9719fa83-b505-4803-b709-146fe5f4a38a · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.180556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.715111Z digest=sha256:6861c76e18056ef969ec48efec4306b6af04d1586b691404842d4661ca4db5b7

Observation f64d7a23-799d-4b28-95d8-91d5f431c613 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.165361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.720146Z digest=sha256:d0703c5baa52ee6df54c40f7a02de2fb205a5a11caf952a2a169161a7b84edab

Observation 27c7109f-22b1-4a40-9825-173789221db5 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.150775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T05:50:17.724549Z digest=sha256:637bff524b318b587bda32e6f00ccea1d452b2f5fbdf96ab1246228acba451d3

Observation 74f1ff39-134b-41a1-b0fa-820359ac3e42 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.728841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.728841Z digest=sha256:66e7128f3003e9ab4dc6590a246af6d29d41677a5130573d3a515770b49cd6d7

Observation a6f882af-01b9-441a-9d80-225faee1280e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.733574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.733574Z digest=sha256:a006bdd2c8f9edcc2b341bb1978dbf3244906820bbf17d60850243ad359e1e8a

Pith citing papers

No inbound Pith citation observations are available.