Pith. sign in

Paper Citation Record · LEDGER

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2505.22045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22045 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:17.616572Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:14.606479Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:20:19.500789Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact5
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 990bc393-2e2b-40b4-a080-f8d8f0809d60 · outbound

This paper cites This task is applicable in ar- eas such as multimedia retrieval [2, 3], assistive technologies for the hearing-impaired [4, 5], and intelligent video analysis systems [6].

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning This task is applicable in ar- eas such as multimedia retrieval [2, 3], assistive technologies for the hearing-impaired [4, 5], and intelligent video analysis systems [6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:23.517742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:13.754361Z digest=sha256:cf3ee887e39c6784d6ff54e53de51d7fbf1d7e79ceb3ba4dd8afb6b79148b910

Observation c12670d2-e9e1-4c7c-a64b-51375d1c0e73 · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:22.382099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.270326Z digest=sha256:7fccc231100bee0a67d5ff84a96596308f23621f0c47f6ef8aa05979af9d5fba

Observation 90bc94fd-7a50-4db7-9b8f-084ecf204c51 · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:22.261447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.364911Z digest=sha256:b6da48a1412a1c20e4d284eb9db60a302819999f411785a834ad115e63faba16

Observation eb936e67-e883-4cf1-b057-e56dbc38e60f · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:22.548312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.135637Z digest=sha256:3e6663373cdf089280a64525360cee8075fbc9c69fde0f01b13d0073474639f1

Observation c793a1bb-f23e-46da-9739-496822bd745e · outbound

This paper cites Video accessibility enhancement for hearing-impaired users,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Video accessibility enhancement for hearing-impaired users,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:21.105960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.414743Z digest=sha256:3db7b3cad5d3a2a8533e37297bdfe9b9d69dfa5b059a2d21706fa8f76b0998c7

Observation e616a5a0-59cd-4e20-a11e-9eba3d580b2b · outbound

This paper cites NowYouSee Me: Context-Aware Automatic Audio Description.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning NowYouSee Me: Context-Aware Automatic Audio Description

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.566942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.541739Z digest=sha256:4a1245663505972a5eaa32b709f8cc4d5cbdb699bb33d883bd9cf9ba1f614f20

Observation adbdb357-3523-47fe-8530-11cd4fb8af36 · outbound

This paper cites Moreover, our system achieves an approximately 6x im- provement in inference speed compared to the baseline while still maintaining competitive accuracy.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Moreover, our system achieves an approximately 6x im- provement in inference speed compared to the baseline while still maintaining competitive accuracy

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:22.071681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.457576Z digest=sha256:c0d4f323d22e257dc560b411fe7424ef291df37c8d84bc5f0a1ae305a826e6f0

Observation 5b3a8d93-2fb8-4921-ad48-6933a4507c29 · outbound

This paper cites Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:20:19.650775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.606479Z digest=sha256:542ec271ded02fcd260c268c82d8b04caad677b0c272aba37263f8a3e3a4ca38

Observation 0ae373ed-7693-4ab2-976d-0b72c193b6f2 · outbound

This paper cites Experimental settings 3.1.1.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Experimental settings 3.1.1

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:20:19.374002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.730364Z digest=sha256:b050f793ae7d6ee40d09ab68db8aa100dbedace05269a845c7de0f616f83a9af

Observation 297b5d00-b720-40bf-9c6c-39e32e83f461 · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.860195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.866290Z digest=sha256:1596284a5f54b505e76b05ddd8a8af9e804e0244c6a60a8473f25aca822d02e6

Observation 7c8cc59d-b511-4993-8a1e-2c094d0b08db · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Automated audio captioning: An overview of recent progress and new challenges,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:21.613178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.990600Z digest=sha256:8bbd579fe6aa42e28938f83f1b22b04bb33c40920f0f426c1c3d968061c9775c

Observation bad65a29-c5d1-40ac-969d-888ac3c0dcfc · outbound

This paper cites Building on this, LA VCap.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Building on this, LA VCap

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:22.840359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:13.858742Z digest=sha256:f336bf0c018e0a27f544db8df0bb1d26dae63651dfdab1dfdd13f3980edfd0a3

Observation 2f3c8141-f216-4f2f-8659-2f6409d76d98 · outbound

This paper cites V ACT [11] proposes an Adap- tive Audio-Visual Attention method that integrates audio and visual information through confidence scores derived from the †Corresponding author.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning V ACT [11] proposes an Adap- tive Audio-Visual Attention method that integrates audio and visual information through confidence scores derived from the †Corresponding author

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:22.692869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:13.995623Z digest=sha256:e9b8e3de14c98fdd97f116ca02e06750e2ece84d492af7c5489cf918585eaa77

Observation f26e6adf-f26c-431c-86e0-c3a924c667db · outbound

This paper cites On Metric Learning for Audio-Text Cross-Modal Retrieval.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning On Metric Learning for Audio-Text Cross-Modal Retrieval

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.957344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.102806Z digest=sha256:cb054c89f16e1c35c9d281a15d98adadfa5f77ddbc2d251d6e10acd8eaed3ddc

Observation 9197c34b-162e-4298-a2b0-d28ae0669ef7 · outbound

This paper cites Separate What You Describe: Language-Queried Audio Source Separation.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Separate What You Describe: Language-Queried Audio Source Separation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.203731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.203731Z digest=sha256:8a43d54dfcc11543215a17d2f2d20eda3c883f63e0fef12e84e49459e05dcd3e

Observation 0432c5c9-bbfb-4bac-a482-b5624e261675 · outbound

This paper cites Dynamic captioning: video accessibility enhancement for hearing impair- ment,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Dynamic captioning: video accessibility enhancement for hearing impair- ment,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:21.362097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.338815Z digest=sha256:07d451410d881a90a8282691e1cab3a285f0bfb70144e576ce8e93f3a7f201e3

Observation 2e695f49-ffc9-48b4-9ca6-e94392156a0a · outbound

This paper cites Towards Diverse and Efficient Audio Captioning via Diffusion Models.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Towards Diverse and Efficient Audio Captioning via Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.647051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.647051Z digest=sha256:e6193a53289bc6ace964fdfbb2b56726021aa55120c0ee2038d4f53815b24c30

Observation b8e6e2e4-5745-481d-abaa-bceab90c8304 · outbound

This paper cites Audio Captioning Transformer.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Audio Captioning Transformer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.267155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.714734Z digest=sha256:742a181c307f7b41cf83d40924b099f9a023fa893f4700da12634642ab0e67e9

Observation 50ab83a6-bbc2-48ba-a227-606ea4b0012f · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Clotho: An audio cap- tioning dataset,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.853445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.853445Z digest=sha256:7b9348937217305a4d69bd85037bb79745285d1429b29f0de112fd09309eeca6

Observation 987c54b5-ee94-4442-89d2-239d455eff1a · outbound

This paper cites Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.006011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:15.968878Z digest=sha256:2e0acfeb6951e0c687dbccccdf0f82d2652d6e85c9c88b76ff24c1c407e2e31e

Observation 829b6f68-f844-45d7-90b6-2c9c107b8d66 · outbound

This paper cites Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.007613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.007613Z digest=sha256:438b36638058190467c7318e2820a5eb547b0e09fd67827464fbf0fd73eca8fd

Observation 8b570a96-c39e-45dc-82f4-d53d388ce3bd · outbound

This paper cites Avcap: Leveraging audio-visual fea- tures as text tokens for captioning,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Avcap: Leveraging audio-visual fea- tures as text tokens for captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.844151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:16.098663Z digest=sha256:f08108155239abe8840313cda11bfcbed179528ecf4c24eb3fcba70a2100871c

Observation 21c84687-edff-4251-a172-5b1d6519ca57 · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.225698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.225698Z digest=sha256:5abc34ed894cf2a9e08dbd400e01f4640871595f3991ee5b707c8d0d9e4edbe0

Observation db42c618-b50c-44eb-af32-601868dfd147 · outbound

This paper cites Attention is all you need,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Attention is all you need,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.407539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.407539Z digest=sha256:02e07f8d6d3e833a8b26a8a94c428159d83d572a7b05b9aebb1b4ba05b3f2f49

Observation 4b5530ac-5d7b-4943-9ce5-9a3b7aae252e · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Contrastive Audio-Visual Masked Autoencoder

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.505610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.505610Z digest=sha256:32692fded7292a41c761d13ccb689f5610254875752f2af54c37d98d8e7fbaf2

Observation c03b655d-b14e-4e42-ba09-491c1f051c61 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.676736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.676736Z digest=sha256:ce70fa69a47c42653de9aa65a0d8e9e262b805d7df5c4a9046d8cfe737273614

Observation 77c30047-e1f8-4dba-9d37-5b52cef44cf0 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Audiocaps: Generat- ing captions for audios in the wild,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.761687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.761687Z digest=sha256:d2a040ee4a72f07f74821834aa2271d090905f148d070817ec4e258b32ca4507

Observation 9c36b8df-cb5c-41bb-95a2-5e1fb082b908 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Audio set: An ontology and human-labeled dataset for audio events,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.839969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.839969Z digest=sha256:2a208600f48e5699058ccdfea5d0c235c5fb64f293207bda003e700713b224d2

Observation bec60f63-820b-4a4f-aea3-fb6c780ad572 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Bleu: a method for automatic evaluation of machine translation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.934236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.934236Z digest=sha256:4821257cb2f88254b747122b040b95948f2f2b3f27a392bd6f877c0430030612

Observation 5a72782a-6566-4ddb-884c-0e8661853b19 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.172488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.172488Z digest=sha256:81761e7d071a6b6566d6beb2a2c96f08b448ea0e555b2674908b5b46040b6c4d

Observation 685df9f8-685a-4bf1-9b69-598c1c271207 · outbound

This paper cites Rouge: A package for automatic evaluation of sum- maries,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Rouge: A package for automatic evaluation of sum- maries,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.246884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.246884Z digest=sha256:9862c78224dde6d8b51656151253be449818d5cce5c3c63baad48862296cbe13

Observation bf2ab7b3-ee15-491a-922a-ec8ebbf9222e · outbound

This paper cites Cider: Consensus-based image description evaluation,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Cider: Consensus-based image description evaluation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.685683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:17.364740Z digest=sha256:fae848987e87add1e736874cf52d5fc0e3ec2e9710cb64f39f82d414a973c914

Observation fbc936b4-d313-491e-9f1a-439a4feb4b53 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Spice: Semantic propositional image caption evaluation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.419907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:17.461337Z digest=sha256:56a6635ad4703aa6a4031e7ae49fd7cb3326d894a245e88a9f5cc33f6c7d387d

Observation 62b9ede9-917c-48ea-a188-2d2857d12da9 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Improved image captioning via policy gradient optimization of spider,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.172010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:17.534077Z digest=sha256:34ff6800e386960d60f5ed5d68ccad3f7f59be8d1629792d1775c664ce7e1c7c

Observation 32c864e5-d2ad-45be-96cf-e0ad68300752 · outbound

This paper cites Microsoft coco: Common objects in context,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Microsoft coco: Common objects in context,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.936515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:17.616572Z digest=sha256:c27f3d75fdbbb3788c2331cf95d39613e71ef5cddecbc30c7d3406fd8c0bc330

Pith citing papers

Observation 5b3a8d93-2fb8-4921-ad48-6933a4507c29 · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:20:19.650775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:20:14.606479Z digest=sha256:542ec271ded02fcd260c268c82d8b04caad677b0c272aba37263f8a3e3a4ca38