Pith. sign in

Paper Citation Record · LEDGER

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.22045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22045 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:17.616572Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.106889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:20:19.500789Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact5
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 990bc393-2e2b-40b4-a080-f8d8f0809d60 · outbound

This paper cites This task is applicable in ar- eas such as multimedia retrieval [2, 3], assistive technologies for the hearing-impaired [4, 5], and intelligent video analysis systems [6].

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning This task is applicable in ar- eas such as multimedia retrieval [2, 3], assistive technologies for the hearing-impaired [4, 5], and intelligent video analysis systems [6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:23.517742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:13.754361Z digest=sha256:685bb5392407bae5d12f53fdae2baccc5b02f0ab643ac2cbf4f1b6d227801739

Observation c12670d2-e9e1-4c7c-a64b-51375d1c0e73 · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:22.382099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.270326Z digest=sha256:1d7d3a73297e7f09df979fd6658e79042a4b5749ab0306a2128e6a787792da31

Observation 90bc94fd-7a50-4db7-9b8f-084ecf204c51 · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:22.261447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.364911Z digest=sha256:06a648e0b26e7b49d9c057382f6dd0d417087baf9f7b3356e29abb88dbbdfab6

Observation eb936e67-e883-4cf1-b057-e56dbc38e60f · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:22.548312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.135637Z digest=sha256:4bc688501b996c6da805c307d2fff6247b2cc4b4f56950223f8e4ea4b6333053

Observation c793a1bb-f23e-46da-9739-496822bd745e · outbound

This paper cites Video accessibility enhancement for hearing-impaired users,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Video accessibility enhancement for hearing-impaired users,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:21.105960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:15.414743Z digest=sha256:0691ce1ba111bcef3489e3989e752c57eab3f234df8ccf0da9b323db4a90fe5c

Observation e616a5a0-59cd-4e20-a11e-9eba3d580b2b · outbound

This paper cites NowYouSee Me: Context-Aware Automatic Audio Description.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning NowYouSee Me: Context-Aware Automatic Audio Description

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.566942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:15.541739Z digest=sha256:af171312ce83e5f08385ef60e490dfed260281e5934bd1258b1aa3cf15b8b9b0

Observation adbdb357-3523-47fe-8530-11cd4fb8af36 · outbound

This paper cites Moreover, our system achieves an approximately 6x im- provement in inference speed compared to the baseline while still maintaining competitive accuracy.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Moreover, our system achieves an approximately 6x im- provement in inference speed compared to the baseline while still maintaining competitive accuracy

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:22.071681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.457576Z digest=sha256:a1dc6210479ae6b64d993a8614a2680f35316955cca38faa8b3e11663532b5f1

Observation 5b3a8d93-2fb8-4921-ad48-6933a4507c29 · outbound

This paper cites Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:20:19.650775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.606479Z digest=sha256:ca1cf2b1873d8b6bca74947034707269877c634a870067f678245d667c9fcf69

Observation 0ae373ed-7693-4ab2-976d-0b72c193b6f2 · outbound

This paper cites Experimental settings 3.1.1.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Experimental settings 3.1.1

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:20:19.374002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.730364Z digest=sha256:a6d9d1197a6420d07257e009ddada7f1a7042dcc124eda87c1412f80e69ba581

Observation 297b5d00-b720-40bf-9c6c-39e32e83f461 · outbound

This paper cites an unresolved cited work.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:21.860195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.866290Z digest=sha256:cf4b5bba59560ec21db9f1935bf77841f9bf1bee55de4e92933c2bfaabc81a68

Observation 7c8cc59d-b511-4993-8a1e-2c094d0b08db · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Automated audio captioning: An overview of recent progress and new challenges,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:21.613178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.990600Z digest=sha256:85366a8ed395c974369fa2d78741e8358714e2a09bd3a17d279ec6eb959e70e6

Observation bad65a29-c5d1-40ac-969d-888ac3c0dcfc · outbound

This paper cites Building on this, LA VCap.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Building on this, LA VCap

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:22.840359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:13.858742Z digest=sha256:4186bfd90ff9c7c790cbb62124acfda23055bf2a6bb26ced055c2644007c98db

Observation 2f3c8141-f216-4f2f-8659-2f6409d76d98 · outbound

This paper cites V ACT [11] proposes an Adap- tive Audio-Visual Attention method that integrates audio and visual information through confidence scores derived from the †Corresponding author.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning V ACT [11] proposes an Adap- tive Audio-Visual Attention method that integrates audio and visual information through confidence scores derived from the †Corresponding author

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:22.692869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:13.995623Z digest=sha256:e2dd18b2ae1a8268b61f86535854206aab5bd69e612d8afc1c15f1b467fcd2ed

Observation f26e6adf-f26c-431c-86e0-c3a924c667db · outbound

This paper cites On Metric Learning for Audio-Text Cross-Modal Retrieval.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning On Metric Learning for Audio-Text Cross-Modal Retrieval

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.957344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:15.102806Z digest=sha256:90aa015884326769d8d86984b5c0b4896594811d2fc427d34af1a8f10012d8db

Observation 9197c34b-162e-4298-a2b0-d28ae0669ef7 · outbound

This paper cites Separate What You Describe: Language-Queried Audio Source Separation.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Separate What You Describe: Language-Queried Audio Source Separation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.203731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.203731Z digest=sha256:c6aa23303f9e6642df3b6c0004f78507731faca62b57172850a7cd6cbd734ff9

Observation 0432c5c9-bbfb-4bac-a482-b5624e261675 · outbound

This paper cites Dynamic captioning: video accessibility enhancement for hearing impair- ment,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Dynamic captioning: video accessibility enhancement for hearing impair- ment,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:21.362097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:15.338815Z digest=sha256:06f0521575d53e97e9a5bb765c7794fc84347e3f240fcfa66316f562aa3b2f6e

Observation 2e695f49-ffc9-48b4-9ca6-e94392156a0a · outbound

This paper cites Towards Diverse and Efficient Audio Captioning via Diffusion Models.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Towards Diverse and Efficient Audio Captioning via Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.647051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.647051Z digest=sha256:c5a325cb59ec1241b8d81ac4b18497fcbc3c94164a8a3057e7649caaa6b760e5

Observation b8e6e2e4-5745-481d-abaa-bceab90c8304 · outbound

This paper cites Audio Captioning Transformer.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Audio Captioning Transformer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.267155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:15.714734Z digest=sha256:21462f21d64518dc109d0dec8ee45003f9d80b33a0b9e4e6026071b3f854ebb7

Observation 50ab83a6-bbc2-48ba-a227-606ea4b0012f · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Clotho: An audio cap- tioning dataset,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.853445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.853445Z digest=sha256:13bc6d5e49d09a18fc7f37300009233f4f891753a68db83b959f1cc210e8fbad

Observation 987c54b5-ee94-4442-89d2-239d455eff1a · outbound

This paper cites Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.006011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:15.968878Z digest=sha256:91c9ee4005482fe8c05125907fab623347ca9fa3a1d9f45e6cea2e57cc5dfcb9

Observation 829b6f68-f844-45d7-90b6-2c9c107b8d66 · outbound

This paper cites Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.007613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.007613Z digest=sha256:5188d09ec5dd70a543f4548a0d808f08cd9f55e1f3f08bdbccf432eac77fa0fa

Observation 8b570a96-c39e-45dc-82f4-d53d388ce3bd · outbound

This paper cites Avcap: Leveraging audio-visual fea- tures as text tokens for captioning,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Avcap: Leveraging audio-visual fea- tures as text tokens for captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.844151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:16.098663Z digest=sha256:997243889329a8158bdead267ca0030f2586da927c95610de13da741a2761dd3

Observation 21c84687-edff-4251-a172-5b1d6519ca57 · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.225698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.225698Z digest=sha256:62e49404fa20735eb502d0ebcf726be60de3277b722639395f5c7985bf09085e

Observation db42c618-b50c-44eb-af32-601868dfd147 · outbound

This paper cites Attention is all you need,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Attention is all you need,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.407539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.407539Z digest=sha256:4fef689732d77a738384ce406fe0ceaf177aaee1b3fd26c565fb617c2205e646

Observation 4b5530ac-5d7b-4943-9ce5-9a3b7aae252e · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Contrastive Audio-Visual Masked Autoencoder

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.505610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.505610Z digest=sha256:bf9751f057133069567ea77f283599a6ab8e5d64ab133509c7afe603a6751f47

Observation c03b655d-b14e-4e42-ba09-491c1f051c61 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.676736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.676736Z digest=sha256:f9a04238bb19f572e12b82cd4becc3984cd74ce75412590a15dcc3bda902fc32

Observation 77c30047-e1f8-4dba-9d37-5b52cef44cf0 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Audiocaps: Generat- ing captions for audios in the wild,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.761687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.761687Z digest=sha256:6d8c1949e210fc0c227a2698941313ae2f4d22c17a3503d8ebe85abb3778c7c0

Observation 9c36b8df-cb5c-41bb-95a2-5e1fb082b908 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Audio set: An ontology and human-labeled dataset for audio events,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.839969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.839969Z digest=sha256:1e8e611b4e40baab9d18263aeb54b72a05bf8c20452c0bd486d3a6f00d77246e

Observation bec60f63-820b-4a4f-aea3-fb6c780ad572 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Bleu: a method for automatic evaluation of machine translation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.934236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.934236Z digest=sha256:6c0d0fb0f370f2bbb9ad7d89ebddffb1fb0bc0b734f3695478195ecd7aee0fb3

Observation 5a72782a-6566-4ddb-884c-0e8661853b19 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.172488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.172488Z digest=sha256:29bb9e84ce4019a049783e0d546f66c29d889fbfd31bcb76d44e50c61e7c525c

Observation 685df9f8-685a-4bf1-9b69-598c1c271207 · outbound

This paper cites Rouge: A package for automatic evaluation of sum- maries,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Rouge: A package for automatic evaluation of sum- maries,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:17.246884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:17.246884Z digest=sha256:40c77bcdf418c971eafe005f5b89bc7ac03b8685d2261dab983bbdcec253a699

Observation bf2ab7b3-ee15-491a-922a-ec8ebbf9222e · outbound

This paper cites Cider: Consensus-based image description evaluation,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Cider: Consensus-based image description evaluation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.685683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:17.364740Z digest=sha256:2d69d968b317175b0d9f49e74093edefa162848ca12bbba61bfe40a50a4714b0

Observation fbc936b4-d313-491e-9f1a-439a4feb4b53 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Spice: Semantic propositional image caption evaluation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.419907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:17.461337Z digest=sha256:4d9871a01cbeaa7f19b39494a2881877afab016ff63e4890aa0d6ab03dac1faf

Observation 62b9ede9-917c-48ea-a188-2d2857d12da9 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Improved image captioning via policy gradient optimization of spider,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:20.172010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:17.534077Z digest=sha256:183631a4b13dd1ba68eecc30c9440a76286bae456c256b913fd57fb9e91db6a0

Observation 32c864e5-d2ad-45be-96cf-e0ad68300752 · outbound

This paper cites Microsoft coco: Common objects in context,.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Microsoft coco: Common objects in context,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:19.936515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:17.616572Z digest=sha256:12d6d52b7ec120c1ea7b08d29fcf938e89301f653081527eefabf47bfa0ae1e8

Pith citing papers

Observation 5b3a8d93-2fb8-4921-ad48-6933a4507c29 · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:20:19.650775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:20:14.606479Z digest=sha256:ca1cf2b1873d8b6bca74947034707269877c634a870067f678245d667c9fcf69

Observation 8db4405e-fa54-4f54-95d9-ff440f26235c · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.106889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.106889Z digest=sha256:37deba14b04f8bf1cbc6dfa2dc3ae8a46fcdc7ed80935a870202d07d0f9f2b26