Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:18:05.048266Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2501.03183.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:18:05.048266Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 78210de9-879c-4326-9946-a11aa4dcc1a4 · outbound
Classifier-Guided Captioning Across Modalities Show and tell: A neural image caption generator,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15ac4e8-e6a1-4262-af59-0d6cce8667b2 · outbound
Classifier-Guided Captioning Across Modalities Deep visual-semantic alignments for generating image descriptions,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c506ad58-7a92-4a3b-9aab-ed9695ec1c51 · outbound
Classifier-Guided Captioning Across Modalities Long-term recurrent convolutional networks for vi- sual recognition and description,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4a8c7b6b-d18e-4c4f-ae28-f286b5dd11b0 · outbound
Classifier-Guided Captioning Across Modalities Show, attend and tell: Neural image caption generation with visual attention,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a05dee41-949c-40fd-a2f9-291058b70830 · outbound
Classifier-Guided Captioning Across Modalities Image captioning with semantic attention,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f391700-c4f3-4a8c-94e7-fd484b72d490 · outbound
Classifier-Guided Captioning Across Modalities Bottom-up and top-down attention for image captioning and visual question answering,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bd93a27d-58cf-46bf-8bfc-ee0e49ecf9de · outbound
Classifier-Guided Captioning Across Modalities Automated audio caption- ing: Describing audio content with natural language,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 322c8105-402b-44f4-891e-1deb1e77f68b · outbound
Classifier-Guided Captioning Across Modalities Audio captioning using pre-trained large-scale language model guided by audio-based similarity,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b14266be-1e21-4dd9-ac0c-be6d3f71119a · outbound
Classifier-Guided Captioning Across Modalities Audiocaps: Generating captions for audios in the wild,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 80f109b1-bc92-4bd0-b6c0-b28397e315b7 · outbound
Classifier-Guided Captioning Across Modalities Audio caption: Listen and tell,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ff2682c8-463d-485b-90bf-5841ec2584b7 · outbound
Classifier-Guided Captioning Across Modalities DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a354f7-3592-4422-97cb-4744973ef479 · outbound
Classifier-Guided Captioning Across Modalities Clotho: An audio captioning dataset,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a6f8b90-d07f-427d-b6d1-edc61257ccbe · outbound
Classifier-Guided Captioning Across Modalities BLEU: A method for automatic evaluation of machine translation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 94d10fb7-a134-4147-a3f2-17505e53cb5d · outbound
Classifier-Guided Captioning Across Modalities METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0b9e62ec-d122-47c7-90e6-78756a178118 · outbound
Classifier-Guided Captioning Across Modalities ROUGE: A package for automatic evaluation of summaries,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e734133b-4db1-406e-b0d3-6c862389dd3b · outbound
Classifier-Guided Captioning Across Modalities SPICE: Semantic Propositional Image Caption Evaluation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 09c650db-31e4-40ca-82b6-21de913e4f24 · outbound
Classifier-Guided Captioning Across Modalities CIDEr: Consensus-based image description evaluation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 66cd22c9-e800-41f3-a1ed-9abfa0cb7167 · outbound
Classifier-Guided Captioning Across Modalities CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe1d891d-a31b-4ac8-864f-e72f57f0b7ed · outbound
Classifier-Guided Captioning Across Modalities ZeroCap: Zero-shot image-to-text generation for visual-semantic arithmetic,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aed2f169-d73e-4828-b27c-650bb3656e7e · outbound
Classifier-Guided Captioning Across Modalities ClipCap: CLIP Prefix for Image Captioning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71fb319e-6ce4-47ba-9ad8-26be86c9db02 · outbound
Classifier-Guided Captioning Across Modalities Prefix tuning for automated audio captioning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1366c640-4e96-4369-84c4-d51dcb07ecf1 · outbound
Classifier-Guided Captioning Across Modalities RECAP: retrieval-augmented audio captioning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8899e979-39e8-4e3b-b82a-ae50ee3904c1 · outbound
Classifier-Guided Captioning Across Modalities Zero-shot audio captioning with audio-language model guidance and audio context keywords
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a395bbe-361e-4761-aa54-790fdd99c23b · outbound
Classifier-Guided Captioning Across Modalities Training audio captioning models without audio,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 926b9a59-d539-4b50-baad-335164637750 · outbound
Classifier-Guided Captioning Across Modalities Weakly-supervised Automated Audio Captioning via text only training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922d3f24-c35f-49e9-8c59-944ca82a7cf3 · outbound
Classifier-Guided Captioning Across Modalities Zero-Shot Audio Captioning Using Soft and Hard Prompts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 658df8bf-b551-42db-aae4-5ebd97215674 · outbound
Classifier-Guided Captioning Across Modalities Show and Tell: A Neural Image Caption Generator,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 89a915bf-2a91-4118-9e1a-f29875a5903f · outbound
Classifier-Guided Captioning Across Modalities Show, Attend and Tell: Neural Image Caption Gen- eration with Visual Attention,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1c3f901d-dfb4-45b3-896e-64a4f55905c0 · outbound
Classifier-Guided Captioning Across Modalities Self-critical Sequence Training for Image Captioning,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1648d6af-eafd-4f0c-a71c-bfd8d129bdec · outbound
Classifier-Guided Captioning Across Modalities Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 88a5efc7-3120-4221-9938-61c06bc2f4eb · outbound
Classifier-Guided Captioning Across Modalities Microsoft COCO: Common objects in context,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 32aff856-5090-4c8f-a15d-d4ea27a5b10a · outbound
Classifier-Guided Captioning Across Modalities Flickr30k Entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 73e2fdb5-7765-4d30-9f24-fd27c1532720 · outbound
Classifier-Guided Captioning Across Modalities BERTScore: Evaluating Text Generation with BERT
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 691c4315-bb22-4fcf-9b97-0e2611d79279 · outbound
Classifier-Guided Captioning Across Modalities CLAP: Learning audio concepts from natural language supervision,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.