Pith. sign in

Paper Citation Record · LEDGER

Classifier-Guided Captioning Across Modalities

As of 22 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2501.03183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03183 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:18:05.048266Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78210de9-879c-4326-9946-a11aa4dcc1a4 · outbound

This paper cites Show and tell: A neural image caption generator,.

Classifier-Guided Captioning Across Modalities Show and tell: A neural image caption generator,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:04.919134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:04.919134Z digest=sha256:bae2a2766d82c38e2a86ff478d78ddfe9bb2f969c81492f6dd64bdd2d8d7ecc1

Observation b15ac4e8-e6a1-4262-af59-0d6cce8667b2 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Classifier-Guided Captioning Across Modalities Deep visual-semantic alignments for generating image descriptions,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:04.923828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:04.923828Z digest=sha256:51385cf2be0d25b773d6240b1d0ffe66b8c453e32bd43b7c7f8d8db73214674d

Observation c506ad58-7a92-4a3b-9aab-ed9695ec1c51 · outbound

This paper cites Long-term recurrent convolutional networks for vi- sual recognition and description,.

Classifier-Guided Captioning Across Modalities Long-term recurrent convolutional networks for vi- sual recognition and description,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.472753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.927640Z digest=sha256:62a0ec96cb4c3503dacf2d2ccc15665b3c44f568148292da389992e6d2177acc

Observation 4a8c7b6b-d18e-4c4f-ae28-f286b5dd11b0 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention,.

Classifier-Guided Captioning Across Modalities Show, attend and tell: Neural image caption generation with visual attention,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.459360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.931781Z digest=sha256:6cfff617d8d314ba5078afc9b5cdb59b069517b4c3d1d324c8c734132d1eefa7

Observation a05dee41-949c-40fd-a2f9-291058b70830 · outbound

This paper cites Image captioning with semantic attention,.

Classifier-Guided Captioning Across Modalities Image captioning with semantic attention,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:04.935997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:04.935997Z digest=sha256:0468896feecffc7f7a117fd05b84cd1e4b8ba1336e26c32e47ab382f5f637ad8

Observation 0f391700-c4f3-4a8c-94e7-fd484b72d490 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

Classifier-Guided Captioning Across Modalities Bottom-up and top-down attention for image captioning and visual question answering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.438905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.939836Z digest=sha256:f50589f4ef3362e5d40ddeae54a0b8bd555ba6467433757efabb0c440b8a487d

Observation bd93a27d-58cf-46bf-8bfc-ee0e49ecf9de · outbound

This paper cites Automated audio caption- ing: Describing audio content with natural language,.

Classifier-Guided Captioning Across Modalities Automated audio caption- ing: Describing audio content with natural language,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.427091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.944049Z digest=sha256:66cee2b41f8878c5c1971ea84859cc3a19bf778b8af04282eb263ed3738b7d6a

Observation 322c8105-402b-44f4-891e-1deb1e77f68b · outbound

This paper cites Audio captioning using pre-trained large-scale language model guided by audio-based similarity,.

Classifier-Guided Captioning Across Modalities Audio captioning using pre-trained large-scale language model guided by audio-based similarity,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.414664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.947529Z digest=sha256:f652a1e22415269cd57e1d382a95e11c89eac535fba004791753e82997043ad2

Observation b14266be-1e21-4dd9-ac0c-be6d3f71119a · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

Classifier-Guided Captioning Across Modalities Audiocaps: Generating captions for audios in the wild,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.403179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.951128Z digest=sha256:830cd354d69c4aa217f7e061e54c9a001b0003ec02d7df8cf9d5e5e0163d9f8a

Observation 80f109b1-bc92-4bd0-b6c0-b28397e315b7 · outbound

This paper cites Audio caption: Listen and tell,.

Classifier-Guided Captioning Across Modalities Audio caption: Listen and tell,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.391796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.954685Z digest=sha256:817d82e09fda0e653a3d1cf3743b8951cb709284eb7f7e58f072f2a34c6e6833

Observation ff2682c8-463d-485b-90bf-5841ec2584b7 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Classifier-Guided Captioning Across Modalities DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:04.958544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:04.958544Z digest=sha256:8a3aa7e16a1f56a0f916f173f912edf557503e08604b0a7bb09110de7cd1890c

Observation 81a354f7-3592-4422-97cb-4744973ef479 · outbound

This paper cites Clotho: An audio captioning dataset,.

Classifier-Guided Captioning Across Modalities Clotho: An audio captioning dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.381034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.962574Z digest=sha256:a06afa82e0290fe696da36a27c1057a790c54b2f80ae68f3aeb0a0b71b2e38ca

Observation 6a6f8b90-d07f-427d-b6d1-edc61257ccbe · outbound

This paper cites BLEU: A method for automatic evaluation of machine translation,.

Classifier-Guided Captioning Across Modalities BLEU: A method for automatic evaluation of machine translation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.369913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.966230Z digest=sha256:3bb64e7b7a0d0dfe261a83c7d141dfbc0dc1a4873b8f787f49779dadc63e91a0

Observation 94d10fb7-a134-4147-a3f2-17505e53cb5d · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,.

Classifier-Guided Captioning Across Modalities METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.358608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.969722Z digest=sha256:2e5f081fab8e8574cd4cccb621cc4cb822f4a355dbb4888d99fac06348872951

Observation 0b9e62ec-d122-47c7-90e6-78756a178118 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

Classifier-Guided Captioning Across Modalities ROUGE: A package for automatic evaluation of summaries,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.346912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.973205Z digest=sha256:1c33f9ca87d19aa308b622987a1e3216583c4e61caae7c129b67c42436a5d956

Observation e734133b-4db1-406e-b0d3-6c862389dd3b · outbound

This paper cites SPICE: Semantic Propositional Image Caption Evaluation,.

Classifier-Guided Captioning Across Modalities SPICE: Semantic Propositional Image Caption Evaluation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.334971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.976898Z digest=sha256:f0d54cc95232f4816aa69bcda12f1b4213bf516d4e1cf2ea411354f5cf0ded38

Observation 09c650db-31e4-40ca-82b6-21de913e4f24 · outbound

This paper cites CIDEr: Consensus-based image description evaluation,.

Classifier-Guided Captioning Across Modalities CIDEr: Consensus-based image description evaluation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.322899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.981184Z digest=sha256:0995ab66d6f194cf95763c2fc217c8a274c64f5b536ecfd0ad7b258df66d3467

Observation 66cd22c9-e800-41f3-a1ed-9abfa0cb7167 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Classifier-Guided Captioning Across Modalities CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:04.984909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:04.984909Z digest=sha256:b68584099a6397541804c8e3c78e358da3846748115d7d749598645c3f5375d9

Observation fe1d891d-a31b-4ac8-864f-e72f57f0b7ed · outbound

This paper cites ZeroCap: Zero-shot image-to-text generation for visual-semantic arithmetic,.

Classifier-Guided Captioning Across Modalities ZeroCap: Zero-shot image-to-text generation for visual-semantic arithmetic,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.311402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.988864Z digest=sha256:4568967a38c154d394be13ac95c7926e9a8b49d502c63306a22c944336b28b87

Observation aed2f169-d73e-4828-b27c-650bb3656e7e · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Classifier-Guided Captioning Across Modalities ClipCap: CLIP Prefix for Image Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:04.992584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:04.992584Z digest=sha256:4b9e9838ca7e660a092379166eaf3163b8df700529e01f62d354e11f42c8f14d

Observation 71fb319e-6ce4-47ba-9ad8-26be86c9db02 · outbound

This paper cites Prefix tuning for automated audio captioning,.

Classifier-Guided Captioning Across Modalities Prefix tuning for automated audio captioning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.298770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:04.996728Z digest=sha256:8dc62f82a1f1c082b0c8fa3fbe770de41299e0c607047fe4cf057a8d19cdd790

Observation 1366c640-4e96-4369-84c4-d51dcb07ecf1 · outbound

This paper cites RECAP: retrieval-augmented audio captioning,.

Classifier-Guided Captioning Across Modalities RECAP: retrieval-augmented audio captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.286869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.000446Z digest=sha256:49aea33fc07bcedc494a4deb480a24fb8010953981c9cf6036876404df67fd9b

Observation 8899e979-39e8-4e3b-b82a-ae50ee3904c1 · outbound

This paper cites Zero-shot audio captioning with audio-language model guidance and audio context keywords.

Classifier-Guided Captioning Across Modalities Zero-shot audio captioning with audio-language model guidance and audio context keywords

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:05.004061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:05.004061Z digest=sha256:c51f8b38038ae40e6f8e6bf7932b5649cc081c9930dc033f0f0ac522c0f7ebd4

Observation 2a395bbe-361e-4761-aa54-790fdd99c23b · outbound

This paper cites Training audio captioning models without audio,.

Classifier-Guided Captioning Across Modalities Training audio captioning models without audio,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.273630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.008543Z digest=sha256:7037b941f934e39629fa0c46ccddcf37a042e4a75726e6acc231c84153109ac4

Observation 926b9a59-d539-4b50-baad-335164637750 · outbound

This paper cites Weakly-supervised Automated Audio Captioning via text only training.

Classifier-Guided Captioning Across Modalities Weakly-supervised Automated Audio Captioning via text only training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:05.012263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:05.012263Z digest=sha256:cc17999a98e2db9585999bc0cc9b93bf15b822bd889db61e18f6fb6435543af9

Observation 922d3f24-c35f-49e9-8c59-944ca82a7cf3 · outbound

This paper cites Zero-Shot Audio Captioning Using Soft and Hard Prompts.

Classifier-Guided Captioning Across Modalities Zero-Shot Audio Captioning Using Soft and Hard Prompts

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:18:05.101683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.016484Z digest=sha256:f887f5b23b792584ff9b9691fe19e99661e4be0a4fc671e532eeaef11f0806a9

Observation 658df8bf-b551-42db-aae4-5ebd97215674 · outbound

This paper cites Show and Tell: A Neural Image Caption Generator,.

Classifier-Guided Captioning Across Modalities Show and Tell: A Neural Image Caption Generator,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.260538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.020540Z digest=sha256:2e2005031c98626ba6a5d8cd40c7495c0beec617c89cffbccb36b0a9bf13f68c

Observation 89a915bf-2a91-4118-9e1a-f29875a5903f · outbound

This paper cites Show, Attend and Tell: Neural Image Caption Gen- eration with Visual Attention,.

Classifier-Guided Captioning Across Modalities Show, Attend and Tell: Neural Image Caption Gen- eration with Visual Attention,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.247832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.024407Z digest=sha256:a523ae35a830ad17de1ea9c6f8d0b4e9a917485b4a7d8dca6a2ac76ee124e2ad

Observation 1c3f901d-dfb4-45b3-896e-64a4f55905c0 · outbound

This paper cites Self-critical Sequence Training for Image Captioning,.

Classifier-Guided Captioning Across Modalities Self-critical Sequence Training for Image Captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.233886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.028373Z digest=sha256:6b1fc2fb88eac62ba3c083e61b7ddc361f2e36f7291c9158e5b30c3bdf858e61

Observation 1648d6af-eafd-4f0c-a71c-bfd8d129bdec · outbound

This paper cites Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering,.

Classifier-Guided Captioning Across Modalities Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.220173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.032219Z digest=sha256:91d1c037343a87f3d63597cb38db4cf1732e8f083db566efbff813aea195da82

Observation 88a5efc7-3120-4221-9938-61c06bc2f4eb · outbound

This paper cites Microsoft COCO: Common objects in context,.

Classifier-Guided Captioning Across Modalities Microsoft COCO: Common objects in context,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.206008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.035896Z digest=sha256:1440cd18d679055f240b1037ab58e924d84396f473539e323c1a6716c9e6dbc1

Observation 32aff856-5090-4c8f-a15d-d4ea27a5b10a · outbound

This paper cites Flickr30k Entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models,.

Classifier-Guided Captioning Across Modalities Flickr30k Entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.191045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.039652Z digest=sha256:279774679b64b011db7df42d17b651c2cd665b3c4cd502727b7f5cbbd8e9abc4

Observation 73e2fdb5-7765-4d30-9f24-fd27c1532720 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Classifier-Guided Captioning Across Modalities BERTScore: Evaluating Text Generation with BERT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:05.043287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:05.043287Z digest=sha256:9ae033505bfc171a2eedf3865a989fc06323fcfa7e2041a455066ec8341bcacf

Observation 691c4315-bb22-4fcf-9b97-0e2611d79279 · outbound

This paper cites CLAP: Learning audio concepts from natural language supervision,.

Classifier-Guided Captioning Across Modalities CLAP: Learning audio concepts from natural language supervision,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:18:05.177813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:18:05.048266Z digest=sha256:b240c4fc8ba037958cda3f201479821bb3bf1ba280c3bfc456d4ce2d71e6b126

Pith citing papers

No inbound Pith citation observations are available.