Pith. sign in

Paper Citation Record · LEDGER

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation

As of 16 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:1909.02489.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.02489 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:54:04.946058Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5271c8d7-f91e-4229-bb64-dd531300903e · outbound

This paper cites Grounded compositional semantics for finding and describing images with sentences,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Grounded compositional semantics for finding and describing images with sentences,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.365653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.835608Z digest=sha256:e15c10a679f20a879bc26e6d744a940f2b55ff7fe43f8f942b2081af9a0ff4e6

Observation edff3a68-14f4-4011-acdb-e385d9c32882 · outbound

This paper cites Explain Images with Multimodal Recurrent Neural Networks.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Explain Images with Multimodal Recurrent Neural Networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.839718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.839718Z digest=sha256:3cb36d22d188f4b530e840d297d2cfc776f1111ebb64bb55940e590fe86a03be

Observation f824b446-a2ba-4477-a309-399d6369fbb4 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Deep visual-semantic alignments for generating image descriptions,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.352318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.843544Z digest=sha256:382fc938693bb4f3896f4ab61cf27869820033ef206d204da1e8ea950a3f6836

Observation 633ac000-b3ec-4262-a14b-d739611ef342 · outbound

This paper cites Show and tell: A neural image caption generator,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Show and tell: A neural image caption generator,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.847349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.847349Z digest=sha256:b66d5b59444cad31f8720ff20bb1922c178c79639bdfd16c153f3fb8a458b54d

Observation a2690025-8471-410f-a58c-92b4edc65378 · outbound

This paper cites What value do explicit high level concepts have in vision to language problems?.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation What value do explicit high level concepts have in vision to language problems?

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.332693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.850276Z digest=sha256:af14fa46186243f838b1cc1c9b1838f889490b57ae832fba5e943d5c85a75440

Observation d859782d-0415-4d35-a1e3-03ff3342d2ca · outbound

This paper cites Learning to Guide Decoding for Image Captioning.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Learning to Guide Decoding for Image Captioning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:54:05.053799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.853093Z digest=sha256:30f901c36af928d42df3385868927f6614271e106088303b3cef455fe2b13e65

Observation 20ce009d-762e-4c23-ab74-59dd47101b1e · outbound

This paper cites Boosting image captioning with attributes,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Boosting image captioning with attributes,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.321435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.856659Z digest=sha256:58496b0ab4e66a91ae44b820cd22c47ac72db67d50cfa9dae38e1bd309a045d9

Observation 50407182-5489-4c32-b4aa-69e8b8ddc515 · outbound

This paper cites Long-term recurrent convolutional networks for visual recognition and description,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Long-term recurrent convolutional networks for visual recognition and description,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.309260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.859353Z digest=sha256:dba695d6e41bb0a268e2705cda160d004aaec801110f40085bea1e9d9d5daa54

Observation 9760106a-97ea-4421-a9e5-df028cecf750 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Show, attend and tell: Neural image caption generation with visual attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.297203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.861995Z digest=sha256:89920ba8e64d78ad6e1b0c1c4036e8643ab1d59cf4c1523372a5beda9f7e7c9b

Observation 90f0e4ca-c44b-48f7-b956-b5107ec9559b · outbound

This paper cites Composing simple image descriptions using web-scale n-grams,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Composing simple image descriptions using web-scale n-grams,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.286263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.864757Z digest=sha256:f4f76ec8b3b1a71d56094615ef3fcb3de3b8e6722d97dde414bbf2ffdd05e7df

Observation f0bd61e3-3ba9-49e2-8976-f136d088106b · outbound

This paper cites From captions to visual concepts and back,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation From captions to visual concepts and back,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.275515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.867570Z digest=sha256:86382010e837370c2ba13795ef7f900c8b4bf3a0220d10a6855d8b5de7eb647c

Observation 8e94d150-d86f-4b1c-bb5c-acea114c07d1 · outbound

This paper cites Image captioning with semantic attention,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Image captioning with semantic attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.264806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.870372Z digest=sha256:78c30e6fe01b6740ec6dac9b5dfd934f7904ed6657dff389134b18174e4878dd

Observation 048f97c5-045c-456d-9af4-4db1a0638267 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Bottom-up and top-down attention for image captioning and visual question answering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.252864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.874210Z digest=sha256:46845f00dba9f3968fe6344c65daba1cfbc18d91d58e59472cbd07750bbf09c7

Observation 033a27d7-e3cf-40c3-87f4-35acffb7d0e7 · outbound

This paper cites Long short-term memory.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Long short-term memory

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.240370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.877858Z digest=sha256:4d76f93e9d0d9ee155d2570ca57c8b0d278f77f1f175ad504b79d1f91c4c2102

Observation f88c1b1b-4ad2-4236-bd21-3ae29400e9e3 · outbound

This paper cites Long-term recurrent convolutional networks for visual recognition and description,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Long-term recurrent convolutional networks for visual recognition and description,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.228583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.881333Z digest=sha256:975829b67d4233085bb8542ba8db4e0b9d6ab3907599475569827fb37b67b7bd

Observation 2638d605-6120-4794-a2fb-bcfd3d7a6e9b · outbound

This paper cites MAT: A Multimodal Attentive Translator for Image Captioning.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation MAT: A Multimodal Attentive Translator for Image Captioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.884978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.884978Z digest=sha256:d51a32204ba154854806f92204de6cff8a5633d0035f43d9b07bdbca7498c7d9

Observation 154d9a95-d587-4474-81a4-98465ffdf8f4 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sentinel for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Knowing when to look: Adaptive attention via a visual sentinel for image captioning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.216613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.888810Z digest=sha256:c7db98a30227ce56b37c7ca177ccfe57d1b3b2efa1037b82064a7af15014b384

Observation 1d71ad26-d6c9-4b8f-a366-ebec72db68df · outbound

This paper cites SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.892361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.892361Z digest=sha256:9839e22e5351f09c048a876dd0a623c2dd22fc66787a255da4f7f7b79395a717

Observation fa9d3f40-5e18-4721-8954-8135f66858e3 · outbound

This paper cites Stack-captioning: Coarse-to-fine learning for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Stack-captioning: Coarse-to-fine learning for image captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.203275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.896236Z digest=sha256:3cca14ed5e643a2446e6e26da9ff82cf494f237f952279a13f880a1fe3d5dd5f

Observation 77508de8-9732-4846-baf8-3a058ad019b5 · outbound

This paper cites Adam: A method for stochastic optimization,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Adam: A method for stochastic optimization,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.899424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.899424Z digest=sha256:a03c2d30af5a86f89b86953f275550dd84f33f90c856e4e3c960286bd51c55dc

Observation 366f9cd0-3c39-4127-afb8-d0dd0220e61a · outbound

This paper cites Self-critical sequence training for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Self-critical sequence training for image captioning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.182907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.901955Z digest=sha256:e1515ecfba628e9fe5059e02d0964d9ad345a85c488853f15d1f8c688b16618b

Observation c9fd4ea3-f525-4ed1-92e9-930085560f67 · outbound

This paper cites Deep captioning with multimodal recurrent neural networks (m-rnn),.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Deep captioning with multimodal recurrent neural networks (m-rnn),

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.171355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.905106Z digest=sha256:5dd5327053d1fd9a439bfb71781e1c617d380f59d9b28061070fcf9be55af546

Observation dcb71e70-1e2d-491e-8d51-eee73a797dcf · outbound

This paper cites Review networks for caption generation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Review networks for caption generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.158688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.908388Z digest=sha256:2a45cdd1ee3bf19928c79aa214fad98d5027ee9ae37a63607303b062fa8bb8c4

Observation d471f483-2115-44d0-a45c-6e80272e0394 · outbound

This paper cites Gated hierarchical attention for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Gated hierarchical attention for image captioning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.146073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.911512Z digest=sha256:93c4a1fc60b5d42f02fed2baf617636fbf6374d5c4bfbf6c1976dddad0fd89f7

Observation 9bc62d48-1dd3-4bcf-bc5f-4f17bd85f076 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Faster r-cnn: Towards real-time object detection with region proposal networks,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.915430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.915430Z digest=sha256:6e913554a6774bda93b4d09c7c0006c3e4546025d6740314130db3e723d2288d

Observation 9e92f6ea-38ab-4575-87e7-deeacaa9e58a · outbound

This paper cites Sequence Level Training with Recurrent Neural Networks.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Sequence Level Training with Recurrent Neural Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.919052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.919052Z digest=sha256:dfae208751aefb9615ea4a1798417df31f7bb361675407cd5149aa425647bf52

Observation 50854ca6-dc1b-40ea-bdd8-f4f10bef00c6 · outbound

This paper cites Improved Image Captioning via Policy Gradient optimization of SPIDEr.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Improved Image Captioning via Policy Gradient optimization of SPIDEr

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:54:04.996732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.922474Z digest=sha256:e674bff6037dfd90e12aa65dab4d5d92f57f28752b2cca9f9cf73580c038c304

Observation e4fe2802-6e49-4871-832a-7b9a4b2a0c06 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.925858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.925858Z digest=sha256:bdcc0d29f3f4d78f18f7aed28ce17b436d7fa776c4b25eace6c9ed46d7fc68e2

Observation fb45ef34-65a3-454a-b256-64cdb27806fd · outbound

This paper cites Microsoft coco: Common objects in context,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Microsoft coco: Common objects in context,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.928673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.928673Z digest=sha256:608d154d2ae0107f0ca306e98b1d5ffbbaae15473de55ad1b50555177a6c7968

Observation 5d886f73-43a7-47a2-989e-5458d4517825 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Rouge: A package for automatic evaluation of summaries,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.118547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.931562Z digest=sha256:1206aba62d44db1aa2fa442d871069d578ed2b4eeac37850ae4192236b950bc8

Observation 722bb495-bb34-483c-bc0c-0173a1f437db · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Bleu: a method for automatic evaluation of machine translation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.934995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.934995Z digest=sha256:a7548f0610a0fc9db448c8652ee6fb33c273ecc6a1a5f1a6d1a82745d988c7e7

Observation 345dc83a-bbf2-4980-ade6-8e25a1a14e70 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Cider: Consensus- based image description evaluation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.938567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.938567Z digest=sha256:e4725a34a980ac253e8422bbe57b5b57d6741d1643cfbb7413c56639c7e56155

Observation 20f4a934-c6f5-4fb9-9660-c95de892fe5e · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.091180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.942130Z digest=sha256:f2f444de60061abdae8f682b55773c177ba0072eb2422254fa8746283e01fc8b

Observation 829302c8-1ebf-4e60-851b-90769b770c6a · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Spice: Semantic propositional image caption evaluation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.079495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:54:04.946058Z digest=sha256:6e71a552775f8c80905cab0feaf3e0ffaa7a9dde4c368a55745067e2934f7ab6

Pith citing papers

No inbound Pith citation observations are available.