Pith. sign in

Paper Citation Record · LEDGER

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation

As of 15 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:1909.02489.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.02489 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:54:04.946058Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5271c8d7-f91e-4229-bb64-dd531300903e · outbound

This paper cites Grounded compositional semantics for finding and describing images with sentences,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Grounded compositional semantics for finding and describing images with sentences,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.365653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.835608Z digest=sha256:fd12206a7fd7e54b132ceab18a694d3ef8ca2e336285ec916311143fb3e42d3f

Observation edff3a68-14f4-4011-acdb-e385d9c32882 · outbound

This paper cites Explain Images with Multimodal Recurrent Neural Networks.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Explain Images with Multimodal Recurrent Neural Networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.839718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.839718Z digest=sha256:43a5fffcb4fe5990415c0a55d95f282eded6c01e05baa0feda6791403d4b3678

Observation f824b446-a2ba-4477-a309-399d6369fbb4 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Deep visual-semantic alignments for generating image descriptions,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.352318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.843544Z digest=sha256:19de8f586c14fae0e411a88db207ea8a65cdfcfd1986658298b06d46e7b59a40

Observation 633ac000-b3ec-4262-a14b-d739611ef342 · outbound

This paper cites Show and tell: A neural image caption generator,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Show and tell: A neural image caption generator,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.847349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.847349Z digest=sha256:a85b07a1bf2cd233c71fb5d36679e750c8551841662b90d4f055b8de5fe99335

Observation a2690025-8471-410f-a58c-92b4edc65378 · outbound

This paper cites What value do explicit high level concepts have in vision to language problems?.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation What value do explicit high level concepts have in vision to language problems?

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.332693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.850276Z digest=sha256:752eefd73a73c6212330519926b0282c5e5984a162c235f834150c533bbe3ece

Observation d859782d-0415-4d35-a1e3-03ff3342d2ca · outbound

This paper cites Learning to Guide Decoding for Image Captioning.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Learning to Guide Decoding for Image Captioning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:54:05.053799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.853093Z digest=sha256:0f5c0038d4b98fd555821c5f4ab1417b491153f5e20966ac65dbb030c6660e2e

Observation 20ce009d-762e-4c23-ab74-59dd47101b1e · outbound

This paper cites Boosting image captioning with attributes,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Boosting image captioning with attributes,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.321435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.856659Z digest=sha256:91f018196c8d9a73e95d354100bcf3261701934ed6f1feafab034b4dca8487f8

Observation 50407182-5489-4c32-b4aa-69e8b8ddc515 · outbound

This paper cites Long-term recurrent convolutional networks for visual recognition and description,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Long-term recurrent convolutional networks for visual recognition and description,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.309260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.859353Z digest=sha256:c5fc84f523666e9afc0ce2a20b3ece77d0265343067dae15a5b502d948ed66b6

Observation 9760106a-97ea-4421-a9e5-df028cecf750 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Show, attend and tell: Neural image caption generation with visual attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.297203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.861995Z digest=sha256:e608ce4a3d9bedc64378d6e4271b67f8ed22bdb821acd6579da4e83d85712e0e

Observation 90f0e4ca-c44b-48f7-b956-b5107ec9559b · outbound

This paper cites Composing simple image descriptions using web-scale n-grams,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Composing simple image descriptions using web-scale n-grams,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.286263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.864757Z digest=sha256:53da0a3d0ce5282e62a6d8ae12b7e91f95fb9832b4e3d1e33091b62c45f3beaa

Observation f0bd61e3-3ba9-49e2-8976-f136d088106b · outbound

This paper cites From captions to visual concepts and back,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation From captions to visual concepts and back,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.275515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.867570Z digest=sha256:09b83aff02a3e1b42fcce240252831898285c4faef87a4bd5645c3c1715b4691

Observation 8e94d150-d86f-4b1c-bb5c-acea114c07d1 · outbound

This paper cites Image captioning with semantic attention,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Image captioning with semantic attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.264806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.870372Z digest=sha256:49455fe1e832df9db1e638752741c58f02396fd62d8033602480206c62b75490

Observation 048f97c5-045c-456d-9af4-4db1a0638267 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Bottom-up and top-down attention for image captioning and visual question answering,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.252864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.874210Z digest=sha256:2302104fc38037f459b19bc75c875cc0d7a9f6ef05c3fe73be667aede56173f2

Observation 033a27d7-e3cf-40c3-87f4-35acffb7d0e7 · outbound

This paper cites Long short-term memory.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Long short-term memory

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.240370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.877858Z digest=sha256:f73be4d99da6b9428ecb52982bd60bf40aa226c9ddcbac695283cc62736a323b

Observation f88c1b1b-4ad2-4236-bd21-3ae29400e9e3 · outbound

This paper cites Long-term recurrent convolutional networks for visual recognition and description,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Long-term recurrent convolutional networks for visual recognition and description,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.228583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.881333Z digest=sha256:fbd83032a4f7a0c7535e10ab8df7390d8776795ae03bb5a7520a4b66dde2b7fb

Observation 2638d605-6120-4794-a2fb-bcfd3d7a6e9b · outbound

This paper cites MAT: A Multimodal Attentive Translator for Image Captioning.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation MAT: A Multimodal Attentive Translator for Image Captioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.884978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.884978Z digest=sha256:907b55806fcabc68b3ffbfef97432478225035f27dcdb6ea4f0fb9091a2078fb

Observation 154d9a95-d587-4474-81a4-98465ffdf8f4 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sentinel for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Knowing when to look: Adaptive attention via a visual sentinel for image captioning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.216613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.888810Z digest=sha256:4f3efe445da632712239471a5900539f2d7014207fd5cc59836806352af2c111

Observation 1d71ad26-d6c9-4b8f-a366-ebec72db68df · outbound

This paper cites SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.892361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.892361Z digest=sha256:88ae258334cc49d9223ca51f59d491d85ac4b72a315d66e39d6671cf3e9e9406

Observation fa9d3f40-5e18-4721-8954-8135f66858e3 · outbound

This paper cites Stack-captioning: Coarse-to-fine learning for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Stack-captioning: Coarse-to-fine learning for image captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.203275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.896236Z digest=sha256:e3813ea1c3fd9ea7982ab3c2650ca2eeff924878dc1568410ce0ba8dce0b1702

Observation 77508de8-9732-4846-baf8-3a058ad019b5 · outbound

This paper cites Adam: A method for stochastic optimization,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Adam: A method for stochastic optimization,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.899424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.899424Z digest=sha256:330a46f77a2b81ed9e5d3b7da2c23e71642d30cf67ab56efab87e158f6182e4d

Observation 366f9cd0-3c39-4127-afb8-d0dd0220e61a · outbound

This paper cites Self-critical sequence training for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Self-critical sequence training for image captioning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.182907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.901955Z digest=sha256:eaa11367195350efbf3d45a2bf2e87eed55740bdd0ce3648add4763ffbf6ddbd

Observation c9fd4ea3-f525-4ed1-92e9-930085560f67 · outbound

This paper cites Deep captioning with multimodal recurrent neural networks (m-rnn),.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Deep captioning with multimodal recurrent neural networks (m-rnn),

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.171355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.905106Z digest=sha256:7cf24b79411ffc9aaed27b94147d98ebeb2bd7a61288b63312386bedffdd010a

Observation dcb71e70-1e2d-491e-8d51-eee73a797dcf · outbound

This paper cites Review networks for caption generation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Review networks for caption generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.158688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.908388Z digest=sha256:a29a2ef8327cedeef5eadf9c44eef378da2ca63244be082abc61ea032d964ba4

Observation d471f483-2115-44d0-a45c-6e80272e0394 · outbound

This paper cites Gated hierarchical attention for image captioning,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Gated hierarchical attention for image captioning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.146073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.911512Z digest=sha256:257235eef1e26bea123fef26fcd26a40a6af7986edbd332e705b1048a2d8da53

Observation 9bc62d48-1dd3-4bcf-bc5f-4f17bd85f076 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Faster r-cnn: Towards real-time object detection with region proposal networks,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.915430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.915430Z digest=sha256:23800d8a8a8a39b3082634956ff4f9f0649255fb222dda3f16e845bbd0698ee9

Observation 9e92f6ea-38ab-4575-87e7-deeacaa9e58a · outbound

This paper cites Sequence Level Training with Recurrent Neural Networks.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Sequence Level Training with Recurrent Neural Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.919052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.919052Z digest=sha256:79ed35d036caaedce6148d2dae87484f72f54047b752620b0c89e42da1b5548f

Observation 50854ca6-dc1b-40ea-bdd8-f4f10bef00c6 · outbound

This paper cites Improved Image Captioning via Policy Gradient optimization of SPIDEr.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Improved Image Captioning via Policy Gradient optimization of SPIDEr

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:54:04.996732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.922474Z digest=sha256:74ee12f01302554d89fc4818c7aec74ec861fdb6c36615c0cd308846aa9cef15

Observation e4fe2802-6e49-4871-832a-7b9a4b2a0c06 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.925858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.925858Z digest=sha256:539f768ca71e0ba8c6362ea5612ee721d2fe7ce315f654d4de5238ef4e7f7b45

Observation fb45ef34-65a3-454a-b256-64cdb27806fd · outbound

This paper cites Microsoft coco: Common objects in context,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Microsoft coco: Common objects in context,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.928673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.928673Z digest=sha256:31427220ef801d70a7c163b4579793185e39357bd03a2f3d727c7dc92f275c71

Observation 5d886f73-43a7-47a2-989e-5458d4517825 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Rouge: A package for automatic evaluation of summaries,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.118547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.931562Z digest=sha256:4d171e1a3a1df514cf8cbec948fc9b05a31f42bd41d9ad19f4bb67ae486b1c59

Observation 722bb495-bb34-483c-bc0c-0173a1f437db · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Bleu: a method for automatic evaluation of machine translation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.934995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.934995Z digest=sha256:c5da57ab0ee80c58b1241146ec183e699f8e04e13f626cde9b59e703aa64f64f

Observation 345dc83a-bbf2-4980-ade6-8e25a1a14e70 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Cider: Consensus- based image description evaluation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:04.938567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:04.938567Z digest=sha256:2ad3b4fbade519e39703f79df7001d9e92ca83eb2742e67effca284e6f53788d

Observation 20f4a934-c6f5-4fb9-9660-c95de892fe5e · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.091180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.942130Z digest=sha256:69553db0c96f949e584fd44cc8b29c2d688cd628d5079e7f142e5380e8bb8ee0

Observation 829302c8-1ebf-4e60-851b-90769b770c6a · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation Spice: Semantic propositional image caption evaluation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:05.079495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:54:04.946058Z digest=sha256:a146128f39fc086713165ffd84b6b89784594eb58e8a4d54640e6b985fba1b85

Pith citing papers

No inbound Pith citation observations are available.