Pith. sign in

Paper Citation Record · LEDGER

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings

As of 23 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:1908.09317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.09317 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:21:49.358294Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact2
  • verified fuzzy58
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d68a81e-813b-4fc0-8937-6e11956cb70d · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Spice: Semantic propositional image cap- tion evaluation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.395385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.050800Z digest=sha256:d9e1e495e5f0976dbd02fe5a9dd98f49bf87964af3bd1bfe3906f15f8b750f0c

Observation 5af452b7-3d16-462c-879f-85ce8327d493 · outbound

This paper cites Guided open vocabulary image captioning with constrained beam search.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Guided open vocabulary image captioning with constrained beam search

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.381155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.055957Z digest=sha256:3b674f4589b7e49935d1828d36f9038ba8b37407f9c89ce2a565b3a6c539a97d

Observation 97bd61fd-f0ce-4a27-be01-0d1a96643300 · outbound

This paper cites Partially-supervised image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Partially-supervised image captioning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.366747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.060608Z digest=sha256:3f3610f030fca17c192fd903dc8b92688f63312e0e146c6e0df407834f8b0032

Observation cd622e6e-eb10-4e62-9da2-049a44506978 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Bottom-up and top-down attention for image captioning and visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.351769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.065362Z digest=sha256:0b66efc500bcf5dbf9458d7f18bc986ba0077c439fca846e8e4ec3ff5dd669eb

Observation 465e539e-eaee-4d12-b385-cb964f062c7c · outbound

This paper cites Women also snowboard: Over- coming bias in captioning models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Women also snowboard: Over- coming bias in captioning models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.336902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.069958Z digest=sha256:532aa64519fe1a33f880d503fe3c205933592e86bffcb497e1c15d706699de42

Observation cd20f5a7-7f92-487b-9e7c-18306e0943e9 · outbound

This paper cites Deep compositional captioning: Describing novel ob- ject categories without paired training data.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep compositional captioning: Describing novel ob- ject categories without paired training data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.323001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.074636Z digest=sha256:2f993e7efeb0b3a2cdc9efba1134cc0376e36667eb8a9a463a2e84ae454e1ddf

Observation 37608bf8-7d38-4074-bf61-dad8e8ef08f0 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Lawrence Zitnick, and Devi Parikh

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.308732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.079192Z digest=sha256:5feada5c6254102f4c60238adebd921e0858459bfb075ea94ec107191569f154

Observation f0696281-d196-425f-ae58-76177ceb549c · outbound

This paper cites Adversarial text generation via feature-mover’s distance.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial text generation via feature-mover’s distance

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.294420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.084024Z digest=sha256:9bc8c6972500067ef2971b9a60aa1804e2ee7bf390aef33682d1ccbe2eb29ef5

Observation a6a209ca-8435-4475-8d61-c1b6eb3e5543 · outbound

This paper cites Show, adapt and tell: Adversarial training of cross-domain image cap- tioner.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, adapt and tell: Adversarial training of cross-domain image cap- tioner

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.280462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.088273Z digest=sha256:ff6913150971676f97b3fe900bcdb2bdefc89fdb209dd9adbaa38b835a2c9980

Observation 37b91ab8-d4c5-4820-a2d9-9baf2e576db4 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.092525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.092525Z digest=sha256:7e91a486ef38be81c751f973593b8310ac14e99429cc63f77017bbc68fc0ce3a

Observation 14fb92d0-7223-468e-b608-9afa051ee48e · outbound

This paper cites To- wards diverse and natural image descriptions via a condi- tional GAN.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings To- wards diverse and natural image descriptions via a condi- tional GAN

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.266341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.097343Z digest=sha256:790fd759e4733fdc4b1fe04702e53d47db3e16bcbef90999ab6b6dc773623d42

Observation a296666d-260e-4bad-a15c-60fbf54f8b2f · outbound

This paper cites Visual dialog.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual dialog

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.251995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.102402Z digest=sha256:897444ab791fd68611b2a1856540b2d53db9d17bbd1fb32741ba6e83c5d26d4e

Observation 9b6c5ffe-db60-4622-ad3a-ac17afd8e0ba · outbound

This paper cites Meteor universal: Lan- guage specific translation evaluation for any target language.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Meteor universal: Lan- guage specific translation evaluation for any target language

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.238471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.106878Z digest=sha256:68bb2105f17abb80f39a99196fc13ca467b36a21878dc7629a770593a1740c34

Observation 96ddcc7f-eeb6-4429-9539-e8ece88393e6 · outbound

This paper cites Long-term recurrent convolutional net- works for visual recognition and description.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Long-term recurrent convolutional net- works for visual recognition and description

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.224589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.111212Z digest=sha256:2e0dbd7ff68e70613eabd2777339244f57db691fae1bb5913a704adf73a70520

Observation 174b51ff-a2df-4233-a99e-92a9490c6535 · outbound

This paper cites Adversarial Feature Learning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial Feature Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.115213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.115213Z digest=sha256:ac5bcdcdce935a58a9e72a80dd405901bb40e3cdf7d4e2ef8ba1f24209af5cc1

Observation d557ae01-89a5-4f1c-8b2f-3ea5e75fa703 · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.119851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.119851Z digest=sha256:4399026a601006578ca0a61415bcd6a904039695e87a25fa21176e4b3f2b0260

Observation aea4e3e0-1b75-4341-b3a1-40ca4affc4dc · outbound

This paper cites From captions to vi- sual concepts and back.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From captions to vi- sual concepts and back

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.209802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.124345Z digest=sha256:97270a54bb00a111e8b4b141c147b8a1433f0dc04d5329b81d9e7cb260fd0ac1

Observation 4866dddf-2f31-4059-b0bd-2177e7d0a151 · outbound

This paper cites Unsupervised Image Captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised Image Captioning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:21:49.464577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.128637Z digest=sha256:716b4255dee8c1e31b055aed963e73db66ee0f6b3c952d3e86179cdb5a2b9735

Observation 79f3acfb-a3a4-4087-bf21-c112bc25ec89 · outbound

This paper cites StyleNet: Generating attractive visual captions with styles.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings StyleNet: Generating attractive visual captions with styles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.194138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.133156Z digest=sha256:7b70d01a0781631b733fa7da134908ca7b0d86979963f03067db24978d8ff962

Observation ce139965-5b8c-46f6-ad4e-724719239783 · outbound

This paper cites Un- paired image captioning by language pivoting.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Un- paired image captioning by language pivoting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.179739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.138236Z digest=sha256:6433761203371a70b38710bafd3884f8bded85825a4bf670316e805dc30d04a6

Observation 5487240a-7548-4ed3-9600-ac34efdb04ac · outbound

This paper cites Improved training of wasserstein gans.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved training of wasserstein gans

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.165547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.142747Z digest=sha256:39fdf21fc2c647ec39724729d3e64e9ee302342d42485dff0e0a1d4a9b7a247e

Observation 711de645-aeae-4efe-b938-683a37a07f0c · outbound

This paper cites MSCap: Multi-style image captioning with un- paired stylized text.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings MSCap: Multi-style image captioning with un- paired stylized text

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.151076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.147579Z digest=sha256:92dfbe0d3606d92f8f13a73996b4ef12a31e1497a437d4770ce5fe63a977f2a2

Observation 7d9848f4-4871-4f17-aedb-d955dd1e28a6 · outbound

This paper cites VizWiz grand challenge: Answering visual questions from blind people.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VizWiz grand challenge: Answering visual questions from blind people

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.134384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.151969Z digest=sha256:3c1b5097958dbb5da7bb3acf0c605c87027d5ff072b9c41caea4d8fa501535b1

Observation f1255e8a-a030-44fd-bf18-fc2e7ac68918 · outbound

This paper cites Deep residual learning for image recognition.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep residual learning for image recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.156222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.156222Z digest=sha256:b6bc5a79f63a875499794aec7b46f4456d50bb8156c3d33d29f65a6842fb0fb1

Observation f7e9902d-867a-47e4-acbc-3be5f12ab9bd · outbound

This paper cites Speed/accuracy trade-offs for modern convolutional object detectors.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speed/accuracy trade-offs for modern convolutional object detectors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.109732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.160518Z digest=sha256:dbea1ed0d67396fba99acc69e5c2c3ea191bc1085c2b4549a2836ee34475d4d9

Observation 7412c2ab-6660-4bc3-b97e-1cbfe57af8b9 · outbound

This paper cites DenseCap: Fully convolutional localization networks for dense caption- ing.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings DenseCap: Fully convolutional localization networks for dense caption- ing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.094595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.164737Z digest=sha256:4abc37dc642c132e8e1448b778fe68f4a266683fb3bc95561234f45d583d71e2

Observation 13c9505c-4874-4445-b13b-17d8a512eed3 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep visual-semantic align- ments for generating image descriptions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.080284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.168914Z digest=sha256:2eefe15f8d80eb7cc74fc66a5a47567a8a1507ffaa3ef9943e87a9396e02e756

Observation e921cc49-5e6e-4c28-bfa2-e04e0d2a58bd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.173388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.173388Z digest=sha256:26ad83978ff42e7e2f6c8903d2bf9389d9d74a6ac31b03742ba0d26ca340b92d

Observation 7256ea6d-0897-4078-a360-1a693a8a1913 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.177786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.177786Z digest=sha256:e24826ab0144ee9a4c795e927abf843b49fa3335d473b94f13d56de266f294d8

Observation edbb9db4-cad3-4413-ad80-ab692e553d16 · outbound

This paper cites OpenImages: A public dataset for large-scale multi-label and multi-class image classification.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings OpenImages: A public dataset for large-scale multi-label and multi-class image classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.065623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.182853Z digest=sha256:be7db7d4a4891b7c092cbae9b5ba59ff8f96cdb2d8e1acf010385a44f2e5957b

Observation 140f1fa0-6291-4298-aac8-482c66e61aae · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.050852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.187321Z digest=sha256:288e2ce5a948384efae7383972ddeb38c48142da1a06daa08bc80fe1456266b5

Observation 71e19b4c-db59-4d3a-8318-29c1d03ef217 · outbound

This paper cites From word embeddings to document distances.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From word embeddings to document distances

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.035102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.191699Z digest=sha256:08f3f6a0a47451d177f8d1dd55f5b4d12d67f147600b3f53b8e859f3f94a820c

Observation e639d1fd-85ff-4d48-91aa-73dce18fb0a8 · outbound

This paper cites The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.195941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.195941Z digest=sha256:696ff0c3a27f9a1c9143113f0a9125eba5ca1670a6f6f09f252f2b43f16ddb4b

Observation 3547e4ff-82bd-4d82-913a-d8777c58e2ce · outbound

This paper cites Unsupervised machine translation using monolingual corpora only.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised machine translation using monolingual corpora only

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.019131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.200592Z digest=sha256:381bc76fec7e8542e67f8fda483fff444a8fd06f4abeadf936fd0f95fcf33315

Observation 87046abb-bc66-47a2-ad59-99ee07ee2c10 · outbound

This paper cites Phrase-based & neural unsupervised machine translation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Phrase-based & neural unsupervised machine translation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.003670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.204978Z digest=sha256:ac0afdc045024486b962107e019d8447165b97b68fc0216a3a314cdd56ea1128

Observation 75b16d00-50a4-4480-94a6-2fd4d930ec3a · outbound

This paper cites Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:21:49.400024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.209513Z digest=sha256:46f0dc2b1fcecd8b61ee20fdf10104937dba26420a0ac59231848ccdac64dba9

Observation 1966479c-1b24-4bb0-9faf-d738c6c39e50 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Rouge: A package for automatic evaluation of summaries

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.989330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.214205Z digest=sha256:0a350f967b665cffe8372675c71e53e88828ee4cc1154a11740bb553a6da5b3b

Observation 11b42c84-79c1-4e7b-ac00-ae6e040add50 · outbound

This paper cites Microsoft COCO: Common objects in context.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Microsoft COCO: Common objects in context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.974238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.218844Z digest=sha256:7c82ece25e63330ab3b98ad1d8314fd3c7def318942d0350acaad8e8d0bb4f87

Observation d0480356-7346-47b4-9fcf-19c9de0e11dc · outbound

This paper cites Teaching machines to describe images via natural language feedback.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Teaching machines to describe images via natural language feedback

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.958902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.223234Z digest=sha256:e536566d7305d1cae9d4985fe3df34c69b80df6722a742d20132ca8d32a08078

Observation 5b174ceb-e1e2-44d7-bef1-d72524f35b61 · outbound

This paper cites Improved image captioning via policy gra- dient optimization of spider.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved image captioning via policy gra- dient optimization of spider

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.228016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.228016Z digest=sha256:27e6b75219d28cd3737d48c89876f86a224902a08b19b4ca37bbf0e86ecd6f38

Observation a4a589ba-ea53-45d1-b4cd-00b88a060639 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sen- tinel for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Knowing when to look: Adaptive attention via a visual sen- tinel for image captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.232199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.232199Z digest=sha256:c16d4f865f2a5834577415f2d575d4e7a5c228e115056173ed7a8019f5f6a7c4

Observation 5297d2d2-0ce3-4cd1-a37d-29822afa472b · outbound

This paper cites Neural baby talk.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Neural baby talk

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.925602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.236314Z digest=sha256:76a84e87f449a4e984840d95a2fbadcb99a6260c904b5caedfdc4e06b0ef9c04

Observation 4646014e-da54-4b69-b8b8-d63ddf7abc43 · outbound

This paper cites Visualizing data using t-SNE.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visualizing data using t-SNE

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.911058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.240435Z digest=sha256:6e933552976a98eaca93e9b4436a2b28ac612e889e8eddbae4034250e93b11bb

Observation 61847bfc-ef07-40d7-99c1-324254d67622 · outbound

This paper cites The stan- ford CoreNLP natural language processing toolkit.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The stan- ford CoreNLP natural language processing toolkit

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.896734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.245049Z digest=sha256:7d4449faff16b77d412ba96a401634475efb13e5fdcbf050533c4978be48c7a4

Observation c8eec48e-6ba1-4a5d-ad17-e8f0b04ea120 · outbound

This paper cites Learning like a child: Fast novel visual concept learning from sentence descriptions of images.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning like a child: Fast novel visual concept learning from sentence descriptions of images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.883100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.249691Z digest=sha256:60c396afd9ad0245b3f6a3d21c2de83a66b8e6401cc421e967af31354f4acd54

Observation caaacc1c-8c40-455b-8dde-0f7985c42798 · outbound

This paper cites Sem- Style: Learning to generate stylised image captions using unaligned text.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sem- Style: Learning to generate stylised image captions using unaligned text

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.868624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.253865Z digest=sha256:7d8afc77f9ee0c62d427fe266c8c290711beb0d26064c1430b7d80f4a8648e49

Observation 081451b6-48d3-42cd-8ff9-c1a9c6202780 · outbound

This paper cites Jointly modeling embedding and translation to bridge video and language.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Jointly modeling embedding and translation to bridge video and language

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.853953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.257982Z digest=sha256:c7505c7d7b848f954352929d0799e80b515917825a0a6c7909bfc20f8ae52497

Observation 210a1364-a49a-4038-a1dd-46f43fbb0cb8 · outbound

This paper cites GloVe: Global vectors for word representation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings GloVe: Global vectors for word representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.839595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.262233Z digest=sha256:89bc5bfabac5e2010f6cbf5122c2b34c5121e245b29511e08c748952699b27a0

Observation 9d83399a-4384-45d3-8d09-8f5d91b25f1f · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.825483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.267346Z digest=sha256:6ed36ba195a9c770ccd6b5a41827b80ee724b4cbaf3675bf2969c2625827a97c

Observation b2437259-aa6f-40e0-a48e-16a13e1e9d77 · outbound

This paper cites Self-critical sequence training for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Self-critical sequence training for image captioning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.810801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.271747Z digest=sha256:4bb9ac5f5b73945e6209e0a2a6ab3e6b5b2ebbe7ec5dd9983054cb7b77b7f56e

Observation 08d40f3f-c859-47e4-af41-b5c0fd0e97c8 · outbound

This paper cites ImageNet large scale visual recognition challenge.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings ImageNet large scale visual recognition challenge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.796526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.276815Z digest=sha256:f122aecc5bcfc384e818808fe7f673219c69abd911ce49820ec302fdc4fd383a

Observation 0fd873de-d811-4278-bf7c-5eb2202282fb · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.781359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.281797Z digest=sha256:eb0b8b0cbc4ad34401e30cb64160489db60a172b849f4e79a1d2609ad452be39

Observation 8d405e70-3988-446d-8598-da7516652da1 · outbound

This paper cites Speaking the same language: Matching machine to human captions by adversarial train- ing.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speaking the same language: Matching machine to human captions by adversarial train- ing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.764925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.286027Z digest=sha256:ab9d72f515e4aedefea1ee48d370727cafb04632f21cc941137b278077a27fb6

Observation 21de770f-2a64-44db-be90-6b321b68c7ea · outbound

This paper cites Deforming autoencoders: Unsupervised disentangling of shape and ap- pearance.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deforming autoencoders: Unsupervised disentangling of shape and ap- pearance

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.750435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.290120Z digest=sha256:0f64202241586e092445d782e932367e54a688a49c57e2f8490e6e02ba2db958

Observation 24d7acb7-651e-404a-b054-6fb0e294b2af · outbound

This paper cites Engaging image captioning via per- sonality.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Engaging image captioning via per- sonality

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.736319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.294410Z digest=sha256:5eae2089a9a1efdbcab45d89fae5747f71d32e38acf4fc260244d7a59a06d9c3

Observation eb488959-56e3-4534-83f7-3e5b4a2a5dcc · outbound

This paper cites Towards text generation with adversarially learned neural outlines.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Towards text generation with adversarially learned neural outlines

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.722240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.298686Z digest=sha256:aba49bea5f94450b717da2c6b6235fe15a92af59c300143dbd8ca4792c6a226e

Observation 3b62f948-9c4f-425c-87a7-e44dc90b407d · outbound

This paper cites Sequence to sequence learning with neural networks.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sequence to sequence learning with neural networks

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.707847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.303007Z digest=sha256:0fddfdb48391887d98a168143f8ec566c21fffeb1aa153eb55e199e3070c751b

Observation 2ca248f7-0463-4f3b-be81-6aca77e5b0f4 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Cider: Consensus-based image description evalua- tion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.307627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.307627Z digest=sha256:f4df4962961ba1e3571819d3f54edb10961dc1e73f5e3a08da41fe649a8e6711

Observation 1e69ab5a-de3b-4ad6-abcc-de758b06f2a1 · outbound

This paper cites Captioning images with diverse objects.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Captioning images with diverse objects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.684154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.311808Z digest=sha256:c6c92d5978b4edb62b5147895d8a7da38a7aa461b3761075e389b1f4f82f64a1

Observation add7915b-047b-4bb5-9fb3-b86bc1d53de2 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show and tell: A neural image caption gen- erator

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.669643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.316206Z digest=sha256:9f3bf6ae520e5ba70b8a1cd1d1fc716f3908ae609a5ac39227a6ce0edef79a25

Observation aba98efa-5dc5-4e67-ac51-3ba198649935 · outbound

This paper cites Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.655861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.320325Z digest=sha256:5198d63b36295f3d024b36343c7c7407959f21f0bcbbd0ec2599cd3f0bef1861

Observation f563f42d-6c22-4729-a433-657eb34d8d69 · outbound

This paper cites Automatic alt-text: Computer-generated image de- scriptions for blind users on a social network service.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Automatic alt-text: Computer-generated image de- scriptions for blind users on a social network service

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.641250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.324482Z digest=sha256:feb1cb2df929900461321f061cf83fe63014e720c3f0c3329ddcbba3b4ac434b

Observation ae702d51-7242-4aeb-a8c3-04d1e674e811 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.626520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.328768Z digest=sha256:d057663dd93e5ebb69814a4bbf67359d8472c378c3320c42afa55c63cd52057b

Observation 3672b3b3-97c3-459d-8afe-c45b03a8cb18 · outbound

This paper cites Review networks for caption gen- eration.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Review networks for caption gen- eration

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.612477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.332844Z digest=sha256:c7b23558ebd3548edef6759a785f2cae129f3787c639908c420d65f5b304880b

Observation 0f0f9291-e594-40ce-a1ad-0e53b8936924 · outbound

This paper cites Incorpo- rating copying mechanism in image captioning for learn- ing novel objects.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Incorpo- rating copying mechanism in image captioning for learn- ing novel objects

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.597659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.337040Z digest=sha256:0775b9f64dfb6dd603a3c5f77270962096929ef7f9bea0a0e3221a3cae54c0a4

Observation d2256260-1f4b-4f0b-ac04-a156ac8daaca · outbound

This paper cites Explor- ing visual relationship for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Explor- ing visual relationship for image captioning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.580762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.341293Z digest=sha256:5354b42fc9c102d0c37620ad66dc8266d99f5e65c064fc4977ae86b6a93a44e1

Observation fcb4a1cd-f968-4957-9b8c-652f58199050 · outbound

This paper cites Boosting image captioning with attributes.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Boosting image captioning with attributes

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.565477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.345606Z digest=sha256:4fc77629e305fc7fd53727c107dbf66edb53cf6caa6deda8d144201c32ff6502

Observation 0aee144c-dc01-4e1f-a6ad-c9d1b1f0af0d · outbound

This paper cites Image captioning with semantic attention.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Image captioning with semantic attention

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.550752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.349967Z digest=sha256:3f00acb7ce02b5901437a2fe908b1ca9f2d71bc063cdd7a4c1be4ea5c14e7d5d

Observation 168df444-f91c-4e67-b062-2d06e4af929c · outbound

This paper cites Dual learning for cross-domain image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Dual learning for cross-domain image captioning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.536687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.354148Z digest=sha256:30af92067dde7d2210d0772e97b62ac6b42f69fc9e87bda196196514df8a61e6

Observation 51603e1c-a047-49ea-af49-27d042c7c6c5 · outbound

This paper cites Unpaired image-to-image translation using cycle- consistent adversarial networks.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unpaired image-to-image translation using cycle- consistent adversarial networks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.522176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T11:21:49.358294Z digest=sha256:541b0126fba53dc54ed7a25bb9f75c9eaafe719a057af59e062f49b4f5c4605e

Pith citing papers

No inbound Pith citation observations are available.