Pith. sign in

Paper Citation Record · LEDGER

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings

As of 23 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:1908.09317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.09317 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:21:49.358294Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact2
  • verified fuzzy58
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d68a81e-813b-4fc0-8937-6e11956cb70d · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Spice: Semantic propositional image cap- tion evaluation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.395385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.050800Z digest=sha256:f965d66836c80f1df437330547477e3137508c66df137e93b40c3628876b54a6

Observation 5af452b7-3d16-462c-879f-85ce8327d493 · outbound

This paper cites Guided open vocabulary image captioning with constrained beam search.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Guided open vocabulary image captioning with constrained beam search

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.381155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.055957Z digest=sha256:dd13c693afe1dc2ea9b7eff8dab402471bb9823802972936694a3335799d0d21

Observation 97bd61fd-f0ce-4a27-be01-0d1a96643300 · outbound

This paper cites Partially-supervised image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Partially-supervised image captioning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.366747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.060608Z digest=sha256:9e137256c74930b083df9467118be1bebf3b259a1305d2fd1d5a5a4a71afb11b

Observation cd622e6e-eb10-4e62-9da2-049a44506978 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Bottom-up and top-down attention for image captioning and visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.351769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.065362Z digest=sha256:f1e65d7bc1a72ab5edac74768191b5bd9fd75f1cd2b46476082dd8b625141a9b

Observation 465e539e-eaee-4d12-b385-cb964f062c7c · outbound

This paper cites Women also snowboard: Over- coming bias in captioning models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Women also snowboard: Over- coming bias in captioning models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.336902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.069958Z digest=sha256:f6aaefb4ca2f1086dde336b839e3e9296cc92760075d1f39ff003b0051b46268

Observation cd20f5a7-7f92-487b-9e7c-18306e0943e9 · outbound

This paper cites Deep compositional captioning: Describing novel ob- ject categories without paired training data.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep compositional captioning: Describing novel ob- ject categories without paired training data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.323001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.074636Z digest=sha256:33ac103e068190f30adaa1f66ff35423d6b3c3b5da3ee9de9c92d2909c5e6c8a

Observation 37608bf8-7d38-4074-bf61-dad8e8ef08f0 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Lawrence Zitnick, and Devi Parikh

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.308732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.079192Z digest=sha256:20df73c04d4834cc958336df85bb08534cbeb9cc3b4d4a6056519053fd6b7357

Observation f0696281-d196-425f-ae58-76177ceb549c · outbound

This paper cites Adversarial text generation via feature-mover’s distance.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial text generation via feature-mover’s distance

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.294420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.084024Z digest=sha256:4d8aaa13141980c23804593559a63ab66bfbeca56043ae337cf23d5cd27ff665

Observation a6a209ca-8435-4475-8d61-c1b6eb3e5543 · outbound

This paper cites Show, adapt and tell: Adversarial training of cross-domain image cap- tioner.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, adapt and tell: Adversarial training of cross-domain image cap- tioner

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.280462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.088273Z digest=sha256:4e80428680da1e4088498e06528a0154eea2623732439e0d3d7fe6710c661921

Observation 37b91ab8-d4c5-4820-a2d9-9baf2e576db4 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.092525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.092525Z digest=sha256:7e91a486ef38be81c751f973593b8310ac14e99429cc63f77017bbc68fc0ce3a

Observation 14fb92d0-7223-468e-b608-9afa051ee48e · outbound

This paper cites To- wards diverse and natural image descriptions via a condi- tional GAN.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings To- wards diverse and natural image descriptions via a condi- tional GAN

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.266341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.097343Z digest=sha256:2553f0ccb2759d36082c1aeafd02c6008abe0d96596e872ee8e979e8752c142b

Observation a296666d-260e-4bad-a15c-60fbf54f8b2f · outbound

This paper cites Visual dialog.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual dialog

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.251995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.102402Z digest=sha256:19ae1650b05e82b22ff775c43ea00378830de8c620a3856f245043a2c05b3994

Observation 9b6c5ffe-db60-4622-ad3a-ac17afd8e0ba · outbound

This paper cites Meteor universal: Lan- guage specific translation evaluation for any target language.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Meteor universal: Lan- guage specific translation evaluation for any target language

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.238471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.106878Z digest=sha256:72420a67bf3c394c71b2329bbaa214584013b5d58367bce2897fe1da517860a2

Observation 96ddcc7f-eeb6-4429-9539-e8ece88393e6 · outbound

This paper cites Long-term recurrent convolutional net- works for visual recognition and description.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Long-term recurrent convolutional net- works for visual recognition and description

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.224589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.111212Z digest=sha256:1774cf1acc7f059c7fb704bb7c489ad8fc0a8f67b6d63952a7607d8305535a31

Observation 174b51ff-a2df-4233-a99e-92a9490c6535 · outbound

This paper cites Adversarial Feature Learning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial Feature Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.115213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.115213Z digest=sha256:ac5bcdcdce935a58a9e72a80dd405901bb40e3cdf7d4e2ef8ba1f24209af5cc1

Observation d557ae01-89a5-4f1c-8b2f-3ea5e75fa703 · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.119851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.119851Z digest=sha256:4399026a601006578ca0a61415bcd6a904039695e87a25fa21176e4b3f2b0260

Observation aea4e3e0-1b75-4341-b3a1-40ca4affc4dc · outbound

This paper cites From captions to vi- sual concepts and back.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From captions to vi- sual concepts and back

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.209802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.124345Z digest=sha256:f008fc47952e9f181bfbb7960c46ee9637ed2fb399663b6d5b5395f10ae677d0

Observation 4866dddf-2f31-4059-b0bd-2177e7d0a151 · outbound

This paper cites Unsupervised Image Captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised Image Captioning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:21:49.464577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.128637Z digest=sha256:b0c8a5aba795348c3cc52f1a1fd4bc9bc4952f504212ac0b50ec1325b48166dd

Observation 79f3acfb-a3a4-4087-bf21-c112bc25ec89 · outbound

This paper cites StyleNet: Generating attractive visual captions with styles.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings StyleNet: Generating attractive visual captions with styles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.194138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.133156Z digest=sha256:005be8593cad472735319b26f200017073d957f6581edcd6900e8c940e000958

Observation ce139965-5b8c-46f6-ad4e-724719239783 · outbound

This paper cites Un- paired image captioning by language pivoting.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Un- paired image captioning by language pivoting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.179739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.138236Z digest=sha256:26660911b72a6f3b99c7431eaccf0723c5c24f947e8b6be4acb8345403664540

Observation 5487240a-7548-4ed3-9600-ac34efdb04ac · outbound

This paper cites Improved training of wasserstein gans.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved training of wasserstein gans

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.165547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.142747Z digest=sha256:dcaff86c19949a7ea115a149605680f3950f81628cac810ef6b85a1533a29137

Observation 711de645-aeae-4efe-b938-683a37a07f0c · outbound

This paper cites MSCap: Multi-style image captioning with un- paired stylized text.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings MSCap: Multi-style image captioning with un- paired stylized text

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.151076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.147579Z digest=sha256:9e742a27c10e82ece27f228fd26dba2e5fd7fbe29e019eba520b64bbb337dec9

Observation 7d9848f4-4871-4f17-aedb-d955dd1e28a6 · outbound

This paper cites VizWiz grand challenge: Answering visual questions from blind people.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VizWiz grand challenge: Answering visual questions from blind people

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.134384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.151969Z digest=sha256:25be3f6c2faed304c73a5e9b1ee216b3f3a68a7f0485283e6dca9f718cf3f47d

Observation f1255e8a-a030-44fd-bf18-fc2e7ac68918 · outbound

This paper cites Deep residual learning for image recognition.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep residual learning for image recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.156222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.156222Z digest=sha256:b6bc5a79f63a875499794aec7b46f4456d50bb8156c3d33d29f65a6842fb0fb1

Observation f7e9902d-867a-47e4-acbc-3be5f12ab9bd · outbound

This paper cites Speed/accuracy trade-offs for modern convolutional object detectors.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speed/accuracy trade-offs for modern convolutional object detectors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.109732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.160518Z digest=sha256:1ba3859442636f920972492a3674af5b1023665d962296e947a7bebd26ae1e5d

Observation 7412c2ab-6660-4bc3-b97e-1cbfe57af8b9 · outbound

This paper cites DenseCap: Fully convolutional localization networks for dense caption- ing.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings DenseCap: Fully convolutional localization networks for dense caption- ing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.094595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.164737Z digest=sha256:6e1d786ee196bd071162a1b5b816b7625785a78654dc6f75a91b3dc9d151319a

Observation 13c9505c-4874-4445-b13b-17d8a512eed3 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep visual-semantic align- ments for generating image descriptions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.080284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.168914Z digest=sha256:95eebc37c14ded0cff315e6b9a0970d907ce59901cecff0aea799f63acaa63d5

Observation e921cc49-5e6e-4c28-bfa2-e04e0d2a58bd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.173388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.173388Z digest=sha256:26ad83978ff42e7e2f6c8903d2bf9389d9d74a6ac31b03742ba0d26ca340b92d

Observation 7256ea6d-0897-4078-a360-1a693a8a1913 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.177786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.177786Z digest=sha256:e24826ab0144ee9a4c795e927abf843b49fa3335d473b94f13d56de266f294d8

Observation edbb9db4-cad3-4413-ad80-ab692e553d16 · outbound

This paper cites OpenImages: A public dataset for large-scale multi-label and multi-class image classification.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings OpenImages: A public dataset for large-scale multi-label and multi-class image classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.065623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.182853Z digest=sha256:2e2537da20dd527bf6fbd70ff056ea6e26902a41befe56e9326cc956cee9390d

Observation 140f1fa0-6291-4298-aac8-482c66e61aae · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.050852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.187321Z digest=sha256:c0609462f290a054dbb391af866c9dcbd869fafd88d7f1d19360d76b8f49f4cc

Observation 71e19b4c-db59-4d3a-8318-29c1d03ef217 · outbound

This paper cites From word embeddings to document distances.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From word embeddings to document distances

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.035102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.191699Z digest=sha256:9ee280a8302727b2f9946545de68f4030452b9cd01f04b125cd3dfd87cdd3b81

Observation e639d1fd-85ff-4d48-91aa-73dce18fb0a8 · outbound

This paper cites The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.195941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.195941Z digest=sha256:696ff0c3a27f9a1c9143113f0a9125eba5ca1670a6f6f09f252f2b43f16ddb4b

Observation 3547e4ff-82bd-4d82-913a-d8777c58e2ce · outbound

This paper cites Unsupervised machine translation using monolingual corpora only.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised machine translation using monolingual corpora only

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.019131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.200592Z digest=sha256:f8c38cacceeaeae615c4c79f219bfa6262aff22e06f18a88c54862036dd41acf

Observation 87046abb-bc66-47a2-ad59-99ee07ee2c10 · outbound

This paper cites Phrase-based & neural unsupervised machine translation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Phrase-based & neural unsupervised machine translation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.003670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.204978Z digest=sha256:e38af9d0296eb4b0e03cb01b4697d32e2bdb7b271a6ad0219311ef46c35283a8

Observation 75b16d00-50a4-4480-94a6-2fd4d930ec3a · outbound

This paper cites Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:21:49.400024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.209513Z digest=sha256:e7f87fcde5eb91b654f950d6208c60320e3aaa1451fc6469d25d40f2a4a358a5

Observation 1966479c-1b24-4bb0-9faf-d738c6c39e50 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Rouge: A package for automatic evaluation of summaries

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.989330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.214205Z digest=sha256:d9a9953da3bcdcb58fbd8c38fae2cfcae5cb36e4b52c62f719030f31ccd19753

Observation 11b42c84-79c1-4e7b-ac00-ae6e040add50 · outbound

This paper cites Microsoft COCO: Common objects in context.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Microsoft COCO: Common objects in context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.974238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.218844Z digest=sha256:6e5901beb3be63ce46453d981acf195162e8cbb496debcfa004e013e6fd5d5e3

Observation d0480356-7346-47b4-9fcf-19c9de0e11dc · outbound

This paper cites Teaching machines to describe images via natural language feedback.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Teaching machines to describe images via natural language feedback

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.958902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.223234Z digest=sha256:54df04a38956a16b5cbb1fd00b2aaae65c0834c3dfe1a25b3e185d867b10a980

Observation 5b174ceb-e1e2-44d7-bef1-d72524f35b61 · outbound

This paper cites Improved image captioning via policy gra- dient optimization of spider.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved image captioning via policy gra- dient optimization of spider

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.228016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.228016Z digest=sha256:27e6b75219d28cd3737d48c89876f86a224902a08b19b4ca37bbf0e86ecd6f38

Observation a4a589ba-ea53-45d1-b4cd-00b88a060639 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sen- tinel for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Knowing when to look: Adaptive attention via a visual sen- tinel for image captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.232199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.232199Z digest=sha256:c16d4f865f2a5834577415f2d575d4e7a5c228e115056173ed7a8019f5f6a7c4

Observation 5297d2d2-0ce3-4cd1-a37d-29822afa472b · outbound

This paper cites Neural baby talk.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Neural baby talk

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.925602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.236314Z digest=sha256:b8c2f084ce2d72108700d52406f2c23c1161fb6ccc3f95a1ed1f2b0791889305

Observation 4646014e-da54-4b69-b8b8-d63ddf7abc43 · outbound

This paper cites Visualizing data using t-SNE.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visualizing data using t-SNE

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.911058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.240435Z digest=sha256:96939dd8862f5c4b9a569965ae9c88eb889aef8f41ddbff8240bc3a45876653c

Observation 61847bfc-ef07-40d7-99c1-324254d67622 · outbound

This paper cites The stan- ford CoreNLP natural language processing toolkit.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The stan- ford CoreNLP natural language processing toolkit

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.896734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.245049Z digest=sha256:b63c1b7b2c8c0b8f1d1726d749d3894d6cc68ef260f7c2005fbe69b5e01d445c

Observation c8eec48e-6ba1-4a5d-ad17-e8f0b04ea120 · outbound

This paper cites Learning like a child: Fast novel visual concept learning from sentence descriptions of images.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning like a child: Fast novel visual concept learning from sentence descriptions of images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.883100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.249691Z digest=sha256:f3eedbe52cea75bf04cbad0830c53550e25cb0a8591f162fc4e958a12ca4ac4b

Observation caaacc1c-8c40-455b-8dde-0f7985c42798 · outbound

This paper cites Sem- Style: Learning to generate stylised image captions using unaligned text.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sem- Style: Learning to generate stylised image captions using unaligned text

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.868624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.253865Z digest=sha256:43ed3bd123387b6c5f5310bdffef4f64afa7370c4f067ad4e4faa95ab50306b8

Observation 081451b6-48d3-42cd-8ff9-c1a9c6202780 · outbound

This paper cites Jointly modeling embedding and translation to bridge video and language.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Jointly modeling embedding and translation to bridge video and language

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.853953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.257982Z digest=sha256:3c0ae04e63adf3bb753d595f2c0a61cb534847332e7ea26931355a8977160ada

Observation 210a1364-a49a-4038-a1dd-46f43fbb0cb8 · outbound

This paper cites GloVe: Global vectors for word representation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings GloVe: Global vectors for word representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.839595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.262233Z digest=sha256:d2ac923a1e1a0f73072fbd8db9f08d484dadf823e1d7355fc7d399f7de0e779c

Observation 9d83399a-4384-45d3-8d09-8f5d91b25f1f · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.825483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.267346Z digest=sha256:19e478e387c6e3bdd262bead3e0ef28315340fcd73bc7bf7b5372196b4c05084

Observation b2437259-aa6f-40e0-a48e-16a13e1e9d77 · outbound

This paper cites Self-critical sequence training for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Self-critical sequence training for image captioning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.810801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.271747Z digest=sha256:6321ed007016f1dd47db6fadf10ad8f1ac11bc11fff7ff61f44565bef7f79da1

Observation 08d40f3f-c859-47e4-af41-b5c0fd0e97c8 · outbound

This paper cites ImageNet large scale visual recognition challenge.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings ImageNet large scale visual recognition challenge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.796526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.276815Z digest=sha256:891c7ee8243b89245bc343034cef75ddc1e2a1406ea5046e0e57f54cba1ffb7c

Observation 0fd873de-d811-4278-bf7c-5eb2202282fb · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.781359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.281797Z digest=sha256:c5c9f668d680f8d00726c985682d76a0133e7bfd113a30082dc28c7dae9b770e

Observation 8d405e70-3988-446d-8598-da7516652da1 · outbound

This paper cites Speaking the same language: Matching machine to human captions by adversarial train- ing.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speaking the same language: Matching machine to human captions by adversarial train- ing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.764925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.286027Z digest=sha256:266eeb1c5b597e7ab1915eac0609892f979aaa9340318fba227e75b058f3356d

Observation 21de770f-2a64-44db-be90-6b321b68c7ea · outbound

This paper cites Deforming autoencoders: Unsupervised disentangling of shape and ap- pearance.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deforming autoencoders: Unsupervised disentangling of shape and ap- pearance

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.750435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.290120Z digest=sha256:7c7b508daedb089508e50e2cbae5069ee29c776ceb6e58d39db84078d8dcf90f

Observation 24d7acb7-651e-404a-b054-6fb0e294b2af · outbound

This paper cites Engaging image captioning via per- sonality.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Engaging image captioning via per- sonality

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.736319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.294410Z digest=sha256:6a333f8bfeed516d662aa0ca771303fe9e57217aaa5a07dacebde3a286f93274

Observation eb488959-56e3-4534-83f7-3e5b4a2a5dcc · outbound

This paper cites Towards text generation with adversarially learned neural outlines.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Towards text generation with adversarially learned neural outlines

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.722240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.298686Z digest=sha256:b9614a3c6e4bee088f30550a0e6645b60833f0a6996e26c01536828b4ec1bab3

Observation 3b62f948-9c4f-425c-87a7-e44dc90b407d · outbound

This paper cites Sequence to sequence learning with neural networks.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sequence to sequence learning with neural networks

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.707847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.303007Z digest=sha256:da912a434552e03dc23cf0286b39a335d0d1959548a1b9fe4b993e38128592f0

Observation 2ca248f7-0463-4f3b-be81-6aca77e5b0f4 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Cider: Consensus-based image description evalua- tion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.307627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.307627Z digest=sha256:f4df4962961ba1e3571819d3f54edb10961dc1e73f5e3a08da41fe649a8e6711

Observation 1e69ab5a-de3b-4ad6-abcc-de758b06f2a1 · outbound

This paper cites Captioning images with diverse objects.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Captioning images with diverse objects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.684154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.311808Z digest=sha256:1a3a695c9c23dbadcf385be234206f2e5f543e4407fbd1fbe142e6e1b1959ec6

Observation add7915b-047b-4bb5-9fb3-b86bc1d53de2 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show and tell: A neural image caption gen- erator

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.669643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.316206Z digest=sha256:c91be30f2803f8ba653fe79283b53b0eb8626f1d09a2d6dc4021d96cd21718fe

Observation aba98efa-5dc5-4e67-ac51-3ba198649935 · outbound

This paper cites Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.655861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.320325Z digest=sha256:8aaa7596143cbba1ade51539269da4e3cfb4cc1f6e7b09bb649632201d6069c0

Observation f563f42d-6c22-4729-a433-657eb34d8d69 · outbound

This paper cites Automatic alt-text: Computer-generated image de- scriptions for blind users on a social network service.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Automatic alt-text: Computer-generated image de- scriptions for blind users on a social network service

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.641250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.324482Z digest=sha256:0f22524eaa0fb95e7938b6056ab739a30fb83460509af26c62abb226a2f09803

Observation ae702d51-7242-4aeb-a8c3-04d1e674e811 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.626520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.328768Z digest=sha256:dea70761a3cebfb089266c4c2bb30e7bbcbf1ef824994c340a691e62e1fddaea

Observation 3672b3b3-97c3-459d-8afe-c45b03a8cb18 · outbound

This paper cites Review networks for caption gen- eration.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Review networks for caption gen- eration

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.612477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.332844Z digest=sha256:7df6cff2326f476dd08c7789c8fb97ed6f420d9d9b7ce606592a96edd2980649

Observation 0f0f9291-e594-40ce-a1ad-0e53b8936924 · outbound

This paper cites Incorpo- rating copying mechanism in image captioning for learn- ing novel objects.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Incorpo- rating copying mechanism in image captioning for learn- ing novel objects

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.597659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.337040Z digest=sha256:220035f18df7ce03c56808bc499aaa4413d7dfe48ca905b76142b62044fbb967

Observation d2256260-1f4b-4f0b-ac04-a156ac8daaca · outbound

This paper cites Explor- ing visual relationship for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Explor- ing visual relationship for image captioning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.580762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.341293Z digest=sha256:a776bf7c270ba6f8e6d36fe4f7491f701716a8cad857a5dda12ff689589a2bae

Observation fcb4a1cd-f968-4957-9b8c-652f58199050 · outbound

This paper cites Boosting image captioning with attributes.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Boosting image captioning with attributes

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.565477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.345606Z digest=sha256:4f89b4fb2180dd81dc247b4dbb1ee295b36c09e7e1bfddeb8048e9a62bee4221

Observation 0aee144c-dc01-4e1f-a6ad-c9d1b1f0af0d · outbound

This paper cites Image captioning with semantic attention.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Image captioning with semantic attention

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.550752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.349967Z digest=sha256:3e31636f1ae0eb0a75deca3f0a33dc7344985379cad2c5965f79baefe36b1784

Observation 168df444-f91c-4e67-b062-2d06e4af929c · outbound

This paper cites Dual learning for cross-domain image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Dual learning for cross-domain image captioning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.536687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.354148Z digest=sha256:cbdd8406cac2103f0550e93c47dca918ad1345af92ce72b6de1bac207bef57d1

Observation 51603e1c-a047-49ea-af49-27d042c7c6c5 · outbound

This paper cites Unpaired image-to-image translation using cycle- consistent adversarial networks.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unpaired image-to-image translation using cycle- consistent adversarial networks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.522176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T11:21:49.358294Z digest=sha256:9fa04cd7cb333e29ac2922b6ef2f3b21dca0c804dde02f06d4417eb68d9bf131

Pith citing papers

No inbound Pith citation observations are available.