Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T12:55:23.500601Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:1908.06327.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T12:55:23.500601Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 05673b87-a82d-4201-9bbe-87ef0e12a7fd · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Bottom-up and top-down attention for image captioning and visual question answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0f11676c-966c-4cd1-bea7-c7516614e435 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Lawrence Zitnick, and Devi Parikh
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 89bd2507-b301-4f47-8689-f1c83873daf0 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks A neural probabilistic language model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d5c0a084-e680-44f1-a8a0-4f8024d0e059 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Enriching word vectors with subword infor- mation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ab2d3489-2386-4ddc-9cdd-7fa02a0ec28a · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Temporally grounding natural sentence in video
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fbe12f15-649e-4df4-9c64-e151f6533914 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Query-guided regression network with context policy for phrase grounding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a8703fea-8a77-4915-a0c0-19963663f8b8 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf2fc75b-c4f4-4940-b986-80fedc702e69 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Supervised learning of univer- sal sentence representations from natural language inference data
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d091ef20-4eeb-486d-bde9-d2b8fe74968b · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks ImageNet: A Large-Scale Hierarchical Image Database
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8e844827-57f9-42a7-ab3a-fe05ba3d444a · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d984e53-44a1-4369-bcc9-af0b77a5c762 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Fleet, Jamie Ryan Kiros, and Sanja Fidler
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 31fb7634-5fc7-4cb4-91f8-e7b64e8ee2a1 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Image caption- ing with word level attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 41083366-a5bc-4ef9-8920-655086707a8e · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks From Captions to Visual Concepts and Back
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e725e87-36f0-47ca-82ae-bc55c3c6cc64 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Jauhar, Chris Dyer, Eduard Hovy, and Noah A
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ac24baa3-5866-4b1c-ad39-21dbeb8fb9ee · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Multimodal com- pact bilinear pooling for visual question answering and vi- sual grounding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 566b3310-f018-4f6b-9d72-1b3013340191 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7e5fef83-c6f5-476a-a334-993db1c1c0b4 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks The IAPR TC-12 benchmark – a new evaluation resource for visual information systems, 2006
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f5ae9acc-ce98-472c-be11-1a13d1a8cb4e · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Deep Residual Learning for Image Recognition
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe59da6-8953-4961-8f77-63d0e0c7e864 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Localizing mo- ments in video with natural language
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1e0d6761-3846-4a5e-ad03-c172169b4887 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Discriminative learning of open-vocabulary object retrieval and localization by neg- ative phrase augmentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4a68563e-2733-42f1-ac22-525f153f626d · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning to Reason: End-to-End Module Networks for Visual Question Answering
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4532e726-6089-4cc2-8703-1788be8043f1 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Natural language object retrieval
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bcebb588-d855-4324-80a5-dbd091c40b33 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning semantic concepts and order for image and sentence matching
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3da144ea-c279-41c1-8b3a-de7ec0cc75de · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks ReferItGame: Referring to objects in pho- tographs of natural scenes
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d74c7a17-770d-44ef-a8e3-d27874c6b790 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning image embeddings using convolutional neural networks for improved multi- modal semantics
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2145237e-f086-4e70-b827-b42a8707e770 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Bilin- ear attention networks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9c04970d-3c76-4a08-a3e2-e2b166fef98a · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Fisher vectors derived from hybrid gaussian-laplacian mixture mod- els for image annotation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f5fff47a-0b68-4a8a-82ad-82aa90750d1f · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks What are you talking about? text-to-image coreference
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5ce4735f-a3da-4ec2-b131-0605beffa216 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dbe27d71-5d1d-4697-bf10-f7844820c790 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Dense-captioning events in videos
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 37fb26c6-171b-4fab-90a1-97027e4e11ef · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 687c6c77-9a4f-4a59-9f7d-2f9370603f42 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Combining language and vision with a multimodal skip- gram model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation df411384-814f-44b4-b702-2073aeab690f · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Stacked cross attention for image-text matching
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 386b0541-5907-45e2-a556-457daa4a8424 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks RNN fisher vectors for action recognition and image annotation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8bf4674e-a7d5-46c6-8522-5375798e22e8 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Microsoft COCO: Common objects in context
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a25c6695-f085-4a41-a142-08bf4f197a30 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Temporal modular networks for retrieving complex compositional activities in videos
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8424463c-16d3-46a9-ab08-3e177bb4403b · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Hierarchical question-image co-attention for visual question answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 05ed0127-a0b4-4ec8-9cad-52819f2bcd92 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Packnet: Adding mul- tiple tasks to a single network by iterative pruning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a7c62b-55d9-40ab-a231-f6312e8a1267 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Linguis- tic regularities in continuous space word representations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 26bcd364-e83d-4d8a-acac-ec59d6279433 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b2db0d93-8c36-423a-97a8-ba227b9e1782 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Dual attention networks for multimodal reasoning and matching
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bd9369d1-8666-448b-95d1-542dfb020ae4 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f140c4de-9280-4c0b-bf53-7002cf5ad21d · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7076aa22-6460-4082-9a09-ad8f74d64fa5 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Paige Kordas, M
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cfd22438-7b2a-40fb-a714-30a569c54c0f · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Arun Mallya, Christopher M
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8aac6b13-e821-425a-834e-bee868f96b59 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Revisiting Image-Language Networks for Open-ended Phrase Detection
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 509d86fb-b9fe-4d8c-9a8a-1d5a455f0f75 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Liwei Wang, Chris M
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 945de334-2d37-4b1a-9142-6b0bd7908fb8 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Faster R-CNN: Towards real-time object detection with re- gion proposal networks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e9e5436-225a-44a2-943e-198234c8e4e8 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Grounding of textual phrases in images by reconstruction
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a313310f-a7d4-49ac-a599-00aeea4854a7 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Training region-based object detectors with online hard ex- ample mining
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 99f44b98-3602-4513-a8d5-3422b4343ac7 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Very Deep Convolutional Networks for Large-Scale Image Recognition
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a131315e-8d8b-4b8a-892d-4cfaf4fde43a · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Svetlana Lazebnik, Alex C
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9818656e-5c95-4988-bb4e-1ffbebaae2ef · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Attention is all you need
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation de8cd408-4458-445d-b9f7-a63d6acca6b0 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Order embeddings of images and language
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 009b7f1a-4551-404e-a700-d42edb4c3bf2 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Se- quence to sequence – video to text
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation df279ad1-d9dd-4e1b-b696-b3e876aa03f6 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Show and tell: A neural image caption gen- erator
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ad58edcd-0668-4e98-8407-2d8bcbd3612c · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning Two-Branch Neural Networks for Image-Text Matching Tasks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7db42155-4366-408f-8a05-aa843f10d37d · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Structured matching for phrase lo- calization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e70d0a93-a244-4885-bca8-0a5c581ae3f1 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks R-C3D: Region convolutional 3d network for temporal activity detection
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bd3d6b8c-2b67-456d-96e2-d2d4b6d66c94 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation da96f9cd-3bbe-4139-b593-a2d04f5f36d3 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f94ca81-2d51-44e9-be28-6dcc154fbb02 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6c272654-be03-4108-9cf3-006b9febd61f · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Berg, and Tamara L
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 903b884c-687b-410f-8545-e29bc97917a5 · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Improving lexical embeddings with semantic knowledge
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a62bea2b-db37-4e5c-9800-fd8e0bd79d6b · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Yin and Yang: Balancing and an- swering binary visual questions
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6d7a8d99-aafe-4639-a5cf-039c1f3f53fd · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Deep cross-modal projection learning for image-text matching
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 68d03680-06a0-42a6-8649-2ed7420b9bad · outbound
Language Features Matter: Effective Language Representations for Vision-Language Tasks Datasets Flickr30K [62]
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.