Pith. sign in

Paper Citation Record · LEDGER

Language Features Matter: Effective Language Representations for Vision-Language Tasks

As of 16 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:1908.06327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.06327 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:55:23.500601Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact2
  • verified fuzzy53
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 05673b87-a82d-4201-9bbe-87ef0e12a7fd · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Bottom-up and top-down attention for image captioning and visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.912558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.047017Z digest=sha256:357b79323ac5b49f31e05a2de56a9d88c84a3baeede5105009035eaffabd62cd

Observation 0f11676c-966c-4cd1-bea7-c7516614e435 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.887641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.055821Z digest=sha256:5b9321d017a9d2bdbaf1efe199e9a6cecd4444ee795d1743a93f79fe1e02194d

Observation 89bd2507-b301-4f47-8689-f1c83873daf0 · outbound

This paper cites A neural probabilistic language model.

Language Features Matter: Effective Language Representations for Vision-Language Tasks A neural probabilistic language model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.867887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.062557Z digest=sha256:d729b3e5fc1c0424d7eabada0bfb72cb639f4cbe7d401074ffef1aa45c5d65f6

Observation d5c0a084-e680-44f1-a8a0-4f8024d0e059 · outbound

This paper cites Enriching word vectors with subword infor- mation.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Enriching word vectors with subword infor- mation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.845577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.070702Z digest=sha256:0b468a2a2eeacca012fa481583ea48e3f6f5f66fe7bff95cc9cfd8cc52a9202e

Observation ab2d3489-2386-4ddc-9cdd-7fa02a0ec28a · outbound

This paper cites Temporally grounding natural sentence in video.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Temporally grounding natural sentence in video

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.828470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.079155Z digest=sha256:807773ed965858d784702566991862311f0ebc68a910f81a716b1a080b3f8ca5

Observation fbe12f15-649e-4df4-9c64-e151f6533914 · outbound

This paper cites Query-guided regression network with context policy for phrase grounding.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Query-guided regression network with context policy for phrase grounding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.808939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.090922Z digest=sha256:f3c529463c3ee85be0dec9cc3620ce19ebbc5db30dac600c23ad3f3bc39e5c5b

Observation a8703fea-8a77-4915-a0c0-19963663f8b8 · outbound

This paper cites Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.100124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.100124Z digest=sha256:1a5bcdc3ea94066853464bf60c14f6fabaf2a021d5d313b770a5f1ee7e6e6f9a

Observation bf2fc75b-c4f4-4940-b986-80fedc702e69 · outbound

This paper cites Supervised learning of univer- sal sentence representations from natural language inference data.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Supervised learning of univer- sal sentence representations from natural language inference data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.791142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.112546Z digest=sha256:b0835be5ed714663d5b2483e4371fa17fbb331c0f4a7e5a6392e6fa18e39cf1a

Observation d091ef20-4eeb-486d-bde9-d2b8fe74968b · outbound

This paper cites ImageNet: A Large-Scale Hierarchical Image Database.

Language Features Matter: Effective Language Representations for Vision-Language Tasks ImageNet: A Large-Scale Hierarchical Image Database

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.772027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.119226Z digest=sha256:3031f3efacd76545bccde0fbc89cfb32071b3f6f087233fd334274ace29aaa1b

Observation 8e844827-57f9-42a7-ab3a-fe05ba3d444a · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Language Features Matter: Effective Language Representations for Vision-Language Tasks BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.125805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.125805Z digest=sha256:5ff1741d5b96cd75a18ee919e7441eab6553d74d629664f459d0f037252573df

Observation 1d984e53-44a1-4369-bcc9-af0b77a5c762 · outbound

This paper cites Fleet, Jamie Ryan Kiros, and Sanja Fidler.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Fleet, Jamie Ryan Kiros, and Sanja Fidler

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.754694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.132127Z digest=sha256:501345f0bdd43f01cb1e2c9d6bd98bdee6f52c72f2bf6b17e031fc015d2315ac

Observation 31fb7634-5fc7-4cb4-91f8-e7b64e8ee2a1 · outbound

This paper cites Image caption- ing with word level attention.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Image caption- ing with word level attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.738823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.139969Z digest=sha256:c1df5eaf4f6ee43f25780b6737c2aca707b10220a41839a967eb8dc019f80d15

Observation 41083366-a5bc-4ef9-8920-655086707a8e · outbound

This paper cites From Captions to Visual Concepts and Back.

Language Features Matter: Effective Language Representations for Vision-Language Tasks From Captions to Visual Concepts and Back

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.145294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.145294Z digest=sha256:edc9e870c936b7ed23e99089d8aa176d7ca6622bbc225718898bd3f100eb4741

Observation 1e725e87-36f0-47ca-82ae-bc55c3c6cc64 · outbound

This paper cites Jauhar, Chris Dyer, Eduard Hovy, and Noah A.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Jauhar, Chris Dyer, Eduard Hovy, and Noah A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.722809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.153830Z digest=sha256:b4cf7c437fdf392f4e35609c8f1e74ce4d87d882d2a52e46eae7d30db0b9acbf

Observation ac24baa3-5866-4b1c-ad39-21dbeb8fb9ee · outbound

This paper cites Multimodal com- pact bilinear pooling for visual question answering and vi- sual grounding.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Multimodal com- pact bilinear pooling for visual question answering and vi- sual grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.704458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.161983Z digest=sha256:ea493746c213a26f2dc9d0fbc5752136fc3d0c4cc4e50fc1e65e5a5a59ab8ecd

Observation 566b3310-f018-4f6b-9d72-1b3013340191 · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.684423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.170606Z digest=sha256:0562a3975f06c23f23f04ca0d1ed4318f55b208388b983534ba1f77831beba4b

Observation 7e5fef83-c6f5-476a-a334-993db1c1c0b4 · outbound

This paper cites The IAPR TC-12 benchmark – a new evaluation resource for visual information systems, 2006.

Language Features Matter: Effective Language Representations for Vision-Language Tasks The IAPR TC-12 benchmark – a new evaluation resource for visual information systems, 2006

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.665763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.176625Z digest=sha256:cec4a92a3dc913bf1600e3e379f80060351f0b59694aea6c1d7578314a730ce7

Observation f5ae9acc-ce98-472c-be11-1a13d1a8cb4e · outbound

This paper cites Deep Residual Learning for Image Recognition.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Deep Residual Learning for Image Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.183630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.183630Z digest=sha256:9ebc31ba98f61b8989bc0042c74e84785dbdd3a3a81cffa5127c44eace89cb8f

Observation 1fe59da6-8953-4961-8f77-63d0e0c7e864 · outbound

This paper cites Localizing mo- ments in video with natural language.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Localizing mo- ments in video with natural language

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.644666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.189964Z digest=sha256:75bdac371d05a4a3a707ea9a05c5dec19cbe96fc6e6dded93ccc43dda4d7b82e

Observation 1e0d6761-3846-4a5e-ad03-c172169b4887 · outbound

This paper cites Discriminative learning of open-vocabulary object retrieval and localization by neg- ative phrase augmentation.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Discriminative learning of open-vocabulary object retrieval and localization by neg- ative phrase augmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.626304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.195987Z digest=sha256:87dc133ee40436e327e8236b221859f6a1a7e39df233bbdf899952dd832b196c

Observation 4a68563e-2733-42f1-ac22-525f153f626d · outbound

This paper cites Learning to Reason: End-to-End Module Networks for Visual Question Answering.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning to Reason: End-to-End Module Networks for Visual Question Answering

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:55:23.659179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.201706Z digest=sha256:80394093f72f1f7cc23ec78d2ceaaea0b0cb6146ff1febd5d607cd86404f41ac

Observation 4532e726-6089-4cc2-8703-1788be8043f1 · outbound

This paper cites Natural language object retrieval.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Natural language object retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.603925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.210194Z digest=sha256:89250a206f39bd0c2a5d36e39ce1fcc772e302437ef29561567c7623dfeecfb8

Observation bcebb588-d855-4324-80a5-dbd091c40b33 · outbound

This paper cites Learning semantic concepts and order for image and sentence matching.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning semantic concepts and order for image and sentence matching

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.582211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.219108Z digest=sha256:83ddc7d93d9ddd57481ac42b151173e81c1d5779328ac2088ab7ff2b4e946c4a

Observation 3da144ea-c279-41c1-8b3a-de7ec0cc75de · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

Language Features Matter: Effective Language Representations for Vision-Language Tasks ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.559771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.227241Z digest=sha256:1bece869f05d880e1c212c0c09047a47d68b5761ec6cd4025353b9b020b1be65

Observation d74c7a17-770d-44ef-a8e3-d27874c6b790 · outbound

This paper cites Learning image embeddings using convolutional neural networks for improved multi- modal semantics.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning image embeddings using convolutional neural networks for improved multi- modal semantics

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.537178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.236533Z digest=sha256:9a629c9ee4b325c559bc0bfebb81e965c16b5edf67ab1732acf298c738558f01

Observation 2145237e-f086-4e70-b827-b42a8707e770 · outbound

This paper cites Bilin- ear attention networks.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Bilin- ear attention networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.513298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.245501Z digest=sha256:3b20a7d435cf1182a855595596a9c7f63279f42bf496a988ac314f1c2f25efb2

Observation 9c04970d-3c76-4a08-a3e2-e2b166fef98a · outbound

This paper cites Fisher vectors derived from hybrid gaussian-laplacian mixture mod- els for image annotation.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Fisher vectors derived from hybrid gaussian-laplacian mixture mod- els for image annotation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.493942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.253884Z digest=sha256:6f80e527e2ce05c9b05973a771639609bf2218b20522be38b45caada61451bc0

Observation f5fff47a-0b68-4a8a-82ad-82aa90750d1f · outbound

This paper cites What are you talking about? text-to-image coreference.

Language Features Matter: Effective Language Representations for Vision-Language Tasks What are you talking about? text-to-image coreference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.475489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.269286Z digest=sha256:c5d22fea88e119e02b1636a55aac3d1c115e0cee70252de2b01aaad0a458d71f

Observation 5ce4735f-a3da-4ec2-b131-0605beffa216 · outbound

This paper cites an unresolved cited work.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:55:24.457860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.276136Z digest=sha256:e0c4ec6488495734ce9136c93c86d83f120d873ecb2fffdf29bf209d0255fe4c

Observation dbe27d71-5d1d-4697-bf10-f7844820c790 · outbound

This paper cites Dense-captioning events in videos.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Dense-captioning events in videos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.439975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.286192Z digest=sha256:82745f2ae39c20afda332773fa8fac4b8d8f8b8a81bf040e21b35edd989b5593

Observation 37fb26c6-171b-4fab-90a1-97027e4e11ef · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.420296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.292583Z digest=sha256:38aa9ad906eba763341bd6209ba2a076fba812d95492612a3226c3af95067548

Observation 687c6c77-9a4f-4a59-9f7d-2f9370603f42 · outbound

This paper cites Combining language and vision with a multimodal skip- gram model.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Combining language and vision with a multimodal skip- gram model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.395183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.298583Z digest=sha256:88e3e91661634d113553675011f0703a04e8de04f71dd56426b1185949cb9da0

Observation df411384-814f-44b4-b702-2073aeab690f · outbound

This paper cites Stacked cross attention for image-text matching.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Stacked cross attention for image-text matching

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.375515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.303367Z digest=sha256:826e8384f7e9ba464b090c0cedeb6c8a4d2ffeaad1adae7644c85b97c9aaf3cc

Observation 386b0541-5907-45e2-a556-457daa4a8424 · outbound

This paper cites RNN fisher vectors for action recognition and image annotation.

Language Features Matter: Effective Language Representations for Vision-Language Tasks RNN fisher vectors for action recognition and image annotation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.356331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.309232Z digest=sha256:220324f2a8a9123ba035856812edc25c922555e5f5fe8b8615abc3ee41d8a263

Observation 8bf4674e-a7d5-46c6-8522-5375798e22e8 · outbound

This paper cites Microsoft COCO: Common objects in context.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Microsoft COCO: Common objects in context

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.333341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.315705Z digest=sha256:e5136c2eebf5f9cef2a2b7ad513f2df7aa940bd0904f2d60935a260c9a6a64a4

Observation a25c6695-f085-4a41-a142-08bf4f197a30 · outbound

This paper cites Temporal modular networks for retrieving complex compositional activities in videos.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Temporal modular networks for retrieving complex compositional activities in videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.314501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.321889Z digest=sha256:18cf485fdee04c05279c58bc40fd5888176b70fa7b952393b0507c1a5396ac34

Observation 8424463c-16d3-46a9-ab08-3e177bb4403b · outbound

This paper cites Hierarchical question-image co-attention for visual question answering.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Hierarchical question-image co-attention for visual question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.295145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.328158Z digest=sha256:4e1d48eeee0bd7f2475808785910bcb9d902ca28c12fa2adea346a8c96dbdd86

Observation 05ed0127-a0b4-4ec8-9cad-52819f2bcd92 · outbound

This paper cites Packnet: Adding mul- tiple tasks to a single network by iterative pruning.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Packnet: Adding mul- tiple tasks to a single network by iterative pruning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.334090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.334090Z digest=sha256:aaf13f388179e5d3ee412293d2b6942605a8a7dcf3dfa331a899f4f8a0810de9

Observation 04a7c62b-55d9-40ab-a231-f6312e8a1267 · outbound

This paper cites Linguis- tic regularities in continuous space word representations.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Linguis- tic regularities in continuous space word representations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.254754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.339772Z digest=sha256:6760233edb9942afa15a7377b136d525f1ee7a4031d75467e4c68d4d2344161f

Observation 26bcd364-e83d-4d8a-acac-ec59d6279433 · outbound

This paper cites an unresolved cited work.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:55:24.233287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.347990Z digest=sha256:99add0dffa11c7bc78f63ce9daad28a9d14e48a0f36317cc0b54778d12a731fc

Observation b2db0d93-8c36-423a-97a8-ba227b9e1782 · outbound

This paper cites Dual attention networks for multimodal reasoning and matching.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Dual attention networks for multimodal reasoning and matching

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.213338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.353879Z digest=sha256:092cd1792d2a04e443ec5a32e4befa7891004f2a9ad9e9f8854f0ea9f60450c3

Observation bd9369d1-8666-448b-95d1-542dfb020ae4 · outbound

This paper cites an unresolved cited work.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:55:24.192144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.359422Z digest=sha256:be5a1b5c7cc9d3d846b7a53894ebeadc19012afbab23a0b8a63bafe619612ff7

Observation f140c4de-9280-4c0b-bf53-7002cf5ad21d · outbound

This paper cites Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.167686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.365317Z digest=sha256:5e0f0749c8b30eb43ae80c71536dc6985cd4129defe24088e553068000bf7118

Observation 7076aa22-6460-4082-9a09-ad8f74d64fa5 · outbound

This paper cites Plummer, Paige Kordas, M.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Paige Kordas, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.142941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.370095Z digest=sha256:90f226a414cf4f4ffe746bf7dbca1309feb9d678ffe914cc826bdf1a7c1750b1

Observation cfd22438-7b2a-40fb-a714-30a569c54c0f · outbound

This paper cites Plummer, Arun Mallya, Christopher M.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Arun Mallya, Christopher M

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.123491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.374819Z digest=sha256:2174c1cbcba492ea6dbec9d1f65e9017154f9c83474f292df6ff0c6bd11d8721

Observation 8aac6b13-e821-425a-834e-bee868f96b59 · outbound

This paper cites Revisiting Image-Language Networks for Open-ended Phrase Detection.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Revisiting Image-Language Networks for Open-ended Phrase Detection

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:55:23.630735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.380013Z digest=sha256:5391fc85150f89a791e733f8ec945316aea10de750d77a8cdb4e29e12dd70b10

Observation 509d86fb-b9fe-4d8c-9a8a-1d5a455f0f75 · outbound

This paper cites Plummer, Liwei Wang, Chris M.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Liwei Wang, Chris M

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.094583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.387038Z digest=sha256:5b0dad08874383fcf64a3369dea043caebdb19bdad239fb84b6ecf867dc5a734

Observation 945de334-2d37-4b1a-9142-6b0bd7908fb8 · outbound

This paper cites Faster R-CNN: Towards real-time object detection with re- gion proposal networks.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Faster R-CNN: Towards real-time object detection with re- gion proposal networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.392265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.392265Z digest=sha256:375048356f0a44b98907bf03be80c100852a468e0e5cd8af9b3e3696ce82161d

Observation 6e9e5436-225a-44a2-943e-198234c8e4e8 · outbound

This paper cites Grounding of textual phrases in images by reconstruction.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Grounding of textual phrases in images by reconstruction

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.061774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.397711Z digest=sha256:848e7accddda200ab3b5882c61bc08f19a9c34a0fdebc2ba4f63c6b33508074a

Observation a313310f-a7d4-49ac-a599-00aeea4854a7 · outbound

This paper cites Training region-based object detectors with online hard ex- ample mining.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Training region-based object detectors with online hard ex- ample mining

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.038650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.403258Z digest=sha256:c21a827ed38c3db7010e31d739ad75add777e82a96ace94a6bd1844888715f08

Observation 99f44b98-3602-4513-a8d5-3422b4343ac7 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.409458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.409458Z digest=sha256:e2204c474a6a473924abe8b8e0aa4a4a383c70472554eb963a48167263fbb954

Observation a131315e-8d8b-4b8a-892d-4cfaf4fde43a · outbound

This paper cites Plummer, Svetlana Lazebnik, Alex C.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Svetlana Lazebnik, Alex C

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.020432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.415882Z digest=sha256:03b7080e202f6022eddba3a7ca680ed035931e49772a020946eeaf41ea43c56f

Observation 9818656e-5c95-4988-bb4e-1ffbebaae2ef · outbound

This paper cites Attention is all you need.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Attention is all you need

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:24.002290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.420889Z digest=sha256:3fc1267d8a64b023f549615b7a3a1000b59661449fe6e39777d463ef84ee272e

Observation de8cd408-4458-445d-b9f7-a63d6acca6b0 · outbound

This paper cites Order embeddings of images and language.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Order embeddings of images and language

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.983503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.426819Z digest=sha256:39d2d3f31a47b0e7f4768f22eb4a4fa9a3ed0ba437a03acb4576f9b386394852

Observation 009b7f1a-4551-404e-a700-d42edb4c3bf2 · outbound

This paper cites Se- quence to sequence – video to text.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Se- quence to sequence – video to text

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.966080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.432884Z digest=sha256:ae6faeea9cc94b5b79c75601af554fe520b8cfb1b6fa05d3da23074947717017

Observation df279ad1-d9dd-4e1b-b696-b3e876aa03f6 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Show and tell: A neural image caption gen- erator

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.946461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.438878Z digest=sha256:f2238a97c40915ba52914eb799fb59c3e372429d410cc78edb09099258b509f1

Observation ad58edcd-0668-4e98-8407-2d8bcbd3612c · outbound

This paper cites Learning Two-Branch Neural Networks for Image-Text Matching Tasks.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Learning Two-Branch Neural Networks for Image-Text Matching Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.444026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.444026Z digest=sha256:650b046fe950762d076efdc306cf96a08da24c45b71d31751899d585658dfd5c

Observation 7db42155-4366-408f-8a05-aa843f10d37d · outbound

This paper cites Structured matching for phrase lo- calization.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Structured matching for phrase lo- calization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.929206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.449539Z digest=sha256:0eb9006464e329e608d5f5af6857d78ed157e0e4897baa845f6ce25c13a89317

Observation e70d0a93-a244-4885-bca8-0a5c581ae3f1 · outbound

This paper cites R-C3D: Region convolutional 3d network for temporal activity detection.

Language Features Matter: Effective Language Representations for Vision-Language Tasks R-C3D: Region convolutional 3d network for temporal activity detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.909379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.455396Z digest=sha256:186c3ec21d5c32727757292198b0a730236bb1c259b5f4ca4c512740dc62e8e3

Observation bd3d6b8c-2b67-456d-96e2-d2d4b6d66c94 · outbound

This paper cites Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.887763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.461729Z digest=sha256:a7793371777b9355d23021fca32384f878c1d2b42f8f4d521daa1b20cb02aeb3

Observation da96f9cd-3bbe-4139-b593-a2d04f5f36d3 · outbound

This paper cites Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-14T12:55:23.467869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:55:23.467869Z digest=sha256:0382d4f03c844854508c403b5aaeb04fa77c86c1e0eaea317adc23ab49d6765b

Observation 6f94ca81-2d51-44e9-be28-6dcc154fbb02 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

Language Features Matter: Effective Language Representations for Vision-Language Tasks From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.862689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.473314Z digest=sha256:e2e1dade38d432e830e9a731f8d022fe40e0a432e531a4535eb679ba5dc791d3

Observation 6c272654-be03-4108-9cf3-006b9febd61f · outbound

This paper cites Berg, and Tamara L.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Berg, and Tamara L

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.840153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.478663Z digest=sha256:5bf5ab7eaf1a2907718eea9979af236f82b286ea129fdaf41dc6ed19c3670fb0

Observation 903b884c-687b-410f-8545-e29bc97917a5 · outbound

This paper cites Improving lexical embeddings with semantic knowledge.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Improving lexical embeddings with semantic knowledge

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.821753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.483556Z digest=sha256:1d043729f871d2b7a3d7358074f56194d843be94e999097b9c07c0dee3510f7a

Observation a62bea2b-db37-4e5c-9800-fd8e0bd79d6b · outbound

This paper cites Yin and Yang: Balancing and an- swering binary visual questions.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Yin and Yang: Balancing and an- swering binary visual questions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.802530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.489209Z digest=sha256:94f0e46323e08cc06ac835e227d295bacbdd823772a265f0944992f2d9af8c31

Observation 6d7a8d99-aafe-4639-a5cf-039c1f3f53fd · outbound

This paper cites Deep cross-modal projection learning for image-text matching.

Language Features Matter: Effective Language Representations for Vision-Language Tasks Deep cross-modal projection learning for image-text matching

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.778533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.495332Z digest=sha256:c9bc4c5a6d9d92f1cecfc0282e7aec7f62a4100c75c33b6adde5e6ebe1352ba0

Observation 68d03680-06a0-42a6-8649-2ed7420b9bad · outbound

This paper cites Datasets Flickr30K [62].

Language Features Matter: Effective Language Representations for Vision-Language Tasks Datasets Flickr30K [62]

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:55:23.751362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:55:23.500601Z digest=sha256:349e47ae511b94aa9aa8e0126183d10153e7806b220b9aacbc44e568a191e7f9

Pith citing papers

No inbound Pith citation observations are available.