Pith. sign in

Paper Citation Record · LEDGER

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 29 inbound Pith citation observations for arXiv:1908.02265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.02265 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:54:41.910373Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:42.788616Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1675
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b8fae9c0-6b9c-4390-b0be-a5f30ffae3f0 · outbound

This paper cites an unresolved cited work.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:54:42.527937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.715981Z digest=sha256:262a09887d3815816c74bb696f4234b584ec07e6c7a2831502255f59133a52cf

Observation 1fa016a7-422d-4cba-8bd5-0e8e5ff667e9 · outbound

This paper cites an unresolved cited work.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.721129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.721129Z digest=sha256:c76c2faa3ba92626b23d72cae0815af8df03c92de28b69be13aaca7623c81873

Observation 405e9517-e5a1-4f83-a976-2fb3e3f4af00 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Lawrence Zitnick, and Devi Parikh

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.505305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.725405Z digest=sha256:dad48cbe71f70dc1a70c521e4c1cb4a50050137d073963c3dcadcab8cbd0c62d

Observation 7a95221d-8e15-4709-b46a-7c8810a23519 · outbound

This paper cites an unresolved cited work.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:54:42.491585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.730077Z digest=sha256:11b5fb8170b42a52aebc089474b9e11baad2301563faf5ac254b914cf64ca176

Observation 3c67ad70-a369-49f9-99d6-58153b0cd641 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.734742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.734742Z digest=sha256:472f3e38ac8b2b9afc125e67baef2737a11d0461a34fb604f02f2a78d51a1a45

Observation e75e8e42-e1ab-4f93-aa03-248e64581f24 · outbound

This paper cites foil it! find one mismatch between image and language caption.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks foil it! find one mismatch between image and language caption

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.477685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.739613Z digest=sha256:303d6f8e1ab98ae8dbd9603e076ad6c9e3e052c7bc64342c939690fdfd3ddea0

Observation 27ab7fd0-c77c-4ec6-93ce-0865b1f768ca · outbound

This paper cites Embodied Question Answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Embodied Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.744340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.744340Z digest=sha256:56247ebf7df59b04a2b7b97fe8677c64777dacb5d86328bfe2a172c3861aadfa

Observation 3589cdaf-866d-4745-9e10-7d0d480c7bb9 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.455218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.748498Z digest=sha256:633cfc4daf554c13f7b0961a31aacaa2665536bef3451d166db6b897a02223c3

Observation d6e2e915-46ed-455d-947e-f9c7b94ce0ff · outbound

This paper cites Don’t just assume; look and answer: Overcoming priors for visual question answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Don’t just assume; look and answer: Overcoming priors for visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.440601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.752931Z digest=sha256:6a1fca3ff853abe12ff3f4c4642834e3242cf8217f3084099c89e0802fd575bf

Observation d98e4f41-c2bc-4011-8377-39a3080dbbd4 · outbound

This paper cites nocaps: novel object captioning at scale.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks nocaps: novel object captioning at scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:54:42.048363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.757219Z digest=sha256:96d6f723b4f15d3bc85c80a0b96dc528f8b7ec432db1e4ce7d7a2e79fc5ab706

Observation 6c76b031-dcbe-45a3-a7f7-899775b61797 · outbound

This paper cites Deep residual learning for image recognition.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Deep residual learning for image recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.761689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.761689Z digest=sha256:b05f873c27bef127c29e5c2f7de095097c486b8ea22bbf3a50573489d8385dbf

Observation 729e08fb-c372-4219-8db9-5453e054af65 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.765854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.765854Z digest=sha256:24b69fce36c2211794a47318febb983655e18f404bbf5686c7ffc6281334c52f

Observation bede3cb3-c9f0-4199-8aea-471431a0a59e · outbound

This paper cites Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.417597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.770843Z digest=sha256:45cbf1fd82bb69dca2a8523d11f215a2825c84bedf6e72c854b9f540747982ad

Observation c4977e9b-8f1f-483f-8838-3a155826a247 · outbound

This paper cites Improving language understanding with unsupervised learning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Improving language understanding with unsupervised learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.403348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.775087Z digest=sha256:a8cd15bb41be6e0dbf3b35c48c5977819c5d60f480367aecd92c3c0fd9fb416c

Observation fd8d8339-eb10-4cec-a6a3-60aa5660ff73 · outbound

This paper cites Berg, and Li Fei-Fei.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Berg, and Li Fei-Fei

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.779434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.779434Z digest=sha256:336fc636702edaf0c2ddd4fa9e7911658fbaa3cd563d0eec2e7c6d05b61ba113

Observation 44c8df1b-c7ff-47d1-ae94-4e94ac6e09fd · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.783733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.783733Z digest=sha256:5ce8091f5fe5b592fa341a79da685b0cf986af5483f9e583775b9b6abf1b3018

Observation 4a931e69-6555-4cb6-bef2-a8d7bb5b2719 · outbound

This paper cites Aligning books and movies: Towards story-like visual explanations by watching movies and reading books.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Aligning books and movies: Towards story-like visual explanations by watching movies and reading books

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.379199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.788156Z digest=sha256:de47486337300d7f4d3267ab31ba422558e209ea5d1f9cc28687cafc2948655a

Observation cc52f0f3-3bb0-4761-8ec2-e57fad484d82 · outbound

This paper cites URL https://en.wikipedia.org/.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks URL https://en.wikipedia.org/

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.364622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.792364Z digest=sha256:96e178038c735d855da7e0c6a4dd23023fa7475f2807a8334d5ced9d3f5a8410

Observation 3d237df6-a9d9-4524-b513-dd37c38d4592 · outbound

This paper cites One billion word benchmark for measuring progress in statistical language modeling.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks One billion word benchmark for measuring progress in statistical language modeling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.349631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.796867Z digest=sha256:be0cd59c762f9dc26d207f0ec14d90ea3dc9f9f07f1c34490e8bb77e9e8e83c7

Observation ea437fef-cd26-4a3b-8feb-dfc3f86daa18 · outbound

This paper cites Colorization as a proxy task for visual understanding.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Colorization as a proxy task for visual understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.335581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.801205Z digest=sha256:571be01f8f30c21fa6c63626d164c3f8b7ee987354eddbc0a148b1de5a1a69c0

Observation 18dcce10-bc79-4977-9c3d-3252c8e6af42 · outbound

This paper cites Shapecodes: self-supervised feature learning by lifting views to viewgrids.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Shapecodes: self-supervised feature learning by lifting views to viewgrids

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.321415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.806094Z digest=sha256:85da395bf97229539f83a80e44b75dacc0bb72d55fb6fbcca82cf09ce23ac6b3

Observation dd8d5424-451d-49a5-b858-18ead5c46d8a · outbound

This paper cites Look, listen and learn.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Look, listen and learn

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.307014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.810443Z digest=sha256:a03b909e7307beda148eaab97674267ac41af2e5f64ae49915386dac2a5a5dc7

Observation 7f047bb8-390d-4482-904e-be35fe57c304 · outbound

This paper cites Learning features by watching objects move.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Learning features by watching objects move

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.290416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.814544Z digest=sha256:a1991bb9071e23d7a64d3472689698449e4e382e5a21a805ee6783b11a85d94a

Observation 99e38133-ad65-4215-9531-35d2080f0e34 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.818952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.818952Z digest=sha256:90c27b68c8c3b5fed9d58c7707a9f710ea14f7380d43d273b8684b154d22a595

Observation 5cebbe19-81d2-4467-a158-ebb002609938 · outbound

This paper cites From recognition to cognition: Visual commonsense reasoning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks From recognition to cognition: Visual commonsense reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.822905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.822905Z digest=sha256:59bebbe6a252edd604437601b5685cdf6d6e81a02fa6e29fdf57a60aa072e782

Observation cd8edebe-f4bb-49fe-a9e0-129127aa1104 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.827134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.827134Z digest=sha256:440658b1f637f6447219feae1da863b6ec42bb239211487c7dbb7819f035c060

Observation aecd10d1-5847-46e2-9aa3-ba19f2aa93cd · outbound

This paper cites Attention is all you need.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Attention is all you need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.831210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.831210Z digest=sha256:7e7cf6c2856a191902f40138c71974bb00a9fa499376ee157b0b1f4c356e8726

Observation ac8f1a2f-9a10-47d5-bc4a-d0afdfe82cf4 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.835507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.835507Z digest=sha256:a044fe56b78b8ccd4e1c9a5622421a4ad896103201302bd9055a15d22a0f42d1

Observation 0e96ad61-628f-4e10-8e03-2e6cf31489c0 · outbound

This paper cites VideoBERT: A Joint Model for Video and Language Representation Learning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks VideoBERT: A Joint Model for Video and Language Representation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.840774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.840774Z digest=sha256:65c30196a972391f75a7a426c90e649d24f4c6b7432ffb98316d38b1ca29505a

Observation 14aab524-892e-41c8-affd-f3078120b476 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Bottom-up and top-down attention for image captioning and visual question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.845410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.845410Z digest=sha256:7bb11fc07d721e1db1fcfa605011d3668025be31fdd13e0eab4d5c9c3b8c85a6

Observation 1fb83a2d-8745-4d42-9ed8-88462481760e · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.230555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.849617Z digest=sha256:601fc7af0159d8bdcfe816163b36129b9ce0a7e564bb0a151c7107920a568c82

Observation e857b7d4-7934-4897-8e8c-1273306c134b · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Referitgame: Referring to objects in photographs of natural scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.216875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.853703Z digest=sha256:b51481fb9543944fca7c9ea0926ea77aef4e00fed7a997699c9237b7597deb67

Observation 7dad33ac-6be7-455a-9c84-8837ddbc74ae · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Mattnet: Modular attention network for referring expression comprehension

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.202206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.858106Z digest=sha256:aa628bef048e8deb9cdd7415f7cf4c5f03d8feb57dac284ea0904272c6945f50

Observation 6a66df2f-075a-4e3c-b7ba-c46c540614d0 · outbound

This paper cites Mask r-cnn.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Mask r-cnn

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.186388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.862141Z digest=sha256:261d7f62ee94f5cdd682291f025b2573c055b702fdfb922500a6666d95a157cc

Observation 7e59628b-d5b8-49b3-8cee-a0b6dfe14169 · outbound

This paper cites Stacked cross attention for image-text matching.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Stacked cross attention for image-text matching

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.172389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.866210Z digest=sha256:bbb1958bd43306eb36fed66ac0519b7f4545a023642ce2538dd5f072ef6cd696

Observation c77166b2-c5d1-4e60-9b1d-7334467c6ed2 · outbound

This paper cites Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.870680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.870680Z digest=sha256:2c423971818c23cb63f5cedfc9570d4addf3b0d2fb362869e4fadca448f303f8

Observation 411ea18a-47e1-432f-b2df-a10478bbd2db · outbound

This paper cites BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.875322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.875322Z digest=sha256:d5996de900a13ab4af8e8d0bbf8216897e4daf683d5d1a27203da04f9962fd13

Observation a1486f96-6943-4e04-b33e-8eddb4b02df5 · outbound

This paper cites Unsupervised visual representation learning by context prediction.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unsupervised visual representation learning by context prediction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.156813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.879689Z digest=sha256:e26ab3575bee9c3c98236c8577d9f50d57c71a06b2900760465dd974540c74ae

Observation 71a9d6fa-5b50-4038-8fb8-8bb3039410a6 · outbound

This paper cites Colorful image colorization.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Colorful image colorization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.142864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.884569Z digest=sha256:d355fdc81fae48beef9f3d96e443e30383a3118ce3bb166521c9a0560d8279cc

Observation e46e166b-584a-4b27-81d6-e61180a2d4b8 · outbound

This paper cites Discriminative unsupervised feature learning with exemplar convolutional neural networks.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Discriminative unsupervised feature learning with exemplar convolutional neural networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.128503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.888833Z digest=sha256:e493f6dabf8ec3aaa893b00d7ac21cdba7e08f66e9808d0de5031ab473d379cb

Observation 6be4d9c9-cc7c-43c8-8403-8b064124c98c · outbound

This paper cites Context encoders: Feature learning by inpainting.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Context encoders: Feature learning by inpainting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.892878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.892878Z digest=sha256:b6df591da33392ed8c03e8dceebe39bc6c4575e911f9ef65dc28f7cfe4fe250a

Observation 0e43164d-2b37-4ba6-9ef9-e83ad7cfeaa0 · outbound

This paper cites Learning image representations tied to ego-motion.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Learning image representations tied to ego-motion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.105716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.897142Z digest=sha256:0a082120eb17a0e52a62a4be3dff561e6afbe0af4c3aeaf42a26ef1765f52261

Observation 90800b48-2a7b-473a-ae9a-2bc9413ba1a5 · outbound

This paper cites Shuffle and learn: unsupervised learning using temporal order verification.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Shuffle and learn: unsupervised learning using temporal order verification

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.091551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.901553Z digest=sha256:238020ff67a153ab29af809bd082e59aa37f483bd99ca71c15e30a3ecbcf35f4

Observation 1418ecdb-7534-4fc7-af88-9bd735f8adba · outbound

This paper cites Cross-lingual Language Model Pretraining.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Cross-lingual Language Model Pretraining

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.905995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.905995Z digest=sha256:a644fb27aedb5a3c25a653a3343242bf599be48dce156820b30cf62f849dbb7f

Observation bfcc0816-f708-4cf1-88f2-24e2ba26c6a1 · outbound

This paper cites Courville.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Courville

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.077360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T14:54:41.910373Z digest=sha256:7b94cf8010d47a187c25ca9c6c944e69165572116af4f21f892769f26e3abb06

Pith citing papers

Observation ff96a687-5a8e-4fbb-971e-7ba9a2fcfec8 · inbound

VisualBERT: A Simple and Performant Baseline for Vision and Language cites this paper.

VisualBERT: A Simple and Performant Baseline for Vision and Language ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:59:37.786807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T23:59:37.636404Z digest=sha256:6a75b4f9cfdc1a595c4b786b7ae5b72b7b274a936920a338375e8d0de62bbb46

Observation fe831b02-d01a-4554-b47e-7eee89b55662 · inbound

Multi-modality Latent Interaction Network for Visual Question Answering cites this paper.

Multi-modality Latent Interaction Network for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.637749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.637749Z digest=sha256:185b3584d9c18789b26c4423a7f4104fa79b478f194ff0b3c3108c0de6284ff5

Observation 22ca41b3-de21-4c32-925f-6933a0f7f51a · inbound

Fusion of Detected Objects in Text for Visual Question Answering cites this paper.

Fusion of Detected Objects in Text for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.367264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.367264Z digest=sha256:5a1d2cd375e1aebcbe22aeb56868fc21f69761fe628fe580eb1305f8e9a2e189

Observation 0ee6b0e0-0cef-45be-b6ad-9436b121f399 · inbound

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training cites this paper.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.266000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.266000Z digest=sha256:e8ef28bf25722407d909d6300190ca54f4c84eec0cc8add0efd3c127464204c5

Observation 18ccb43a-3cd4-4824-9e2b-eec14122377e · inbound

LXMERT: Learning Cross-Modality Encoder Representations from Transformers cites this paper.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T12:22:25.590947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:22:25.590947Z digest=sha256:6d95258ced177b3231f0e551d7b387477e39064ce2a254d8616ece2597c86208

Observation ea78c704-dae8-4dad-b4a5-705cab9935d2 · inbound

VL-BERT: Pre-training of Generic Visual-Linguistic Representations cites this paper.

VL-BERT: Pre-training of Generic Visual-Linguistic Representations ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:19.107745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:42:19.107745Z digest=sha256:62d0486b157f51c38c8214e36c5cb079892b27eaa50421ab5549972553a1dc6a

Observation dde256fc-9f6d-475a-9589-44e54120cb74 · inbound

Text and Code Embeddings by Contrastive Pre-Training cites this paper.

Text and Code Embeddings by Contrastive Pre-Training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:24:12.008090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T19:24:11.907204Z digest=sha256:fcab6372f8770cce854e331342a7f4b173eb3f9ceecf10194607540318ab1c66

Observation 2afcd137-29e2-48ca-ab18-d7b7ebbb7240 · inbound

A Comprehensive Survey on Visual Question Answering Datasets and Algorithms cites this paper.

A Comprehensive Survey on Visual Question Answering Datasets and Algorithms ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T18:54:50.223310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:54:50.223310Z digest=sha256:6cca62e4b28255f979ebb255eb9afc60b85d407c2905a0b89dfe965dd69b08e6

Observation 33a53c64-c86b-403a-a2be-a5762bb1a7d3 · inbound

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models cites this paper.

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:52:57.089001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:52:57.089001Z digest=sha256:9212a5bb37639b1988d9ba843ef7bb9788ec4110b091fa24aa10029b0ece0b84

Observation f9d4fabe-2f81-400e-8c38-39e7c695f6e8 · inbound

SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment cites this paper.

SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:31:54.596220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:31:54.596220Z digest=sha256:c5b63baeb51a4ce6463b7dafc17903621bfce9545b86cebd213b21a46305e517

Observation c76cf31f-1383-49d6-8c21-87bd74ed90a1 · inbound

Multimodal Multihop Source Retrieval for Web Question Answering cites this paper.

Multimodal Multihop Source Retrieval for Web Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:24.403171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:24.403171Z digest=sha256:99b5ced8afff54754f585928a535e99559b2fd827d85fd5038d0853ac42759a8

Observation d91dae2a-3456-4a2f-ad03-b7eae883678c · inbound

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models cites this paper.

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:47:58.978874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:47:58.978874Z digest=sha256:f1859015caa9416cc0ee1c53f8eac8ab9bd1b7a77e0a1cb0d1bef1555732a6bd

Observation 883a7e11-56e0-4ad0-bde5-c29ede8dd36a · inbound

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering cites this paper.

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:51:53.740176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:51:53.740176Z digest=sha256:5202f6123466a58141b5a292737ae96d62e3a271d4f69751d074aa13d446e5e7

Observation 5b875199-5513-4653-99ad-c25625dfa6a5 · inbound

Performance Analysis of Traditional VQA Models Under Limited Computational Resources cites this paper.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.323443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.323443Z digest=sha256:4842e12ea2e7a92f76307ed10be10740dd439ceb91dbf439862aef13ff065403

Observation 87ea1a9c-bf48-4b70-a47a-df80579b30b7 · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.315864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.315864Z digest=sha256:cb21258c7563abee82f0a856ebbfa2c7fcdd2b60741dfe1ffd675f93c9a24275

Observation 2f6bb367-4920-4ce2-9860-cea82207a6a0 · inbound

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI cites this paper.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.788616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.788616Z digest=sha256:29674edd5fac019b7afa26b8d86fa705ba33a086e901dabf2d9c827301e88b4f

Observation 16d433d1-594f-41b7-90f4-22b6da9fb43c · inbound

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models cites this paper.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.198790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.198790Z digest=sha256:f321d654b015097c49c9c6c06559350309bb65d8f755b19b594867704ca9cb49

Observation 18d9656d-e6ea-45f0-9bac-d593c441cf48 · inbound

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval cites this paper.

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:40.942765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:40.942765Z digest=sha256:7fd51fffd2580f357ea45d8a67927b0f9d63db30ff48be5103e868c869c97a2c

Observation 8519c36c-eef2-429e-aa49-9d5a12ae840e · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.773647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.773647Z digest=sha256:890c76f55ea94980e9db8f9f32c5d68d0bbaa1d3b01ac599f15c52fa5cb030a3

Observation 18a56846-eb1f-4f9e-8a8b-53be205dc05b · inbound

On the Resilience of Underwater Semantic Wireless Communications cites this paper.

On the Resilience of Underwater Semantic Wireless Communications ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:13.121047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:13.121047Z digest=sha256:78ea56e67b7801e6c2a42ad94302f45d204b6db135b3733ff88bf29709028b26

Observation 233de5ba-3820-4e13-b61c-6c4c11479083 · inbound

Representations in vision and language converge in a shared, multidimensional space of perceived similarities cites this paper.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 6241

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.540405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.540405Z digest=sha256:db29c41e138355b978e88557812f5dcfb61c156e5c540efac3dc8d5c62829c5b

Observation 691bee43-3208-4384-888e-71ed7341d524 · inbound

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation cites this paper.

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:42:20.512268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:42:20.512268Z digest=sha256:6ecc8423428305b26494a5639ed6c02f22ff05239e272d26980401d5e376deec

Observation e585f74d-cd12-4a2c-a89d-fea58f28007f · inbound

From Image Captioning to Visual Storytelling cites this paper.

From Image Captioning to Visual Storytelling ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:39.329615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:39.329615Z digest=sha256:46d5514f354a327172f0e2347d95ad9084723ce760a00484023f80690d37ed9b

Observation 96d59a51-e9ae-4059-99ad-9723eee7f085 · inbound

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images cites this paper.

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:55.502242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:49:25.227568Z digest=sha256:e5df2bfece4cb00b25c1dbb534a3e33057c3c9e226c15da82bbaf64e75fc0225

Observation 09885d11-81f9-41f3-bb38-1cec11183d43 · inbound

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments cites this paper.

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:44:57.769037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:40:33.082748Z digest=sha256:31861be186aafb1fe072772d9e4c01035f4526909437462f86c1c51b4e29264f

Observation 4dfd39e3-fc51-48b1-9eb4-04611f821c0f · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.805103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:94eaf60d0568b4d0fd8b1ab6709c0310655f665ca03b96168a3f53aa819d66f5

Observation 878188b0-7386-4b73-9b97-e531469b0240 · inbound

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction cites this paper.

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:58:54.073345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T01:56:35.760918Z digest=sha256:26598d409f1d1ca17ab113f7b99e0b635de91628af2e6529d4c7f45b33213f11

Observation e1deed52-e357-4079-9b7e-30134d9c6f08 · inbound

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset cites this paper.

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:33:06.216866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T04:32:39.678440Z digest=sha256:02683d6cc58f927b60f91e4483224d6d2bbc6e1abf8da0c8735f4dc42dd6c112

Observation c2c49c0a-a86a-4b2d-82cd-01515858c4fb · inbound

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery cites this paper.

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-12T00:26:30.762901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:26:30.762901Z digest=sha256:eaf0fd0d84437c055214caf602761c36ed1909400129fda0c049654ae44477c0