Pith. sign in

Paper Citation Record · LEDGER

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

As of 18 August 2026, this Paper Citation Record lists 100 of 143 outbound references and 0 inbound Pith citation observations for arXiv:2506.11515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11515 v1

Coverage vector

measured 100 of 143 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:08:48.226962Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 143 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b787420d-6acc-4f61-a52a-79299a9128bc · outbound

This paper cites ManagerTower: Aggregating the insights of uni-modal experts for vision-language representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ManagerTower: Aggregating the insights of uni-modal experts for vision-language representation learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.096271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.096271Z digest=sha256:23ab630e703b7fb1a7a6dd6e6b4bf0dceb1ead5e6efc605088a5b1ccdc44e1ac

Observation 55f9426a-ef5c-4aa8-b023-99165f154c40 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.202223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.202223Z digest=sha256:f348c51d31c6d4472e0e10c82b77cb51635fb319a40a1798f4f3cbba42a31ca4

Observation ebc1cac8-cde5-463f-b1a8-406c2422b85d · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.416327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.416327Z digest=sha256:a950bf74ab0545f53d785d4c2b8f89da20863dc9db9314f717c57386f16c3c0f

Observation 3a06e115-c7a9-4c85-ab83-083584a3d28c · outbound

This paper cites A corpus for reasoning about natural language grounded in photographs,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A corpus for reasoning about natural language grounded in photographs,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.710099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.710099Z digest=sha256:7cfa7bb5a483b0c71ebcd3a7489845f61b6a08c575ae4b6a4225b5019ce6ed62

Observation 62e00204-edb0-4c9f-8b1e-48ec10634540 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.832464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.832464Z digest=sha256:39c8ded415060fedd35704d1358cee515b1518b86677b0dc0713c1d4cf33da20

Observation 4deea2d7-3081-4629-977c-7375b2fc0ce8 · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An empirical study of training end-to-end vision-and-language transformers,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:33.933376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:33.933376Z digest=sha256:db7d1a2d89308cdc05a3c7e2eca86020ff01d3f9b59ad38b565c1e033323077b

Observation 6b114773-63c4-4e72-bb35-62f12f64839e · outbound

This paper cites Bridgetower: Building bridges between encoders in vision-language representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Bridgetower: Building bridges between encoders in vision-language representation learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.047268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.047268Z digest=sha256:c835904dd89861c073dbf481d64923b8667c7da22c6623b1687c9a935c0e998e

Observation f1d71738-f08d-4cd5-9398-f9bf73b4f60b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning transferable visual models from natural language supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.138970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.138970Z digest=sha256:63e24611fd8804329f792aa1cde27de542c69914dee81de063a6bc15e7006e0c

Observation 5a635d12-cd8e-4d0e-a1e3-a20123271c14 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.269423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.269423Z digest=sha256:697bccee8050e589cd6c0ba15d845868c9ad708907a642439ad098e074fdf590

Observation 1497da84-d879-4b24-a45b-d1310f0bc138 · outbound

This paper cites Learning deep transformer models for machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learning deep transformer models for machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.396539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.396539Z digest=sha256:0af4bf537c1af37a6e4cb3ab9df553ff12292d2f4637d5458c937eab384942b5

Observation 3f8768f4-df16-48b4-ae99-fd24685256d8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.510959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.510959Z digest=sha256:08e7d0b7763152beeb2edc9c9025adf4cb2e0764341fd4f8f995d6030f5f6539

Observation 17f631e1-3a02-48f3-ad3c-0edf39ec72bb · outbound

This paper cites Llava- next: Improved reasoning, ocr, and world knowledge,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava- next: Improved reasoning, ocr, and world knowledge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.681245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.681245Z digest=sha256:d44cd86d79a56509fb37ae7db88b7e18ebd95a8ccfc2d75e65c141f5b79e6d4e

Observation 2f78c9a6-e55e-4637-a4a3-8d5d209ae51f · outbound

This paper cites How much can CLIP benefit vision-and-language tasks?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs How much can CLIP benefit vision-and-language tasks?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:34.848451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:34.848451Z digest=sha256:c5a09d80b4c85f07e091c3bfbc0e7d8bbaa404fe89b7e830f30a02ba79ffc202

Observation fe89ca05-e8d3-472f-ae87-bf3686995f4a · outbound

This paper cites UNIMO-2: End-to-end unified vision-language grounded learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO-2: End-to-end unified vision-language grounded learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.012205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.012205Z digest=sha256:2111088fc1de3f41b97413ef143434bf2ebf9293d3e0e01ca40304061b51702b

Observation 83e54d8e-4a8a-4f7b-9af6-5cad87292bb6 · outbound

This paper cites Neural machine translation of rare words with subword units,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation of rare words with subword units,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.146927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.146927Z digest=sha256:316bcbaf8ae65431d4658305a9985730dc3c565189d815c2c1bcb03472871a5c

Observation c02c04b3-7ca1-4bca-b887-8375e818dd1f · outbound

This paper cites Language models are unsupervised multitask learners,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.256696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.256696Z digest=sha256:d767e9e355207fc8ff5333f1abc7d91744084bc2595257050eaa58330e244c40

Observation db02e517-fb23-4300-a9be-3426bf0d409d · outbound

This paper cites Attention is all you need,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention is all you need,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.432776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.432776Z digest=sha256:7354bcd51b6e5683fed50d528357350a875b3db4d2982b232cd469a8f0981e4f

Observation 2aa3a712-7ab3-44c4-a016-f21eb83b3039 · outbound

This paper cites Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.541595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.541595Z digest=sha256:59460693ec98186e9c8325c76cb61e2719951367f0e9fb7399603e035a4bb457

Observation 127b0b7e-b728-4454-9243-9d426aeea31c · outbound

This paper cites Multi-layer representation fusion for neural machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multi-layer representation fusion for neural machine translation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.664720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.664720Z digest=sha256:2ddede6917a2754d459e7c423591bb5a8b421648fdff8a1e729306cc26f3a9d1

Observation 6e228351-5890-459e-a8cb-e664beffcaa7 · outbound

This paper cites Multiscale collaborative deep models for neural machine translation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Multiscale collaborative deep models for neural machine translation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.819400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.819400Z digest=sha256:fa4b673ce2e642b21dead003c7f0c96935a9abc5fdca4f1ee51bc68a8daffb26

Observation 168de691-262c-4021-983a-4525d727306a · outbound

This paper cites Layer Normalization.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Layer Normalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:35.926027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:35.926027Z digest=sha256:f0403908b2e4d032025cfe7ae5daaa2bd10f448635b7cc2749d9fdbe5a0cbd2c

Observation d3f72292-993e-42a2-8b69-78960073aa47 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.069126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.069126Z digest=sha256:15bc508f10f6507b05007e5027247a33bf1e4b8381d3030ae7c2f34301cccf10

Observation a6364764-2666-4314-a977-3c149ba81cd8 · outbound

This paper cites Decoupled weight decay regularization,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Decoupled weight decay regularization,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.245435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.245435Z digest=sha256:b0ebae2065126c54c365c851b176be0c471eedfea41275831a8ca1dddda1e51d

Observation 13501472-651d-4356-b2bc-8d4511f2f927 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.380900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.380900Z digest=sha256:7ec0d479b56546072742b6934817b0a668d2ddd58b0eca05cd44653cffe1c35d

Observation a8d542ec-e519-46e3-a9d0-bbe5db76fa07 · outbound

This paper cites Uniter: Universal image-text representation learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Uniter: Universal image-text representation learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.538152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.538152Z digest=sha256:c0a66b4f33619984e210e961547f8c063299b5c0e481fe694294ee7a48a19c9b

Observation 4559d4da-8f84-45c6-b95c-2d17fde19ed3 · outbound

This paper cites UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.679187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.679187Z digest=sha256:47550f2a66cf33c33d69e8a42165b066c2591f4c544f25bc8716899bbca67cd6

Observation 6aa74da3-3b4e-460c-bd91-263168634441 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Align before fuse: Vision and language representation learning with momentum distillation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.787717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.787717Z digest=sha256:167a164e047c0168e0f04caa9ccab5cae97d24bf31a768770cd975f6ac8cea66

Observation 40ec964d-4a3b-43ff-9eb0-ac8e6074e064 · outbound

This paper cites Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:36.927145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:36.927145Z digest=sha256:a99eea92c1b64b3d1ac78182935ff9cfba7bb866d74606e568bd43f85938cd14

Observation 0e1b9acf-fcc9-438f-9f15-94e42948a820 · outbound

This paper cites Simvlm: Simple visual language model pretraining with weak supervision,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Simvlm: Simple visual language model pretraining with weak supervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.043824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.043824Z digest=sha256:e3e0f849c2be66804d257d39c3bd06a304b0407e7289a904491749d1ee517345

Observation 9d31b259-9a5c-4d4e-97d5-8f750e5c074a · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.184911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.184911Z digest=sha256:33bb6aeb76814d95721119718beef609c7a0a03933b9e612fe27bd40f729e55a

Observation 81f9cc6f-3c60-48a9-8c62-b536775451df · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.280478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.280478Z digest=sha256:5b502bab977442ec40d46fb0b66a1c38c9bfe0c3b874206d2dea3702ecee5832

Observation ab0fd178-91fb-40f4-8f7b-13bfa4a7129b · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Im2text: Describing images using 1 million captioned photographs,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.369147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.369147Z digest=sha256:d1521e623ff9339e695428656647a2d6d38fae4448e282bb3cfc5f8f559924af

Observation 16c6ed4c-a231-4626-9caf-a5e85204bc85 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.503386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.503386Z digest=sha256:ad915c10b83e926f8df9fd9eaca62fe434313fb7ffa144337f89c6c541a87173

Observation 43234c42-5705-49f9-bb51-c0a6d5b9edd7 · outbound

This paper cites Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.643209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.643209Z digest=sha256:e2376b5a91ac4bafbcf6f083141e475d3c5cf88202ca07bc429e9f68417c50ec

Observation 2559ceb1-fa12-4997-bc03-069586b28428 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Improved baselines with visual instruction tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.753062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.753062Z digest=sha256:bd2e14fb09ea06f71a88a3e8054745f3fd9437898d8ddcec776588a16ed38261

Observation 67f49981-e54c-4b78-a46e-cc375d64580f · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Monkey: Image resolution and text label are important things for large multi-modal models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.870902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.870902Z digest=sha256:3a5f1bba3906324cc541b6b2c82fa848349af14b8584019c405eb5e1d4634a87

Observation 54ecea55-9c8f-4520-b022-d83c20fe3c04 · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Docvqa: A dataset for vqa on document images,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:37.996940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:37.996940Z digest=sha256:9e06682ecbe2dfd54babf769d471c6c6276ab2091234be8eff1929ffe5db3802

Observation 537332fd-f14c-4b01-bebc-63fb4571fe5e · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.138298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.138298Z digest=sha256:d7c6317b676ea0d2a5ce078e71af492b863c81e2c2b6c890a069706550748bbf

Observation f0ed166f-4887-4415-ab39-a07af0c8721d · outbound

This paper cites Sigmoid loss for language image pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Sigmoid loss for language image pre-training,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.349219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.349219Z digest=sha256:0891ded3c60c3409e3b577645bc00fa7829b28846e7c446635a8b8bfd1cea11c

Observation 44a4699c-f096-4e34-83f5-f5e54cef2ca5 · outbound

This paper cites Qwen2 Technical Report.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.535201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.535201Z digest=sha256:645d0f0e83074688462afff15fecb6719f3527d98ee0171c88f388bb04771a6f

Observation 9d5ec591-45c1-41cb-b6d8-244f741b30bb · outbound

This paper cites Exploring Plain Vision Transformer Backbones for Object Detection.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring Plain Vision Transformer Backbones for Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.678581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.678581Z digest=sha256:dbddb6e3ed4c1c2b9054d4f7651b059c1aca9992eebfee0c4ffa082f9b91c8bb

Observation 9c6490ac-a365-4e03-a021-585cda9746f2 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.805151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.805151Z digest=sha256:2cf620848fade0c1db8e967ad84e662576f6f0bf24cbe67b6ce81e3f5df1db9b

Observation 8158c46c-bef8-4e83-8312-21b3dc5bdd76 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OK-VQA: A visual question answering benchmark requiring external knowledge,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.963496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.963496Z digest=sha256:d36c15ab820c563ec3bfa59410990c0c8ae31b1b4b340324caf7f1ea74da5899

Observation d4c7b5df-bf1e-4fdf-a001-c90e39924fb9 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs GQA: A new dataset for real-world visual reasoning and compositional question answering,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.071175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.071175Z digest=sha256:d8717d8c1e2a8f7cbe782d8c03bea1fe439fc65214263f21524d34b24feefd20

Observation be1fb631-5e0e-400e-b7d2-6f4ae267f566 · outbound

This paper cites MM-vet: Evaluating large multimodal models for integrated capabilities,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs MM-vet: Evaluating large multimodal models for integrated capabilities,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.186859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.186859Z digest=sha256:0140fb87ee8014d82489e93556b74d93ac0b628376983248854171b4157b0e16

Observation 0f6f0d9e-1324-4eec-b3ec-8182e7455207 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.324730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.324730Z digest=sha256:4abeb0ef48737996349dc54a7baa9a8ddbbd2890952e4a16470a16e09f888489

Observation 28546be3-8d16-4c27-8c88-47030076110e · outbound

This paper cites Grok-1.5 vision preview.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Grok-1.5 vision preview

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.541834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.541834Z digest=sha256:caa76522010ab02566f63da5abcd07aad069f81a666188a6c13c34462f465afa

Observation db613073-00ff-4358-82df-5a1e2594cff4 · outbound

This paper cites Towards VQA models that can read,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Towards VQA models that can read,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.732658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.732658Z digest=sha256:cb1746f0cc4e6cad37ef953d7b67b5abb4778be3d8f49bddf40b2e2b7129289c

Observation fd5ae75d-1c63-4d54-b5d4-335a3171eb90 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs ChartQA: A benchmark for question answering about charts with visual and logical reasoning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:39.947386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:39.947386Z digest=sha256:6d74d85e233307a4068ab2a548fe912a2d875c90d9b0bd039cdef6295962d77b

Observation 0d760691-3514-4c2a-896d-d4a8a8034796 · outbound

This paper cites Infographicvqa,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Infographicvqa,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.137978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.137978Z digest=sha256:1048e51f90c0a6c81a3e4a5d5c7c2044a292188d0e46e3a4144387cd5f53ed05

Observation d1350f9c-11ba-457c-9adc-0a6cd3b1d957 · outbound

This paper cites A diagram is worth a dozen images,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs A diagram is worth a dozen images,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.270291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.270291Z digest=sha256:bb1df8e947b174b8b5c03736a4f89e2a504460e49d28e6a922e64b2563d8a715

Observation c9896961-3bb1-4c9c-bdeb-be0b8c47cb4f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.485663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.485663Z digest=sha256:c2f06be9cae22591021f4ddce72b64477d47f1729c7fa5d7784b011c93d367b6

Observation 86e7b22c-7d89-46c0-956b-294c214ac865 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.621152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.621152Z digest=sha256:c7197fcddf69a058874cd1185314e30b89ee37d30be66054797bdc5a146a53e1

Observation 2f75a396-e90c-4cb4-90c5-f2a010df1e03 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.791974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.791974Z digest=sha256:99b6d2c6c99aa42479a99259c69519a558120f5ec8b8a161b3d76bba7822e9df

Observation 791fa363-476e-4220-b480-1297520c747b · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llava-next: What else influences visual instruction tuning beyond data?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.922650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:40.922650Z digest=sha256:a1851a028e9381c9de8748b6a805d28053dfd9af5252511e6ac507afd8585cab

Observation 1b0053e0-2ad7-40d6-9e37-341dea810c36 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.016343Z digest=sha256:aa2b5c94fb0fb4e513a54cb1f666676cc80435b5179bc5e69513941cb5876109

Observation ccfd809d-005b-4e6b-9e68-8651e9771796 · outbound

This paper cites Visual instruction tuning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Visual instruction tuning,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.133126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.133126Z digest=sha256:b7fe95b7e30a8c0d4cfbfef7e85d156d6f0bd8b159716edd972e1309f6b196b6

Observation 630611ec-21aa-4254-a63b-4364f039a30e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.228490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.228490Z digest=sha256:95af5820c460ba389ad7f7c4239177d52e43c57f1ed13f68b6046383de4e3c1f

Observation 7a46b53b-59b1-4e10-a856-f00b73567ab5 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.334971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.334971Z digest=sha256:2fbc4d9bc603365defee49333ee83ea20de6f4547f8a1ee40c199678b3b4c38e

Observation 1a13bfda-229f-4e88-b2f1-ff9484da7785 · outbound

This paper cites Neural machine translation by jointly learning to align and translate,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Neural machine translation by jointly learning to align and translate,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.454860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.454860Z digest=sha256:1a56dc81978813be289f23f6d128387e4ea7cd5d07b6f2769099709bedcf4837

Observation 78ea5f86-41b0-4a53-9cb8-04e90e66d748 · outbound

This paper cites Revealing the dark secrets of masked image modeling,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Revealing the dark secrets of masked image modeling,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.599170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.599170Z digest=sha256:4485e60d5b2bb4ac064690ff889f1c671384a3102a8e7587d592aff009dcc1d2

Observation 50e4abef-6a5a-42c5-a7d4-42205f4297a5 · outbound

This paper cites On information and sufficiency,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs On information and sufficiency,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.718447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.718447Z digest=sha256:8cd0f777bc536c7a55b7f782910efce23cf09d4c9648091a462a5518f7094f23

Observation 11507107-d0ac-4d9d-be1d-a4aa00ff6b15 · outbound

This paper cites VL-BERT: pre-training of generic visual-linguistic representations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs VL-BERT: pre-training of generic visual-linguistic representations,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.750278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.750278Z digest=sha256:7a4fb677e646802160a500e11d16f7e1b1bcc44a7e825652b18eb8b3806990fd

Observation c21d2ddb-66ff-4d8d-a3b1-8f6fe4d4b2db · outbound

This paper cites Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.831669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.831669Z digest=sha256:118d15824a82248e959e7c9f5fcff6a9a3c227cb4e28e7a96231deb6e1c0de2d

Observation 9185042d-2240-4ff6-bab4-27b2067d0c14 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oscar: Object-semantics aligned pre-training for vision-language tasks,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:41.945825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:41.945825Z digest=sha256:6c7dd4cce1e03f1399960756987340195625c237af11249242fa79e824fb7f49

Observation b0e88f0c-1d92-4bd9-9257-708632af24b9 · outbound

This paper cites OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.037290Z digest=sha256:d983ab158443368d3c3daba3a1f7e2e37674b4f443e67273eff001fa22cb5a7a

Observation 4f4663b1-9867-4080-86a5-7bb1fd1cc08e · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:42.307549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:42.307549Z digest=sha256:19914474e59c1d585d782954757edc63a89d8530593215103bee002d1d99d13b

Observation 1b21a005-6ee6-4ecf-9f2e-dfa787d4f2f0 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:43.340094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:43.340094Z digest=sha256:03de1299f46e29ccbe7f18df621d0dfde445b53739ca3a34ad6dbab8ce51c902

Observation 860a30ee-b50a-4881-90a6-4fac86b0df49 · outbound

This paper cites Exploring vision-language foundation model for novel object captioning,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring vision-language foundation model for novel object captioning,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:43.896751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:43.896751Z digest=sha256:370e0c1353cf954bfe44926f3f35a2cd6321c80650373be76a309ed68532e24e

Observation 0d2e4955-30c7-4ba3-b795-3ef241e06712 · outbound

This paper cites Unsupervised domain adaption harnessing vision-language pre-training,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unsupervised domain adaption harnessing vision-language pre-training,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.003055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.003055Z digest=sha256:72ee5aefedabb82daae0c79ae1d0c0be27abcce1ae3633f44d61894b813bff18

Observation 9850c635-261e-4a42-996e-4a495732adf6 · outbound

This paper cites Do vision transformers see like convolutional neural networks?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Do vision transformers see like convolutional neural networks?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.141844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.141844Z digest=sha256:25e2830ef5e9a19c748f469836c481864821e4ef1da4a57df9b0ba402c670ef6

Observation 4e491e1b-f238-4429-bc18-38adbcdc299f · outbound

This paper cites Intriguing properties of vision transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Intriguing properties of vision transformers,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.267295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.267295Z digest=sha256:90434cbc0e04e7cfc8b3bbcc4cb7f8793317cca17d4d95280896d320aba32ab8

Observation 20dc7104-e1f1-4a72-ade2-e2ddc88cbc52 · outbound

This paper cites Dissecting contextual word embeddings: Architecture and representation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dissecting contextual word embeddings: Architecture and representation,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.396287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.396287Z digest=sha256:9466a0c01788b354ed3395d51b1ac10f2eec1566233b15fec4302f4e3230d16a

Observation 0937db5d-0aa4-47ad-872d-540f20f90442 · outbound

This paper cites Linguistic knowledge and transferability of contextual representa- tions,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Linguistic knowledge and transferability of contextual representa- tions,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.506110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.506110Z digest=sha256:602f15668f0e3f4052aa9dca359558aa13e79102e76684df75c28827f1b60ede

Observation 0d159b50-7879-4346-a4b9-4b3d6b6049ed · outbound

This paper cites What does BERT learn about the structure of language?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs What does BERT learn about the structure of language?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.638326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.638326Z digest=sha256:fc45240b71c54a9b5e57847a7987d60ec4999ab8cea9ccaf1200c2680a86e673

Observation f7a40f94-b25b-4c85-be23-f6d1a752f0cd · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.742392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.742392Z digest=sha256:ca61125fe430498b9273b4c2be839bea4f80150e667cf0d9eb07fc830e0ac550

Observation 82ead94a-b173-4027-b219-aa7e4f40108e · outbound

This paper cites Feature pyramid networks for object detection,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Feature pyramid networks for object detection,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:44.850613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:44.850613Z digest=sha256:60624212ba501b33eb584efccd497586469733f40f6e79b05e7034ddd4a1fb6a

Observation 5491a66c-cb13-49a4-befe-b504ac064462 · outbound

This paper cites Densely connected convolutional networks,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Densely connected convolutional networks,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:45.304746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:45.304746Z digest=sha256:2ec7078c024918ed03636e9bef21bb8005d348d417336462279d393ce58faeea

Observation 4ddf0667-663a-41c7-8681-a534e57566c2 · outbound

This paper cites Deep layer aggregation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep layer aggregation,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.716660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:45.913961Z digest=sha256:5a2646e171b0dbdc7b31d6746e2dda908a42d7e80c6877426ece0b3701f7ada7

Observation 81b656ed-2771-43e3-85fb-4b7858b5b7c5 · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transformers,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Segformer: Simple and efficient design for semantic segmentation with transformers,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.489616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:46.004678Z digest=sha256:379f61376936116c6b807c10853a6c721837d7873b21a5b33fd88bc3aa70865b

Observation 87038d9d-e3d1-497c-818f-f7eee285c3da · outbound

This paper cites Clsr: Cross-layer interaction pyramid super-resolution network,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Clsr: Cross-layer interaction pyramid super-resolution network,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:07.169806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:46.141424Z digest=sha256:335c7328bf81af741f74fcd81a5684857cf13667ccda282a99ca07a49d856f8b

Observation 8239b6d0-30e7-47e8-acc0-fd8a21ec2a71 · outbound

This paper cites Attention-based layer fusion and token masking for weakly supervised semantic segmentation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Attention-based layer fusion and token masking for weakly supervised semantic segmentation,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.932055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:46.248925Z digest=sha256:9c57d7183a0dc846447d3893736a2ffb53a6afcfeea40e3cbddb48a794f87d2e

Observation f0638c7e-7665-4b3a-bd03-f89a1ae50506 · outbound

This paper cites Artificial- spiking hierarchical networks for vision-language representation learn- ing,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Artificial- spiking hierarchical networks for vision-language representation learn- ing,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.661838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:46.355704Z digest=sha256:05db6480274ba39240cc3dfb94c02b803f7816a6a9ce6ff8b6f5a633d5703e0a

Observation b1667396-7ca9-447c-87c5-2d97a56a6271 · outbound

This paper cites Deep contextualized word representations,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Deep contextualized word representations,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.341513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:46.461060Z digest=sha256:ef46504bdf2a73de38a5e944c98852b4c563c045e6c0997092255261ff3d9f7f

Observation fc0fba08-fdf4-4305-a1d7-a34552aad9cd · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Coarse-to-fine vision-language pre-training with fusion in the backbone,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:06.033250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:46.583941Z digest=sha256:162f21ae3eb03b1a5255458363f27d24cf85cd1b62d3e0209c97fb253acfb990

Observation 9ea500f7-b2be-4f3c-a87b-5624dbf3905b · outbound

This paper cites Dense Connector for MLLMs.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Dense Connector for MLLMs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.709519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.709519Z digest=sha256:9163d124af0b94f87f718f50e4c8ce034439cf25dc11e9cf12bc32d20fd50a5e

Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:46.824918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:46.824918Z digest=sha256:7f65e7f9449c68ff244fb9ec8b44221ebbc0fdddcb372985ad25ef0c8faf04d8

Observation 2a35e3cd-6c1f-489a-9ae9-bbf020cf6cb9 · outbound

This paper cites Language models are few-shot learners,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Language models are few-shot learners,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.754466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:46.918627Z digest=sha256:a66c4d0570a130c192bd3ead7e843a30511299fc4160a050e244cd24238b3e30

Observation 3e4e945f-0795-4c44-a2fd-3049636038b5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.068817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.068817Z digest=sha256:302bfcbe81bd1011463b962d9872edf192657638cb8b936a869aae52c82ab31a

Observation 5b98f0a0-33cd-4afc-b40b-2d8e3fbbd78e · outbound

This paper cites Large Language Models Meet NLP: A Survey.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Large Language Models Meet NLP: A Survey

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.189430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.189430Z digest=sha256:83d132f5b028966770400d74d5b196f262f5f4b07aeb0bfa983304763529e9f1

Observation 33ec15ab-e7f7-4152-8e89-546a98610b0d · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.526388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:47.273636Z digest=sha256:f2cae78651379d78878fa42c2a9fd6cf1c642c76c274012227409b8407b61a09

Observation f50a89ec-fb71-4975-87dd-9660679b2255 · outbound

This paper cites Introducing our multimodal models,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Introducing our multimodal models,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.298926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:47.359818Z digest=sha256:981ce0e5facbb83b14543b5fdc31d7425952a7613ba16739f0423b92f3ce462e

Observation ffe0c48d-d30c-467d-9802-8039d28c4610 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.490864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.490864Z digest=sha256:4bd5cc713bf593208cb061776f39561866a5939c29197175dcf80fa106433b42

Observation 0ed14fee-6450-4306-8821-431c65915b8a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.591621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.591621Z digest=sha256:bc060a93693c0e4af4afcaec7168572b427b8c342a75fdd57cd06e5d6d08a00b

Observation f2c93c76-ca12-46eb-8d29-e800355ccd76 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.692121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.692121Z digest=sha256:90e546d061e0dae0a39db314cbd468bb2a110b82e2e17b93bfb77360c100d6fc

Observation 30c172e8-53e3-4dc4-9254-b71a08b7a65e · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:47.795386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:47.795386Z digest=sha256:3a1ac8306de8af380727aa04486b1aa4315ee238ecad756467a14b921da3a3ac

Observation fe259a9d-cf04-425f-93f8-ba3e61506701 · outbound

This paper cites When do we not need larger vision models?.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs When do we not need larger vision models?

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:05.061995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:47.882134Z digest=sha256:b34281ee21947742cbfcf5fda8b31b6327a7e5bd5afb5b08c2ff1ffa3535639c

Observation 55ef876c-a123-4673-beb6-714e832d43db · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:48.010046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:48.010046Z digest=sha256:97f172fb9ed4a056016c8bf9b05a914af227647e480a7b14172c095f80443bce

Observation 3aac154c-821b-4d07-b583-e03cd288f6a1 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Honeybee: Locality-enhanced projector for multimodal llm,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:04.816456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:48.091130Z digest=sha256:f2c1adf118475bb65a64e238d3a8c2937aec5582c531bb447ca3d7cbd4fdb8cd

Observation 0815b0f9-2125-4277-8e8e-d6211aecd15d · outbound

This paper cites Unified language model pre-training for natural language understanding and generation,.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Unified language model pre-training for natural language understanding and generation,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:09:04.564950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:08:48.226962Z digest=sha256:dad6f28364d94f552f6f59bbdcae64ec90014d7216a90e7f5c445d20bfce441d

Pith citing papers

No inbound Pith citation observations are available.