Pith. sign in

Paper Citation Record · LEDGER

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training

As of 17 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2505.08971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08971 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:28.151834Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact5
  • verified fuzzy9
  • unresolved82
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b4a69d45-0675-4a6f-9889-01abd5f98526 · outbound

This paper cites Pixtral 12B.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.672537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.672537Z digest=sha256:f6d3f39b1c99901fc67b955e4e209a569f5b7b48178c2ed8229cda3440124dbe

Observation eac79578-b33d-4d61-8055-54d5347c4c7a · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.678003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.678003Z digest=sha256:c054357ae74e26922b6a65ad23643aa65bb9088107bf71cd45fe0a45c71b6537

Observation 2e4b0177-ebec-4c49-bb28-d1620663d0b2 · outbound

This paper cites MINT-1T: scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MINT-1T: scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.681895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.681895Z digest=sha256:cb1dba1ab866473bb196a11f726ccc94cb307dc5947da027d8a505edcc04cd30

Observation 8b7b35db-ab7d-4346-bb4a-ec2459e40c71 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.685443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.685443Z digest=sha256:e9fc2d2b46c1ac8abb14d21cd39d1c091190cc8e4a3cb59de6900a3a71e7b4d2

Observation bf650b6e-bd5b-42eb-8b00-a0ef908fd3d7 · outbound

This paper cites Language models are few-shot learners.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.689864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.689864Z digest=sha256:213453a0538f54ff690e2fc49b2a8fbc00035a862f37a9acbfb98bb4f0c3a98c

Observation a1748332-e4c4-4340-ab26-ebe6c57d6913 · outbound

This paper cites Lundberg, Harsha Nori, Hamid Palangi, Marco T´ulio Ribeiro, and Yi Zhang.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Lundberg, Harsha Nori, Hamid Palangi, Marco T´ulio Ribeiro, and Yi Zhang

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.694115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.694115Z digest=sha256:f99bd6771d4d1fbe9cde70a01fec2c359b4338d5e119b2db1976d67a18476c26

Observation 193d4a2f-c0a6-4676-bdd2-cc379ab3f6dc · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Sharegpt4v: Improving large multi-modal models with better captions

Reference 8

Resolution
verified exact
doi, observed 2026-08-15T21:49:28.536110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:27.702072Z digest=sha256:3f70d857984dbfc0001a0fe5320ee357c00df91070239e9982918470b830c581

Observation e5337ba4-b7db-4866-a6ed-8045cbbe4d53 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.706161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.706161Z digest=sha256:a834ca717d0c3e31333724902eee7af2e8b54f34c03bd33d813480cf1991c843

Observation 7d6176b4-2d18-4521-bfd1-43c0e6ddca68 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Sharegpt4video: Improving video understanding and generation with better captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.710172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.710172Z digest=sha256:67e4d2ef7046a1294a15b57e7f0f5f44bfedbf46730e989f20c46e298b6143b8

Observation b900c1c4-cf50-4c71-a589-a6197200a49f · outbound

This paper cites CompCap: Improving Multimodal Large Language Models with Composite Captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training CompCap: Improving Multimodal Large Language Models with Composite Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.714083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.714083Z digest=sha256:97e96cf4f723df339ee1e7fc93737df385a0cb3aa058551f92ca3cf211598f67

Observation f3f217c6-9abe-4c43-ab6a-a8756bade8e0 · outbound

This paper cites DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.722824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.722824Z digest=sha256:b97487d48c8d2e4c8c0259699727bfd07691effc3dba913ce54c7732f6dcf32d

Observation ef02f1db-390f-4a24-b7e8-23c29d79ab26 · outbound

This paper cites ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:49:29.099720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:27.726826Z digest=sha256:5582bcce0817a550fafc76342a467d7ce69de0528c75894247746519b0db7153

Observation 80f5c340-a661-4d8b-9aea-0324e68270d6 · outbound

This paper cites Scaling Laws for Predicting Downstream Performance in LLMs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Laws for Predicting Downstream Performance in LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.730622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.730622Z digest=sha256:587a45cd93ab7dbd6e22b28798f0089dbb607185873e8bb952e93b81e7bbfad7

Observation f7a0b0cc-42e7-4957-8077-93ec3a6cb291 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.734866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.734866Z digest=sha256:e59ee6f7c6d8f9b9605eaecf645f0524f62590e6911c11bb4b7e502961e5c7c3

Observation 79b26b88-b6ec-4573-8418-6c9a459c59cf · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.739146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.739146Z digest=sha256:36e444209165e732a3f2a2fb8e7236298c6fb18ade82e686436b7ba53d7042f4

Observation 5f997c8d-10f6-4d55-aa56-1ef07d717643 · outbound

This paper cites VisionArena: 230K Real World User-VLM Conversations with Preference Labels.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training VisionArena: 230K Real World User-VLM Conversations with Preference Labels

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.743070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.743070Z digest=sha256:ca0d8ae7100e49b9785fd7e4d263e11f77e4033915a8fd805290c51b61e99894

Observation 365ae205-0775-4f22-812c-fb6b119112ac · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.747119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.747119Z digest=sha256:723ccaeb7f8a7fb971ff4b43c1bd38bdd8b04f7126bf00ba73cfafd0a804a4ad

Observation eb3e2227-14db-4f52-a364-5876d877d964 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training NVLM: Open Frontier-Class Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.751226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.751226Z digest=sha256:dbb91406ce11efaa785eb9c2336803669c1a6795ebf42c792357831d832e4778

Observation a44695a5-20f7-4c90-9b42-69e27ac5eb05 · outbound

This paper cites Unveiling encoder-free vision-language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unveiling encoder-free vision-language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.755310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.755310Z digest=sha256:c5eb2e50489f81abd56d107d5fc392af426c91bcbaafa843605d2d40f03c2c1c

Observation 82f4f70d-6151-44d4-a90c-b8a976f052a3 · outbound

This paper cites A survey of vision-language pre-trained models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training A survey of vision-language pre-trained models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.759085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.759085Z digest=sha256:decc609c5d318d1e77fe3d6c89ccb8b4de22036b1a9343f38ae2219b6292a98c

Observation 2638719a-6663-4c36-8752-37b15a7b659b · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.762903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.762903Z digest=sha256:d247092671734c1c06521b575facefd4a2c901f9f768b4ed2d3f2ee29e657384

Observation 318baf8d-a093-4204-af64-0384b002044f · outbound

This paper cites The Llama 3 Herd of Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.767272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.767272Z digest=sha256:e70c2bd2838e4a6fdc70b5a09eba3db48b23ca72ed8c25c77f8e098b8606543a

Observation 17d48c46-1b18-46bd-8c29-9480716cab80 · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.771127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.771127Z digest=sha256:ed311e209c6744386e62e1fa2a520f0ca06a860c4df5e7381474505199c9f042

Observation 2b3c34e9-828b-4093-9126-33a9a2bd0920 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.779889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.779889Z digest=sha256:8290f07461bc9837aabf9118eef2b5027da292ad140a52e13d9dd7bb5f7d7aa8

Observation ce61c136-8d4b-4514-8c5e-204acd8d0cc5 · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.775568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.775568Z digest=sha256:c1223140322e516515f282cfd1952522e99b0a095fe69d415d50387316814bfc

Observation 6d789315-dd49-4844-a28b-1c0980a06322 · outbound

This paper cites Making llama SEE and draw with SEED tokenizer.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Making llama SEE and draw with SEED tokenizer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.787389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.787389Z digest=sha256:6c70e1c61c850d6858164271547c57e2b96d84d0e1ad9dfe9cc09d4e19cdf843

Observation 261dcc71-7fd3-4ff6-a3db-4693e805c0ce · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.783967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.783967Z digest=sha256:ccd902da4d1166135fd90d49a25e2b9d149636c46bc371434ef03d7a2d8b1bc5

Observation 3e8ba83b-2c37-4ea6-813b-205a423b23d6 · outbound

This paper cites GneissWeb: Preparing High Quality Data for LLMs at Scale.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training GneissWeb: Preparing High Quality Data for LLMs at Scale

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:49:29.045263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:27.799078Z digest=sha256:f6600e0cc4608a52f9efaffcbf80d049636d7f5c7938ceed108384e960284c74

Observation 2c1f37b1-9fae-4d97-b5b4-a15a1b0d5d06 · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.791070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.791070Z digest=sha256:6f2fbdad77ad7b1fa410a1aae20dd12e45b5cf1ebc431d2a1649ae7a41df99d9

Observation 90741552-0327-4828-a6bd-6b6511beaa7a · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Exploring the frontier of vision-language models: A survey of current methodologies and future directions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.794662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.794662Z digest=sha256:fdbdccebd86fcb5017a8d72647528248e076f032b00ae47a9a19c868eaa3960c

Observation 02a9fbaa-accb-46e2-9ae0-697d8c48db7f · outbound

This paper cites InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.810010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.810010Z digest=sha256:06833c0aab1d09511bb30fde768d8bd98a1f810e93a2391ecf5c02ffb0787969

Observation f7140819-8ced-4d7d-b98a-4f199578fae2 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.802928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.802928Z digest=sha256:4b53a1a8c593d0dfdd059f8fb04f2e46944aa6167bcdfb442e29541f7ef3e8fe

Observation e98f92c9-47e8-4a08-a328-0c1cdccc6de9 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.806428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.806428Z digest=sha256:d8f64d200a64aedbb7033b3e76b0c541dbc6acd738a281ec875edf31c40c017c

Observation fd34ff79-fe5c-48bb-81c4-b443b111b954 · outbound

This paper cites Scaling Laws for Neural Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Laws for Neural Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.822792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.822792Z digest=sha256:4161d956ecda79f9358ddd15e467b637e563901a81d0537a9c544e60bd0f3f53

Observation b1cc2334-1a52-4202-9562-d6522531ad34 · outbound

This paper cites Compression Represents Intelligence Linearly.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Compression Represents Intelligence Linearly

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.814220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.814220Z digest=sha256:f54c6b391554ce2d80fa4151a7e69b1e415ffe36f7babec8f61ec26c75df9b2c

Observation 5f83a591-95c1-4821-8c14-0b91a272b4e2 · outbound

This paper cites GPT-4o System Card.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training GPT-4o System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.818250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.818250Z digest=sha256:2d605f75b337bdda642e9b3483156a7979c6f4c31163593e571053affc9a1a65

Observation a3172c1e-ea2e-428d-96a7-4e2d469dff55 · outbound

This paper cites The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.833950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.833950Z digest=sha256:a5b0949087f75e5f0a9e79beb0f50caed11f64d83fa1d44fb4dead3d2fc732da

Observation 15b72efd-41e5-4a00-861f-fecff345027b · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Grounding language models to images for multimodal inputs and outputs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.826874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.826874Z digest=sha256:cc3d99ff8bb1fd277a03f966c79ae0dd492be619dd75cf483d869af70c782b40

Observation 876999fe-af79-4977-ba31-ae9939d891db · outbound

This paper cites RL with KL penalties is better viewed as Bayesian inference.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training RL with KL penalties is better viewed as Bayesian inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.830463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.830463Z digest=sha256:f1fe22564be9c9901838e567d238a6e020f48374382b91745d6a9abee0dba2b9

Observation eb459e6f-d094-4a22-b529-c4d5196e416d · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Datacomp-lm: In search of the next generation of training sets for language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.844873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.844873Z digest=sha256:c06acfe3e49b3fa348054d30c2aa40862c53c7a29dd0655bb233d52691f0ca2d

Observation 6f610eba-bf88-417c-9484-6b7154780b3f · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.838051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.838051Z digest=sha256:36077586c7669b12073652024be1046627206ec28c6cf34e7c6bbb2d44a1e23d

Observation c0160033-8a5e-4766-9c5b-caf5b1c93210 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Seed-bench: Benchmarking multimodal large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.841739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.841739Z digest=sha256:eca76f924338e5764a2ae86add2ba06e9b3495b8c532f99f82f5a2afeffcb38a

Observation 7480b147-b897-4d8e-9dd4-b9379c5240ac · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Silkie: Preference Distillation for Large Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.856076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.856076Z digest=sha256:380d1849942a40bd71449ae0fb975b7b5ff1125b7cc071974cc968e206990ccb

Observation 4d435741-251e-484a-a5fa-5145de9af93e · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.848093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.848093Z digest=sha256:f8c45817da23a2e1c1d83e69e6293643584d12afa87e987d48b3cd3e34d7bbbb

Observation 83152920-85f9-464a-a8b2-9e820abe7f2d · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.851400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.851400Z digest=sha256:2a8d8b5d2708ccadc12b7fba4eda3c7d8f9707ce2961d5732d03f9c51285284e

Observation e389e1cd-a794-40d7-bfc2-84c10a3bb9d7 · outbound

This paper cites OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.867920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.867920Z digest=sha256:dbced8c37b800d50e7c127d61a215eef5654ca350203ca3714f178c81a40b48d

Observation 64b974cf-03f0-4957-9ade-a54420a8ddf9 · outbound

This paper cites Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.860976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.860976Z digest=sha256:cd39efcf2cc304b5e0e87f42cb323399bf93490047d99efcd47a99c75c692ff7

Observation 0e1943bf-0a4e-4196-943c-f2405ecb5dc8 · outbound

This paper cites Visualbert: A simple and performant baseline for vision and language.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Visualbert: A simple and performant baseline for vision and language

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.864520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.864520Z digest=sha256:3e454984729d9d47c1a70e11138ba6d960b79e96109724b7417baddc04bad4cc

Observation 030a7c39-0ea4-48c8-985e-e6137171ad2d · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.879585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.879585Z digest=sha256:f0cdea91f9ab3847a1564971d35d7a99f7b3c9545c37dc82b5bf6c3408437216

Observation 5f3ca946-0d59-4336-ab7c-e5c123507aef · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Evaluating Object Hallucination in Large Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.871672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.871672Z digest=sha256:117e1d7a13119b9679c4c8b52175a569d0c9dcb368e1faab579e8ba55d36a86a

Observation a24ff640-5790-44ae-81e3-8638138857b3 · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.875302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.875302Z digest=sha256:aa5f832ae216f39a48bf38cf85cb04d093cd163affd9d9ca877614d53c85363b

Observation d0b3e2b7-8bed-41e6-810d-ead96c0b64d5 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.994447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.994447Z digest=sha256:430f228ac8eaa8bbf2447f1e55217a1a037d7917cc8b85d4e5e3655693b151b6

Observation 55e44184-c986-4471-9ab1-d201a7d5b174 · outbound

This paper cites Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.884393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.884393Z digest=sha256:ea6fb5618cbbbc3f6ef065ce73ee23a6dd832ac0959455079b1d14103941312c

Observation 64af87a0-4391-48d9-a7b1-402833b212ec · outbound

This paper cites Microsoft coco: Common objects in context.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Microsoft coco: Common objects in context

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.373240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:27.888232Z digest=sha256:9ab3959917e4f767d8bc8b622e87b34671902878f9e2949a6416db94d8d8effd

Observation 0f72c81f-683a-4b5a-8cb7-c7b016a21ec2 · outbound

This paper cites Visual instruction tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Visual instruction tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.360488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.007067Z digest=sha256:96ec9822849a7a2d0a31ebdcbb6e80f19dd99edc4e962a8ed09ae7b790d4add5

Observation c8d007ba-729f-4e2b-ba97-02809fd98630 · outbound

This paper cites Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.999473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.999473Z digest=sha256:11f00f9b1aaa93804d215645bfe15faae19c2b609095f8af483591dd4c79dfc8

Observation 74826612-12bf-4a03-a56c-c979fb637aa1 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Improved Baselines with Visual Instruction Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.003360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.003360Z digest=sha256:2c7827be84e3b5b8e5c1334160cb8e290a6f43d5b1de82f626c30f5bfafbc757

Observation 25e0a611-7da3-4a9a-a5f3-78cf6b062332 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.018040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.018040Z digest=sha256:637d52b74a39c38ff31e3dc90f5ace7dbedd90d6a3b622eccda5f094e07de228

Observation 1c4e6354-915d-4aab-ade1-b9e1fd6fde6e · outbound

This paper cites Diving into Self-Evolving Training for Multimodal Reasoning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Diving into Self-Evolving Training for Multimodal Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.010620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.010620Z digest=sha256:79b31e1a5317d70b95668046c8975a72a1a8bdd4f52894a6322beecb0c404a74

Observation abcf93ad-1c9a-4057-a401-855473f41316 · outbound

This paper cites MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.014362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.014362Z digest=sha256:2cb90daddf9431e726460d6dfb7db97bb126986800a01bff38b176f2fe6eef0c

Observation ebb61c47-cd18-4434-9020-7256098dcdf9 · outbound

This paper cites DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.035412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.035412Z digest=sha256:c5b8b1b0c1fc6b6fe7e4a0d42837a03af719a6ae5d03e77fe3fd7950c409fa6f

Observation fd2b56e1-36e0-4c13-ba33-2fda71e1fd25 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.022589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.022589Z digest=sha256:18bfc90f34b820a89ba47a9b8f1cda8267c510891a2803428635f2eed260f46d

Observation 13a1a594-15ed-44fe-968a-0f57024e18d4 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.027700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.027700Z digest=sha256:f978afffe1ed839f73c2da3f4cbc9dd7c9a4cb27196e50d39072d1cbcd35c5f0

Observation 96804bae-efba-4945-9a92-aa6c61c3e9ae · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.339876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.031470Z digest=sha256:3c75bb8993817b757639370823f37d9bf5275216148b0d4e009dd833f7939ffa

Observation 85f6e33c-f575-48ce-8f74-8ac8a5a68c97 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Learning transferable visual models from natural language supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.051509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.051509Z digest=sha256:4b4b44530f0fe6ce802a16bed57999ce527c222a2c5610af718c14e8fd8257ba

Observation d6cbddcc-b0f9-4c4f-87ed-ae55c6cbc287 · outbound

This paper cites MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.039274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.039274Z digest=sha256:32939f5040b3afc79831298e7feb8c5be0ae524ea3b8855276b09c924bec9d7e

Observation c1d25b96-0c5f-40ef-85d1-08513a34091f · outbound

This paper cites Openomni: Large language models pivot zero-shot omnimodal alignment across language with real-time self-aware emotional speech synthesis.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Openomni: Large language models pivot zero-shot omnimodal alignment across language with real-time self-aware emotional speech synthesis

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.043623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.043623Z digest=sha256:a2ac1688c944400fabd5d25cc15a2000eec8e11ad15a5f54d575d04618214374

Observation bb22e9f9-d6f7-4ba8-aab5-1e175efc7fce · outbound

This paper cites Policy opti- mization via importance sampling.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Policy opti- mization via importance sampling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.047320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.047320Z digest=sha256:4c8833fc54feca62b84743fc08e5dcdf8772f3ce186fff33db85e04d56f7a2bb

Observation 1dbcd1a2-a988-47a4-a2d8-dd151e8fc9d5 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.066102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.066102Z digest=sha256:883433961bc43775124269fa2ab9d83c2e593376649f0401a351d15447b3cb38

Observation b9abe44e-0b76-4db6-ac12-ddaf0d0f5925 · outbound

This paper cites Shaker, Salman H.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Shaker, Salman H

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.055287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.055287Z digest=sha256:48f247f806f8360aa3bef1e8fbcd738913abfb94223d82c3985133a8b1e3c204

Observation 3879d285-9344-4d6b-bf72-5688cc28d5f0 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.058792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.058792Z digest=sha256:cb0aede973b3d9f113c7428809164da5e41505001c2985f4584cf58d56681e8b

Observation d6bea31c-8c9a-4b51-8623-3ed4649d3763 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.062362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.062362Z digest=sha256:756beec5b3874b783fd89cdbbb2890997f0b06b0b1b5fd84db4e74212a2b5e96

Observation cb89ef5c-69ad-403c-aafe-b27ca3640b4e · outbound

This paper cites Emu: Generative pretraining in multimodality.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Emu: Generative pretraining in multimodality

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.290113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.080798Z digest=sha256:b60b3f4d91d8fc652f41ea32a9ea66459bccd49fabe544fadfba5d154ae1b583

Observation 5b5cbdee-8ca9-4930-91d5-bebdfe1845d3 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.303589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.069756Z digest=sha256:3ddc1c0c88b8e3e2f22fd3ac1bfc38c94442db557fdc8b37c324b85282354213

Observation 360d5162-a319-46d7-8474-878065e5879a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.073076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.073076Z digest=sha256:21e63a983f9c0335347ba5d1c44ad6247ff74e99521afe399cc8f177835953c2

Observation 3da53c17-8322-42eb-b5d4-d3185b89ee71 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training PandaGPT: One Model To Instruction-Follow Them All

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.077016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.077016Z digest=sha256:aefc179fc21a8073bb89e0710afb53423b328be95c918e9c5dddece8f38163f6

Observation 8ff4b12a-1022-4dd7-a36f-6792416b7a70 · outbound

This paper cites OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.277670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.095926Z digest=sha256:e79bf72e2683256b1a364243d6bb117280a662f82cf01c29348622139a336460

Observation 2df2a628-0da3-4ae6-bef0-8615714bda71 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.084067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.084067Z digest=sha256:9d19baa4aa23bc26d01d999676f51f856087e6ec1246bab0dae41636f208da01

Observation 84b0381f-ec30-40e2-8293-c9407eb28d2f · outbound

This paper cites HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:49:28.700017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.087787Z digest=sha256:7dc4feda93fd5e0e208199085576885688fc7a7614736b3560aa3be89641d326

Observation 3335f5c4-4051-491a-b313-b5ff7f2c1351 · outbound

This paper cites Reconstructive Visual Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Reconstructive Visual Instruction Tuning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.091627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.091627Z digest=sha256:8075f5ed6ed439d838d272088723ff57929c4ad09406834d5e4aaf3dc2719875

Observation ad5553e3-56c8-4a2e-89a1-af392a36dd7a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.110794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.110794Z digest=sha256:ff55d1badb60dd90dd92f75467b4145c48c3d340b44c72d7f57e82953a2edded

Observation 1d93d7d5-4d49-4725-bcc1-7909cbddcb40 · outbound

This paper cites Scaling Pre-training to One Hundred Billion Data for Vision Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Pre-training to One Hundred Billion Data for Vision Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.099780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.099780Z digest=sha256:ab0500a1ebf9e718b7d71808fcc24e3c5800a81a1ea3b5ce8d58227f8862a44a

Observation d9749ed0-4144-48d1-ad61-70dc9e887f97 · outbound

This paper cites SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.103319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.103319Z digest=sha256:1a8c11e3d7f995e23c576cbf1b57b740007f7782ef0325c0f58ca1fadb03029d

Observation 0d837735-db69-4a1c-a801-adf9a07ce6de · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.107044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.107044Z digest=sha256:bcf8cf3ea157534e4f3d91943c5e75db4b10e4c893fdd3203e56e277f7d27b8c

Observation d5f5d2ea-f6fa-4feb-a59e-cef1e23e31cd · outbound

This paper cites Knowledge-augmented few-shot visual relation detection.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Knowledge-augmented few-shot visual relation detection

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.257203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.125115Z digest=sha256:9e52ac563df7aa61b715cc687e2251dcf5df0efd84dfd8731a781e94f3e5d0a5

Observation eed933be-e5d8-4304-81af-4f50737727fe · outbound

This paper cites Qwen2.5 Technical Report.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Qwen2.5 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.114329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.114329Z digest=sha256:6812b63b30c5f2ebb8ffbee18334437562711dbc7626b7228da063b7a5a0d424

Observation 290d4613-fbd5-4354-b505-ccf4a1af24e1 · outbound

This paper cites Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.118433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.118433Z digest=sha256:b03f11caadeb5cee75524a4db9e25ea549a6f2c67df873d4dd3c5c4c6d5e45c6

Observation f3c8e3e5-187a-4f61-b571-83731f8e9d7a · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Capsfusion: Rethinking image-text data at scale

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.121972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.121972Z digest=sha256:06bbb2dc52ef3f2a90dac49f6ec75b77794ded22038db5105936162bb8e0ef64

Observation 5713fdfb-d4ab-4fe3-8ad7-32ea69cf88df · outbound

This paper cites OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.140238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.140238Z digest=sha256:ac1268f912fbcc80b2c134bd6b6b4021ce6503dfb5e2e160ed7289c41a8cc27c

Observation 048a03d7-ca78-4126-a150-c8fcb0b526ad · outbound

This paper cites Xing, Xiaodan Liang, and Zhiqiang Shen.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Xing, Xiaodan Liang, and Zhiqiang Shen

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.244964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.129074Z digest=sha256:d2e9b40d716b0e544c5cbfa5bebb895bdeb8538c534b4562be454d786359c8e5

Observation 3aea2af3-136b-4e7b-94b9-a7548592bc4a · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Vinvl: Revisiting visual representations in vision-language models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.132406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.132406Z digest=sha256:d8f5d20b96c110fe33f50ac820b014625bbf742d8b477066c3d346c0a8b5e035

Observation 443d7519-58aa-4878-8b9a-b7cf2ec0812a · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.136196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.136196Z digest=sha256:06608081aee46076b82a42573014f701c7c0e7a48cca1882d74b84d7f3427450

Observation e80c6efc-31ac-41fb-aad6-ac2ca9b4bde8 · outbound

This paper cites Minigpt-5: Interleaved vision-and-language generation via generative vokens.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Minigpt-5: Interleaved vision-and-language generation via generative vokens

Reference 95

Resolution
verified exact
doi, observed 2026-08-15T21:49:28.240603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.144082Z digest=sha256:36a656a2702b26c8043428f3451cc638a3b1f76a88866aaab5905fe180b7ad0c

Observation 0b3a4518-f49f-4699-99a1-4af729b51b37 · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.232823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:49:28.147999Z digest=sha256:f07ea341bc05293b73be52ab97b257a0368798a80c1891a7fd25614d65c01c3a

Observation 87d92ebd-c559-40de-af3c-bd1e0548a290 · outbound

This paper cites Generalized decoding for pixel, image, and language.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Generalized decoding for pixel, image, and language

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.151834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.151834Z digest=sha256:c13066232efd09a963a57245db8e01935628578957b46796ab8d1ba046f374cd

Observation d08c95db-92d6-44b6-a2db-cdbd44d42d01 · outbound

This paper cites CompCap: Improving Multimodal Large Language Models with Composite Captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training CompCap: Improving Multimodal Large Language Models with Composite Captions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.718480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.718480Z digest=sha256:814079d45fd03e4bef9784b94b95e2339d99a738d3d4e8ed0e7b8fd02f6e56ff

Pith citing papers

No inbound Pith citation observations are available.