Pith. sign in

Paper Citation Record · LEDGER

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training

As of 17 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2505.08971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08971 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:28.151834Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact5
  • verified fuzzy9
  • unresolved82
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b4a69d45-0675-4a6f-9889-01abd5f98526 · outbound

This paper cites Pixtral 12B.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.672537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.672537Z digest=sha256:5fff01cb5520781b6fe1c8f8b4532fc2389c60853e75067aa0b2107ca6a2af69

Observation eac79578-b33d-4d61-8055-54d5347c4c7a · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.678003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.678003Z digest=sha256:6578f65b66d8e0a10c2259f1e3c785c18bc5e197a30cd1963190e39b31a0f555

Observation 2e4b0177-ebec-4c49-bb28-d1620663d0b2 · outbound

This paper cites MINT-1T: scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MINT-1T: scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.681895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.681895Z digest=sha256:09dd53c5ba038651f9a40263b6ffc7c74d4703e1b834c44f47cb3a45a72f8209

Observation 8b7b35db-ab7d-4346-bb4a-ec2459e40c71 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.685443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.685443Z digest=sha256:b434613e15d343831a30a3325cd357e27bc34a70b8645a407e141d5c6f004083

Observation bf650b6e-bd5b-42eb-8b00-a0ef908fd3d7 · outbound

This paper cites Language models are few-shot learners.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.689864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.689864Z digest=sha256:0c81895b5da873782dd9d7ad4b17233e7e48e16f11eb628a60a2e8e86272a5eb

Observation a1748332-e4c4-4340-ab26-ebe6c57d6913 · outbound

This paper cites Lundberg, Harsha Nori, Hamid Palangi, Marco T´ulio Ribeiro, and Yi Zhang.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Lundberg, Harsha Nori, Hamid Palangi, Marco T´ulio Ribeiro, and Yi Zhang

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.694115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.694115Z digest=sha256:80dc1fc124c0db5be9646592bd72211abc8e0024ed1bcf1e552491022014163c

Observation 193d4a2f-c0a6-4676-bdd2-cc379ab3f6dc · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Sharegpt4v: Improving large multi-modal models with better captions

Reference 8

Resolution
verified exact
doi, observed 2026-08-15T21:49:28.536110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:27.702072Z digest=sha256:23bb61ab39632a731a44912aac70787f6b95c0afb9464b6e934c6e0013b7b20e

Observation e5337ba4-b7db-4866-a6ed-8045cbbe4d53 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.706161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.706161Z digest=sha256:515fc125b212c00fd97264c4c66ef03a0cb4600af332275947637790d15feeb2

Observation 7d6176b4-2d18-4521-bfd1-43c0e6ddca68 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Sharegpt4video: Improving video understanding and generation with better captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.710172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.710172Z digest=sha256:fdae849620a7196173f0684d91f687a2ba9c1a9665b9560d127f9ba918bf651c

Observation b900c1c4-cf50-4c71-a589-a6197200a49f · outbound

This paper cites CompCap: Improving Multimodal Large Language Models with Composite Captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training CompCap: Improving Multimodal Large Language Models with Composite Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.714083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.714083Z digest=sha256:5c575aa54d58f6ae724830b09b31f1aad601d5b53a54a2a4a3cad8dada960fd8

Observation f3f217c6-9abe-4c43-ab6a-a8756bade8e0 · outbound

This paper cites DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.722824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.722824Z digest=sha256:6fc8300438bf8d149a67eb0a4d5fc4044d4585c658bb3427188a37140a49c9d9

Observation ef02f1db-390f-4a24-b7e8-23c29d79ab26 · outbound

This paper cites ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:49:29.099720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:27.726826Z digest=sha256:29a0535ce79c1c2241ffc125cb0247ade5f40c855d7c71cb520560a27b0a5589

Observation 80f5c340-a661-4d8b-9aea-0324e68270d6 · outbound

This paper cites Scaling Laws for Predicting Downstream Performance in LLMs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Laws for Predicting Downstream Performance in LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.730622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.730622Z digest=sha256:db1dc5309bbf659a382bf50573aa2ce99bf38eea53304a614f06d5859cbba2e5

Observation f7a0b0cc-42e7-4957-8077-93ec3a6cb291 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.734866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.734866Z digest=sha256:6726fd579c8693487623dc5a99420f28c255012bc8a8970b850dd907b152d5a6

Observation 79b26b88-b6ec-4573-8418-6c9a459c59cf · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.739146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.739146Z digest=sha256:5904b7d235060813fabc9aebfcf884a0d86632b4ae47d0a329bd0ef6b1460b18

Observation 5f997c8d-10f6-4d55-aa56-1ef07d717643 · outbound

This paper cites VisionArena: 230K Real World User-VLM Conversations with Preference Labels.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training VisionArena: 230K Real World User-VLM Conversations with Preference Labels

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.743070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.743070Z digest=sha256:27135073a063816b269c584eacb283f69e9f8f5ab8f7feed4525fd960e13fe6a

Observation 365ae205-0775-4f22-812c-fb6b119112ac · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.747119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.747119Z digest=sha256:cf2b2d52f61edf9623d4d4f3d57dc066d9d785ddd89eea3613d49f5bd43fe7a6

Observation eb3e2227-14db-4f52-a364-5876d877d964 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training NVLM: Open Frontier-Class Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.751226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.751226Z digest=sha256:8821091182ae333d6643ca1909a444ce150c12a16a26a79be3acdf17bc5c1876

Observation a44695a5-20f7-4c90-9b42-69e27ac5eb05 · outbound

This paper cites Unveiling encoder-free vision-language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unveiling encoder-free vision-language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.755310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.755310Z digest=sha256:b65d4112f8c82faf705824c61ffa000376d38b465848838a2e18522725f01113

Observation 82f4f70d-6151-44d4-a90c-b8a976f052a3 · outbound

This paper cites A survey of vision-language pre-trained models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training A survey of vision-language pre-trained models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.759085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.759085Z digest=sha256:2712624b492e3ef9a52f8d2a6d580cfc6da2163a566da9698705c2ba548380db

Observation 2638719a-6663-4c36-8752-37b15a7b659b · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.762903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.762903Z digest=sha256:a8c4df876cca683825050d36858395de4fcef925d3443a7750a40b8c238630a6

Observation 318baf8d-a093-4204-af64-0384b002044f · outbound

This paper cites The Llama 3 Herd of Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.767272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.767272Z digest=sha256:56133183651a1d0a6e678e950a5a0615666350d07e1e33d1e2c476374f5d432a

Observation 17d48c46-1b18-46bd-8c29-9480716cab80 · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.771127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.771127Z digest=sha256:d53b83c333a4c2c11771a78b9832b1ed8ea619c05c71ba1d35f63c0b9e1570ed

Observation 2b3c34e9-828b-4093-9126-33a9a2bd0920 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.779889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.779889Z digest=sha256:a1192ccf14fffb9eb8c9f38783430baa7fec11d6f2f82e74d90e11609976a3a5

Observation ce61c136-8d4b-4514-8c5e-204acd8d0cc5 · outbound

This paper cites On Pre-training of Multimodal Language Models Customized for Chart Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training On Pre-training of Multimodal Language Models Customized for Chart Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.775568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.775568Z digest=sha256:6b04d72379189e2312cdc015776100cd07347557df87766093837ebbf712afdf

Observation 6d789315-dd49-4844-a28b-1c0980a06322 · outbound

This paper cites Making llama SEE and draw with SEED tokenizer.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Making llama SEE and draw with SEED tokenizer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.787389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.787389Z digest=sha256:2c17d22b1e0c54091379d33bcf36820b804656133828c5440d36ac8a085a723c

Observation 261dcc71-7fd3-4ff6-a3db-4693e805c0ce · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.783967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.783967Z digest=sha256:c16b6ff2c4d8869410f06d8dd1cb676acad73079c9e71386e3ecbc526e8ce8ae

Observation 3e8ba83b-2c37-4ea6-813b-205a423b23d6 · outbound

This paper cites GneissWeb: Preparing High Quality Data for LLMs at Scale.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training GneissWeb: Preparing High Quality Data for LLMs at Scale

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:49:29.045263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:27.799078Z digest=sha256:3c799695e18c864a73316f0fefb2922d72ab9f2ce48d6bb4b42a0a7029051c4e

Observation 2c1f37b1-9fae-4d97-b5b4-a15a1b0d5d06 · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.791070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.791070Z digest=sha256:d558894d624b221b1e9b621759ad1115a9e8d732301fa7337290830e9cc25279

Observation 90741552-0327-4828-a6bd-6b6511beaa7a · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Exploring the frontier of vision-language models: A survey of current methodologies and future directions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.794662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.794662Z digest=sha256:1f75f606ad3e5802bc41fc05974aba40be5005dc699f1ccd6c19ab776388dad9

Observation 02a9fbaa-accb-46e2-9ae0-697d8c48db7f · outbound

This paper cites InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.810010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.810010Z digest=sha256:b4c992d3e64e1d28d44864a97ed77823d00c1c941475ef67b1a81517b327d193

Observation f7140819-8ced-4d7d-b98a-4f199578fae2 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.802928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.802928Z digest=sha256:cca001f281c3503a5c22ac68b93627cc5944413cf0483f96e8c1a6edb589e892

Observation e98f92c9-47e8-4a08-a328-0c1cdccc6de9 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.806428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.806428Z digest=sha256:18bb7c9f7cb434d8a49998d58d05e986f71ff0735502a5ab72233f2399b5670c

Observation fd34ff79-fe5c-48bb-81c4-b443b111b954 · outbound

This paper cites Scaling Laws for Neural Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Laws for Neural Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.822792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.822792Z digest=sha256:276ad06dca27f8292db8cc805083b4320a5d0503e6c6380596beeb0a2d56d355

Observation b1cc2334-1a52-4202-9562-d6522531ad34 · outbound

This paper cites Compression Represents Intelligence Linearly.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Compression Represents Intelligence Linearly

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.814220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.814220Z digest=sha256:cef044c84bac19b31b99b476c824f7d06996db8971b95ae5726a9ecc43c5146e

Observation 5f83a591-95c1-4821-8c14-0b91a272b4e2 · outbound

This paper cites GPT-4o System Card.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training GPT-4o System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.818250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.818250Z digest=sha256:b4f1f7a7f71779d2685207664c553e2f8a89e38f5494452ef93f635b276a2477

Observation a3172c1e-ea2e-428d-96a7-4e2d469dff55 · outbound

This paper cites The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.833950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.833950Z digest=sha256:4296a70e08f790c44151f91058f465b0a08381f982d14acb834ba732007321fa

Observation 15b72efd-41e5-4a00-861f-fecff345027b · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Grounding language models to images for multimodal inputs and outputs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.826874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.826874Z digest=sha256:91dd9521f59c3b1e89d56e4ad2fc75a5a8684338e5c98d97575b5fc9aab5fe9d

Observation 876999fe-af79-4977-ba31-ae9939d891db · outbound

This paper cites RL with KL penalties is better viewed as Bayesian inference.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training RL with KL penalties is better viewed as Bayesian inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.830463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.830463Z digest=sha256:bcfa856fd647fd6c9af913e29d3cc281fbd572f2663b9f30410ef0b319bec108

Observation eb459e6f-d094-4a22-b529-c4d5196e416d · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Datacomp-lm: In search of the next generation of training sets for language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.844873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.844873Z digest=sha256:e137a9b59cbf28c8a5717872bd250f692393d6cc975bc90845d172a0e33a24f2

Observation 6f610eba-bf88-417c-9484-6b7154780b3f · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.838051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.838051Z digest=sha256:574059fda7bf4e1b10e337941bdcc3e7d67d63e08c39a922768eaeb4c76731fc

Observation c0160033-8a5e-4766-9c5b-caf5b1c93210 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Seed-bench: Benchmarking multimodal large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.841739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.841739Z digest=sha256:f550da77f672bda84d32c85e360ae90aa58e5ecdff15efd25706416d065a0185

Observation 7480b147-b897-4d8e-9dd4-b9379c5240ac · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Silkie: Preference Distillation for Large Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.856076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.856076Z digest=sha256:93370f17137ff19b5ac5952a47fcc79d5c10e9ff744dbfb1d9d7aeac9db48d82

Observation 4d435741-251e-484a-a5fa-5145de9af93e · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.848093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.848093Z digest=sha256:ca05b6ef1f367edc4c0fe230b63c510339218b0dca15e673bd154c4c21f4499c

Observation 83152920-85f9-464a-a8b2-9e820abe7f2d · outbound

This paper cites an unresolved cited work.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.851400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.851400Z digest=sha256:60bca6c29317c00a653f05b0a6197bef7cdd947df7200079d8abde38606307cb

Observation e389e1cd-a794-40d7-bfc2-84c10a3bb9d7 · outbound

This paper cites OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.867920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.867920Z digest=sha256:89151d655897d79858ad8e8123b3f4c8e03154ded119ea2a39cc07fa796e179a

Observation 64b974cf-03f0-4957-9ade-a54420a8ddf9 · outbound

This paper cites Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.860976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.860976Z digest=sha256:24a0b0c7cf851040901d5a54d97e162708f91111f2816e750049819e3515c3c5

Observation 0e1943bf-0a4e-4196-943c-f2405ecb5dc8 · outbound

This paper cites Visualbert: A simple and performant baseline for vision and language.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Visualbert: A simple and performant baseline for vision and language

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.864520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.864520Z digest=sha256:24291e7411272eb543faf9ef6093998da433f9f88d2bea767fd21f69fda66b9d

Observation 030a7c39-0ea4-48c8-985e-e6137171ad2d · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.879585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.879585Z digest=sha256:21cd654a1da8927e9e471cd1d49f237a2b3fecb06820bf9ec480e7eae31e8bab

Observation 5f3ca946-0d59-4336-ab7c-e5c123507aef · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Evaluating Object Hallucination in Large Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.871672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.871672Z digest=sha256:85872dec0e111723a5e2c941dfa1e9c297d307d988e1fb926289607c2b705092

Observation a24ff640-5790-44ae-81e3-8638138857b3 · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.875302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.875302Z digest=sha256:4a244b3f28ae7ec12be356d112dcfd9625ded4606af748084832ef736edeb09e

Observation d0b3e2b7-8bed-41e6-810d-ead96c0b64d5 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.994447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.994447Z digest=sha256:dc6ff67603bbc64a6788daa2dc9d11ef22271375a5814b6d50cc6286feffc209

Observation 55e44184-c986-4471-9ab1-d201a7d5b174 · outbound

This paper cites Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.884393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.884393Z digest=sha256:ba756361916dbe9727c618a5376d873b9219bdf5788b6c5073ad3f5be6393b59

Observation 64af87a0-4391-48d9-a7b1-402833b212ec · outbound

This paper cites Microsoft coco: Common objects in context.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Microsoft coco: Common objects in context

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.373240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:27.888232Z digest=sha256:9786695ec90d52c12046cbf981f413fa291e6d251c3666c699daca0c6675e6e0

Observation 0f72c81f-683a-4b5a-8cb7-c7b016a21ec2 · outbound

This paper cites Visual instruction tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Visual instruction tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.360488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.007067Z digest=sha256:58d209d3d5f9e8a3d5e771e45903fbb1be0fc3bbfd6189c2ed6ab26b7e027f18

Observation c8d007ba-729f-4e2b-ba97-02809fd98630 · outbound

This paper cites Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.999473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.999473Z digest=sha256:ceb21050460bb00993124db70af337501bcfb42c2a8d1cef7e084443cd729283

Observation 74826612-12bf-4a03-a56c-c979fb637aa1 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Improved Baselines with Visual Instruction Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.003360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.003360Z digest=sha256:7a3e70a4af99bf3eb7b553233c87a5a8a4774c4b5baa13645d3ab55c20884dc6

Observation 25e0a611-7da3-4a9a-a5f3-78cf6b062332 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.018040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.018040Z digest=sha256:cffd21ac4c241ad6bd066a13772f55e4323b1a629a458e79edc886b833cdc756

Observation 1c4e6354-915d-4aab-ade1-b9e1fd6fde6e · outbound

This paper cites Diving into Self-Evolving Training for Multimodal Reasoning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Diving into Self-Evolving Training for Multimodal Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.010620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.010620Z digest=sha256:8f77d3180b1eb7e509ef84add8a2df23f9f68be087622d53d79845ecec5ca541

Observation abcf93ad-1c9a-4057-a401-855473f41316 · outbound

This paper cites MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.014362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.014362Z digest=sha256:679f5370e28909dbd28e166c82cbf7c1a830eff5891c6c23ac1ce42ebffba3d6

Observation ebb61c47-cd18-4434-9020-7256098dcdf9 · outbound

This paper cites DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.035412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.035412Z digest=sha256:65b743f2ee2ee98ab62829f413e6af68c5716c1c499d40878af646d3e9641043

Observation fd2b56e1-36e0-4c13-ba33-2fda71e1fd25 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.022589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.022589Z digest=sha256:94401974be3a2729902d69af261ff7659b0a5b5e831f86b7a051f46696ac988d

Observation 13a1a594-15ed-44fe-968a-0f57024e18d4 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.027700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.027700Z digest=sha256:2da8133a35afd5ec48da2ed78e6be8d336213f4657eeee1a91959587770d7a92

Observation 96804bae-efba-4945-9a92-aa6c61c3e9ae · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.339876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.031470Z digest=sha256:7fb7570741d5e3b170b6b2e6fc873ee2da57ffb431bdba103850896e1fdd4748

Observation 85f6e33c-f575-48ce-8f74-8ac8a5a68c97 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Learning transferable visual models from natural language supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.051509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.051509Z digest=sha256:469d1ba981b7d75e2804f72c53ff0ad14c3a35e58045a36724299494f04e121b

Observation d6cbddcc-b0f9-4c4f-87ed-ae55c6cbc287 · outbound

This paper cites MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.039274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.039274Z digest=sha256:1de94e32f5190c538438667af31a199d88a4566caf3e6b2b805a6cc53f029c52

Observation c1d25b96-0c5f-40ef-85d1-08513a34091f · outbound

This paper cites Openomni: Large language models pivot zero-shot omnimodal alignment across language with real-time self-aware emotional speech synthesis.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Openomni: Large language models pivot zero-shot omnimodal alignment across language with real-time self-aware emotional speech synthesis

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.043623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.043623Z digest=sha256:8b120742c44c31d316b782f27dee8dbe4d11205d3f9e0ecf9612328c4636e0ea

Observation bb22e9f9-d6f7-4ba8-aab5-1e175efc7fce · outbound

This paper cites Policy opti- mization via importance sampling.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Policy opti- mization via importance sampling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.047320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.047320Z digest=sha256:53fff07da077614be55c05bcdc9cf29c62a1d1d00482cb73fc2d684206fa2bfe

Observation 1dbcd1a2-a988-47a4-a2d8-dd151e8fc9d5 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.066102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.066102Z digest=sha256:67cef67daac01c0ddaff2c6f409c2efb7688f2d53db2966dae351527036b033f

Observation b9abe44e-0b76-4db6-ac12-ddaf0d0f5925 · outbound

This paper cites Shaker, Salman H.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Shaker, Salman H

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.055287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.055287Z digest=sha256:2e703dc897f7b4d4b3ab6198cbadef2ca08689d5c8c789d0042f3f720e7065fc

Observation 3879d285-9344-4d6b-bf72-5688cc28d5f0 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.058792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.058792Z digest=sha256:0feead1104408155a90ea96bd7c87d1b611c619a2b1c195941a33cb9a404eeb7

Observation d6bea31c-8c9a-4b51-8623-3ed4649d3763 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.062362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.062362Z digest=sha256:8b2beee8a9078315507220b3821f1e16c3259746f3ef3957ea28d364f41268b8

Observation cb89ef5c-69ad-403c-aafe-b27ca3640b4e · outbound

This paper cites Emu: Generative pretraining in multimodality.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Emu: Generative pretraining in multimodality

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.290113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.080798Z digest=sha256:92a05770f9e227e95e135787d03da44a9b1d3c9bc82da4681820ab1c9a306bff

Observation 5b5cbdee-8ca9-4930-91d5-bebdfe1845d3 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.303589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.069756Z digest=sha256:6b0ff52b3e557e0913721c07b1f043a0f042e988415ac96ea9c919f7f92f47ae

Observation 360d5162-a319-46d7-8474-878065e5879a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.073076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.073076Z digest=sha256:7e05e4cb302249aaeb527401c7689b06068271614550bb16bd2e9dbfd18359ac

Observation 3da53c17-8322-42eb-b5d4-d3185b89ee71 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training PandaGPT: One Model To Instruction-Follow Them All

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.077016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.077016Z digest=sha256:bf1e445c37d349aa4abaddf2a5cfc44522515d4bc731d1e604d9011fb482ec6c

Observation 8ff4b12a-1022-4dd7-a36f-6792416b7a70 · outbound

This paper cites OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.277670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.095926Z digest=sha256:83123bef620d987f903735d39db5561f7ebf3774993539e336a596648ab6fc19

Observation 2df2a628-0da3-4ae6-bef0-8615714bda71 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.084067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.084067Z digest=sha256:42ff8e3289db7bd8f8ebc0a6255526b608cf4a8f227cbbebbdae2d620be70368

Observation 84b0381f-ec30-40e2-8293-c9407eb28d2f · outbound

This paper cites HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:49:28.700017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.087787Z digest=sha256:7afeb205ed30cfe2c5918fdbc429a2600b807f3b6ffb52c4ef06cab7087cd898

Observation 3335f5c4-4051-491a-b313-b5ff7f2c1351 · outbound

This paper cites Reconstructive Visual Instruction Tuning.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Reconstructive Visual Instruction Tuning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.091627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.091627Z digest=sha256:eed423462544065d5e761554fd378a2d7d60d49aaa81aa8ee134034d528edd21

Observation ad5553e3-56c8-4a2e-89a1-af392a36dd7a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.110794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.110794Z digest=sha256:bff9e62c5257bf6c2ea3b72086e155022793ed2def6dd0a80593ce43688b7130

Observation 1d93d7d5-4d49-4725-bcc1-7909cbddcb40 · outbound

This paper cites Scaling Pre-training to One Hundred Billion Data for Vision Language Models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Pre-training to One Hundred Billion Data for Vision Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.099780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.099780Z digest=sha256:cbd6a4672449ad1f4d1d31adcd0930ad2019cb919dbea0a94f76d0eddaef53b7

Observation d9749ed0-4144-48d1-ad61-70dc9e887f97 · outbound

This paper cites SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.103319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.103319Z digest=sha256:935ffeb8fa9849529253d66f24d8ef5fb3362f2a700e21a96b7407d21c2e5155

Observation 0d837735-db69-4a1c-a801-adf9a07ce6de · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.107044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.107044Z digest=sha256:518bb72fa0efcc797ded52e1ea1f8625e473d39183396cc14fb36eaabdc39aae

Observation d5f5d2ea-f6fa-4feb-a59e-cef1e23e31cd · outbound

This paper cites Knowledge-augmented few-shot visual relation detection.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Knowledge-augmented few-shot visual relation detection

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.257203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.125115Z digest=sha256:2416178a7bef7377902773363e4432f1b4b4bd4a6f850ee0004dfb6ac31dadbf

Observation eed933be-e5d8-4304-81af-4f50737727fe · outbound

This paper cites Qwen2.5 Technical Report.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Qwen2.5 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.114329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.114329Z digest=sha256:62c9748a39ea990855f613ea53ef2f97cc408c77c71ef54e9d851b3526a65e83

Observation 290d4613-fbd5-4354-b505-ccf4a1af24e1 · outbound

This paper cites Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.118433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.118433Z digest=sha256:46bf43b2a157ea37e6273278ec57b191c0e23472838286d4e241aae59306b2ff

Observation f3c8e3e5-187a-4f61-b571-83731f8e9d7a · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Capsfusion: Rethinking image-text data at scale

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.121972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.121972Z digest=sha256:634c013cd44528f037857f55396d03b10dd4c75d83337a90277c570897d643fe

Observation 5713fdfb-d4ab-4fe3-8ad7-32ea69cf88df · outbound

This paper cites OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.140238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.140238Z digest=sha256:04362d099c4fa4bab67d8a574a52324ee836147c6316f2be5b13f19d11b6443f

Observation 048a03d7-ca78-4126-a150-c8fcb0b526ad · outbound

This paper cites Xing, Xiaodan Liang, and Zhiqiang Shen.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Xing, Xiaodan Liang, and Zhiqiang Shen

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.244964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.129074Z digest=sha256:2a1e80d15fbb660640020765034337e5b7425d33952483689ab54933706c2061

Observation 3aea2af3-136b-4e7b-94b9-a7548592bc4a · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Vinvl: Revisiting visual representations in vision-language models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.132406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.132406Z digest=sha256:d590a67d5eec97cffa289c5ccb08b7b6f5146a8559794f62ecf35db865142ec5

Observation 443d7519-58aa-4878-8b9a-b7cf2ec0812a · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.136196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.136196Z digest=sha256:d4ac03296e3d89c3bbe6bb5c2b959b1983b1a7b50ab884bc2a664d0707c9805b

Observation e80c6efc-31ac-41fb-aad6-ac2ca9b4bde8 · outbound

This paper cites Minigpt-5: Interleaved vision-and-language generation via generative vokens.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Minigpt-5: Interleaved vision-and-language generation via generative vokens

Reference 95

Resolution
verified exact
doi, observed 2026-08-15T21:49:28.240603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.144082Z digest=sha256:7b3387c64377c01ae3f7189e060ec1fe69006fd67719ac002395c7cead5b1a78

Observation 0b3a4518-f49f-4699-99a1-4af729b51b37 · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:29.232823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:49:28.147999Z digest=sha256:a22fbb0c357339348ceb24981c272563a4cb0063b39dfdf07cdb36bc3ada3db5

Observation 87d92ebd-c559-40de-af3c-bd1e0548a290 · outbound

This paper cites Generalized decoding for pixel, image, and language.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Generalized decoding for pixel, image, and language

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.151834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.151834Z digest=sha256:cdcc4b8ece53118a2f801e44f011312f701cf7e5c098a2481b37c1fa8ee3e562

Observation d08c95db-92d6-44b6-a2db-cdbd44d42d01 · outbound

This paper cites CompCap: Improving Multimodal Large Language Models with Composite Captions.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training CompCap: Improving Multimodal Large Language Models with Composite Captions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.718480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.718480Z digest=sha256:36fd3d90798bae861d2234faf4dd2d085c6742e95f0d8234fb142e22081cc67c

Pith citing papers

No inbound Pith citation observations are available.