Pith. sign in

Paper Citation Record · LEDGER

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.20156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20156 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:51:42.939404Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact6
  • verified fuzzy19
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e6b01c7-ae0a-4c47-aef8-eb42e524c7a7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.792055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.792055Z digest=sha256:60851af615729c6b86f0f69f7e0527c39cf9246f286cbdf3ef32d5cacac9399f

Observation e00b4115-514b-4eee-95e4-1bb26842c2bc · outbound

This paper cites Wang et al., Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution ,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Wang et al., Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution ,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.777817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.797227Z digest=sha256:f1adfe59ac83d1288bd51bc5eaf6e8953fddc45b4f68750a63f352ee4a61ab24

Observation 183d1dc8-1e3a-4d5c-be39-1e6d9576aa0c · outbound

This paper cites GPT-4 Technical Report.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.805659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.805659Z digest=sha256:f5d4aee15549146a2cf720ed88fc8db434ab20c71fd75c8b3eed1462f5ba9cf7

Observation c2c8b985-9a1b-4bb4-b3dc-5917fd79b9e6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.809672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.809672Z digest=sha256:7b15a4a0d58b6e09beae545054f92a560e98d1aaa9ce69115642e2cc7ca6fc23

Observation b1318520-c311-4073-a0d8-7985c669838e · outbound

This paper cites LLaV A-onevision: Easy visual task transfer,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality LLaV A-onevision: Easy visual task transfer,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.765837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.814331Z digest=sha256:8a60020ab983eb66674f38c17a2defb5265b0a55a935d21f62e892684fd1f40f

Observation 73b8b5eb-d6ea-4da7-9035-8b5ca9246bfb · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.818265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.818265Z digest=sha256:e3c3e007966b60fa62d22c51de5879e4ba86da5d4682f0d83e0ccf2e77b8f6ea

Observation 336b6437-7ad8-4a96-aab0-9fafdb0b7ece · outbound

This paper cites Masry et al.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Masry et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.753778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.822523Z digest=sha256:aa8e5071e5e506741b15bb9c3a6443fe60fd6c69fa4710a5039e3fdbd8362e67

Observation 1d86c0e1-6bd8-4667-8769-8edaf95e3ad4 · outbound

This paper cites Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.830659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.830659Z digest=sha256:c4cc7017ab62fed50f9fcfa9e10e3f4a66b5d4da79018101d27f45d2e7c056df

Observation 6b92f115-f401-4d14-ad10-ef1761793569 · outbound

This paper cites Mplug-owi2: Revolutionizing multi- modal large language model with modality collab- oration,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Mplug-owi2: Revolutionizing multi- modal large language model with modality collab- oration,

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-15T17:51:43.543200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.835619Z digest=sha256:3bd4a1168c5aa54cf7413998de382cf1e3d1f730ea523b4273c42704747c14f1

Observation 78af3003-449a-47cc-91f8-2db9ee342ba2 · outbound

This paper cites Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts,

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-15T17:51:43.480475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.839683Z digest=sha256:b094460899791be186625200534420d9ae33cc706e0f76927b201b1a4fb73a07

Observation 473a62d3-b1fd-4d9b-b6e9-383c663df6d5 · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality What If We Recaption Billions of Web Images with LLaMA-3?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.843903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.843903Z digest=sha256:c2f2384d4a0774d9ff3531c8a99507df37a9f6debb0696e22401cb0db66bf9f0

Observation 34c429a6-d137-4f20-96d0-0f6be4310713 · outbound

This paper cites Obelics: An open web-scale fil- tered dataset of interleaved image-text documents,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Obelics: An open web-scale fil- tered dataset of interleaved image-text documents,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.741786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.848435Z digest=sha256:c512e125da1fecb47d491cc3958cbf9bb1f72c662c3dd92af674cb61c90bc093

Observation 9e104414-5404-4c31-9018-ac5f49acf053 · outbound

This paper cites Bai et al.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Bai et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.729804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.852105Z digest=sha256:b84f77ac50409e9dda9277ef0623b004f0e6549b2b3b5c170034288319cdf315

Observation 0e493f98-bf3c-4af0-b44d-12a222dfb98d · outbound

This paper cites Towards efficient visual-language align- ment of the q-former for visual reasoning tasks,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Towards efficient visual-language align- ment of the q-former for visual reasoning tasks,

Reference 14

Resolution
verified exact
doi, observed 2026-08-15T17:51:43.158871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.861283Z digest=sha256:ac41b9b48226f23c3f20d1c81fd420eab2c3f50724c0809928384ba3555808c2

Observation 6e260517-0b60-4f18-a560-8cb159f743b1 · outbound

This paper cites Training language models to fol- low instructions with human feedback,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Training language models to fol- low instructions with human feedback,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.717929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.865159Z digest=sha256:4aead06b6410e1704e9b1b515c7af37a504c6cd16444b58397532bc8b19adc65

Observation 54f10874-2138-4c03-9c82-14b38eb9cb4b · outbound

This paper cites Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:51:43.176356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.856037Z digest=sha256:94a0174b0ed8eadf992c9371715d9217fc6755bda885b0974fd0e5d8150c827c

Observation 2ec930d8-ce85-4785-bf34-a4aa211b036f · outbound

This paper cites Lima: Less is more for alignment,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Lima: Less is more for alignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.693170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.872904Z digest=sha256:a241b282f0f088794712959f4bfc4dca9e67be8b88277904921ed8ece30fcfd8

Observation 00ccbdd9-15cd-48ed-8bc3-9d6e4dc55e88 · outbound

This paper cites Agarwal and D.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Agarwal and D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.680958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.876778Z digest=sha256:90c47c2b4344d6e979f0b1cc1178115ed2840e2e06c01a6d018e87427fa98db3

Observation ed77cb0e-c99b-4a29-b816-52dd411f3a29 · outbound

This paper cites Visual instruction tuning,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Visual instruction tuning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.704893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.868943Z digest=sha256:8534ffa285fdf74670e0694d5d8e77ec24bb15dc0442052c3da92cf61b863b24

Observation 9d9c0583-18e0-48b0-a1a4-381266f45c07 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A Survey on Hallucination in Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.888839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.888839Z digest=sha256:1b2b8f7637126536e15f607a268f4996603bcb35257faf94f37f2ce9c145175f

Observation c5d58d02-0272-421e-b5ab-3d5a4059181f · outbound

This paper cites On the origin of hallucinations in con- versational models: Is it the datasets or the models?.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality On the origin of hallucinations in con- versational models: Is it the datasets or the models?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.669088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.893005Z digest=sha256:b08d9a017a0cfbd1c5998d1f4c3caf92a8a9bc9f2d00962496e427825d80335b

Observation 4b5aff2c-9fb3-422d-887f-aa0e71838b34 · outbound

This paper cites 48550 / ARXIV.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 48550 / ARXIV

Reference 22

Resolution
verified exact
doi, observed 2026-08-15T17:51:43.145209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.880779Z digest=sha256:70f64606d61684b5e9632659de55af84aee8583829a9d90d13e0c3df153ec5e2

Observation a5b05a76-f4e4-4ee9-8d8f-eba98670ca9a · outbound

This paper cites Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.884682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.884682Z digest=sha256:4c1d62674f04f3ad0aa7e50ac399544a2e098807c8a46b0bb76e0c05eeda848d

Observation 9137c9a2-2555-463d-9124-98d294a67a3f · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language mod- els,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality MiniGPT-4: Enhancing vision-language understanding with advanced large language mod- els,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.644770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.905045Z digest=sha256:1c64ea3bd04dc308a331a21545a2a587d27fb9ec5c441589cb4ca7e9fe5145dd

Observation b65c4bde-955a-4669-bced-69f5fb7df324 · outbound

This paper cites Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.632535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.909037Z digest=sha256:e9e1ee5ca2580eacaa150e6cbaecd71181c6529e0bbf59db1e8af92bc99b95e6

Observation f842be4f-5590-4ba2-a228-f4a8114b7da0 · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.896673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.896673Z digest=sha256:16d48018508cc60cd6e630faa842bc45dff1706b0b424dce5fb41b7c9502888f

Observation 86302a80-cbc6-4715-ac71-616b76a72c2a · outbound

This paper cites Alpagasus: Training a better alpaca with fewer data,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Alpagasus: Training a better alpaca with fewer data,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.656968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.900720Z digest=sha256:db33c8f848f13c0232fd4d54576842e616e3ce90aa9c9b95c8d805348b7bc935

Observation b7cf5628-da59-48db-85ee-ea0ce844e3bc · outbound

This paper cites Language models are unsupervised multitask learners,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Language models are unsupervised multitask learners,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.607479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.920881Z digest=sha256:aa6ae0006d48f6101e9ac27f90998026c63c7483cb354ac0e4a8a49f1553af0c

Observation 4cd92eee-129b-4164-9e4e-b0d6cd0874c2 · outbound

This paper cites A frustratingly simple approach for end-to-end image captioning,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A frustratingly simple approach for end-to-end image captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.582723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.928389Z digest=sha256:505a51a1fad4b3dfec8bef3020d2266c4c29cd6ac824370281d02063204659a3

Observation 9b75eceb-4c74-4217-bd8a-9a4b8eece9c6 · outbound

This paper cites Learning transferable visual mod- els from natural language supervision,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Learning transferable visual mod- els from natural language supervision,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.620012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.912892Z digest=sha256:c627b583c06ac591ab8f46fad191b524e15ded13870c8149ba8b1f9529edaee9

Observation 5a46c7e5-1d23-4691-88a8-cc08dfe4f5cf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.916590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.916590Z digest=sha256:11f7364ad6b2024bf7cdf0bcc81bb64978767abba529b7725f6c87cd794dee45

Observation ec5e046e-f3f5-49b7-8e2c-91e7c739fae6 · outbound

This paper cites Available: https://cdn.openai.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Available: https://cdn.openai

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.595346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.924572Z digest=sha256:97c985264e8aa86a0394a7cd97e90c87b9fece83597153c1b177e8110d7d59e5

Observation a81406a3-565c-48ea-904b-866e719a8904 · outbound

This paper cites A survey on enhancing image caption- ing with advanced strategies and techniques,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A survey on enhancing image caption- ing with advanced strategies and techniques,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.569802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.932021Z digest=sha256:145567d705565782ff09f562bbe7eab7192a5944e475a3ab94e84e39efa5b7e4

Observation 2a3bc4dd-9bb9-48eb-ad34-b5cb2aa1b190 · outbound

This paper cites SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.555538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.939404Z digest=sha256:b52210668d93b49ee0191710bbb56a23775ea571a25f5de4bf9fc41e2e5954c9

Observation 57702283-bbd4-4e3a-b763-cefa40cf6a70 · outbound

This paper cites 32604 / cmes.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 32604 / cmes

Reference 1506

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T17:51:43.399739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.935633Z digest=sha256:c2999e3ad6336db154aae2eb71d2d6b6601d2889a39eccf347af2d7016f16c3b

Observation a99d609a-3c26-45d0-afd2-016f4732510b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.801311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.801311Z digest=sha256:f201f3561f8ee0a22b8cffc8657e1c95498f1fcfc67a3ba2b44c7662199c2ca2

Observation bb5a050c-5f47-489a-928f-c215d8208db3 · outbound

This paper cites 48550 / ARXIV.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 48550 / ARXIV

Reference 2025

Resolution
verified exact
doi, observed 2026-08-15T17:51:43.260921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:51:42.826469Z digest=sha256:dcb1ea8f5312beeabc9685f8c68c89ffa5b4b6ab941fe2ef72129531e83d135f

Pith citing papers

No inbound Pith citation observations are available.