Pith. sign in

Paper Citation Record · LEDGER

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2505.20728.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20728 v4

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:34.081389Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T06:34:56.032634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T21:11:15.085382Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c399578-31b0-470e-9f41-1b4bc570dda5 · outbound

This paper cites online" 'onlinestring :=.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.566175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.566175Z digest=sha256:75bd3f83b41618037a3f24d5a38fa4feb282ef0d72b59287447757b7e12e90cf

Observation 9d151fad-124c-448c-802b-279daa1da518 · outbound

This paper cites write newline.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.634571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.634571Z digest=sha256:139874a6cc02c2fbb9f234b463722c4f27cfce951e0dba2025856da006c42b2b

Observation b7040f26-b49f-46db-9f98-0ba1569ab0bb · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.732806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.732806Z digest=sha256:48dccbc8b70f99389e806db92c4eba8f71a3aafc213746b6a5f4e7b2270fbe5e

Observation 752de2f1-8c9d-4fd0-b9d1-e72e290aab7a · outbound

This paper cites GPT-4 Technical Report.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.813020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.813020Z digest=sha256:17d29ead668ca2ca5243648b647275295b10392e38ba9cb559b9d5bc69fec5a1

Observation 915d26d0-19a8-4d7f-8af0-09413637133e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.868504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.868504Z digest=sha256:db5bf6209ce8ad61302dfda05626be9d4f6ba6e8a73449ad7442ef0e8003544f

Observation a998c525-c809-401b-a93c-920b0ea44347 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:30.952646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:30.952646Z digest=sha256:b82dcff9005317d7eff65c87536035f8c1e5b1dda750e531ac5b22f7ebe1d709

Observation 5d9339a7-f16c-43db-b718-064548533b5c · outbound

This paper cites Qwen2.5-VL Technical Report.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.034463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.034463Z digest=sha256:0e0cc9054e799678ba6a9d280efe3d1006840f72ccf2d1103d3bd6ba4582edf6

Observation 18f9b066-a05a-4a05-8efe-ba769070c060 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:36.299628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.115810Z digest=sha256:08411b80aab7a65113521a0e3aab196a0022276124b6f1c5c8814a29e92e38ec

Observation 83229649-630f-4331-9b98-ad31b15475fc · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:36.156624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.165621Z digest=sha256:3a4a87b5fb2fd9bc525d7c55ab42e0ac46f571f271ce7249a84c7afa013adf6c

Observation 0b6d4f18-972e-4d8c-b198-074d5dd7cb52 · outbound

This paper cites Aya Vision: Advancing the Frontier of Multilingual Multimodality.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Aya Vision: Advancing the Frontier of Multilingual Multimodality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.236941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.236941Z digest=sha256:6e5c7260f2778793034d0645a05c96d39b9a5d41ea6953110d1276bec65a3845

Observation 88c75eb4-923e-4bb7-bbb5-8cb2c3f30022 · outbound

This paper cites Kimi-VL Technical Report.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Kimi-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.301410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.301410Z digest=sha256:efe5f7ffee37dff21cbb2872e2669cb87b5b8e601b7203086e392c4dfcda70c4

Observation d67f9f9f-cb3a-43ab-87b8-be3144fcc46f · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.970655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.377792Z digest=sha256:8acd66e2335a2c82fa3b1d64ae2a2e2bcb99c2548feda974a062956c12ebfb3a

Observation 4b860dae-a052-4a8b-bf93-b134355876e7 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.444781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.444781Z digest=sha256:2eb828770f33ec0bbc8d575524664526053cde8df42d87322b9ed1edc8f6231a

Observation 1e2bdc44-4582-4ee6-866a-cbbed726afb1 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.779086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.516823Z digest=sha256:5370509bda65bc91eeeee0b6f6b93a70913e4e208db526bf543d66362b8f6b58

Observation 56b1f494-4cca-4dd8-8989-6add8b57b1e8 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.590650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:31.595020Z digest=sha256:fea5aa7d74ded343cf4b29e3ac9b27ba15078d3f62967ab23a46499f0c64df71

Observation 06300119-8468-4e4b-b66c-f7cbd0fe8a1f · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.686596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.686596Z digest=sha256:3a0af10b8b722b88e669024415ef6cf3c30de75229d6b8d9ebb9c666ecf9eafb

Observation 8d8cc75a-6a49-4d13-b764-86b738695d19 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.769097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.769097Z digest=sha256:d308edfd72c7b9f9ce2cfe078842502f4469cfd499dd58fb75de540d0aebf175

Observation 1751a488-47f6-463a-95ee-3434fdd5a170 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.864485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.864485Z digest=sha256:27bfd9af02a445df2672adef3d0feb1cbdbef2a2852fc41f5bcca9b059f0961a

Observation 98096c57-9bd8-444f-a72b-8ac62487f656 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.905490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.905490Z digest=sha256:0eee0d6f7947714e7cc45187cb271739f11e76cacc3e05ef99b62eb446374d26

Observation 8835969b-b7aa-46d7-ab34-527f0d1550ba · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:31.978776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:31.978776Z digest=sha256:d0020cec3592bbab6b5ff08b8516746bcfc15bbb922df08be17aba2d530ade02

Observation e95c8114-bbb1-485f-a17f-6ce584870e4e · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.068720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.068720Z digest=sha256:ea2a6de852d0003b8f24e4dac90373ff6fddb121c9edd9aa9a06b987b545a5e5

Observation 001d6cb3-83f1-49ba-98c9-d03602fda910 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.334929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.171204Z digest=sha256:7ea83598f8d8a84e2005348e652cc9d2527c9196cc695a907537586b858cb1ee

Observation 601b543a-3e8e-40a9-ba87-49474f2fca34 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.138981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.242418Z digest=sha256:97ed5481156f84fce0a306eaf7fbb8e82883e470148a82ffbb5f048db7ba1f2b

Observation 0f972e68-e574-42ee-a6ed-21d647e6ff52 · outbound

This paper cites CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.315784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.315784Z digest=sha256:ce063bb05311c45584733c9bc766ba5a1b57447a2df00d660db1643e67b212f1

Observation ead65715-e649-4dca-b20c-e1f20b6c4f0f · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:35.032708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.389330Z digest=sha256:e460ae0d0440a1e4f8be2ea2c546ac03386b6d981186e4c9da36c075ada0f9dd

Observation 5c97cc13-0618-4740-880f-af2c42d94a96 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.470962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.470962Z digest=sha256:964dffdf2d74fd95e4b5e74c2b2453ecd4cd73a8401046af43e05df899db8c1f

Observation ffbe223e-d9c3-4ae9-9117-c5fa30c22be1 · outbound

This paper cites VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.535373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.535373Z digest=sha256:4d99b2ddba9f25ee2ac0729121184170d60366c9529c3bb68032ad3081c5a8df

Observation dfb006dd-2a19-4373-a98d-e6b997626f22 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.608633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.608633Z digest=sha256:de04e0bd7dbd71046ed3341ca4c4aa66a17f69dadd248881149f2f36d038ad33

Observation 874f6b88-88ce-4ec7-a840-424a733f8385 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.676073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.676073Z digest=sha256:4daa23e8dcf81e9dbab797e05451254f9f344ffee1b9b25de95d334be1e415ab

Observation 4eae3f36-1e9a-4657-b9a4-c3b516b3d359 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:34.893955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:32.758845Z digest=sha256:d7b0a331d0c85fb33135013f317a73151e08aefd8d5689658bac21d86010c883

Observation 62794210-eb1e-45bc-b7fd-b271d7ac7dc9 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.851572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.851572Z digest=sha256:cadc4715fd0a3e5234c7198ba4039919228668b743b83a38a2525c5fbe6a7b60

Observation dc85e25c-5131-46e0-a594-a8fb0bacfe77 · outbound

This paper cites VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.927998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.927998Z digest=sha256:a220b98058f8153dada24dc25c08e8d72a5eec865e2d8752466920a161e9602f

Observation fc5d705f-b0e8-4de9-a506-d73312103c29 · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.010530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.010530Z digest=sha256:904f0e68479f7d44da668fccfb69b6d3f8c2b2b352ac9558348fb5e70a45e44a

Observation f462880d-35a6-456c-a477-9290ac0be834 · outbound

This paper cites LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.120737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.120737Z digest=sha256:e7155e0f4359a4c46596c40187786b3d4c23cf4fae75527dea4e62676ca917e5

Observation d5a87d20-629c-4fed-a2ca-1aa7cc8a503d · outbound

This paper cites Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.214128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.214128Z digest=sha256:6fb66db0974706e57cc9eb53a2b581e1f7140389b28279014420351f458945a6

Observation cf46d1fe-c7ad-4a60-ae0d-05f05bc2541a · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:50:34.743151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:33.357115Z digest=sha256:f3b0ae0465def188bdd6b68019afd9fb997bfdfe7265f16d78cf21a231e8d9e6

Observation 5594654b-632a-47ac-b866-0c18d131b687 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.477473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.477473Z digest=sha256:7cbbc3d296dd7e043715289bcaa90797221a923a7003ca313afda96c9b21903c

Observation 537d74e7-7f34-4249-9156-1de218b34bdc · outbound

This paper cites CogLM: Tracking Cognitive Development of Large Language Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CogLM: Tracking Cognitive Development of Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:50:34.334219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:50:33.577929Z digest=sha256:d37bfca9be6cc66be7d4b31311660953cde2933e6738aa8625868ff2f909c7e9

Observation 1633c790-48b4-4a8f-ad8c-b011cab5840c · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.703965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.703965Z digest=sha256:515ba53afc1824fd6ac85f7374d25a4123c3c68b50b63f8eb9bb01a4261c09d4

Observation ae6ea074-f45a-4264-af88-2e2f6d623eb0 · outbound

This paper cites an unresolved cited work.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.819280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.819280Z digest=sha256:2c5ec3f193ee3de02cdbfec10db7c7df9770e4fbf8f99243d4930b4582348dd9

Observation 01d53ced-5ee7-4b5e-b343-a35ca024c4f8 · outbound

This paper cites Redundancy Principles for MLLMs Benchmarks.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Redundancy Principles for MLLMs Benchmarks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.938748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.938748Z digest=sha256:d37975bd24971a9c9bf2c7c55140d5a31e31138da0f7ca27e4ebefab7ac19741

Observation d9652d4f-5803-404f-bfbe-04cbc5d6017a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:34.081389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:34.081389Z digest=sha256:1055e36e148faf836971c0773622e0f11e25868ac9552c73a387f86e23f46e4b

Pith citing papers

Observation af191167-4ede-47cd-87e2-f3dc88dfe650 · inbound

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction cites this paper.

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:15.133893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T06:34:56.032634Z digest=sha256:5bbee7786ee9b1ab26941fc6be2490b5e3e14d857e7494686e8a31467106e593