Pith. sign in

Paper Citation Record · LEDGER

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.07584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07584 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:22.245545Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4476f292-0ae6-45be-aa36-0a19615192ed · outbound

This paper cites 2022 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2022 , publisher =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:23.035059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.069655Z digest=sha256:1bc24ad3f0a28773d1305065ca385b736c329fb3a11a07c6e2c5f96edf973ba6

Observation 64d35d65-af2f-4476-9bc7-8600396687eb · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.078210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.078210Z digest=sha256:fb8b4fe4ba69ff4b462acbfc8505b24a404195543393de64b6365c10ce9a7606

Observation cd3324ef-fff3-4469-8f7b-967e10675454 · outbound

This paper cites Lawrence and Girshick, Ross , booktitle =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Lawrence and Girshick, Ross , booktitle =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.990804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.088566Z digest=sha256:f0be39d73276b05a3973a6c4276afc1100f6ffeaa233244d6cd1b9d3291d3fdd

Observation 51e95402-573c-4d2e-8e2a-c49209618063 · outbound

This paper cites 2024 , doi =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2024 , doi =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.969167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.096685Z digest=sha256:ab862051783958783c7f30a0380594ff3357c398c9f1f4eb6898e3835c2f33d5

Observation c5f6c67b-10d0-45bf-a23a-ddc118cb8283 · outbound

This paper cites 2025 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , publisher =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.946307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.103389Z digest=sha256:221b8775e88fcd9cf199ec29b44bb5d86fdbff2bf0d0effa61aaa914899c71c2

Observation b4265b22-c5f0-4ae5-bf4f-0d0147d6d303 · outbound

This paper cites Journal of Data-centric Machine Learning Research , volume =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Journal of Data-centric Machine Learning Research , volume =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.928124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.108889Z digest=sha256:6c6fa60923fb19ee595791ec2e5132b0302398cbad7664934dbe8d28e54f9fed

Observation 387f56a9-c392-48cb-9e1a-109170f7b229 · outbound

This paper cites 2019 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2019 , eprint =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.114861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.114861Z digest=sha256:fa305d1299e364ac970937f6d11f31ec73ee676788914bdf023fa569615071f4

Observation 16cf6ca3-e2f6-4e97-8a37-0cf641434036 · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.121965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.121965Z digest=sha256:e5d180804516f40b112a840210cef494ab81c27bcdc29bad2d08dd9d7ab850a7

Observation b6432445-2633-4b02-b7c0-d69144babb8c · outbound

This paper cites MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.128677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.128677Z digest=sha256:aa94c6c9a125457a499c377b41e1ffa970f468bac05f38dfdc0dc1e0af7f3363

Observation 59df4741-b015-4241-ad15-04b1e257e67f · outbound

This paper cites 2024 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2024 , publisher =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.889531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.134634Z digest=sha256:130953a95334b2887434f4cb303243364d4d738aacfe437dcacb2e05a895f785

Observation 2ead2ad2-c935-42e2-a944-dc8ad8a058b9 · outbound

This paper cites 2025 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , eprint =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.861618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.142290Z digest=sha256:af7c492c545eac128262752636d53ae703a16900f70ca2a4697c77337ee6cffe

Observation 6964239f-fadf-4c63-bfa3-2830d73f31de · outbound

This paper cites PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.148366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.148366Z digest=sha256:ad3f05a526398a8b856aa7b57baa209eb9f5be8425e12a184e75d7219baccc73

Observation 6fdbe709-ec3d-4eae-a840-88dabe50aeec · outbound

This paper cites 2026 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2026 , eprint =

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.842026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.156416Z digest=sha256:83a62ba5490dcd479bc48bb957d6e494a095d29717022897b7f2250ade75628c

Observation 7c5160d0-5f51-4c7f-994d-825831750f7c · outbound

This paper cites NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.162543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.162543Z digest=sha256:e6c70e3c390df38514c8b5414502ba38129bf7787408dafb77b509de4afc0d08

Observation 0ae47376-8cfc-49ae-b668-44c7468c4ec9 · outbound

This paper cites 2025 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , publisher =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.820733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.170552Z digest=sha256:17f73a405553dc37c526e6f85970b10ee084ca347821cea2c4aeb32b3376fd1b

Observation 801da51e-d348-486c-acc6-532317dd6cee · outbound

This paper cites MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:36:22.517208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.176532Z digest=sha256:a98700743ae4e3c9ee30e99661e65fac569947bde4acb9a9465387baf3e620e9

Observation eae8b963-b47e-41b7-a638-46c48763c3fb · outbound

This paper cites 2602.08367 , archivePrefix =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2602.08367 , archivePrefix =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.184820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.184820Z digest=sha256:74dd4bb149de53c8eddda754c0f7d38f6b7b8262e20d5fc059a4de0b4f2d6a9c

Observation bb30cc29-e875-43cd-ab43-ff33ec22c586 · outbound

This paper cites 2026 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2026 , eprint =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.801065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.191335Z digest=sha256:7cb29aec96801bfee64b4d0b3003249389fe6a18b752cee767715327516c6d95

Observation 9f7642af-c173-4bb1-86d9-bf0125a3b057 · outbound

This paper cites ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:36:22.351158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.197364Z digest=sha256:15060897ed2ce3dc7c982754cb115dab9ff5124a2a6765a96fe29266afb9250f

Observation 8cb0bcc6-4c1a-4a03-878e-d14bcfffb816 · outbound

This paper cites 2025 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , eprint =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.781866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.204494Z digest=sha256:2a7cb08dcad4a5bfb8bb391e33f4b7ef9f451c6c5da96a5a8fa318f51e7eabdd

Observation 34c41c24-873b-457e-9a62-02262ed09492 · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.211047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.211047Z digest=sha256:2cecb38de05563555af120c78d8b5b17391eecc8a8e94bcb12e62eddeffbf793

Observation d87097a0-a98e-4758-9d20-9779e2db31aa · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.216429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.216429Z digest=sha256:a5826990136bd61922d240857399b1d502d28dbcde6ce481f8583d899dd45139

Observation 2577fa03-b8e9-4dcb-83bf-5749fee3e8a5 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.742909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.222214Z digest=sha256:8bfaeaeb4c527ef8edd28388cda9a5048c654da1352b66706075ae9277d01e58

Observation f19be4b3-7942-493d-84b7-0c85d85eb92d · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Advances in Neural Information Processing Systems , year =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.226802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.226802Z digest=sha256:11c5cc3df4117ee47edb30c982a76586386a04f578b74f4cb9e49c2ff6f2e126

Observation 9c4f1d09-8883-45f5-8ae2-e9bb2c002866 · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.232109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.232109Z digest=sha256:bf82f70be7dd528d4e816ce1e8933477c8a26f750db42d70d8be639ff5d0c3a1

Observation 84ad65ce-7676-410b-8f3d-1f48132d4a42 · outbound

This paper cites Complexity of Computer Computations , editor =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Complexity of Computer Computations , editor =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.686421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.237385Z digest=sha256:9559055fbb01c8d4e8997e0b3c38f3effab6b94f45e778c305bfc0ddca6ed43a

Observation d99f49fb-09d8-4954-bc74-e5ae756b4aba · outbound

This paper cites 1979 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 1979 , publisher =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.664591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.245545Z digest=sha256:fa019fa734cabc20df92e93a15cb3b8e1563277cea03f6544f71094743b5b086

Pith citing papers

No inbound Pith citation observations are available.