Pith. sign in

Paper Citation Record · LEDGER

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2504.15485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15485 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:29:01.175213Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:32.315784Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T16:44:56.123044Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy31
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f31bd67-203f-4f7d-88e3-cbd31f32813e · outbound

This paper cites Llama 3.1 model card.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Llama 3.1 model card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:02.034118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.901917Z digest=sha256:cca2240c3e2fdc48c1851e6a6cc497c6e627fe511f78d64b5c676840d914c04e

Observation 6499a395-ef09-4816-90d8-3637fc2a5575 · outbound

This paper cites UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:00.908466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:00.908466Z digest=sha256:b3d827870df72277082020bc6aab663a55d922f5abd587c0f15695ff8708f89f

Observation c776e55d-2837-410e-bbb0-7cadd595de7e · outbound

This paper cites CountGD: Multi-Modal Open-World Counting.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting CountGD: Multi-Modal Open-World Counting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:00.914353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:00.914353Z digest=sha256:d3e287560de1e30e299aa8fc2f1dd5df50ff9677fedf18fcea200199e17a0b08

Observation fc7c5e2b-8c77-4c54-a770-dda9cf0f1db8 · outbound

This paper cites Vqa: Visual question answering.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Vqa: Visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:02.019502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.920057Z digest=sha256:e9a11a9633d9fe2c680b1c4e8883b0b144824bdb70c4951ddd78fe97e7665279

Observation 6846673f-be1f-4b85-9144-d502784f0bc6 · outbound

This paper cites Image amodal completion: A survey.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Image amodal completion: A survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:02.005506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.926262Z digest=sha256:f45b551a04eff1dcbc4168f0f5485e753bd6292878c982a74cfacd57140f35ce

Observation d14a7d27-4ea5-46cd-a8ec-0c5c370ef6b5 · outbound

This paper cites Open-World Amodal Appearance Completion.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Open-World Amodal Appearance Completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:00.931495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:00.931495Z digest=sha256:5ce012a82c76481b55ce27c286630a7873434328f2ec434a4d438ecb4a8dc2d2

Observation c968a922-a7d0-43da-8583-453e9b970118 · outbound

This paper cites Text to image model arena, 2025.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Text to image model arena, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.990689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.937750Z digest=sha256:c3499f977a41d09d15a982e110088e4943c1f4ab0644bdff4c8a2589b20b1d01

Observation 23a56670-7aec-4c5f-8b6a-45c8f415ea98 · outbound

This paper cites Partially occluded object detection and count- ing.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Partially occluded object detection and count- ing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.976430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.943730Z digest=sha256:4bfb2df79519419d4f3780f23a4132f53c4abebc9c02104d7b7cac4f5a197823

Observation cb31961b-808d-48f3-a39f-a22f7d99613f · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:00.949346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:00.949346Z digest=sha256:93dd22a4a979ee14801c2f068c27f28cc7dfad79f88431552fcdbbc66cbb64d9

Observation 41fe2455-d86c-42db-9f8c-585c9f72f1ef · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:00.955941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:00.955941Z digest=sha256:7d6c2be1dcf78c41dd36539bd6032057bbe34915d136c5d5f6e47bb30fbc2bd1

Observation 471c5f52-e430-42b9-95f1-1c6d803dacf3 · outbound

This paper cites The coefficient of determination r-squared is more informa- tive than smape, mae, mape, mse and rmse in regression anal- ysis evaluation.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting The coefficient of determination r-squared is more informa- tive than smape, mae, mape, mse and rmse in regression anal- ysis evaluation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.961827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.962584Z digest=sha256:7eaddaaca66bfaed82bf125f368bbb0288227a511c70ae118b1df8931a53a5eb

Observation 45b0d63f-6371-42f7-b4b7-8ee6f5251cb1 · outbound

This paper cites How frequent are numbers? Language & Communication, 31(1):27–37, 2011.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting How frequent are numbers? Language & Communication, 31(1):27–37, 2011

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.947714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.968483Z digest=sha256:c6396993b10808e3e7167ccb8771ef375ed499762455e436280ab6a10d1fecaf

Observation aef7b08c-cec9-435c-bfe1-41ecec495e7d · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:00.974177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:00.974177Z digest=sha256:6957bc2e916c3574ca1275d4c01440b68887415959dfa539508af5bc18bb9483

Observation 7dfb5193-b9c1-4c1c-91dc-bbb667780387 · outbound

This paper cites Multi- branch segmentation-guided attention network for crowd counting.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Multi- branch segmentation-guided attention network for crowd counting

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.933775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.980304Z digest=sha256:82648acb2499f1e0d88cbb7e09bf8374da8e2c1974efb52f43870dc8e234d5b8

Observation 0d786f96-1eec-433d-8679-4ad315aa6664 · outbound

This paper cites A pragmatic view of accuracy measurement in forecasting.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting A pragmatic view of accuracy measurement in forecasting

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.919721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.986225Z digest=sha256:956686f0068729c6af3cf38e46c2ad6213bad640fa5a00660c347c769e5bc2df

Observation 6d78ec9b-9a67-4011-9db0-d02b7147ce83 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.905314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.992501Z digest=sha256:59b1f7b4c3cdd8b86eeaac2775ac8b21f94d0f5492ffd7a3513f4fe3e8a6be43

Observation 4f53422c-6e23-4feb-98e3-19eb44128216 · outbound

This paper cites Recurrent world models facilitate policy evolution.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Recurrent world models facilitate policy evolution

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.891371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:00.998361Z digest=sha256:ea46965bbde89373d686bacb16ca6bd93b9bb8cc341b9cd955687a4972aee61e

Observation 404abf88-149c-48f7-85e5-5c831ac726b0 · outbound

This paper cites Countnet3d: A 3d computer vision approach to infer counts of occluded objects.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Countnet3d: A 3d computer vision approach to infer counts of occluded objects

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.877312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.003771Z digest=sha256:56200e86f0cc9bcc6f09957d5f463550be46db31d3334de71465f667429a6379

Observation 1db48b16-d25b-470f-b93c-03bdebf3cd9f · outbound

This paper cites Organi- zation in vision : essays on gestalt perception.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Organi- zation in vision : essays on gestalt perception

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.863277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.009254Z digest=sha256:0a1dc29fa18f7e0a7ce49936a3f37094f83ee07d2859d753780a4ad3a7c6edb7

Observation 08df761c-23e5-40c9-8be6-a27e018fff66 · outbound

This paper cites Are Deep Learning Models Robust to Partial Object Occlusion in Visual Recognition Tasks?.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Are Deep Learning Models Robust to Partial Object Occlusion in Visual Recognition Tasks?

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:29:01.446550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.015132Z digest=sha256:50eee6ca8f48629f47d72d294ec987d080beb7caea9bbb61a6fa9ef63ceefa16

Observation 66b90470-0cf3-42cc-be60-74c5a908b42b · outbound

This paper cites an unresolved cited work.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:29:01.848574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.020974Z digest=sha256:d0d8d83aec97833f0e32b8df3a100cea577aa98c0cafcdc0d0952597c9c4be96

Observation eb936311-59f6-4a9a-a854-c8c57a70f60b · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.026211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.026211Z digest=sha256:ddc021dfa6bf89a4de9183880f03b37e748f43504ddb3f7371252e95841213c0

Observation a290d531-097d-44be-aa37-50b8ac296e44 · outbound

This paper cites Crowdclip: Unsupervised crowd counting via vision-language model.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Crowdclip: Unsupervised crowd counting via vision-language model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.834691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.032325Z digest=sha256:44a7dd7531a97c642e915e88fdf991646aa1736c7f926b532f5bf42a81ccc1ac

Observation db2caff6-265b-470e-aa48-598244cd92b8 · outbound

This paper cites Visual spatial reasoning.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Visual spatial reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.820486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.037944Z digest=sha256:f27bf4a2669e44309857180588d4fad27ea28d71e3a4d87ef7b673ffd3a2863a

Observation 86eb0e07-47ad-4b9a-96a8-143d74cf52a3 · outbound

This paper cites Optimum design of chamfer masks using symmetric mean absolute percentage error.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Optimum design of chamfer masks using symmetric mean absolute percentage error

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.806803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.043304Z digest=sha256:88b9ac879d24b9718d8f4ed280c9b8429c67043a0d3e4bfe46523feb4148e545

Observation 4a81c6e7-f02c-4416-95a5-9e8cd3ab01d3 · outbound

This paper cites Spartqa: A tex- tual question answering benchmark for spatial reasoning.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Spartqa: A tex- tual question answering benchmark for spatial reasoning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.791821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.048783Z digest=sha256:bae4e8e69193fee2cc72e4f23e11011e66e2b348abbc35273aa48fdef28b63a4

Observation bd50a9ad-b90c-4263-a8d1-5ab1322668a1 · outbound

This paper cites Neuronal representation of occluded objects in the human brain.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Neuronal representation of occluded objects in the human brain

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.776120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.054128Z digest=sha256:a6c0e11605eaeffc244231bd65bb79a69f8163b8d7b589939229b1867d2307a1

Observation 9a3ea8cd-264d-48b5-9cf6-a3178ca4eb34 · outbound

This paper cites Hello gpt-4o, 2024.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Hello gpt-4o, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.761366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.059669Z digest=sha256:4f70bfdd446fc78c3339783822f8c01301489985aa66dd4c5743904bd2468bf0

Observation 352c322f-0a46-4be9-8794-f126c3db1c54 · outbound

This paper cites Openvlm leaderboard.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Openvlm leaderboard

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.747360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.065089Z digest=sha256:b9ffb14d7ba7485e6fc7e975ccf684d059d25020b40dd74886d6a7d0bdf84138

Observation eeba0727-1b0d-48fd-9a85-334c0d069f5d · outbound

This paper cites Development of modal and amodal completion in infants.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Development of modal and amodal completion in infants

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.733165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.073007Z digest=sha256:109f99f0d653e889da4b3fafa59053b4748284ec6079650e29a52ae459125666

Observation 3c94c717-76e7-4d0d-a49c-d6ed29561c96 · outbound

This paper cites pix2gestalt: Amodal segmentation by synthesizing wholes.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting pix2gestalt: Amodal segmentation by synthesizing wholes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.718689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.078650Z digest=sha256:91ba54f01dc7fa5d5a975c7304b7ff90ea4823217b5660b38df0c00586cd9c24

Observation b60186fa-f33d-4f53-b6e1-4f7c59db9b4f · outbound

This paper cites Is Temperature the Creativity Parameter of Large Language Models?.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Is Temperature the Creativity Parameter of Large Language Models?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.083875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.083875Z digest=sha256:eb07e8eb80859394426e01be30e1160b3553b9f411ce5e1257db317b3cf944cc

Observation 69323239-fae5-4061-a84b-38469b2a8c6b · outbound

This paper cites Lvlm-count: Enhancing the count- ing ability of large vision-language models.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Lvlm-count: Enhancing the count- ing ability of large vision-language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.089146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.089146Z digest=sha256:474101423195d4245ded762d5cd99528a695b71d8bf9cc6ff73aa1e1f0e16e34

Observation d8e74d68-641c-4462-ab1a-a961132b9137 · outbound

This paper cites OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:29:01.281300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.094190Z digest=sha256:58fb17e5d770bec3e3cf619a8850b04c1a84ccc5129560625072783261e7c3c7

Observation 706b7659-1b93-4291-9592-09540f894774 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.099402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.099402Z digest=sha256:0b58c99745f7d1dddd811b4cd1010dc33d63ac331c4c1e9c945315ccba53cec1

Observation 0b803602-9b66-418e-835f-b9f4d8dd977a · outbound

This paper cites GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.104745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.104745Z digest=sha256:cee885a65387f897dc28feead66a41a54c5d88d85cf654f884c559e3ecbd12c8

Observation 11ee0401-5b64-4942-aa9b-45f79300784f · outbound

This paper cites Learning to count everything.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Learning to count everything

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.694020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.110028Z digest=sha256:708690c2f344399ecf1fb769fa7a826e07c2d9ba5f2a4420c131da6dd80fd3df

Observation f847fe79-1dae-44a9-a70a-e700f10963e9 · outbound

This paper cites Mask guided gated convolution for amodal content completion.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Mask guided gated convolution for amodal content completion

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.680057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.115351Z digest=sha256:5de1c17da434b1f354543545723a1da61a14be26464bb1bce2a8217af4ed5e7b

Observation 0ddd6478-86dd-4d7d-b813-1f81b48882df · outbound

This paper cites A cor- pus of natural language for visual reasoning.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting A cor- pus of natural language for visual reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.664181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.120543Z digest=sha256:c3fd7f582f9d86938edc44ef9fc0679cae558e62f53a19b7c240a4a5a9d24e57

Observation f27f5309-6a1e-4093-bda8-687e96f7a299 · outbound

This paper cites an unresolved cited work.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:29:01.648208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.125751Z digest=sha256:ba1a16d54e3d64574be018272a06ce43bf150ad87991fbb27dd014a0d0da06bb

Observation c5dee602-87b3-4022-ae44-acd4e8ab93e0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.130709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.130709Z digest=sha256:0e0c5c2fac6ade927718fb972e0d954bc395f1179e2f42c6e9b60b0502c7ed40

Observation 31b297b8-3407-4f5b-b568-0eb633295175 · outbound

This paper cites Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.135882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.135882Z digest=sha256:1e2c958438f7429af31fb984083b132c5e1b00173f252f1385abab33f5e6c90e

Observation 46df3c93-5ee8-42fb-bc57-102584265e0b · outbound

This paper cites Occlusion robust wheat ear counting algorithm based on deep learning.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Occlusion robust wheat ear counting algorithm based on deep learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.631770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.141328Z digest=sha256:72a7db9cf944575865daeec8cd2a77c203d2aad823c9f52e85a479c5a02874f8

Observation 31e8c6e0-60c6-436d-8431-49cc9738e2fe · outbound

This paper cites Dual- branch counting method for dense crowd based on self- attention mechanism.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Dual- branch counting method for dense crowd based on self- attention mechanism

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.614682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.146682Z digest=sha256:2a5ee5b91f1f885550f782cb60a7ce5501da677a7af683378c9feb4f62bcfddf

Observation 9cafec2d-51d5-4c78-aeb6-b85e7bd5c8ca · outbound

This paper cites Children’s understanding of counting.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Children’s understanding of counting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.599014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.151811Z digest=sha256:d87d5f0bfc063bd06f8d0e9eb797863c061eab081ef3d70c5fade430e2f41def

Observation 1480534f-aed0-489b-9db3-13a090094f9d · outbound

This paper cites Amodal com- pletion via progressive mixed context diffusion.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Amodal com- pletion via progressive mixed context diffusion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.583139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.159215Z digest=sha256:7a0e7f6c9bde2707bae282322fb74ddda4056dc24a764d3c9fbf74c01b1dab4e

Observation 20e52af8-35a5-4f22-ace5-0e13340980cb · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:29:01.164596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:29:01.164596Z digest=sha256:5ccc9d031174ccf89d104b1b4c61fdd33cea4a2566daf88c853566cf946ac1a2

Observation b5099834-2171-4fbd-a9c5-bc6bc5e53d90 · outbound

This paper cites Multi-branch progressive embedding network for crowd counting.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Multi-branch progressive embedding network for crowd counting

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.566952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.170049Z digest=sha256:6962bc3ae744e0d441e8dc282916dd523d0692b82b241ad60c4159f7b9c69f03

Observation 44685ef7-9712-47f9-8abc-972ae342f9f8 · outbound

This paper cites no”, the images were immediately discarded. If the model output was “yes.

CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting no”, the images were immediately discarded. If the model output was “yes

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:29:01.551241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:29:01.175213Z digest=sha256:846855aa4237ccea85f0f45623c3d2b8ab4e9657003ecd6e0d0baada499e4765

Pith citing papers

Observation 0f972e68-e574-42ee-a6ed-21d647e6ff52 · inbound

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models cites this paper.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:32.315784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:32.315784Z digest=sha256:56d3f930d3bc94c88a4cca9da1122ddee65fc3c36b88cc29b7251214bacf3ab9

Observation 89ed4db9-1ac1-465e-9632-203f668c3226 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.167072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.167072Z digest=sha256:5f8247c10427d0485ba2400244ad5057a4751f4eb6c1b0319bee1726863dd911

Observation b1d2c762-7cdd-4a35-a45b-60977dcf6494 · inbound

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates cites this paper.

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:07.618065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:07.618065Z digest=sha256:a1c77f8349b32479dcd2c10e098052fd33714f3053c0d2a9a26de4ea6f7c4af9

Observation 64a2ad40-79a8-4e5a-b180-c7fa679bd364 · inbound

Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications cites this paper.

Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T14:53:02.245247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:53:02.245247Z digest=sha256:4fad5e8622442edd77926dd15a21822401bfac4d55022b5f8661c80103322b48

Observation 818d1e8a-e5a3-4816-b936-dd8ead34c7a9 · inbound

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations cites this paper.

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T23:42:44.515158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:42:44.515158Z digest=sha256:385410eb2e9360d2f8d54da052dce09660abf2106e55ed8823bf75166c790233

Observation 20a09eda-3925-4e68-8f68-0ab48bdece20 · inbound

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving cites this paper.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:15:22.749650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T05:10:32.522453Z digest=sha256:ddbdd0d4a0709e3440c3718d0594c5beb14561f37fb3487362892cb01cece2e8

Observation e513a6c6-598c-42cc-946a-022c513cc61d · inbound

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving cites this paper.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.124516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:56bc6f7a62fad40ccfd9e87b939e1a4de20321cdcec557d66397e1734c6185fa