Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:29:01.175213Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2504.15485.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:29:01.175213Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:32.315784Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T16:44:56.123044Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9f31bd67-203f-4f7d-88e3-cbd31f32813e · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Llama 3.1 model card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6499a395-ef09-4816-90d8-3637fc2a5575 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c776e55d-2837-410e-bbb0-7cadd595de7e · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting CountGD: Multi-Modal Open-World Counting
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc7c5e2b-8c77-4c54-a770-dda9cf0f1db8 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Vqa: Visual question answering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6846673f-be1f-4b85-9144-d502784f0bc6 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Image amodal completion: A survey
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d14a7d27-4ea5-46cd-a8ec-0c5c370ef6b5 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Open-World Amodal Appearance Completion
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c968a922-a7d0-43da-8583-453e9b970118 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Text to image model arena, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 23a56670-7aec-4c5f-8b6a-45c8f415ea98 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Partially occluded object detection and count- ing
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cb31961b-808d-48f3-a39f-a22f7d99613f · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41fe2455-d86c-42db-9f8c-585c9f72f1ef · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 471c5f52-e430-42b9-95f1-1c6d803dacf3 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting The coefficient of determination r-squared is more informa- tive than smape, mae, mape, mse and rmse in regression anal- ysis evaluation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45b0d63f-6371-42f7-b4b7-8ee6f5251cb1 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting How frequent are numbers? Language & Communication, 31(1):27–37, 2011
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aef7b08c-cec9-435c-bfe1-41ecec495e7d · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfb5193-b9c1-4c1c-91dc-bbb667780387 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Multi- branch segmentation-guided attention network for crowd counting
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d786f96-1eec-433d-8679-4ad315aa6664 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting A pragmatic view of accuracy measurement in forecasting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6d78ec9b-9a67-4011-9db0-d02b7147ce83 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f53422c-6e23-4feb-98e3-19eb44128216 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Recurrent world models facilitate policy evolution
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 404abf88-149c-48f7-85e5-5c831ac726b0 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Countnet3d: A 3d computer vision approach to infer counts of occluded objects
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1db48b16-d25b-470f-b93c-03bdebf3cd9f · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Organi- zation in vision : essays on gestalt perception
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 08df761c-23e5-40c9-8be6-a27e018fff66 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Are Deep Learning Models Robust to Partial Object Occlusion in Visual Recognition Tasks?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66b90470-0cf3-42cc-be60-74c5a908b42b · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eb936311-59f6-4a9a-a854-c8c57a70f60b · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a290d531-097d-44be-aa37-50b8ac296e44 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Crowdclip: Unsupervised crowd counting via vision-language model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation db2caff6-265b-470e-aa48-598244cd92b8 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Visual spatial reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86eb0e07-47ad-4b9a-96a8-143d74cf52a3 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Optimum design of chamfer masks using symmetric mean absolute percentage error
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a81c6e7-f02c-4416-95a5-9e8cd3ab01d3 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Spartqa: A tex- tual question answering benchmark for spatial reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bd50a9ad-b90c-4263-a8d1-5ab1322668a1 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Neuronal representation of occluded objects in the human brain
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9a3ea8cd-264d-48b5-9cf6-a3178ca4eb34 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Hello gpt-4o, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 352c322f-0a46-4be9-8794-f126c3db1c54 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Openvlm leaderboard
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eeba0727-1b0d-48fd-9a85-334c0d069f5d · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Development of modal and amodal completion in infants
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c94c717-76e7-4d0d-a49c-d6ed29561c96 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting pix2gestalt: Amodal segmentation by synthesizing wholes
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b60186fa-f33d-4f53-b6e1-4f7c59db9b4f · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Is Temperature the Creativity Parameter of Large Language Models?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69323239-fae5-4061-a84b-38469b2a8c6b · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Lvlm-count: Enhancing the count- ing ability of large vision-language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e74d68-641c-4462-ab1a-a961132b9137 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 706b7659-1b93-4291-9592-09540f894774 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Learning transferable visual models from natural language supervi- sion
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b803602-9b66-418e-835f-b9f4d8dd977a · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ee0401-5b64-4942-aa9b-45f79300784f · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Learning to count everything
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f847fe79-1dae-44a9-a70a-e700f10963e9 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Mask guided gated convolution for amodal content completion
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0ddd6478-86dd-4d7d-b813-1f81b48882df · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting A cor- pus of natural language for visual reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f27f5309-6a1e-4093-bda8-687e96f7a299 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c5dee602-87b3-4022-ae44-acd4e8ab93e0 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b297b8-3407-4f5b-b568-0eb633295175 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46df3c93-5ee8-42fb-bc57-102584265e0b · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Occlusion robust wheat ear counting algorithm based on deep learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31e8c6e0-60c6-436d-8431-49cc9738e2fe · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Dual- branch counting method for dense crowd based on self- attention mechanism
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9cafec2d-51d5-4c78-aeb6-b85e7bd5c8ca · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Children’s understanding of counting
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1480534f-aed0-489b-9db3-13a090094f9d · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Amodal com- pletion via progressive mixed context diffusion
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 20e52af8-35a5-4f22-ace5-0e13340980cb · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5099834-2171-4fbd-a9c5-bc6bc5e53d90 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting Multi-branch progressive embedding network for crowd counting
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 44685ef7-9712-47f9-8abc-972ae342f9f8 · outbound
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting no”, the images were immediately discarded. If the model output was “yes
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f972e68-e574-42ee-a6ed-21d647e6ff52 · inbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ed4db9-1ac1-465e-9632-203f668c3226 · inbound
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d2c762-7cdd-4a35-a45b-60977dcf6494 · inbound
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a2ad40-79a8-4e5a-b180-c7fa679bd364 · inbound
Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 818d1e8a-e5a3-4816-b936-dd8ead34c7a9 · inbound
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a09eda-3925-4e68-8f68-0ab48bdece20 · inbound
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e513a6c6-598c-42cc-946a-022c513cc61d · inbound
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.