Pith. sign in

Paper Citation Record · LEDGER

Do large language vision models understand 3D shapes?

As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.10908.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10908 v5

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:32:59.803624Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3034be33-4458-4393-be03-3d806614462b · outbound

This paper cites Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6].

Do large language vision models understand 3D shapes? Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.672489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.617247Z digest=sha256:d8fd27cb9a4c942a53c01769638dc44d0d1f9ae2375f4cff2da3f758c5f226d8

Observation 0a98dca0-2506-4489-b256-8723388d0a50 · outbound

This paper cites The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials.

Do large language vision models understand 3D shapes? The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.655761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.623477Z digest=sha256:c7c6acae253391344970ece723469dfbcb3767793ef02057f206a6d5260379f4

Observation dc14ceb1-d4df-4d9f-8fd7-075eb122aa50 · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.637457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.629201Z digest=sha256:47228d3bf62137d2eefc23369c88f836d2efec1fa640567e1e5fe7b56b062d12

Observation a36ba2db-b468-4525-b700-c15386f9024f · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.617632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.635538Z digest=sha256:2923eecab15f1012629dcb9d8927786a1c5d158722dc19ac9bf5f814d45f9670

Observation 501796c8-e85a-4d6b-9aea-6ed22d964a8b · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.598836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.640895Z digest=sha256:6a3dd1eac7fa114290b0dac84addc73e0df29345b5ece6f1503417cb52325e96

Observation e686b0bb-aee8-43e1-9337-462d38ce1b52 · outbound

This paper cites Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition.

Do large language vision models understand 3D shapes? Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.579705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.646636Z digest=sha256:d9b67a8b39a6ff9e3470c3522bacf8c7c2e962570c7e4745b8ad0961d33f09a0

Observation abc6e0d9-52d2-4f76-9914-571fe80d2593 · outbound

This paper cites Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter.

Do large language vision models understand 3D shapes? Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.562114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.653051Z digest=sha256:6a446be5fa51aa7024fee4838dff0821a626240f0c0107e849f0650abb6d690f

Observation 6368588a-5d7d-42a8-9a86-55ea9f8c417d · outbound

This paper cites 2) Has a different orientation compared to the object in panel A.

Do large language vision models understand 3D shapes? 2) Has a different orientation compared to the object in panel A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.544037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.658106Z digest=sha256:f70dd304f912854e945ce563abc81e3b00410d392091539ed878b52cd1ae574f

Observation 680de0eb-c39b-4424-bf54-467d886c0059 · outbound

This paper cites Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1).

Do large language vision models understand 3D shapes? Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.525021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.663124Z digest=sha256:050866eb2926fa219fbd9bfdbee3bec09a6b9ac7f032fda16928ae7db71f8609

Observation 8a5c8630-6636-4aac-a126-99e5bc22f2eb · outbound

This paper cites These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12].

Do large language vision models understand 3D shapes? These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12]

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.503699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.668844Z digest=sha256:737a2b60d43ed937895cf75fa229a569d8de89abc4ac0563617fad52cc3f87f4

Observation 2e6bfac0-3978-4ddf-a952-d83db8758918 · outbound

This paper cites Code used to evaluate the models on the benchmark available at this URL.

Do large language vision models understand 3D shapes? Code used to evaluate the models on the benchmark available at this URL

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.477553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.674052Z digest=sha256:d35e8d03f82fd89671b23a1fe59caa281c1674ed8acab734f4d46310ac21513e

Observation 7c8821e9-3f3f-4f04-8fd5-2cb6bb35516d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Do large language vision models understand 3D shapes? Flamingo: a visual language model for few-shot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.458914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.679122Z digest=sha256:38864baa529211cd94764c6ddd527b37e7b3644800baaf3467bd5fb053c05dfc

Observation 50ea74b2-1b65-4e03-8042-7766b9cbe2cd · outbound

This paper cites GPT-4 Technical Report.

Do large language vision models understand 3D shapes? GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.684975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.684975Z digest=sha256:458099fc57bd35954546827106dc6ec2abf9c485a991726c8cbfb698deabec5f

Observation a9899b30-3309-445f-8474-9cff7f7ab3f7 · outbound

This paper cites Real-world robot applications of foundation models: A review.

Do large language vision models understand 3D shapes? Real-world robot applications of foundation models: A review

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.441207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.692018Z digest=sha256:cbceba071928d7c988a560232ea4b1e3a52d3366d1cbd8475030eec544243074

Observation 0ab174d4-2dda-4e41-8d9f-03f3439000f4 · outbound

This paper cites Approaching human 3D shape perception with neurally mappable models.

Do large language vision models understand 3D shapes? Approaching human 3D shape perception with neurally mappable models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.700235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.700235Z digest=sha256:cfd7d334b33c6a50f39030c69136ad85e9eeb3719866e44b0c9c0b86dbd53c11

Observation c95c0070-cf76-4ea2-a227-819545c5a26f · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models.

Do large language vision models understand 3D shapes? Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.421219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.706021Z digest=sha256:168c257231364b5a78ccabcd32e038940616416fbe4d5fd4697c826e1bb4899e

Observation b51dd177-3386-4663-a34b-04f75132c30f · outbound

This paper cites VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information.

Do large language vision models understand 3D shapes? VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.711784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.711784Z digest=sha256:145596bd6769601d51da24f7d375050c7ca8bf3171c2a452066db5598a927261

Observation 19fba42e-ae53-433a-832d-7bb50df268e8 · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

Do large language vision models understand 3D shapes? Can We Talk Models Into Seeing the World Differently?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.717893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.717893Z digest=sha256:157794583333830ad15def6cd1405d62afb7f7b28b5d77b9016ac1ee29991da7

Observation 9b2d850c-06bd-4181-be6d-f1398fa68724 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions.

Do large language vision models understand 3D shapes? Exploring the frontier of vision-language models: A survey of current methodologies and future directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.723920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.723920Z digest=sha256:79282e091ba57c7d3427c99e8c5f1a0f32fd35ef2902dcaa7bfd3c20a91ab21b

Observation b98382ba-30bc-42b8-915d-d86fbd198ffc · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Do large language vision models understand 3D shapes? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.729389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.729389Z digest=sha256:fd413e9003d77cf407589b31d3e9affafe206b676fff98bd21e57396bf3e0611

Observation 89992980-f643-4bcd-9190-ff61916e22ca · outbound

This paper cites Can 3D Vision-Language Models Truly Understand Natural Language?.

Do large language vision models understand 3D shapes? Can 3D Vision-Language Models Truly Understand Natural Language?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.735368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.735368Z digest=sha256:7923d31075bb5e346aaeedf76a7ce60397e7533bcb7228fa26e00bb8b4ff3612

Observation 7767142a-5c19-4f6a-a0aa-7a0d04da4b3c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Do large language vision models understand 3D shapes? Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.741436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.741436Z digest=sha256:1103092a4c291ae3225e237313af0d380a8a22c075d36c3844508885b1fbad19

Observation 7ed68ce4-b4cd-4062-aa72-90accc550ae2 · outbound

This paper cites GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?.

Do large language vision models understand 3D shapes? GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.747061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.747061Z digest=sha256:b81a5bf021d6132b5b2882b5cd9305da79e5101b62c1bffa20760b03dab528fd

Observation ec3a7477-e733-413f-8262-f7e521146d38 · outbound

This paper cites A general protocol to probe large vision models for 3d physical understanding.

Do large language vision models understand 3D shapes? A general protocol to probe large vision models for 3d physical understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.384059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.752641Z digest=sha256:b0cc527bd40e1e8e18cab3f559e513474a56b2dafd7c1c7af5903da0d55289be

Observation f96fe009-55fb-4756-a18d-1f54a1fb6376 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Do large language vision models understand 3D shapes? A Survey on Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.757593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.757593Z digest=sha256:d6272cdab616afb8e10160f22aaddcd52ae58ab96ed395d50c0fbca409ebc6ae

Observation 4765c8e0-0de6-46d9-9039-00709859032f · outbound

This paper cites When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models.

Do large language vision models understand 3D shapes? When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.764351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.764351Z digest=sha256:7a80c2464216c7abcc4fa6111bfdbaf8ec4d42941f6993b4cbf0a23ac7435fa0

Observation add2d72e-3b04-461f-bfcc-1d9e03910ff2 · outbound

This paper cites Language-Image Models with 3D Understanding.

Do large language vision models understand 3D shapes? Language-Image Models with 3D Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.771499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.771499Z digest=sha256:62d23632bd75d950aa7759ce4627fd27374c8b580ab63821be1b0de574e06359

Observation 56fee059-30fe-4cdb-86e0-03dacf78109a · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

Do large language vision models understand 3D shapes? Shapellm: Universal 3d object understanding for embodied interaction

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.362487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.778276Z digest=sha256:323ba472439de2d3e80aa638f8e2098bb8adb660048a482c42ba1d331e23e968

Observation e5721b71-513f-4a10-bdab-8aa799b477dc · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Do large language vision models understand 3D shapes? Objaverse: A universe of annotated 3d objects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.345067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.786119Z digest=sha256:e8bf70c36e396bb76da9b9bfb3ec78fda05be1d9a54172ab97bf2102cc28374a

Observation 2af7571c-6ed6-4cee-82fc-5a8b44b5bc94 · outbound

This paper cites Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods.

Do large language vision models understand 3D shapes? Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.792476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.792476Z digest=sha256:8b19e3f1b064ffa1070a7da857f69d7ded732a26faec1ac7888ea8a2741e4ec2

Observation bf7155e3-a1fc-48f5-ab46-ed90f98c84ec · outbound

This paper cites Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation.

Do large language vision models understand 3D shapes? Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.325661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.798061Z digest=sha256:5c8f2b6845b1659ab9bafab8e3b5550b208ac9446b0cecb9c9bb0f117c4a4001

Observation cc5021e5-b612-4a76-af9f-6b71d0dfdd6a · outbound

This paper cites Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain.

Do large language vision models understand 3D shapes? Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.307386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:32:59.803624Z digest=sha256:a24fe100f43e5be2297068c388988b692c15703fb6bc9c7fa011bd7d521ab44a

Pith citing papers

No inbound Pith citation observations are available.