Pith. sign in

Paper Citation Record · LEDGER

Do large language vision models understand 3D shapes?

As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.10908.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10908 v5

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:32:59.803624Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3034be33-4458-4393-be03-3d806614462b · outbound

This paper cites Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6].

Do large language vision models understand 3D shapes? Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.672489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.617247Z digest=sha256:d6a10e070eccb4a4d4cde8e94eea6d180a0475fcecd46cb99b624ff65c250274

Observation 0a98dca0-2506-4489-b256-8723388d0a50 · outbound

This paper cites The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials.

Do large language vision models understand 3D shapes? The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.655761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.623477Z digest=sha256:30c63c6074b8081dc66b4ae8694f5e3ef937650db55fc572d34e1f7219f843bb

Observation dc14ceb1-d4df-4d9f-8fd7-075eb122aa50 · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.637457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.629201Z digest=sha256:f7ceb0f1a5caf138a7b17c533116a54f1425713c3dcfa5acff9c03148f394d56

Observation a36ba2db-b468-4525-b700-c15386f9024f · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.617632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.635538Z digest=sha256:6e879cd31038b3fff201acc6e7388b3941cb4af9285e85733745b0a17baee61d

Observation 501796c8-e85a-4d6b-9aea-6ed22d964a8b · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.598836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.640895Z digest=sha256:c00f6f42cbe8d31efb56bd71b215bd9f5227b847046208bcffe46a8b2d525a02

Observation e686b0bb-aee8-43e1-9337-462d38ce1b52 · outbound

This paper cites Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition.

Do large language vision models understand 3D shapes? Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.579705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.646636Z digest=sha256:c7695cf641a1a01b2a0c7c89e68b1d74fc9c3a917c8e886d32b42c21883deabb

Observation abc6e0d9-52d2-4f76-9914-571fe80d2593 · outbound

This paper cites Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter.

Do large language vision models understand 3D shapes? Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.562114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.653051Z digest=sha256:e48d7ac92e4f026608bca979ed91f22f80fe921f035906390969c5ccef3d551f

Observation 6368588a-5d7d-42a8-9a86-55ea9f8c417d · outbound

This paper cites 2) Has a different orientation compared to the object in panel A.

Do large language vision models understand 3D shapes? 2) Has a different orientation compared to the object in panel A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.544037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.658106Z digest=sha256:573d6cc4a7fda3165eb352ec790743a2ea28aa3edc4579cd7e1a3e8c2813fd53

Observation 680de0eb-c39b-4424-bf54-467d886c0059 · outbound

This paper cites Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1).

Do large language vision models understand 3D shapes? Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.525021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.663124Z digest=sha256:51f7a0ebe8ae8f9c22f288317d55d4ac8735908fb6e6a401c0569a20dcd66433

Observation 8a5c8630-6636-4aac-a126-99e5bc22f2eb · outbound

This paper cites These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12].

Do large language vision models understand 3D shapes? These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12]

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.503699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.668844Z digest=sha256:29a53c667b12e06ca198f79b969a7ed9c00bce9c46327b66b0ca48e65451c4a4

Observation 2e6bfac0-3978-4ddf-a952-d83db8758918 · outbound

This paper cites Code used to evaluate the models on the benchmark available at this URL.

Do large language vision models understand 3D shapes? Code used to evaluate the models on the benchmark available at this URL

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.477553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.674052Z digest=sha256:8d2a81bcd2662c14cb47ce721b99e9921f1b6f96d5eba659586b194ae4c742c5

Observation 7c8821e9-3f3f-4f04-8fd5-2cb6bb35516d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Do large language vision models understand 3D shapes? Flamingo: a visual language model for few-shot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.458914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.679122Z digest=sha256:90670fafb75660cb9653f31ba6a213dcc92ab4b7cfe31f1d6a08b4af3c5ef2d5

Observation 50ea74b2-1b65-4e03-8042-7766b9cbe2cd · outbound

This paper cites GPT-4 Technical Report.

Do large language vision models understand 3D shapes? GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.684975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.684975Z digest=sha256:26bf0467682a4246c6b4b7d49d84b527984066eb6b6b22707e1f0162913dd5e8

Observation a9899b30-3309-445f-8474-9cff7f7ab3f7 · outbound

This paper cites Real-world robot applications of foundation models: A review.

Do large language vision models understand 3D shapes? Real-world robot applications of foundation models: A review

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.441207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.692018Z digest=sha256:7602dbcca0bae995321a7ea7b29d5f2891619e9ff606c208bae85b9804bb9d69

Observation 0ab174d4-2dda-4e41-8d9f-03f3439000f4 · outbound

This paper cites Approaching human 3D shape perception with neurally mappable models.

Do large language vision models understand 3D shapes? Approaching human 3D shape perception with neurally mappable models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.700235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.700235Z digest=sha256:6d8d1cf4a5ad4face1bc97226c1fb80b0a5d86cb3f892a99724c22b93a49a2c3

Observation c95c0070-cf76-4ea2-a227-819545c5a26f · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models.

Do large language vision models understand 3D shapes? Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.421219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.706021Z digest=sha256:9e294b02e8b048505714e11cd1f61984c70e41b30c50a2100ab15bbc5bce621b

Observation b51dd177-3386-4663-a34b-04f75132c30f · outbound

This paper cites VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information.

Do large language vision models understand 3D shapes? VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.711784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.711784Z digest=sha256:1bb1134bf1217790d50b4c489c1357f1f545b09fde6be5d04e95234e75e6fe4b

Observation 19fba42e-ae53-433a-832d-7bb50df268e8 · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

Do large language vision models understand 3D shapes? Can We Talk Models Into Seeing the World Differently?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.717893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.717893Z digest=sha256:1cd8615c8cd946ac4868e8e062b9f0c41b1278d109767a39f70bc873b0f60a88

Observation 9b2d850c-06bd-4181-be6d-f1398fa68724 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions.

Do large language vision models understand 3D shapes? Exploring the frontier of vision-language models: A survey of current methodologies and future directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.723920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.723920Z digest=sha256:4e9e472aa7408135f5452e4bda8aa6950f12cb0cef7d93ee713db2218a93b4c7

Observation b98382ba-30bc-42b8-915d-d86fbd198ffc · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Do large language vision models understand 3D shapes? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.729389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.729389Z digest=sha256:927f23bf3149be6eed49b5a97d01149fd9368b6a5b9934ad50467a3c3f302cc4

Observation 89992980-f643-4bcd-9190-ff61916e22ca · outbound

This paper cites Can 3D Vision-Language Models Truly Understand Natural Language?.

Do large language vision models understand 3D shapes? Can 3D Vision-Language Models Truly Understand Natural Language?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.735368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.735368Z digest=sha256:5c7eec83af684b8e0253806685d1c112c5a0a76de3e55431e492ff2f2ebe1290

Observation 7767142a-5c19-4f6a-a0aa-7a0d04da4b3c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Do large language vision models understand 3D shapes? Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.741436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.741436Z digest=sha256:42cb87f180cd99e439ee8b95bb6a5bf208c5453d94a984ef0001cdc7c598b5a6

Observation 7ed68ce4-b4cd-4062-aa72-90accc550ae2 · outbound

This paper cites GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?.

Do large language vision models understand 3D shapes? GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.747061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.747061Z digest=sha256:efeebd3648c5d652090d80f2ec475d9a8196d47df7576bd7a3ad1854864b1bdb

Observation ec3a7477-e733-413f-8262-f7e521146d38 · outbound

This paper cites A general protocol to probe large vision models for 3d physical understanding.

Do large language vision models understand 3D shapes? A general protocol to probe large vision models for 3d physical understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.384059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.752641Z digest=sha256:5eab5c47b7c158ed39a43572b15bdaf9ccc58d8b652304e99b6d37edecc7fd83

Observation f96fe009-55fb-4756-a18d-1f54a1fb6376 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Do large language vision models understand 3D shapes? A Survey on Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.757593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.757593Z digest=sha256:28b546adadc9454e1e8a58b88cc64a6f9e9fee2530294fda7a36d6a976694cb1

Observation 4765c8e0-0de6-46d9-9039-00709859032f · outbound

This paper cites When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models.

Do large language vision models understand 3D shapes? When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.764351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.764351Z digest=sha256:84b6543d44a3e3a061a714c9121c3d8fb2a5cfb813d7b361a70838d663dd233f

Observation add2d72e-3b04-461f-bfcc-1d9e03910ff2 · outbound

This paper cites Language-Image Models with 3D Understanding.

Do large language vision models understand 3D shapes? Language-Image Models with 3D Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.771499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.771499Z digest=sha256:5201f97b29d78a0b7049f841b77cc0a230071aec2646438db2f603ad6d23bd9b

Observation 56fee059-30fe-4cdb-86e0-03dacf78109a · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

Do large language vision models understand 3D shapes? Shapellm: Universal 3d object understanding for embodied interaction

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.362487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.778276Z digest=sha256:aa1c26b8d6827f034655701588ecb8697513bba655286755b9536a9d5b1f89f4

Observation e5721b71-513f-4a10-bdab-8aa799b477dc · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Do large language vision models understand 3D shapes? Objaverse: A universe of annotated 3d objects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.345067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.786119Z digest=sha256:fc9ec12b7139dd5dae3d14bca171d21a603f4a877c64f90c5faa3c05c6d3c5e0

Observation 2af7571c-6ed6-4cee-82fc-5a8b44b5bc94 · outbound

This paper cites Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods.

Do large language vision models understand 3D shapes? Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.792476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.792476Z digest=sha256:609ebfcaf27006520e066c402ba55320861d91ad94c348d57fe568c074446e9f

Observation bf7155e3-a1fc-48f5-ab46-ed90f98c84ec · outbound

This paper cites Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation.

Do large language vision models understand 3D shapes? Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.325661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.798061Z digest=sha256:0da350684a8561fc372d13a015bb02cb34aae2109c6267bfd3c899e3ac54ebce

Observation cc5021e5-b612-4a76-af9f-6b71d0dfdd6a · outbound

This paper cites Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain.

Do large language vision models understand 3D shapes? Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.307386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.803624Z digest=sha256:f33cf7f7a3ef073a8ab2d5b6f9420f9b87584473c8fd22a731f8103ee17b256b

Pith citing papers

No inbound Pith citation observations are available.