Pith. sign in

Paper Citation Record · LEDGER

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

As of 14 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 4 inbound Pith citation observations for arXiv:2506.23329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23329 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:19.061479Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.147646Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:46:19.349816Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved48
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70399b4d-b6e2-4b73-94bf-bad20488e404 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-OneVision: Easy Visual Task Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.055004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.055004Z digest=sha256:bb7cb56b103f7d4d3984af50adb1ddb28dc9cbf467fe6071c5431f282ffa9551

Observation e399d4ea-d573-4ec4-8742-5321185e750e · outbound

This paper cites GPT-4o System Card.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GPT-4o System Card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.097608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.097608Z digest=sha256:867204ec7bc30656838793293090c619b44ff9c6886732fdc9d013ad1fcd6480

Observation c1bf531f-f280-4fa3-83ef-1b4ba3d07b95 · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Claude 3 Model Family: Opus, Sonnet, Haiku

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.169741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.169741Z digest=sha256:0763c2002610dea7117559e4a1304ae7eccd938c57e04ef62a85c961d982fb8e

Observation c5f399e7-11ab-498b-a352-6b5424b05bb8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.224504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.224504Z digest=sha256:f40097a1673014eefb1df969d53846e9d6a532f46af94a41eaf4eb92bf1c4a6b

Observation b4095e18-2038-4477-a16d-c90e47d0c61d · outbound

This paper cites Qwen2.5-VL Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.285409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.285409Z digest=sha256:f6c554c011b3d0b1f0479eef42587403aae347c5791fd7ccbc94553c948c4dd2

Observation d393dec2-7886-404f-bd47-c76fdd7b1f14 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.354626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.354626Z digest=sha256:bd4df7e1c9389e602addc5264a36465328c584cec144dbc68e17cad8968c572f

Observation 623e3470-0483-491f-8878-f676cb8fedf9 · outbound

This paper cites Perceptions as hypotheses.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Perceptions as hypotheses

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.396632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.396632Z digest=sha256:e075dc88e5deccebf65f0abb56a1704130822bef8530de308ba0d7d394613f3b

Observation b67c78c0-8f4d-430b-9a8a-c6a05d6a7535 · outbound

This paper cites Yuille and Daniel Kersten.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Yuille and Daniel Kersten

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.441337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.441337Z digest=sha256:ee00265dcc7d1333b6b7f618992d9010e222f1fe5efd3e095ffb1a1c3f9680f6

Observation 4ba89f9d-b7fd-48cc-94e5-6035ea2f9f20 · outbound

This paper cites Efficient and robust analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Efficient and robust analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.504617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.504617Z digest=sha256:f2299fbca82f7c5dd8487dd8cf46cf5d136a35f868b532a7b45f9e86e0b27412

Observation 04895bc9-16e4-42ab-b6cb-fc311ebde2ea · outbound

This paper cites Bever and David Poeppel.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Bever and David Poeppel

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.569097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.569097Z digest=sha256:a2e8ba141ec0b956182c1a43b0fbafcae4b94e21913e14197ed1f3483114dc69

Observation 2bd2439d-548f-4935-b8eb-ebc1323b56a8 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Toolformer: Language models can teach themselves to use tools

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.626028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.626028Z digest=sha256:0590d0e0995b68b8601e9ed097bc64723b4bf9acb3ab8a727df0d6cd3779f8f0

Observation 41d94738-46fb-42ca-aa07-daad572b27a0 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Visual programming: Compositional visual reasoning without training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.687738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.687738Z digest=sha256:8c7f0ecde547bc582fea3a5ea2217df96f5b0343a7d3c9c61644dd8bc7679c1e

Observation 2371c8ed-4db0-4785-b740-808dc0eddae7 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Vipergpt: Visual inference via python execution for reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.750586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.750586Z digest=sha256:2a9a26671c8c11a221ffb95d6e6fa5ad2b0f3a7918928fabd22da3afd9bbe8a1

Observation 759e214b-494b-47d0-a7ff-29d41b221fb4 · outbound

This paper cites Gorilla: Large language model connected with massive apis.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Gorilla: Large language model connected with massive apis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.564130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:11.803713Z digest=sha256:0096efb4e82930c783b98f4b0b6728a389775ba62d1135b3c7cbc866c4448eb7

Observation 58910aac-d9e8-4948-9f22-a5cf9a3c51dd · outbound

This paper cites Blender - a 3D modelling and rendering package, 2016.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Blender - a 3D modelling and rendering package, 2016

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.281964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:11.874949Z digest=sha256:6c828fdea1aeeb85b5becfa038f6bd84b52e921f052480d16680365f760b9664

Observation 4f94bb30-7534-47ce-b4a6-8f2c24fe64e2 · outbound

This paper cites https://github.com/modelcontextprotocol, 2024.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering https://github.com/modelcontextprotocol, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.013539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:11.913246Z digest=sha256:ca05b4fdcbb7a19b9770fcf6e264a0f8930ebb4e5a8c94e442094a1a60fa2989

Observation ecc096df-bc12-4ee8-b717-625b2d4f4528 · outbound

This paper cites Blendermcp.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Blendermcp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:26.613500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:11.977613Z digest=sha256:58b368c3202e43c32e753fe2e98ed46f6eb0daf46d0281fd52628d37e73f2048

Observation c343af7a-bbde-476a-9168-91e73e1b1192 · outbound

This paper cites Embodied question answering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Embodied question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:26.204239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.022798Z digest=sha256:5971ba6ff889f2de7873123553b1872e79c4e28c1681fe912dfb4f2c7182bd68

Observation 24c50811-3954-4c4c-8c9a-3443545a9ca4 · outbound

This paper cites Multi-target embodied question answering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Multi-target embodied question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.083938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.083938Z digest=sha256:74217023a4e748b09d46ca6eeb2e288c6814f3fd4a66ba792aaed410dc644b29

Observation 5b4e46d7-ef6f-4292-8bc7-a0adc8407f6b · outbound

This paper cites 3d concept learning and reasoning from multi-view images.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d concept learning and reasoning from multi-view images

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.777892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.138262Z digest=sha256:90c497ad16e51ee31cc6ded651f4d77429c51ff19ec7ea21fbc6e8e52df7262a

Observation 9a512965-3142-40a5-a3af-38bb88cef8f6 · outbound

This paper cites One step at a time: Long-horizon vision-and-language navigation with milestones.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering One step at a time: Long-horizon vision-and-language navigation with milestones

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.671215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.194541Z digest=sha256:0ec1413c36bedda78ecdbac841b9f6f10d5e28b4604d291a4a1ee9eb54d990ae

Observation c6d80a6f-3acd-410b-b3bb-eb48b0cd720d · outbound

This paper cites Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.241831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.241831Z digest=sha256:1952aa2e1478f153ee1bda8d71ae7877c61aad5a1355620f33e9f47db9986252

Observation de9ccc8a-466a-4a37-afa9-fe5ffb61f905 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.279337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.279337Z digest=sha256:33af8e5485ec8b68891df7e05496701a1069dad296a3ff78cb082c035f5e2a19

Observation e9a6c378-4606-4592-94b9-abeafdd0c1ae · outbound

This paper cites Barrow and Jay M.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Barrow and Jay M

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.549695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.293443Z digest=sha256:ba3ae96651212cff6031df3fe27ccf3118700341d53f34f72e9854c0293bc980

Observation 05dad22f-a4c9-4d5e-b07d-273a93a6fc3a · outbound

This paper cites Soft rasterizer: A differentiable renderer for image-based 3d reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Soft rasterizer: A differentiable renderer for image-based 3d reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.418707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.327295Z digest=sha256:5e96d3ddfaac78e56b942812d4da2cbcfd0bb7baea794e0ebab85d55e5920594

Observation 47e3950a-d8e2-4c22-ac40-8243dea9bbdb · outbound

This paper cites Carr, Jonathan Ragan-Kelley, and Frédo Durand.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Carr, Jonathan Ragan-Kelley, and Frédo Durand

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.281806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.364961Z digest=sha256:640b2fb711c9764d4b006cf904cc420a73c6a34bf9be0328596a2185948dc86a

Observation 03cb4dd2-5833-44e6-a443-1ded4a1cbcc2 · outbound

This paper cites Differentiable vector graphics rasterization for editing and learning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Differentiable vector graphics rasterization for editing and learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.156345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.397916Z digest=sha256:ac12bac05e009c841d5e59047f6375f0375892fe34d05486766bbab7f3cbfdee

Observation fdf89410-523c-4f03-82e4-6e26fc4b84b2 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.044121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.474675Z digest=sha256:f92cf5e703b74a4303c7cd0e6fec098c8ee47ecf326a1574c1dd3789a3f49b73

Observation 185ef9c7-49ac-46b2-984d-870a0eff56e2 · outbound

This paper cites Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.925554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.567066Z digest=sha256:c25dc81f3ea06c2f36a833ec11547599b5d5d4d6fcd47ab6bdbfcf162671ffff

Observation 170cb4bb-be8d-4702-b529-96c80dc09769 · outbound

This paper cites V olume rendering of neural implicit surfaces.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering V olume rendering of neural implicit surfaces

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.826114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.634446Z digest=sha256:997623dad07981b6a8d35e4010bcf9eb731d3ceb89993665cae145152779348f

Observation fe13a610-3f42-47f0-be62-c42423e06960 · outbound

This paper cites Synsin: End-to-end view synthesis from a single image.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Synsin: End-to-end view synthesis from a single image

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.683222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.704011Z digest=sha256:ed9977c01681887602fc8aa4cf29351ff00580a4d26e093636bcb1662a05985e

Observation a40ed8df-92b3-4d8e-9c4f-eb74c9b4898b · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d gaussian splatting for real-time radiance field rendering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.597407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.762820Z digest=sha256:43617cabff55efdcf417014fdb9c1dd07bb35464af4424159994ead99ffba285

Observation 120bdca9-0787-411c-af13-41bf6860ac8f · outbound

This paper cites Path-space differentiable rendering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Path-space differentiable rendering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.822049Z digest=sha256:34c6089da15d37f09390d6f975111679668bafd6624ab82feb81144d7d40c02d

Observation 4a82b5d2-f0fa-4dcd-90ad-0e7dafa35cb0 · outbound

This paper cites Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning, 2025.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.322002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.878226Z digest=sha256:d9e8d9ca07af1883861abf6134754a703e02185f74df05f7d30dc6cd19529f7f

Observation baa76daf-d0cc-4046-867a-c8c4e18cb39f · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.924934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.924934Z digest=sha256:b7d5015b8c1a48794c8434f42ee8c380417067b9018ee9b97d68ce4bc34a32af

Observation af2108fa-3480-4223-889d-1ddbf54662d7 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scanqa: 3d question answering for spatial scene understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.123602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:12.978758Z digest=sha256:e417e79e500745816c678ef3855e99bce05f20d84a964ef1a5de5535f2e2c473

Observation 0fdde439-780f-4367-abce-abcb0b2ba4ac · outbound

This paper cites 3d-vista: Pre- trained transformer for 3d vision and text alignment.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d-vista: Pre- trained transformer for 3d vision and text alignment

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:23.895150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.066768Z digest=sha256:bd325982bfde544bb659acaa8f8fd7200ce59f0aec8fc1d13e9717e2aff2e70e

Observation 804d0774-75fd-4430-b308-0a344c7c0b98 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.125119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.125119Z digest=sha256:4045e62c000355642a29cc420afb941f1eb8900d2aa3456c2000774a881a556c

Observation 069e343b-baa2-47c1-ab21-fa116d0be223 · outbound

This paper cites an unresolved cited work.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:23.716990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.176031Z digest=sha256:a9d9b0fb5a19c1a0c67abf8740c4c657aff1db556aeb558a8ca48841d99f6bdd

Observation 8a82e53a-4cde-4c52-8b39-2781b3ce7187 · outbound

This paper cites an unresolved cited work.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:23.514382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.257778Z digest=sha256:417fefa31db8c498790772c566ff652df783a03ed9c64a569a7311fff6a6e9d6

Observation 91a71c99-fc19-4bfb-9743-85f859313aca · outbound

This paper cites Scalable 3d captioning with pretrained models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scalable 3d captioning with pretrained models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.307929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.307929Z digest=sha256:2b611ee3a60802618421c516d154b8cc89dd033e16a91fd30977f74729783384

Observation f8e232af-a3db-4d52-a83a-5d822a3dd8a3 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Openeqa: Embodied question answering in the era of foundation models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:23.340070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.367178Z digest=sha256:3427f009af2d79d0547c67d127f7620c5b211a3967e17ae73cd28dfb304e4ce3

Observation 5364b1bc-adb6-4a12-a480-5b8ade7cb1cb · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:23.158643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.427092Z digest=sha256:604e970bfa7b8f472a02179275baa61c6adf928996c7f9403b844226213cac56

Observation 79c150fb-dcb2-4dda-9b34-946805354b5f · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.964985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.473093Z digest=sha256:e782d59662658a04df6b0aea46a268d6826a8cf5b715bfb62c0670595f4aa188

Observation 2243b079-f024-4250-87c2-d590e88832e9 · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.738013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.520082Z digest=sha256:f96efb94813fc375f7e212c0cc73256b6f62dbc3ac9824a81d3802ec8b61c2cc

Observation 13c423ae-e694-4b4d-9483-84ce12cea5bc · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering ConceptFusion: Open-set Multimodal 3D Mapping

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.629446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.629446Z digest=sha256:ddf78b4633d211737fd86a636230d7d947984c821f3bc72b8a29a0545c226d23

Observation 0854f686-091f-4bf1-9322-c38faaea3951 · outbound

This paper cites Context-aware entity grounding with open-vocabulary 3d scene graphs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Context-aware entity grounding with open-vocabulary 3d scene graphs

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.553902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.689722Z digest=sha256:33d8e7b8ab20eabb32287f73a328514362e7f6bd23364f6b0f4f576da1bef21f

Observation 09505462-0c34-481b-9759-994420a4609d · outbound

This paper cites Tenenbaum, Antonio Torralba, Florian Shkurti, and Liam Paull.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Tenenbaum, Antonio Torralba, Florian Shkurti, and Liam Paull

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.310278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:13.772311Z digest=sha256:ed9ec9e208e986db00cc54bf324ac1776234a7a599fdd1e7c31464c297893013

Observation 5d0d8d06-ed63-48ce-ab2b-6ad2bebe8242 · outbound

This paper cites 3D-LLM: Injecting the 3D World into Large Language Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3D-LLM: Injecting the 3D World into Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.851642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.851642Z digest=sha256:632763ded978d856ea40493cd4341ff29f3f547734df78f89b4d41c15c64fe44

Observation af0980fc-8aae-4b20-8ca6-af1ebf75083a · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.961478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.961478Z digest=sha256:0865d064958de1fd929528532a9737c5267beedbcf6b670eb0a4cedbc55af84c

Observation 477b6121-dda0-4230-a3af-7ebb78398485 · outbound

This paper cites 3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:14.100701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:14.100701Z digest=sha256:3b3f316db9611fb948db9f47d4e3b2259d48aa9d520f1d2c6b9f814bbca7ffff

Observation aa631a06-2c28-496b-934d-0e48b671585c · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Openeqa: Embodied question answering in the era of foundation models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:14.285145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:14.285145Z digest=sha256:4cedf81a981db6410d50daf54ff64c021e2885102cb38064b02a2102f7925067

Observation bf2e2b69-3430-422d-9178-84203d7d112f · outbound

This paper cites Kulkarni, Pushmeet Kohli, Joshua B.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Kulkarni, Pushmeet Kohli, Joshua B

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.056195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:15.001949Z digest=sha256:678bc8cf5d13b473e9e2f8fd2f4b62a218b6eead881fc2cc145c91344e80048d

Observation 92033c46-13f7-4403-9222-3acc43d758a9 · outbound

This paper cites Learning to infer graphics programs from hand-drawn images.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning to infer graphics programs from hand-drawn images

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.901078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:15.895509Z digest=sha256:aa8e33014a99b9099f6b2ac816aca7a52a4da19186c31d51ed41fb0f275490d6

Observation ab25332c-a816-4f5b-8b68-f98bd8ce40d5 · outbound

This paper cites Learning to infer and execute 3d shape programs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning to infer and execute 3d shape programs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.706626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:16.005434Z digest=sha256:832ac0ce17cc043141c18f4ad4c56c4daf1395f21958adbf3d28c28c3ab9d7e8

Observation 370dc2bc-3632-4845-839c-20e7bf03bfeb · outbound

This paper cites Shapeassembly: Learning to generate programs for 3d shape structure synthesis.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Shapeassembly: Learning to generate programs for 3d shape structure synthesis

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.534646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:16.170005Z digest=sha256:3da2604bf629ed33be9784d3e7203070e43d54acf7b57238c3221643c01b3dc9

Observation 616311f7-9e01-49cd-bbf1-2bee4be39133 · outbound

This paper cites Scenecraft: An llm agent for synthesizing 3d scenes as blender code.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scenecraft: An llm agent for synthesizing 3d scenes as blender code

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.372486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:16.338669Z digest=sha256:df2e1dd9eaf0818e0f37b2fda2e204a9094d0b1cc34eae527f6158d34ac7eeba

Observation a1d01f54-7b0d-4d4d-ac16-f51c5f1babf8 · outbound

This paper cites The Scene Language: Representing Scenes with Programs, Words, and Embeddings.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Scene Language: Representing Scenes with Programs, Words, and Embeddings

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.453656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.453656Z digest=sha256:e26fa397964e831f0918016436b6e64346fcd0a13e3d348a3f9b94f128dc4ac6

Observation a6f75e8e-096c-4c4f-82dd-00bfbd5e88a2 · outbound

This paper cites 3D-GPT: Procedural 3D Modeling with Large Language Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3D-GPT: Procedural 3D Modeling with Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.602426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.602426Z digest=sha256:6350f5dafdea7a50753ca02df0759b582000692f6221f5dfb77a8c32fd651a50

Observation c0e1084c-6cc2-4184-af33-3f98af61bc67 · outbound

This paper cites SceneMotifCoder: Example-driven Visual Program Learning for Generating 3D Object Arrangements.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering SceneMotifCoder: Example-driven Visual Program Learning for Generating 3D Object Arrangements

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.700146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.700146Z digest=sha256:e8d767453ca4d5391325307c64b9c3084e821a45b038742e6be37485820de518

Observation 2abac4fd-99b5-4555-9d70-d6c9bc572221 · outbound

This paper cites L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.816617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.816617Z digest=sha256:765d6ae72dce920ac5e309cd07896c09de2c1af005b304a275286a3bf61fa381

Observation d6757f77-95e2-4fef-8b62-63e1a9ee112f · outbound

This paper cites Creative agents: Empowering agents with imagination for creative tasks.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Creative agents: Empowering agents with imagination for creative tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.883085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.883085Z digest=sha256:fcc0b4d559acf91291edbaa6115f45d33eafcaad342300ce17d5bbcd36189712

Observation 79bd3400-c7c6-46cb-b80a-a8beebee358c · outbound

This paper cites Scenex: Procedural controllable large-scale scene generation via large-language models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scenex: Procedural controllable large-scale scene generation via large-language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.193430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:16.958079Z digest=sha256:a3c252183d10d73e743f67ede7ef7352fc286131ba1805cc5ad285dfb6e9a0d1

Observation c70a0e45-e2d4-4063-9d60-a6cd36e02871 · outbound

This paper cites Program-guided image manipulators.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Program-guided image manipulators

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.997606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:17.034681Z digest=sha256:70222977acf7af2dfc2aab80a8375e0df34bd68b59efcff14af4441d8628e8c4

Observation 56d960d6-db08-4b58-8524-19e5526477d4 · outbound

This paper cites LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.103306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.103306Z digest=sha256:aaafd12ef0c9674df9406c134cd960f0251347400ecbd602b7ebdecc27f5cfde

Observation ef41b7f0-8b70-486f-bd19-089ce59ad090 · outbound

This paper cites Virtualhome: Simulating household activities via programs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Virtualhome: Simulating household activities via programs

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.787871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:17.197318Z digest=sha256:fdbdb91799c00e497d23f4f73bdf58044ab849dc90114485fbacfeb1cf72a656

Observation bfd141a2-30a4-4447-96ec-8d97ac1962a1 · outbound

This paper cites Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.637410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:17.287964Z digest=sha256:1b981aa5fe2ec68e48e52ce5219b7e669fab46a22bbb85a92025252421a8d631

Observation a65d344b-a34a-4ad8-8d9e-6228213936f2 · outbound

This paper cites MONet: Unsupervised Scene Decomposition and Representation.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering MONet: Unsupervised Scene Decomposition and Representation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.363407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.363407Z digest=sha256:f4f4013c42a189ff91a3131b32cc945636ae75e589bc6c3a5e77b7d6a357374e

Observation 6d8a2049-58fe-4748-bb8a-c73c40b2b546 · outbound

This paper cites GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.455262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.455262Z digest=sha256:e043ca7b235bd50a40af3d7661c7f4183431378ed5369d4222ff80e69313fe58

Observation a4f98124-b414-415b-a5d4-c4825010d2c2 · outbound

This paper cites Giraffe: Representing scenes as compositional genera- tive neural feature fields.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Giraffe: Representing scenes as compositional genera- tive neural feature fields

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.430818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:17.543148Z digest=sha256:8f0a7db07b8a989d5f04e3b0e227732fb9d7d53e94dc3d9d6819903c12c7b812

Observation e66a6f52-80df-424c-844f-adb73f893c23 · outbound

This paper cites A simple neural network module for relational reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering A simple neural network module for relational reasoning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.210637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:17.637002Z digest=sha256:8df660c321d26166eec94c2308ca532aeedba09bfcc881a646d07dcae16a9503

Observation 0c3d53fe-8b34-406e-a0f3-bcdcb4df24c8 · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Compositional Attention Networks for Machine Reasoning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.725712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.725712Z digest=sha256:b806d1f8e679013032a575fee60479071c34031a93b118882f2de8a0c8b5c66d

Observation b22eba6e-6dc0-4ade-97a4-8535b2058c59 · outbound

This paper cites Learning transferable visual models from natural language supervision.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning transferable visual models from natural language supervision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.799376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.799376Z digest=sha256:f64f33bd2b43b035fb01e9fa107a539b1575c5838096d25c4380abbac7e0b0bc

Observation e8e258f8-9e15-4947-b421-1b5902e24e86 · outbound

This paper cites Segment anything.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Segment anything

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.870461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.870461Z digest=sha256:4053e27383712afe98501676f4d68355e79b19a89b8af699ff5966481fc74467

Observation 3eda3f9c-55df-4792-a0c4-65ce2a8f4966 · outbound

This paper cites GPT-4 Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GPT-4 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.945267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.945267Z digest=sha256:2bcbde531311c7c55f15ea4c3594d697dbabedde12b9c96354a997c628531274

Observation 0d9e3326-48fc-4d65-ad73-61666236654d · outbound

This paper cites The Gemini 2 Model Family: Google Deepmind.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Gemini 2 Model Family: Google Deepmind

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:19.994411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:18.015298Z digest=sha256:b28be22750e21d2cad0a1e0d9df02db69e879e17ded47374b73225d08e4184f7

Observation 3c804746-c846-4535-9ac2-932819b67dd7 · outbound

This paper cites The Grok Model Family: xAI.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Grok Model Family: xAI

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:19.789753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:18.083136Z digest=sha256:defc569ba9f3afaad56e161d8bf00eeef90cde47d4988c054788adb7e8be425b

Observation fd58f6e1-3696-4dba-9f93-904ae09d03a6 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.167303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.167303Z digest=sha256:0752e2505c14c85c969dd8f49ce6a5bd4cb0cc3b1a456f7263b68608c81946d4

Observation 70584297-2efd-4ea7-a61b-8ecc834cb15a · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, 2024.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Llava-next: A strong zero-shot video understanding model, 2024

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:19.606170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:52:18.247348Z digest=sha256:38bdb4631740e9bc5007d8fb9dc99dc6e3625acd28753d39f92b0e3515b43ac6

Observation 4ba9f124-075a-4a11-962e-08cfa35bf935 · outbound

This paper cites The Llama 3 Herd of Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Llama 3 Herd of Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.335284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.335284Z digest=sha256:645a5d2c176b70e77cf8ba01102828a5853b9dc71a6cd99a28bd8a2477ac8a9b

Observation ddad7a53-b8f2-4202-8256-79c24da2b0b9 · outbound

This paper cites H2OVL-Mississippi Vision Language Models Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering H2OVL-Mississippi Vision Language Models Technical Report

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.427568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.427568Z digest=sha256:27958402d9923258290365a7249720d62fe4787dd5b248b8f1c51a9041dd2f48

Observation 33af74a0-f683-449d-ad39-85af9eb13b2c · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.492837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.492837Z digest=sha256:5b3bce3b423d77964d54f58147320d2c522956fc07e7d15bd7fa343aee5b7ba1

Observation f76e75a6-afe5-4cc8-b853-38164d5b91a9 · outbound

This paper cites Pixtral 12B.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Pixtral 12B

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.563315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.563315Z digest=sha256:233821ffe42177c697fcc9bca21140eb79c7a53fd9074fa788701ee7ace6b599

Observation 84e3669f-09fa-4e89-bf38-aecdd5b59f7a · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.651138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.651138Z digest=sha256:e4e398f6c6d1dffae4085205379e0f243212e9ee5149508c506ce32509ae4e82

Observation 66574c34-677a-4271-8694-6f96025fe335 · outbound

This paper cites OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.730268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.730268Z digest=sha256:26866ca66f70d0de8f7823cffb3e3de5542502a6fe7db6d9477366adcf846019

Observation b05f4901-3c8d-45b0-beb2-d76b513b1823 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.821199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.821199Z digest=sha256:e8dade0b8829f3c62a283f1aee43d1b7244d53796ef476b1077cd8e93a73c0e2

Observation 862d55c2-a190-4aa5-983e-591ccce7ca14 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.924233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.924233Z digest=sha256:4b4cbfbd888841d80e4f2a968ff3e43788b79bd4454efcf44143f93995f03fdf

Observation b1dbb342-9b92-4151-94cc-b8c1d253e240 · outbound

This paper cites Qwen2.5 Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Qwen2.5 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.990164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.990164Z digest=sha256:9a0d97d29495b8342d239fbd8d4c240dd9514a767574b20d0896824d820b29cc

Observation bbb341f6-1094-4204-bbad-c1ad333a3a6b · outbound

This paper cites Mistral 7B.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Mistral 7B

Reference 90

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:19.061479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:19.061479Z digest=sha256:5d7ea4a61073740043cb0bf8e5d00c14e300b85a2e556af3a3d8cd2165dbf784

Pith citing papers

Observation 58f1947f-c76c-4444-8630-cb9fb76e6631 · inbound

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cites this paper.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.147646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.147646Z digest=sha256:9daf6807e8af38bd26f9ca281c330500e62b29a88ffacd0cbbf3a61d65a22298

Observation 2a8e50f9-e90a-41bd-9177-58400071a09e · inbound

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis cites this paper.

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:03.448702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:36:33.149921Z digest=sha256:a25cd2e83b2d4eb7b83e91c58e0007eaff0098bdaf73f3e7f6574190d4cd1c48

Observation d5b28e77-12f4-45cc-8b14-f7dcdfcd832d · inbound

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models cites this paper.

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:46:19.352179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T15:04:01.773923Z digest=sha256:75a6f8f71255e9f45920fb1f9524803e32642b925332e28a4976f37a7f877085

Observation 239479b0-2d39-4cf4-b714-2e34cd9b6a29 · inbound

IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning cites this paper.

IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T01:11:59.721089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:11:59.721089Z digest=sha256:5ca4b3aa220b99a4307a2027d87ec0d2bb4448921636d5951c182c463691be74