Pith. sign in

Paper Citation Record · LEDGER

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

As of 20 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 25 inbound Pith citation observations for arXiv:2505.02836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02836 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:46:00.688693Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:13:37.518119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:09:37.256195Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e39183b6-3b5a-4a6c-8aa5-04f986831572 · outbound

This paper cites Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.427918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.427918Z digest=sha256:735d910f10578a9c3771ac99bcfa140712d8ce24978640d16f93bfae3d23e332

Observation 8d0f2b66-11e7-4c3a-a52b-71b972135940 · outbound

This paper cites Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.433408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.433408Z digest=sha256:e8fac0c03f3b427559060c088e649859c4990ed69d4a4cce068c072cbeb008e0

Observation 686103c4-9afa-4454-9fce-34ad462a010d · outbound

This paper cites Learning spatial knowledge for text to 3d scene generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Learning spatial knowledge for text to 3d scene generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.534244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.438498Z digest=sha256:7b8de7984f6a41f4df4e33a25866e421ce57b204505a5858d8c96a905148c958

Observation cc3a0045-b9c0-467b-aaae-0250f165621b · outbound

This paper cites Secondpose: Se (3)- consistent dual-stream feature fusion for category-level pose estimation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Secondpose: Se (3)- consistent dual-stream feature fusion for category-level pose estimation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.443590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.443590Z digest=sha256:5be79626b06c37e933393b44ed18fdc537feba5ea8736a1c60fa108f62f4624f

Observation c599146d-adb1-4516-9294-a18c944b9b36 · outbound

This paper cites Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.448345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.448345Z digest=sha256:3d7b0c26916548aaa22a45df5e420c587ab1da2a7a14037ae71d8ee1d6458994

Observation 734502df-a7e8-4db0-8728-692fe62e92a0 · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Procthor: Large-scale embodied ai using procedural generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.508644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.453471Z digest=sha256:38cf6c60ab843653104ad3b23cf1a2aea0e08a36472036485c6504ffa2f3d83d

Observation f2d12e95-80b9-4685-a289-71b22c9de74b · outbound

This paper cites Phone2proc: Bringing robust robots into our chaotic world.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Phone2proc: Bringing robust robots into our chaotic world

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.493711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.458005Z digest=sha256:5f9b98642d27f9301b72589439f164424d97d4accb05c1b39d569d9ec6fde7d7

Observation 3baa9fa7-fcb9-48b5-b40e-4ed0225fcc73 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Objaverse: A universe of annotated 3d objects

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.478193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.463051Z digest=sha256:44258b5e154d196d1180699b20af1c08bca1660b8ef309b5b64a368d5d6d29e3

Observation 1dbdc5b4-90bc-4513-a474-debb6857b192 · outbound

This paper cites Roma: Robust dense feature matching.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Roma: Robust dense feature matching

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.461028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.467459Z digest=sha256:e8a5de221fe9561f165e062c49ad39a9db0e16f917afe6c0797b1a0039696125

Observation 29b7e13a-c431-42c4-b6a8-10bfe4c66c0e · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.445571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.472070Z digest=sha256:ed829587f62b1469cef856f857c1e62e86cc1110f4e79ab22853ba05a44bd170

Observation 0bdf0fae-2176-47be-9e45-fa11483ae749 · outbound

This paper cites Scenescape: Text-driven consistent scene generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenescape: Text-driven consistent scene generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.430364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.476599Z digest=sha256:e9fb98f73e097385d5b2daa9da6dfc330028036b0a72771b29e0f095029adef5

Observation 6912fc79-bedf-43a1-89c3-066a6cb85caf · outbound

This paper cites 3d-future: 3d furni- ture shape with texture.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation 3d-future: 3d furni- ture shape with texture

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.414634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.481357Z digest=sha256:b1f5769d5bc6106f0a752513810a2e8ab019a032e939e062bcdcbbaaf363109b

Observation 884935c8-33d2-46e8-b2dc-472039b342cf · outbound

This paper cites ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.485632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.485632Z digest=sha256:a002366b12988623d54a85a5e2a145fd04f0899ea75d1457ff40a171b22663c2

Observation 5ec0f0c4-eecc-4b16-bacb-b3a3f1dd6cd5 · outbound

This paper cites Text2room: Extracting textured 3d meshes from 2d text-to-image models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Text2room: Extracting textured 3d meshes from 2d text-to-image models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.399192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.490450Z digest=sha256:58d826b8762040851570942cb0a4e9841881cc5e8d86d92b880d6e43742cfa7a

Observation a3aec485-98ec-4d27-8753-92b82068fcec · outbound

This paper cites Scenecraft: An llm agent for synthesizing 3d scenes as blender code.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenecraft: An llm agent for synthesizing 3d scenes as blender code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.494669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.494669Z digest=sha256:df5aad7b01106633dc0a627d0e7ad0c440ff7386fa7141cfaaf0da11b16944cf

Observation 38c1b6d4-d274-41b6-9685-a2ecf388381e · outbound

This paper cites GPT-4o System Card.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.499364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.499364Z digest=sha256:477d76ceae1e9750a164d095f7ebab91c30dd38442c2210b9bc390e6c54f75ff

Observation e22eef73-10de-4a20-ac16-f4ff3d092254 · outbound

This paper cites Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.374746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.504096Z digest=sha256:e1d9192d999ae28915574c941307dfa83ff87e9d402de6132b5da45e16c4d210

Observation bb0d66e1-0446-4196-a93a-3e97cd5a5478 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.508809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.508809Z digest=sha256:fd2f8af2ef26a2525033e46baea38ddf2585fee26a896ab463622dc5630a2eeb

Observation 1711031c-8e6d-4648-a83f-b689878bc674 · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Beyond the nav-graph: Vision-and-language navigation in continuous environments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.513869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.513869Z digest=sha256:96c656c7a160c1beac8697743dbf6fd89f7a9aabaa6edb7fb24d31c53cf0286d

Observation bea5f99f-8b9b-450e-9a6c-b8ec68cda999 · outbound

This paper cites Scenecraft: Automating interactive narrative scene generation in digital games with large language models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenecraft: Automating interactive narrative scene generation in digital games with large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.350361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.518657Z digest=sha256:6aeb89ece59c17616418438712d8652f10ebecc8a5e02fc39bc47a758013836e

Observation e127134a-a824-4808-bc7d-64b34f334854 · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 ev- eryday activities and realistic simulation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Behavior-1k: A benchmark for embodied ai with 1,000 ev- eryday activities and realistic simulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.523185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.523185Z digest=sha256:c304727bfc7c522c263f863f31997d76cb9babf2801185cd2e4c081c200595d6

Observation f23b76c8-e24d-46f2-a33f-c9dd7b7eba13 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen 14 image encoders and large language models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Blip- 2: Bootstrapping language-image pre-training with frozen 14 image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.325201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.527689Z digest=sha256:8a0cf1efffee680c5b833fb919f683b54efbbaf32e02778a4d5ba7159fe01106

Observation 2ea097fc-9af3-43fc-b2e0-986124f288cc · outbound

This paper cites Grains: Generative recur- sive autoencoders for indoor scenes.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Grains: Generative recur- sive autoencoders for indoor scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.309856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.532304Z digest=sha256:5aa736f0b3b0b64e5bba0284541b99cd19e326a5cecf2d9d82cd16383e895d11

Observation 41cf1899-5181-4d00-a72f-790c151566f0 · outbound

This paper cites Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.295297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.536749Z digest=sha256:83f1bfc11b448e5922cdf004ba50d5729a9e844eaaf553160c898d20237c136d

Observation 1ca3121f-7f22-4de1-96ee-403841fdfdfe · outbound

This paper cites InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.541434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.541434Z digest=sha256:50c8ceb56cf98e407905cbbeff2dd8b513a05334295161ad481d21402182aa8f

Observation a4d22872-d6d0-46fd-a396-79019b37761b · outbound

This paper cites Genusd: 3d scene generation made easy.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Genusd: 3d scene generation made easy

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.280210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.546207Z digest=sha256:575e756ae08a762081ef677bce1b700502a3963b064bbaf23fef3d5f63e60095

Observation 9520ca1a-b3ac-4df1-b2af-32e2998a1938 · outbound

This paper cites Eval- uating text-to-visual generation with image-to-text generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Eval- uating text-to-visual generation with image-to-text generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.263916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.550596Z digest=sha256:bc074a39a20e92c73eb3b572c389fa3eb83abdd4e33fba9281d92c8e60ed66c7

Observation 6b4bd02d-fc9f-4d9f-954c-65aa7dbb715f · outbound

This paper cites Eval- uating text-to-visual generation with image-to-text generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Eval- uating text-to-visual generation with image-to-text generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.246880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.555303Z digest=sha256:8b4d0978b90302b840060fb846ebae8888a8952619cb6a0c8541b14b77e5a61f

Observation 4403f879-94c0-4c53-87d3-046b3fe9f53f · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.230960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.559737Z digest=sha256:27a142450541530384661e8083ede488b832e0e515f193ac3c4dccea2da81e91

Observation 18b14a8e-3c74-4372-bc8c-6a5aea98cd5c · outbound

This paper cites End-to-end optimization of scene layout.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation End-to-end optimization of scene layout

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.215834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.564191Z digest=sha256:43a669203e11a952e3a11de5722dce784621e7e055a435c11f30e933ec8359e9

Observation f767d047-803f-419f-92fb-f36f5c59821e · outbound

This paper cites Warp: A high-performance python frame- work for gpu simulation and graphics.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Warp: A high-performance python frame- work for gpu simulation and graphics

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.201150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.568666Z digest=sha256:8bbaf27f1e6b4177e05ddae3646324078bc1dd13cb33754c48023ec50db5b758

Observation 5499666d-a9b9-4393-b058-1a35e86065d5 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.573094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.573094Z digest=sha256:efdbf6de9efc84235a8b8a7a303935f96684b04c6be7cc6662c26bc0c507c118

Observation d702ca41-f367-4d30-ac7a-e8bb66ed1a0a · outbound

This paper cites Sceneteller: Language-to-3d scene generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Sceneteller: Language-to-3d scene generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.185950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.577945Z digest=sha256:0a05f69be69058d7ac69a4b4eed12577444937a64a23eb95c88acccaf7fa26d0

Observation 0d8798b5-8f5c-4d45-901c-9e7db6e398fb · outbound

This paper cites Atiss: Autoregres- sive transformers for indoor scene synthesis.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Atiss: Autoregres- sive transformers for indoor scene synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.582381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.582381Z digest=sha256:124bf70e4a3beec14b8705f2122c7ef34aa8dd1129c80d38dadd000f4f30b431

Observation 8d7cae58-513b-480f-8ac6-0280cde88938 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Pytorch: An imperative style, high-performance deep learning library

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.587767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.587767Z digest=sha256:90da1205d6b3274752217cffc4beb5f9721fd8bda8ccdafb425ddbe92d86f8d3

Observation d88d00f3-c3c8-452c-b8d9-dc85155222c6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.592392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.592392Z digest=sha256:adb588bce91d0c2370f25c36db301e6fde17aacd1aa66b1cf3dce0c12bfba852

Observation 27a30eb3-1573-42a2-bc20-c1ffee7e8827 · outbound

This paper cites Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.140983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.596875Z digest=sha256:312864e2d75afead330fe3f35edf7689a6b70665dd178a18ef3fa84e9fdbdbab

Observation 5d5195f3-3016-4302-8371-b5b69dbd24e0 · outbound

This paper cites Accelerating 3D Deep Learning with PyTorch3D.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Accelerating 3D Deep Learning with PyTorch3D

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.601729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.601729Z digest=sha256:f0ed8ca2a6d04b1471925fbaa904531a965dcd1fe598d5dedbfac70f537d29bd

Observation 9a7f72a1-2741-40cc-b36e-36058ab1e097 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.606616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.606616Z digest=sha256:bed3245f1e682d28f0814f0be3aa70518a55caa12622754a3512426a004db9f4

Observation eda677e5-5435-4db7-9840-a55a897ab99a · outbound

This paper cites LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.612079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.612079Z digest=sha256:51ac10347692101f28707b7199de0e3aaf323840ed9488545dcd420ad2feac26

Observation f91fcb89-3807-4832-8888-cf720ebc6c16 · outbound

This paper cites Diffuscene: Denoising diffusion models for generative indoor scene synthesis.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Diffuscene: Denoising diffusion models for generative indoor scene synthesis

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.125500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.616769Z digest=sha256:5f18019eebe845568ee65fc3729d69cc6ad75e662869d9231f596fcf62cb9085

Observation 0226a29a-7dcc-413e-832e-624c5c745b55 · outbound

This paper cites Planit: Planning and in- stantiating indoor scenes with relation graph and spatial prior networks.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Planit: Planning and in- stantiating indoor scenes with relation graph and spatial prior networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.109369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.621850Z digest=sha256:3e3c592868c020b26acd38ce72d2842e1a2c7977110377da94bf24ca52926cc6

Observation 9f9d1ccd-319f-4036-a8ed-0f817e3767e3 · outbound

This paper cites Sceneformer: Indoor scene generation with transformers.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Sceneformer: Indoor scene generation with transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.093355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.626554Z digest=sha256:2bf74989701fc06871685e9000954901fe59ec00c1fc8e6d5f4b42509d69ac3a

Observation 0bd09022-012e-407a-9c1e-35e9acdd0f89 · outbound

This paper cites RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.631699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.631699Z digest=sha256:3e4b36c4787e7c0b44a3c2b9ce850876533ef959d9586056d49b8cabcd53a67c

Observation 48a97787-726b-4479-8a5c-ea4fb01d055e · outbound

This paper cites Architect: Generating vivid and interactive 3d 15 scenes with hierarchical 2d inpainting.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Architect: Generating vivid and interactive 3d 15 scenes with hierarchical 2d inpainting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.077615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.636551Z digest=sha256:50c81b7b567f95b5e51aa6e024ae6947b1ae1ea585e7c8af8ddccca299af090a

Observation 329e26e4-99eb-4d54-8e0d-ee8e4f0d8219 · outbound

This paper cites Reconfusion: 3d re- construction with diffusion priors.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Reconfusion: 3d re- construction with diffusion priors

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.061269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.641087Z digest=sha256:be9ef804d175249142946bf42215eee4bb4fd9f2fb57b696b5baf3431921d25e

Observation 6cb4c01c-6f37-4522-8138-d4db96f93df5 · outbound

This paper cites Physcene: Physically interactable 3d scene synthesis for em- bodied ai.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Physcene: Physically interactable 3d scene synthesis for em- bodied ai

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.045517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.645972Z digest=sha256:e07edb76cb35a4d0299fd302f5fab65f7b079d3192ad4e7e3f72f1cb7aad0f89

Observation c269309d-abd9-4e4a-b9d8-a4f926e29e82 · outbound

This paper cites Holodeck: Language guided gen- eration of 3d embodied ai environments.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Holodeck: Language guided gen- eration of 3d embodied ai environments

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.029965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.650570Z digest=sha256:40f514d3b72e8d8087a9f3b1af16ae07b6defbe2ed2d9cda515880bbb5321eef

Observation f1a77a23-8ef8-466c-8c60-8308639681e7 · outbound

This paper cites WonderWorld: Interactive 3D Scene Generation from a Single Image.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation WonderWorld: Interactive 3D Scene Generation from a Single Image

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.655062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.655062Z digest=sha256:eba23657cc181b6524aa7633d380d08f96eb489f9698c77af12d869b5b57c9d2

Observation 1c83ff4d-58fb-46ab-a5f7-3ec22870a1ab · outbound

This paper cites Wonderjourney: Going from anywhere to everywhere.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Wonderjourney: Going from anywhere to everywhere

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.659772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.659772Z digest=sha256:b65eda8486bc37834042f82ce1a691f77d3432ecf7ad753fec3025ea20663b3d

Observation 86453762-8db4-49c6-8959-d4e8443694e9 · outbound

This paper cites Telling left from right: Identifying geometry-aware semantic corre- spondence.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Telling left from right: Identifying geometry-aware semantic corre- spondence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.004290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.664049Z digest=sha256:60cb1f4d8f9ad6e59904259ad25e21ebb444f5ba3e4f4444c3119a4dd0169a54

Observation 4bc4cd6f-dc3c-426c-8b19-6ba5a8d203a6 · outbound

This paper cites Text2nerf: Text-driven 3d scene generation with neu- ral radiance fields.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Text2nerf: Text-driven 3d scene generation with neu- ral radiance fields

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:00.986549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.669451Z digest=sha256:56d95cb0f7da23436312cf46814d0b388a183cb4e6189cac9512e234de79b819

Observation 030a6a4a-9f77-4122-a2f4-64e7f07793c1 · outbound

This paper cites Zero-Shot Scene Reconstruction from Single Images with Deep Prior Assembly.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Zero-Shot Scene Reconstruction from Single Images with Deep Prior Assembly

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.674162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.674162Z digest=sha256:ea920c0803dd826a39f1982efb63317d9adad83887a16d458b9ae68f8a27cd55

Observation 9b0a44c7-46e4-4a01-9a18-de8412241f5e · outbound

This paper cites SceneX: Procedural Controllable Large-scale Scene Generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation SceneX: Procedural Controllable Large-scale Scene Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.679009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.679009Z digest=sha256:8fca9d1e5903fe7c23025e2ec4a73da7a6ad3beb68ebe11648f2f2b6719dfe80

Observation 885cbb08-0a9f-4312-b9ab-12a88fd6b039 · outbound

This paper cites GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.683788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.683788Z digest=sha256:b8399212d9fde34f8627f2a890d3e581353d6b656a9f1f4b7b6a2315f06790d2

Observation 98425961-cc68-497b-a9c7-7ccc4002016e · outbound

This paper cites Scenegraphnet: Neural message passing for 3d indoor scene augmentation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenegraphnet: Neural message passing for 3d indoor scene augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:00.971298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:46:00.688693Z digest=sha256:c287db262aac62bad956db7d71d01921112323245abb880bb720652ef50e7328

Pith citing papers

Observation 5b32ae89-6483-435b-9982-647d8edc2a7b · inbound

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding cites this paper.

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:43.104278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:43.104278Z digest=sha256:3a9d687603599f1c952d39e034651b81761173738387aaf766281c02eb767b52

Observation a31b2567-2df7-46dd-ab93-3d480503aff5 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:32.850883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:32.850883Z digest=sha256:0c3592c700ab5d23bd476c1c2c377f72a4af91cf9b5baf0eec90c04cee461574

Observation f97302bc-978c-48e6-a955-c770b23c387b · inbound

Video Perception Models for 3D Scene Synthesis cites this paper.

Video Perception Models for 3D Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:59.729438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:59.729438Z digest=sha256:888b7fda06f1123b62128b83ca424736d943bad7c4cf27c40a33c12a0da400c9

Observation a4081951-88c8-47fd-88d2-e67f97dafcec · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:42.305835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:42.305835Z digest=sha256:611ed99cab90791b5b0f66ed7a94e02ebb1ff7b2be17a326f228f6263b730a5c

Observation c0daa79a-4e47-481e-a554-80e5d332ed29 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.368585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.368585Z digest=sha256:7152ec8d4232b2ca212b68c475bc09ae1523f0960c56489af2bbb283d6f860a2

Observation 85383ed5-aad4-4b7c-b011-5beca377713b · inbound

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation cites this paper.

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T19:19:20.992326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:19:20.992326Z digest=sha256:135e93915724f33eea4bb4dc073351dff436e3463d016c0afefb2330931510a5

Observation 36e4d326-5360-4fd9-9be8-22eb11c93a07 · inbound

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs cites this paper.

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.648939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T15:07:54.930273Z digest=sha256:688ceeb9aee74521c6dad821ea946d03631482e78bc888f8016a4909a875ba87

Observation 26543a02-6378-46d7-9a85-33f1673708c1 · inbound

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning cites this paper.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.560832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:389649f256b4f684a0f5fa758cda9577c5cf84fc397d3d77a77967831ffb2845

Observation e1a06504-1b43-4f3b-9a11-30a8982adfca · inbound

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models cites this paper.

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:05.932363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:34:15.916685Z digest=sha256:f6b39792fc1b784855bd35188072ed3bfc58afd277730679cec31dff6b555c7c

Observation e693439b-16c8-4d60-83b0-4fc46678f9d3 · inbound

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis cites this paper.

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:11:03.657154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:36:33.149921Z digest=sha256:4cee4e45b0663eed5d84ebf29e3fa8a8e14e9ced25cac3de4be3a0ace9184af0

Observation 3d6782b4-b96a-4ba2-bdf7-ad05b5150363 · inbound

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development cites this paper.

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.232494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:34:58.288245Z digest=sha256:48969ea237c1059da2e4a5ffcc03ad6366bb8995b1c1201bf34c00eea0f792c1

Observation 057d08e1-e449-482e-bb80-fe41c9e25c24 · inbound

Repurposing 3D Generative Model for Autoregressive Layout Generation cites this paper.

Repurposing 3D Generative Model for Autoregressive Layout Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.732757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T08:09:00.779456Z digest=sha256:2cbae61873a3cdac0f9084de2df39adb9e46971b7884ef212cd88efa45c4c406

Observation f5417681-3361-451a-9452-ae11b24206a1 · inbound

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion cites this paper.

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:36.978937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T09:22:10.651922Z digest=sha256:477fce6d243ebbeb2a3fba6582727b2e959aff68e779dd67a6c1dbf251504905

Observation 05f27cc3-095e-4f35-800e-ad8400f5a2e8 · inbound

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation cites this paper.

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:06.580065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T03:22:30.012610Z digest=sha256:893f3e1624b9eb36836ae65ae718f7197b1d88e4c8e72adcfb4cd760f326032f

Observation 4112b671-dff6-4520-b981-7c75a03eb637 · inbound

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models cites this paper.

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:59:36.508177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T04:55:40.370796Z digest=sha256:2dabadf9dfa6cc68520daf5be1e336b7c8e77fb589b3db4af9d7a6487bc2537e

Observation d46764da-2d11-4041-85f1-8298ebb0421b · inbound

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation cites this paper.

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.323641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T23:11:59.282627Z digest=sha256:3df2132144cb02664d2e1d3e3c7370cdd0ca6447e973de06edf2f1c710dc3477

Observation 975ec25c-94f1-4a7a-a71b-d301f9dc1282 · inbound

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image cites this paper.

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.366072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T10:59:17.553098Z digest=sha256:b51db2a984bc917a6b93c0879110e477a20d13bbe2f7cf137734b7a6f5963678

Observation 8ffe7006-363f-4c1e-ba7a-92183b4b77d6 · inbound

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration cites this paper.

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:26.723110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T19:08:08.329817Z digest=sha256:75b01e8d6e59e77323004a01d1dac7bf262e95a10f836293ddb3165e86e2f7c5

Observation ee327f72-b816-4837-b399-d85f8f7a3d3a · inbound

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction cites this paper.

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:09:37.258248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T14:50:16.317301Z digest=sha256:9212457035e504a1689a9dc3a7176a1343eb846498e27869c5ffef4076ad00c3

Observation d0407fcf-3f38-47fc-a0ff-d6f490c6f095 · inbound

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation cites this paper.

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:14:25.460428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T08:12:30.693936Z digest=sha256:306b712621af7271712ea18fbbee484ee78fb36797b6429df55ca419c8d09e85

Observation 117dbc72-dd2b-4e11-9aad-123f302ebc8f · inbound

One Video, One World: Turning Monocular Video into Physical 4D Scenes cites this paper.

One Video, One World: Turning Monocular Video into Physical 4D Scenes Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:45.028897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T05:38:19.541391Z digest=sha256:c1234cf6d360c403d9c18874c6c67de7eeaabe5bff9533404ebcf7f64d2bcff6

Observation 9b69de31-9219-4245-be32-20415ad32f18 · inbound

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image cites this paper.

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T22:29:10.666143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:29:10.666143Z digest=sha256:9ea082e58c3725fa9f1fb7a9161d9d0af095d0c8fa8c06d7200e139bb7e5b905

Observation 0973aaf3-5216-40b5-b911-1f59fe8edfde · inbound

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis cites this paper.

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:57:31.940710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:57:31.940710Z digest=sha256:c76b284ba1d1d0374777c684cf65b1fece179d7159fb9c9ae1ae68639fac4973

Observation d469ef02-472c-40dc-9508-54510118a5f4 · inbound

Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation cites this paper.

Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:56.584297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:56.584297Z digest=sha256:6c76ebc0537ded2f4c9e9579e6163e59a0caac24b83aff37b3f6de72b4f2bc4d

Observation 59fb88cd-8d6b-4934-a980-eb680a56c41f · inbound

WorldClaw: Agentic 3D Open-World Generation at Scale cites this paper.

WorldClaw: Agentic 3D Open-World Generation at Scale Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:37.518119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:37.518119Z digest=sha256:4fb4a37aa6452fdd4a1f53121b2680831469d17bc510ecbb40d869a8f987328e