Pith. sign in

Paper Citation Record · LEDGER

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

As of 20 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 25 inbound Pith citation observations for arXiv:2505.02836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02836 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:46:00.688693Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:13:37.518119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:09:37.256195Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e39183b6-3b5a-4a6c-8aa5-04f986831572 · outbound

This paper cites Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.427918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.427918Z digest=sha256:735d910f10578a9c3771ac99bcfa140712d8ce24978640d16f93bfae3d23e332

Observation 8d0f2b66-11e7-4c3a-a52b-71b972135940 · outbound

This paper cites Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.433408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.433408Z digest=sha256:e8fac0c03f3b427559060c088e649859c4990ed69d4a4cce068c072cbeb008e0

Observation 686103c4-9afa-4454-9fce-34ad462a010d · outbound

This paper cites Learning spatial knowledge for text to 3d scene generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Learning spatial knowledge for text to 3d scene generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.534244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.438498Z digest=sha256:e537017b6d87c5781900439f0e8a696476a52cf6b27ee1dd7cbc7c38406bc46f

Observation cc3a0045-b9c0-467b-aaae-0250f165621b · outbound

This paper cites Secondpose: Se (3)- consistent dual-stream feature fusion for category-level pose estimation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Secondpose: Se (3)- consistent dual-stream feature fusion for category-level pose estimation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.443590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.443590Z digest=sha256:5be79626b06c37e933393b44ed18fdc537feba5ea8736a1c60fa108f62f4624f

Observation c599146d-adb1-4516-9294-a18c944b9b36 · outbound

This paper cites Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.448345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.448345Z digest=sha256:3d7b0c26916548aaa22a45df5e420c587ab1da2a7a14037ae71d8ee1d6458994

Observation 734502df-a7e8-4db0-8728-692fe62e92a0 · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Procthor: Large-scale embodied ai using procedural generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.508644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.453471Z digest=sha256:b6eb030c44bd4301c76f7b213643be116ad65a4469f920e332c926c4d3bcb355

Observation f2d12e95-80b9-4685-a289-71b22c9de74b · outbound

This paper cites Phone2proc: Bringing robust robots into our chaotic world.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Phone2proc: Bringing robust robots into our chaotic world

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.493711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.458005Z digest=sha256:cc4c3ecbda8f7390e2a96d0c1d81897a79a2040ce5e19b289d9993893fb041a5

Observation 3baa9fa7-fcb9-48b5-b40e-4ed0225fcc73 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Objaverse: A universe of annotated 3d objects

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.478193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.463051Z digest=sha256:a1baf679c48ddaee992c801cd0e187b2141c893793b17cf0efc8678a74b4c84f

Observation 1dbdc5b4-90bc-4513-a474-debb6857b192 · outbound

This paper cites Roma: Robust dense feature matching.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Roma: Robust dense feature matching

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.461028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.467459Z digest=sha256:6797871b50015f6675e611601c4fabc1c11f618df2b660e386abf14e3a1b5b37

Observation 29b7e13a-c431-42c4-b6a8-10bfe4c66c0e · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.445571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.472070Z digest=sha256:4cff2b0ea1238b9445ffafd70a2f616d40d365c0decec20d29c7f5fecb72390c

Observation 0bdf0fae-2176-47be-9e45-fa11483ae749 · outbound

This paper cites Scenescape: Text-driven consistent scene generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenescape: Text-driven consistent scene generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.430364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.476599Z digest=sha256:fa6bc36cd2d95f69dd89f13109a3c79ef8a925714e41ea051ca50d81da265106

Observation 6912fc79-bedf-43a1-89c3-066a6cb85caf · outbound

This paper cites 3d-future: 3d furni- ture shape with texture.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation 3d-future: 3d furni- ture shape with texture

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.414634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.481357Z digest=sha256:5ede88330bab2c2c916c897db9c5b57c586eae717ec626e8ef4bbd45648aaa35

Observation 884935c8-33d2-46e8-b2dc-472039b342cf · outbound

This paper cites ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.485632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.485632Z digest=sha256:a002366b12988623d54a85a5e2a145fd04f0899ea75d1457ff40a171b22663c2

Observation 5ec0f0c4-eecc-4b16-bacb-b3a3f1dd6cd5 · outbound

This paper cites Text2room: Extracting textured 3d meshes from 2d text-to-image models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Text2room: Extracting textured 3d meshes from 2d text-to-image models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.399192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.490450Z digest=sha256:947897482eb796c1a00285e90300cda0daea2e66453d8c3a1ead41d8f2707670

Observation a3aec485-98ec-4d27-8753-92b82068fcec · outbound

This paper cites Scenecraft: An llm agent for synthesizing 3d scenes as blender code.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenecraft: An llm agent for synthesizing 3d scenes as blender code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.494669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.494669Z digest=sha256:df5aad7b01106633dc0a627d0e7ad0c440ff7386fa7141cfaaf0da11b16944cf

Observation 38c1b6d4-d274-41b6-9685-a2ecf388381e · outbound

This paper cites GPT-4o System Card.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.499364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.499364Z digest=sha256:477d76ceae1e9750a164d095f7ebab91c30dd38442c2210b9bc390e6c54f75ff

Observation e22eef73-10de-4a20-ac16-f4ff3d092254 · outbound

This paper cites Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.374746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.504096Z digest=sha256:5711293ec865da16c2a7db23630d30d5a4ad2d3db29f769266c968dcc5c2e8b9

Observation bb0d66e1-0446-4196-a93a-3e97cd5a5478 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.508809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.508809Z digest=sha256:fd2f8af2ef26a2525033e46baea38ddf2585fee26a896ab463622dc5630a2eeb

Observation 1711031c-8e6d-4648-a83f-b689878bc674 · outbound

This paper cites Beyond the nav-graph: Vision-and-language navigation in continuous environments.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Beyond the nav-graph: Vision-and-language navigation in continuous environments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.513869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.513869Z digest=sha256:96c656c7a160c1beac8697743dbf6fd89f7a9aabaa6edb7fb24d31c53cf0286d

Observation bea5f99f-8b9b-450e-9a6c-b8ec68cda999 · outbound

This paper cites Scenecraft: Automating interactive narrative scene generation in digital games with large language models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenecraft: Automating interactive narrative scene generation in digital games with large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.350361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.518657Z digest=sha256:576153f2e0351e42a1dd931e81a7e457e442cf1c552b29fdcd0418ecc6641de7

Observation e127134a-a824-4808-bc7d-64b34f334854 · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 ev- eryday activities and realistic simulation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Behavior-1k: A benchmark for embodied ai with 1,000 ev- eryday activities and realistic simulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.523185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.523185Z digest=sha256:c304727bfc7c522c263f863f31997d76cb9babf2801185cd2e4c081c200595d6

Observation f23b76c8-e24d-46f2-a33f-c9dd7b7eba13 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen 14 image encoders and large language models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Blip- 2: Bootstrapping language-image pre-training with frozen 14 image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.325201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.527689Z digest=sha256:0400888fff2c0d5f6f8d1dbd82f6ace0b813d165cc5480c391b38aa72716a619

Observation 2ea097fc-9af3-43fc-b2e0-986124f288cc · outbound

This paper cites Grains: Generative recur- sive autoencoders for indoor scenes.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Grains: Generative recur- sive autoencoders for indoor scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.309856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.532304Z digest=sha256:de20146530ad51488b623f04b7043b5b8a638d8f7a09d8050f782e9636738c26

Observation 41cf1899-5181-4d00-a72f-790c151566f0 · outbound

This paper cites Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.295297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.536749Z digest=sha256:f8bde7058bccd3f826ed32e3aecf335af2d662a74a2523fa512d565ea0e0e801

Observation 1ca3121f-7f22-4de1-96ee-403841fdfdfe · outbound

This paper cites InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.541434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.541434Z digest=sha256:50c8ceb56cf98e407905cbbeff2dd8b513a05334295161ad481d21402182aa8f

Observation a4d22872-d6d0-46fd-a396-79019b37761b · outbound

This paper cites Genusd: 3d scene generation made easy.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Genusd: 3d scene generation made easy

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.280210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.546207Z digest=sha256:8399b3b9d8432cd6ee2702f3ecaa23083a08e38c51756ba8cbb71948e1174a7a

Observation 9520ca1a-b3ac-4df1-b2af-32e2998a1938 · outbound

This paper cites Eval- uating text-to-visual generation with image-to-text generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Eval- uating text-to-visual generation with image-to-text generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.263916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.550596Z digest=sha256:086f0042e3e0d6d4c8f42b69209623f17fd89a188ce7e3f2f2d795b23de90ed3

Observation 6b4bd02d-fc9f-4d9f-954c-65aa7dbb715f · outbound

This paper cites Eval- uating text-to-visual generation with image-to-text generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Eval- uating text-to-visual generation with image-to-text generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.246880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.555303Z digest=sha256:6e3fa34aca10a6d5e3d3b84f4e718f6070a5686f6209dc84d7977b9d7ecb07ff

Observation 4403f879-94c0-4c53-87d3-046b3fe9f53f · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.230960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.559737Z digest=sha256:e87de2195c69b7c9ce6c138fdb5880256db9861ec121c982c8ac12f7d5662599

Observation 18b14a8e-3c74-4372-bc8c-6a5aea98cd5c · outbound

This paper cites End-to-end optimization of scene layout.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation End-to-end optimization of scene layout

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.215834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.564191Z digest=sha256:477246342789eef19f8f499f715c4a5b2ac60471e8beb77171739db2e82f470e

Observation f767d047-803f-419f-92fb-f36f5c59821e · outbound

This paper cites Warp: A high-performance python frame- work for gpu simulation and graphics.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Warp: A high-performance python frame- work for gpu simulation and graphics

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.201150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.568666Z digest=sha256:a4cbcb300880cbe06e39e3fc61fedbbb0dea6aa67a6bb9e942a8b53237f28105

Observation 5499666d-a9b9-4393-b058-1a35e86065d5 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.573094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.573094Z digest=sha256:efdbf6de9efc84235a8b8a7a303935f96684b04c6be7cc6662c26bc0c507c118

Observation d702ca41-f367-4d30-ac7a-e8bb66ed1a0a · outbound

This paper cites Sceneteller: Language-to-3d scene generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Sceneteller: Language-to-3d scene generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.185950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.577945Z digest=sha256:c14450a2cc5dd8be9182c55c157de3b38c358b88c69c14d6f85b0000da69bc1f

Observation 0d8798b5-8f5c-4d45-901c-9e7db6e398fb · outbound

This paper cites Atiss: Autoregres- sive transformers for indoor scene synthesis.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Atiss: Autoregres- sive transformers for indoor scene synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.582381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.582381Z digest=sha256:124bf70e4a3beec14b8705f2122c7ef34aa8dd1129c80d38dadd000f4f30b431

Observation 8d7cae58-513b-480f-8ac6-0280cde88938 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Pytorch: An imperative style, high-performance deep learning library

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.587767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.587767Z digest=sha256:90da1205d6b3274752217cffc4beb5f9721fd8bda8ccdafb425ddbe92d86f8d3

Observation d88d00f3-c3c8-452c-b8d9-dc85155222c6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.592392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.592392Z digest=sha256:adb588bce91d0c2370f25c36db301e6fde17aacd1aa66b1cf3dce0c12bfba852

Observation 27a30eb3-1573-42a2-bc20-c1ffee7e8827 · outbound

This paper cites Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.140983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.596875Z digest=sha256:0b23387b949c0b646970cd89861f78994c65a081cc837ebda67703efef69ada1

Observation 5d5195f3-3016-4302-8371-b5b69dbd24e0 · outbound

This paper cites Accelerating 3D Deep Learning with PyTorch3D.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Accelerating 3D Deep Learning with PyTorch3D

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.601729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.601729Z digest=sha256:f0ed8ca2a6d04b1471925fbaa904531a965dcd1fe598d5dedbfac70f537d29bd

Observation 9a7f72a1-2741-40cc-b36e-36058ab1e097 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.606616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.606616Z digest=sha256:bed3245f1e682d28f0814f0be3aa70518a55caa12622754a3512426a004db9f4

Observation eda677e5-5435-4db7-9840-a55a897ab99a · outbound

This paper cites LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.612079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.612079Z digest=sha256:51ac10347692101f28707b7199de0e3aaf323840ed9488545dcd420ad2feac26

Observation f91fcb89-3807-4832-8888-cf720ebc6c16 · outbound

This paper cites Diffuscene: Denoising diffusion models for generative indoor scene synthesis.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Diffuscene: Denoising diffusion models for generative indoor scene synthesis

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.125500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.616769Z digest=sha256:3dec977b8706f4eb8b86f6ef0302c21c6fa32ab02c1f2e01190569d5a31866b2

Observation 0226a29a-7dcc-413e-832e-624c5c745b55 · outbound

This paper cites Planit: Planning and in- stantiating indoor scenes with relation graph and spatial prior networks.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Planit: Planning and in- stantiating indoor scenes with relation graph and spatial prior networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.109369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.621850Z digest=sha256:3590af6a719419acf6e067353648827211828515866291b5f7f54e6481d442ff

Observation 9f9d1ccd-319f-4036-a8ed-0f817e3767e3 · outbound

This paper cites Sceneformer: Indoor scene generation with transformers.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Sceneformer: Indoor scene generation with transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.093355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.626554Z digest=sha256:5c0e6f59021d317acffb928c6f9579e4bde49a8b622a394471f976b9c2a3454e

Observation 0bd09022-012e-407a-9c1e-35e9acdd0f89 · outbound

This paper cites RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.631699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.631699Z digest=sha256:3e4b36c4787e7c0b44a3c2b9ce850876533ef959d9586056d49b8cabcd53a67c

Observation 48a97787-726b-4479-8a5c-ea4fb01d055e · outbound

This paper cites Architect: Generating vivid and interactive 3d 15 scenes with hierarchical 2d inpainting.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Architect: Generating vivid and interactive 3d 15 scenes with hierarchical 2d inpainting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.077615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.636551Z digest=sha256:26b83f17b523f4dac56f879fdaa3267a4aa47ea042e674f742e62356be3bd077

Observation 329e26e4-99eb-4d54-8e0d-ee8e4f0d8219 · outbound

This paper cites Reconfusion: 3d re- construction with diffusion priors.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Reconfusion: 3d re- construction with diffusion priors

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.061269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.641087Z digest=sha256:742d2a6695f341e65958bfefabd17edba8dfc77949e77ca07c11019cc956d068

Observation 6cb4c01c-6f37-4522-8138-d4db96f93df5 · outbound

This paper cites Physcene: Physically interactable 3d scene synthesis for em- bodied ai.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Physcene: Physically interactable 3d scene synthesis for em- bodied ai

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.045517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.645972Z digest=sha256:322d106107c632c57298ccdfc5fdf6b4233f1704863c671402d0590ddc53ce1b

Observation c269309d-abd9-4e4a-b9d8-a4f926e29e82 · outbound

This paper cites Holodeck: Language guided gen- eration of 3d embodied ai environments.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Holodeck: Language guided gen- eration of 3d embodied ai environments

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.029965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.650570Z digest=sha256:66ef0d5e294389327ea77e59939fddd7b16ba5641cce26cd88325dbdd0f70a11

Observation f1a77a23-8ef8-466c-8c60-8308639681e7 · outbound

This paper cites WonderWorld: Interactive 3D Scene Generation from a Single Image.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation WonderWorld: Interactive 3D Scene Generation from a Single Image

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.655062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.655062Z digest=sha256:eba23657cc181b6524aa7633d380d08f96eb489f9698c77af12d869b5b57c9d2

Observation 1c83ff4d-58fb-46ab-a5f7-3ec22870a1ab · outbound

This paper cites Wonderjourney: Going from anywhere to everywhere.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Wonderjourney: Going from anywhere to everywhere

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.659772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.659772Z digest=sha256:b65eda8486bc37834042f82ce1a691f77d3432ecf7ad753fec3025ea20663b3d

Observation 86453762-8db4-49c6-8959-d4e8443694e9 · outbound

This paper cites Telling left from right: Identifying geometry-aware semantic corre- spondence.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Telling left from right: Identifying geometry-aware semantic corre- spondence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:01.004290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.664049Z digest=sha256:4e082493613484b27d5e35111bc3d0d3ad1b8f2cef03fca040ad72c8cc828eef

Observation 4bc4cd6f-dc3c-426c-8b19-6ba5a8d203a6 · outbound

This paper cites Text2nerf: Text-driven 3d scene generation with neu- ral radiance fields.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Text2nerf: Text-driven 3d scene generation with neu- ral radiance fields

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:00.986549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.669451Z digest=sha256:f0d56576da798bd6c7ae5424f40c93c8f80565f7fe2b2710191ad8770b13e2ad

Observation 030a6a4a-9f77-4122-a2f4-64e7f07793c1 · outbound

This paper cites Zero-Shot Scene Reconstruction from Single Images with Deep Prior Assembly.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Zero-Shot Scene Reconstruction from Single Images with Deep Prior Assembly

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.674162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.674162Z digest=sha256:ea920c0803dd826a39f1982efb63317d9adad83887a16d458b9ae68f8a27cd55

Observation 9b0a44c7-46e4-4a01-9a18-de8412241f5e · outbound

This paper cites SceneX: Procedural Controllable Large-scale Scene Generation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation SceneX: Procedural Controllable Large-scale Scene Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.679009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.679009Z digest=sha256:8fca9d1e5903fe7c23025e2ec4a73da7a6ad3beb68ebe11648f2f2b6719dfe80

Observation 885cbb08-0a9f-4312-b9ab-12a88fd6b039 · outbound

This paper cites GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:46:00.683788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:46:00.683788Z digest=sha256:b8399212d9fde34f8627f2a890d3e581353d6b656a9f1f4b7b6a2315f06790d2

Observation 98425961-cc68-497b-a9c7-7ccc4002016e · outbound

This paper cites Scenegraphnet: Neural message passing for 3d indoor scene augmentation.

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation Scenegraphnet: Neural message passing for 3d indoor scene augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:46:00.971298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:46:00.688693Z digest=sha256:27dd5f4a9dfca6bbf64d3da151d609ab7ba854cd5f055f3d6f2b7379c9fef10a

Pith citing papers

Observation 5b32ae89-6483-435b-9982-647d8edc2a7b · inbound

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding cites this paper.

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:43.104278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:43.104278Z digest=sha256:3a9d687603599f1c952d39e034651b81761173738387aaf766281c02eb767b52

Observation a31b2567-2df7-46dd-ab93-3d480503aff5 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:32.850883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:32.850883Z digest=sha256:0c3592c700ab5d23bd476c1c2c377f72a4af91cf9b5baf0eec90c04cee461574

Observation f97302bc-978c-48e6-a955-c770b23c387b · inbound

Video Perception Models for 3D Scene Synthesis cites this paper.

Video Perception Models for 3D Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:59.729438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:59.729438Z digest=sha256:888b7fda06f1123b62128b83ca424736d943bad7c4cf27c40a33c12a0da400c9

Observation a4081951-88c8-47fd-88d2-e67f97dafcec · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:42.305835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:42.305835Z digest=sha256:611ed99cab90791b5b0f66ed7a94e02ebb1ff7b2be17a326f228f6263b730a5c

Observation c0daa79a-4e47-481e-a554-80e5d332ed29 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.368585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.368585Z digest=sha256:7152ec8d4232b2ca212b68c475bc09ae1523f0960c56489af2bbb283d6f860a2

Observation 85383ed5-aad4-4b7c-b011-5beca377713b · inbound

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation cites this paper.

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T19:19:20.992326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:19:20.992326Z digest=sha256:135e93915724f33eea4bb4dc073351dff436e3463d016c0afefb2330931510a5

Observation 36e4d326-5360-4fd9-9be8-22eb11c93a07 · inbound

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs cites this paper.

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.648939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T15:07:54.930273Z digest=sha256:89d17bd0ef8673d9247cdde2f5791397952e7f29522953e7c9419056bfe1d745

Observation 26543a02-6378-46d7-9a85-33f1673708c1 · inbound

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning cites this paper.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.560832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:9dcbfc7040688228aa9f5a247795af46560d063a165121796e7bdca66d6a346a

Observation e1a06504-1b43-4f3b-9a11-30a8982adfca · inbound

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models cites this paper.

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:05.932363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:34:15.916685Z digest=sha256:4dd527ede00c280830d2c44909e081d1cd4ffd40f8b09a4952acc52ef6a20a1e

Observation e693439b-16c8-4d60-83b0-4fc46678f9d3 · inbound

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis cites this paper.

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:11:03.657154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:36:33.149921Z digest=sha256:a37e05c12337d545fe39291d06edc13272e1f4d87db0c70dbe260fc3232e8a0b

Observation 3d6782b4-b96a-4ba2-bdf7-ad05b5150363 · inbound

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development cites this paper.

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.232494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:34:58.288245Z digest=sha256:3d7f206b30b677c740d4a4ce3b21278f32164cb62c124cc0181dd19210ff7c7e

Observation 057d08e1-e449-482e-bb80-fe41c9e25c24 · inbound

Repurposing 3D Generative Model for Autoregressive Layout Generation cites this paper.

Repurposing 3D Generative Model for Autoregressive Layout Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.732757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:09:00.779456Z digest=sha256:db7c238a0def16b9df89e3b15be5ee1707461eae05ab9e1f91e94a9965717a98

Observation f5417681-3361-451a-9452-ae11b24206a1 · inbound

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion cites this paper.

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:36.978937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T09:22:10.651922Z digest=sha256:e94aaf570d7e8f75dfe7419d1700397f0d028e001195e256b73c8a2cd354505d

Observation 05f27cc3-095e-4f35-800e-ad8400f5a2e8 · inbound

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation cites this paper.

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:06.580065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T03:22:30.012610Z digest=sha256:6d2b3f5b021dd91c5ea53277da339f2ddc6f27a043ba90de10834db1ab8a3140

Observation 4112b671-dff6-4520-b981-7c75a03eb637 · inbound

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models cites this paper.

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:59:36.508177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T04:55:40.370796Z digest=sha256:bcd02a510f3c3c4c9cf408a0a0a4a26be9c0073e13bbbe178d53ed92fded682d

Observation d46764da-2d11-4041-85f1-8298ebb0421b · inbound

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation cites this paper.

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.323641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T23:11:59.282627Z digest=sha256:c93941f585afddce8a1fd373c9969f5824a070fa17dcc92a0248ac19495fc657

Observation 975ec25c-94f1-4a7a-a71b-d301f9dc1282 · inbound

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image cites this paper.

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.366072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T10:59:17.553098Z digest=sha256:f3f36c963afd2e61fcd16e2c8aa7c97fc10593cf837b99fb7351bfe0d0d6323f

Observation 8ffe7006-363f-4c1e-ba7a-92183b4b77d6 · inbound

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration cites this paper.

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:26.723110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T19:08:08.329817Z digest=sha256:3c8644bf80b002986669cfcaf14b0f3b8d3deab75ce5fe1cbd7e3fb853d58892

Observation ee327f72-b816-4837-b399-d85f8f7a3d3a · inbound

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction cites this paper.

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:09:37.258248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T14:50:16.317301Z digest=sha256:71d465315c06fa88795336e1719902e03dcc3046c833edabbf0827c6f3f2601f

Observation d0407fcf-3f38-47fc-a0ff-d6f490c6f095 · inbound

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation cites this paper.

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:14:25.460428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T08:12:30.693936Z digest=sha256:129867c2b651f3d73aaaa18bfadad85c8ddf344f7cd19fdb76463adb6b9dcd64

Observation 117dbc72-dd2b-4e11-9aad-123f302ebc8f · inbound

One Video, One World: Turning Monocular Video into Physical 4D Scenes cites this paper.

One Video, One World: Turning Monocular Video into Physical 4D Scenes Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:45.028897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T05:38:19.541391Z digest=sha256:021e3e31052f692febeb45cedb959803a6606e652119654078b72306e869f9e9

Observation 9b69de31-9219-4245-be32-20415ad32f18 · inbound

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image cites this paper.

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T22:29:10.666143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:29:10.666143Z digest=sha256:9ea082e58c3725fa9f1fb7a9161d9d0af095d0c8fa8c06d7200e139bb7e5b905

Observation 0973aaf3-5216-40b5-b911-1f59fe8edfde · inbound

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis cites this paper.

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:57:31.940710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:57:31.940710Z digest=sha256:c76b284ba1d1d0374777c684cf65b1fece179d7159fb9c9ae1ae68639fac4973

Observation d469ef02-472c-40dc-9508-54510118a5f4 · inbound

Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation cites this paper.

Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:56.584297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:56.584297Z digest=sha256:6c76ebc0537ded2f4c9e9579e6163e59a0caac24b83aff37b3f6de72b4f2bc4d

Observation 59fb88cd-8d6b-4934-a980-eb680a56c41f · inbound

WorldClaw: Agentic 3D Open-World Generation at Scale cites this paper.

WorldClaw: Agentic 3D Open-World Generation at Scale Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:37.518119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:37.518119Z digest=sha256:4fb4a37aa6452fdd4a1f53121b2680831469d17bc510ecbb40d869a8f987328e