Pith. sign in

Paper Citation Record · LEDGER

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

As of 20 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 5 inbound Pith citation observations for arXiv:2507.06484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06484 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:09:30.946930Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:11:00.840340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:56:34.752902Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 757fdf72-0fd8-42f4-8226-9a5f310f6c49 · outbound

This paper cites GPT-4 Technical Report.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.060519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.060519Z digest=sha256:c3ebcb471206ea13c2ec8ac2b243f246e942852f6a5e836c41ff091647fb2d75

Observation 56ab8fd7-a601-4246-b55f-bec25da8fa23 · outbound

This paper cites Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.147440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.147440Z digest=sha256:472ef102b6c16db28ca667e63b0e04fd2b4c585d5105eebff711378f14c836f9

Observation 3e3f0ed2-7c48-42d1-a7a2-3999f5c443ca · outbound

This paper cites I-design: Personal- ized llm interior designer.arXiv preprint arXiv:2404.02838,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds I-design: Personal- ized llm interior designer.arXiv preprint arXiv:2404.02838,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.251909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.251909Z digest=sha256:f73765d9a134f27fc4a39a7f287d965e57096bca4c4a762f4d689eb468620de4

Observation 22b496ac-4bef-4dfc-a5d2-e0b4a47a2a15 · outbound

This paper cites Learning spatial knowledge for text to 3d scene generation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Learning spatial knowledge for text to 3d scene generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.332637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:25.319361Z digest=sha256:fdfce7af0a85484d6fce92207dde07e1cf87f42b3d50df6dcdea8a9ed63915df

Observation 40790c0b-4ea2-4831-8d58-c1c6cbee8dec · outbound

This paper cites SceneSeer: 3D Scene Design with Natural Language.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SceneSeer: 3D Scene Design with Natural Language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.395511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.395511Z digest=sha256:583febdfc43d8959a1446c7227fe1f332698c8c92281e17f168ac1ef9ec73731

Observation 82379160-45a2-46f9-b488-2652080775ae · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.194302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:25.490051Z digest=sha256:8070567eff9bd542e1eb1477320bedcf48a91dd96c1457d3b5a46dea205d1b7b

Observation b60dd022-5bcc-4f0e-bfd5-3d99bb6bc0da · outbound

This paper cites URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.571972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.571972Z digest=sha256:6d951dd8cdfc2a4299630cdca05a96ab17406f5f0ce2ae57d80bb5c693881e45

Observation aa4cd893-691c-4236-96ff-78c5eb4d968e · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.638480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.638480Z digest=sha256:457cef577bca74788de72502a993272e651d29fb9482d66547f70a0bb6354827

Observation 5e35302d-176e-4e4c-af07-ec4ccdfc46a8 · outbound

This paper cites Wordseye: An automatic text-to-scene conversion system.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Wordseye: An automatic text-to-scene conversion system

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.160901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:25.731961Z digest=sha256:f38c57c30d1c044da9e9a647a3d6c61dcf07bf2dc525e251bbd3249df1250d7d

Observation 0bb1a632-f669-463d-8509-c1b34df47882 · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.Ad- vances in Neural Information Processing Systems, 35:5982– 5994, 2022.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Procthor: Large-scale embodied ai using procedural generation.Ad- vances in Neural Information Processing Systems, 35:5982– 5994, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.048234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:25.878293Z digest=sha256:2940207c77a7e895de6ba01e176e3ff2fb5135880c7b4be96f069397586e3d6d

Observation 06aed207-0e9d-41d1-96a6-bb47042fd37c · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Objaverse: A universe of annotated 3d objects

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.866121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:25.968577Z digest=sha256:d8885eecf36ab99b77701cd46f8ab3550a9337f61c833ff6d74330f5d51fa1b8

Observation 4f681eb1-6109-42ce-b1ff-eddb7344277b · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:26.052398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:26.052398Z digest=sha256:4375fd5b2f643f7c85704f1da46d6087399cfa91ececcd0f34ae1ce920ab1e24

Observation b692d7ec-d7ff-4038-bd19-782c9f7bbc46 · outbound

This paper cites ImageNet: A large-scale hierarchical im- age database.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds ImageNet: A large-scale hierarchical im- age database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.702757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:26.151933Z digest=sha256:f928a49a4989368fc6a85a2e1b205130d2c2d352e3345a1b5c3c7f33afb6d99a

Observation d4d23b0e-ff5c-4ecd-8eba-2997edfc0767 · outbound

This paper cites Disentangled 3D Scene Generation with Layout Learning.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Disentangled 3D Scene Generation with Layout Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:26.240368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:26.240368Z digest=sha256:5c90beab5685f28028f7c1f696d0100bb146cf93e6c7891759f04625a77895b6

Observation f04a234c-3512-43dd-a05c-32384fc3740b · outbound

This paper cites Diffusion360: Seamless 360 Degree Panoramic Image Generation based on Diffusion Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Diffusion360: Seamless 360 Degree Panoramic Image Generation based on Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:26.363654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:26.363654Z digest=sha256:77c4bf58be3768ffcbfabd45ff69e5327f73772d5d6931ced8106b35f91561e0

Observation 94d2422b-0527-4e9a-acfc-f877a188431a · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.Advances in Neural Information Processing Systems, 36, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Layoutgpt: Compositional visual plan- ning and generation with large language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.644267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:26.459243Z digest=sha256:bb4dce3c6c5e77811c01dc9ee3e9cc0aef97a20f792dfd691d4df7d623b85b6c

Observation 9c076fbb-9d0b-460a-aea4-3cf6dbdf70b7 · outbound

This paper cites Example-based synthesis of 3d object arrangements.ACM Transactions on Graphics (TOG), 31(6):1–11, 2012.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Example-based synthesis of 3d object arrangements.ACM Transactions on Graphics (TOG), 31(6):1–11, 2012

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.565521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:26.586131Z digest=sha256:d5212ddc54686fbd6f86c24645cb797de5e1cc6dee40728b0dc42b234b797b1d

Observation dda20418-a8e1-4c05-b2de-3dbee326ae00 · outbound

This paper cites Any- home: Open-vocabulary generation of structured and tex- tured 3d homes.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Any- home: Open-vocabulary generation of structured and tex- tured 3d homes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.391203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:26.694531Z digest=sha256:ab304779b4bb033f8360ea7c3e34a91be5a27af6e47027bff5a81e9e0469ad02

Observation 12864e82-d745-4f54-8ff6-76f09672ce55 · outbound

This paper cites Rel3D: A minimally contrastive benchmark for grounding spatial relations in 3d.Advances in Neural Information Pro- cessing Systems, 33:10514–10525, 2020.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Rel3D: A minimally contrastive benchmark for grounding spatial relations in 3d.Advances in Neural Information Pro- cessing Systems, 33:10514–10525, 2020

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.207373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:26.771115Z digest=sha256:0404933ba61755019f3de63f04696dcea1ffee3c5bf2910a2a1af81b4edd973f

Observation 8fc49ff0-f424-47f4-8370-fbfdad7a956c · outbound

This paper cites Text2Room: Extracting textured 3D meshes from 2D text-to-image models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Text2Room: Extracting textured 3D meshes from 2D text-to-image models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.922274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:26.844439Z digest=sha256:52282a3d3b1d0b402fa6168f27d4971e699e2de865a7b18fd728e16d9a638b32

Observation 73641011-fd9f-4384-8daa-bb76568dc2b9 · outbound

This paper cites 3D-LLM: In- jecting the 3D world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds 3D-LLM: In- jecting the 3D world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.695667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:26.935072Z digest=sha256:581a45e58ba1fef8cc49934ecbf9d4bcc8ea2f69d148ebbba6af67b0719eb381

Observation 628e9190-c971-4d58-a671-8761e57772cd · outbound

This paper cites Scenecraft: An llm agent for synthesizing 3d scenes as blender code.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Scenecraft: An llm agent for synthesizing 3d scenes as blender code

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.559575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:27.045961Z digest=sha256:ce9e376a6972ecef9ac33abc687f630b39296750d7c65f7c7a3d188dce279c88

Observation 6e0d37c2-cbfe-42a5-8923-7c1e5662c247 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Large Language Models Cannot Self-Correct Reasoning Yet

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.114885Z digest=sha256:f8f54fc6c043873a0cb3289b19f47397442a0fc4928fcfca2e9f81a7402b64bb

Observation 400aedfb-5b36-4ce2-be5b-4d624cbe5ae6 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Training Language Models to Self-Correct via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.212385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.212385Z digest=sha256:400a22be2758607b08cdc3ec9b34d5088faa3b5ac5b51ecd86b9fd0c699c6295

Observation 97f7ff7f-42c6-4291-b3e2-842bfb087700 · outbound

This paper cites InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.266519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.266519Z digest=sha256:6f799031c14d5b817ee8f437893228a0bb59da3a9be4df34ae8e03e6a72f9f5b

Observation 8c6b66b2-27c5-46ea-9dd3-d96b532ec92b · outbound

This paper cites Microsoft coco: Common objects in context.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Microsoft coco: Common objects in context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.354684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.354684Z digest=sha256:f8f29691bb32539157ad0cc6491f0c1c09ea58ac0f30ecaa031e332794e09a93

Observation 649621b1-eddf-4348-9ed5-86f25b3237cf · outbound

This paper cites Material palette: Extraction of materials from a single image.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Material palette: Extraction of materials from a single image

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.404330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:27.417306Z digest=sha256:a9e0d165f5cca2fcee8282982863b8752d304f1373fcd2ad396777ff3cc44966

Observation c6654d33-4ca3-4af9-8118-56fa931bdbab · outbound

This paper cites Language-driven synthe- sis of 3d scenes from scene databases.ACM Transactions on Graphics (TOG), 37(6):1–16, 2018.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Language-driven synthe- sis of 3d scenes from scene databases.ACM Transactions on Graphics (TOG), 37(6):1–16, 2018

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.331089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:27.482068Z digest=sha256:a3dda7797a14b977286e69d05b2b87399b9dd284e1415b7f74edf8c5d25a8aff

Observation ecdc951c-3b1f-4264-8d72-89b6ac9ace65 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.557334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.557334Z digest=sha256:696fe199b178d580dfe02b0f1753b608c6be7cc424eb2797faab234e0fe6f80e

Observation f0a61184-4834-46a2-80c6-f0dc2d834564 · outbound

This paper cites Edify 3D: Scalable High-Quality 3D Asset Generation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Edify 3D: Scalable High-Quality 3D Asset Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.659339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.659339Z digest=sha256:4dadf208b64842a7fbc394c64644df335e97ec5ecce7ae7c3905674ad080e213

Observation d5d1dbc3-e9e4-4b1a-a653-ef2e27248c83 · outbound

This paper cites Atiss: Autoregres- sive transformers for indoor scene synthesis.Advances in Neural Information Processing Systems, 34:12013–12026,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Atiss: Autoregres- sive transformers for indoor scene synthesis.Advances in Neural Information Processing Systems, 34:12013–12026,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.762897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.762897Z digest=sha256:89385eb186c4110df4d544f86e893996512c2aa0fe088437463c40193bd84d72

Observation 4166ae79-89be-490f-a2ca-63425b49735e · outbound

This paper cites Advances in data- driven analysis and synthesis of 3d indoor scenes.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Advances in data- driven analysis and synthesis of 3d indoor scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.204046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:27.847004Z digest=sha256:a46825b5dcd24c474d1754f06c73fa0bee669bcc7bff6fece7ebb4f85d68b8cb

Observation ac336ab7-8668-4344-ac3f-b09e09c91a16 · outbound

This paper cites Compositional 3d scene generation using locally conditioned diffusion.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Compositional 3d scene generation using locally conditioned diffusion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.091543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:27.896624Z digest=sha256:5b68191fb4a67b15e131b969328d9eeda99a969d37b970b22a3bb4062472fd27

Observation 6ed5852a-8429-4350-bd96-029b7ddbe82e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.959566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.959566Z digest=sha256:a89eb6e528ed233a92bbc4442b204e801afb1f2f301aa52901fcc0dfc2d345f8

Observation 45cd910c-6ffe-4d68-8f44-932d3a284a16 · outbound

This paper cites Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:09:31.267876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:28.043773Z digest=sha256:71c94e82b79ac05855851583a32d3efbeca2bcdbce46aa889ff6ee997ef39eaf

Observation 9924f5f2-4431-475e-af85-12b9231f8995 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.095933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.095933Z digest=sha256:8d296887c39964ebf56c0ec0d3176f1057afb5c0ac9ba3bdc9a77247b83dd9f9

Observation 22cfc7e8-f557-498f-ac5c-98c37baee8f6 · outbound

This paper cites Susskind.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Susskind

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.195189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.195189Z digest=sha256:ba39b00439600a60b533575de98c5159521fdc4029760f52dedaed520f4143c3

Observation 68c8327f-2d4c-433f-85f0-231279d08c57 · outbound

This paper cites Object Hallucination in Image Captioning.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Object Hallucination in Image Captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.253128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.253128Z digest=sha256:a615452519bed5b4e8b02e9863991aa755ee2ec9d15b3ebfaedb61e57cea569d

Observation df90daee-1c8d-4296-a659-20304ca21b12 · outbound

This paper cites Controlroom3d: Room gen- eration using semantic proxy rooms.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Controlroom3d: Room gen- eration using semantic proxy rooms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.335751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.335751Z digest=sha256:35923da22c053d4dc831a006ad5700c187fadb603414969dd245bdcbf533f3ce

Observation 378a64ea-5841-49a3-ad39-8a4a146a5199 · outbound

This paper cites Real-time automatic 3d scene generation from natural language voice and text de- scriptions.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Real-time automatic 3d scene generation from natural language voice and text de- scriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.928249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:28.456742Z digest=sha256:d4597e3c13831a17c0ecabe820e3db76b6002b3f4e41f05497da982533e6675c

Observation 5941c165-779d-4786-9f71-77114fbce186 · outbound

This paper cites Horizonnet: Learning room layout with 1d represen- tation and pano stretch data augmentation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Horizonnet: Learning room layout with 1d represen- tation and pano stretch data augmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.777795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:28.580659Z digest=sha256:0a2702f84eb2d9b55d8f1c67c4df4ebc6b8717c4e96dcae0ad2ac9165b8a4c22

Observation 189c25ce-7354-44c2-8df8-900b39caf607 · outbound

This paper cites LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.680768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.680768Z digest=sha256:e412f445bc58bd34cf9ec3ee40515a6070ef5f17bd5cfe7a1a3d0633e9361e56

Observation 31197d64-302c-4235-afef-3475e3b4b745 · outbound

This paper cites Factorsim: Generative simulation via factorized rep- resentation.Advances in Neural Information Processing Sys- tems, 37:87438–87472, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Factorsim: Generative simulation via factorized rep- resentation.Advances in Neural Information Processing Sys- tems, 37:87438–87472, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.665930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:28.768156Z digest=sha256:a1d9c90cc8e69f34702c8e53332fe49da299963be5647ac35c62e2e9a1c680b5

Observation f236d975-a3cd-4f93-b853-7c59bce4e246 · outbound

This paper cites Partial-view object view synthesis via filtering inversion.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Partial-view object view synthesis via filtering inversion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.571398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:28.869353Z digest=sha256:76e15a510cc3b0f0c019931437fc6cdb7712915d0acdc67c80e45cb47e4f7b2f

Observation 61e176e0-4271-4ed5-ac51-1d454af5b741 · outbound

This paper cites Lgm: Large multi-view gaussian model for high-resolution 3d content creation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Lgm: Large multi-view gaussian model for high-resolution 3d content creation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.460132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:28.965667Z digest=sha256:5433937a5e1776f356eb2b7fb2d00c7bf11edad8546629cdf34dacd3153096ad

Observation 0a787edc-b8d6-4976-adc7-9cb6b0b6db5f · outbound

This paper cites DiffuScene: Denoising diffu- sion models for generative indoor scene synthesis.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds DiffuScene: Denoising diffu- sion models for generative indoor scene synthesis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.329825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:29.068290Z digest=sha256:e0b10ceb0edc05c7c70e31fbaccba09e48a9f809b7d002f2153cb94fc93079f7

Observation 5fb754f0-75b4-4bf3-bd21-ee0657327702 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.197304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.197304Z digest=sha256:b17dedaf629ee3ad7b5d0e2d8263301b4d0e4cdba9507323488509e0f4faf602

Observation 3c692faa-d191-487c-b32e-a6db1f62d42f · outbound

This paper cites Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.291799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.291799Z digest=sha256:71863d60f864836473677494c06550f29d34c684fcc8eea6ed03686dbccfb633

Observation 36be4184-03cd-4347-a035-b2a495046f64 · outbound

This paper cites RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.396038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.396038Z digest=sha256:148bcff9740fc77ccacb572776578b5e3c4810163a3f6ecd0670bd84831927e6

Observation faed3a08-26c4-42f0-bddc-03853aa784b0 · outbound

This paper cites Architect: Generating vivid and interactive 3d scenes with hierarchical 2d inpainting, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Architect: Generating vivid and interactive 3d scenes with hierarchical 2d inpainting, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.206158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:29.490104Z digest=sha256:6b943f2242d001a644d0366aee2a07d38f0cb83a44b33dc6a42ad15d23a9bfa5

Observation d4598024-9fcb-44ce-b65b-52ff61dc3d11 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.093312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:29.607181Z digest=sha256:222c0c656b8bf7accdce38e00f7c6ea90bb72a60a35667acc9a99c40a6b9e69e

Observation 685240fa-fe05-4b7f-be0e-61bef0ae514d · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.717119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.717119Z digest=sha256:fc19e0b4eccfffb253b0034a6f5fd95c279a280a212e34d9fbfda1f08e52bdf3

Observation c76dbc7a-d1d5-4e2d-ba4e-b6f1566e51e7 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.817253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.817253Z digest=sha256:31497efe004736bc1d829227a60e55c3403a3db9d9b0e5ef1559a120041d049e

Observation ff5bf40f-6a21-4fa3-8ea0-bf0eaa3b2871 · outbound

This paper cites Physcene: Physically interactable 3d scene synthe- sis for embodied ai.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Physcene: Physically interactable 3d scene synthe- sis for embodied ai

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.914444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:29.929380Z digest=sha256:ae873b71715aa3d105cc86cfbda52d945b4344de8bd5276cd3cecb9b05b48754

Observation 32381829-480e-4e4d-b2bf-515812bb4628 · outbound

This paper cites Holodeck: Language guided gen- eration of 3d embodied ai environments.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Holodeck: Language guided gen- eration of 3d embodied ai environments

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.707862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:30.026747Z digest=sha256:290065cc24febc6a190b483ae825a6ba351ca6625368d721c58f1f4607a88e89

Observation d65d69dd-45a4-4cc0-889a-bf0928bf6c8f · outbound

This paper cites The clutterpalette: An interactive tool for detailing indoor scenes.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds The clutterpalette: An interactive tool for detailing indoor scenes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.521254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:30.106167Z digest=sha256:9b13e6292d2fdafc45e2401ac104df14c20d556cc4f33097b90c3a529a514f12

Observation e7f87d50-a65e-4b5d-aeff-8e8484634865 · outbound

This paper cites RLF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional hu- man feedback.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds RLF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional hu- man feedback

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.329272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:30.186377Z digest=sha256:7c782951a1571d0f6ebc159008355edc0614d2df1d797750e5c173163dd386ab

Observation e4dc9c68-d7aa-4626-8768-134154e5c463 · outbound

This paper cites Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets.ACM Transactions on Graphics (TOG), 43(4):1–20, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets.ACM Transactions on Graphics (TOG), 43(4):1–20, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:30.324696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:30.324696Z digest=sha256:ec072bd36aff1af5c9df03e0b9d95d113904aec35744e8441b405f1e3fdc9f02

Observation 7f84a514-1eeb-472b-84de-235ba23992e0 · outbound

This paper cites SceneWiz3D: Towards Text-guided 3D Scene Composition.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SceneWiz3D: Towards Text-guided 3D Scene Composition

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:30.414347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:30.414347Z digest=sha256:ab87e10d9b90b4a23f5884be2cdcbf1a5bb6a5d18e68cd1a941f39c2154442ac

Observation a2e42970-38ae-4611-a5ed-21dbd4680c14 · outbound

This paper cites DreamScene360: Uncon- strained text-to-3D scene generation with panoramic gaus- sian splatting.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds DreamScene360: Uncon- strained text-to-3D scene generation with panoramic gaus- sian splatting

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.122004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:30.558770Z digest=sha256:4542fa4769f1cba90a8b6645527d81ae5a2077a3f269bdf95654a5fdd1a7f21b

Observation c99a6018-6cf9-4342-94d3-ed9242d464bf · outbound

This paper cites GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:30.649221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:30.649221Z digest=sha256:b8f72e6b4a999c682d5166721d11d993366dd38dbf082a9e6ccb52c5f97fea20

Observation 4c5e14bb-7747-43ea-9a01-5b16d95a5721 · outbound

This paper cites placements.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds placements

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:31.918596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:30.753677Z digest=sha256:8b40ecf47c00e9d75a089f50c673ee60c760d12b6ca6ad3c80ae5f65c80820ba

Observation 68d92fdb-d18e-4240-bd51-c81cf5b9e13a · outbound

This paper cites an unresolved cited work.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:09:31.791694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:30.834018Z digest=sha256:ae865f09b28aabc27172ee3206fdd46c07c684a4be075b9ada65eeb5c58ce650

Observation e54cadc9-f2ba-4df3-a054-1a519a52ea5c · outbound

This paper cites receptacle objects.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds receptacle objects

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:31.585806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:09:30.946930Z digest=sha256:902d54fe90b4c2f9fb1d7571c2a852c065b90cfe851e2de5cbee1d6b928b8a30

Pith citing papers

Observation f4882744-93ec-4c58-b0f6-95ce98fe92b5 · inbound

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning cites this paper.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:51:00.591108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:4df6cf0f0da59e07a07f8beaef4b1f3d88ec1d7cfa9f2873a31eeabc8a490d01

Observation 7b8f4dcd-56b7-4aaa-859a-2ba50a53d977 · inbound

StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics cites this paper.

StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:28:26.411745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:24:26.557732Z digest=sha256:523e89d8e216b3089f1adfd2d45178a54f639cffea5652c42001d36b49ba598a

Observation ce742356-f2b1-44fe-a795-630313bf90dc · inbound

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis cites this paper.

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:53:14.996405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T11:49:53.113415Z digest=sha256:7ad82fe779d1c9f046d999c94c0db75e351675cd7ae095c483c50b234ca3b711

Observation adb0c979-3899-46e7-aaa7-e10b964af779 · inbound

Function2Scene: 3D Indoor Scene Layout from Functional Specifications cites this paper.

Function2Scene: 3D Indoor Scene Layout from Functional Specifications 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:12:46.730083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:11:00.840340Z digest=sha256:36abae78c25b21533224de034b91f8b8da15dd32c7fbc53daa7e20b3363fd3bc

Observation 667ff8d3-a95f-4923-85bc-f27ea0d09903 · inbound

PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification cites this paper.

PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:34.754685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T09:34:34.220944Z digest=sha256:99e9e817c3527cd563c92e9d9c30de32ffa50ec18ce42bce2a89ec2f2f8020d4