Pith. sign in

Paper Citation Record · LEDGER

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

As of 9 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2511.19119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.19119 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:38:58.579422Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6171e5a5-9779-4182-b1c1-21820f9db851 · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.657295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.657295Z digest=sha256:ee1cf04b1bef66c8ae7a469bec557c5b686b17f585abe0c052550bd74af9c8b2

Observation 94e10a9f-6b82-451b-a339-8de016a6fb1e · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.706106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.706106Z digest=sha256:cb1d2578aa5cdbbcdc09ab794548e1af8d33792c0b522e04711ac1eead69da64

Observation 12207c0e-8757-48fe-8568-fd461fa8b6be · outbound

This paper cites Qwen2.5-VL Technical Report.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.781179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.781179Z digest=sha256:6a4a5576a2cb8fbad8462c413e495134980caed42deba5f7da1b1a95d4e1ea5e

Observation 710d92b2-ac15-4abf-bc0b-12b491aa1962 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.831472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.831472Z digest=sha256:c492f4cfb41eaa66de9e9ac5e002a27f49654a527a22074f3c2850dacd04aac9

Observation 0fd9d04f-37b2-42d8-a1e5-62fe5678f162 · outbound

This paper cites Omni3d: A large benchmark and model for 3d object detection in the wild.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Omni3d: A large benchmark and model for 3d object detection in the wild

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.864656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.864656Z digest=sha256:ff4546f863b1e105f14b0c4f8a4394ed01e60e59752d32f59b95f790e3ae436a

Observation 89fc6ba7-bae3-419e-9de8-09a4738a8d4a · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.925300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.925300Z digest=sha256:5460d6340ff9be99541ded319e1ab34b12018b9854b3b32f1db1c62349ee9be3

Observation 1c086a39-011b-49a0-9149-adc2640e85c3 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.954815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.954815Z digest=sha256:94253fe59f9351ac26606a1a647c446f6ffd617b0b681a73b778958a883145b3

Observation 6c6b597c-a254-49e1-ab03-f7a18a027b21 · outbound

This paper cites Perception before reasoning: Two-stage reinforce- ment learning for visual reasoning in vision-language mod- els.arXiv preprint arXiv:2509.13031, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Perception before reasoning: Two-stage reinforce- ment learning for visual reasoning in vision-language mod- els.arXiv preprint arXiv:2509.13031, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.974631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.974631Z digest=sha256:41ce161a31ee314b04fc67a03b08a994025fb4b5b83466c73821083de939df4d

Observation a8942ee0-7dbf-45b9-ac42-d7d87e28b24f · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision-language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial- rgpt: Grounded spatial reasoning in vision-language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.015330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.015330Z digest=sha256:adc1b0143256d0b9786d43b9b8c52eadb3c1ff573f7d1987a3bf56f773e6578b

Observation 508972a0-ed4b-495d-ad5b-d2c5e2b04dab · outbound

This paper cites Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.028226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.028226Z digest=sha256:5cc83b0381273a4e2edab53b8473411d016b43986650fb6b529cdec130974f80

Observation e73763e9-b8bb-4338-ad19-4d2eadd1c8f7 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.094974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.094974Z digest=sha256:a276f5417b607fb96ddb2f83df238e876872dc7a47d0f8138c67dd5235eb38ce

Observation 72fc96a9-1852-4a86-a0bd-ed1dcd958dde · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.151754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.151754Z digest=sha256:be38b4d794a223dec8c5ece87204c9c2a029a98687e0427d6bb4272b2364ffdd

Observation eee8f01c-22c9-4336-b728-07c975bf8000 · outbound

This paper cites Surds: Benchmarking spatial understand- ing and reasoning in driving scenarios with vision language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Surds: Benchmarking spatial understand- ing and reasoning in driving scenarios with vision language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.184590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.184590Z digest=sha256:d2b4771e2b294e023be03d4cce1d5870cc1808dd2475d24c51ead884de82e3d2

Observation ab6ab7ac-37fd-4510-8642-cfb12d71f050 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3d-llm: Inject- ing the 3d world into large language models.NeurIPS, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.239896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.239896Z digest=sha256:3f64533a9a4a61bdc2dbc07f4013ca2f3ef56e0b3ace7dff670432e4e173bc7f

Observation 909bc398-84c5-4dd4-b423-b1ccc8b10f08 · outbound

This paper cites What’s ”up” with vision-language models? investigating their strug- gle with spatial reasoning.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images What’s ”up” with vision-language models? investigating their strug- gle with spatial reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.287693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.287693Z digest=sha256:770a2a3ce3b891cb8495d2751dd869f8f002f44eb43830080c800db39a704395

Observation 346095bd-61b2-48d3-81e5-e9db8400d8db · outbound

This paper cites Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, and Minhyuk Sung

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.318595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.318595Z digest=sha256:fbaf85c5d7653aa9f0e776e05a371a9a20f5a500e02cda4a19d4debbb3eea4c2

Observation 532c478c-b7da-43c5-85d3-6ad7c014b912 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Seed-bench: Bench- marking multimodal large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.361460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.361460Z digest=sha256:399ad262851f15abab57f78b874fbe81d8d1c842ba42ab8660041a7f7aa5493b

Observation cee1622d-630d-4358-9580-72b51b92f0c4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.411826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.411826Z digest=sha256:6dd3cb0d4e1f0add5dbdcca2b6728c52e79f36d6bef492ecaf6382ad9c8c623c

Observation 95c1dadc-6fda-4c81-8d38-d9f20f4b0ee3 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.448060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.448060Z digest=sha256:d7bc567780a615680b01e16d11843a621194fb4a9eb1ce0337a944a36b44caf5

Observation cd0f1fc6-39d2-4a64-b2c2-0c78d649dc57 · outbound

This paper cites Visual Instruction Tuning.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.503949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.503949Z digest=sha256:24aadb384f3ab5b11ede40d6c130c49f6b6bd937286156de3395385a228b901d

Observation 8761fb1f-1515-494e-9010-dd67e17c3095 · outbound

This paper cites Spatialladder: Progressive train- ing for spatial reasoning in vision-language models, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatialladder: Progressive train- ing for spatial reasoning in vision-language models, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.545210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.545210Z digest=sha256:5e17e704bb9beeb7ef5113747a59c7ffb76accc283dd3202c5cbf8f81c235e00

Observation f54e9786-4ca9-4ced-b037-d896d0b231c2 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.596804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.596804Z digest=sha256:dc78f01508861b47becedbaf28b9e1a43447b96bfa67b98d5639929d4ad205d7

Observation a2d028f0-fe31-4c10-8379-6933e4834c96 · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.643951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.643951Z digest=sha256:82ff18b9c6ed7ac843582f8686542d209a7cf9cc0b36d132c2032d21083abf8d

Observation 65d6b3ed-76fc-44a4-bc01-4fb57d4847a8 · outbound

This paper cites A Novel Multi-Agent Deep RL Approach for Traffic Signal Control.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A Novel Multi-Agent Deep RL Approach for Traffic Signal Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.718590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.718590Z digest=sha256:d1b31d3a17430dcc370bab49a41dbfe7eddfc48f770d401a447857c02ad055b9

Observation 7def5476-984c-4f25-ab84-09af786d2853 · outbound

This paper cites Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.770035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.770035Z digest=sha256:49f247342a7773c36de957753c40f4196fd6a0c652e1f283a0883b91e9fb0621

Observation e9a6c7f1-deac-4565-833b-e8f48d1b5405 · outbound

This paper cites Grounding dino: Mar- rying dino with grounded pre-training for open-set object 10 detection.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Grounding dino: Mar- rying dino with grounded pre-training for open-set object 10 detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.855026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.855026Z digest=sha256:486863f7f6de93bc36026c718a5f6c258c78ba52778b70a803e205d77d68e3d0

Observation 06e758fc-b272-4fee-b172-b22316866e94 · outbound

This paper cites 3dsrbench: A compre- hensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825, 2024.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images 3dsrbench: A compre- hensive 3d spatial reasoning benchmark.arXiv preprint arXiv:2412.07825, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.901480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.901480Z digest=sha256:d8a6528ae40703e20c0c084c6f156c6579ebca476f6c391df75c0dc9a31991c9

Observation 809fd4d9-e63b-4352-8c61-e370e45e891c · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Sqa3d: Situated question answering in 3d scenes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.967821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.967821Z digest=sha256:dec4aae0d6bcd07900091d595cef5979362f9fa47b82d593834ca5f99eb12b1f

Observation f5c72476-ad3f-498d-b895-80649dd2dc72 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.020436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.020436Z digest=sha256:4dca3de177b65ed8f758237388f3c3b97f2ca6e5c1a54564d041c708a516f6b6

Observation 21436383-2142-4cfc-a7ad-c80fe38dfc3b · outbound

This paper cites GPT-4 Technical Report.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.091709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.091709Z digest=sha256:70a292b468ed952cb11123257d7d5ba8dd5686a7efcf2eb4958b936dd83b5800

Observation d48f1012-52dd-4854-80a4-4ad5e0383e39 · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Shapellm: Universal 3d object understanding for embodied interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.095771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.095771Z digest=sha256:0a931af55c0c6d3843faeb0aceaa2bafb4912421a3acb5a99b08f05be4bbc743

Observation 50e00068-8775-49a4-9a94-2dbaf0723a35 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.102240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.102240Z digest=sha256:a97c6a0ccc0ac71bec3df3d7bdf574a45ff835c69d4b110653466deaccb01068

Observation 1e11018a-2d3c-4ed3-8394-c1f711d24bc9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.157395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.157395Z digest=sha256:3f54ff694f10dc54f31e9bfcd1bb779589ac66f93bb98e45541137685b3a7f21

Observation ec6506ca-7cad-48d3-9bda-080642b9689e · outbound

This paper cites Space3D-Bench: Spatial 3D Question Answering Benchmark.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Space3D-Bench: Spatial 3D Question Answering Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.248164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.248164Z digest=sha256:d76969dbe00978194b59d8ef94e9e8b48d95ee83656b8510b30e3119a7649b5f

Observation b37d4bd3-ca58-4b9d-8dc0-095031994996 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.274268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.274268Z digest=sha256:174f5030ce68b878483e3745d9f80ab79342b9064015507bf6fd77732bb38035

Observation 0ed634ff-f6e0-4037-8e01-a0acee736260 · outbound

This paper cites NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.380114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.380114Z digest=sha256:e1b909d9f3d6c7d9bfd85b6a483ef5c996ac36f3978c17a4aeac1b56d03292c7

Observation e059e4b6-49a5-44ef-bc97-727183fc8bdf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images LLaMA: Open and Efficient Foundation Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.472453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.472453Z digest=sha256:845c3ed503880c170e50a209880bf30695f8db64c4174af7dc1b2ddfcbcf384a

Observation df08ed63-aad8-4ade-8dee-f858468399a6 · outbound

This paper cites Cross-modal pro- jection in multimodal llms doesn’t really project visual at- tributes to textual space.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Cross-modal pro- jection in multimodal llms doesn’t really project visual at- tributes to textual space

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.591800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.591800Z digest=sha256:7a0fea1ae96b48f26fc6e7767da5aafc9e5f7e2b89594f8d9adf73182c6d3847

Observation 2e583bfd-8301-48ac-8aa6-092681b606bb · outbound

This paper cites Learning 3d semantic scene graphs from 3d indoor reconstructions.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Learning 3d semantic scene graphs from 3d indoor reconstructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.673894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.673894Z digest=sha256:f8f30cc406a867de172cd9b70e01c9a5af99a81fc2958a0f28e5872bbbd59378

Observation a4a5cb51-d025-44e9-b5f3-a00c59a9361b · outbound

This paper cites Is a picture worth a thou- sand words? delving into spatial reasoning for vision lan- guage models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Is a picture worth a thou- sand words? delving into spatial reasoning for vision lan- guage models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.735634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.735634Z digest=sha256:beeca58191a5f0cb5ce2ec557ac72e0d902a4adb91e16dc92ca54552fadc433e

Observation 7acbe468-8842-49d1-99d4-ca4dfe6618ae · outbound

This paper cites Spatial 3d-llm: Exploring spatial awareness in 3d vision-language models,.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial 3d-llm: Exploring spatial awareness in 3d vision-language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.777803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.777803Z digest=sha256:f4a86165f1f83a0322adb51f50ac42b0bbfbbf02aa71214daa9b4b8cf6e13d18

Observation 04f4922f-24b2-44f6-bfe3-9389f52a9375 · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.844553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.844553Z digest=sha256:a0193ec018015f5683bb7457b11737768442c187a4f0a1052402a188c4d79c33

Observation 0373d929-38c4-4c28-a8b6-d3554f1e5416 · outbound

This paper cites Pointllm: Empowering large lan- guage models to understand point clouds.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Pointllm: Empowering large lan- guage models to understand point clouds

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.910137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.910137Z digest=sha256:a1bf4beb5d19be87dc43969204fa9aabf01bdff1a5aa48c1d49517a83031c417

Observation 45090bbc-c81e-4083-83e5-efdd44782115 · outbound

This paper cites Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:56.976381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:56.976381Z digest=sha256:e06f446c2f13daca5bcc6411667a95675c1581de01dd99666949748b89335f0d

Observation 4ec212f5-83fa-4cbb-89f0-582422b9b7eb · outbound

This paper cites Open-vocabulary object detection using cap- tions.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Open-vocabulary object detection using cap- tions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.124463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.124463Z digest=sha256:651818e64414884654ea8b94afad91567d53c3974421a74eff9a62d0442eb518

Observation 09b0e65d-2a0c-49e7-8238-0a0cd844b2c7 · outbound

This paper cites How to enable llm with 3d capacity? a survey of spatial reasoning in llm, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images How to enable llm with 3d capacity? a survey of spatial reasoning in llm, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.258520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.258520Z digest=sha256:21ed2cef00b7a5bc2d30a65ce6494355f017258d7d250eb8146c68584a31e28f

Observation 1d04de49-7abf-40cc-ad8c-ccefd429fd3c · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.314772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.314772Z digest=sha256:3ec2b141cfbf8bfcc62e36d9143d416125c2022f08fa0ab5d289b37b89b920fc

Observation 904bbdd2-6d8e-42bc-9365-6868345a91a9 · outbound

This paper cites Spinbench: Perspective and rotation as a lens on spatial reasoning in vlms, 2025.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Spinbench: Perspective and rotation as a lens on spatial reasoning in vlms, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.422175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.422175Z digest=sha256:9053c87622aa7030dc3f52e3741af68e67e52a1cb7e4cc13e91ee00008f69ae9

Observation 2a6ba5e5-489d-4e13-b4e2-6678b7a104cf · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.553330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.553330Z digest=sha256:bd78164f4d54c967ae0f123f6282b6769df56ff54aa7e5486ccf127ed42d9b88

Observation b94a55ac-cab6-4579-8625-16871717526a · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.596020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.596020Z digest=sha256:478ba9d16c25487c20fd7db12205c472a78777dd794d63d17f4889afb5bb2186

Observation f65d7cbe-43bc-4dbc-a2d9-6e2a7126a51f · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.652883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.652883Z digest=sha256:9175f8ebda44baa1a2238391bf23105242b23f61ed04aa2130b41068fb1247bb

Observation 9d293d6a-d7db-4986-a1e2-6b7bd44fcead · outbound

This paper cites A detailed description of each level and its cor- responding tasks is provided below.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images A detailed description of each level and its cor- responding tasks is provided below

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.761326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.761326Z digest=sha256:8fb2347bd326bdb1c2b4d0820fce16df50ad96fe594ee0bd8af8d7a74d601b60

Observation ff698a18-5d01-4deb-9db5-359f450867ae · outbound

This paper cites Specifically, after obtaining high-quality raw data through filtering, we generate image captions and construct scene graphs to serve as our underlying database.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Specifically, after obtaining high-quality raw data through filtering, we generate image captions and construct scene graphs to serve as our underlying database

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.853194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.853194Z digest=sha256:076459d7282330fe50eb16cdf64b012205f7aa5df101461b17c990643cc821cc

Observation f1e64af2-c666-458b-a4b7-1e450b6963cc · outbound

This paper cites - {variation_instruction} - The rewritten question MUST include: (a) A brief motivation clause describing WHY we need this information, consistent with the motivation hint.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {variation_instruction} - The rewritten question MUST include: (a) A brief motivation clause describing WHY we need this information, consistent with the motivation hint

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.909081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.909081Z digest=sha256:1b153bb3bac1f53ff6c0079e46ae0f72ed9315fb6d90af838233362267752492

Observation 591aa810-2306-491a-b5ac-fbdda1f876b0 · outbound

This paper cites - {answer_constraint} - {task_extra} - You MUST NOT flip yesno.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - {answer_constraint} - {task_extra} - You MUST NOT flip yesno

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:57.975775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:57.975775Z digest=sha256:ccedb1503e67b3d511887b1973daf490c50ef5319d1416b1c736c8e686d5efc6

Observation 39a5c178-0457-4ead-b97f-b774b381fd20 · outbound

This paper cites thinking.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images thinking

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.030558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.030558Z digest=sha256:df54ff3e9fb065b90f56dabad3f8abaf29b5077a7afb926d9b7c56d0e4bfc372

Observation 7ab6194a-8858-4083-bc70-cf1d59d2870e · outbound

This paper cites Scene In- formation adds global context, 2D Visual Prompts improve local grounding, and 3D Bounding Boxes deliver the largest gains through explicit geometric structure.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Scene In- formation adds global context, 2D Visual Prompts improve local grounding, and 3D Bounding Boxes deliver the largest gains through explicit geometric structure

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.090253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.090253Z digest=sha256:8b69d84dc71a3230644ac07f556adc22ac70db58ef3cced90f347cb608f72027

Observation b027e5a3-9197-4b17-93b2-27c3c38441ef · outbound

This paper cites For the 3D bounding box information, each object is rep- resented by its center coordinates, spatial dimensions (size), and orientation expressed as a rotation quaternion.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images For the 3D bounding box information, each object is rep- resented by its center coordinates, spatial dimensions (size), and orientation expressed as a rotation quaternion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.171986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.171986Z digest=sha256:9563f8b6261a977b451aa57f1b22ab4c5bad99a2e69d3ae4610f90b6b4158ecd

Observation d6e8e8b5-5ad3-4d78-8169-69fca60a2b03 · outbound

This paper cites All inputs are processed with the officialQwen2.5-VLprocessor, which supports dy- namic image resolutions up to 262,144 pixels (512×512).

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images All inputs are processed with the officialQwen2.5-VLprocessor, which supports dy- namic image resolutions up to 262,144 pixels (512×512)

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.262429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.262429Z digest=sha256:9083bf2e0a5a9495c974f0724c7e562b8806b7a5a3fe170736b96c0c976fe3d5

Observation 41a9de49-fc95-4dfc-8e96-9391fe0f4693 · outbound

This paper cites - Use the format <obj>...</obj> to describe your mapping between textual entities and object IDs.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images - Use the format <obj>...</obj> to describe your mapping between textual entities and object IDs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.358583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.358583Z digest=sha256:86f0d4af2acd410d1220180d6eda69f1d8475db8b7d08448f3883135d5e991d1

Observation bb5c9776-80e7-4524-8b8f-cb623da6ca9e · outbound

This paper cites an unresolved cited work.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.429529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.429529Z digest=sha256:d18dc1a577f3062d588f3b9861d6f684d7014135edf3fb3103b99a4406637643

Observation 5a1117f2-1d56-48f4-9935-77abe67ff869 · outbound

This paper cites Your output format should strictly follow: <obj> Object mapping: - entity_1 object {id_a} ({caption_a}) - entity_2 object {id_b} </obj> <reasoning>.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images Your output format should strictly follow: <obj> Object mapping: - entity_1 object {id_a} ({caption_a}) - entity_2 object {id_b} </obj> <reasoning>

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:58.579422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:58.579422Z digest=sha256:3c48d976f03190770c9d1684105dd95d25f9e073c8d37f9e35e921c41f180525

Pith citing papers

No inbound Pith citation observations are available.