Pith. sign in

Paper Citation Record · LEDGER

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

As of 13 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2506.23120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23120 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:53:04.127385Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:13:52.383043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 325774cd-e2b0-4437-9db1-1771425d731c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.793154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.793154Z digest=sha256:9635797f683d2273f35ef151e4ecd80fadfb77151c970d686cb5fff3dd2d970f

Observation 6a7764da-d893-49aa-80d6-879fc4821c98 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.729965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:57.848548Z digest=sha256:229be1ee09535444ed9ed57e93bb7389e40847245cce0c0522cb69fb198518ee

Observation ec84b564-9542-4b2a-b53d-af68d50282d6 · outbound

This paper cites Semantic scene com- pletion via integrating instances and scene in-the-loop.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Semantic scene com- pletion via integrating instances and scene in-the-loop

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.576806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:57.941641Z digest=sha256:1597cf2fa16e50ca27a3bdcfc37ecc8b9a739b8f9740b2ecb23e0d9a63c40feb

Observation eb38820f-994a-4424-a259-5980f3e22e0d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.406809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:58.068312Z digest=sha256:f2ddeba527a3e4f110fcb2943f8889c384ba6bcd3c4cf6bd0e4ee572829484a1

Observation 7525dbd6-533c-4a95-babf-e4d55a965115 · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Language conditioned spatial relation reasoning for 3d object grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.238738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:58.206308Z digest=sha256:798e66f9a0b1aa55e58ca242584bef936cf63f634995fb8c0ddeb62d98d9f4f4

Observation 5b1a3463-d775-4faf-b2e6-b2f6a048ebe2 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.061588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:58.293363Z digest=sha256:49e4484b44579ccc02f63688483b1fee2ada114117902dcc1799f740f767794f

Observation becd4c34-3a8b-4b1e-9027-ac8dbaaa3d1b · outbound

This paper cites Mppnet: Multi-frame feature intertwining with proxy points for 3d temporal ob- ject detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Mppnet: Multi-frame feature intertwining with proxy points for 3d temporal ob- ject detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.902145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:58.381758Z digest=sha256:e9754f3ea65e1b1262ec843224c53f61a99a8afe94421caafa701513e752fe06

Observation fd2f6363-7d1b-458d-a749-37362235b4e2 · outbound

This paper cites Trajectoryformer: 3d object tracking transformer with predictive trajectory hypotheses.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Trajectoryformer: 3d object tracking transformer with predictive trajectory hypotheses

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.687509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:58.480676Z digest=sha256:8d6acdc84e6005e2a3ecb72fa16a871c8371acebf52ea1dd839e6912194f98f7

Observation 1c299235-7d13-492f-b05f-0ee47acad1ef · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Grounded 3D-LLM with Referent Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.607061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.607061Z digest=sha256:184668d70becaf93bf5faed2e88bd9325ab9597b6e82bac4ba6152412133ce65

Observation a1c83751-31cb-426a-ac9f-b69ab4ad7c73 · outbound

This paper cites Reslt: Residual learning for long-tailed recog- nition.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Reslt: Residual learning for long-tailed recog- nition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.513529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:58.701391Z digest=sha256:d415f1710854b7270546d79ab6e113d593cee6fd907200b4e99aa666bb6ae853

Observation 458a1831-5f2a-48fc-803e-39cdd608875a · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.811072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.811072Z digest=sha256:975023d5bb1459fb0069f5e4606d278dbef403908947c7e554799b162b34d3e9

Observation 94ac9a92-49d9-42d3-8d0f-78cc8cc7c625 · outbound

This paper cites MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:04.929386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:58.917530Z digest=sha256:9d15b28fe8d9eccf93a4d34c0f4b25d3588a249343cfe6199f515f546129dbe7

Observation e112cc66-d919-4cc3-b7ea-c1c26b72a5df · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.052695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.052695Z digest=sha256:fddcbb3227ce33c5646f2e0420458f901dad58783b75fdc868db6eaf58c01198

Observation 90af0a97-5bb2-413d-9280-9877aa54cf5e · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Pla: Language-driven open- vocabulary 3d scene understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.155141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.155141Z digest=sha256:3363de74eaf7d42c4b43f95e6eb1407d7e1c59b17ef6a5b8415c95872b6f2687

Observation de2e4887-707f-4e28-96b1-bdf9f352fea9 · outbound

This paper cites The llama 3 herd of models, 2024.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation The llama 3 herd of models, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.360728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:59.243701Z digest=sha256:9e91f374ca0cbd3e0bc19086bbd937fd2a04f6f5c11539b62c0712a420ba54aa

Observation 44a72786-e440-4241-b3d0-a9436f35572d · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.369243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.369243Z digest=sha256:94640afe9cc63cc32b0de6aa4826610c8bb285af676a2161617744be937d3ec6

Observation ad57e631-7596-4529-8f48-cc1ce6ee5a4d · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation ImageBind-LLM: Multi-modality Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.473536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.473536Z digest=sha256:84392bf4c4e910fc60132f74ed7f61f182e26c333a1d1a44ffad2d30ec1666f6

Observation 97ce80bc-d7d0-4338-a09c-467d0f2c3547 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation 3d-llm: Inject- ing the 3d world into large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.146921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:59.578599Z digest=sha256:88ffbcfebe52e7d69d8f0d308254d38ffda438d4995232a7a706842d3a1dbea0

Observation 0f9edb34-9b72-4b97-8b5d-91ae6ec18f30 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.681855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.681855Z digest=sha256:b1dc9f259b947ade7ca1b7e40e00d3f1945923ff6a2e41e0c81f769b64fe4c08

Observation 34d95fbc-e937-4715-9210-663522ad7705 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation An Embodied Generalist Agent in 3D World

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.769267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.769267Z digest=sha256:e2bfdd2c1764285be4282e218c051a9d8a36102740b844bbf8a7b825a63a0e72

Observation 4b4618bd-cdea-4d2c-828d-1afa9f1c645b · outbound

This paper cites Multi- view transformer for 3d visual grounding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Multi- view transformer for 3d visual grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.946706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:59.867853Z digest=sha256:c9450eed35ad1d709bec891c01e368dd299b257ccb832f53864b78be309d279e

Observation 62e5aba4-dde3-4665-b802-8ea3177c3189 · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:04.545494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:52:59.994335Z digest=sha256:f68017abc039d2c83e3f68e5962078cb3864c15d5d09cf9e954805ae1a94d248

Observation ef4f5f50-f464-48c0-82e9-bf42cd0c62ed · outbound

This paper cites Guided point contrastive learn- ing for semi-supervised point cloud semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Guided point contrastive learn- ing for semi-supervised point cloud semantic segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.761403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:00.083655Z digest=sha256:97cfd751df66f6a157bbc1836c7a8a0aaadf4dc4805721866d124966326e7781

Observation 831e018b-d7b8-4946-89c5-d2f9cd460c8e · outbound

This paper cites Segment Anything.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Segment Anything

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.204418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.204418Z digest=sha256:c2729e6636e79d7e653fb2a9e91a966c418a117dd1138a1fcb2ddc7c4b5f066f

Observation 7188b26a-7abb-4603-90da-ee7f35813e5a · outbound

This paper cites Oneformer3d: One transformer for unified point cloud segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Oneformer3d: One transformer for unified point cloud segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.583338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:00.306050Z digest=sha256:e4ee385f484d2109582d66d491291e53dddcce86289b23f89d7c23fbd71515d7

Observation 76ace732-993f-481d-af74-7ab8261c1a13 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.403341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:00.437200Z digest=sha256:d915f009ca8358079e570548135d2bca31143b982449b6bfaf347164022e0dc6

Observation d9653dec-3380-4c78-8316-8b21d5210f57 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.558586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.558586Z digest=sha256:16b9bd2d1250e352744ce78c2a8d47de532a6b0b6927defa71f59a694033461e

Observation b4becde9-e5fd-46d0-80c8-0fe3f452cab6 · outbound

This paper cites Large-scale point cloud semantic segmentation with superpoint graphs.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Large-scale point cloud semantic segmentation with superpoint graphs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.158422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:00.644188Z digest=sha256:c948b42feba5ed6de9e92e6712ab9b969d5ea6a529931e0ee9d50ed2c498836a

Observation dc14eb7e-771a-4c7f-b789-4daefd08ace7 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.967553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:00.717881Z digest=sha256:9bc5d6b581ab1d99a6949c9ba5e6c541720dc4db02ff1b2b71f1206b3045bb7c

Observation 3ea31733-91bf-48ee-8dce-ef838f65b9e2 · outbound

This paper cites Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.799961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.799961Z digest=sha256:5e2714ba114701c44ec6514d343070075aa1964ea8dc64666177299900fc4482

Observation c4d0f4aa-fc8b-469b-bf92-4d7ef8153b75 · outbound

This paper cites Decoupled Weight Decay Regularization.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Decoupled Weight Decay Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.901821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.901821Z digest=sha256:f953b753e571613269055c19934c5829e072cc48c08f8b66ca081a75ca72b199

Observation 27d46f3b-d777-4076-b839-5e64e9a65f93 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SQA3D: Situated Question Answering in 3D Scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.971816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.971816Z digest=sha256:97d41718976888100b0ba479f5a5e8904107a1497f73cbdb1d3bc43f7525a5fe

Observation 146b51d3-c45f-413a-90dd-24a29d730431 · outbound

This paper cites An end-to-end transformer model for 3d object detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation An end-to-end transformer model for 3d object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.753791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.046044Z digest=sha256:d0803334e69ab0356183f2adda9316f6e4d47f017e3799fd2c38de3706104215

Observation dd071416-a327-43c7-a1db-a6996a849c5f · outbound

This paper cites Boosting few-shot 3d point cloud segmentation via query-guided enhancement.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Boosting few-shot 3d point cloud segmentation via query-guided enhancement

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.575644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.149234Z digest=sha256:dc570ca18afae7b28a0f60349bbed8bc4868ce92d402e9821298929e76e7b0ce

Observation 50ced8bf-31e5-4b03-8fc6-e2c65b4f4d2d · outbound

This paper cites Hierarchi- cal dense correlation distillation for few-shot segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Hierarchi- cal dense correlation distillation for few-shot segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.421226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.255571Z digest=sha256:c33731c2c4d1254e7f5c43474f79c676783cebcd6a508e44c30506e68c04f814

Observation f7cae70b-0dc2-4b0c-8eb7-ce7ca5f1d669 · outbound

This paper cites Scalable Language Model with Generalized Continual Learning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scalable Language Model with Generalized Continual Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.348484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.348484Z digest=sha256:fc6d2559249c28c0036d7bd8f90061f5f233927b2690e172d8929cf354155446

Observation 584d8855-8324-4fd3-aae2-bdc8b758c3ba · outbound

This paper cites Oa-cnns: Omni- adaptive sparse cnns for 3d semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Oa-cnns: Omni- adaptive sparse cnns for 3d semantic segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.173646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.438300Z digest=sha256:a4e74162cd10ef58148b026c38a810c0a636be3dc31f16a88f056e6642756938

Observation 0d582eb0-37e5-4246-a451-0eb88695eccc · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Openscene: 3d scene understanding with open vocabularies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.538579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.538579Z digest=sha256:e441e9b060a5efc29c16cc3c72e05af2dba4bbb20cc918efb5ed8ac07931b6aa

Observation 56dc775b-f6a2-446e-a8c0-172202b6e50c · outbound

This paper cites Explore the potential of clip for training-free open vocab- ulary semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Explore the potential of clip for training-free open vocab- ulary semantic segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.983916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.634815Z digest=sha256:85c40cff27b1db4b1dfbcabdc6b9cf0d557af380133d12dad857a83991d9ec88

Observation 61f5a869-bb7a-4f88-b675-271d3b0fe13b · outbound

This paper cites Pointr- cnn: 3d object proposal generation and detection from point cloud.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Pointr- cnn: 3d object proposal generation and detection from point cloud

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.816844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.702626Z digest=sha256:d8556f8e2b9ca429cd9b7198f9eb7f8c41e569968415db925a8a448e83f70c0f

Observation 1096ace6-d48f-4595-b0ac-f2ab5e1a36b7 · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.780508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.780508Z digest=sha256:116188795d495b33c3d22ff9dc23ae98489609211936354f1154ac7637c139ff

Observation 78a01afb-911b-4a4b-b494-590f26f7df4c · outbound

This paper cites Mind the interference: Retaining pre-trained knowledge in parameter efficient continual learning of vision-language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Mind the interference: Retaining pre-trained knowledge in parameter efficient continual learning of vision-language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.628744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.855126Z digest=sha256:954a50aa0903545cd866d9d15e323eb8c090d02aabe93ca60423742e33f24d13

Observation 934c49e3-811d-49b1-b6ca-78cb5612802d · outbound

This paper cites Learning shape-aware embedding for scene text detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Learning shape-aware embedding for scene text detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.428096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:01.946814Z digest=sha256:e12d36cd2682126f2204cc0910bf690bea8846778f9bf8a88beb7c3bd4066c4a

Observation 0c26d983-be1a-410b-ae62-43726fbc61fc · outbound

This paper cites Prior guided feature enrich- ment network for few-shot segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Prior guided feature enrich- ment network for few-shot segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.189030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:02.042318Z digest=sha256:576eec5387eebfed8bfad3bcd9e789f29add621d7cdf704c43ed0bc2f14d081c

Observation 8e314eef-7b3c-45b6-b717-432117460287 · outbound

This paper cites Adaptive perspective distillation for semantic segmen- tation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Adaptive perspective distillation for semantic segmen- tation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.980659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:02.106679Z digest=sha256:25937ea6bcd158b4425f560beb18bbd02acaf0f992c64674628cba03bc26c982

Observation 17d04d7b-54fd-4bcc-8dec-367477883c72 · outbound

This paper cites Generalized few-shot semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Generalized few-shot semantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.786638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:02.173893Z digest=sha256:7de3f9c94733b22054faa7ee8d4c799fd77d73ba22b4d0edf9bf6e91f77aa108

Observation 8cc5ac61-fe44-428f-8d39-d6f24a46680b · outbound

This paper cites Learning context-aware classifier for semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Learning context-aware classifier for semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.577862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:02.264287Z digest=sha256:edd9b8d34945064aa022a03787031caad18125c27321f41991537bd30521a66b

Observation 04974178-964d-4fc1-abcf-d397d4456184 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.361077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.361077Z digest=sha256:773d658d06e0c8637c8bc47e72f1a4107496bcda28d9e61f02f140c217890c89

Observation 19030596-2323-44f8-bf0a-ff83d7969880 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.436049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.436049Z digest=sha256:c05206f6063cdd5090aeccdb73db741ea54e55361045f4421646a8ce97d14efc

Observation 8e24a87e-fa6e-47ec-80aa-350393141259 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.513697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.513697Z digest=sha256:cfdf96e379fc38ff9dc41bd6d4ad47308c1ea78f7a1eedd8435d215c2653724e

Observation cd0acc59-f07b-445c-a6af-dd175fe92d31 · outbound

This paper cites Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.433475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:02.621278Z digest=sha256:a121e57b54aa74f925b5e8e8363a74bb722e9920f3d02d99cbab5f2519acee91

Observation 222944e9-9eef-4452-891d-bcb190a9e218 · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.670233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.670233Z digest=sha256:11e43d6b8e4b0b25f001553275ee8d33dabfca332e5a7ce40670f370d311eb3a

Observation 691d61e1-8cbe-46c6-bf68-9bd49f55c352 · outbound

This paper cites Unified language-driven zero-shot domain adaptation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unified language-driven zero-shot domain adaptation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.286190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:02.785477Z digest=sha256:bb993675aba997464283afcfaf2ed2c28e35b60a261d08543d684f3dd38dac15

Observation e364c948-8f20-4ccd-9c55-f72f539d2bca · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Visionzip: Longer is better but not necessary in vision language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.155714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:02.857736Z digest=sha256:ae7be466c5c85397f4350ef87eedc2748b7598508b6a69c1c3b34bc087a7d0eb

Observation 59b157f9-8fa1-4d26-9496-c327aeed8fc7 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.923647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.923647Z digest=sha256:965d05349f93c226ffd8c8a8c23daf4975b2343681e115b3de2f649bea149b65

Observation 5ab31b1a-6074-49f5-a2cf-f4ec5ac7f4da · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.977100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.977100Z digest=sha256:6ee9512d3fd7f77455088e0eb95a5c4d50c52c752c10d4778f00a7bc15c0cdab

Observation 944d160a-1c9a-4f48-a492-c43445372d40 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation OPT: Open Pre-trained Transformer Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.034745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.034745Z digest=sha256:d95ce3f5b608f6ed9075a77d6aa3c91569d4d6bea21479a2fc71daca7a85616a

Observation 1ea536c1-6748-486c-aed5-8131ee72c7df · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Multi3drefer: Grounding text description to multiple 3d objects

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.978931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.093626Z digest=sha256:9364e2ac2a58a96f12a3e6d5b93dd5f44c62c8dcbcd8d80190a7d3d97b096f46

Observation ce0c7d61-939f-4525-b2a1-c6a915a0a248 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.839030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.159558Z digest=sha256:f9de8e6a573762aeaf30a41f80da8960e0e3bfdb54d1154fa6ba19d7d67f6471

Observation 5fd12392-a455-441a-8630-a817b0b433fe · outbound

This paper cites ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.225641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.225641Z digest=sha256:3f98cdd43a5fe1f5894fc60b43ae645951207380906b8363bc31b85a1e7a80a6

Observation 8820c6be-a73e-4ae8-a85b-38ae4e588bf0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.285974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.285974Z digest=sha256:ed386b27e53b2753b4cfe472b9e174ff9d534f173a58d5af0ce47ddcdaa5315b

Observation c264643a-7b76-4de0-9b77-f692ca3febba · outbound

This paper cites in the right of.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation in the right of

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.641534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.364743Z digest=sha256:524ec13f175f43bff0da60e3f87566a9818d499bffdd01c85ff015805a3ae81d

Observation 79c14c01-dd42-438c-8312-73dd76869e8b · outbound

This paper cites Focus on detecting and distinguishing objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Focus on detecting and distinguishing objects

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.491329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.415149Z digest=sha256:bba25e17fb8729fc1024e15ee1fea7c5a453daa57aab60d3dd5b71087e74d03f

Observation 325f6a96-a76f-4951-8e96-bad0ba2a8387 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.344472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.466751Z digest=sha256:509d612b632640fc6c3951b4caac576140fedc127ba08aebc5539a0994cb604e

Observation 05db9656-4a26-484c-add9-192c02bf0113 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.151275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.522108Z digest=sha256:755c204e387912b67e92c42b928a9d5f788dc0e805e1440cdcbb99a77b2acc04

Observation 560ad566-91ec-441d-af03-e42af7682409 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.004162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.581218Z digest=sha256:e738044718dfe7dde5be80e5871fc9ce993c5e653b9f4604cc7e7e9eeab3b375

Observation 61cf7d28-9358-4867-85b6-8c1c07519fab · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.821554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.628077Z digest=sha256:e26282eb402ebb9c68873911297ceb5978904d72047504b81b1f7983e1edd16f

Observation ac8e90e1-e941-4283-8a6f-0a8799e076db · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.673168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.675226Z digest=sha256:9bbbf0408352cb6754f54f65658aaac49de37a967fc908e85cf9c60b7dcafed3

Observation ba3b2e56-4854-4d61-ad78-6809cc68d482 · outbound

This paper cites For multiple objects of the same category, use [<ins_id>][<ins_id>][<ins_id>], etc.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation For multiple objects of the same category, use [<ins_id>][<ins_id>][<ins_id>], etc

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.463313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.739329Z digest=sha256:617512546db9feb9c7d5f491ba260745d8f863787a92b67d46044b748952794b

Observation 90115f2d-f5d1-4533-8061-275982f0466d · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.134037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.785678Z digest=sha256:2251f57dc9605ca53511d28fd3503d1c33dc801631a91a14156bd22dabd334f0

Observation 83c98ba2-e58d-42af-8cde-eca333407c73 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:06.791230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.839144Z digest=sha256:628d0a4bb403a729cfccaa9478513f58a8e70fe63feafdf985811dfebfbc56c3

Observation 684938be-0073-4770-8639-a5573a11fc08 · outbound

This paper cites Don't output anything else.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output anything else

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.518179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.898635Z digest=sha256:3c25725c890000ff954e180924c30d808dc965618f0b9d1f752e7b0b280f17e8

Observation dfe122b5-8d40-4404-9fa1-d59bf2f82354 · outbound

This paper cites Analyze the properties and spatial relationship of objects in the scene.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Analyze the properties and spatial relationship of objects in the scene

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.182178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:03.961199Z digest=sha256:d897edef3bc5e902e98abc75b4cec88e39977b602487b0125b4fe58b000e6b40

Observation 555da127-ea1a-47e9-93b8-43a22e8f0a3c · outbound

This paper cites Don't output anything else.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output anything else

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.893817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:04.025687Z digest=sha256:89bc8a9935c4e9e7659ad4c993f2f0ae842ed3317ad6319996e1b00aec03db78

Observation 1dc3a3b7-e30e-4d0a-a3b0-e4bf029bbf62 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:05.705709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:04.088203Z digest=sha256:cd69950587fc31c7b4402b17512114cf8e39de493a0e772f6e924d25299d7c2e

Observation 0c966068-2eff-4233-b2eb-aea3fb3858e4 · outbound

This paper cites Analyze the description and identify description-related objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Analyze the description and identify description-related objects

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.526378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:04.123579Z digest=sha256:db4c612159f178373e95d8a1170058a14a6f219cb1b21d536932fcd1b33cf1ee

Observation 4ab6c0e7-5250-42a7-a73c-9fe22efa05d4 · outbound

This paper cites Don't output any other things.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output any other things

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.267420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T21:53:04.127385Z digest=sha256:98328a2a2bcc238a140f5bc406f864d321a31ae2c2eb60978be95adaee83df5a

Pith citing papers

Observation ba066641-729c-4b3d-b6c6-854f6ebf7940 · inbound

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation cites this paper.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.383043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.383043Z digest=sha256:fe3b1293fac6c5cdad089cb117942cca39a03fc41dbefcb85693c22fd30d4174