Pith. sign in

Paper Citation Record · LEDGER

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

As of 19 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2506.23120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23120 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:53:04.127385Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:13:52.383043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 325774cd-e2b0-4437-9db1-1771425d731c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.793154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.793154Z digest=sha256:958b579bdd808c6a64999bfb7d7d9b02ba75b0d5f2b6015b753659a16c30239b

Observation 6a7764da-d893-49aa-80d6-879fc4821c98 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.729965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:57.848548Z digest=sha256:1b53b5f04eca23eb60cefa47754941c4cffa76b8bb5c8f27bbfb6cc8dc736191

Observation ec84b564-9542-4b2a-b53d-af68d50282d6 · outbound

This paper cites Semantic scene com- pletion via integrating instances and scene in-the-loop.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Semantic scene com- pletion via integrating instances and scene in-the-loop

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.576806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:57.941641Z digest=sha256:624763ac54b7b6597de6458f3128a0571c1eef169a4afca084eb97aba849bcc9

Observation eb38820f-994a-4424-a259-5980f3e22e0d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.406809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:58.068312Z digest=sha256:45e3d6c0a5287819d4aa1930cf41d52fdf190a5155b4be5067a8ce23f6a707bf

Observation 7525dbd6-533c-4a95-babf-e4d55a965115 · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Language conditioned spatial relation reasoning for 3d object grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.238738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:58.206308Z digest=sha256:d56229b926f662d4736e90a5b9360bff4f6510a20503f0400b98f964e458bad4

Observation 5b1a3463-d775-4faf-b2e6-b2f6a048ebe2 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:14.061588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:58.293363Z digest=sha256:ae9b397fb7c305f5516e3307610f2e46f8c929a4f43e2b233993ae41a95b5aa1

Observation becd4c34-3a8b-4b1e-9027-ac8dbaaa3d1b · outbound

This paper cites Mppnet: Multi-frame feature intertwining with proxy points for 3d temporal ob- ject detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Mppnet: Multi-frame feature intertwining with proxy points for 3d temporal ob- ject detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.902145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:58.381758Z digest=sha256:dbcd76fe1574dc4d7d73cbf5696e80a837ad0dfdc9851644188f24aa01834d85

Observation fd2f6363-7d1b-458d-a749-37362235b4e2 · outbound

This paper cites Trajectoryformer: 3d object tracking transformer with predictive trajectory hypotheses.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Trajectoryformer: 3d object tracking transformer with predictive trajectory hypotheses

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.687509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:58.480676Z digest=sha256:f218280b941a9bde6d6deae7be6493b557c2cad4227d11aeeffe6d01e11a6792

Observation 1c299235-7d13-492f-b05f-0ee47acad1ef · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Grounded 3D-LLM with Referent Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.607061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.607061Z digest=sha256:c8046ad995a0d68c8e85352fd3a6e2df4514ebeaab1c8c08411abf8e2b34863a

Observation a1c83751-31cb-426a-ac9f-b69ab4ad7c73 · outbound

This paper cites Reslt: Residual learning for long-tailed recog- nition.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Reslt: Residual learning for long-tailed recog- nition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.513529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:58.701391Z digest=sha256:b46f04120f58eb0cd49ec15bea27b873a7ac3760ec997f21286130b7114e6a34

Observation 458a1831-5f2a-48fc-803e-39cdd608875a · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.811072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.811072Z digest=sha256:526ede595d2e7feb3afd35d2e7061d59a6f8b2d1e327da61d8da4cf5ad01fb46

Observation 94ac9a92-49d9-42d3-8d0f-78cc8cc7c625 · outbound

This paper cites MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:04.929386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:58.917530Z digest=sha256:3d18c65b1aa09e5811363fab93c8c91e5bc17617389b42d489e5bd64d311c315

Observation e112cc66-d919-4cc3-b7ea-c1c26b72a5df · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.052695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.052695Z digest=sha256:9558578bca72dcf5f12a7f437f36861c96faa71a8980b6d2457c6cdcd678e9eb

Observation 90af0a97-5bb2-413d-9280-9877aa54cf5e · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Pla: Language-driven open- vocabulary 3d scene understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.155141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.155141Z digest=sha256:478af1d03d14af805ccce8fe15e5cf30360f42628f9bf293dd2dfbe49d78bdc4

Observation de2e4887-707f-4e28-96b1-bdf9f352fea9 · outbound

This paper cites The llama 3 herd of models, 2024.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation The llama 3 herd of models, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.360728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:59.243701Z digest=sha256:33326df0ec997acf3c8667d78bd58cc7b6fb0157482711bede70dffadc112b09

Observation 44a72786-e440-4241-b3d0-a9436f35572d · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.369243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.369243Z digest=sha256:76f589cbbb22122bf024af56f1da6a2d962539ae8e787afe9cfe32fe5b364c0c

Observation ad57e631-7596-4529-8f48-cc1ce6ee5a4d · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation ImageBind-LLM: Multi-modality Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.473536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.473536Z digest=sha256:8419e07b15323afd30aec1843b65ae21986e165f3982df2e18a381565208f3dc

Observation 97ce80bc-d7d0-4338-a09c-467d0f2c3547 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation 3d-llm: Inject- ing the 3d world into large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:13.146921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:59.578599Z digest=sha256:0d2936aef57015d0ee4373b59df15c53b310309578bbfac409abf4b78b6d4529

Observation 0f9edb34-9b72-4b97-8b5d-91ae6ec18f30 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.681855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.681855Z digest=sha256:2a5e4c04f46d70ff9c636f41100ab96071ec59cc87badd0fea741388387efdc5

Observation 34d95fbc-e937-4715-9210-663522ad7705 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation An Embodied Generalist Agent in 3D World

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.769267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.769267Z digest=sha256:fd32c3d797d940754d54bb47ed5fb2da6ecd146cba6fa52c7bd0a627094495b7

Observation 4b4618bd-cdea-4d2c-828d-1afa9f1c645b · outbound

This paper cites Multi- view transformer for 3d visual grounding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Multi- view transformer for 3d visual grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.946706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:59.867853Z digest=sha256:e834942f429aaccb54083777da40c0849fab38eb9fae6cb325d534f01587779e

Observation 62e5aba4-dde3-4665-b802-8ea3177c3189 · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:04.545494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:52:59.994335Z digest=sha256:55cea6fc6cd4ac0ccbe152ea18cec4f05f43875442b229cd3b8fed49a89f323e

Observation ef4f5f50-f464-48c0-82e9-bf42cd0c62ed · outbound

This paper cites Guided point contrastive learn- ing for semi-supervised point cloud semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Guided point contrastive learn- ing for semi-supervised point cloud semantic segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.761403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:00.083655Z digest=sha256:0e71e3c4ed54a3aac4484e78b086941f96101ac3360d1754b19531a4e22ae398

Observation 831e018b-d7b8-4946-89c5-d2f9cd460c8e · outbound

This paper cites Segment Anything.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Segment Anything

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.204418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.204418Z digest=sha256:d0859b6362f368bb6c18332dcde10a8589825a622a76638d4747c0fe9b84c4b9

Observation 7188b26a-7abb-4603-90da-ee7f35813e5a · outbound

This paper cites Oneformer3d: One transformer for unified point cloud segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Oneformer3d: One transformer for unified point cloud segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.583338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:00.306050Z digest=sha256:0f496e09ae923ee3b139a2026a75d03c1fcf845ea17b677553a7f8523aa75a5e

Observation 76ace732-993f-481d-af74-7ab8261c1a13 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.403341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:00.437200Z digest=sha256:e5d016f97d4bc47b9973c3c2a7f109c954ff06f88553db721be6ecc62900a2ba

Observation d9653dec-3380-4c78-8316-8b21d5210f57 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.558586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.558586Z digest=sha256:e224a7a7cd96b9e7efb65260b6049434a3bc2c0eff79b53b236d7010687becdb

Observation b4becde9-e5fd-46d0-80c8-0fe3f452cab6 · outbound

This paper cites Large-scale point cloud semantic segmentation with superpoint graphs.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Large-scale point cloud semantic segmentation with superpoint graphs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.158422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:00.644188Z digest=sha256:3797c1b8ae1d19c15924b06d0cf4f10fe269389407e2c20c612e7e2c37cd9f73

Observation dc14eb7e-771a-4c7f-b789-4daefd08ace7 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.967553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:00.717881Z digest=sha256:ec8df6cf4bf7db55cec42d6faae3938b02b20759454954db1e3ccb5f7cfaae2e

Observation 3ea31733-91bf-48ee-8dce-ef838f65b9e2 · outbound

This paper cites Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.799961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.799961Z digest=sha256:3c042f57d337015a71bd67bbf1b9bd7520e107420837c39f74f28bdd88cda734

Observation c4d0f4aa-fc8b-469b-bf92-4d7ef8153b75 · outbound

This paper cites Decoupled Weight Decay Regularization.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Decoupled Weight Decay Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.901821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.901821Z digest=sha256:76173ca89f6a11dbe9da06c6e3419b3e9115a2995295ab02efbf7fa0a2815880

Observation 27d46f3b-d777-4076-b839-5e64e9a65f93 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SQA3D: Situated Question Answering in 3D Scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.971816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.971816Z digest=sha256:4fa0854543c28e7f3bd8210235b12283f9cf6179420407a26c89afe5521953de

Observation 146b51d3-c45f-413a-90dd-24a29d730431 · outbound

This paper cites An end-to-end transformer model for 3d object detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation An end-to-end transformer model for 3d object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.753791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.046044Z digest=sha256:a69835b44cabae32ab476a6e5effb02a353c2a02f52bb847435377f8da74793f

Observation dd071416-a327-43c7-a1db-a6996a849c5f · outbound

This paper cites Boosting few-shot 3d point cloud segmentation via query-guided enhancement.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Boosting few-shot 3d point cloud segmentation via query-guided enhancement

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.575644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.149234Z digest=sha256:4624479065da12b0247ea7031d1d437fd1a84e5f64725650c6727a7ae23e6c9e

Observation 50ced8bf-31e5-4b03-8fc6-e2c65b4f4d2d · outbound

This paper cites Hierarchi- cal dense correlation distillation for few-shot segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Hierarchi- cal dense correlation distillation for few-shot segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.421226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.255571Z digest=sha256:1c521553c53828b7a383cc28d093fc592ae255c8b09fde27703a4c7e1fe4eee8

Observation f7cae70b-0dc2-4b0c-8eb7-ce7ca5f1d669 · outbound

This paper cites Scalable Language Model with Generalized Continual Learning.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Scalable Language Model with Generalized Continual Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.348484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.348484Z digest=sha256:cebb1689fc117839778f998d25363e03ade9a7397a4bdb62dd15a0e3dab04a75

Observation 584d8855-8324-4fd3-aae2-bdc8b758c3ba · outbound

This paper cites Oa-cnns: Omni- adaptive sparse cnns for 3d semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Oa-cnns: Omni- adaptive sparse cnns for 3d semantic segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.173646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.438300Z digest=sha256:367484eb05dfb4a9fa1fcc536bddad239962d3cc47167ef2b410f0b52ab549c5

Observation 0d582eb0-37e5-4246-a451-0eb88695eccc · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Openscene: 3d scene understanding with open vocabularies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.538579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.538579Z digest=sha256:7a7e03d5ee795cf8d24c33e897f79b1874d116679344e683f0b54440aaeaff4f

Observation 56dc775b-f6a2-446e-a8c0-172202b6e50c · outbound

This paper cites Explore the potential of clip for training-free open vocab- ulary semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Explore the potential of clip for training-free open vocab- ulary semantic segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.983916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.634815Z digest=sha256:e88b101abf3e17f2e33eeb998dd0da30734a61b11bcaf69672b393b57c5e186a

Observation 61f5a869-bb7a-4f88-b675-271d3b0fe13b · outbound

This paper cites Pointr- cnn: 3d object proposal generation and detection from point cloud.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Pointr- cnn: 3d object proposal generation and detection from point cloud

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.816844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.702626Z digest=sha256:cffa8ebe775fe33bb0b7a1ebf75fa8adecefce9ccc0f73e465ffee4a8a8326f0

Observation 1096ace6-d48f-4595-b0ac-f2ab5e1a36b7 · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.780508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.780508Z digest=sha256:f91ea4db7c2e7fb0e44bc91c85f1866ddbbbdec578bfc70d99d4c484f301f009

Observation 78a01afb-911b-4a4b-b494-590f26f7df4c · outbound

This paper cites Mind the interference: Retaining pre-trained knowledge in parameter efficient continual learning of vision-language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Mind the interference: Retaining pre-trained knowledge in parameter efficient continual learning of vision-language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.628744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.855126Z digest=sha256:1201d7a40975ea481d899144c49a38d73be19b4a4b9542445bdc78153dd75fe7

Observation 934c49e3-811d-49b1-b6ca-78cb5612802d · outbound

This paper cites Learning shape-aware embedding for scene text detection.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Learning shape-aware embedding for scene text detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.428096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:01.946814Z digest=sha256:b5327fef76cc10b5e31b5a24a02a68ebdb67064d29c07fc16243f76601a13c5e

Observation 0c26d983-be1a-410b-ae62-43726fbc61fc · outbound

This paper cites Prior guided feature enrich- ment network for few-shot segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Prior guided feature enrich- ment network for few-shot segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.189030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:02.042318Z digest=sha256:9b9b2d58d6c8e39eff35ec8f689efb446f56d324e34613efae1125d17f239c16

Observation 8e314eef-7b3c-45b6-b717-432117460287 · outbound

This paper cites Adaptive perspective distillation for semantic segmen- tation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Adaptive perspective distillation for semantic segmen- tation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.980659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:02.106679Z digest=sha256:17dcd83c69bb17a44f5c0446f0f75e78aa221fdc0a45b0111e95329ec5b64ac2

Observation 17d04d7b-54fd-4bcc-8dec-367477883c72 · outbound

This paper cites Generalized few-shot semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Generalized few-shot semantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.786638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:02.173893Z digest=sha256:1a7cc59fcb54e8b273e16f85ef2c5f410b9306a0c9fd4b66615a2a7e6085214e

Observation 8cc5ac61-fe44-428f-8d39-d6f24a46680b · outbound

This paper cites Learning context-aware classifier for semantic segmentation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Learning context-aware classifier for semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.577862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:02.264287Z digest=sha256:0303d16b9f5eebb3268d4f919a149cb92b683c8fa04d04b376883ceff406b5aa

Observation 04974178-964d-4fc1-abcf-d397d4456184 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.361077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.361077Z digest=sha256:42d38f32f873344ed530944305fd2e989e8b52518601fc48baae85adaf020139

Observation 19030596-2323-44f8-bf0a-ff83d7969880 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.436049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.436049Z digest=sha256:0bd6d0ff27c6fe1bbc044f6aaa43ab89af4bec9b4474350bde8ad266b65c9391

Observation 8e24a87e-fa6e-47ec-80aa-350393141259 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.513697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.513697Z digest=sha256:2266d337faf7db0f3bd630d620161afe7d4b62060f96199c734d852772a8b15c

Observation cd0acc59-f07b-445c-a6af-dd175fe92d31 · outbound

This paper cites Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Llm- grounder: Open-vocabulary 3d visual grounding with large language model as an agent

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.433475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:02.621278Z digest=sha256:e0e0c932244416d81b5b73dfed66468c63aa9bd918df4c6eecf06d23dcc1d61c

Observation 222944e9-9eef-4452-891d-bcb190a9e218 · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.670233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.670233Z digest=sha256:1cd07ce325f93b73763d43280ef9c0205211413808ead550b4564fe0a0ab500a

Observation 691d61e1-8cbe-46c6-bf68-9bd49f55c352 · outbound

This paper cites Unified language-driven zero-shot domain adaptation.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unified language-driven zero-shot domain adaptation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.286190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:02.785477Z digest=sha256:94113171d2629afb5c3934a23e23608d44f724c9903eae2b1f6a9f5870ac7192

Observation e364c948-8f20-4ccd-9c55-f72f539d2bca · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Visionzip: Longer is better but not necessary in vision language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.155714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:02.857736Z digest=sha256:9fdd7249edc1df0d1177b1d167bf2769893b583788cdd04e14b293f22d2355a3

Observation 59b157f9-8fa1-4d26-9496-c327aeed8fc7 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.923647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.923647Z digest=sha256:20a5b681b5b42e349647b32c19318823a0060c35d3eaab197613d24eff582bda

Observation 5ab31b1a-6074-49f5-a2cf-f4ec5ac7f4da · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:02.977100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:02.977100Z digest=sha256:b8d98965f052de25a3ba9c2185cb2d92ffde694b8a983f67e8e17a16fa8a1dc6

Observation 944d160a-1c9a-4f48-a492-c43445372d40 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation OPT: Open Pre-trained Transformer Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.034745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.034745Z digest=sha256:86a7b56bc45d805bcf0a2daa0e24f1ec79fa7e3be2da4e66f92c9bf5b3b26458

Observation 1ea536c1-6748-486c-aed5-8131ee72c7df · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Multi3drefer: Grounding text description to multiple 3d objects

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.978931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.093626Z digest=sha256:c9115db03ffaf0f12740d3177326ecced39d5b4e259e13bba1818cfcc93198ed

Observation ce0c7d61-939f-4525-b2a1-c6a915a0a248 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.839030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.159558Z digest=sha256:af1b6fae8ffa40a0b7f2cf9fde3b20efbd4f44656570b4434c2d7e64caac2439

Observation 5fd12392-a455-441a-8630-a817b0b433fe · outbound

This paper cites ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.225641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.225641Z digest=sha256:eeb80e8cba1806e6b8816e51a9b07bf63579c13a86ebee78dc67c23d4d2f410f

Observation 8820c6be-a73e-4ae8-a85b-38ae4e588bf0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:03.285974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:03.285974Z digest=sha256:834d6158e232f987a1f5a8ab0892134f3f8cc0e9f0171dcd89748004869d6978

Observation c264643a-7b76-4de0-9b77-f692ca3febba · outbound

This paper cites in the right of.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation in the right of

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.641534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.364743Z digest=sha256:8ff7b0a47471bde7ac59c2dae1363173fb9be1d4a442f75ccc93d769a79f0c83

Observation 79c14c01-dd42-438c-8312-73dd76869e8b · outbound

This paper cites Focus on detecting and distinguishing objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Focus on detecting and distinguishing objects

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.491329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.415149Z digest=sha256:1b544136337f052e3c625bdf235718b55e497eff09ae5d4417c8024aa904e35f

Observation 325f6a96-a76f-4951-8e96-bad0ba2a8387 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.344472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.466751Z digest=sha256:ef4333f2c0019bebea5b3264c22a919759a0df4fb181c6b516de8b93bfa906b8

Observation 05db9656-4a26-484c-add9-192c02bf0113 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.151275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.522108Z digest=sha256:cb6beae5e1e489a8315f9d02ad21c0f10ed8d0f1ab50c111eef297e9ae838544

Observation 560ad566-91ec-441d-af03-e42af7682409 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:08.004162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.581218Z digest=sha256:6037aa22a86363427f4b0699b0a73b6e19def6d49f50c0260a62dbfef30dcbff

Observation 61cf7d28-9358-4867-85b6-8c1c07519fab · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.821554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.628077Z digest=sha256:e81ff7a28c77a9c2d0df2f4e4f13e9bfb445bf8b1fb099b71193541f7ce97d9d

Observation ac8e90e1-e941-4283-8a6f-0a8799e076db · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.673168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.675226Z digest=sha256:2409255d35124ba3a72f391d60013f6c179cf0aa607ef5a697b72bc6ce2bbc7d

Observation ba3b2e56-4854-4d61-ad78-6809cc68d482 · outbound

This paper cites For multiple objects of the same category, use [<ins_id>][<ins_id>][<ins_id>], etc.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation For multiple objects of the same category, use [<ins_id>][<ins_id>][<ins_id>], etc

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.463313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.739329Z digest=sha256:e17749135854f8861cd4e1b024f54c8bc05b5b3e1f5a70f5bc7a08c574766d60

Observation 90115f2d-f5d1-4533-8061-275982f0466d · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:07.134037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.785678Z digest=sha256:3ef61a5fbba9117e2a12d8a4592b8c49f0960617a36fa76430177fe1b62e7254

Observation 83c98ba2-e58d-42af-8cde-eca333407c73 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:06.791230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.839144Z digest=sha256:b7222515bcd12cabd3a272e6cf9142617331ef216119750b99c3899320dfa8ef

Observation 684938be-0073-4770-8639-a5573a11fc08 · outbound

This paper cites Don't output anything else.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output anything else

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.518179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.898635Z digest=sha256:47e73e580022955ae0d8c725760d178123be115566a2e194dd1f6d507795c99c

Observation dfe122b5-8d40-4404-9fa1-d59bf2f82354 · outbound

This paper cites Analyze the properties and spatial relationship of objects in the scene.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Analyze the properties and spatial relationship of objects in the scene

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.182178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:03.961199Z digest=sha256:460dc15d566b580895914a7752ed84dbf34a3fc86af6fdf25f7104540de1b9f8

Observation 555da127-ea1a-47e9-93b8-43a22e8f0a3c · outbound

This paper cites Don't output anything else.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output anything else

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.893817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:04.025687Z digest=sha256:7b18654a7b1a1b117384ac216f2eb965f0933f03519e00cd8c1c72c660692f41

Observation 1dc3a3b7-e30e-4d0a-a3b0-e4bf029bbf62 · outbound

This paper cites an unresolved cited work.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:53:05.705709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:04.088203Z digest=sha256:f888a3a86d7514f844fac694f9da14d6c67d10af3d61c642e1640b7c05352118

Observation 0c966068-2eff-4233-b2eb-aea3fb3858e4 · outbound

This paper cites Analyze the description and identify description-related objects.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Analyze the description and identify description-related objects

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.526378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:04.123579Z digest=sha256:3a431c5538d0b9407ea6613c36105d582bb309d7fbaaa8f1408c062acb05f933

Observation 4ab6c0e7-5250-42a7-a73c-9fe22efa05d4 · outbound

This paper cites Don't output any other things.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation Don't output any other things

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.267420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T21:53:04.127385Z digest=sha256:790bd1363251d60d8ad699482b933f8e61bc2a7df4577110bb7f244659d65f72

Pith citing papers

Observation ba066641-729c-4b3d-b6c6-854f6ebf7940 · inbound

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation cites this paper.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.383043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.383043Z digest=sha256:4eff0d06a207bf4cfdb2f439398c7a3802de199fd0a91e3ae0b0842b7eb7b82e