Pith. sign in

Paper Citation Record · LEDGER

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation

As of 14 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 0 inbound Pith citation observations for arXiv:2412.02402.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02402 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:35:20.663843Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

92 of 92 outbound references displayed

  • verified exact11
  • verified fuzzy44
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66c673e3-7849-474e-94de-fca8aa8d2b86 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.231480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.231480Z digest=sha256:8ab43faa6ccf06670f79a27967a4a0d606d658b0d6cf3eec21c5961ddeefa9e1

Observation 86418cd4-de27-40e3-badd-8b3a9bff1457 · outbound

This paper cites Localizing moments in video with natural language.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Localizing moments in video with natural language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.237165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.237165Z digest=sha256:fe8713f0266c7779a3a8d4bf45ef90476ce888d5de7b9d7737746726d6e02a33

Observation 93102e63-0f06-4130-8ba7-e079449efc61 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.242135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.242135Z digest=sha256:c587860279d60bf7601f83c3ba52eafced4a7dacbc474c5b4ba21ba19ef7bc93

Observation ef71f953-536f-4a36-ad73-df001c704b25 · outbound

This paper cites WeaQA: Weak Supervision via Captions for Visual Question Answering.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation WeaQA: Weak Supervision via Captions for Visual Question Answering

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:21.091139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.246860Z digest=sha256:039d42934b52898870d2d79cb3f2ed6297eea5f7c0c6ceb48a710c53e4c889e0

Observation 3a6efffc-3544-44a3-a3dc-f0e247487913 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.252192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.252192Z digest=sha256:0b0ea50437d7d788ccf2d13426c4ee27ab161786fa3a9e7f52ace3f2f2fdd63a

Observation 65261456-3502-4684-a126-cbbf2bf7ae4f · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Language conditioned spatial relation reasoning for 3d object grounding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.257074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.257074Z digest=sha256:671153e3c0e251cbd9e65cf2af391c2586670907bbf9f1793a08cf343c9ede7f

Observation f2705799-0cea-40c3-87ac-b5c63327b8c6 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.262126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.262126Z digest=sha256:b311363952bc43d89dccf92b15d721753a1ee576f9890f7a6de37b7dca17f5f1

Observation c55aecd9-e080-4c3b-a533-938c503cb105 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.266674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.266674Z digest=sha256:299d178513533a9c74c5f5d7452a97f57dfa01175e8ef089a6308de6c1e776f2

Observation 64fdb21e-e3e2-408e-828f-efe4e26819f7 · outbound

This paper cites Weak Supervision and Referring Attention for Temporal-Textual Association Learning.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Weak Supervision and Referring Attention for Temporal-Textual Association Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:21.052476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.271471Z digest=sha256:8f97c8606c121d9179fe5fef79e2bff8f559086bbbec439046ce2e7866c08451

Observation 0b181c94-4aff-4b1e-9827-09d26aee646c · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.276661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.276661Z digest=sha256:220adb8eeeaf5a5f466f5d19a4fd3e9cf7de1fb49586bce9e8dbc72119ab129c

Observation ccacbda3-33ca-4d30-a573-22c608da1381 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing, 2024.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.281463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.281463Z digest=sha256:dd5a831218abe2f798820b885312d45708a0bb63068f74c91d759fe3330fd521

Observation 5f36c49f-a69c-45f6-9919-79725e0eb90c · outbound

This paper cites Enhancing video-language representations with structural spatio-temporal alignment.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Enhancing video-language representations with structural spatio-temporal alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.286155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.286155Z digest=sha256:00d60e0203a8d373cb94afed845bc9f8e7c982c4eaa2b7dc120b0a351246dd31

Observation 5d7c917a-d0e2-4e00-97b0-1688810e1c8f · outbound

This paper cites Free-form description guided 3d visual graph network for object grounding in point cloud.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Free-form description guided 3d visual graph network for object grounding in point cloud

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.290172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.290172Z digest=sha256:92b067218b83838f7bab90465b23b6503d0a498433f5612a3888ac069edf2bb5

Observation a072464d-e775-4f20-a086-1fc5fd323162 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.294196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.294196Z digest=sha256:d7a4e05c93bcd8f4732dd5b5610f63edf03d15545c0541ed5e07991e195191f8

Observation dd6ba9dc-c8f1-4de6-b1f3-92eb9eb59cdd · outbound

This paper cites Structured multi- modal feature embedding and alignment for image-sentence retrieval.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Structured multi- modal feature embedding and alignment for image-sentence retrieval

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.298756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.298756Z digest=sha256:6606e5b6d273761cb40bac90564b2048389cadd7058b0c8ceef21c86b8333a24

Observation 5240e487-12a5-4aee-a80c-b9c08f1d5c58 · outbound

This paper cites 3shnet: Boosting image–sentence retrieval via visual semantic–spatial self-highlighting.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 3shnet: Boosting image–sentence retrieval via visual semantic–spatial self-highlighting

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.302786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.302786Z digest=sha256:147ce88cb0887e98493da69ca695900cd0ac0214b42c251a305b5cdb940fb135

Observation 8df2d7c9-d48d-4016-93a4-10e8a36c739f · outbound

This paper cites Vqa-lol: Visual question answering under the lens of logic.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Vqa-lol: Visual question answering under the lens of logic

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.948690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.307081Z digest=sha256:fc2cd1d57474ca7a95aec646517cec295d2f13203b42f39d2dcdc4a22b8bd3b3

Observation 897f0df5-a29a-4306-a367-ee33934d29fc · outbound

This paper cites Person re-identification method based on color attack and joint defence.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Person re-identification method based on color attack and joint defence

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.932779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.314186Z digest=sha256:0b7986f383e5852db78e9dabc6803ec02a5946fd1d506d8e374f83e64c82a6ed

Observation 4c7deca1-cbb8-4cee-8693-e6a5b4dfa186 · outbound

This paper cites 3d semantic segmentation with submanifold sparse convolutional networks.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 3d semantic segmentation with submanifold sparse convolutional networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.319004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.319004Z digest=sha256:f22eaf85ab5fb0574c305959949f6b335dfc3b286bb1590ba0753a3ed576fa90

Observation dfc35995-6698-40b7-b170-4d0f3c04438d · outbound

This paper cites Segpoint: Segment any point cloud via large language model.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Segpoint: Segment any point cloud via large language model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.906666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.323821Z digest=sha256:e710034998f40f9e25bf98de0b217bda6dd5ca51e3375bb3308803a3ca8413bd

Observation d407bdaf-be1f-423e-a986-861936cedd2a · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 3d-llm: Injecting the 3d world into large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.328792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.328792Z digest=sha256:c192e4e67c89f65de3864272df26512c34588f62d0cce984f9874159f7158a68

Observation 9ea74b37-bbbf-49b7-9edf-769a8dc1b8c2 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.333277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.333277Z digest=sha256:03ae2c81a7a5ce8e3a78c3460cc72d34b5f49ee54e9c401ff50ad439396aa1b6

Observation 691a22b6-8e62-41fd-9020-fb9b3ad7ce03 · outbound

This paper cites Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.338471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.338471Z digest=sha256:0bab3eb613228aa35677555e8034624d889ac5afc6267666a47def96f9a83769

Observation e1e7b209-f73d-4c38-a180-d146645478ac · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Text-guided graph neural networks for referring 3d instance segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.343905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.343905Z digest=sha256:1462d0dc90923e76b00d0e1ea65bcd54ec032d3ac21e361a215b89e67ffe2e68

Observation 86865ac9-c6c7-45cd-b682-100b786c5381 · outbound

This paper cites Referring image segmentation via joint mask contextual embedding learning and progressive alignment network.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Referring image segmentation via joint mask contextual embedding learning and progressive alignment network

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.348573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.348573Z digest=sha256:b922290f9a17d7becc6d9e96f266395507e45d17caf1c91da0b6e514a56733a1

Observation 7ff30ec0-7027-4525-b672-9fdc7dc4c104 · outbound

This paper cites Bottom up top down detection transformers for language grounding in images and point clouds.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Bottom up top down detection transformers for language grounding in images and point clouds

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.353513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.353513Z digest=sha256:7b5c37dfc22eaa2fcc73641d33e17c5e2c8084beb0e3f09649e6ba7edd555ce0

Observation 5e9ce923-4eec-41b9-b426-0ecfd825ae0e · outbound

This paper cites Comprehensive multi-modal interactions for referring image segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Comprehensive multi-modal interactions for referring image segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.857343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.358251Z digest=sha256:5bb75b1b26359a7169aa5371434943ec34733a0e9101d14383137f4e2c410457

Observation 5ecf5cb1-3442-460f-bd97-e2179ad1acd9 · outbound

This paper cites Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:20.984655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.362722Z digest=sha256:abb25f1c4dfadf044b0da870918cbb0368db1c1e92e89e4ae0011dca1469b418

Observation 5b6d9b74-b99c-47cf-8df3-40cb84a9646a · outbound

This paper cites Flexible visual grounding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Flexible visual grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.367678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.367678Z digest=sha256:63363817133be5659b80ef9d4ec8275cb0aaa68175f13033f1a35337283e55bb

Observation 23b67e82-1e64-4f4f-8748-6df33ed8ba79 · outbound

This paper cites Stratified transformer for 3d point cloud segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Stratified transformer for 3d point cloud segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.841054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.372326Z digest=sha256:6eba3b485cb824cb0afc2604231b6957fef189fc2780857db30d8b7aa4b3ba54

Observation f9134daf-ecde-4535-8558-0a223070e87b · outbound

This paper cites Mask-attention-free transformer for 3d instance segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Mask-attention-free transformer for 3d instance segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.823421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.377072Z digest=sha256:523c12b8e81eb02a7497691a3a1b34c07592582d23032f89d10e49b533d86928

Observation a75278c8-ea1f-4ae7-ac20-58ca0c0423ab · outbound

This paper cites Large-scale point cloud semantic segmentation with superpoint graphs.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Large-scale point cloud semantic segmentation with superpoint graphs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.806638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.381838Z digest=sha256:56d3aea638d073fd1b26db74fe139b1551bf12a2ed65177bcfb9f8906e3c3ccf

Observation 1e2f7944-cbe1-4dc6-a99f-c6202e6f3995 · outbound

This paper cites Weakly supervised referring image segmentation with intra-chunk and inter-chunk consistency.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Weakly supervised referring image segmentation with intra-chunk and inter-chunk consistency

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.790357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.386542Z digest=sha256:07868abfc48a53e2ed16d27a870c1149eb692c8c8572fc6159b396a1f590eb72

Observation 5200bb1a-b612-42a5-b091-564413e63cb1 · outbound

This paper cites Fully and weakly supervised referring expression segmentation with end-to-end learning.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Fully and weakly supervised referring expression segmentation with end-to-end learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.773055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.391780Z digest=sha256:b87bd6a7ac461d848fbc6612fc39065934092bdb5c29f5e27b988ba6fe137d4c

Observation abc6104c-4c98-47d9-91da-7d4ebacc9137 · outbound

This paper cites Fine-grained semantically aligned vision-language pre-training.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Fine-grained semantically aligned vision-language pre-training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.755948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.396223Z digest=sha256:7747d4910349ab0c87cba72451e1730c6774d049df9b4b7b41012ff4e60829b7

Observation 2de1f385-75fe-4565-8529-1ea59c171791 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.400784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.400784Z digest=sha256:d96794bcfe3da34f4435c217821bd9561faa736d073b386a5bc2b6c0bfc6747d

Observation c3a78231-fd75-4047-869c-a2614e1e8b47 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.405459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.405459Z digest=sha256:06dd64f495bbf50fcd52ebf67d588dba4ea1f39ff9216dec0984777d81a15494

Observation eff90e3a-9b6e-4419-815b-3f703f76ecce · outbound

This paper cites Toist: Task oriented instance segmentation transformer with noun-pronoun distillation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Toist: Task oriented instance segmentation transformer with noun-pronoun distillation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.717881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.410222Z digest=sha256:b55c10a4966442de103af4a52e1e21bdef95b3e03e9c46e6ae592ee9b3646d2a

Observation 3a6f97ec-6e78-4aa3-9678-072252225644 · outbound

This paper cites Understanding embodied reference with touch-line transformer.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Understanding embodied reference with touch-line transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.700486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.414906Z digest=sha256:050a57429f4a638cb29ccfe64966820e3b30047facc7c3da022a60f82faa58f4

Observation 525d8460-5579-44a8-9e64-32b9638cd338 · outbound

This paper cites Transformer-empowered invariant grounding for video question answering.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Transformer-empowered invariant grounding for video question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.682439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.419412Z digest=sha256:27e1f10c1dadcf54c68c5bc4b9317b0768f462fa61acb83c437b8fd582f516ae

Observation a61a852c-095a-44bd-91eb-17b7c7fdc2b3 · outbound

This paper cites Laso: Language-guided affordance segmentation on 3d object.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Laso: Language-guided affordance segmentation on 3d object

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.665764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.424006Z digest=sha256:91b0ccc845df155bb652529b02f86cb2798ac7f03957661bb9982b7a6834865c

Observation e03c661d-88ee-49b9-8d26-1c9e0b5ba7ef · outbound

This paper cites Instance segmentation in 3d scenes using semantic superpoint tree networks.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Instance segmentation in 3d scenes using semantic superpoint tree networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.649250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.428757Z digest=sha256:2ee97779e22ac039eee33bffc1988c804367105af699691f5380f4c07f4be70f

Observation 401b6487-1e10-4102-b747-2d0b70773dc8 · outbound

This paper cites A Unified Framework for 3D Point Cloud Visual Grounding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation A Unified Framework for 3D Point Cloud Visual Grounding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:20.962563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.433380Z digest=sha256:a0d0f671aec14c2b4390ee2342d72389ef8d111bc48346bc859e7961f566288a

Observation cc8be356-efdd-4477-8666-45810174251e · outbound

This paper cites Referring image segmentation using text supervision.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Referring image segmentation using text supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.632219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.438339Z digest=sha256:679e8db71352f3560f2383d25e8180c33626b13fd6cc840f7f03c2fb87ac7f37

Observation ad626953-6202-426f-91e9-e41633bb18d0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.442554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.442554Z digest=sha256:886496cefd5a50397bab5002e522a74010b55246a17cf1fed4794c709c4ed5aa

Observation be229806-f0f1-4985-ada9-86e792f6e113 · outbound

This paper cites Scaneru: Interactive 3d visual grounding based on embodied reference understanding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Scaneru: Interactive 3d visual grounding based on embodied reference understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.615114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.446933Z digest=sha256:c6f34c6c615cbe7f8f364ac7c8401bbbf5a1ac14a13bddd3cfbff40a75a3de0d

Observation f8f36734-64ce-4e7e-b9f4-7cc6589206c5 · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.451182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.451182Z digest=sha256:43e2fece9ccc30e9c9111e8563a1e5859df1de5ba99546db9c4024ab458c06fc

Observation d77d5778-af78-4928-b193-6f69ca26310e · outbound

This paper cites The stanford corenlp natural language processing toolkit.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation The stanford corenlp natural language processing toolkit

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.586741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.455490Z digest=sha256:b0f9039ef42fe967e90454949c3ff13dba8ba0b191ff555d0054ad9752c06d1e

Observation cfc0ccf2-800f-4554-ba21-1b64175489b5 · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.459876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.459876Z digest=sha256:749479825cc5468abb95e0e56bf86604476a1e4a72e7ddc4c7164ffa58e0ec7e

Observation 5586e9d9-e49f-45ac-aefa-cb2cf100d86e · outbound

This paper cites Weakly supervised video moment retrieval from text queries.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Weakly supervised video moment retrieval from text queries

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.559697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.464165Z digest=sha256:936fc5fe29fe0c27c8a1e47c5c65168e496aed0286d641b91fbd3fda3a569c75

Observation 62bd7a5c-4078-4e29-8bda-43fc4ea8c32e · outbound

This paper cites Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:20.926325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.468888Z digest=sha256:6e0aa576857a70ff2a6486c46426a3aeb05cbb9eec19bf09ec6601bbc28e454e

Observation a818af2c-e3fb-4f91-91c5-c9471079b479 · outbound

This paper cites Deep hough voting for 3d object detection in point clouds.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Deep hough voting for 3d object detection in point clouds

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.542229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.473978Z digest=sha256:b5a581cdcb62889f3ca4094e361c62af2450154071f577b7b9ebde1e2265781b

Observation 8387d02a-0b29-4b4a-88af-94d117592b61 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.525560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.478744Z digest=sha256:d2f2177e3d5cc9f2b96d8c294eb6240b505be7392a44a1404dd07fbddcd987b0

Observation 0d4b4a55-8d99-4297-a8f9-931ded1ef39a · outbound

This paper cites X-refseg3d: Enhancing referring 3d instance segmentation via structured cross-modal graph neural networks.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation X-refseg3d: Enhancing referring 3d instance segmentation via structured cross-modal graph neural networks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.507677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.483640Z digest=sha256:9b37d1773c8ed4dfa228b69c58abb453afe9bab8414cc31df8fb529427e80a24

Observation 338825f4-7d69-4a2c-aa16-491c5f84b807 · outbound

This paper cites Learning transferable visual models from natural language supervision.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Learning transferable visual models from natural language supervision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.488477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.488477Z digest=sha256:7892c49473e066eecf7307dea51218f56dfc9da1024e3feb83184fb481f392d9

Observation 3e890971-ef3d-4bd0-bad6-5816dabfeefc · outbound

This paper cites You only look once: Unified, real-time object detection.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation You only look once: Unified, real-time object detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.493295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.493295Z digest=sha256:b5f50c41f99d43ef93428fa5bb4d85e9b9e6d41f7624df63371124aa2cc75ba0

Observation cf4a6b87-c014-4444-8323-8ef14e958b4f · outbound

This paper cites Mask3D: Mask Transformer for 3D Semantic Instance Segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.498039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.498039Z digest=sha256:7d03abd968bcea67e6bfb1775328313da013f5e4996252e46b348d333b722f98

Observation 444bc12b-700f-4626-8c0f-217496f9a719 · outbound

This paper cites Mpnet: Masked and permuted pre-training for language understanding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Mpnet: Masked and permuted pre-training for language understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.469640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.503343Z digest=sha256:d6f5bb7587499400f1b3f1418e5cc318334175d22db8e50f143e90445321fded

Observation 5b8c2274-ed08-4afc-bd08-7f279f839f40 · outbound

This paper cites ReCLIP: A strong zero-shot baseline for referring expression comprehension.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation ReCLIP: A strong zero-shot baseline for referring expression comprehension

Reference 59

Resolution
verified exact
doi, observed 2026-08-11T23:35:20.731696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.508347Z digest=sha256:e1d177d875bbe029dfc2adce065dfae0d9d43f9a7d0290a5e23411694da8d5a8

Observation 24502735-7495-43bd-a1bc-ad5f5dce8b80 · outbound

This paper cites Superpoint transformer for 3d scene instance segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Superpoint transformer for 3d scene instance segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.453210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.513161Z digest=sha256:ac68e2925ec59f4877c97cf370e23df160b89d000c0b68b91fd4efc875e412de

Observation 6da5567f-2150-4860-bd4c-c5646015dad6 · outbound

This paper cites Text augmented spatial aware zero-shot referring image segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Text augmented spatial aware zero-shot referring image segmentation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.517871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.517871Z digest=sha256:456436a1fc7cd3f677d3b2ccce8f2670b345568585d7bad7a77b8ec4d0c6b7fe

Observation 476a9e87-7043-404f-ac7e-9bb720dfc74d · outbound

This paper cites Interpretable Counting for Visual Question Answering.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Interpretable Counting for Visual Question Answering

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:20.887589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.522490Z digest=sha256:ad5335c5e40c44036262088c8fbbf4346fa77610c6ae1aa9907c31b7ac83f184

Observation 0740510a-5569-468f-935b-f2a0b1e7cb39 · outbound

This paper cites Attention is all you need.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.527617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.527617Z digest=sha256:74d99112869632b915fdadc4e2fdb9f723d1fc995af0a9090bc01e3ccbfed344

Observation 613c5d7b-3253-4f5b-8695-fbcf3c40f772 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.532072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.532072Z digest=sha256:5fce5b5dc3fbdadcbf38f47b748c5e45505ac1ebd6e8891271af1c3becf00f2a

Observation d7df6c1e-4204-4312-8ea8-84228356509b · outbound

This paper cites 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:20.851094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.536982Z digest=sha256:77e79cf553d2b4485128bca8bef62445c6675fba1b1bc8f22a8c9d31887749c1

Observation ede7de22-f7a1-4904-9b6b-ac5363ee29b4 · outbound

This paper cites 3D-GRES: Generalized 3D Referring Expression Segmentation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 3D-GRES: Generalized 3D Referring Expression Segmentation

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:20.829312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.542038Z digest=sha256:4daac76fbd1e7644944499ab895e4ac702b72a9b7b8e89f26308262b916d77a1

Observation c75f2d21-f628-43ba-a8a8-52e06e9695b2 · outbound

This paper cites Rethinking and improving relative position encoding for vision transformer.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Rethinking and improving relative position encoding for vision transformer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.426059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.546980Z digest=sha256:5cb9a57a43ba9cfc6f2990367ef5fb92d91c62793f6c26c2ddaf309e12b7e3e5

Observation 5cc55593-1a95-40cc-ba82-8f53107d540f · outbound

This paper cites Towards Semantic Equivalence of Tokenization in Multimodal LLM.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Towards Semantic Equivalence of Tokenization in Multimodal LLM

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.551468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.551468Z digest=sha256:0cea90d9a00fb23aff46987595c04b8452645f4fdf677772dfcb80616b3dd321

Observation ad106fa6-16d3-4087-9cb3-533c3fe885a2 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Next-gpt: Any-to-any multimodal llm

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.408387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.556187Z digest=sha256:0cfe8a82e891cfd95d75541c4fc612b5d2ca76d49f61858737a68790966d804d

Observation 9eff1af7-21a2-49c2-8c9e-546f7c1ce603 · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.393168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.560544Z digest=sha256:6aa0736e1db0d8823cbf5a4b3838bf21e8c35c291e26d4363e94bc18dfef36af

Observation 059ce938-98dd-4ae4-95f1-973f9f4bfa34 · outbound

This paper cites A Unified Framework for 3D Scene Understanding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation A Unified Framework for 3D Scene Understanding

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:35:20.790252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.565037Z digest=sha256:25fc58cc58ec6071dc700f3cab3f557268e723cc5caefe59bf2cbca842ddc6f6

Observation 438d85db-fb73-4b6b-9456-8d093fe80274 · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Sat: 2d semantics assisted training for 3d visual grounding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.378358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.569392Z digest=sha256:6c0050f64d2d7a6238ac0d69e68e497eb0bc01553efd08fdf01ad4a2a897c90a

Observation 7e80e01f-b43f-476c-bce1-a5182a3aa5ab · outbound

This paper cites In- stancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation In- stancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.573330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.573330Z digest=sha256:126cca95397e4a4757e7c1d3276bd3567aed13fe57e738e4fb1b4742cef1d690

Observation d33a0356-a7ef-4a71-a945-043ce0987598 · outbound

This paper cites CK-transformer: Commonsense knowledge enhanced transformers for referring expression comprehension.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation CK-transformer: Commonsense knowledge enhanced transformers for referring expression comprehension

Reference 74

Resolution
verified exact
doi, observed 2026-08-11T23:35:20.703936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.577618Z digest=sha256:0bf575af14ec456cd6e66e8707b6e036273fda20c1bae79f18adaaa716ae1fc9

Observation 86a2e0f9-8981-4c1e-b7df-990ee68c6a58 · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.581722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.581722Z digest=sha256:050bc889dd9268de51a16c123a3f354bab19a415bf2ee4359e37303ef78b61fc

Observation 0fdcb3aa-f801-4fb5-91cf-7c23597f5216 · outbound

This paper cites Towards Learning a Generalist Model for Embodied Navigation.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Towards Learning a Generalist Model for Embodied Navigation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T23:35:20.585860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:35:20.585860Z digest=sha256:f2cd14f8cba58f9db34accaf7ea97a0fed3571f4ba91bc03d07177ea34730ded

Observation d54b8337-8471-4c91-a165-0ab361378200 · outbound

This paper cites left”, “right.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation left”, “right

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.342961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.590132Z digest=sha256:cd7d11328d5f6957bf0daf4fd31e62336a841e2e877068639db5e460f874f0bd

Observation 4cc595ef-4e5c-4217-a288-46daf24a8e36 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.326917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.595708Z digest=sha256:be73c6e6d377303a56e3c983bce4b33b783666cd659426ece7ac532b3447e2a6

Observation de3711d3-b46c-4c83-b6bf-241876dd6d00 · outbound

This paper cites Limitations.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Limitations

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.311479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.600631Z digest=sha256:9e5a3b98a2b6cf57c6fa16ef7c4e9ad794b53210c0990e3f54d400d8c3f27744

Observation 59c0f4a0-4694-41b7-b8c4-15d84606ef57 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.295825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.605602Z digest=sha256:5be901c73000ab6bd80fb5f7788e611abae01a04104f71d23a2bf66b34128c68

Observation d6c2dc30-f5d8-488c-ae09-38aad1bbf276 · outbound

This paper cites Detailed experimental settings are provided in Sec.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Detailed experimental settings are provided in Sec

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.279306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.610880Z digest=sha256:4700fbf21317a3b9d6dd0c93f3ed52a9e09097c094ccfcfb0b58865aa6129a79

Observation 9c7ff5a0-ce16-4925-b393-4d7f88552dbf · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.263315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.616330Z digest=sha256:76d7cc1ffc00b6754770313a85b42158e3a3238461b367f7c5fbc6358bb548f9

Observation d7a9b8d7-2e54-414b-942f-c7960e97ee15 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the paper does not include experiments

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.247433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.621089Z digest=sha256:cdf9fccee7b2a43425937a1e3528556cf14e142324e7f8938636e25113a56a6b

Observation b486fab9-93d7-4c22-9610-d7ae331c1cc3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the paper does not include experiments

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.231667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.626328Z digest=sha256:558f2b471fcc18195519200c11d8e25aa0a38bcbe5d933c76f538d6bb22c502b

Observation 93ee714e-964b-475e-9f53-dbaa90981f14 · outbound

This paper cites 4 and Appendix.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation 4 and Appendix

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.216461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.630884Z digest=sha256:5b7a8adbbbcff395d498afe1c647c584b8b69a2c9d8382b787e95eb275163f30

Observation ecc6d0c8-f297-42da-b618-b0ccba2d105e · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.200530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.636245Z digest=sha256:5112b767cc12aee1b58316d2f9c2020f5cfdf216d57d38e093d01cd2a5a98415

Observation 91270bbf-d720-496c-b11d-6b9d2c9ed8fa · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.184463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.640947Z digest=sha256:8b9aa9d6d04473abfaa2de175c177848f11e32ca96ac5bf4d9e0559da10d682c

Observation 5ffc036e-e8eb-412e-a69b-100558a344ae · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the paper poses no such risks

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.167682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.645494Z digest=sha256:5aebe69d356218ebe99f36258a46ad56c8e1be8cdf6d331945a762d145b145eb

Observation 1e1d8fec-e327-4ffe-ad79-99b3cab1b7fa · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the paper does not use existing assets

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.152549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.650172Z digest=sha256:bff338e526789b1666d684af0d20808c757e88af0480b5e4a2a4a35fa2291129

Observation 782fdb6f-179a-49c2-9882-02215be6d888 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the paper does not release new assets

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.138405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.654903Z digest=sha256:2b22db34d183116281eaba64f63504a5747a1bdb82a5b7ac41f74344a64263a4

Observation 3245ba21-ac48-4dea-8810-0c24557492e3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.122814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.659462Z digest=sha256:01fd31b18aa58f32d438ea78746f5dcfa4535de0dc600f6d6244a176c9ff403a

Observation 51bc7f0f-8e6b-4b00-86cb-ec6946077897 · outbound

This paper cites Guidelines: 27 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation Guidelines: 27 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:35:21.107264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:35:20.663843Z digest=sha256:b92c34947f49aa1d9b02746e1ce16f1cf974f808f7ced18fa800755d88038c3c

Pith citing papers

No inbound Pith citation observations are available.