Pith. sign in

Paper Citation Record · LEDGER

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

As of 11 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2501.07819.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07819 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:39:12.809687Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:14.473013Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:45:24.456612Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86537c9f-e7a5-4eb3-9b3e-f488dcd4dcec · outbound

This paper cites Chatgpt: Optimizing language models for dialogue,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Chatgpt: Optimizing language models for dialogue,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.460397Z digest=sha256:aba2372257e4d940fc775639714ea09334715f99793d2c59d5facb4d95b83117

Observation 9c87d6d7-194d-4d0b-bc4c-e48f9b4be5f9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.465858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.465858Z digest=sha256:cac3b0a0d3c500b8e8ebbc974e4585060e4b8b723fe464977e1a093bb25fa53c

Observation 4a7c1dc8-ba53-4303-b5b0-eb004f24549e · outbound

This paper cites The Llama 3 Herd of Models.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.471078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.471078Z digest=sha256:90b8a6e2bc49a9c23c3ee6f86bf098ab6294519387c15b619fd4b209afd704ed

Observation 10c6979b-6176-466d-acbd-8629b503412b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.477063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.477063Z digest=sha256:972233382f0b30886b35e667aeafee43c7fc567c4e907b8c43f997925ab1b4ae

Observation 00713a15-c5e2-4d3d-83c9-7a4288d67de6 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.970572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.483252Z digest=sha256:14e97cb88a79a49ff08e575d540b6bcb263c9fa58310b731d7683cbe8b49ded8

Observation 4ce9ec64-8924-40b1-b6f9-e35cca63b96b · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.489257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.489257Z digest=sha256:919544ee30db7a38ef0358d19bfe8811ed5823021bd74c11cc80452d8804f294

Observation 25316d5b-6cd3-4a81-9267-ad2b14393f95 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Improved Baselines with Visual Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.495609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.495609Z digest=sha256:98256975c3543f4147d3861cae55f8cba6eba3693e1e28888188a0ea89853aac

Observation 54232287-b629-49b0-9e68-f757043a9a7e · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.501353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.501353Z digest=sha256:e6c0e6fa4e197b17c60169d9a4a76e0e3f08e966fe2d9a530c9e094c2f10e638

Observation 882b9a20-4a37-4b81-b3f2-b8c0f70adf50 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Scanqa: 3d question answering for spatial scene understanding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.953513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.506248Z digest=sha256:96bb9512057a1b97f587d60faf2f8b0451102823d33997f8f4c3b6a954cab56d

Observation a9548686-ba0b-4b81-b9d0-bb565b686606 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.511670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.511670Z digest=sha256:bcf3e9272a940f133811e3c9a3339ec8ade931d3621ec6874d0e36e5b4180756

Observation 0db83eb1-3ce6-41d2-bf9c-925ac1f187f3 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.517167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.517167Z digest=sha256:c3d5ece918d06b6ce05e514461ba480ef31393a75bb3d9d96f8e7201ada0cf26

Observation 911bfe9a-57be-4a5c-8eda-440d9954f8d6 · outbound

This paper cites Deep hough voting for 3d object detection in point clouds,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Deep hough voting for 3d object detection in point clouds,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.935076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.522660Z digest=sha256:9d3e817ff9efbd050afc11536893da19826367a243b22f551a0dbf0d989c7b9a

Observation a4988914-e146-436c-984a-5b43be4ccc57 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding 3d-llm: Injecting the 3d world into large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.527226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.527226Z digest=sha256:e6266db7fff4633327a1ec4d144c81b0bc8c39b35c156fc06a3e25e7f04bbe8e

Observation ee3aac66-f8e7-40e5-8e73-9779bf3879af · outbound

This paper cites Segment anything,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Segment anything,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.907356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.533751Z digest=sha256:7c3b258245a3157494041159219f453361ca11c14b2bda21e32f9ea59355ffe0

Observation 75a1cecc-f363-4e4d-871c-366da6dbc5d8 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Masked-attention mask transformer for universal image segmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.890127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.539498Z digest=sha256:30ab0b769cfd80524bb5c1326b3c33f71927d153bfc27fbb369c4773cee98717

Observation abb9dd29-7bf7-41cf-accf-bb0bdf76d39d · outbound

This paper cites Training language models to follow instructions with human feedback,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Training language models to follow instructions with human feedback,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.544476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.544476Z digest=sha256:70f40ab7fdea83dab9812ce56f2f031e2246fce37b85d6fbe8143c600094002d

Observation 6e4a6011-888d-492b-9274-6208796adb90 · outbound

This paper cites An end-to-end transformer model for 3d object detection,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding An end-to-end transformer model for 3d object detection,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.860073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.549839Z digest=sha256:b9a6deadaf2da51254f8a4b131f86cbe850b554865f20c05f9dd6b723c1b74e9

Observation 67d4dd9f-7124-46bb-9445-7cb3bbb22906 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.843023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.555589Z digest=sha256:ee99956548803e1bd18665570dcf65a7df1a00efd15e3c8637263fff7a1e1040

Observation 40b8c5a7-a891-41b9-993e-4fe940c5c8e7 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.561196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.561196Z digest=sha256:d63404c9af036b36c40be25b13a90cb7dc26f7a1ddd83e102a3146622b0d0698

Observation 30ec9318-0d55-4ac6-94ed-7f53fb0d6833 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.566721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.566721Z digest=sha256:a4a94d031c6e909a96a9978859d83d6fdb5f572b669ed40c9881431d1476e255

Observation 76cad131-e458-4a6f-a361-83fd307fa6a0 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Rio: 3d object instance re-localization in changing indoor environments,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.806435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.571422Z digest=sha256:116d401661ca7bad40c2c10ac61bf5b1122a3206c6fd12534c82dbe21a67aea4

Observation cc31d721-d77e-4517-a99e-e737e9b2a136 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.790129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.576311Z digest=sha256:41a6b77d6ba6d242e1ee7eda465f96391a5068ada4d77047c050b7627bf4f6a3

Observation 0ae896f7-fc7f-4c93-91e3-2cbb38a8dd87 · outbound

This paper cites Improving language understanding by generative pre-training,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Improving language understanding by generative pre-training,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.581307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.581307Z digest=sha256:a0e8e95b4c66a541b080fa7b1582b2036430e281542b8245989a4b176c8b388e

Observation da4c0956-d103-42e3-9ca8-86c7b227b352 · outbound

This paper cites Attention is all you need,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Attention is all you need,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.586551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.586551Z digest=sha256:c808c0119481f1c3020b68fba489e996b341dbaaa0cd13d5ad3bbcdfb8f0ce0e

Observation f5dbffb5-a6f5-4d57-ad0f-4e15ee6b49e7 · outbound

This paper cites Language Models are Few-Shot Learners.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Language Models are Few-Shot Learners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.591535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.591535Z digest=sha256:d99fbcf222b6d5346b6c0c3110760bc9547aed6357a1e28b976e4fd0e928c909

Observation cb4d0a92-6165-4aeb-9219-f20695ae9012 · outbound

This paper cites Palm: Scaling IEEE TRANSACTIONS ON MULTIMEDIA 12 language modeling with pathways,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Palm: Scaling IEEE TRANSACTIONS ON MULTIMEDIA 12 language modeling with pathways,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.751537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.596464Z digest=sha256:fa8eee631e556552c09626605a8fbf07821dda4b259f72bbe8b2c5154cf31bae

Observation 1fd794bd-8bd1-4963-8aed-041f62183640 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Baichuan 2: Open Large-scale Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.601023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.601023Z digest=sha256:7ee5592483e7f910082328f682ceaa17cac7768fa0e4cd069e9c7d6c762445a3

Observation 1f536e8c-4164-41ed-8283-950e810934fa · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Flamingo: a visual language model for few-shot learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.607083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.607083Z digest=sha256:f19ef417e4ff31dd337e4ed79279b0fb7c34b5bc73fb1d8af7fca7a692c1ab7b

Observation caaab393-4c06-4497-b3f3-6061c181e667 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.612153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.612153Z digest=sha256:172fefa4b022043c20164b9bbb618e778513f5bf60e43767f66620e831bef10e

Observation 5c7cb602-e538-4525-b269-1e726bc37eeb · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding OneLLM: One Framework to Align All Modalities with Language

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.617254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.617254Z digest=sha256:48e3c1c4ac3e22379fd4f13e7bbcf7333498c6da01f1fe35dff2c8bfd104b98e

Observation c2b8a48b-ad0d-4ca3-bfdb-530e479272af · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Scanrefer: 3d object localization in rgb-d scans using natural language,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.725263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.621969Z digest=sha256:23c16e1fa6de66194f4b28d8c2e2ebab2ef54b2799261c41b6685ca98f3478a5

Observation 911dd5a4-4baf-46cb-b6c8-95473bfd0bbf · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Scan2cap: Context-aware dense captioning in rgb-d scans,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.709117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.626561Z digest=sha256:82bc2ff3cb0fa5ffb3518f74fc2805829d47b0e3e2a6176f9f787629e6c6eddf

Observation c5b17490-9bf3-4fdb-a87a-868bbcd0b204 · outbound

This paper cites Proposalcontrast: Unsupervised pre-training for lidar-based 3d object detection,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Proposalcontrast: Unsupervised pre-training for lidar-based 3d object detection,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.692413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.631225Z digest=sha256:48db73fe587570a527b5eaf621ba57b4e88d812818528d7aa88966be727f1b9a

Observation 6b07d920-cd08-4cb2-a785-23756e104843 · outbound

This paper cites Graph neural network and spatiotemporal transformer attention for 3d video object detection from point clouds,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Graph neural network and spatiotemporal transformer attention for 3d video object detection from point clouds,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.675655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.636279Z digest=sha256:574d6f7cc1039abf1f4e5aa3fb52352016a817d01dbcb4d56fa56c905ee35f83

Observation 8a1fc6ef-ad7c-4e90-85ed-171f1f181541 · outbound

This paper cites Is-fusion: Instance-scene collaborative fusion for multimodal 3d object detection,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Is-fusion: Instance-scene collaborative fusion for multimodal 3d object detection,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.658802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.641014Z digest=sha256:baf82ca03fd80a14959ab6f46b88850f8d1224d3225afd34c09a3504baedd0be

Observation 5d0485a7-4762-46db-bb92-6ca5cd33ec80 · outbound

This paper cites Lsk3dnet: Towards effective and efficient 3d perception with large sparse kernels,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Lsk3dnet: Towards effective and efficient 3d perception with large sparse kernels,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.641281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.645802Z digest=sha256:e658bcb138054ab919db09bea469f88db46e2ee3ede356df773d53bfa6cb1348

Observation 8944b8aa-cb8d-408c-a235-f165081895d1 · outbound

This paper cites Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.650267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.650267Z digest=sha256:64f31d17697f6884f26be611bafc7b1ffde10f2299c4fb0944fa6a5d5fe52fc6

Observation fe59aff1-0f92-4177-867f-3237c083a435 · outbound

This paper cites Clustering based point cloud representation learning for 3d analysis,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Clustering based point cloud representation learning for 3d analysis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.622864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.655833Z digest=sha256:9633ed4523796ad9d348ff016621aba23fd8f1a71e70c4f88a282b047a35e02f

Observation cf5e707a-e2e3-4ed6-8816-3c514a1cde6a · outbound

This paper cites Weakly supervised 3d object detection from lidar point cloud,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Weakly supervised 3d object detection from lidar point cloud,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.603492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.660428Z digest=sha256:e84132c183fdde9115faa79384b608ef52d35f0351ca2323b9dd6391f6c69ac9

Observation eb0cf1c9-3fa3-4b9b-9c30-52b079489d09 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.665081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.665081Z digest=sha256:0f3d3ca3daaa6553e5b8a6782974f0fcae5366b6f700cf04741301d87eed23cf

Observation b9690402-a39d-4b5e-b218-a55365a7d2d2 · outbound

This paper cites ShapeLLM: Universal 3D Object Understanding for Embodied Interaction.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding ShapeLLM: Universal 3D Object Understanding for Embodied Interaction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.669702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.669702Z digest=sha256:794de65bc44575b65e5099743bdb970bba32794454b540f6d3e05bc6d7c25f4f

Observation 199f1a3b-46a0-4550-9002-399de16f956c · outbound

This paper cites Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.586955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.674278Z digest=sha256:d6566c424dd4c5838f893735b68a9e661c1bd597c9d576a701b0dfc207034bac

Observation 7146b3f7-ccfe-4720-bd8b-5d861a27bb73 · outbound

This paper cites Openshape: Scaling up 3d shape representation towards open-world understanding,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Openshape: Scaling up 3d shape representation towards open-world understanding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.569847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.678905Z digest=sha256:c49c8c3aaabe514adffc1599aec6f46a651bb1ddaa4c49e30f38e05cef87af18

Observation 11682bad-c7ed-437e-a279-7bb85f95b4d8 · outbound

This paper cites Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.553376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.684822Z digest=sha256:3397caa817e0e2e4023da79f0b816eb20fd60397f526938222472f2109f73e43

Observation 84f41382-658e-4b23-9bbc-0f1133c29f1c · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.689677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.689677Z digest=sha256:565402bd3417f1ea58d3fc6489378a5a5016a165b233622718efe2647225e586

Observation 43aef20e-db89-4dff-8531-be3e8189dbf2 · outbound

This paper cites BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language Models.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.694961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.694961Z digest=sha256:588416b7a3b6ac9256c930a6e7816f9dd9df60b2b148f63f212388f70eb8ae2e

Observation 2961d74e-47ac-4eb1-aefa-c5cf3c4b02df · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.699950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.699950Z digest=sha256:5e059c9dac3f41c043b36879e69bc1116502c5c54cabfd07b2a10204f49a7f12

Observation e072ceeb-fb24-4ee5-8742-ac729be88d55 · outbound

This paper cites Visual instruction tuning,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Visual instruction tuning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.537382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.705345Z digest=sha256:bbb91422a86ece6e7523dbf513508b018d37a59f57b6bd2ffe974ec3a8f5cdcc

Observation 51a66ac0-39ab-465c-86b7-5954b17470f3 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.711385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.711385Z digest=sha256:aac231652cac5163dd8bc665c21e6fc280fd9353f2f2be663643129f1766d7bd

Observation 489413b1-2a61-445c-aeb1-fbb29118e12b · outbound

This paper cites Language models are unsupervised multitask learners,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Language models are unsupervised multitask learners,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.717610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.717610Z digest=sha256:f29f3f97689697bd63afa2ac2cfa8d608af0ce466659842ee0c868e774b8f852

Observation 27ae5591-3cc4-40d7-8fc5-4dfe36296ef3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Learning transferable visual models from natural language supervision,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.509451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.723595Z digest=sha256:14077c7ffefad71531f82c4590014ad014f08711dc011456ab46a8ac5d15f200

Observation b03ea6e5-2791-4900-a1b2-af686b0e5f11 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding 4d spatio-temporal convnets: Minkowski convolutional neural networks,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.729151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.729151Z digest=sha256:5c75148b695802809a5030751abd0fd51933266891c7dc6bc64c053d8d667d7e

Observation eb37209c-9864-44f0-96fd-5607582103c0 · outbound

This paper cites Point transformer,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Point transformer,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.477999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.734456Z digest=sha256:c07b77243f8ce676735cf31fb3ddc05e585fd1d8fff43e2e57b8e93383c5206f

Observation 9dbd8e82-1bb1-4f20-b1c0-e73f73c3c873 · outbound

This paper cites slam: Dense slam meets automatic differentiation,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding slam: Dense slam meets automatic differentiation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.459257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.739314Z digest=sha256:bf6b8e96e6183a3c46307a3ee8f2dd97b3cae7f6e97eb83467a217d9ae4e53c2

Observation 5287a31b-8269-4532-9396-18bd72f82627 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.744650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.744650Z digest=sha256:183fb76abb49d936b063f7f689b07143deae8926b5397fe2d1f2883ea442f2dd

Observation 92a66cc1-1d39-49e0-aa03-98ddec7f179b · outbound

This paper cites Context-aware alignment and mutual masking for 3d-language pre-training,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Context-aware alignment and mutual masking for 3d-language pre-training,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.440400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.750530Z digest=sha256:cc5b3da41a0df822f7577d0f27098065f4ce0c09bf90f0aece721b67afb1160e

Observation a7b91f80-1c30-452a-b690-2b3b0e639472 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.422876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.755718Z digest=sha256:28bc98685539c768b2a499877e910de51aee56b77b6fec36d20b7f048987a610

Observation 01d1941d-f9f0-4048-aa9b-c0a58889375c · outbound

This paper cites Cider: Consensus- based image description evaluation,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Cider: Consensus- based image description evaluation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.403734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.760661Z digest=sha256:4dfb18e1458696a48bdc3ded83c49f6975c969338a3932b09dd4d5287a048fd1

Observation 2888d537-fda7-40f5-806d-4849acac4193 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Bleu: a method for automatic evaluation of machine translation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.385071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.765569Z digest=sha256:df4d2a56c266cfaebd140bf2cd9078efc326fab658e9273cb2da26ade1350c75

Observation 5c93929f-3281-45a8-b720-fba181e57d35 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.366738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.770202Z digest=sha256:620b2534ee7d912c160b543fd77025fdae2a1ecb2494f36af58f6c4cc6015e69

Observation b33b90a1-ac3f-4f2d-9fc1-8d37aa5a8a1e · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Rouge: A package for automatic evaluation of summaries,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.774546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.774546Z digest=sha256:b61ed427d146f865f258bc2c1a496ed58575d824c7ebb02a57cb49f4ca01f7d2

Observation 98e26045-219e-4bfe-be49-e6b22eecc441 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Mask3d: Mask transformer for 3d semantic instance segmentation,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.779101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.779101Z digest=sha256:3803a70698adc35dc244534158ad0669407b2aa4d8fdb04c6cf6143e6f2e2af6

Observation 812c78b1-b1b7-4895-a80b-d5fd3c69dd3a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.783741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.783741Z digest=sha256:3c423894f96fc3b773f7bea0fece02c8fa697afbe0fc54d2f97f392d9fbf159c

Observation 836d19e6-d68b-4282-8b6c-9606a30d2992 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.315861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.788269Z digest=sha256:495dafa938fceff1617b69e067df08f6fbad808492eab9b2c73c729fb7d37488

Observation 9b6f0acb-dab6-4861-9b44-8b2625642794 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Scaling Instruction-Finetuned Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.793736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.793736Z digest=sha256:63898a67881dae8cf6f86aa3e5d665c03dd28a2de500d83c17a2208f9ecaa8e9

Observation e46bb015-5a37-4b73-94d5-2350e39a3335 · outbound

This paper cites Decoupled Weight Decay Regularization.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Decoupled Weight Decay Regularization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.798622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.798622Z digest=sha256:dbe33af59a328a84936859ebd7c456240695de605c726fb81c4d6048fe450f6c

Observation d894135c-e36f-4193-8e0e-be7883d1cb40 · outbound

This paper cites V otenet: A deep learning label fusion method for multi-atlas segmentation,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding V otenet: A deep learning label fusion method for multi-atlas segmentation,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.297387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.804140Z digest=sha256:fd9e1fbf25ffb4aa82b74a4ebeaef2a4b23aa18224eda2dc8b08f0c7a0d581d4

Observation 87d23797-1061-4368-a0bb-6ba93a11f832 · outbound

This paper cites Deep modular co-attention networks for visual question answering,.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding Deep modular co-attention networks for visual question answering,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:39:13.278368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:39:12.809687Z digest=sha256:f93942b8e311169e1f6fe192f68736fd80802e0435f2c00b81a50129e2ed08c9

Pith citing papers

Observation 58d23528-1edd-4483-a4f0-a17c2c013d05 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.473013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.473013Z digest=sha256:88918384b6d3c4a3b2edfe915e45d3a78ec587b2005d372c1fd461095dcaa90c

Observation 8e4e8dc4-e183-471f-a647-449bd00eab2b · inbound

PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets cites this paper.

PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:18:05.122134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:18:05.122134Z digest=sha256:8097c2a08533689991e619fc867da8e09c785c216f74f6e28c85b9b8658e3b8c

Observation cc07f7e3-e66e-40f2-bf53-9bebee77e807 · inbound

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models cites this paper.

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:45:24.459884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T22:43:16.761970Z digest=sha256:8b96fa5cb5447e4f6ff6a2cf817de3545c4ac2d04ddd1f6e0d10e29265aef770