Pith. sign in

Paper Citation Record · LEDGER

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

As of 9 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2507.09334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09334 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:03.566081Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f48ccb24-04c1-4af9-ab85-115adccb60b5 · outbound

This paper cites GPT-4 Technical Report.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.124609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.124609Z digest=sha256:ff34d13573b90ef44c776f1678e4a6f7468e6e3a1f2895c15df647a394a3b263

Observation c97d3428-4d4b-4ca9-bfc9-53a764938a49 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.157375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.157375Z digest=sha256:f89d8d104645471e00d6158aea394f2b56389e6df3ca207d23c489433991db4a

Observation 451f4ec4-0adb-4693-b0e8-cf0c196b7a00 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.223606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.223606Z digest=sha256:2b2570290e0e6b9a162f56df0676a227c6d9aba211b54051627b712fa1547262

Observation 5ed70302-e3d8-48dd-b8b9-60d04bc948aa · outbound

This paper cites Token Merging: Your ViT But Faster.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.302019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.302019Z digest=sha256:a7df2f1f508fccbfbdcf2347eb13bba0cb43d862680b0fff0bf0c849150c23b8

Observation 3b25b53c-588f-4af3-9966-29ef4776d0ad · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.364155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.364155Z digest=sha256:d00ba6808789f1857ae966172cee55d6c55e4ea86629e7c18054d9d8e6ddb14a

Observation e4cbb8fe-1c56-4b2f-8079-d6820a086803 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.422074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.422074Z digest=sha256:bdfe0305275bf7e2719fd48159ef122fe8a151ad2137333f463a89d2fec49e69

Observation 772f2383-4452-4096-8cc5-fa90cab7567e · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.504730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.504730Z digest=sha256:2def9132c395304b3f6378d69b3ee74e003260f0defc6c19a9b4da5a13ef73eb

Observation 0e65442a-8ef8-44e2-86c1-fa56cf9150ee · outbound

This paper cites Efficient Large Multi-modal Models via Visual Context Compression.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.561451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.561451Z digest=sha256:9647810110a0598fcc116595096b4e067ed288803a0f95c9611cdd8b5170e5f2

Observation b05cc5d5-aee2-4d4d-885f-6acb86921161 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.638753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.638753Z digest=sha256:b62a1f9b27062a48280bda8dd0ce26aa2c4de3cebe122d4f42925c0f85e374ab

Observation 015c0e41-aed1-4605-bda8-030c0e18aaa1 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.697380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.697380Z digest=sha256:08519108d981114b2be3c4ecbbb0ca3955dfa224813c66c0a0049a66681067a2

Observation 07554a33-1893-4957-9184-ecb0773e65cf · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.759815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.759815Z digest=sha256:67be67eaa48e3acee3d5e6637ff41030771e9ec31a1e135626010ef59efffd33

Observation 3c7f5ec5-5aab-427a-bf75-90febf63009d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.817271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.817271Z digest=sha256:b0d6be02833cc7c52aa62834e97dff9cdeaae69640bb676bfa967bff01a972e3

Observation 428ab502-aaaf-46c8-9e4d-2c368183afa6 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Grounded 3D-LLM with Referent Tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.856085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.856085Z digest=sha256:a266eee6bf00f10d4d031132b57256163203440119a70b7fbcaefc496ed1d331

Observation 58363efd-d264-40bf-8056-4fe131cdcc46 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.914716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.914716Z digest=sha256:da5bf87ea5d55452f94ac2ea64d00273be2411138bc5f88579f7e26d7de43305

Observation 7854f2f5-0054-4eab-884b-f903fa972108 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.972389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.972389Z digest=sha256:c4b56dbf427ef784ce6089dee8fb8871f42c34564dab86c1c09e4e2641b5d7f3

Observation b81c9a9c-0bd5-4dd9-aa3b-85195db93ebe · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.036403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.036403Z digest=sha256:232a094d867237d1a5e77d9483f0e3ed3cc5fd0e29119a6fd55782405f4957e2

Observation fa1e8111-6433-4fed-b898-564526b8756d · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.100588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.100588Z digest=sha256:dcb05c822dec19d19dbfdcd21455ba0b911189556e3bbfc0c351e437bd79e388

Observation fc7c0d4e-de84-4322-8dfb-1d57500b0b96 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.151702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.151702Z digest=sha256:44f4cbe8957100fa5f677349273d2fbaf977a8e900f8535c5d4d475d733da466

Observation c344f0a6-47bd-4900-b813-e89346807cdc · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.210209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.210209Z digest=sha256:80d4f6784d43ea5431a32c2b9a6f58a707ed182d2bdbce3d7ad345132b74e24e

Observation 9261afb5-5d0a-4769-93d5-a5c5a9b0cb68 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.260129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.260129Z digest=sha256:c0da9707db0f5d1191fb6735a73778af7aeef11ddfc614daaa670935f86d7724

Observation 68ae865f-2237-4e64-b034-6bb324ef0489 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.310302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.310302Z digest=sha256:7e5b965f42fb78d4207bf1f77e3abe0fc35eee33e110f164b19b2bb5880e8006

Observation 18c0dbd6-b05f-4c15-a4df-c38a6f616566 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.359288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.359288Z digest=sha256:aaf818349283be059343db7a70a55c4ef9349706603a6154f1b9dcabe6b47da1

Observation 11b29fae-b0ba-4098-9893-4791cfeeef98 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding ImageBind-LLM: Multi-modality Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.427373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.427373Z digest=sha256:492efdd17e618504caa656fd03466398f730cfc580b3bbe6c8f1fee18b57e5d6

Observation dac39ce8-5bc7-49d5-bca1-aad175bc498d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.538187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.538187Z digest=sha256:c556520aea77299de519e7987a4f787148f4f47467fe09242b79d7aa19ddbb7c

Observation 12dbf14c-b7ad-45d6-a642-dda8bf02b171 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.785935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.785935Z digest=sha256:87454ca65bddd55b3aacf6e43ba34b231ca183ab20fc3582a752581d4e42587a

Observation d3b9fa8c-0657-4b68-ac8d-f777b35998c1 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.896886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.896886Z digest=sha256:1b367c4f0353f6cac10b002c51f4253edd5915c8b0b84ca9874cf68428ec4d88

Observation ee98cb67-0a3a-478b-9812-a7ad333c6f60 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.333211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:02:59.961573Z digest=sha256:308bdcade33747bc9e10412c69f6e777808e935632e4afdf312ffed531467afe

Observation 3eb360f5-3190-4106-8bdc-1133acb30210 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.981076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.981076Z digest=sha256:44b5fed00883c71789238cac2d0bb020d76aa2c2718dc1377b729343500b0d53

Observation d929c75e-a43d-4e5b-a955-1e35386283a2 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.320759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:00.100464Z digest=sha256:6188b641d0221c6e28cd8b1348c69a22e8973891d1b4adec5bbb8ca23f6f86b8

Observation 14f17569-c587-46af-bf64-25ab994e7398 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.188535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.188535Z digest=sha256:f6201562d1c846cc01869136e17a394ca4d4b1e5a7c683bfb7fc104fcd58e5f3

Observation 6d6927c8-a06d-4f62-82f4-2feec4d9382f · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding An Embodied Generalist Agent in 3D World

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.275263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.275263Z digest=sha256:f9dba261bac5c414fdf5e9828fcfaf5661737d87c670be5416ddbb1e7c2e2da9

Observation 31965077-7e67-48fe-aaa2-ced25d9f7d75 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.312278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:00.390648Z digest=sha256:e8ff8584c71d1b240177da85640e66966662e38a557efcfcfb799f80aa015438

Observation e7bd4f78-18c4-45ac-8ee2-59ab9ed91133 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.515248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.515248Z digest=sha256:2864befef9122b1f9d81753658df1e7a8d5f20461c71f24c055bd10af791652e

Observation 943b7901-d968-4229-9a6e-ea74284ac571 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.299584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:00.648257Z digest=sha256:4224195916db3cea26cc36a18f31e479cc3f12fc0c00a63e926e5b7e93260cef

Observation 4d279c18-dd29-4371-a6dd-b9afe731d86f · outbound

This paper cites What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.712937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.712937Z digest=sha256:bd89ab83f6f919cb3919058bfe5896100058d9908929c9dfc5761ef900d2b9e1

Observation 263f6650-c2bf-406b-bdf2-e5da196521b9 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.291211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:00.796668Z digest=sha256:f0a3357782fea46afbb4f0962b41f14e994625720aad7ba4c84dd3b6b21a4014

Observation a944f717-4782-49d0-8b83-340b38a70338 · outbound

This paper cites Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.870564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.870564Z digest=sha256:def853c7398298659ab8d88804ab72d98f826a9e2733cb71c97544be93b88b9e

Observation 7346c2f8-6b98-4f16-8a59-a2848d296f0d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:00.965145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:00.965145Z digest=sha256:75b3bf7e5548418c19216007516d5a3ae13fbd5aa7edd4edbdff08e3a8bdcbac

Observation 58936284-d05d-4d50-9943-35cc6971894e · outbound

This paper cites RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.186306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.186306Z digest=sha256:f83b2d5771cb0560c55615ff4dc6d7ee470f35e76073702a230fc51833ffa05b

Observation 0c424190-ddff-486c-a6ef-16aaf8b4e4ef · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.298449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.298449Z digest=sha256:766238932cb174fcc9e4a048a2e246ad102c0e2f35c65c4d0592f1ed38b1e09a

Observation 7d46be39-c1eb-43b0-a664-74208764c257 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.384970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.384970Z digest=sha256:2215a1a74ed44e8147fc40ea7269cf39ce3545504c0fbea024e340a369898991

Observation ba871690-ba73-4293-b9cb-ef2c964ae505 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.459589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.459589Z digest=sha256:62a05c7485909befc21c5b824bbf77e04fa464d6e9d40dc73e418cae8ec4af59

Observation 8b21d098-658d-4f82-8300-bd55089c689f · outbound

This paper cites Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.514560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.514560Z digest=sha256:b3d796baff985fd8a9fbd9db72171d3b16853734b2b833f456ef31513fdd7bcd

Observation d5ab953b-5a41-4c58-abc4-31cbd716dbe0 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.268643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:01.598313Z digest=sha256:244c4360be473fd80d022b99fcec2ff214fb271e0c0c7963a6b3ae6dc84ef2ba

Observation 4209fa98-41a5-4b90-8146-f2305a0afdd0 · outbound

This paper cites A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.736901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.736901Z digest=sha256:ab92ecc00431fbbcec2ef9e3f2e43bc0b730fd9725cd68ab0b71897b22d555fe

Observation 63d6651f-e8da-4f2f-b415-e3cce06f15ba · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.862985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.862985Z digest=sha256:5383dfe696c09eb88b3d23f0dbcc3e61d5aa5ab298dd437220bd76eb0b520924

Observation aaad7d22-b811-41a8-bde0-57f758a96a4a · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.260407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:01.958274Z digest=sha256:17b4787c54b68a2fedd0bb47f26ee9e7c94f6d9bee287d9f9fc042627e68aa80

Observation c4dfc515-b507-46d5-bb4a-5c2bee356090 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.096446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.096446Z digest=sha256:cdd4d5545bb05178e9815b9eea88384a7a0f65ae2b1795b9abda4bd181b9b164

Observation 8e495855-3c2f-4e86-871d-c49d21bf92f7 · outbound

This paper cites Decoupled Weight Decay Regularization.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Decoupled Weight Decay Regularization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.177529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.177529Z digest=sha256:76920ff3935ba211b7f4f0f48bcd4a12ad19d3f81fa5ee020b7536458961cd04

Observation c05b8ad3-7462-4d43-a348-e4e2897b91a0 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.275383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.275383Z digest=sha256:4ceadc0edbe3edbfcfb16fcdd54c3b193d08a77ebc466ed70d039d3e46a02d30

Observation cc8d66a7-de7c-48f2-ab20-092c5b8e0252 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.247401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:02.387663Z digest=sha256:1cdf8591af7e92cb2dd00c235096d0bf5a9cecfaf117ca40a098165615c940ed

Observation c6747194-ffa5-4fad-a468-8e298098a2ef · outbound

This paper cites DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.529236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.529236Z digest=sha256:3d8036409ce84d3a578b09bc2d5d8cd0b1d8091aea906e72856980f5dcc8d9c1

Observation b4fb18f8-b5c6-409b-bd09-a7820b23a8e8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.638809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.638809Z digest=sha256:b65871daecf4817d89672f82c9e5cdc6b263d5361ebe6ee81b6f06ebbe3aa0ca

Observation aa459827-2401-489d-9602-c3cc74856acb · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.775854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.775854Z digest=sha256:2bb759edb7a938360974b205a17c4a3b6edffb2cc4923e100aeeb927c49763e1

Observation f00de20a-bd16-44f9-a5d5-e95b21d78b99 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.234376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:02.909116Z digest=sha256:8515baa3b2aa5f7c66e80d9598bd3404f4bc6899d2b3c88e1ee26df42c17d05f

Observation beb01344-99fc-4f67-8d51-9643a7b40d2b · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.225873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:02.982828Z digest=sha256:779036c8ad4a8ab90b8b4efd137131a8b7680f8d9ee114e9905e08d965aa66f1

Observation 10c1911f-5be5-43f9-a8e1-23036f06c7d8 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.159380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.159380Z digest=sha256:21e3cc8411f2df094cb137e20cb26ea80ec591c84d2154a4f416afd824da735f

Observation 85fc4b70-9b34-4241-9f9d-26410c2934b0 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.419032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.419032Z digest=sha256:7df85e5fcf71ab23af985dbf5b66c2f446eea8d91f06a1e2e3d7b82727161247

Observation d3d5e37a-1175-423d-9e07-11cfb0186ccc · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.196313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.476300Z digest=sha256:f422cb8e0645b9ba73ac6e2000f862f4cfee078347745c57246d982bb1469dd8

Observation 1d2efe2b-79bc-4dd7-8171-3abf4037646f · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.479044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.479044Z digest=sha256:f75ede22977b831115ef781b0b6aefb6628d3da3c88cf46b35d1cd826b3e7892

Observation 826bb000-010a-4fe9-af72-5ff6515c331f · outbound

This paper cites CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.482002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.482002Z digest=sha256:a9d7480c947bad0009b46a77cf5472068c4732921ea2d26a43c16fbb7a36260a

Observation f625af16-3a32-43a3-b587-62a26cd8d262 · outbound

This paper cites Advances in neural information processing systems 34 (2021), 13937–13949.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Advances in neural information processing systems 34 (2021), 13937–13949

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.207795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.473632Z digest=sha256:fd891b3a82d130e7540b72fb99f57580b4edf0311071aded0af5e815d92597f6

Observation 8e66426c-04bb-4b6e-ba8d-17fa51174131 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.488481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.488481Z digest=sha256:3caf5b26acb3bdf190eaea53b20bfc3343bf304d5f60be0fd023f391e072390a

Observation fd9a1140-07e0-493f-9aba-660ef3bda262 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.491561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.491561Z digest=sha256:5d732bbc0a9adb7e00089ca649c2f7425624f4868801a06a750b104d557336be

Observation 8cd510ac-b256-40bb-8011-9c7fe44dc19e · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.494966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.494966Z digest=sha256:2416213c88e91f2d8e8457fb8348e58a604f488e6ab3de2bd1172c9ce8f89331

Observation a7e98b14-bf4d-4ba0-b598-48ecc44be977 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.485061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.485061Z digest=sha256:c1605aaa016de6bbff0fae682ae7f8410afe598c04225e5f6d46556a0515cd0e

Observation 12439054-5c02-4e68-95df-1195896b9604 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.163793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.500422Z digest=sha256:036f2bf5eabba64d6187cfc8077697a778096b12293c008ba0cee7981effae4d

Observation 31a64dc9-6dae-47d2-8735-d4d256d72e78 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.506443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.506443Z digest=sha256:e8ba2f27f6e40d48ead19b9f88f5b286bea18fc5d21fd07b93f45054ab551801

Observation 1c0768b6-6793-4731-a6e9-a2dbe0c99a8a · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.509381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.509381Z digest=sha256:18b849db7c52bb878574004af519bd8b14eae5d9676ffe9cee91a26c7409f01e

Observation 6db3fe3b-5dfe-4b00-85e9-0b3513d74123 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.171716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.497751Z digest=sha256:215f555e334b4fcee5f3f4319b729d366f006381cffe7ad8680904fafc40d944

Observation b402ff32-e7ce-435a-9dfc-b27fa2980217 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.135222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.517025Z digest=sha256:c7373c0f9304d64e768631cda639492076e350f222fc5bdf9d6d98e9faea3607

Observation bf08dcb9-4566-433e-b52b-f72d9f749cab · outbound

This paper cites [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.503424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.503424Z digest=sha256:5d52e0723f36506bcf4e64c3a2a3a8d2c7d012d6e6f0474c28d38a01dcfcb202

Observation 710f3385-21c6-4bd4-b355-9f7a54843c1a · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.521866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.521866Z digest=sha256:c2a756c959d8c8d7613430578f6dd22af1c78781b545e48a530b01bd63017a0e

Observation 9d18b4bf-9bd0-4340-9b6c-470f3b9efc82 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.524477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.524477Z digest=sha256:01a3482a8e0c251a11bd6c836cf86b73925205887754509331f1d588f51cced2

Observation 093a4a1a-8679-4475-b912-ffb8858f64e3 · outbound

This paper cites In Findings of the Association for Computational Linguistics: NAACL.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding In Findings of the Association for Computational Linguistics: NAACL

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:04.144483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.511790Z digest=sha256:5d6465c3a288ee297042d5fdce17f72c0c56f94127042cc33545287996eb2bbb

Observation 787ddb93-6028-4bca-8479-89bbd6852216 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.514224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.514224Z digest=sha256:a08b267d445a3d0fa8b764b44c70746830d3207c594de9a73305620577012e58

Observation 1643fde7-2768-4b46-bc1c-8a6a60bbdb1c · outbound

This paper cites 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.534576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.534576Z digest=sha256:0299c2f90f5c82ac8a22ca79b2bba848d1ca5bea2818e8cbaa9d8f5aefb96a5f

Observation 69cbbbb6-fb80-4d49-9b2c-da94686edd61 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.126834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.519321Z digest=sha256:ac78e2cc7d3003da5626faa48907e1f10bca605ec4fc709157e42b6514dfed22

Observation 4583d938-0a81-440b-aa6e-02f54282bb68 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.097685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.542291Z digest=sha256:24c3be516700fef5af9eaf2d26311247447c0aaefffee9ef8c5ecd95775b433d

Observation 35ef04ae-baf7-481c-b4bf-7ee6ea162144 · outbound

This paper cites Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.544777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.544777Z digest=sha256:9d08d8038f4d7a2407c272583c187c5de65e4c9004a4f19d9871dec2342d2f4c

Observation c5403498-287f-4bc9-8c20-0ef179aba5b5 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.118907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.527088Z digest=sha256:0fa4e8e5c72cfc1e063ed34326ab902deb2d1bf7efacd95079dc2bf48e159f2d

Observation 84c13542-3f1b-49c4-b010-e4b1694cf24c · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.529576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.529576Z digest=sha256:6419d336ac6e0bb04474059ba761f9a7c1de588599ab68c5a767264f49752025

Observation 3b558d08-a2d2-449c-80fc-1505b8cbf49d · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.111097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.532227Z digest=sha256:efc8a24751b48d505c7c0c075241f4b7614df7241bc7a39f660ab492c9248bd4

Observation b344fd13-b7d1-46a2-af72-80b526c10748 · outbound

This paper cites AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.555597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.555597Z digest=sha256:e6bc44d553171c5b29cc8dec1ecce8c669df026769db2cdf60db214d1a08c481

Observation ebf472a7-2538-4504-9b01-4bad846c8f82 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.537492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.537492Z digest=sha256:c26b18ae28da31aa8c76607324e76db9fc35ee2e99b82c7d83d6e34de3e671fb

Observation 098af109-893c-4b64-a642-5475658e6011 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.539721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.539721Z digest=sha256:888fdac735afa8e4a8733b196d11cd210565c307c5582ec986a3020b5ade5bc8

Observation 688744cc-7d1a-49a3-a462-de351c6c7125 · outbound

This paper cites an unresolved cited work.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:03:04.088401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:03:03.563680Z digest=sha256:301a5067976000bfa8b3b90bf8bfe23bd24429c5ac41fdba32ca0abcda8059cd

Observation bded14ac-604b-402d-8a64-1dfa3143d598 · outbound

This paper cites A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.547382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.547382Z digest=sha256:bf8be8396afc8af346a4f59959c23a1cc38dcede1b52fbc3b1b5218a7b661621

Observation 87499f39-b287-4cd0-b263-0a6eadf54e95 · outbound

This paper cites Dynamic Diffusion Transformer.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Dynamic Diffusion Transformer

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.550076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.550076Z digest=sha256:abc568afc030f14c76475e887daed50a79c577904b6e813d1632cf8c41d3e3f1

Observation 6a42d758-339c-4ce3-842c-429e342f8015 · outbound

This paper cites Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.553163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.553163Z digest=sha256:8dcb89055a9473f2b1b2425dd94658604f9e27cdc313967029b5d6ee49c8d3c2

Observation 5276eb9e-4006-4c13-852f-11657fbe3d35 · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Uni3D: Exploring Unified 3D Representation at Scale

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.558444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.558444Z digest=sha256:5e0bf937c6311796bf970cebe6b938e1ddf593e2e725af6ad34611564ef44be2

Observation 88aba4f5-63ef-40eb-bcbc-6f0cab7a4d4c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.561057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.561057Z digest=sha256:11421d044c200a52d025dc919140722da1068000ca7c6ca17be6824a39af0787

Observation b6237571-f163-467e-89ff-16f506e8ad34 · outbound

This paper cites ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.566081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.566081Z digest=sha256:ce38002c0372aaaeead1be7a4b167428c0445251b159d8c9b9a0e7673173ff8e

Observation ae7110bb-e2d5-42b2-8559-7285a281dae2 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 11 (2021), 7436–7456.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 11 (2021), 7436–7456

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:59.680523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:59.680523Z digest=sha256:91ae29fb77f07b1a8c0c21f3943b000b630ffc803691828d63b05066e8b85ecb

Observation 65fb1d5d-d32b-4f50-9781-e7455cb29b16 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.029960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.029960Z digest=sha256:6d8657b824c104d50c80ab5d031a3e27c47c38c4dc08f95164a1734a99e26af3

Observation 79870942-2f3b-4997-919b-9adaf1f622a7 · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.297222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.297222Z digest=sha256:0e060d69780ce6506da56d704741768c332fa8ec337fb91cd93b23a06dcf8dcb

Pith citing papers

No inbound Pith citation observations are available.