Pith. sign in

Paper Citation Record · LEDGER

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2508.17595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17595 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:09.859176Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13a425ea-cabe-4d18-9c8f-d328ba2295ec · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.425566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.724943Z digest=sha256:9aa37ae3c72d8f99662805b5a301dcd1ba701d5a67f5ca661b12b2dd0449aa00

Observation a693fb44-c00b-4473-bdee-174a06ca1aae · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision-language models.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Spatial- rgpt: Grounded spatial reasoning in vision-language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.407806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.730624Z digest=sha256:deff0662ed405aa087e8b2775d04051a16f561eaacc3f70ba861bd22f5639c54

Observation 5f32692a-f101-4fee-a90b-a1b6bcbbc49b · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.736041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.736041Z digest=sha256:a0d89aa893f8a0a1143fde451645c295ad2cbf77f682232544f6ac3ca0972465

Observation e0418e18-c407-440a-a0df-efafedc7d4d9 · outbound

This paper cites Fusemoe: Mixture-of-experts transformers for flexi- modal fusion.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Fusemoe: Mixture-of-experts transformers for flexi- modal fusion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.387153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.741439Z digest=sha256:e923ae33322fd2fdabbe2ffda10291f3d33899406f109edcb40ca138521400a3

Observation 53d91718-43ab-4106-908a-f24f64b5211e · outbound

This paper cites Multiply: A multisensory object- centric embodied large language model in 3d world.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Multiply: A multisensory object- centric embodied large language model in 3d world

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.361885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.746747Z digest=sha256:da00f4bf0b5fb7d49a8e8dc539af80dd6cababe7c3d6c6f2207665dbe05927b0

Observation ee59531c-505a-474a-9225-12b090e87e3a · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.344898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.751482Z digest=sha256:0488624341eed6aa405b2d71b50f38afcf7f41c77d9d44b92e3c081d17fef6f0

Observation 715c92b2-a6a4-4655-b77f-327de8a955cb · outbound

This paper cites Jordan and Robert A.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Jordan and Robert A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.316607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.756907Z digest=sha256:30935ebdc0c3f5cc5633a0f2db7cab08e64fa5e6a8c0096fc00a7f5eee84909a

Observation f6badd95-d893-4aad-8ea6-a1bd72969f74 · outbound

This paper cites Chang, and Manolis Savva.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Chang, and Manolis Savva

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.292549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.761344Z digest=sha256:21d12832f816f4f352fe4d20e781429695a04aaabb6a9dd0005d7b3ee10979ba

Observation 3f7d3547-81db-4fe4-a64f-e979f8ba1835 · outbound

This paper cites Sparse mixture- of-experts are domain generalizable learners, 2023.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Sparse mixture- of-experts are domain generalizable learners, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.264764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.765896Z digest=sha256:036357a4963b0805a9e8227492352bb68bf02a8d76d827b1642c282b14564c66

Observation 2724bc3e-3750-4f21-b47d-0347ba17e308 · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Shapellm: Universal 3d object understanding for embodied interaction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.770372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.770372Z digest=sha256:f4c5dbd05cdb229d71eea30df2a7cbbc712d118e057e32130c08eda10a498830

Observation 7f5e2d44-26b1-4470-9a43-05d4f8d085fd · outbound

This paper cites Gpt4point: A unified framework for point-language understanding and generation.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Gpt4point: A unified framework for point-language understanding and generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.233071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.775396Z digest=sha256:3cb430aa5e8c4288baf563c2eecce7364ef3579c5cec6e9b09c4f53152c64e4c

Observation 24acd481-5d72-453e-a989-3d9fd703fdee · outbound

This paper cites Learning transferable visual models from natural language supervision.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Learning transferable visual models from natural language supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.208589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.780162Z digest=sha256:3747438920c08ff4ef54731181036f509bb26c4ead528e9a21923bf3fa476320

Observation f73a2fa5-567a-4dbb-84db-052a23f17d30 · outbound

This paper cites an unresolved cited work.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:07:10.188461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.785362Z digest=sha256:d0ddcf28ca6ef119b4ddafa58b9723ced1bf6de12dd96525fc777cf48fabd0aa

Observation 2a77a542-21cd-4354-bf39-0a6d2849be32 · outbound

This paper cites Vi- sion transformers for dense prediction, 2021.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Vi- sion transformers for dense prediction, 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.166335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.790035Z digest=sha256:1943700199193ece69da4a34319143cf4fb3b3e85370e06ff90a975ed49cb6c7

Observation 02d569d2-dbda-466e-864e-5818c7333c82 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.794766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.794766Z digest=sha256:3dd4419b35191001877bd5e7460d89ac841df440f21fd474bafa99c2ae7aa39c

Observation d804690d-9b08-43e8-99c0-20ec5e3ee2ce · outbound

This paper cites Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.799816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.799816Z digest=sha256:68776a25e2a00b9b067a0cdaa9d2c07440d3e4d13c435c19d6561fe54982b4b6

Observation 0aab04aa-d72d-422f-80f7-dda01a989a25 · outbound

This paper cites an unresolved cited work.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:07:10.133677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.804311Z digest=sha256:db55a0f5bc26d81cf9b9f15d74a32de794f1d76290d2cb3c9029f017cc65c913

Observation 940ea8e0-4bdd-49e8-bb14-0fdf162b5e12 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.809207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.809207Z digest=sha256:eef3a190dbc3cc0a8d081795f004cebed45a48ba9018b9287f40a08a8a915f99

Observation fba8d501-18e3-406c-a043-f593a8fea8d3 · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Pointllm: Empowering large language models to understand point clouds

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.814047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.814047Z digest=sha256:24807c95a7b250f07b6e4e9d0875defa1d9e493a1302105a69bc07dc177f6da3

Observation 9f8d59fc-f8fe-4668-88a2-f8836c1e8ae2 · outbound

This paper cites Qwen3 Technical Report.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.819128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.819128Z digest=sha256:3629e93260f77b75e45431d073024de748ff8e6b8d2fe17c1632e07eab9e513d

Observation 0ef245bd-565d-493a-b113-6dc16de3fb31 · outbound

This paper cites RoboPoint: A vision-language model for spatial affordance prediction for robotics.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints RoboPoint: A vision-language model for spatial affordance prediction for robotics

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.097315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.824467Z digest=sha256:6c5c41fd470f4414671326cbd8030c1b0523cbccf6b5ff259d3c04c3bcc4b84c

Observation e1ba79ab-a99f-439d-9ae7-090dc7b80a6d · outbound

This paper cites CounterCurate: Enhancing physical and semantic visio- linguistic compositional reasoning via counterfactual exam- ples.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints CounterCurate: Enhancing physical and semantic visio- linguistic compositional reasoning via counterfactual exam- ples

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.075430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.829131Z digest=sha256:1e0aa7488ec85a5f81622378daea043c4bd996bca6d85f31d0573f3a8ce9b3c0

Observation baff45ea-33b3-4324-850e-03b46679490e · outbound

This paper cites Pointclip: Point cloud understanding by clip.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Pointclip: Point cloud understanding by clip

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.833896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.833896Z digest=sha256:f77eedfcd2a40e6f49337bfce49d6cf0b8788f9ebdded953716206765c652bb6

Observation 3459d38e-c90d-4178-9ab2-1d9b097dae74 · outbound

This paper cites Agent3d-zero: An agent for zero-shot 3d understanding.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Agent3d-zero: An agent for zero-shot 3d understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.044370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.838882Z digest=sha256:050fa448fead0d79f1d4d63f2c3d61bbecf296e2c963c5e71bb78b6096b1b715

Observation 67fc2bc3-27b0-4f3c-ab1f-274ad0d5b818 · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Uni3D: Exploring Unified 3D Representation at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.844137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.844137Z digest=sha256:1aa899dec84753bd78715b17463ca352825e355c747b26fd69636370d7b62acf

Observation 5358cf3b-e0cd-4e85-a541-bad5a4f793e0 · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.027399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.849489Z digest=sha256:56f30161bbc8d53b89580450b69cd0b6e58b8e54f8718f5a34791e0a8001d089

Observation eb5a55ea-172b-4544-977b-fbe6895be27e · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.854311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.854311Z digest=sha256:d7ffcceac54c435ac9a53ef08d316439a473cba9a29a57068b429141f98273df

Observation c720176f-8496-45ca-b223-50f8b038d0ed · outbound

This paper cites Point- clip v2: Prompting clip and gpt for powerful 3d open-world learning.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Point- clip v2: Prompting clip and gpt for powerful 3d open-world learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.006525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:07:09.859176Z digest=sha256:89720e483bb758f2f14da14699d4f07fa85c6cb10a652c8a17dac36a0a18ac24

Pith citing papers

No inbound Pith citation observations are available.