Pith. sign in

Paper Citation Record · LEDGER

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2508.17595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17595 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:09.859176Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13a425ea-cabe-4d18-9c8f-d328ba2295ec · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.425566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.724943Z digest=sha256:db043008ed0af09a674a4ee81c06b5c501d917287490c26d3d89424fa8014911

Observation a693fb44-c00b-4473-bdee-174a06ca1aae · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision-language models.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Spatial- rgpt: Grounded spatial reasoning in vision-language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.407806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.730624Z digest=sha256:74fd8d08709c0605ba05a42f79ca5cb45275709afad84b027df163692b7be07d

Observation 5f32692a-f101-4fee-a90b-a1b6bcbbc49b · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.736041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.736041Z digest=sha256:a0d89aa893f8a0a1143fde451645c295ad2cbf77f682232544f6ac3ca0972465

Observation e0418e18-c407-440a-a0df-efafedc7d4d9 · outbound

This paper cites Fusemoe: Mixture-of-experts transformers for flexi- modal fusion.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Fusemoe: Mixture-of-experts transformers for flexi- modal fusion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.387153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.741439Z digest=sha256:be8246ba8500bb09d8161f15e659ece110d23c089f65fda91784d0e253166e7f

Observation 53d91718-43ab-4106-908a-f24f64b5211e · outbound

This paper cites Multiply: A multisensory object- centric embodied large language model in 3d world.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Multiply: A multisensory object- centric embodied large language model in 3d world

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.361885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.746747Z digest=sha256:b4b47361dbe3c5a732e6bff8d69844b1841dbb77c4c6a2e39019d991367958d1

Observation ee59531c-505a-474a-9225-12b090e87e3a · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.344898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.751482Z digest=sha256:d746ea2f4e2dd623ddcda2798185026b57af90472e0b5f27cd37c59ff8d63423

Observation 715c92b2-a6a4-4655-b77f-327de8a955cb · outbound

This paper cites Jordan and Robert A.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Jordan and Robert A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.316607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.756907Z digest=sha256:cb12f2b5e47a0bd0ba8e2f96638c311b6e963987d69ec3ff4947ec9391cb1987

Observation f6badd95-d893-4aad-8ea6-a1bd72969f74 · outbound

This paper cites Chang, and Manolis Savva.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Chang, and Manolis Savva

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.292549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.761344Z digest=sha256:f5630d017d4874b2ae76b437602158da57984a6c2048f66ede20aafbe050546b

Observation 3f7d3547-81db-4fe4-a64f-e979f8ba1835 · outbound

This paper cites Sparse mixture- of-experts are domain generalizable learners, 2023.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Sparse mixture- of-experts are domain generalizable learners, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.264764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.765896Z digest=sha256:6ec51e97633c924faf806d1df1ea146c9af797edfba3b332779522f26a6d1cf2

Observation 2724bc3e-3750-4f21-b47d-0347ba17e308 · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Shapellm: Universal 3d object understanding for embodied interaction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.770372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.770372Z digest=sha256:f4c5dbd05cdb229d71eea30df2a7cbbc712d118e057e32130c08eda10a498830

Observation 7f5e2d44-26b1-4470-9a43-05d4f8d085fd · outbound

This paper cites Gpt4point: A unified framework for point-language understanding and generation.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Gpt4point: A unified framework for point-language understanding and generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.233071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.775396Z digest=sha256:a666415a22513511a98cc2b58c3339398da8a873839540c65fcc787ae76abcbb

Observation 24acd481-5d72-453e-a989-3d9fd703fdee · outbound

This paper cites Learning transferable visual models from natural language supervision.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Learning transferable visual models from natural language supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.208589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.780162Z digest=sha256:246e00b6b89b4c54cbddfb8a1de81e229a42ef6d36561808d269645c2232fe05

Observation f73a2fa5-567a-4dbb-84db-052a23f17d30 · outbound

This paper cites an unresolved cited work.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:07:10.188461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.785362Z digest=sha256:b529f2c46dc452eeb61312bc717f316738891c3719de30b74545e99177c525bd

Observation 2a77a542-21cd-4354-bf39-0a6d2849be32 · outbound

This paper cites Vi- sion transformers for dense prediction, 2021.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Vi- sion transformers for dense prediction, 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.166335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.790035Z digest=sha256:e2d5ed3b7a12ecf96356c317cfa6b6d1e45b28419fad414ca393c658613d118c

Observation 02d569d2-dbda-466e-864e-5818c7333c82 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.794766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.794766Z digest=sha256:3dd4419b35191001877bd5e7460d89ac841df440f21fd474bafa99c2ae7aa39c

Observation d804690d-9b08-43e8-99c0-20ec5e3ee2ce · outbound

This paper cites Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.799816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.799816Z digest=sha256:68776a25e2a00b9b067a0cdaa9d2c07440d3e4d13c435c19d6561fe54982b4b6

Observation 0aab04aa-d72d-422f-80f7-dda01a989a25 · outbound

This paper cites an unresolved cited work.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:07:10.133677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.804311Z digest=sha256:cbb8fc4c453601aeafac46a585324952d5e49109b0b198530a5a38ed3701feb4

Observation 940ea8e0-4bdd-49e8-bb14-0fdf162b5e12 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.809207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.809207Z digest=sha256:eef3a190dbc3cc0a8d081795f004cebed45a48ba9018b9287f40a08a8a915f99

Observation fba8d501-18e3-406c-a043-f593a8fea8d3 · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Pointllm: Empowering large language models to understand point clouds

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.814047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.814047Z digest=sha256:24807c95a7b250f07b6e4e9d0875defa1d9e493a1302105a69bc07dc177f6da3

Observation 9f8d59fc-f8fe-4668-88a2-f8836c1e8ae2 · outbound

This paper cites Qwen3 Technical Report.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.819128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.819128Z digest=sha256:3629e93260f77b75e45431d073024de748ff8e6b8d2fe17c1632e07eab9e513d

Observation 0ef245bd-565d-493a-b113-6dc16de3fb31 · outbound

This paper cites RoboPoint: A vision-language model for spatial affordance prediction for robotics.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints RoboPoint: A vision-language model for spatial affordance prediction for robotics

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.097315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.824467Z digest=sha256:0e580e6f8cc049e87ade15387c19a9c33fcfe0dbb6624a8f56fbac5e28745be4

Observation e1ba79ab-a99f-439d-9ae7-090dc7b80a6d · outbound

This paper cites CounterCurate: Enhancing physical and semantic visio- linguistic compositional reasoning via counterfactual exam- ples.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints CounterCurate: Enhancing physical and semantic visio- linguistic compositional reasoning via counterfactual exam- ples

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.075430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.829131Z digest=sha256:0f0748def412c4e93e69698875dc61d5ca2009b471c73af2b52745694c500eac

Observation baff45ea-33b3-4324-850e-03b46679490e · outbound

This paper cites Pointclip: Point cloud understanding by clip.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Pointclip: Point cloud understanding by clip

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.833896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.833896Z digest=sha256:f77eedfcd2a40e6f49337bfce49d6cf0b8788f9ebdded953716206765c652bb6

Observation 3459d38e-c90d-4178-9ab2-1d9b097dae74 · outbound

This paper cites Agent3d-zero: An agent for zero-shot 3d understanding.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Agent3d-zero: An agent for zero-shot 3d understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.044370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.838882Z digest=sha256:6061a49609cd45f3e9aeb450b6b05231764dcb49142c3a1b107a810bb852e6d1

Observation 67fc2bc3-27b0-4f3c-ab1f-274ad0d5b818 · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Uni3D: Exploring Unified 3D Representation at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.844137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.844137Z digest=sha256:1aa899dec84753bd78715b17463ca352825e355c747b26fd69636370d7b62acf

Observation 5358cf3b-e0cd-4e85-a541-bad5a4f793e0 · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.027399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.849489Z digest=sha256:b2ca4de9dbe3fd74c825059ffd1cb69a3ebcba6157d4a81c894b9e2db5108eb7

Observation eb5a55ea-172b-4544-977b-fbe6895be27e · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:09.854311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:07:09.854311Z digest=sha256:d7ffcceac54c435ac9a53ef08d316439a473cba9a29a57068b429141f98273df

Observation c720176f-8496-45ca-b223-50f8b038d0ed · outbound

This paper cites Point- clip v2: Prompting clip and gpt for powerful 3d open-world learning.

TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints Point- clip v2: Prompting clip and gpt for powerful 3d open-world learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:07:10.006525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:07:09.859176Z digest=sha256:5c1161d1f35f37e726d80ab5375135570a6d6f053da745f94c275d5e33c98bfa

Pith citing papers

No inbound Pith citation observations are available.