Pith. sign in

Paper Citation Record · LEDGER

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.15054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15054 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:23:03.822300Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4afc75ac-6588-4b67-9378-2ffc5ef46892 · outbound

This paper cites GPT-4 Technical Report.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:00.973913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:00.973913Z digest=sha256:1ea3d666b5cc0c0af1c72f389861e51592468b73f6413c270de2df734e607ff8

Observation f649197a-b618-408a-87f8-e57038564a1e · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Grounded 3D-LLM with Referent Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.210993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.210993Z digest=sha256:36743514f26741e4754fe3a0c0f3a9c09acc17e6d43b171ae553a20bdd6b9d43

Observation e8b896e1-377a-4bf6-b5de-c6c188cc415e · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.362939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.362939Z digest=sha256:077ca8acb546f1dff87ff8034abd6f1810431a4dc673da5aba543781e18be6d5

Observation 4675c47a-7b5a-4ab3-b24f-83a706477a0e · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.530176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.530176Z digest=sha256:50ae82e294aaf6c9ce37a2f7fd0980dfcca1bead7a7bc05868eb5ebc5fc0bdf6

Observation 05707b40-b89d-40bf-8e03-39d0298d5072 · outbound

This paper cites GPT-4o System Card.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.696722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.696722Z digest=sha256:3487a1f9a7ba89564271028c21aedd2b277014084d32ba5185a0473400c7a29f

Observation 87e64744-d2a1-4358-87b8-233278902a8f · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.780480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.780480Z digest=sha256:64dcb01ebc31f4ebaea18b9ca7a93fa805818eefd1a729c451b314ff54904781

Observation 5e88b47d-566b-43b7-9ad8-e1851429c6fc · outbound

This paper cites Thinking with Geometry: Active Geometry Integration for Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.948198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.948198Z digest=sha256:61b02aa38161297d497ec3caa49ccf663aa5f5990c2b92a5132bfca5d2e57ccd

Observation 2f189430-0882-4be4-8ef6-5b005ae1de66 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Depth Anything 3: Recovering the Visual Space from Any Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.051394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.051394Z digest=sha256:f54c5eac7ec6a9dd89bcd097eb7781ec0abc32ce089b9a723fa48504eee1afb5

Observation 4a7fe89a-a414-4b7c-9969-fa9de197a021 · outbound

This paper cites Trace anything: Representing any video in 4d via trajectory fields.arXiv preprint arXiv:2510.13802, 2025a.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Trace anything: Representing any video in 4d via trajectory fields.arXiv preprint arXiv:2510.13802, 2025a

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.211902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.211902Z digest=sha256:bb25530a994ac8bc60078a5f46aad8428693797620bb90543318aeac3855457d

Observation bf8cbbc5-f5d1-4423-802f-21ab9d50271a · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.322345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.322345Z digest=sha256:68ce75e4952eea775238ce0eac86949d091f16dcde9db21cfa5c086700819b78

Observation d3b06a84-d677-4a0a-9a70-d483a1a880cd · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.432874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.432874Z digest=sha256:7d4a160ca221e1842e3c433f3f485a78d555cf2e1b9b13da40e563ed0200a28a

Observation edf82e13-76f9-433c-902f-45a4153ecd22 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.519901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.519901Z digest=sha256:e01200b82a836984b5a23b759bf787f43451ff95c86bd1f0eeed5921ffdcaca5

Observation 2cac0151-ae00-4379-8ced-26b44dd6e06d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.675732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.675732Z digest=sha256:a52df6fd4635dc4f669ee34d2e48add72257bdb4eb4e952dad3f528cf8d21911

Observation 21e08e6c-d6e8-46fa-afc1-3e2ea1de8f1b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Wan: Open and Advanced Large-Scale Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.835659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.835659Z digest=sha256:747c42c2f3f5f3c6f458b17d02d1a0e7b9ee4e74ba52468e1ee72e80c418d578

Observation ce1e85c9-3eed-47d0-9f27-d9e51c7d7277 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.003520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.003520Z digest=sha256:500b32d586bf54f198bf4f7d6d4acbaf947205c637392d624069ad8a0fefa8db

Observation f91d52c3-b8c2-47b3-a065-f8004863312e · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.164292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.164292Z digest=sha256:65f3e734f00baed9a8880cd61db441ac3c9a6ba54025f1f33f49387ce726f750

Observation a94d5c6e-44f2-4f55-b3f5-195fd52981b0 · outbound

This paper cites Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.278174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.278174Z digest=sha256:37947d44a35b20c87f7ede8e1f825f08f5c624ad2ce62d45ccc6f06f8af2194f

Observation 7437ff82-5dab-4f3b-bb6f-57ef258a53b1 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.437035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.437035Z digest=sha256:947536209080175f55bdade5406339ac49c1ea1833eee085df6b795014cf0982

Observation a305829a-724f-4163-949f-21405c8dc6f3 · outbound

This paper cites SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.547032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.547032Z digest=sha256:a0220ede0cf6eb24bed92390c313533b542c2d84d5debcc8892a018e88f15c4c

Observation 65a174ab-bebd-457d-9885-22696f41c3b0 · outbound

This paper cites Long Context Transfer from Language to Vision.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Long Context Transfer from Language to Vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.658491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.658491Z digest=sha256:23bd28b378328f14f5f030e342f210f0d6e83e53c04b9322c056d3932272a539

Observation 9ffbde95-a8de-47a6-b3f6-8568ef4d81d3 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d objects

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.740289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.740289Z digest=sha256:83a91a033d78ac91482045bebee396395da2d3a26cf0b712fb3b0ac9aeeb3b84

Observation c7f2432c-0cb8-4974-abe2-f3bc6d064307 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Unifying 3d vision-language understanding via promptable queries

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.822300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.822300Z digest=sha256:54bf1f30caaa766012a05a69665895d44c96f764063cbb243e2c93fdae3d2ce1

Observation c891bda5-50ec-4552-b0e0-f6f66ad68479 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.864475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.864475Z digest=sha256:f2e7ef9d40deb5a0f14bf3992c8f24b0d0838d9812c3fff19c9e564cad033db5

Observation c4ec22de-4f42-4007-8a85-dc08415a53dc · outbound

This paper cites Seeing through imagination: Learning scene geometry via implicit spatial world modeling.arXiv preprint arXiv:2512.01821,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Seeing through imagination: Learning scene geometry via implicit spatial world modeling.arXiv preprint arXiv:2512.01821,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.120140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.120140Z digest=sha256:573d72727d16ac3e994a462178799c9a60da07611c5930670f9527664fcf7db5

Observation 99ea6e02-6a1d-4020-adf5-02c703efe42a · outbound

This paper cites Qwen3-VL Technical Report.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Qwen3-VL Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.048543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.048543Z digest=sha256:5f2ef5a322c02d4b8f01185a59c0d89477699301aa4524ec517416425adc4079

Observation 71d197b1-e280-443c-b1c5-c20763002696 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.arXiv preprint arXiv:2506.01946,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding 3drs: Mllms need 3d-aware representation supervision for scene understanding.arXiv preprint arXiv:2506.01946,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.613098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.613098Z digest=sha256:abc75d6d39f7e568fa88744e0a1b055faf7ddfa7e282c6338d2707253d61f1ce

Observation 4ba8b9c3-c4ab-43e9-9f7c-f875897ca14e · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.446109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.446109Z digest=sha256:384a9d540b3d872a74bf4177c18009b9702c119cfa2fc32b10fcd5713c0e080b

Observation 3a027346-4e82-4a16-966d-da6358e1aa4e · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.277839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.277839Z digest=sha256:c9c0b7de2f58f79a37f99c9d34ab0750665007a1bac965453138ebe99083e570

Observation 8c4a1fa3-70d6-499e-b42b-d96f56b54eb2 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.979540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.979540Z digest=sha256:a7942dac4d0dfb3aa1ad9ac3823d3aa9a2c7a9475966ad17cf1b0ef1100965d9

Pith citing papers

No inbound Pith citation observations are available.