Pith. sign in

Paper Citation Record · LEDGER

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.15054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15054 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:23:03.822300Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4afc75ac-6588-4b67-9378-2ffc5ef46892 · outbound

This paper cites GPT-4 Technical Report.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:00.973913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:00.973913Z digest=sha256:66a48938f8da54a296ca7afae9a95bbc807e7921921018b20c2a2a87dd68fd73

Observation f649197a-b618-408a-87f8-e57038564a1e · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Grounded 3D-LLM with Referent Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.210993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.210993Z digest=sha256:59570cb8d6c8658accb38b95e42b8cf6c9cf4c062c17889fa8860d9ff5ad89a9

Observation e8b896e1-377a-4bf6-b5de-c6c188cc415e · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.362939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.362939Z digest=sha256:a3359482acf3e6c7c90b47e5e293613bcb7e5eb101bac480e4a6281ad6a0e7da

Observation 4675c47a-7b5a-4ab3-b24f-83a706477a0e · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.530176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.530176Z digest=sha256:faa8a69d6e6f52298493f1c00b3b4bb7607e82b73dd855bc5e09feddd2909ab3

Observation 05707b40-b89d-40bf-8e03-39d0298d5072 · outbound

This paper cites GPT-4o System Card.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.696722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.696722Z digest=sha256:0aa4d153da77434cbfd0a886d224df9783170da658646b407a26e4634bdf6e9e

Observation 87e64744-d2a1-4358-87b8-233278902a8f · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.780480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.780480Z digest=sha256:bcccffb5c4dcb7be56016d726fc6b92dfdb3f4a164f80e4fe584f0edfbf8bbe1

Observation 5e88b47d-566b-43b7-9ad8-e1851429c6fc · outbound

This paper cites Thinking with Geometry: Active Geometry Integration for Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.948198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.948198Z digest=sha256:3c309597d1d31a5a4396528de4019a79ff4dcc281b05772a639ecd017b893d08

Observation 2f189430-0882-4be4-8ef6-5b005ae1de66 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Depth Anything 3: Recovering the Visual Space from Any Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.051394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.051394Z digest=sha256:d132c61b9f6b719731a6efcc35e7a327ddd767259741fd6fad9d10bf900c6dc7

Observation 4a7fe89a-a414-4b7c-9969-fa9de197a021 · outbound

This paper cites Trace anything: Representing any video in 4d via trajectory fields.arXiv preprint arXiv:2510.13802, 2025a.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Trace anything: Representing any video in 4d via trajectory fields.arXiv preprint arXiv:2510.13802, 2025a

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.211902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.211902Z digest=sha256:fcf40a60f9064d7ea9d1931544d68543c2f9bf9d984f9c5e1986c00d59de7151

Observation bf8cbbc5-f5d1-4423-802f-21ab9d50271a · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.322345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.322345Z digest=sha256:9c295f773bb44cc1ceb65fa6864e248790721d9d40817bcbfef897c9487d2c45

Observation d3b06a84-d677-4a0a-9a70-d483a1a880cd · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.432874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.432874Z digest=sha256:694c4222c1397f9b2e995794639dc74e4aa371493ebd42053316459305442198

Observation edf82e13-76f9-433c-902f-45a4153ecd22 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.519901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.519901Z digest=sha256:b2262ea82851cc73c0531d08be4ce50ddff5988d6bde0748cc9a612dd47f04f2

Observation 2cac0151-ae00-4379-8ced-26b44dd6e06d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.675732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.675732Z digest=sha256:e155bac36738fc48a39339b9f60da6bf7d62826ded325c1c58144fe631d5ce3f

Observation 21e08e6c-d6e8-46fa-afc1-3e2ea1de8f1b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Wan: Open and Advanced Large-Scale Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.835659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.835659Z digest=sha256:c3f81d58624cedd6153477e5b773fc99787a541a0620ffae37ca9f8a20fcd42f

Observation ce1e85c9-3eed-47d0-9f27-d9e51c7d7277 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.003520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.003520Z digest=sha256:9b7648b93ac50a7db72d458813357db8fff88e304e64aa2537b63593b88232a1

Observation f91d52c3-b8c2-47b3-a065-f8004863312e · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.164292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.164292Z digest=sha256:1d01aea459a7b83ec5f1e3a38d4a264568ddb1e2116cbaf22709ca8d3c0ca542

Observation a94d5c6e-44f2-4f55-b3f5-195fd52981b0 · outbound

This paper cites Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.278174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.278174Z digest=sha256:b95e206fb850562d76d445d984187f865ac1722d7d77af548e9a9f6ffffd1b1a

Observation 7437ff82-5dab-4f3b-bb6f-57ef258a53b1 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.437035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.437035Z digest=sha256:9bda26eb8c06db54552e600ea5fc9c25b74d31a458d1e8796c778a352b52e85b

Observation a305829a-724f-4163-949f-21405c8dc6f3 · outbound

This paper cites SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.547032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.547032Z digest=sha256:e893662541b2670811ece06e1221f23475325cabbc0551728b3d2b8da72a4751

Observation 65a174ab-bebd-457d-9885-22696f41c3b0 · outbound

This paper cites Long Context Transfer from Language to Vision.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Long Context Transfer from Language to Vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.658491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.658491Z digest=sha256:2b0407c1cb08476663056342a2bb8ede3446cfb2793922ae3982e657355632a7

Observation 9ffbde95-a8de-47a6-b3f6-8568ef4d81d3 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d objects.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d objects

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.740289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.740289Z digest=sha256:6e0c67ae924bcd4bb3a49e75ec24f55bca7816aaab45deb0bf01a34508098743

Observation c7f2432c-0cb8-4974-abe2-f3bc6d064307 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Unifying 3d vision-language understanding via promptable queries

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:03.822300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:03.822300Z digest=sha256:01550ac5e3b4fa37367382360100353612d3c25eaf80c7679767b0de5feeeced

Observation c891bda5-50ec-4552-b0e0-f6f66ad68479 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.864475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.864475Z digest=sha256:68ffc57910f105c72d8d6228f49d4e1af4043252ca15b7215098c852e6107558

Observation c4ec22de-4f42-4007-8a85-dc08415a53dc · outbound

This paper cites Seeing through imagination: Learning scene geometry via implicit spatial world modeling.arXiv preprint arXiv:2512.01821,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Seeing through imagination: Learning scene geometry via implicit spatial world modeling.arXiv preprint arXiv:2512.01821,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.120140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.120140Z digest=sha256:083470cf034d8378df84ba643dbee63e805112a38da1cbe97aa970c65754e93f

Observation 99ea6e02-6a1d-4020-adf5-02c703efe42a · outbound

This paper cites Qwen3-VL Technical Report.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Qwen3-VL Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.048543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.048543Z digest=sha256:e96c7920bdb5926ee040f25c341cf0a257217d7e4a2d83a1b0b50dab8b6744c3

Observation 71d197b1-e280-443c-b1c5-c20763002696 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.arXiv preprint arXiv:2506.01946,.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding 3drs: Mllms need 3d-aware representation supervision for scene understanding.arXiv preprint arXiv:2506.01946,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.613098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.613098Z digest=sha256:ea997ee086f127bbb3c17bb3a127405719932e70d132b7a7ece8a1f2929f6228

Observation 4ba8b9c3-c4ab-43e9-9f7c-f875897ca14e · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.446109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.446109Z digest=sha256:4968bf9c04b9609fecf618793d956bc3f18c8ae9618d8077d3ebc8bb9fce9cee

Observation 3a027346-4e82-4a16-966d-da6358e1aa4e · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.277839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.277839Z digest=sha256:1bb5bff5528eef94bcee7186e080711d8c6a3afb20da97cf505836a17bfba539

Observation 8c4a1fa3-70d6-499e-b42b-d96f56b54eb2 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.979540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.979540Z digest=sha256:1d95f96ff3bfef5624024540cee2d9c5a355bfdd9cfffe6e03f5468db7bd6eef

Pith citing papers

No inbound Pith citation observations are available.