Pith. sign in

Paper Citation Record · LEDGER

Warehouse Spatial Question Answering with LLM Agent

As of 10 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2507.10778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10778 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:29:40.087802Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2ed1f8c-2a8d-4720-84e0-5316589660b5 · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Warehouse Spatial Question Answering with LLM Agent SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.536084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.536084Z digest=sha256:869adc75f88daeca29a8e71702b8629feb35063101ae5b0f7d3efc66ece4ec5a

Observation 8546b2a5-113b-4722-93f7-8bacdf1ba08e · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Warehouse Spatial Question Answering with LLM Agent Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.601314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.601314Z digest=sha256:9d889e7fa9b35c271a0b18debef91805ee43909e0ab00a6c8b1afe04bf56946d

Observation 454134af-8b06-4101-84ee-8f8796bd87a7 · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision-language mod- els.

Warehouse Spatial Question Answering with LLM Agent Spatial- rgpt: Grounded spatial reasoning in vision-language mod- els

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.903708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:38.675477Z digest=sha256:f0af9907b2219c99e76db6b25e5845ce6affc4bc1e8ab2e6911f6c2a78605511

Observation 82dec2ed-f7d3-4dc7-b385-6488c149875e · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Warehouse Spatial Question Answering with LLM Agent Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.751572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.751572Z digest=sha256:8e9aa579fc058834a41228458bbbe5c5915f7e6cc908dde8fe770041f4f1ae8d

Observation b2839409-843e-4a79-8ea7-8a2741ae333b · outbound

This paper cites Deep residual learning for image recognition.

Warehouse Spatial Question Answering with LLM Agent Deep residual learning for image recognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.866567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.866567Z digest=sha256:b4f8696ac5639bb6ae637d0241b08395276cccd397fe474eceaaac44859abafb

Observation 91eac91c-5b43-4613-918e-ed4b68643fa5 · outbound

This paper cites ToSA: Token Merging with Spatial Awareness.

Warehouse Spatial Question Answering with LLM Agent ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:29:40.213090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:38.951111Z digest=sha256:76edc2ae655484da477823c11689c189c2e919cd0369ef375aab8a2937f1460d

Observation 9ab03bd5-fe61-43cb-9a85-e8ea79c7a311 · outbound

This paper cites Zero-shot 3d question answering via voxel-based dynamic token compres- sion.

Warehouse Spatial Question Answering with LLM Agent Zero-shot 3d question answering via voxel-based dynamic token compres- sion

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.636975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.046690Z digest=sha256:0f47e0e220f817155e59fa191b808a0dca32efa6cccddba8ef532b092ac960b8

Observation bea3ddf5-48cb-4953-a0a4-9a820e2facfb · outbound

This paper cites Embodied agent inter- face: Benchmarking llms for embodied decision making.

Warehouse Spatial Question Answering with LLM Agent Embodied agent inter- face: Benchmarking llms for embodied decision making

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.484786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.147806Z digest=sha256:e2e46e2bb68797c22399c10b63d7459bb5bfe952626a354664296f623237fa59

Observation c0645438-3102-44c6-be13-edce72644785 · outbound

This paper cites Seeground: See and ground for zero-shot open- vocabulary 3d visual grounding.

Warehouse Spatial Question Answering with LLM Agent Seeground: See and ground for zero-shot open- vocabulary 3d visual grounding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.340612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.240988Z digest=sha256:b14f8c1a5ceec6b5b6dd19ad86450813682fb310688233a57187633bcc02750c

Observation 95ded4ad-249c-406d-b44a-c81945463e3c · outbound

This paper cites Focal loss for dense object detection.

Warehouse Spatial Question Answering with LLM Agent Focal loss for dense object detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:39.339118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:39.339118Z digest=sha256:65541848b2090f3283d916ad147a58b8d75f641e4047cc7706adb8830ba67921

Observation b91428a8-af37-48e5-8393-4014bd586534 · outbound

This paper cites an unresolved cited work.

Warehouse Spatial Question Answering with LLM Agent Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:29:41.171583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.441137Z digest=sha256:9c8e92420764c87882fab0f727be47d75863e2b37d27098dbfdf9d63ad6de375

Observation 5fab6fd2-c228-419e-a7fa-02f60e1b36c3 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Warehouse Spatial Question Answering with LLM Agent Videoagent: Long-form video understanding with large language model as agent

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.004358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.531046Z digest=sha256:caa8f6d06bf0dbfe2ef04e90e094c8ba6a9fbcb1056af562b5ea8f4fd6675cd5

Observation 0eb380ba-f903-4438-8451-a846145b6362 · outbound

This paper cites Vlm-grounder: A vlm agent for zero-shot 3d visual grounding.

Warehouse Spatial Question Answering with LLM Agent Vlm-grounder: A vlm agent for zero-shot 3d visual grounding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.857991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.602826Z digest=sha256:f4f798bbac8d7b340dc351c7a29f72938abcd15da598bbfaf4a62bcf2fcef2ad

Observation fdd43a2f-1baa-469b-91e6-839abcf414f5 · outbound

This paper cites Fouhey, and Joyce Chai.

Warehouse Spatial Question Answering with LLM Agent Fouhey, and Joyce Chai

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.724684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.706694Z digest=sha256:0645bf186e3c70ac2d04fc2ec5def8065c854510703bf2345ed2733e148a0b68

Observation 8860a391-0e75-410b-af83-022e58746082 · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

Warehouse Spatial Question Answering with LLM Agent Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:39.837375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:39.837375Z digest=sha256:e988ab4b15ab55be511f28be2556691ec49379a06c1b756a2d5d3c59c4d63e7e

Observation 1d12a078-4ada-4793-ac93-cdb943b2dae1 · outbound

This paper cites Agent3d-zero: An agent for zero-shot 3d understanding.

Warehouse Spatial Question Answering with LLM Agent Agent3d-zero: An agent for zero-shot 3d understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.590524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:39.930124Z digest=sha256:0d06abac3b149f5d9192630cd996b1d16e6f0144692d73efbfa82c1cc04c6004

Observation a4424dfe-9a38-4859-b3f6-9f00496230c5 · outbound

This paper cites See and think: Embodied agent in virtual environment.

Warehouse Spatial Question Answering with LLM Agent See and think: Embodied agent in virtual environment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.473966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:40.024063Z digest=sha256:c8328142df88be87f8c5b0eff26bb64349b4ad59174f2440a9c90c3c5eff6f32

Observation bf329b4c-22fd-4676-8123-d4de0cb898f8 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

Warehouse Spatial Question Answering with LLM Agent Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.349978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:29:40.087802Z digest=sha256:9b14bd63107c10c075bcdaf78010ad31b17ca8715bc27e77147b181c72db5449

Pith citing papers

No inbound Pith citation observations are available.