Pith. sign in

Paper Citation Record · LEDGER

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning

As of 17 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2604.10506.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10506 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:53:52.775348Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T03:13:07.963343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T03:14:31.636817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0f9d5f2-4935-4446-8273-5f7f78ed5520 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:09:24.919963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:3f48f1aa29182601a494170aa82510c5c2fceb58a10ea5fec316ca4448f4e54e

Observation 39d35e26-3360-410a-8325-7721fc702ca9 · outbound

This paper cites Surpassing Cosine Similarity for Multidimensional Comparisons: Dimension Insensitive Euclidean Metric.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Surpassing Cosine Similarity for Multidimensional Comparisons: Dimension Insensitive Euclidean Metric

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.794011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:de3ca51bdd061b55648ab46cfef179e4b6cbdf203868c87b99694defe44c7e88

Observation 804dfea6-33a5-482d-8986-eef411c84a4e · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:20:31.604247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:be43623b36540f5e355379e8d9f1059ec307441bff9382bd0e2ae71a3d976a18

Observation 9e6523ea-ac96-4fe2-9022-b7d9c3b5fddc · outbound

This paper cites On the modification and revocation of open source licences.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning On the modification and revocation of open source licences

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.822915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:5127bbefb417ee5a786c924cabc93201a86557e5c53fd9685595119b2368ceb1

Observation 2b981b61-9835-4d9f-ad5f-fac46f34d757 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning NVILA: Efficient Frontier Visual Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:41:01.759826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:e8d62c81f023cd03ff05d9714e0f0e594cbc99be65bf99ba3ccbb997ec36703b

Observation c1f6976b-ab0d-48b3-9060-0279b1ac953a · outbound

This paper cites Decoupled Weight Decay Regularization.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Decoupled Weight Decay Regularization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:41:01.855632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:e33cab6a6c10fbe055f63cfc5d09d3a7a548b4bd8b9f97d14b1b7957e7875807

Observation 1aacd13c-734c-4295-91bc-01ac0155a1c6 · outbound

This paper cites Composing Bridges.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Composing Bridges

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.896580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:863cf997d09236bf731430f6e043e14a23661b512b7e6747ea66866f68df6442

Observation 6a4c00ee-46b3-40e3-8923-71ae8a9a5fbd · outbound

This paper cites Local vs. Global Interpretability: A Computational Complexity Perspective.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Local vs. Global Interpretability: A Computational Complexity Perspective

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.785215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:9a17e004b4056aaccacc8d02ef160f6f3ef26fe31aec270a8e69514ec6380e50

Observation 883bd1eb-44d5-42f6-af33-99d11bbb71a6 · outbound

This paper cites Relation between the keV-MeV and TeV emission of GRB 221009A and its implications.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Relation between the keV-MeV and TeV emission of GRB 221009A and its implications

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.838513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:3bc72a02dbdff25217e778d84710762ca4f76b8322c294c1010bf81caf0a22ee

Observation 25f711df-b3a4-4442-a632-63840a22a4b8 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:09:32.945221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:559e388bad796d255a364bf6c4f56b9a9e06cfce6fb030b0a22fb6d17c89734c

Observation deb80c75-d03d-45a0-88dd-ebb85acb1217 · outbound

This paper cites Novel $H(\mathrm{sym} \mathrm{Curl})$-conforming finite elements for the relaxed micromorphic sequence.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Novel $H(\mathrm{sym} \mathrm{Curl})$-conforming finite elements for the relaxed micromorphic sequence

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.903053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:cb92385c6c9522691c62593861e711aedcec8ee6d3a056d5adefaf40f93ab535

Observation c5baa0f4-c316-431b-a70e-765a69a39388 · outbound

This paper cites Complexity of Multiple-Hamiltonicity in Graphs of Bounded Degree.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Complexity of Multiple-Hamiltonicity in Graphs of Bounded Degree

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.890769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:c0802412eb81cc13bb154ae225dea1ddd5aadd39775777d017920b7cf7d66234

Observation 2bde0eae-ac20-40cc-90b1-de24c00cd0b9 · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-26T02:02:26.295966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:15a2c81188981bff70623a96c604ba13403c58f2d1e05f5f51f53baf861e2e17

Observation 51135fcb-62cc-4db5-ac2e-48761b49e915 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:a3e60aa2533dbfc8c303d31e850c3e8e050f63ea217f03563e99be65816b7cc6

Observation 6f6f27bb-e7d2-4e8a-8ae7-ca3dc3495693 · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:41:01.809714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:c67c2369dad7204c1ab2bb47e9eeb692c1b151a65479ae2afb67fca7d8818bb3

Pith citing papers

Observation 59ab3e84-c3b9-4ca9-9af7-0014f97de02f · inbound

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment cites this paper.

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-08T03:14:31.638468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T03:13:07.963343Z digest=sha256:65f85ab19f3663f302715d3ffb2fea83ef56c401e58d4a9f1432fa3cda328ada