Pith. sign in

Paper Citation Record · LEDGER

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 12 inbound Pith citation observations for arXiv:2412.00493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00493 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:26:19.763700Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:59:11.071380Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:00:51.358037Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved21
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8877d8d5-9b7b-4c2f-a68c-b79cd714b710 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.528763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.528763Z digest=sha256:03aba6dcc1fb5cc618a44d4bf3b8fa47d2fa9201e9f1a6fa2a3e9f34b43589a8

Observation f390bf75-f74b-47ff-9b99-688b48be521b · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.484056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.533863Z digest=sha256:dd9be6696db0bd2c647903a9fe3001c2a2269cf45c34e9cb786593916d916aff

Observation 7fcd4a09-f3fc-4b44-ab77-dce2bc4a78b8 · outbound

This paper cites 3djcg: A unified framework for joint dense caption- ing and visual grounding on 3d point clouds.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3djcg: A unified framework for joint dense caption- ing and visual grounding on 3d point clouds

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.470793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.538516Z digest=sha256:d3086a2c4e74a99466b082413053220a5848ceeec5ec544c8fa46305a0a6a1dd

Observation e685faa3-2175-4ab5-b915-ce2ff4f6bee9 · outbound

This paper cites Guibas, and Fei Xia.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Guibas, and Fei Xia

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.457559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.542906Z digest=sha256:7f2111e44f16e3e19d403a4a0f7f21bcb2e7e776631e86e4b4207f13d7c80057

Observation b0c42af2-40ad-4956-a903-0094a211f815 · outbound

This paper cites Chang, and Matthias Nießner.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chang, and Matthias Nießner

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.444356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.547045Z digest=sha256:474c22f55d8125660541af3b14cc6b2e7302527f41f72809192d886046894286

Observation b3d87e49-2dad-4b71-8073-7582904abb67 · outbound

This paper cites an unresolved cited work.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:26:20.432156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.551280Z digest=sha256:75a1a2a32997768e5f736907cd26e52b337359a3338d79602ce4c5e040216eb8

Observation c2e5edaa-2613-4029-82c5-9f74a70e139f · outbound

This paper cites an unresolved cited work.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:26:20.419515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.555094Z digest=sha256:002b7db60cc6e977378a8e1adbf20f9e21425d9fd8031c56b7201eda6c7cdb8a

Observation bebb84cf-978f-4307-b76a-9a3283f00fc2 · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Language conditioned spatial relation reasoning for 3d object grounding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.406810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.559497Z digest=sha256:571980ccbe9eee703fbc40f6ca0ebce0601222280d091432e22225b451fee2cf

Observation 3d27d9f4-8102-4ca0-bb97-b9b222df0427 · outbound

This paper cites LL3DA: visual interactive instruction tuning for omni-3d un- derstanding, reasoning, and planning.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding LL3DA: visual interactive instruction tuning for omni-3d un- derstanding, reasoning, and planning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.394364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.563326Z digest=sha256:2630b20c514e5a73295f5cd2ab0a9cdd336370f14af2b8d7db053a86e66da650

Observation 64981b8e-d522-4f13-bfcb-628ef9b5f471 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Grounded 3D-LLM with Referent Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.567697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.567697Z digest=sha256:b70aca721d8ce6d2c72f75cf08f6f55fce4f4b4effe9bdfe96fedefc795629a0

Observation 8a73b7cd-cad0-4227-b25a-ed94845beb48 · outbound

This paper cites Intern VL: scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Intern VL: scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.381105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.572980Z digest=sha256:e299ba477fc845cc4c5b66819ea4a79e9599a74c963f44cded3f89b296faa59f

Observation f7b8c156-85f8-49e1-8e0d-a9c16700aee4 · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas A.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chang, Manolis Savva, Maciej Hal- ber, Thomas A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.368506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.577793Z digest=sha256:5615659c369a849eebcbc73a51755646620e40237d9a58abb6851b0b2b199d28

Observation 8f6c05eb-9542-444b-88b4-bca6494f8448 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding An image is worth 16x16 words: Transformers for image recognition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.581745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.581745Z digest=sha256:9aed4e9a860773bb737e58a1e414ff904c734fc5ffa3c00de91581693a06b8b3

Observation 2cec3e35-4029-4ede-9680-12b6169a77db · outbound

This paper cites The Llama 3 Herd of Models.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.589672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.589672Z digest=sha256:320fd02ae178a8786fd5a89ff5dbab004f80fe413d67415e46ee89ec69ee6bee

Observation fb9650f3-8338-44f2-95a8-698d8cf91340 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.593868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.593868Z digest=sha256:01262602774a30bf49678edab25af1d19f168075be83c93487b327035a998b27

Observation c3ff64a6-c90a-4a73-9ece-323b7e22b0a3 · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3d-llm: In- jecting the 3d world into large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.335860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.598085Z digest=sha256:fb7522dd9310806307a117865005aecc682709ee55fa59c8ed9b4202dfe4a5d8

Observation f82e33e3-62e5-412a-8276-a6204f0f24b7 · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.602076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.602076Z digest=sha256:6b62f9eb35624d72f465b2427e3655a6bbdc9662d8669893d00ae5b7df647040

Observation 599d0e1f-4988-4d43-b4fa-89a4aa9731b5 · outbound

This paper cites An embodied generalist agent in 3d world.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding An embodied generalist agent in 3d world

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.324038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.606516Z digest=sha256:826e95557bb64f5e677c51a9d237727b8f1be04d9939581be9c9d3baccba0852

Observation b9a8c148-edc9-43d7-a551-e7be0aadae39 · outbound

This paper cites Multi- view transformer for 3d visual grounding.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Multi- view transformer for 3d visual grounding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.299841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.614299Z digest=sha256:906656e11c2b2a5a15c328e5c06cc9b5ded3347e09c52d1134bd304001845ab1

Observation 12b43e78-d485-486a-899d-56e04ce7f700 · outbound

This paper cites Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.618414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.618414Z digest=sha256:a5416890fce66e5ea92f54a971a7fa0588e4f3331368f53127ce851bb790ecd5

Observation f4d9ba2e-890f-4984-92e0-21f24b917066 · outbound

This paper cites The budgeted maximum coverage problem.Information Process- ing Letters, 70(1):39–45, 1999.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding The budgeted maximum coverage problem.Information Process- ing Letters, 70(1):39–45, 1999

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.288134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.623700Z digest=sha256:7f22bfafa019e5a3fb9a43991d3b7e817bcc4a503d8d29902772991d69dc7707

Observation c3da25e1-c635-4943-889b-e8c17fdaddf3 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.627405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.627405Z digest=sha256:fec221a43ca46263798df1472024c247e42daa842e6dd43ad79322df49ea4d15

Observation 90534326-cfc1-4102-81f9-e41489960c3e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.631614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.631614Z digest=sha256:8d744d28585155c95687ae8270c7b64873e1bb3a672ec4538eb7f9f66539404e

Observation 68e9609b-219f-4beb-8668-26cc1a514994 · outbound

This paper cites an unresolved cited work.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:26:20.275826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.635864Z digest=sha256:badb7fdf8a94337acee6f3ed7fb1144f79b95e0cc15998dcd76791b890ea9b02

Observation c9a8f55b-b79c-4c41-ace9-3cb9e484c3e3 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.263002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.640985Z digest=sha256:e5391948fd5ab473b73c842768597214184a5343c669b93be55fef4ef8b2c7f1

Observation c14c733d-09e3-42b6-83dd-9547ef8d1f6a · outbound

This paper cites VILA: on pre-training for vi- sual language models.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding VILA: on pre-training for vi- sual language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.251113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.645440Z digest=sha256:179923a7eeb097dd0b3bf401e7ff1dd49aec4eee7cbcb8fbdd9635d426c91bcd

Observation a12a8f55-5848-4ce5-998d-15c55039cc4e · outbound

This paper cites Multi- modal situated reasoning in 3d scenes.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Multi- modal situated reasoning in 3d scenes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.238774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.649121Z digest=sha256:d1723989416d3a99f54d69d54fe5eac4ef98f5eb07e909f922108cad24349d15

Observation 79d48281-b306-48e1-ab6a-3d2d66438a15 · outbound

This paper cites Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.652755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.652755Z digest=sha256:4829eb2d364c41ad0d8410e5aa6288ed8e733fd99c459bd777dd9ee5dd8d3ac0

Observation d599c639-00b2-4896-b576-4dc0849cfe42 · outbound

This paper cites Visual instruction tuning.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.226343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.657190Z digest=sha256:37610800d802c9f0ab5ad935956c29383bc499131dbf42739c3271bae91216d9

Observation e9ab469b-1570-4b36-a5ef-5d045d1d3b1d · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.660952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.660952Z digest=sha256:e8e0ccaeb773b9dba50c495bbc4bcb547fb4e84e813f69f69a2d53d7a7e46cfc

Observation b61d4049-8934-4e47-beb4-c54b5f41605a · outbound

This paper cites SQA3D: situ- ated question answering in 3d scenes.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding SQA3D: situ- ated question answering in 3d scenes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.213754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.664761Z digest=sha256:aa09d8600b93ae736d9761bb6579d93256cea9517181cabd05f6b40c8ace81f8

Observation 3f945e42-0fe6-46fa-9aa8-0eb9c9519b90 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Openeqa: Embodied question answering in the era of foundation models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.201043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.669415Z digest=sha256:502b50ee4cdc63592002d33a3de1ad43a3df84a1b811c8c99e1db4b61c272742

Observation f5b47376-fa21-4202-8631-d992952119f1 · outbound

This paper cites GPT-4 Technical Report.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.673528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.673528Z digest=sha256:0327da88a65d12085fdc2acd4a68e219e85d2af9eb37c38119d3e6adf5542d52

Observation 086660a2-e875-4904-a7e4-89324de6cea1 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.189324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.677690Z digest=sha256:afe544d43eb0032068cb866788e232c5e3f98d3d3b6ae99e7d6331c3b8c9ac3f

Observation bac8b147-e9c3-4e62-885b-61a83f3a8f49 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Representation Learning with Contrastive Predictive Coding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.681759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.681759Z digest=sha256:1b222740efcfa86e8ae9bd6abc6fa213113883d7d591eaa6028164ffa5d62303

Observation 98b54314-aae9-4b88-abc0-04277f916161 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.177490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.686682Z digest=sha256:614afdee61f5c5dbbb88e2c867c4388429633a2ae41cc43a533d6a890b601f59

Observation cff1c11f-7e9b-4164-81ff-c6a424e112c9 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Lawrence Zitnick, and Devi Parikh

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.165280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.691509Z digest=sha256:faa101666bcc828aabbd8fb39654b8ffdb2e9377677838123256b6875384fcc6

Observation fbce2f75-bd48-45c4-8f80-f34edc7488c7 · outbound

This paper cites RIO: 3d object instance re-localization in changing indoor environments.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding RIO: 3d object instance re-localization in changing indoor environments

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.151929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.695353Z digest=sha256:ef214c1949367667ec8f6ac77cff55b8791c77dbadb0000dfcf65b58d4cc3c2b

Observation c65460b7-88a2-436c-bc0e-8cb8f8c5e46b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.699005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.699005Z digest=sha256:806880a9900445b462734b1046927243faade670c6677fbb755828ce218b9e01

Observation 0db3730d-6d90-4267-8779-d8710368023b · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.703017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.703017Z digest=sha256:58a6d6ebd7f899899699561d0f01e22aaa1d2751a7be2f1d65ed1be08cb6906b

Observation 08ccbae5-c6fc-403c-bf7d-f4681a58639c · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.139782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.707464Z digest=sha256:e0be5e5861fbae2b60bf0f862b1a6167813445830c8e7e9e90b4c9befea5e941

Observation 08991898-4268-4499-a3a7-b6eaba3c4f16 · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Pointllm: Empowering large language models to understand point clouds

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.128160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.711484Z digest=sha256:73d3b06cc28149b22cec53058887b64313c904699f5bec3fbca9df9eb5e294f7

Observation dbc7b6b9-57b6-4383-98cd-9f99f5584892 · outbound

This paper cites Qwen2 Technical Report.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Qwen2 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.715463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.715463Z digest=sha256:1f00a0de8cc3d12f1156b95df1013d39e3a18187daae97d0cdff880458a371d7

Observation 36c9e5d1-6395-421b-a21b-7d656022ced1 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.116197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.719377Z digest=sha256:b7217e8596af713b1257569a4f3bfa37ab7a03c1a547d86c194da98b1e1bcd26

Observation 861e3871-867e-4a4b-9e5c-337351f765a7 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.103316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.723081Z digest=sha256:644175f42718d67c4022842428de96da728dbd9e5c93e2d960196be05ff49892

Observation 0db287ca-307d-4243-8995-3398c5513ad7 · outbound

This paper cites an unresolved cited work.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:26:20.089161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.726746Z digest=sha256:9fd8033b266f89ef89e8690053b45534a4d2fee37b8c7f6c8b55f9db16c93a28

Observation 18522524-38f7-4c96-bc66-9c1e775d4560 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Video instruction tuning with synthetic data, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.077621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.730587Z digest=sha256:52a8720c6a5b35711cef591aa458834f64998b37613496144d0651383c2d3bb8

Observation 62b66678-86f0-4874-8129-8420e6eaa992 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.065206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.735053Z digest=sha256:79c62fb188e03b0b8a8803b1c20aa8e859ea54c77cae1a66cc71379707c18620

Observation 39cc2573-ba74-45a5-9013-3404aef4e92c · outbound

This paper cites Towards learning a generalist model for embodied navigation.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Towards learning a generalist model for embodied navigation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.051470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.739504Z digest=sha256:72c9c071fc51aa30e340317995e98d5d57216f078afe3a539dc78e30eb6dbafb

Observation 676546bc-40b0-4bcd-bcc4-bb54fa735bdb · outbound

This paper cites Scanreason: Empowering 3d visual grounding with reasoning capabilities.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Scanreason: Empowering 3d visual grounding with reasoning capabilities

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.037525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.743443Z digest=sha256:bb575807a3e10f2c4b62a533137ebc8da74182cf967d1bb2130b4e042e21f933

Observation 88242134-1269-491b-a8a9-efc1291d4763 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:19.747590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:26:19.747590Z digest=sha256:4580eb20c157ed9902f9501cdf4a8796b012760e448a36ca3db1efc85f212de1

Observation a6a56442-6815-442b-b74e-271e088112a5 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.024539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.751825Z digest=sha256:8c01d9ea2ad5a5bbfa13f93b6ee5f9b3ec973717110ecc83b2bdb7bf4bfb7477

Observation f09a9d05-930a-4648-a5fa-651c141dc30f · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unifying 3d vision-language understanding via prompt- able queries

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.010380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.755795Z digest=sha256:225ed0074f76e61e39cbb0302118da6fb0ba5e09ce211074b64732e489ad50d5

Observation 14b9ca5b-87fa-4197-9d54-6a645cba941e · outbound

This paper cites sos” and “eos.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding sos” and “eos

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:19.996535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.759496Z digest=sha256:c3d7daf3f26e2d64ee884375e4a3b3043ea36d2ec127e59016ef03da990a54fd

Observation 02f8bf64-0933-4f8f-842b-ff07a356618b · outbound

This paper cites ZT” denotes zero- target, “ST.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding ZT” denotes zero- target, “ST

Reference 57

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T05:26:19.981988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.763700Z digest=sha256:a32548a01070ea64c0838caca2ebed4d4c7e9cf504c7bbd391d62d00de71404d

Observation a89ac1e7-e156-434d-a51d-0dd70c3ce81e · outbound

This paper cites an unresolved cited work.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding Unresolved cited work

Reference 2021

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T05:26:20.348445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.585725Z digest=sha256:f3c9ca1bac4b7b1877f63d2bd0afa4e507d6d14e9dfa9cf33f7accfa5dadcd48

Observation f3ac7b5b-73ec-45fb-ab0a-464dc426e83d · outbound

This paper cites 1, 2, 5, 6, 13, 14.

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding 1, 2, 5, 6, 13, 14

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:26:20.312225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T05:26:19.610352Z digest=sha256:0e022e9034b298f55584db386fe6a00909ca94dc31d4523f9acbb1d578b16a26

Pith citing papers

Observation 0c04a790-97a6-474a-b317-50b3f1a42402 · inbound

The Internet of Large Language Models: An Orchestration Framework for LLM Training and Knowledge Exchange Toward Artificial General Intelligence cites this paper.

The Internet of Large Language Models: An Orchestration Framework for LLM Training and Knowledge Exchange Toward Artificial General Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:04:34.122621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:04:34.122621Z digest=sha256:8edeef2f1a0ac69161a84f8df9a843ab6f9883cefa0298d7a58004c3e00df88c

Observation 950c6378-f069-49eb-9012-56ffe44cb7b6 · inbound

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation cites this paper.

Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:59:11.071380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:59:11.071380Z digest=sha256:f6828c830f3bc5239a89fe3ccd089cc418bb563cc3805a0c76ad9e5c46f64cb8

Observation 8f1dc900-6b44-4a4c-97c7-1fb7e79d48b9 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.768759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.768759Z digest=sha256:59ba6b075dd8f13dfbd72681ffed73b2b0ca98ac7900df4eb3e9b20306beb7c1

Observation 31acac97-89dd-432e-855f-aef3ea07a7bd · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.494993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.494993Z digest=sha256:b5bf4528620b6ceb0bf01df1e01eabfb034c199e6c0364d616c882849d7ccddb

Observation 0c60e752-4b9b-46f1-8d89-b0217515b7c5 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.872348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:d70daa4293af3f16ce0c2151c191bf4f9d6e8c164ded7ab7ef075e141fc71834

Observation a4c06c78-5058-4645-811b-0778b0d90f73 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.361125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:d90d08bc31e81fba89eee91255faa4623067e8074a00a6691443d3c3b3745a8b

Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.802988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.802988Z digest=sha256:808260dde635a310604245dc2fc195fd0d2e20a4980bdd73d62beb3199f518ce

Observation 2d0cd3b5-f8a4-4fb0-95d8-c5b13fa3842e · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.704510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.704510Z digest=sha256:26bceb8131c94474d0c532db0c35918e179127f4ce9306dca97c61cec6a5c48d

Observation 9f03306a-48ba-4fb2-813d-0922a73fea03 · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:51:19.088377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:d51e6bc5db3d1d9e8c48c56827d1dc3ec1921d9a4f19711c0f1704b46cb83d9b

Observation 6f6f27bb-e7d2-4e8a-8ae7-ca3dc3495693 · inbound

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning cites this paper.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:41:01.809714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:b1a691a2725e192a39180c4c69d23176b7cb1e7836f692e62d019cb2e347d514

Observation 40688cb6-5a6f-412d-afde-7fe87b94e9ca · inbound

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering cites this paper.

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:50:49.625343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:50:49.625343Z digest=sha256:0981cbd8af7b664a7412e36bbae1b895b1ed72db29e43b5646cd8c061cbb3cd6

Observation a34fcf9c-3f7f-4b4b-b317-abf2803f0958 · inbound

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models cites this paper.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.383190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.383190Z digest=sha256:4ff4929f60354300bb6ec5193e3c042f69124b44632c56754784b6bad459065c