Pith. sign in

Paper Citation Record · LEDGER

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

As of 12 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.14497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14497 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:59:18.438342Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87102886-f169-4b84-a307-920d8b8cf547 · outbound

This paper cites Vqa: Visual question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Vqa: Visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.673764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.673764Z digest=sha256:f1303761a71b8a9ae6902c7b0a6661fea68b77412fbcd43a4c9c16cfe31cbf93

Observation d1baaa2c-29cd-459b-8cd8-0bd917856433 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.709112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.709112Z digest=sha256:7d1214de4fde8e31084cd72484be374b7e8396f307798d5855b9a2b8c79caee5

Observation c6563afa-6cfa-45d2-b7a5-fa893b1a44db · outbound

This paper cites Where did i leave my keys?- episodic-memory-based question answering on egocentric videos.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Where did i leave my keys?- episodic-memory-based question answering on egocentric videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.760465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.760465Z digest=sha256:889b3694b42e11fa0da9213f0010deb6712e2268a88b3383c766ea1227c7abd1

Observation ba5afb2f-81d0-4d25-bcfd-99e37fb528cc · outbound

This paper cites Ad- abins: Depth estimation using adaptive bins.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Ad- abins: Depth estimation using adaptive bins

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.813459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.813459Z digest=sha256:613363eabf8ee4662f3ca30f88a74f41394467e0c69ec7fd0ae5763f38f5e72c

Observation 53c3138e-a0d3-48de-b8b1-ae1269a4eec2 · outbound

This paper cites Local- bins: Improving depth estimation by learning local distributions.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Local- bins: Improving depth estimation by learning local distributions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.889620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.889620Z digest=sha256:ec52ea38bac759beb90762e3da7c522737a259965d8df3ddf84df1731c30ea40

Observation c913e4c3-f841-42a2-a44e-5bb888d908cd · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.973132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.973132Z digest=sha256:e8d50beb2019b3f3aae41436989b537698b1699dbb3ddb45019ad8afe639f91b

Observation b00dd320-b97d-4d9c-b48c-de6115cbe39c · outbound

This paper cites Scene text visual question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Scene text visual question answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.051371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.051371Z digest=sha256:e8fbf161523f7847c6760675586f2ab3a63988b91dcd1b5e5d0168619d830366

Observation 0904920c-fad8-4d7d-861f-109b43a2e7b4 · outbound

This paper cites Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.110105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.110105Z digest=sha256:3f99f6f693ac7bbcd3d493d2476ba41a4ec1f151e1687a0b261ad1c8183429b4

Observation eb49d6b0-ea75-4ba5-ad42-13c9129251b5 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.160109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.160109Z digest=sha256:caf8dd9b8886f3f2e9625d87937a3f04edbe51e3fac26b180321740c62a6f469

Observation edc05bd7-0390-45f0-82c1-7042fad95dfe · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.211777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.211777Z digest=sha256:5cda4d4a0e395fd2f672d34bda8a4bf168954f63445ddb3f92965533a6bc849e

Observation 4e551c8d-b69d-4ceb-a583-a5fa3b571b01 · outbound

This paper cites Egothink: Evaluating first-person per- spective thinking capability of vision-language models.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egothink: Evaluating first-person per- spective thinking capability of vision-language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.261373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.261373Z digest=sha256:dd730f898f1d391781d7f4610ee3afa88223d8cf7fbb5750e1498f1cd02e8fff

Observation 7200e93a-5def-46fe-bcb1-3cace70790b4 · outbound

This paper cites In- structblip: Towards general-purpose vision-language models with in- struction tuning.Advances in neural information processing systems, 36:49250–49267, 2023.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation In- structblip: Towards general-purpose vision-language models with in- struction tuning.Advances in neural information processing systems, 36:49250–49267, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.307319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.307319Z digest=sha256:4b8c7f51f76896c9c6e22e439b3fcaf113d474fcaa73522a7002469cef96f1e8

Observation 7c6fbf91-1078-4122-86ee-f7e220cad36c · outbound

This paper cites Egovqa-an egocentric video question answering benchmark dataset.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egovqa-an egocentric video question answering benchmark dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.366901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.366901Z digest=sha256:b07ff6c003f843473530dad524acd5291b27b7bc6a1c5235a4085153abc604aa

Observation 64d09369-925b-4cb3-955f-c0ca4032c1ef · outbound

This paper cites Deep ordinal regression network for monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Deep ordinal regression network for monocular depth estimation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.370801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.370801Z digest=sha256:4603e0e3d622affbe598da337a3d1f7af8812202f6df49063cb43aa6e8d07e0b

Observation a68823fd-3ddf-4389-a26b-3b300c8eedd2 · outbound

This paper cites Geowizard: Unleash- ing the diffusion priors for 3d geometry estimation from a single im- age.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Geowizard: Unleash- ing the diffusion priors for 3d geometry estimation from a single im- age

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.386825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.386825Z digest=sha256:bdfeb5a27843b9c562f9d72b13d276166ed5a5d43d439467663b24d5acc7fdf3

Observation a0f02a9a-0189-407b-8ee2-f0ac16bb1249 · outbound

This paper cites Unsu- pervised monocular depth estimation with left-right consistency.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unsu- pervised monocular depth estimation with left-right consistency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.496125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.496125Z digest=sha256:d6ca7a8f70d9bc1933751cca450b00130115e2523dffa1fa45906e7b73a85822

Observation 8dfe3c77-c5b6-4d0e-9487-613d8a9f0a94 · outbound

This paper cites Digging into self-supervised monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Digging into self-supervised monocular depth estimation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.635625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.635625Z digest=sha256:543aba9594a4740bf470fc97aba2b66e69222ca2f74128ff8eff0a83fe4ce23f

Observation bfd9642e-3138-4ad4-90fd-e6b1130a8b85 · outbound

This paper cites Depthfm: Fast generative monocular depth estimation with flow matching.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depthfm: Fast generative monocular depth estimation with flow matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.715378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.715378Z digest=sha256:5623a4c74d85c9e5436e3a37916460a024b65d97f48bf9fc7d54ef870f6a897b

Observation b6a72b5d-c784-499e-b672-ae41c181adca · outbound

This paper cites Towards zero-shot scale-aware monocular depth es- timation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards zero-shot scale-aware monocular depth es- timation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.800762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.800762Z digest=sha256:b2c28898005a961ca51d67cf3dbe8b25fb8e10ef39dc9f5e1cb906a5d0fe9294

Observation 1043def6-f760-4ee6-a3ee-efca78ed658f · outbound

This paper cites an unresolved cited work.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.915084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.915084Z digest=sha256:4de91b02182b079396620b2de7de8c04b23582b94fffe69a7561bea736afb81c

Observation c3d7ac6d-e563-47cd-95ef-b71a75f18254 · outbound

This paper cites Ego- taskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Ego- taskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.989903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.989903Z digest=sha256:941d8d14be10544d81fdfe671d53f2d8408ca13ff8be6809255f75eb7e638895

Observation bedfa933-1617-41d8-be95-6468ee25ca17 · outbound

This paper cites Repurposing diffusion- based image generators for monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Repurposing diffusion- based image generators for monocular depth estimation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.054865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.054865Z digest=sha256:a1da10ec33c781bba8c51eb3bc3b882e0ee6b595fa558052a58b39cfc98c8082

Observation aac8e7bc-8820-4dba-a97b-c697b76ec942 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.130687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.130687Z digest=sha256:4704ded76555fc17e0ad7ba5d4ef3ceb2e23931f656f805068497ac824934780

Observation f2837398-d58e-4ca5-9547-b1366468f428 · outbound

This paper cites Egocross: Benchmarking multimodal large language models for cross- domain egocentric video question answering.arXiv preprint arXiv:2508.10729, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egocross: Benchmarking multimodal large language models for cross- domain egocentric video question answering.arXiv preprint arXiv:2508.10729, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.237961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.237961Z digest=sha256:0ccbd64f794f4b88de16f1a06f819e925c997f3510a3f4c961b90530d5ecad4c

Observation 3ee1175e-1636-415f-a0b1-67519baae3d8 · outbound

This paper cites BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.314451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.314451Z digest=sha256:1d35dc1e88cce79156d962ee297004b0031f0bb40c5d06a7078694e4246534e2

Observation 8b4ab995-c0aa-4481-b14c-79666e8cd61d · outbound

This paper cites Patchfusion: An end-to-end tile-based framework for high-resolution monocular met- ric depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Patchfusion: An end-to-end tile-based framework for high-resolution monocular met- ric depth estimation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.396078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.396078Z digest=sha256:840e6e3c326428b249aa941043ca15107b467e997de9fe41cfff1a8b4fe52395

Observation 4e457fda-1d07-4c99-af11-3dc7964ebf5b · outbound

This paper cites Unibind: Llm-augmented unified and balanced representation space to bind them all.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unibind: Llm-augmented unified and balanced representation space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.465821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.465821Z digest=sha256:a3a9b9c9e450a165b147d61fd9d4b27413fd1a31c3803cd99e798a8471bae0a3

Observation 63523dd7-515a-4f88-a91d-3a270c51eab6 · outbound

This paper cites Realrag: Retrieval- augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Realrag: Retrieval- augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.610317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.610317Z digest=sha256:95356b0b6a5dc9111df9f4a9e983bee3316dd0c8a10f56838cb53ce16efbef91

Observation f9a25bdb-33de-4010-8f4d-d7ec315278b4 · outbound

This paper cites Single image depth estimation: An overview.Digital Signal Processing, 123: 103441, 2022.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Single image depth estimation: An overview.Digital Signal Processing, 123: 103441, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.686243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.686243Z digest=sha256:9ca25aa8f3befbbe150a30bbcef4fbb79c45adfd8a64a80f156cf83b2b471c10

Observation 75867d58-7767-4496-b5a4-60dd58376bac · outbound

This paper cites UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.809240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.809240Z digest=sha256:370581e4d2a3aca58f0f43129f3bdabae73b6e8e2c8ef42ecaf179ca860706a7

Observation 19aa8d47-db44-4ab3-8735-4d50f542d48e · outbound

This paper cites Towards robust monocular depth estimation: Mix- ing datasets for zero-shot cross-dataset transfer.IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637,.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards robust monocular depth estimation: Mix- ing datasets for zero-shot cross-dataset transfer.IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.895759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.895759Z digest=sha256:c59a5263a17734f21ddb0f501e3c750de362421108d2a6ded4263bdb1e0d3783

Observation aa72a3f3-6776-49ce-a8db-4fe5854726ee · outbound

This paper cites Monocular depth estimation us- ing neural regression forest.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Monocular depth estimation us- ing neural regression forest

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.960864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.960864Z digest=sha256:1c604a0de0e5e784694f5bd22a6b22bf6011c80eda023a5a8b32cb532e021021

Observation e09c48aa-acf9-4216-90f2-c1c3b39690db · outbound

This paper cites Towards vqa models that can read.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards vqa models that can read

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.077636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.077636Z digest=sha256:a0b6f23f74a547eba14e5388659359af3ec5f3d94baf9036ee1dadc9d3be5915

Observation 8f5c2a54-2c1c-46d1-a50e-8ff3b43a3dcf · outbound

This paper cites AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.186399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.186399Z digest=sha256:528d4baa5a8e651ffb22feb6d9a2fb467aff4a6704f4446a34aed5bcd2b9918d

Observation 58287337-48c1-4af8-bd02-c2ce6333dd5b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.304733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.304733Z digest=sha256:26e5b54e0551ebe9d07816a08062eaa4683922d2cdadf228fad1378ce574a46f

Observation be175290-bc13-4f19-a30c-748ab95edd3e · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Emu3: Next-Token Prediction is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.406676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.406676Z digest=sha256:ace5b09d05e6796acff192e0040b35d14b479c1c37d29fb60898835cfd720c27

Observation c0f015b3-0808-4241-8a97-b5a6fd790104 · outbound

This paper cites Fastdepth: Fast monocular depth estimation on em- bedded systems.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Fastdepth: Fast monocular depth estimation on em- bedded systems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.515648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.515648Z digest=sha256:8c0550cb948d82501757c5b4d5ebcd36a73057c4916e0e7d8e900538e30a5019

Observation 029520bc-4b9d-41fb-a7a1-c1ca069ee1b2 · outbound

This paper cites Visual question answering: A survey of methods and datasets.Computer Vision and Image Understanding, 163:21–40, 2017.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Visual question answering: A survey of methods and datasets.Computer Vision and Image Understanding, 163:21–40, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.647380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.647380Z digest=sha256:b36b442c462493244cdf5f58fb3521627ab17194a5dfe4c8afb05143d0eca646

Observation b3001eee-3042-4fda-be0e-89e8fe23942b · outbound

This paper cites Egolife: Towards egocentric life assistant.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egolife: Towards egocentric life assistant

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.731562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.731562Z digest=sha256:2beffa454467eea2f6b4b70364e63bf9c456363c2be8fb0f5fb55ef6987b192f

Observation a0246b27-dfdb-4ffc-9442-702193a6e75b · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37:21875–21911, 2024.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depth anything v2.Advances in Neural Information Processing Systems, 37:21875–21911, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.802293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.802293Z digest=sha256:1b20051c8e6003870dcd6847ad62b4bf5ceb8e00dae1f9a18158dd81afb8d989

Observation 9a8cce54-98b9-488c-ad2a-3b209436c2fe · outbound

This paper cites MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.934979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.934979Z digest=sha256:eba5ae2fef26bdddcd00670f47291670accbb293f9b37b28bcef4bfd97563a66

Observation 967597eb-dbee-467e-b868-02cb603a993a · outbound

This paper cites Metric3d: Towards zero- shot metric 3d prediction from a single image.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Metric3d: Towards zero- shot metric 3d prediction from a single image

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.058550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.058550Z digest=sha256:ce90dd915465f9892ae384ca0d8eda0ecc23b6ace337cd5c69a66ce487ce378c

Observation 857d1fd3-24c7-4d7b-ad72-d2e06c774bbf · outbound

This paper cites NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.177460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.177460Z digest=sha256:02156f8dfb850f2f5f30a118d9bffed284c876d6358197d3d6ad49bd39755ba5

Observation ace98f0a-0237-417a-a0ad-8f059242d5c8 · outbound

This paper cites Egonight: Towards egocentric vision under- standing at night with a challenging benchmark.arXiv preprint arXiv:2510.06218, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egonight: Towards egocentric vision under- standing at night with a challenging benchmark.arXiv preprint arXiv:2510.06218, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.266673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.266673Z digest=sha256:6e644c762d280aacd63dd9b297e234126f6883194f6299598c27d36b24b18b41

Observation 9400028d-e145-47f0-830b-fc6eae73cfb5 · outbound

This paper cites Egotextvqa: Towards egocentric scene-text aware video question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egotextvqa: Towards egocentric scene-text aware video question answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.334406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.334406Z digest=sha256:aba8813d90be0378b3de944d0a5cc536a4714ad97ec3afe2a099b0a0ed6e2682

Observation d5169957-d4f4-42c6-8280-353bf109e023 · outbound

This paper cites ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.438342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.438342Z digest=sha256:b65fb787070f301f03398dc8d66f3f249ef2d80c1c6427731848be80f5b6caab

Pith citing papers

No inbound Pith citation observations are available.