Pith. sign in

Paper Citation Record · LEDGER

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.14497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14497 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:59:18.438342Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87102886-f169-4b84-a307-920d8b8cf547 · outbound

This paper cites Vqa: Visual question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Vqa: Visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.673764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.673764Z digest=sha256:755ed86f985f52ed70cc31ebd02b39362eb2758f82169f4b73c16b92e260cc28

Observation d1baaa2c-29cd-459b-8cd8-0bd917856433 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.709112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.709112Z digest=sha256:e7830dd7e9201b5acd6bffc23aefa4589bd57b92bd374d0ee7f9dbb20c6ac9cf

Observation c6563afa-6cfa-45d2-b7a5-fa893b1a44db · outbound

This paper cites Where did i leave my keys?- episodic-memory-based question answering on egocentric videos.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Where did i leave my keys?- episodic-memory-based question answering on egocentric videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.760465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.760465Z digest=sha256:8d71039454c047f50861e23c95bd31c3bc966611eef0e5d3e2a14ec6946ea2b8

Observation ba5afb2f-81d0-4d25-bcfd-99e37fb528cc · outbound

This paper cites Ad- abins: Depth estimation using adaptive bins.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Ad- abins: Depth estimation using adaptive bins

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.813459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.813459Z digest=sha256:3e7b5d0ec525ddb3784decc366da0ff4a27e72ff7207ffbe31384cbf0c97076f

Observation 53c3138e-a0d3-48de-b8b1-ae1269a4eec2 · outbound

This paper cites Local- bins: Improving depth estimation by learning local distributions.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Local- bins: Improving depth estimation by learning local distributions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.889620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.889620Z digest=sha256:c5fb2bb156ce1f3342cd3c0067ae236d6650db7d8f62512915c5bad7e81c7f4d

Observation c913e4c3-f841-42a2-a44e-5bb888d908cd · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:14.973132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:14.973132Z digest=sha256:24830366835db9f9b009721c0305bcf0bd726032d02e075934ef8174a0b57f72

Observation b00dd320-b97d-4d9c-b48c-de6115cbe39c · outbound

This paper cites Scene text visual question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Scene text visual question answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.051371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.051371Z digest=sha256:6e1515bc4a70c89be876ad20623899bd2aa75b15ffbc9dfbf20c834ac9ab0166

Observation 0904920c-fad8-4d7d-861f-109b43a2e7b4 · outbound

This paper cites Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.110105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.110105Z digest=sha256:b1a6bfd1d81658008c9d0012ace278b30a09f2c3c311ed1485ad1d656dfbe5a3

Observation eb49d6b0-ea75-4ba5-ad42-13c9129251b5 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.160109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.160109Z digest=sha256:5e26dfe51929e4d98f9059a3913acab5caf7c126bf366bb2aafa38c8e38b0191

Observation edc05bd7-0390-45f0-82c1-7042fad95dfe · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.211777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.211777Z digest=sha256:631be7254248dcad666c0ce8f2e1a767b2a8f6a3e8617e7b52c4cd2b9d8d8183

Observation 4e551c8d-b69d-4ceb-a583-a5fa3b571b01 · outbound

This paper cites Egothink: Evaluating first-person per- spective thinking capability of vision-language models.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egothink: Evaluating first-person per- spective thinking capability of vision-language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.261373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.261373Z digest=sha256:2160837a2dd59346ada1d94bbef87138615ec514e237659d7a74c2eadc73693b

Observation 7200e93a-5def-46fe-bcb1-3cace70790b4 · outbound

This paper cites In- structblip: Towards general-purpose vision-language models with in- struction tuning.Advances in neural information processing systems, 36:49250–49267, 2023.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation In- structblip: Towards general-purpose vision-language models with in- struction tuning.Advances in neural information processing systems, 36:49250–49267, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.307319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.307319Z digest=sha256:2af9371f69581c2c833cf6bca7e9895e2f8eac20e1f0389d6fc669acccad78cf

Observation 7c6fbf91-1078-4122-86ee-f7e220cad36c · outbound

This paper cites Egovqa-an egocentric video question answering benchmark dataset.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egovqa-an egocentric video question answering benchmark dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.366901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.366901Z digest=sha256:97995bc75e729543bf82e24b971ae7a6150949a85f60b2432b3934b21acdc1bc

Observation 64d09369-925b-4cb3-955f-c0ca4032c1ef · outbound

This paper cites Deep ordinal regression network for monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Deep ordinal regression network for monocular depth estimation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.370801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.370801Z digest=sha256:2772c6347d6c1f1e5c0780d86516b1cc8ff9d629a743d440dd94442ee5254a20

Observation a68823fd-3ddf-4389-a26b-3b300c8eedd2 · outbound

This paper cites Geowizard: Unleash- ing the diffusion priors for 3d geometry estimation from a single im- age.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Geowizard: Unleash- ing the diffusion priors for 3d geometry estimation from a single im- age

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.386825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.386825Z digest=sha256:62d5b0e90f10ac9ab774f3c8f3ec10360b61f08461420579ea55ca093fe29e73

Observation a0f02a9a-0189-407b-8ee2-f0ac16bb1249 · outbound

This paper cites Unsu- pervised monocular depth estimation with left-right consistency.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unsu- pervised monocular depth estimation with left-right consistency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.496125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.496125Z digest=sha256:bded72e60a8a7c5eea274f22a07531fe102090be794cc25b723108dce73b48df

Observation 8dfe3c77-c5b6-4d0e-9487-613d8a9f0a94 · outbound

This paper cites Digging into self-supervised monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Digging into self-supervised monocular depth estimation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.635625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.635625Z digest=sha256:7d3ca2e58f8c69a89d1c975edde7372ec072cc7780fafd16d4294c66733af5e8

Observation bfd9642e-3138-4ad4-90fd-e6b1130a8b85 · outbound

This paper cites Depthfm: Fast generative monocular depth estimation with flow matching.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depthfm: Fast generative monocular depth estimation with flow matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.715378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.715378Z digest=sha256:48794587e9cbd93b75872ed0d4d2daa0120facb3c42f3bf67a8a3591eb3c5788

Observation b6a72b5d-c784-499e-b672-ae41c181adca · outbound

This paper cites Towards zero-shot scale-aware monocular depth es- timation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards zero-shot scale-aware monocular depth es- timation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.800762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.800762Z digest=sha256:00bf3b0dc747aa80c8db8f8aafa8bf9e8ae00cf33e29543eacd590fd4aac718e

Observation 1043def6-f760-4ee6-a3ee-efca78ed658f · outbound

This paper cites an unresolved cited work.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.915084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.915084Z digest=sha256:49c919f6f1760878df4838d7f5cf9f9ad3d91a26337e9d595720608bf3ed31bb

Observation c3d7ac6d-e563-47cd-95ef-b71a75f18254 · outbound

This paper cites Ego- taskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Ego- taskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:15.989903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:15.989903Z digest=sha256:251b5b28159804def0d7e580eba7da0d2a81e9919a5d58923c1b1b7c104fe731

Observation bedfa933-1617-41d8-be95-6468ee25ca17 · outbound

This paper cites Repurposing diffusion- based image generators for monocular depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Repurposing diffusion- based image generators for monocular depth estimation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.054865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.054865Z digest=sha256:7942ea2673dadab0707e8390eb94a3478b5ef0fb355889a9ebdd6f2cc9836344

Observation aac8e7bc-8820-4dba-a97b-c697b76ec942 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.130687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.130687Z digest=sha256:59fae8a9ee5ac8adda38519cd25bcc1ff7fd4797317c8d6e5a9349af91dc825a

Observation f2837398-d58e-4ca5-9547-b1366468f428 · outbound

This paper cites Egocross: Benchmarking multimodal large language models for cross- domain egocentric video question answering.arXiv preprint arXiv:2508.10729, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egocross: Benchmarking multimodal large language models for cross- domain egocentric video question answering.arXiv preprint arXiv:2508.10729, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.237961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.237961Z digest=sha256:c585205bdbf48d1a90117f9573b32370620ba6481ca94d72f91d76a9efeb236e

Observation 3ee1175e-1636-415f-a0b1-67519baae3d8 · outbound

This paper cites BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.314451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.314451Z digest=sha256:b7a84b0d307b3ebc2f8b1f4f1e4092b146b35cafbbbaa5c4a7856eaa08983d4a

Observation 8b4ab995-c0aa-4481-b14c-79666e8cd61d · outbound

This paper cites Patchfusion: An end-to-end tile-based framework for high-resolution monocular met- ric depth estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Patchfusion: An end-to-end tile-based framework for high-resolution monocular met- ric depth estimation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.396078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.396078Z digest=sha256:105b5c970514cbae41a84203b8dfc9ef4cc3b106f5b233f9816fe6fa56b5b1ea

Observation 4e457fda-1d07-4c99-af11-3dc7964ebf5b · outbound

This paper cites Unibind: Llm-augmented unified and balanced representation space to bind them all.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Unibind: Llm-augmented unified and balanced representation space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.465821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.465821Z digest=sha256:e3c2127ed700382f30095a5e1f7ce2be572f2401a734929a1e06aa05f652b0b4

Observation 63523dd7-515a-4f88-a91d-3a270c51eab6 · outbound

This paper cites Realrag: Retrieval- augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Realrag: Retrieval- augmented realistic image generation via self-reflective contrastive learning.arXiv preprint arXiv:2502.00848, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.610317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.610317Z digest=sha256:d8f5c83d7adb0140d2297a64da8642d14c9141fddddaec4e07dd3dc5e550a9d7

Observation f9a25bdb-33de-4010-8f4d-d7ec315278b4 · outbound

This paper cites Single image depth estimation: An overview.Digital Signal Processing, 123: 103441, 2022.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Single image depth estimation: An overview.Digital Signal Processing, 123: 103441, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.686243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.686243Z digest=sha256:cf112f2a468d227faada983351651adb8e14c1d3a2db7f1a090550471e6d1fc9

Observation 75867d58-7767-4496-b5a4-60dd58376bac · outbound

This paper cites UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.809240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.809240Z digest=sha256:8bfc7f9cdd3118a480882e5c900b54c6baec6cb157551e2c4178b3d217ae94e1

Observation 19aa8d47-db44-4ab3-8735-4d50f542d48e · outbound

This paper cites Towards robust monocular depth estimation: Mix- ing datasets for zero-shot cross-dataset transfer.IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637,.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards robust monocular depth estimation: Mix- ing datasets for zero-shot cross-dataset transfer.IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.895759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.895759Z digest=sha256:144f78e98aac40c75c2fb29ffb247259850a833bf73031cf1efafb0201ae3ade

Observation aa72a3f3-6776-49ce-a8db-4fe5854726ee · outbound

This paper cites Monocular depth estimation us- ing neural regression forest.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Monocular depth estimation us- ing neural regression forest

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:16.960864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:16.960864Z digest=sha256:76bf379e59e02b4ddddbd1b672adc8d39551d02be65d0f4a0f4555471b0f9bb9

Observation e09c48aa-acf9-4216-90f2-c1c3b39690db · outbound

This paper cites Towards vqa models that can read.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Towards vqa models that can read

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.077636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.077636Z digest=sha256:7f64c6c3f7a10ef3ae4b509d5452981353c2c5862edc2640d8a449c39ec6fad7

Observation 8f5c2a54-2c1c-46d1-a50e-8ff3b43a3dcf · outbound

This paper cites AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.186399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.186399Z digest=sha256:1b3bce299ecbd71367fa821a32e62889d5dd8c229f9d85ec6f9fa2c5ab30ce11

Observation 58287337-48c1-4af8-bd02-c2ce6333dd5b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.304733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.304733Z digest=sha256:efdd128db0755710921ed4ded5344b2052f5274d6112f68e2adb3ec6cef74fa6

Observation be175290-bc13-4f19-a30c-748ab95edd3e · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Emu3: Next-Token Prediction is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.406676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.406676Z digest=sha256:c71b935421e7230f136957aed0c859faac65651f1a05cc8d09785c68699db4fd

Observation c0f015b3-0808-4241-8a97-b5a6fd790104 · outbound

This paper cites Fastdepth: Fast monocular depth estimation on em- bedded systems.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Fastdepth: Fast monocular depth estimation on em- bedded systems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.515648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.515648Z digest=sha256:8c8993f7b997241d028a1a08bf56dde50dc29a4f752ce72405ccb2ba20df4a82

Observation 029520bc-4b9d-41fb-a7a1-c1ca069ee1b2 · outbound

This paper cites Visual question answering: A survey of methods and datasets.Computer Vision and Image Understanding, 163:21–40, 2017.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Visual question answering: A survey of methods and datasets.Computer Vision and Image Understanding, 163:21–40, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.647380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.647380Z digest=sha256:162a9d21d71907e682a1d19037156b59df6c38e356c4e7c5eb8ba19f91113bf6

Observation b3001eee-3042-4fda-be0e-89e8fe23942b · outbound

This paper cites Egolife: Towards egocentric life assistant.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egolife: Towards egocentric life assistant

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.731562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.731562Z digest=sha256:3e7bd0e22e49a70f52471043d4b24801caa8b74783cbd24d7d47c6769ee1aeee

Observation a0246b27-dfdb-4ffc-9442-702193a6e75b · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37:21875–21911, 2024.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Depth anything v2.Advances in Neural Information Processing Systems, 37:21875–21911, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.802293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.802293Z digest=sha256:2b08e3dd00d768516b921c6ee73c72e99337e988990b360381f8593c2cf42643

Observation 9a8cce54-98b9-488c-ad2a-3b209436c2fe · outbound

This paper cites MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:17.934979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:17.934979Z digest=sha256:9e1fb4484c0eb9d64261aad9986400a1a1f7a8cff76a7cd128bafe2e05a39fcf

Observation 967597eb-dbee-467e-b868-02cb603a993a · outbound

This paper cites Metric3d: Towards zero- shot metric 3d prediction from a single image.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Metric3d: Towards zero- shot metric 3d prediction from a single image

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.058550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.058550Z digest=sha256:6cdd85d813eb49fb9b8dda72fae5a3c9b59c147677a082216755e8c16a9a5259

Observation 857d1fd3-24c7-4d7b-ad72-d2e06c774bbf · outbound

This paper cites NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.177460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.177460Z digest=sha256:174a9d8edf9ce39950f509e5aa4a6a5a3658e12f924fd9cb0f13854868aab383

Observation ace98f0a-0237-417a-a0ad-8f059242d5c8 · outbound

This paper cites Egonight: Towards egocentric vision under- standing at night with a challenging benchmark.arXiv preprint arXiv:2510.06218, 2025.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egonight: Towards egocentric vision under- standing at night with a challenging benchmark.arXiv preprint arXiv:2510.06218, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.266673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.266673Z digest=sha256:91441fa696962b660ca856e26dd4138cea35475bb427fd018f24ec82a7744858

Observation 9400028d-e145-47f0-830b-fc6eae73cfb5 · outbound

This paper cites Egotextvqa: Towards egocentric scene-text aware video question answering.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation Egotextvqa: Towards egocentric scene-text aware video question answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.334406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.334406Z digest=sha256:971b312f4fca75717b5adc4efb7a2b231245fa413c947f0d574b7e3f79820bf2

Observation d5169957-d4f4-42c6-8280-353bf109e023 · outbound

This paper cites ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos.

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:59:18.438342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:59:18.438342Z digest=sha256:c5d6f657c43c0a7e9af50ce9df5a80aecd6bddfad1392417c05696ab17d20539

Pith citing papers

No inbound Pith citation observations are available.