Pith. sign in

Paper Citation Record · LEDGER

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos

As of 13 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 1 inbound Pith citation observation for arXiv:2411.09145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.09145 v4

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:07:40.137416Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:13:42.052386Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:20:26.104422Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc871b51-c499-4303-bab3-dacf8162cb8a · outbound

This paper cites Map-free visual relocalization: Metric pose relative to a single im- age.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Map-free visual relocalization: Metric pose relative to a single im- age

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.722354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.722354Z digest=sha256:09e1f986ebd65060f9038c4342c8ac1f808621aafb3c4103d83bc77ec9826a61

Observation 26c7e916-6c5c-4731-957b-d3c2f3e253a5 · outbound

This paper cites Affordances from human videos as a versatile representation for robotics.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Affordances from human videos as a versatile representation for robotics

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.727486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.727486Z digest=sha256:33c0bf7b6f31c32b214cf51e69b3c00dc21ad4010fee4022db63c0163ad771ee

Observation 62af5fcf-ab0f-49d1-a46f-fd0ce6277d14 · outbound

This paper cites Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.732033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.732033Z digest=sha256:64336a50a21f9dfad117a0efd9aaf7decfc62943f5dc2138b67522e87c43c15b

Observation 4cc3440b-6392-450e-9c64-c491634de4f5 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.736384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.736384Z digest=sha256:de96ba8c900d48c126e5000d62f81a1ccfe06d36786a1068a12878e522d2c31c

Observation 9065cc68-ab5f-4f63-8b2a-e69e1d536228 · outbound

This paper cites Unsuper- vised scale-consistent depth and ego-motion learning from monocular video.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Unsuper- vised scale-consistent depth and ego-motion learning from monocular video

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.740804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.740804Z digest=sha256:68003418f296a4c9b9a52fdf7573d7ff75208e0f64b8f84d38bc27bcedce294a

Observation 833882f7-bd0c-4c9c-9582-8167e6ca95a9 · outbound

This paper cites MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.745048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.745048Z digest=sha256:d73dedc390532073305a216e31660f3838a987ea1891d7fdda782fa96f4975db

Observation 1bad3f8f-7ece-4ea4-b145-a79d8cb0839d · outbound

This paper cites Orb-slam3: An ac- curate open-source library for visual, visual–inertial, and multimap slam.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Orb-slam3: An ac- curate open-source library for visual, visual–inertial, and multimap slam

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.749526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.749526Z digest=sha256:4772f0915abda9bda9ac22dba6ec745788e80a916debe0ea363c3e96396d6f5b

Observation a1b3097b-edce-47ed-8e39-7db7f4ac045e · outbound

This paper cites Leap-vo: Long-term effective any point tracking for visual odometry.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Leap-vo: Long-term effective any point tracking for visual odometry

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.753033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.753033Z digest=sha256:2d87cf5fe1adbca30fe2fc6f0e83b56d7a8a114d789a1a0d047966c1d65a59b0

Observation 9407b8ef-acd2-40a5-9e72-155f7d6edcf2 · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Vision Transformer Adapter for Dense Predictions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.756515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.756515Z digest=sha256:9569b3a3dff7da116602ae3d4d0f39f5fb5e30cf5da20502d9e0dcbb388a33e7

Observation 0bdcb2d6-179f-4774-8467-20c23a8c49d7 · outbound

This paper cites Treating mo- tion as option to reduce motion dependency in unsuper- vised video object segmentation.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Treating mo- tion as option to reduce motion dependency in unsuper- vised video object segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.761226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.761226Z digest=sha256:4ad1fd57cd569ee6d9775040dc4a44e2efdf1f3df862bc9a415a2eeb720c4508

Observation b5c51ec0-2d47-4095-b1e3-1b9dd05a276e · outbound

This paper cites Deep global registration.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Deep global registration

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.765117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.765117Z digest=sha256:c4e7fc73bd5b0ffd3a1c147e393e1fd5dbb359097734f9d76f76f900e0392de6

Observation 4f4aaee8-351a-4009-982b-7127eae75a6f · outbound

This paper cites DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.769174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.769174Z digest=sha256:faca3e45469d20f5e540a9693c8f805635bbf2c5ba379fa7f292b4987bfc7421

Observation e15d13c9-ea4b-4ef7-868c-8e97b7952971 · outbound

This paper cites 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.773462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.773462Z digest=sha256:e13ef49c03c7997922bfb5ee62154308c8c1233afce2530a993cc959b24eab83

Observation e5f0714b-1fc7-40da-a2af-1a2b722cf532 · outbound

This paper cites Global structure-from-motion by similarity averaging.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Global structure-from-motion by similarity averaging

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.778342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.778342Z digest=sha256:314c7b25cdd6c18993aa562c29ec12228766b3a0a9fd5b2374bf47c2d4544bb5

Observation f4c9634c-7571-4226-bbd5-e64b9624be02 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.782785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.782785Z digest=sha256:0a08048267914055aa348d324fbbb5251315f1f707c507d20274ee7f42bd95cf

Observation ed8293ed-8ba4-4990-9f28-013ee4d55557 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.787070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.787070Z digest=sha256:f507aaedb274d3826841171de9829d8fb1420e6ca3989155305a1f4591980cae

Observation 8c0911ff-fafe-4898-9ab5-e5bee2ac9ded · outbound

This paper cites MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.792383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.792383Z digest=sha256:f621f542c07f09f05d04a6c211f4fa6a37282e784d9e388191c42331d91da341

Observation c3d7c593-8126-46ac-bf5a-19ab458f272f · outbound

This paper cites FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.797146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.797146Z digest=sha256:4ccac601fb8f5ca4d21f2bae0225727cf5972a82ef2348c1009229b7e023eab5

Observation 3292ec59-0f45-4b18-9a05-210ee33f3a1b · outbound

This paper cites Light3R-SfM: Towards Feed-forward Structure-from-Motion.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Light3R-SfM: Towards Feed-forward Structure-from-Motion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.801035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.801035Z digest=sha256:2c1a797b2c51ab4590e15620313b2833bba1fc8acb06baf7c1d25ac841d564ef

Observation 35be13ad-6993-474c-9e69-42332e07921e · outbound

This paper cites Arctic: A dataset for dexterous bimanual hand- object manipulation.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Arctic: A dataset for dexterous bimanual hand- object manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.805305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.805305Z digest=sha256:2a4189467189a8c851e72698bcc0ba99912f4f1227087e5be8d0ddcf5f4866d8

Observation ae25698e-2d92-4cd6-9f0b-0dd4687260c8 · outbound

This paper cites Kaolin: A pytorch library for accelerating 3d deep learning research.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Kaolin: A pytorch library for accelerating 3d deep learning research

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.809042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.809042Z digest=sha256:90d9655757a76c705f34c2914601ec2cca2d3cd0ff458fd6b0bcf1396f3c82b6

Observation 1c7d14fc-90b6-45a4-9c5e-40ecf0138573 · outbound

This paper cites First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos First-person hand action bench- mark with rgb-d videos and 3d hand pose annotations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.813130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.813130Z digest=sha256:0db62dec2c5cf878be3eb73a4c9deaba745e6ae02369b289a37091b9eb8c618c

Observation b3cc91b5-5adf-4e96-9c31-f6075fc1ac7c · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.816847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.816847Z digest=sha256:30871da95439974b160daaa4a368635bf36ad805ab88c94021ed4ea23bfe9757

Observation 019b77d5-f088-438a-a5e7-c825514ff200 · outbound

This paper cites Deep relu networks have surprisingly few activation patterns.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Deep relu networks have surprisingly few activation patterns

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.820503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.820503Z digest=sha256:91c2abf5417bc6815e2dea7bf5564ee9b1b571a87eae356ca4a30deb3f40382d

Observation 2b4b57a4-5191-499a-ac66-8d0c8262d8ca · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.824475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.824475Z digest=sha256:7a65a5c95affd54e6de7c34c49f63b1cdea859c623ed8cbe5b278cc883e164dd

Observation 13a25523-1650-48aa-9a76-c70de8e6d72c · outbound

This paper cites Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.829236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.829236Z digest=sha256:d13d7028e8400d35722d6941dde34fdf6139f0c4ad58c9cd844c9b804337cb29

Observation fa2ebabf-c41f-448f-82fe-c1b11a3272c3 · outbound

This paper cites CoTracker: It is Better to Track Together.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos CoTracker: It is Better to Track Together

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.833351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.833351Z digest=sha256:d6b5dc41d21ec5dc136ec38cfa1f21a6fddca9e19e228c8ddfcc433ac72c4556

Observation 0cad78d6-a241-470d-ac6b-00b0a4c4eb52 · outbound

This paper cites Fast encoder- based 3d from casual videos via point track processing.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Fast encoder- based 3d from casual videos via point track processing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.837229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.837229Z digest=sha256:a7fd7a3dfa4f9599b42499d22e1006147f698eacf5d5016bee4f804c690830d7

Observation fa4ded9c-035d-4aff-8d63-78e54952c664 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos 3d gaussian splatting for real-time radiance field rendering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.841062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.841062Z digest=sha256:878b1885b71e8f5dcd60077f8aac92e00a124c910872ffe61f94d1f428dffb8d

Observation d62c97b7-dcb3-4e44-b01e-b24637a62dd4 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Adam: A Method for Stochastic Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.845079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.845079Z digest=sha256:90176480db4139514a97b63b572eeaee561612ab48b5827514ec28f5537f0e4d

Observation 7d0781c7-f5fc-4083-b656-ae365e645e0f · outbound

This paper cites Segment anything.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Segment anything

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.849213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.849213Z digest=sha256:8339ec231f5eb6eaa1c11e4afca23009deac27b7f50422149414fce8abf498b6

Observation ad4f867f-7741-4622-a6c9-c46622d340f3 · outbound

This paper cites Ro- bust consistent video depth estimation.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Ro- bust consistent video depth estimation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.852914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.852914Z digest=sha256:78fadbc13fd27cf2c741e6cfa839c73a03e255b8c6e7c7bae7e69b73e166e17b

Observation e267ad21-5df7-4496-b9e5-875670612683 · outbound

This paper cites TAPVid-3D: A Benchmark for Tracking Any Point in 3D.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos TAPVid-3D: A Benchmark for Tracking Any Point in 3D

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.856645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.856645Z digest=sha256:df46952269f102d94b1cbaf3a0f4a937a707ecc48f88a9ef2d21cf823a35d6bb

Observation a581af7c-ea2a-49d6-b74e-c2aeb18a21e4 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos H2o: Two hands manipulating objects for first person interaction recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.861008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.861008Z digest=sha256:bd8917423e6c8b94a81f8dd0159eb658e3bd6f1b7a3374aba0dbaa2ff460715e

Observation 771f7275-ead0-4751-930e-73dcd92293d1 · outbound

This paper cites MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.864731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.864731Z digest=sha256:41756bcb9b20e07450db67200a2c0c5da29784bb079ce0e3ab2f730d6c915be0

Observation 410c6c70-28ca-4277-9c74-ad7d54ac8563 · outbound

This paper cites Grounding Image Matching in 3D with MASt3R.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Grounding Image Matching in 3D with MASt3R

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.869186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.869186Z digest=sha256:8c2838c682bee306ba72d1b23a7be2426cf3a24b5cfcd541393079bcfc068404

Observation 6bbdfc7a-5bd7-4253-934f-717519d1ca98 · outbound

This paper cites Egocentric pre- diction of action target in 3d.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Egocentric pre- diction of action target in 3d

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.873375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.873375Z digest=sha256:a41ac73ac47e9a7f7c27176627eb4b09db35906e5e287b5296a1d1cac4d20b46

Observation 360c94b3-c92c-473a-bf4c-538f9bf39895 · outbound

This paper cites MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.876857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.876857Z digest=sha256:a947a12322acc019db5a862656390b93d15d1fbb9b8eb68f8d9743d5efbfb42f

Observation fb0ff72c-6d51-4d1e-a093-473f27697463 · outbound

This paper cites Feed-forward bullet-time reconstruction of dynamic scenes from monoc- ular videos.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Feed-forward bullet-time reconstruction of dynamic scenes from monoc- ular videos

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.881571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.881571Z digest=sha256:75aa77a079de6cd10f64b3db607952a7ce065412a6aebad7e462ae951b56de8d

Observation 7e37d1ab-36be-4a71-aa3e-5b105c8587d5 · outbound

This paper cites Few- shot parameter-efficient fine-tuning is better and cheaper than in-context learning.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Few- shot parameter-efficient fine-tuning is better and cheaper than in-context learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.886105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.886105Z digest=sha256:a324637f981b98805060bc6d693381c8996e410b7ec1716f09c60913edc115af

Observation ba609dcf-acb6-4e2b-9b98-02be8df5d7ab · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.890781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.890781Z digest=sha256:61b9f630267302e8641e7826770b407179a26919f5294c2f86b90deacf103074

Observation 1dc9a8f6-7f05-422c-8669-29f0d104494c · outbound

This paper cites Joint esti- mation of pose, depth, and optical flow with a competition– cooperation transformer network.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Joint esti- mation of pose, depth, and optical flow with a competition– cooperation transformer network

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.895402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.895402Z digest=sha256:ba86a2c3ebe933e23befb7cc0f1affdd6616cd60c2957264b8243cd0508f3aeb

Observation 5cf17ea6-3ca9-440a-9034-3e63a67ba67b · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.899433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.899433Z digest=sha256:c7bc1ba81af09ec30dcef9b492ab0db6d2e689c3f36b28ea0dfa48e631c18776

Observation 01478555-27c0-44ca-88c7-39f22bf8443b · outbound

This paper cites Robust dynamic radi- ance fields.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Robust dynamic radi- ance fields

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.903113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.903113Z digest=sha256:fe52a2be273ea06fa2f1483ef22d85ed5f7823c2c28c219711ca752aa8395d20

Observation fd752f31-870b-4b56-b35c-98f93301cdfa · outbound

This paper cites Align3R: Aligned Monocular Depth Estimation for Dynamic Videos.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.907353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.907353Z digest=sha256:2fad8fdf6f22cd1ae6ca8b1d75e4f3b2f4d29926c1ddecf8375cce954a78a750

Observation dfd32bd4-5a49-416a-9492-7b11a166636f · outbound

This paper cites Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.911685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.911685Z digest=sha256:285828895bfb802b65fc2c368a268ab5072238770bc65f095a788cdb2c3b3fe7

Observation d117891a-ff1d-4d7a-802a-02b937b76f84 · outbound

This paper cites Object scene flow for autonomous vehicles.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Object scene flow for autonomous vehicles

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.119232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.916040Z digest=sha256:c6122463af6d3d1009c17a4b601e82fa34a59bf822ab0e51e10c7584cf3fef04

Observation 4f9548d2-61f9-4a74-9ce5-ae3b0080225b · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Embodiedgpt: Vision-language pre-training via embodied chain of thought

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.107864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.919907Z digest=sha256:e47b7a29af2da53eb215cd88ffef8c0d50d3cfb70471b81d5bd7b9c5e463b7fe

Observation bd638edc-3ab8-44e7-921e-f00c6555a8cf · outbound

This paper cites Orb-slam: a versatile and accurate monocular slam system.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Orb-slam: a versatile and accurate monocular slam system

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.924466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.924466Z digest=sha256:30f6f3509304911e33d9ef6b7270ab42685ea94d6c4b3b4d3b518ff88b847b88

Observation 90def821-ff91-4fc9-a4d2-b9668621b832 · outbound

This paper cites MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.928959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.928959Z digest=sha256:773825d767945d277b0f095cf098a8416409f192809f884046216ef20d920479

Observation a3ffc945-c7e4-4239-baa0-232cc9c90bd8 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos R3M: A Universal Visual Representation for Robot Manipulation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.932997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.932997Z digest=sha256:7d53331e433553a5b0bb33a331284419fb903d13ec955a59c0b6130a17ee0191

Observation 37649b4a-a0dd-49c2-95d0-313d21ccfe1c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.937434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.937434Z digest=sha256:639324c0466523119f1101a553dce8b799bc1b7031f040a1181c778d551ce5b5

Observation c0ee9492-3e9d-4978-9050-c15b7d10b3f7 · outbound

This paper cites From 2D to 3D: Re-thinking Benchmarking of Monocular Depth Prediction.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos From 2D to 3D: Re-thinking Benchmarking of Monocular Depth Prediction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.942006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.942006Z digest=sha256:863ce50653ece10cc782af7fd07e0ff8e37f848a6ff5fa76bfaa7f73b75cce35

Observation a9ae9f06-fcad-4853-8292-49a4ced0968f · outbound

This paper cites Aria digital twin: A new benchmark dataset for egocentric 3d machine perception.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Aria digital twin: A new benchmark dataset for egocentric 3d machine perception

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.088733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.946552Z digest=sha256:43a728853d0709b058abccb5fd7d76f0b7b957b227d82494bc517185cd4160c1

Observation c5a6e2b5-5371-4552-99c2-617ae639307a · outbound

This paper cites Re- constructing hands in 3d with transformers.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Re- constructing hands in 3d with transformers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.077778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.951025Z digest=sha256:f3291285eb7f6291f8213660efd19cec31bbf442f29de74b12dd738c6ee1f32c

Observation d49c8d1c-1fcd-4133-9c2c-9130058fbc12 · outbound

This paper cites Unidepth: Universal monocular metric depth estimation.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Unidepth: Universal monocular metric depth estimation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.066349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.955421Z digest=sha256:d5229c7591c325f18ef76102a0391474e543eb08bbf49e297a0b5665f22ed348

Observation 56050f13-69cd-4ce0-aaa0-2c134170a9a0 · outbound

This paper cites WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.959374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.959374Z digest=sha256:769831fe8de34c9c9fd889029b50495060d1be9458d00459542b7426447e2907

Observation 6458bab8-fcf3-4ce5-87d4-ac5f6ef1f5c4 · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.054606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.964439Z digest=sha256:ea56656b5aedaaaa94933be8c488fcef66c1512bee98fac12645dd67352fcc3c

Observation a123a616-77b1-4ee9-bf5e-00fe23221957 · outbound

This paper cites Affordancellm: Grounding affordance from vision language models.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Affordancellm: Grounding affordance from vision language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.968803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.968803Z digest=sha256:0b05ebd6407e204bcfecd08f8a2f8ac6233a7e2bf6acf9ab18d6414de9c2fa99

Observation 53e1322a-f990-46de-8941-22f71717d591 · outbound

This paper cites From one hand to multiple hands: Imitation learning for dexterous manipula- tion from single-camera teleoperation.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos From one hand to multiple hands: Imitation learning for dexterous manipula- tion from single-camera teleoperation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.034084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.973093Z digest=sha256:dcbb7411c374e104d38634e32f67a05ec270accdf3d7b045a779032dcfecf7f0

Observation 6d3558cb-8b35-404e-8c73-0d423e7ed2f9 · outbound

This paper cites Deep learning- based depth estimation methods from monocular image and videos: A comprehensive survey.ACM Computing Surveys,.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Deep learning- based depth estimation methods from monocular image and videos: A comprehensive survey.ACM Computing Surveys,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.021719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.976711Z digest=sha256:0149bd12d963293a4738511412385f1ad80328cb74459469d5438ff54c7f6671

Observation 2b4e34d6-8d0b-4262-9a6c-9cd953a919bb · outbound

This paper cites Nerf- slam: Real-time dense monocular slam with neural radi- ance fields.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Nerf- slam: Real-time dense monocular slam with neural radi- ance fields

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:41.010196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.980151Z digest=sha256:54788f88c8c7541c0862b82dd9f458e76641cedea84ea3508edc29851b7daf1a

Observation 64366714-3340-4ea4-9237-bfffea7ee635 · outbound

This paper cites Identification and cor- rection of flying pixels in range camera data.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Identification and cor- rection of flying pixels in range camera data

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.998289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.983646Z digest=sha256:7290f2b00e963b58485fe3ae8f8f220abe91223c8f7ab0922aa84252e65fb808

Observation e8f264e2-97a4-4a00-9f6e-be3485b35b25 · outbound

This paper cites Structure-from-motion revisited.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Structure-from-motion revisited

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.986795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.986988Z digest=sha256:be4911d7543d47d30f33cb294999142666103b9d6b80e0f591ab310711a937ef

Observation 261e3b0a-db13-4d3b-8861-33545c495bd4 · outbound

This paper cites Pixelwise view selection for unstructured multi-view stereo.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Pixelwise view selection for unstructured multi-view stereo

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.990585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.990585Z digest=sha256:69da73d750405935e84017ef9118d0403fecd1588c9a03f09a225e83151f9624

Observation 0146b440-b9c6-4b03-88d6-57e184c6ac07 · outbound

This paper cites Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.969211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:39.994388Z digest=sha256:4120293e0a77c2c217494958f72f804abfe67444893097552cca9c0108c3ba22

Observation 32cea04f-46f4-4216-8269-8190c63af2b8 · outbound

This paper cites Understanding human hands in contact at inter- net scale.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Understanding human hands in contact at inter- net scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:39.998249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:39.998249Z digest=sha256:c2aa2fbb347052855102edb7235a10ab2b94b1cd29d5f2d490e18bec877f28eb

Observation b18c3a36-56f6-4411-9d71-decce3f49de7 · outbound

This paper cites Motion-based object segmentation based on dense rgb-d scene flow.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Motion-based object segmentation based on dense rgb-d scene flow

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.950347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.002167Z digest=sha256:b5bd96d3c397b518e4116061d83227f755a0e52971719804f9b3594880075c7c

Observation d9096ce8-f580-4d19-8bfc-969a377b5f3a · outbound

This paper cites Swindepth: Unsupervised depth estimation using monocular sequences via swin trans- former and densely cascaded network.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Swindepth: Unsupervised depth estimation using monocular sequences via swin trans- former and densely cascaded network

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.938047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.005743Z digest=sha256:0cb589bef35cd79cc0b15d3b046bcf97b6f852332d8117217ad94bb967ade81c

Observation 76066f33-1564-49f8-ac58-76d48c533607 · outbound

This paper cites FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene Flow.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene Flow

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.010002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.010002Z digest=sha256:d1f47939acfe38cc431e461c6adb963b041930c37e539cf0979a214caae7b133

Observation 4e6e7d76-16df-41e5-973c-9dfc83ff6167 · outbound

This paper cites FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.014191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.014191Z digest=sha256:0cd3553a6ee91486031a939a7958c6e8f100a1fc28c6d9d442aa0fb7f47d0f83

Observation 82a7edd6-6f8a-4efb-a9d2-cd7501766a42 · outbound

This paper cites Kick back & relax: Learning to reconstruct the world by watching slowtv.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Kick back & relax: Learning to reconstruct the world by watching slowtv

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.925633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.018770Z digest=sha256:8832fddb2fe80ce019f4fd5338fbe9017dd26792f4827b112807d968528aa063

Observation fda441e7-d31b-4993-a376-bfe52e239b3c · outbound

This paper cites Kick Back & Relax++: Scaling Beyond Ground-Truth Depth with SlowTV & CribsTV.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Kick Back & Relax++: Scaling Beyond Ground-Truth Depth with SlowTV & CribsTV

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.023727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.023727Z digest=sha256:d0e3362713a9d9f35a12e55cb6b202c6819c5ebde706ccd2ca396a970d29a0a3

Observation 7f551f1d-ad97-4523-99c2-1ccba4e3bd5c · outbound

This paper cites A unified transformer frame- work for group-based segmentation: Co-segmentation, co- saliency detection and video salient object detection.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos A unified transformer frame- work for group-based segmentation: Co-segmentation, co- saliency detection and video salient object detection

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.913083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.028807Z digest=sha256:2976730f012e1037667bce0e312b96cb64fc79d62aed048b5ef1afda9780146d

Observation 3e1ea8e8-1847-4d64-8e09-16be8973f8ae · outbound

This paper cites 3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos 3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.032635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.032635Z digest=sha256:6e9e556483e5ed891765e74b4734301b7671c001dd4181965d1557ac6ec8557c

Observation 9a30a8a0-d893-495f-b6a3-bfa0c3f21d25 · outbound

This paper cites Dynamo-depth: fix- ing unsupervised depth estimation for dynamical scenes.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Dynamo-depth: fix- ing unsupervised depth estimation for dynamical scenes

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.894609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.036984Z digest=sha256:4515030e95b4b1483d48427e381e8f7a69d468955e4040d4b040bcbc6f896cf0

Observation 6e54c878-c6ec-4f94-bafe-311748d59852 · outbound

This paper cites Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.040954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.040954Z digest=sha256:87cecc8f658e90fe839924a40149a7abf946d228631fc245cd27176c761f06f6

Observation 72f32c76-1d00-4967-aa7a-dc851eeec0bd · outbound

This paper cites Cutting edge, flagship camera with intelli- gent feedback and resolution.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Cutting edge, flagship camera with intelli- gent feedback and resolution

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.862680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.049599Z digest=sha256:abb1e5d5cbf28aebf0f0d13faec3114e58b0040576429c2a7981fb6230e416e1

Observation f643168a-a8de-4260-a3c0-17f713032efe · outbound

This paper cites Attention is all you need.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Attention is all you need

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.850627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.053366Z digest=sha256:83b956089747c56cc1a82bee72ee072510079ae498c8627fd7c07c9db5967d56

Observation ec255f23-e6be-4892-b366-36486d49a1c3 · outbound

This paper cites Three-dimensional scene flow.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Three-dimensional scene flow

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.837299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.057335Z digest=sha256:4bf3cdb1b2f5929429183cb53431f0c4f63e5e8e400eb8055c6808db55323e0c

Observation adccae37-4488-41c4-a8ee-9ac42fc14826 · outbound

This paper cites 3D Reconstruction with Spatial Memory.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos 3D Reconstruction with Spatial Memory

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.061138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.061138Z digest=sha256:c96990768ffb0a592925b736688cc78a92c8ee14796039b54d2895230253350f

Observation 3d66aab6-8614-482b-aeee-a46ed14eb667 · outbound

This paper cites Vggsfm: Visual geometry grounded deep structure from motion.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Vggsfm: Visual geometry grounded deep structure from motion

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.825954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.065433Z digest=sha256:d0539313ad0d203035ff89eb012548b463e829b9a8bce78a5aa176a4a6292707

Observation 94413edf-b1e1-4f10-8258-4c030f14e467 · outbound

This paper cites Shape of motion: 4d reconstruction from a single video.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Shape of motion: 4d reconstruction from a single video

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.069360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.069360Z digest=sha256:648171fedf59c6846e3b51eeb71a26a82d3186758aba73fbee4033b3a6a51337

Observation cf6e0524-e375-4568-af28-06c2946f83f3 · outbound

This paper cites Continuous 3D Perception Model with Persistent State.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Continuous 3D Perception Model with Persistent State

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.073445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.073445Z digest=sha256:d7dc2ddb4c31a946565ebc5fcc8eb6b56a86ab89739582d724c0ca5b8bf7c7d8

Observation a60015aa-3f49-48dc-88f5-bc7ec2655e6b · outbound

This paper cites Pov-surgery: A dataset for egocentric hand and tool pose estimation during surgi- cal activities.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Pov-surgery: A dataset for egocentric hand and tool pose estimation during surgi- cal activities

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.814207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.077918Z digest=sha256:63ffa89498d11575b36098a144380edcff7294bd2a40837b2d9096d5fa315948

Observation b53426f0-28df-447c-aa8a-e6b6383b7121 · outbound

This paper cites Dust3r: Geometric 3d vision made easy.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Dust3r: Geometric 3d vision made easy

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.801928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.081702Z digest=sha256:a24553c01ef9929fdd9815f7e49d6e3da25dcb3c4dd902cfafa7c3e2234df4aa

Observation c399aa9f-7d85-4bea-a234-cae103c88bd2 · outbound

This paper cites Tar- tanvo: A generalizable learning-based vo.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Tar- tanvo: A generalizable learning-based vo

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.790517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.085044Z digest=sha256:6c3fcd13f03fa5624e584e0005d725e2e90cc77f73dd7ae89046d1f5f99ebf31

Observation fcd80a5a-f3f3-490a-b3dc-aa92790581ce · outbound

This paper cites Neural video depth stabilizer.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Neural video depth stabilizer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.089720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.089720Z digest=sha256:6280ea36981f60709a191012d08a1833efe932bc0a932abcf99b1561a79a82cf

Observation 0aefd348-9a0c-4673-b6b1-1a4f21b1140a · outbound

This paper cites Egocentric video comprehension via large language model inner speech.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Egocentric video comprehension via large language model inner speech

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.771100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.094031Z digest=sha256:5f1e0951ddb68e9bbbe3b475ca728c1a596d2ff319691faeb527d077ce8efdb9

Observation 4f0cd970-e03d-413a-845a-b909bef233f3 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Any-point Trajectory Modeling for Policy Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.097465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.097465Z digest=sha256:72ab4bfe51d7ff37ee3f3c82707e4ea96732d11c47c2c149f24ac64c2965772b

Observation 8281ac46-3f53-4c46-906d-71ccccd1b9c5 · outbound

This paper cites Moving Object Segmentation: All You Need Is SAM (and Flow).

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Moving Object Segmentation: All You Need Is SAM (and Flow)

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.101194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.101194Z digest=sha256:5f8132368ec9329fe9bc289983862a488969beb0c127ea26848e0235d35d2d93

Observation ca0a3096-93fd-4acd-a139-2891745a6dc3 · outbound

This paper cites Gmflow: Learning optical flow via global matching.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Gmflow: Learning optical flow via global matching

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.105977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.105977Z digest=sha256:86b195c22177c7125c70dddca2946fe892a64a8552d603bba5e77bcff7919a8e

Observation 52750bf8-f045-4394-878a-3cb63653b417 · outbound

This paper cites Depth anything: Un- leashing the power of large-scale unlabeled data.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Depth anything: Un- leashing the power of large-scale unlabeled data

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.752271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.110069Z digest=sha256:853d522eca3ef9c14c6db26e0604e02a9b3d0775bca6bd7023cb128172d0b779

Observation c75ec479-e73d-4e24-94ac-81ffbbd43533 · outbound

This paper cites Every pixel counts: Unsupervised geome- try learning with holistic 3d motion understanding.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Every pixel counts: Unsupervised geome- try learning with holistic 3d motion understanding

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.740780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.113725Z digest=sha256:c8f4537576b05913943b111c89842b40e2666b6308152b54163e5cd99348c46f

Observation 905bc388-9413-4c4c-b35f-d404585947b6 · outbound

This paper cites Met- ric3d: Towards zero-shot metric 3d prediction from a sin- gle image.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Met- ric3d: Towards zero-shot metric 3d prediction from a sin- gle image

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.729430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.117312Z digest=sha256:941e3175e00c1b333da05f25713aceaa7588a0aaf6b519b651b399f620f6c3e2

Observation 1d52056b-c92f-4d1a-95b7-543183594a55 · outbound

This paper cites General Flow as Foundation Affordance for Scalable Robot Learning.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos General Flow as Foundation Affordance for Scalable Robot Learning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.121084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.121084Z digest=sha256:8303971666f213cc3e0604f5a7c962713671dd0df5c21bd38c9e98981367f954

Observation e63e0849-55fd-436c-b399-f28d2880287c · outbound

This paper cites Recent trends in 3d reconstruction of general non-rigid scenes.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Recent trends in 3d reconstruction of general non-rigid scenes

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.716617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.125318Z digest=sha256:97a4e3f135da55925dc27464362c45274038a6c14297d03b90af91384df43418

Observation 50a9096a-73c7-4311-8bb8-c8c35eba7bdf · outbound

This paper cites EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.129364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.129364Z digest=sha256:826d3c57b76826aafe6ec4813d4527801f7be477e2c415fd7e5ee30b5da293df

Observation 073fd12c-dfd5-4ea0-8c74-0882fa75dcb2 · outbound

This paper cites MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:40.133398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:07:40.133398Z digest=sha256:4500da344544278aec758adf8c69f69f9d88aa03252ca40164ddc15ab82871b8

Observation 787ee4c0-3e36-4aaa-b5f9-553c447c0cf8 · outbound

This paper cites Fine-grained egocentric hand-object segmentation: Dataset, model, and applications.

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos Fine-grained egocentric hand-object segmentation: Dataset, model, and applications

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:07:40.703780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:07:40.137416Z digest=sha256:123cae7773d1ca1a1a570c1d8fad59acb52d059abd46326bc04e2eb4c5362f7f

Pith citing papers

Observation ed3defec-a7ea-4e41-a872-0303ba81f2a5 · inbound

Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective cites this paper.

Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos

Reference 188

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:26.107182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:13:42.052386Z digest=sha256:71877ad152ab9ad00db3a900629e565648331b617100807dafd338c97ee065d2