Pith. sign in

Paper Citation Record · LEDGER

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2507.15833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15833 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:24.845540Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:37:18.332035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:26:17.693248Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce980063-9b2b-461b-838a-d2ef1010ad05 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.700155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.700155Z digest=sha256:c915db8932cacd701eca0af580a3488cfe5eed1da2063b19977b08527720539e

Observation 8d036f2d-56b9-4d0f-a3c2-c84d3e005baa · outbound

This paper cites ALOHA Unleashed: A Simple Recipe for Robot Dexterity.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers ALOHA Unleashed: A Simple Recipe for Robot Dexterity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.704197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.704197Z digest=sha256:43c5cd11cc6e50f1e61e67fdd2fdc235111fe5df578ce16f28cbb4c068f9bd82

Observation ac3e2aa1-72c6-4dce-af6e-7e9098bc9efd · outbound

This paper cites InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.707646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.707646Z digest=sha256:967e2918333207c02a260cfd12a65fbc8911a9cda1d88402212d33e17466b1f9

Observation 8017729d-db14-493f-a309-65fcb4b4572b · outbound

This paper cites Vita: Vision-to-action flow matching policy,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Vita: Vision-to-action flow matching policy,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.711214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.711214Z digest=sha256:9af47cc613fac6ff0b2793aff3355dcd416b7df14d66315fdc8b1672b0860f3d

Observation 4ef3337a-8f3a-4766-8946-6f089d06e628 · outbound

This paper cites Generalizable Humanoid Manipulation with 3D Diffusion Policies.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Generalizable Humanoid Manipulation with 3D Diffusion Policies

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.714604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.714604Z digest=sha256:d6eff74b92d7b9cba20c199ae071f0bd4a67901e1ddb1b016dca223a849af330

Observation be4fd690-f188-45c2-ad0f-80bce9964087 · outbound

This paper cites Flow matching imitation learning for multi-support manipulation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Flow matching imitation learning for multi-support manipulation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.614565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.717772Z digest=sha256:c067469317798f7c4cc049b59a83dd90de03974e0f20cdf74ae4084dff5c7285

Observation 9d405417-9ffc-45f0-9c01-85dd22b99060 · outbound

This paper cites HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.721012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.721012Z digest=sha256:588bad4936e705a770f7ede096caccde1956c2e8c916d4ecd615daec4d23321f

Observation 7b63a1f5-d7fc-4df5-8bdf-4adad98b9637 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via ac- tion diffusion,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Diffusion policy: Visuomotor policy learning via ac- tion diffusion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.724047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.724047Z digest=sha256:cb998b3d18b200ce44977396e16a1cd2cfa69981b66e53487295f4c9c0f1c7d8

Observation 38b966fb-cf15-464d-af2d-ef33f778f2a0 · outbound

This paper cites The foveal confluence in human visual cortex,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers The foveal confluence in human visual cortex,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.597089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.727042Z digest=sha256:9ef5f6022217e687293f3071afbb04b13eb7322e8f33a0027c745b8b5453cb60

Observation a00962be-6c6e-4613-9abd-f1b73dba7680 · outbound

This paper cites Embedded foveation image coding,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Embedded foveation image coding,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.585592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.729740Z digest=sha256:a1506bc7914d619685ac169ea1be049c1c7bbdd37c60beaefa71d52e0aa259dc

Observation 3cf9e613-3c93-448c-a293-fcc2cae307f0 · outbound

This paper cites Object detection through search with a foveated visual system,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Object detection through search with a foveated visual system,

Reference 11

Resolution
verified exact
doi, observed 2026-08-06T15:26:24.880287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.732580Z digest=sha256:e33fc381840ae4547ffe33933bf657676d4eeb78265431c49d95143bb4b0b6b2

Observation a434271e-8c86-4fdc-8ea3-b511b3c1fc0c · outbound

This paper cites Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.735889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.735889Z digest=sha256:071fd83920a8ea6474db4346e0dff75e77859fb0c771765e4532744cfe444834

Observation 7c8085e4-1413-4275-9c96-7767ec56c83f · outbound

This paper cites Open-TeleVision: Teleoperation with Immersive Active Visual Feedback.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Open-TeleVision: Teleoperation with Immersive Active Visual Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.739499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.739499Z digest=sha256:5425e7ceb0c5a3e51436c2d5c32174bc9b413516786dbe7bfd0dd52e654da6c5

Observation 2625bef8-c72f-4b74-b1f6-4bb7ede446be · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.742657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.742657Z digest=sha256:a5a616addf9277d2cce3952515180463b87f7a552d2209190892d24401284cbd

Observation 27dd0c2d-8413-437c-af29-e4eefe655da4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers DINOv2: Learning Robust Visual Features without Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.745877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.745877Z digest=sha256:320cb7ae0b508ba79d6fe4497477c10643188a0b9f7901069568761a13e963be

Observation 7bcea48b-5d35-4a42-8f3e-845284cd1aed · outbound

This paper cites Data Scaling Laws in Imitation Learning for Robotic Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Data Scaling Laws in Imitation Learning for Robotic Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.749095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.749095Z digest=sha256:18c4e4ee487f815bb880fadfe59dc9dc1ade2c53d3117d19a622d07056f9f33e

Observation f14a265d-03eb-4514-bb7d-7197b5033bbc · outbound

This paper cites CAGE: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers CAGE: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.752532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.752532Z digest=sha256:398e358f70eff7eb437549de0121cc56b3f2c1adfbb8c5dc1811aade3cebab27

Observation 88fd0ff6-82b8-463e-91c8-1287cee38fa0 · outbound

This paper cites Segment this thing: Foveated tok- enization for efficient point-prompted segmentation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Segment this thing: Foveated tok- enization for efficient point-prompted segmentation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.574949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.755787Z digest=sha256:fb0886671908ccb156ad93d48d01c8265bc44d57917662e7ef197889399ab151

Observation 5aca11cf-b514-43f6-9e04-287915fc0983 · outbound

This paper cites Non-extensive distribution of human eye photoreceptors,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Non-extensive distribution of human eye photoreceptors,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.563037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.759024Z digest=sha256:d35b4083221717750d35a528c8b8637e654d20551d87ed3eb7817d8ac81f496b

Observation ad4aa7f2-ebb3-49c5-9666-77ede44a1721 · outbound

This paper cites Marr,Vision: A computational investigation into the human repre- sentation and processing of visual information.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Marr,Vision: A computational investigation into the human repre- sentation and processing of visual information

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.762941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.762941Z digest=sha256:5bb695ae675724e1140a781da823d8c984b1b6f579823b547293c2b352ae7d64

Observation 916aa16d-60e9-4fe8-aebd-b97c28eeae05 · outbound

This paper cites Recognition-by-components: a theory of human image understanding.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Recognition-by-components: a theory of human image understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.544217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.766260Z digest=sha256:0f28bbbb2ef59148999d2a495e5855e51ab080a3fa11553464423a8fed54c48b

Observation 9fcb47e2-20db-4f1d-9282-afb2d5973d9d · outbound

This paper cites Neural mechanisms of selective visual attention,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Neural mechanisms of selective visual attention,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.533996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.769210Z digest=sha256:2028edf6b3dbb12e591d37fab808d9e95a5e63f6321c830f2fb3d52efc3328b6

Observation c218dcc9-e614-44e8-88ed-5fedcde6beda · outbound

This paper cites How does the brain solve visual object recognition?.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers How does the brain solve visual object recognition?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.522885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.772234Z digest=sha256:ec3eba66a718d00e974d12b90b488c7ea73e3f6635dc1d19911212b2e0e5849b

Observation 669436bb-8971-43f8-9dda-8ad4d82dd3a0 · outbound

This paper cites Performance-optimized hierarchical models predict neural responses in higher visual cortex,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Performance-optimized hierarchical models predict neural responses in higher visual cortex,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.511051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.775393Z digest=sha256:ce5a234b0bad4178ab9347e4b9b866fb64092caa1bfeb4d6819bb366b4063bfe

Observation 810fd713-c875-4ffb-842d-55542a02b1bf · outbound

This paper cites Context mitigates crowding: Peripheral object recognition in real-world images,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Context mitigates crowding: Peripheral object recognition in real-world images,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.499154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.778460Z digest=sha256:da285e3cc356c17c8649e460bfb61c2ae44770630e45ffa797ce067652ef5e2a

Observation 47c9fc71-658c-4f85-9121-f8b073862ab9 · outbound

This paper cites Emergent Properties of Foveated Perceptual Systems.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Emergent Properties of Foveated Perceptual Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.781610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.781610Z digest=sha256:a15da4e3a9afa97204f2421c15407b8bd94728addae27375360b3261da021786

Observation 2e748845-bfc4-43c3-898b-e83baf2d5dfe · outbound

This paper cites Biologically inspired deep learning model for efficient foveal-peripheral vision,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Biologically inspired deep learning model for efficient foveal-peripheral vision,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.485670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.784603Z digest=sha256:650efd5ce3deefb23a0764777123ee6139f4659eed2edcb3c497524348b4c4eb

Observation 951fb8fd-cdc0-49b5-9e8e-b5c02e4615d6 · outbound

This paper cites FoveaTer: Foveated Transformer for Image Classification.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers FoveaTer: Foveated Transformer for Image Classification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.787335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.787335Z digest=sha256:e3534a4bf76900658665dd66f6cce4bdd1e9d8793adbc18f756e27d961035601

Observation 428f0ab1-805a-4257-a633-35a4991b71a0 · outbound

This paper cites Peripheral vision transformer,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Peripheral vision transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.474845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.790467Z digest=sha256:fa1111fe712aaff0c70ccc4f9bd9b518b10e8ff23ea9e91c5160cdce0c81e7fc

Observation 017d69eb-54fb-4671-bf3f-860f74b47cf3 · outbound

This paper cites An eye-gaze tracking system for teleoperation of a mobile robot,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers An eye-gaze tracking system for teleoperation of a mobile robot,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.464210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.793147Z digest=sha256:833fee4d3373e5de068d9a4c47151a203fb8e41d03c81d99e25ee01cba693f02

Observation 65a6cf5a-6c50-4924-a0ae-8bed3a150269 · outbound

This paper cites The human gaze helps robots run bravely and efficiently in crowds,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers The human gaze helps robots run bravely and efficiently in crowds,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.450860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.796342Z digest=sha256:ad954951563a34b325334abde5ad4af386df6b7f0f92cf5251112b31bb287fb5

Observation 848ce0ca-81f4-4fa7-b6e0-0d1a9c074a5e · outbound

This paper cites Gaze-based intention estimation: principles, method- ologies, and applications in hri,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Gaze-based intention estimation: principles, method- ologies, and applications in hri,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.799515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.799515Z digest=sha256:24e34907ff60280629fb87effe56e1d4feddf41e88e2c72b9de2a3e71b16bacd

Observation 026380ae-af67-45ea-ae5e-f616a70dfa6f · outbound

This paper cites Visarl: Visual reinforcement learning guided by human saliency,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Visarl: Visual reinforcement learning guided by human saliency,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.432702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.802352Z digest=sha256:c1e7db81ed25301900a685e324cd6551f7893cc7bb07fa2393daedf22a22af13

Observation 4f67658a-517d-46cd-8fad-00f374018f91 · outbound

This paper cites Eye, robot: Learning to look to act with a bc-rl perception-action loop,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Eye, robot: Learning to look to act with a bc-rl perception-action loop,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.805436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.805436Z digest=sha256:c886f792b13be2b8796ce5c14c4b18b19f4e3c5df21746b34fdbbfd4f2a2f909

Observation 92743b18-ef07-4803-b0fb-ebea9843e7a2 · outbound

This paper cites Using human gaze to improve robustness against irrelevant objects in robot manipulation tasks,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Using human gaze to improve robustness against irrelevant objects in robot manipulation tasks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.422245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.808245Z digest=sha256:3b3c43aca22aa451799b6bcdbac359055898f31b689e0e822c908e0317cd18ea

Observation 3ed6b318-5ca3-4629-bb99-58e8307b70e8 · outbound

This paper cites Gaze-based dual resolution deep imitation learning for high- precision dexterous robot manipulation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Gaze-based dual resolution deep imitation learning for high- precision dexterous robot manipulation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.410990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.811096Z digest=sha256:32cd7f3fd5eea9e21ea1c1011c112bac95d767f4bdd4434ee72167e6b2b18b76

Observation 77f10773-2874-44ed-b89f-4febcfb395da · outbound

This paper cites Multi-task real-robot data with gaze attention for dual-arm fine manipulation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Multi-task real-robot data with gaze attention for dual-arm fine manipulation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.398050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:26:24.814402Z digest=sha256:70f3504965ca0c8ec2ab72cc5b0d18c8c9a80d867e96d3e084e35923908e51f0

Observation a05b25a7-80d9-4da6-91a4-00f8699dfb19 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.817423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.817423Z digest=sha256:4143bc66f81084bbfb497e66b019b06212676f4c12567df79637b70c1020c7e1

Observation 0f509454-2616-4c63-be45-7f611e147945 · outbound

This paper cites Affordance-based robot manipulation with flow matching,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Affordance-based robot manipulation with flow matching,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.820442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.820442Z digest=sha256:5fdc16c53ecd5688023901bfeab1581df89de81a3146dca9ac27016efa6fa799

Observation 006f5ce2-f174-4492-a967-d330aec1a913 · outbound

This paper cites Flow Matching for Generative Modeling.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Flow Matching for Generative Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.824013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.824013Z digest=sha256:5b2f998f835890b49bfc8ce1f9a3fe5448e42d77bfd45ea57f18f5d995e1d8bc

Observation 829c3ee5-503a-448c-95cb-91bae424411f · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.827374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.827374Z digest=sha256:23f9653e579014715b65b8286361a3b4627080023157d6d85e20592b4b458522

Observation 99343bc7-9761-41f6-9d8d-d17aae6e07e2 · outbound

This paper cites Scalable diffusion models with transformers,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Scalable diffusion models with transformers,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.830907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.830907Z digest=sha256:3bb2fd4894b955daef5f15cc9d293cbef743ed736306740fccafa823af0ec171

Observation 28d675ee-7aec-49f7-9bc2-044a2206edf3 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Movie Gen: A Cast of Media Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.834469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.834469Z digest=sha256:bafd64f5ba4ebcb8e31f0be9befcddb2f0ff92d07f503c255ca564c447a8b705

Observation ebb6e3eb-131c-4335-aa2e-53f694833eeb · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Imagenet: A large-scale hierarchical image database,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.837714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.837714Z digest=sha256:bf9a24ae679bb661e0b9e0dd1658ec219da259d7762efa7ede1cc8d1aa3ce24b

Observation 4e8816bd-ffcd-49f3-b1e1-9c38201f351c · outbound

This paper cites Masked autoencoders are scalable vision learners,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Masked autoencoders are scalable vision learners,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.845540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.845540Z digest=sha256:8102be7688dbc5834a8324467b8f6c1c36ac159c593d3a9305273f8a748d6529

Pith citing papers

Observation 400d48b8-c366-4ddf-b290-5bc82097a836 · inbound

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data cites this paper.

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:18.332035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:18.332035Z digest=sha256:f6d0b6374516c21e3d4cc0034c808688da67dacca58af7f479d5817b35d75222

Observation faa22e20-d41f-4fc6-8fbb-7668f984ac9f · inbound

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations cites this paper.

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:30:47.668896Z digest=sha256:2a0a8ba357cb2e4e9d15ce260d56e07bc1f8fee09d21db98059a9cd38ec2005c

Observation 2418b343-508b-4533-b05e-0840b352c181 · inbound

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games cites this paper.

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:26:50.491037Z digest=sha256:6483cb0e3b7a5cfd0af5fbc9eab46249dc7b22ad836356832a7d2724ff249e98

Observation c2fdc816-f159-4253-8e4d-045e05a1d8b8 · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:57999960ee153f74280393d2403dafab3451a23df56e268adde2b4120912422e

Observation 8931d48f-eab7-4ddb-b650-a5167fc2ec83 · inbound

Policy-based Foveated Imaging and Perception cites this paper.

Policy-based Foveated Imaging and Perception Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 179

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T15:23:48.688620Z digest=sha256:3e54dd7aff24bf182f53cc81cc24ea339f347032f93ccf5a2e701a7838bdd84b