Pith. sign in

Paper Citation Record · LEDGER

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

As of 12 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2507.15833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15833 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:24.845540Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:37:18.332035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:26:17.693248Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce980063-9b2b-461b-838a-d2ef1010ad05 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.700155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.700155Z digest=sha256:c915db8932cacd701eca0af580a3488cfe5eed1da2063b19977b08527720539e

Observation 8d036f2d-56b9-4d0f-a3c2-c84d3e005baa · outbound

This paper cites ALOHA Unleashed: A Simple Recipe for Robot Dexterity.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers ALOHA Unleashed: A Simple Recipe for Robot Dexterity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.704197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.704197Z digest=sha256:b9d0059a3d23bd87cc83f6a5b641e644be5bc851855c6533145671107ad468c7

Observation ac3e2aa1-72c6-4dce-af6e-7e9098bc9efd · outbound

This paper cites InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.707646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.707646Z digest=sha256:e981684a6f7934c06fe306b8785b20bd8a0a61aa81fad65a9e50e8854e130171

Observation 8017729d-db14-493f-a309-65fcb4b4572b · outbound

This paper cites Vita: Vision-to-action flow matching policy,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Vita: Vision-to-action flow matching policy,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.711214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.711214Z digest=sha256:9af47cc613fac6ff0b2793aff3355dcd416b7df14d66315fdc8b1672b0860f3d

Observation 4ef3337a-8f3a-4766-8946-6f089d06e628 · outbound

This paper cites Generalizable Humanoid Manipulation with 3D Diffusion Policies.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Generalizable Humanoid Manipulation with 3D Diffusion Policies

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.714604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.714604Z digest=sha256:d6eff74b92d7b9cba20c199ae071f0bd4a67901e1ddb1b016dca223a849af330

Observation be4fd690-f188-45c2-ad0f-80bce9964087 · outbound

This paper cites Flow matching imitation learning for multi-support manipulation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Flow matching imitation learning for multi-support manipulation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.614565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.717772Z digest=sha256:61105335b78adc7a6d5c7280f4e4cc858d4cfa070b0b270f4df3f404d43120c9

Observation 9d405417-9ffc-45f0-9c01-85dd22b99060 · outbound

This paper cites HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.721012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.721012Z digest=sha256:588bad4936e705a770f7ede096caccde1956c2e8c916d4ecd615daec4d23321f

Observation 7b63a1f5-d7fc-4df5-8bdf-4adad98b9637 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via ac- tion diffusion,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Diffusion policy: Visuomotor policy learning via ac- tion diffusion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.724047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.724047Z digest=sha256:cb998b3d18b200ce44977396e16a1cd2cfa69981b66e53487295f4c9c0f1c7d8

Observation 38b966fb-cf15-464d-af2d-ef33f778f2a0 · outbound

This paper cites The foveal confluence in human visual cortex,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers The foveal confluence in human visual cortex,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.597089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.727042Z digest=sha256:4953cc28156ab9f46920e68762d61c6d3b6868c4ad2b2a89729cbf00eaa3a9ec

Observation a00962be-6c6e-4613-9abd-f1b73dba7680 · outbound

This paper cites Embedded foveation image coding,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Embedded foveation image coding,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.585592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.729740Z digest=sha256:9126478705c77168c173eaa7ca276d2cea91a4357ff9cbe9534088627feb9fc7

Observation 3cf9e613-3c93-448c-a293-fcc2cae307f0 · outbound

This paper cites Object detection through search with a foveated visual system,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Object detection through search with a foveated visual system,

Reference 11

Resolution
verified exact
doi, observed 2026-08-06T15:26:24.880287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.732580Z digest=sha256:9523537151284cde93af9762d673e8c5db7865220192d0fd44f05fab65e68920

Observation a434271e-8c86-4fdc-8ea3-b511b3c1fc0c · outbound

This paper cites Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.735889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.735889Z digest=sha256:4a134ce1a3404c40ae5598f2115f76a96253d83e49a66f7c5c8ae11a221ca290

Observation 7c8085e4-1413-4275-9c96-7767ec56c83f · outbound

This paper cites Open-TeleVision: Teleoperation with Immersive Active Visual Feedback.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Open-TeleVision: Teleoperation with Immersive Active Visual Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.739499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.739499Z digest=sha256:5425e7ceb0c5a3e51436c2d5c32174bc9b413516786dbe7bfd0dd52e654da6c5

Observation 2625bef8-c72f-4b74-b1f6-4bb7ede446be · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.742657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.742657Z digest=sha256:3cac5ac31cc53185dd770dad28015bc70c68dc03a3967db233ac7e50eb39a252

Observation 27dd0c2d-8413-437c-af29-e4eefe655da4 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers DINOv2: Learning Robust Visual Features without Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.745877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.745877Z digest=sha256:da38645be594d9c89d1f9698ba10742e211e4e8e2c46fdc5e7a912a31d57348d

Observation 7bcea48b-5d35-4a42-8f3e-845284cd1aed · outbound

This paper cites Data Scaling Laws in Imitation Learning for Robotic Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Data Scaling Laws in Imitation Learning for Robotic Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.749095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.749095Z digest=sha256:18c4e4ee487f815bb880fadfe59dc9dc1ade2c53d3117d19a622d07056f9f33e

Observation f14a265d-03eb-4514-bb7d-7197b5033bbc · outbound

This paper cites CAGE: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers CAGE: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.752532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.752532Z digest=sha256:7e67d8d9eceefebe4d5a879e07041d3b3c5aa5eab5a3b40105d7e6233914e45f

Observation 88fd0ff6-82b8-463e-91c8-1287cee38fa0 · outbound

This paper cites Segment this thing: Foveated tok- enization for efficient point-prompted segmentation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Segment this thing: Foveated tok- enization for efficient point-prompted segmentation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.574949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.755787Z digest=sha256:272470e8e755fa7852ee6d34301540075f422a59e04f3a3315dfaee8966f6ece

Observation 5aca11cf-b514-43f6-9e04-287915fc0983 · outbound

This paper cites Non-extensive distribution of human eye photoreceptors,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Non-extensive distribution of human eye photoreceptors,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.563037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.759024Z digest=sha256:2c9832795e47921bfbeb6d6082ae5739d8c63fb59c30d862e9f6ee9eeef92533

Observation ad4aa7f2-ebb3-49c5-9666-77ede44a1721 · outbound

This paper cites Marr,Vision: A computational investigation into the human repre- sentation and processing of visual information.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Marr,Vision: A computational investigation into the human repre- sentation and processing of visual information

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.762941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.762941Z digest=sha256:5bb695ae675724e1140a781da823d8c984b1b6f579823b547293c2b352ae7d64

Observation 916aa16d-60e9-4fe8-aebd-b97c28eeae05 · outbound

This paper cites Recognition-by-components: a theory of human image understanding.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Recognition-by-components: a theory of human image understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.544217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.766260Z digest=sha256:dc5dd6d94d32f687e6c610f2fd71332c89ba762bab91526583849f23c5501073

Observation 9fcb47e2-20db-4f1d-9282-afb2d5973d9d · outbound

This paper cites Neural mechanisms of selective visual attention,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Neural mechanisms of selective visual attention,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.533996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.769210Z digest=sha256:5faba0449b3f344cd06895886305f0bab2665af4bdb1894b9c59a31a383f4ddd

Observation c218dcc9-e614-44e8-88ed-5fedcde6beda · outbound

This paper cites How does the brain solve visual object recognition?.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers How does the brain solve visual object recognition?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.522885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.772234Z digest=sha256:883e4f9b7c8b6a3975ab51fffa861d83c9a99160b6d543e197bd255339054c9c

Observation 669436bb-8971-43f8-9dda-8ad4d82dd3a0 · outbound

This paper cites Performance-optimized hierarchical models predict neural responses in higher visual cortex,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Performance-optimized hierarchical models predict neural responses in higher visual cortex,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.511051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.775393Z digest=sha256:7b991b0afbe0a3154b7ab6672e5a1c470592a59790ecc4b9e7af8cca9e9748de

Observation 810fd713-c875-4ffb-842d-55542a02b1bf · outbound

This paper cites Context mitigates crowding: Peripheral object recognition in real-world images,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Context mitigates crowding: Peripheral object recognition in real-world images,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.499154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.778460Z digest=sha256:e76c4fa38633cae4be7854751b091e2d980d00f43c5b5095ecd1fddd6935f8e5

Observation 47c9fc71-658c-4f85-9121-f8b073862ab9 · outbound

This paper cites Emergent Properties of Foveated Perceptual Systems.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Emergent Properties of Foveated Perceptual Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.781610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.781610Z digest=sha256:a15da4e3a9afa97204f2421c15407b8bd94728addae27375360b3261da021786

Observation 2e748845-bfc4-43c3-898b-e83baf2d5dfe · outbound

This paper cites Biologically inspired deep learning model for efficient foveal-peripheral vision,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Biologically inspired deep learning model for efficient foveal-peripheral vision,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.485670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.784603Z digest=sha256:cbebd332909cd717995250c6e138897d18e013db111cc2bd26a454e2bbb0fbfb

Observation 951fb8fd-cdc0-49b5-9e8e-b5c02e4615d6 · outbound

This paper cites FoveaTer: Foveated Transformer for Image Classification.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers FoveaTer: Foveated Transformer for Image Classification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.787335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.787335Z digest=sha256:c2f222cb364c99125119aad6418093be5bd290a9382beff8478064a6036aec2c

Observation 428f0ab1-805a-4257-a633-35a4991b71a0 · outbound

This paper cites Peripheral vision transformer,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Peripheral vision transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.474845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.790467Z digest=sha256:0cb7c3a98d7cc44520350cf408999fb1c4ca6d9f0f2853316027321ba12708ee

Observation 017d69eb-54fb-4671-bf3f-860f74b47cf3 · outbound

This paper cites An eye-gaze tracking system for teleoperation of a mobile robot,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers An eye-gaze tracking system for teleoperation of a mobile robot,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.464210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.793147Z digest=sha256:f1400992575c392486dac587846cf36c1b30ab63aa5821521743007cb03ad8fa

Observation 65a6cf5a-6c50-4924-a0ae-8bed3a150269 · outbound

This paper cites The human gaze helps robots run bravely and efficiently in crowds,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers The human gaze helps robots run bravely and efficiently in crowds,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.450860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.796342Z digest=sha256:d022cb22b12e9ba7b415954233ba3e2395b70a79776d770c94e71f8f9e0bc928

Observation 848ce0ca-81f4-4fa7-b6e0-0d1a9c074a5e · outbound

This paper cites Gaze-based intention estimation: principles, method- ologies, and applications in hri,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Gaze-based intention estimation: principles, method- ologies, and applications in hri,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.799515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.799515Z digest=sha256:24e34907ff60280629fb87effe56e1d4feddf41e88e2c72b9de2a3e71b16bacd

Observation 026380ae-af67-45ea-ae5e-f616a70dfa6f · outbound

This paper cites Visarl: Visual reinforcement learning guided by human saliency,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Visarl: Visual reinforcement learning guided by human saliency,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.432702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.802352Z digest=sha256:bbc66cb956db065ccabb3ec8f447830cc55d6d65e407d15f3d58e830ce7974b1

Observation 4f67658a-517d-46cd-8fad-00f374018f91 · outbound

This paper cites Eye, robot: Learning to look to act with a bc-rl perception-action loop,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Eye, robot: Learning to look to act with a bc-rl perception-action loop,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.805436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.805436Z digest=sha256:c886f792b13be2b8796ce5c14c4b18b19f4e3c5df21746b34fdbbfd4f2a2f909

Observation 92743b18-ef07-4803-b0fb-ebea9843e7a2 · outbound

This paper cites Using human gaze to improve robustness against irrelevant objects in robot manipulation tasks,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Using human gaze to improve robustness against irrelevant objects in robot manipulation tasks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.422245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.808245Z digest=sha256:7eaf3877d8be374ac5afcdf6c7baa682229fe8dcbb8bd494b03fc0004694c0a7

Observation 3ed6b318-5ca3-4629-bb99-58e8307b70e8 · outbound

This paper cites Gaze-based dual resolution deep imitation learning for high- precision dexterous robot manipulation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Gaze-based dual resolution deep imitation learning for high- precision dexterous robot manipulation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.410990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.811096Z digest=sha256:e473fcd187ffc57acdf9a252b158a5c6bbcf436e42bc29ba63b1b0b9b204f75e

Observation 77f10773-2874-44ed-b89f-4febcfb395da · outbound

This paper cites Multi-task real-robot data with gaze attention for dual-arm fine manipulation,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Multi-task real-robot data with gaze attention for dual-arm fine manipulation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:25.398050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T15:26:24.814402Z digest=sha256:d81ec1ad75aca85a975ac052926f7dbfcf3725d332ab07ba23e9eea5ad1f7dfc

Observation a05b25a7-80d9-4da6-91a4-00f8699dfb19 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.817423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.817423Z digest=sha256:4143bc66f81084bbfb497e66b019b06212676f4c12567df79637b70c1020c7e1

Observation 0f509454-2616-4c63-be45-7f611e147945 · outbound

This paper cites Affordance-based robot manipulation with flow matching,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Affordance-based robot manipulation with flow matching,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.820442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.820442Z digest=sha256:5fdc16c53ecd5688023901bfeab1581df89de81a3146dca9ac27016efa6fa799

Observation 006f5ce2-f174-4492-a967-d330aec1a913 · outbound

This paper cites Flow Matching for Generative Modeling.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Flow Matching for Generative Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.824013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.824013Z digest=sha256:5b2f998f835890b49bfc8ce1f9a3fe5448e42d77bfd45ea57f18f5d995e1d8bc

Observation 829c3ee5-503a-448c-95cb-91bae424411f · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.827374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.827374Z digest=sha256:23f9653e579014715b65b8286361a3b4627080023157d6d85e20592b4b458522

Observation 99343bc7-9761-41f6-9d8d-d17aae6e07e2 · outbound

This paper cites Scalable diffusion models with transformers,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Scalable diffusion models with transformers,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.830907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.830907Z digest=sha256:3bb2fd4894b955daef5f15cc9d293cbef743ed736306740fccafa823af0ec171

Observation 28d675ee-7aec-49f7-9bc2-044a2206edf3 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Movie Gen: A Cast of Media Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.834469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.834469Z digest=sha256:8dc3e527c214161a9ad01fab1963876ff823ac488a73b0399ee227112f0f94b8

Observation ebb6e3eb-131c-4335-aa2e-53f694833eeb · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Imagenet: A large-scale hierarchical image database,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.837714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.837714Z digest=sha256:bf9a24ae679bb661e0b9e0dd1658ec219da259d7762efa7ede1cc8d1aa3ce24b

Observation 4e8816bd-ffcd-49f3-b1e1-9c38201f351c · outbound

This paper cites Masked autoencoders are scalable vision learners,.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Masked autoencoders are scalable vision learners,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.845540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.845540Z digest=sha256:8102be7688dbc5834a8324467b8f6c1c36ac159c593d3a9305273f8a748d6529

Pith citing papers

Observation 400d48b8-c366-4ddf-b290-5bc82097a836 · inbound

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data cites this paper.

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:18.332035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:18.332035Z digest=sha256:e4901a33f747e3636242dde5f74f8cfd948e951108a8ac7acbff0b743795ad7b

Observation faa22e20-d41f-4fc6-8fbb-7668f984ac9f · inbound

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations cites this paper.

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T11:30:47.668896Z digest=sha256:74e1234b8941592c8810c9144f132cd168d4a3d9b58cd5c3fcce35b8ca688513

Observation 2418b343-508b-4533-b05e-0840b352c181 · inbound

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games cites this paper.

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:26:50.491037Z digest=sha256:8543128be5b9f92ed979161e48519f436ac87377fbe2a745a745b9acb2720ead

Observation c2fdc816-f159-4253-8e4d-045e05a1d8b8 · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:d22c74547959b00f58fb012c5a365fa0243b045c2cc69a57d4ad28c98ed8b98d

Observation 8931d48f-eab7-4ddb-b650-a5167fc2ec83 · inbound

Policy-based Foveated Imaging and Perception cites this paper.

Policy-based Foveated Imaging and Perception Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Reference 179

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T01:20:46.955469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T15:23:48.688620Z digest=sha256:4e5c4c8e806be8d3c3986f9e9d71bc3c977d3905d6ed0aa70c9291ac876eb145