Pith. sign in

Paper Citation Record · LEDGER

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

As of 13 August 2026, this Paper Citation Record lists 97 of 97 outbound references and 7 inbound Pith citation observations for arXiv:2412.21080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21080 v1

Coverage vector

measured 97 of 97 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:07:14.748483Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:37.248378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.306268Z

Reference resolution

97 of 97 outbound references displayed

  • verified exact0
  • verified fuzzy56
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd6ddc5a-13a4-4b4a-ab2c-4f79252f26e1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.331971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.331971Z digest=sha256:6e3bc09137a7c41d674643d12b5e62203e67a8cc0baf3ec00a9f2366b63548b6

Observation cfcf406a-9ecc-4c4b-b4b3-162a9bb23f32 · outbound

This paper cites Swim- master: a wearable assistant for swimmer.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Swim- master: a wearable assistant for swimmer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.337134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.337134Z digest=sha256:3756db9d1ea40252d985d74a2890a40b61424568ce4348655a55c79aeaa535be

Observation fa0e6e6d-4753-4892-bef7-46e60c23a757 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.341548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.341548Z digest=sha256:9a14b4ff3998a3ac05c9c10d5a0f791021bb50ba9a7ad2b6b4c4ed65e94204c9

Observation d301eb7f-de99-49f7-8b55-c29500649ccf · outbound

This paper cites Analysis of the hands in egocentric vision: A survey.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Analysis of the hands in egocentric vision: A survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.346187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.346187Z digest=sha256:429d3e4d56a97113a5d08fe07b93687841f6424c979dff9ee5ae718ace48c9dd

Observation 9fe67223-53fb-4d19-a7d8-973674cb4640 · outbound

This paper cites an unresolved cited work.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.350491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.350491Z digest=sha256:ab87c19093fe8fcb78e15a9c50b4375cc2d459ee25479d6780e125997fa0353d

Observation 8455371b-fdc7-433d-915e-7ec5c1a59c15 · outbound

This paper cites Internlm2 technical report, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Internlm2 technical report, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.354875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.354875Z digest=sha256:7b869aa85d426d0bd589dec45d6e28b4aba2c412b03570bd827e29fc4bb50e1c

Observation 6f2717d3-f7b1-46d3-8f65-7ff8d8ec4e97 · outbound

This paper cites InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.359740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.359740Z digest=sha256:ac9db41a58d534310cda4bc9989eea16303d83811a0c3e9b76b6af48346a145e

Observation d80c5428-67c9-4371-9864-88b495c6c5d9 · outbound

This paper cites Dcan: improving temporal action detection via dual context aggregation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Dcan: improving temporal action detection via dual context aggregation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.364280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.364280Z digest=sha256:eec7253d69e01dd17bed4761c9d1b49b33f33245fa52bed184cad891e11a4ba2

Observation d5eb3f21-fe80-4eee-9c0d-550b576a4a92 · outbound

This paper cites Elan: Enhancing temporal action detection with location awareness.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Elan: Enhancing temporal action detection with location awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.368251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.368251Z digest=sha256:85e9b972094415fc5b3bfc0c01dab6eb4cb6270ea6b22b314a872c7ef5c2f5d4

Observation bbaa22a6-33f3-40a5-a81b-6b74f22d5da3 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoLLM: Modeling Video Sequence with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.372376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.372376Z digest=sha256:b6d3551486ee46042543002121fca45fd6d34ae7662a705b6ecb175b9eda0dea

Observation cac52273-6630-49e3-be8e-ddcac21a80af · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.376930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.376930Z digest=sha256:2c6290ae075a1b001e0915b0f15bd58dcbb45b72200f6977a15a68caf8dccafb

Observation c6578085-a73b-4727-b924-867d9c0aec1c · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.381559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.381559Z digest=sha256:863b1edb181ef826dcf4c3c1cc6c27121cc5eb4982a1e3a89544adcb2a0c0e7f

Observation dc325425-3ef9-43cf-bd16-228429bf1ae4 · outbound

This paper cites Gatehub: Gated history unit with background suppression for online action detection.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gatehub: Gated history unit with background suppression for online action detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.385876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.385876Z digest=sha256:13e1bbf9fa86d3ea3936804d0752a1dff550a48af40dcee90891f6159b107ba8

Observation 966d4664-06ae-4ffa-bcc8-e76b8ceae919 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Videollm-online: Online video large language model for streaming video

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.389760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.389760Z digest=sha256:ddedbb9b8dd5ed5e657be86377806122a7f389fdcac30b98965dd5c26f95f48a

Observation 647ee9ad-6757-45c9-95c8-dbc54f48f491 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.393831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.393831Z digest=sha256:823307383892e3a2c40763e83e8ee47000d267a0617ebd06233964a58fabfd23

Observation f4e0bdab-c45e-44c5-a275-d2b9e7c40fe5 · outbound

This paper cites gSDF: Geometry-Driven signed distance functions for 3D hand-object reconstruction.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model gSDF: Geometry-Driven signed distance functions for 3D hand-object reconstruction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.397896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.397896Z digest=sha256:96a902d14d563de6339fc9c8b349e4fcb0b38b14da76a2f32825f10a26b81e0a

Observation 4a4005a2-c553-4284-93bf-4d199e027499 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.402030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.402030Z digest=sha256:de05f6a541ee367261e37965988b14383edaf5d94c0bf3bbf8b6b1c3f28fb245

Observation 32aca28f-97e8-4598-b5e7-61f0a41f8529 · outbound

This paper cites Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.406367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.406367Z digest=sha256:f11c5883d8b19a5856099381a58ee82e8d34536fdeb07123bcb803eb7422a744

Observation 20985824-880b-43c4-9a11-69fe87a6f4e5 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Scaling egocentric vision: The epic-kitchens dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.410760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.410760Z digest=sha256:848acf626f58d5ebc3780c04e05dda1363af60eaf72ebca6044511937a7d53d8

Observation cb665245-fe7b-4fe7-8099-28367790350a · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.415273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.415273Z digest=sha256:c6075ba222724b876c25269292f07edf5c338f0ce76ead1e065867874bc02023

Observation 637ce2fc-f4db-4fe4-b052-248e35cac2f5 · outbound

This paper cites Wearable reasoner: towards enhanced human rational- ity through a wearable device with an explainable ai assistant.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Wearable reasoner: towards enhanced human rational- ity through a wearable device with an explainable ai assistant

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.420252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.420252Z digest=sha256:931a8d0ef59b67e836fe1dc12729aa73dd90f891db5df4c86234417c8ce7e694

Observation 07606b05-82c5-4b80-9ec1-0f787131951e · outbound

This paper cites Summarization of egocentric videos: A com- prehensive survey.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Summarization of egocentric videos: A com- prehensive survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.424394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.424394Z digest=sha256:d45449ff6498522a885d2243f6e5907f002ceebb07c4534cd1e1f52ada2b6149

Observation 07b030a9-e035-4a6e-85e2-8c3ec4ba9eee · outbound

This paper cites Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.428452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.428452Z digest=sha256:171cdbff104dba5e44ba712e543b0c7ebcb6cadbe461c5f873f08ac9b09d18ef

Observation 578d1f27-e4aa-4a22-a3ec-4589f6a157d3 · outbound

This paper cites Learning to recognize objects in egocentric activities.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning to recognize objects in egocentric activities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.433221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.433221Z digest=sha256:8a8595a6c88e4c4af7ab12dc18017caed727f41d7b5b8bd43e664fd8f241ec0e

Observation a5e069c1-68bd-4df0-9b8a-1e6ff126d489 · outbound

This paper cites What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.438325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.438325Z digest=sha256:57500142a571a3c15eaef5967fc36df5f008a8e6e8f54563e9176b22c5b8b00e

Observation f462ad46-9f37-4d30-a231-4daebe4d2813 · outbound

This paper cites What would you expect? anticipating egocentric actions with rolling- unrolling lstms and modality attention.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model What would you expect? anticipating egocentric actions with rolling- unrolling lstms and modality attention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.845089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.442431Z digest=sha256:0343ca04dabba4ceb3569679d01002ce5e7c15474b7fa9e665cc6bdc2f813e4b

Observation 89edebc3-d597-43e0-aebc-d6e2b6e32a72 · outbound

This paper cites Unsupervised video summarization via relation- aware assignment learning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unsupervised video summarization via relation- aware assignment learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.831336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.446939Z digest=sha256:281d2400a7ded4323e979514aae54ceaef2ea3f9dc4cc103bf88278609a4b067

Observation 40f8dee7-3026-4cc8-9e26-4bc195da9efb · outbound

This paper cites Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.817776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.451258Z digest=sha256:6330286ab706d9fe39aaf1a8ba07e69b4df7901574de2a66b79f0dccefec181b

Observation 3a10ba8a-273e-4a78-b4f4-5f6d5393a6c9 · outbound

This paper cites Ego4d: Around the World in 3,000 Hours of Egocentric Video.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego4d: Around the World in 3,000 Hours of Egocentric Video

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.804685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.456268Z digest=sha256:92e2d93723a0cb6b54278a42a5845508ef48bc6280b81d17d8d27cc3b74cdcd7

Observation 9710b1ce-8982-4893-abfb-0f3e58fa7d4f · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.460542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.460542Z digest=sha256:626ab91a47205402a23a2782cc100414884161393f73df21524d64e2e2290b9f

Observation 04aaf6b1-9e8b-47a9-aef3-457d7cc7bc63 · outbound

This paper cites an unresolved cited work.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:07:15.791032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.465095Z digest=sha256:ca3eb5d5f02c0a7b37458210569c2491f65cb1d463209176ab22099d5d6b741f

Observation 5f98856a-bfa2-4c60-aa2e-3370688c9bb9 · outbound

This paper cites Pre- dicting gaze in egocentric video by learning task-dependent attention transition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Pre- dicting gaze in egocentric video by learning task-dependent attention transition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.777844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.469409Z digest=sha256:d817e97483c5971410da571b284d143fb0623174e58a76df8bafbdc48bde000f

Observation a32170cf-d6d8-4c90-a14e-28271db4de5b · outbound

This paper cites Mutual context network for jointly estimating egocentric gaze and action.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Mutual context network for jointly estimating egocentric gaze and action

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.764659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.473685Z digest=sha256:7676da878aeba4164cbdcf7d9671bd6d3931f138cc16da9ab2c4e321d0bc66dd

Observation cb7f878e-9aff-409b-a534-9e67225431f0 · outbound

This paper cites Improving action segmentation via graph-based temporal reasoning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Improving action segmentation via graph-based temporal reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.751203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.477883Z digest=sha256:395981cfade9a2ab125f0eb8dd9863642a7c63e6f9d8d51172604f0e959de10a

Observation 5b59425e-f21e-47eb-b220-cfc4b4b14382 · outbound

This paper cites Compound proto- type matching for few-shot action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Compound proto- type matching for few-shot action recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.737841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.482706Z digest=sha256:6bf061ae0a0132bd79ba968a6db9ebc2ab2bc472793b7c2858926c6956cacbd5

Observation 57d405e9-0d19-42da-8818-f7f8fbfd345e · outbound

This paper cites Weakly supervised temporal sentence grounding with uncertainty-guided self- training.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Weakly supervised temporal sentence grounding with uncertainty-guided self- training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.724760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.487063Z digest=sha256:4265dd9b193c7fa2d3d95d58ee934cc652e92212dc024d75329937b26bf7ab4d

Observation 01eab2a6-f1b0-4539-8199-7abbb6f3379d · outbound

This paper cites Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.710985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.491334Z digest=sha256:dc4fcf084c5badac6ad3e0c6f5e5ca0924d859ff945fb59d9af80806f820552b

Observation 5f22b972-d73c-4811-af51-e593c90182bb · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VBench: Com- prehensive benchmark suite for video generative models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.697712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.496055Z digest=sha256:c3b92b7818757693ac481bd24e6a9c2de0098776e145550bea8e9c915932d3fa

Observation f9607eb2-48f9-4602-b167-d3d1df282013 · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.500140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.500140Z digest=sha256:51a527f7d1362c5fb761843b1e841b23cd82c63689ba31f0ee04ecb46124e3bd

Observation 5043d211-f39b-4b40-a2c1-030ab139a71d · outbound

This paper cites Towards intelligent wearable assistants.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Towards intelligent wearable assistants

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.684460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.504426Z digest=sha256:7675b79b5b08e2cf2ac2a7b3c9b6ff5040fc5e4c0abbad791d7c22877db018f2

Observation 8a9609b7-fd27-44df-89d7-959602087926 · outbound

This paper cites Demonstrating tom: A de- velopment platform for wearable intelligent assistants.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Demonstrating tom: A de- velopment platform for wearable intelligent assistants

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.671108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.508558Z digest=sha256:2e86df487f820346c5fc58253f1541cb24413b6d668b6ecf591374a0742483ca

Observation a2c1f624-2e61-4721-85f2-6f3069ba72f2 · outbound

This paper cites Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.657731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.512646Z digest=sha256:e98466b97505a1731724f1cb3ac8c0f993a8f7133229f569b745726eddd5bd00

Observation 37a921ba-fa10-4e25-899a-45e7e0fa65ab · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Epic-fusion: Audio-visual temporal binding for egocentric action recognition

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.643272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.516836Z digest=sha256:c3206e287816abeac1cfdca6ed315675bdd8ba9befb0ce90c02b177f0b772a45

Observation 7f0aeb99-c015-42ee-b426-06c7b4d2d990 · outbound

This paper cites Time- conditioned action anticipation in one shot.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Time- conditioned action anticipation in one shot

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.628326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.520886Z digest=sha256:667158cadb47e6bbcab289841cb5388a0885cde0526462a8187981a2a4be3a34

Observation 6f55c37f-ffd7-4ebf-bccc-598ff346f20c · outbound

This paper cites Learning to discriminate information for online action detection: Anal- ysis and application.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning to discriminate information for online action detection: Anal- ysis and application

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.613474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.525005Z digest=sha256:5585f7ce780c7d79648194e8493a636aaaec9614fc79b0c4738e994dd37af81b

Observation 033e8250-9af6-4c50-bc38-847af0fefbc8 · outbound

This paper cites Ego-body pose es- timation via ego-head pose estimation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego-body pose es- timation via ego-head pose estimation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.599451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.529118Z digest=sha256:c56a240a7b1efd2b3d29752fd2a65ecedeaab875476894a6179e5c4cb93d2a95

Observation ab47944e-5706-4a95-8f8d-40b98bedb0f2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoChat: Chat-Centric Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.533205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.533205Z digest=sha256:a0e80459c82235a7d76fdd492dd3bde3cf6cb5d466ca6a24c108f44141ec3816

Observation d3818a83-dd62-4b37-8639-268bd91027d6 · outbound

This paper cites Jointly localizing and describing events for dense video captioning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Jointly localizing and describing events for dense video captioning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.585995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.537667Z digest=sha256:20f77dbec657e8c35f004723e40034124a7cd0318bde65b57bac5767e2e2af9b

Observation 555b75c5-1459-4a1f-b9ae-1e22cf884672 · outbound

This paper cites Egocentric video-language pretraining.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egocentric video-language pretraining

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.571544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.541881Z digest=sha256:026773e02492b5729d1c56551aea31444a0e2a096082c3fd916f97b579d31617

Observation 8b7c6741-4a86-42e0-aee9-fb177c7aa4c2 · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Univtg: Towards unified video-language temporal grounding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.557145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.546735Z digest=sha256:77739dd767af73bc0fcf3158b8947e764a93516579556fdb7176e67da3c431a1

Observation 497a0e72-4612-42d2-9410-fd80b9bbd3c7 · outbound

This paper cites Visual instruction tuning, 2023.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Visual instruction tuning, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.551203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.551203Z digest=sha256:bdffaac6aa94b6417321bf55022702a1f9b39a1991ef1e1e9b39444bdd1a4e33

Observation f691b753-7a40-4925-8a99-9d714bcb9a69 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.555241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.555241Z digest=sha256:69a6299a7865713dda007ff5eda29976781bb1961945e1cf9cac28ecc86fd895

Observation 8ce5f307-e0e7-49f3-8b7b-845bbfca09bc · outbound

This paper cites Video summariza- tion through reinforcement learning with a 3d spatio-temporal u-net.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Video summariza- tion through reinforcement learning with a 3d spatio-temporal u-net

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.524779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.559366Z digest=sha256:5eec2108d19f866629277d9a73ceb2f9c701df451a39ffe3781a0d369d1b4982

Observation 342c998e-cb91-489a-a796-84db9e2d074e · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Latte: Latent Diffusion Transformer for Video Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.563844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.563844Z digest=sha256:db6500ee0ece212760f5f7ed21892e22d1e9ac3e6af3e9148c237e7eb0fbbded

Observation fe3883d7-c724-4e51-9add-9f19b4d31f53 · outbound

This paper cites Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.568307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.568307Z digest=sha256:65fe7ebaac4b8b2fba826f157caf9d64acf417d839a25f18a7bf277850565b87

Observation 32dc1b55-0211-4429-8a17-76bb31f0c17b · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.510814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.572712Z digest=sha256:101643c337154b9c7c67fec2e73932be2cb644a0c7b4390f071b24df40bb770c

Observation 6cd379de-9fe6-4d5c-9a60-6585048f0de1 · outbound

This paper cites Integrating human gaze into attention for egocentric activity recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Integrating human gaze into attention for egocentric activity recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.496469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.576732Z digest=sha256:bd0a30b8f818ba796f3aec88f2ebca0cd4b740d1a3bcaf2f62f85001e5c83aaa

Observation ed6c5aa4-8cde-446b-82d6-450e3f548839 · outbound

This paper cites Learning affordance landscapes for interaction exploration in 3d environments.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning affordance landscapes for interaction exploration in 3d environments

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.482815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.581234Z digest=sha256:b5c50b2ceb4303979d20f6358bfda8d25755b68f6d15a1ae174e2957359da163

Observation 2d62998d-ebd2-4b6e-bd18-f88fb293ecff · outbound

This paper cites Egoenv: Human- centric environment representations from egocentric video.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egoenv: Human- centric environment representations from egocentric video

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.469334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.585398Z digest=sha256:e8fa8d4233938194ed31660359a13236e6e88292463691ab2f2f95696d46ccd2

Observation 2e359030-db96-472c-a8fb-ec4463a52fb7 · outbound

This paper cites Gpt-4v(ision) system card.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gpt-4v(ision) system card

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.589425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.589425Z digest=sha256:334bb24526fcca6ecd40ee63df9dc485cea6531adcf2766c4030ea43bc4f1e23

Observation 938812b2-fc71-4fcc-9b59-d0733978175e · outbound

This paper cites Hello gpt-4o.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Hello gpt-4o

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.447088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.593620Z digest=sha256:3b1a20c11ea7a8b7e0f3b31cd313c2d3acab5304f16699d0d0472ccf31c11087

Observation 7b5445b0-eb61-4d24-b845-e537c137950a · outbound

This paper cites Actionvos: Actions as prompts for video object segmentation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Actionvos: Actions as prompts for video object segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.434344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.598767Z digest=sha256:89a1ce80dfadd01a3bb2b2bebffb9c14ce8a1feb34871ecd2938c9769b3808e8

Observation 2a983fef-1a55-4d00-bcfc-90257592f88d · outbound

This paper cites Wear- able augmented reality system using gaze interaction.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Wear- able augmented reality system using gaze interaction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.421551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.602833Z digest=sha256:f196fa2209f090dfef12ed24a9bdcce9471aff9e1d12ac3b7732e8564d975925

Observation 4afb15f6-053e-426f-b60c-2f45d49edb71 · outbound

This paper cites Deep learning-based smart task assistance in wearable augmented reality.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Deep learning-based smart task assistance in wearable augmented reality

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.408485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.606907Z digest=sha256:9e8e08dbe2cda6a178b52c4c13d0394cc6f69d46dd027d3dbe453240e9a0ccb2

Observation d4308c20-0da5-4c3f-84bc-367fcd7924b8 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.610984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.610984Z digest=sha256:5dfdb77b54f63483c670c0fefab08ce3b682943647a71bbbb2da1f9c1d92320f

Observation 8420b244-53af-4be1-a537-30ec777c7384 · outbound

This paper cites An Outlook into the Future of Egocentric Vision.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model An Outlook into the Future of Egocentric Vision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.615737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.615737Z digest=sha256:52c45217c803ad85ebb40cd3e0d07fba3d40d1eca7e52aa8058e084ed3136bf5

Observation e3853078-fcfb-40fc-8cc1-f4888ae96575 · outbound

This paper cites Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.620320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.620320Z digest=sha256:6cffa7ce1493999fc459fa01d9afbbfb5dc6d6bb095493d31068e80f66c89a96

Observation 57f3df04-9bd8-4ea8-b715-bfa70ff40e5d · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.395284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.624620Z digest=sha256:366f81094936639bd3366b18a0d43ec14bf6d93c570369b6c31caeb65e7e8cbe

Observation 8ac2dbfe-8937-49b8-abff-77075bfa34ab · outbound

This paper cites Streaming long video understanding with large language models, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Streaming long video understanding with large language models, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.382002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.628691Z digest=sha256:839f124649bfe7a40b88dd166f4846887912ccaa931b88277da8c7bfc562d78d

Observation 5b4a8357-e9e3-4b64-9660-1b69d82fe91a · outbound

This paper cites Sener, D.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Sener, D

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.368150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.633110Z digest=sha256:a021d6fcf964ce757fcb5ce2d8d71904a27c2df5fbec0960bbb7494275dcc554

Observation 3d42fe93-c5a6-4946-b197-a4ad4620c5b4 · outbound

This paper cites Understanding human hands in contact at internet scale.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Understanding human hands in contact at internet scale

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.355037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.637413Z digest=sha256:cd5925b8f46608c8251c3f70122f5a3320769012ad4099bcb7fd8d52c1794192

Observation 0e27b10e-9437-4c19-b883-ebf2237264d5 · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.641633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.641633Z digest=sha256:15a6a8a81897b945ace4c7a0086f786f85ba37f09972c0b905e769e75c78c0c7

Observation 77f1364a-4339-4342-9e57-e2ed89ed533c · outbound

This paper cites Towards diverse paragraph captioning for untrimmed videos.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Towards diverse paragraph captioning for untrimmed videos

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.342012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.645828Z digest=sha256:0dc84837a92f5c7c6b76414843553e24ebe4fec8900557f997fc00b6a436edbf

Observation 4e0cec2e-e935-4890-9a0a-3a188af385ec · outbound

This paper cites Ego4d goal-step: Toward hierarchical understanding of procedural activities.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego4d goal-step: Toward hierarchical understanding of procedural activities

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.329069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.650645Z digest=sha256:38ee361d028f8b9caa3b787d96a23551f0f467cd155f08452ba28ea324110699

Observation fb3115e4-9c67-4df4-8146-ae45bb403574 · outbound

This paper cites H+ o: Unified egocentric recognition of 3d hand-object poses and interactions.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model H+ o: Unified egocentric recognition of 3d hand-object poses and interactions

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.316166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.654852Z digest=sha256:dd8b51a8a2b209451f7b474a55f9ea21c9063c26cbb01ba395cee550ed20f267

Observation 5561dfa5-0ae8-417f-b133-8c3e6663bf47 · outbound

This paper cites Memory-and-anticipation transformer for online action understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Memory-and-anticipation transformer for online action understanding

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.302082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.658845Z digest=sha256:adde7938430d0f7c8d32342ec3dc8b0b334b2020d63b7c166468569b0ef84243

Observation 6e38406f-8142-4d7d-aebe-d4fd380e192e · outbound

This paper cites Scene-aware ego- centric 3d human pose estimation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Scene-aware ego- centric 3d human pose estimation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.288716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.663095Z digest=sha256:d46a272edecbd3af8e68e30b7f055ad6b94611c959e9d57f352739182ccda798

Observation 48a2e583-2d61-44fd-a5fc-bc6e31e7b6bd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.668469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.668469Z digest=sha256:8ecef688ed552f6edf791f8acdaf468896d96651415151d5c5f17238d04bbd12

Observation fa3e5392-a256-4d5e-b7b3-539d03256832 · outbound

This paper cites an unresolved cited work.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:07:15.275336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.672801Z digest=sha256:4a882266817d7fdaa5c9aca42a793b35dc5f090e65b3d06959040a056fb4347a

Observation c7ba0f18-9f9e-46ca-ba13-7896ae059b56 · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Lavie: High-quality video generation with cascaded latent diffusion models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.261999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.677185Z digest=sha256:3f2eb6dd9908375c88f8d41f546193a62849493e87c5fd247fd160b8a7c29d20

Observation 6f18991b-1848-4299-9076-7ae89dc2e674 · outbound

This paper cites In- ternvideo2: Scaling foundation models for multimodal video understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model In- ternvideo2: Scaling foundation models for multimodal video understanding

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.248022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.681282Z digest=sha256:e264decdc5ee3948beb5680c0fdac017933616ad4f94d164a6b05bc31b201328

Observation 1c5e0a52-90b1-47f4-a653-29aa26d6f1ad · outbound

This paper cites Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.234972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.685720Z digest=sha256:0685d721dd29a8afe610fb0c5e3263034bc3321f8ba015b6ce94433cc414d623

Observation 7f1babe2-6713-49e2-b355-bf6c3774d7d8 · outbound

This paper cites Gaze-enabled egocentric video summarization via constrained submodular maximiza- tion.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gaze-enabled egocentric video summarization via constrained submodular maximiza- tion

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.220521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.689868Z digest=sha256:fb84f48811f634047eba0a720c97a4ec51e6854e73f9c88ee92a954061da126f

Observation 60f48183-1dc4-4dc3-afd6-b972d1a6b482 · outbound

This paper cites Retrieval-augmented egocentric video captioning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Retrieval-augmented egocentric video captioning

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.207403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.693756Z digest=sha256:8cbb55d7294fbc465c3ebc8a030f6f3358f472ba241704d99f7d9aff2f8f1675

Observation 17f451e4-2bd5-44c9-b122-156c80b385b0 · outbound

This paper cites Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.194360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.697840Z digest=sha256:b021e364fd92bb88f45bd7fd5f4760eaf7c67f7e0be1a3e19ce8190b5c2c9761

Observation 751d0146-664a-4d0c-8e48-52cf0d3f0c64 · outbound

This paper cites Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.180905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.702716Z digest=sha256:baf960e1982e07f68029116864e98f7933422bf0752badc18d1edeae2352006c

Observation f6c3e92b-cf47-4dd9-9c3b-97c06c6c5c89 · outbound

This paper cites Deco: Decomposition and reconstruction for compositional temporal grounding via coarse-to-fine contrastive ranking.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Deco: Decomposition and reconstruction for compositional temporal grounding via coarse-to-fine contrastive ranking

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.166916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.707238Z digest=sha256:ba8a913c0fd4a6f5c383367ce89c9e63742e5fb421decdf0990c14bf145207d3

Observation 6ad907d8-7b73-404d-8b38-99d84344bd5a · outbound

This paper cites Basictad: an astounding rgb-only baseline for temporal action detection.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Basictad: an astounding rgb-only baseline for temporal action detection

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.153204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.711279Z digest=sha256:ab9081261a80d5c40ef09f33c1ab65cd84ff3372552e6058f265902e50fca033

Observation c35436ff-01b6-4ef2-ac97-f289e8add419 · outbound

This paper cites Flash-vstream: Memory- based real-time understanding for long video streams, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Flash-vstream: Memory- based real-time understanding for long video streams, 2024

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.139248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.715438Z digest=sha256:440c8fb64c93e46e103ea643f8957519408048fda5ad5676dc8dcd9dc52a20b8

Observation b81ca1b8-ad46-4f13-82ec-552fdac0b5e5 · outbound

This paper cites Fine-grained egocentric hand-object segmentation: Dataset, model, and applications.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Fine-grained egocentric hand-object segmentation: Dataset, model, and applications

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.125890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.719771Z digest=sha256:6df32fb6c93a3e987dc71fd443a39f6c1f3eaf872dfbeb56aa112925b36fceb7

Observation b52f9a23-813b-4b13-9d2d-4d2c24a6c2c6 · outbound

This paper cites Masked video and body-worn imu autoencoder for egocentric action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Masked video and body-worn imu autoencoder for egocentric action recognition

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.110023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.724292Z digest=sha256:b5d73ce6bc5607a6f7b7d9333408b639c70970cd8e79aa7721440f57a70bf731

Observation a38889f8-ece3-444c-9287-1e18ec78ffa5 · outbound

This paper cites Internlm-xcomposer2.5-omnilive: A comprehensive multimodal system for long-term streaming video and audio interactions, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Internlm-xcomposer2.5-omnilive: A comprehensive multimodal system for long-term streaming video and audio interactions, 2024

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.095032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.728181Z digest=sha256:9de09458af579acacc8d314ee31519985bfb09dccc04e7c71fed0c5442defaee

Observation 7a6bd842-ceb9-4c81-9175-3e6b937781aa · outbound

This paper cites Ego- body: Human body shape and motion of interacting people from head-mounted devices.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego- body: Human body shape and motion of interacting people from head-mounted devices

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.080915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.732377Z digest=sha256:3ad72fcfbb52f4fa2fa144407c6c585253c48c68a076c3532781f55c532bf433

Observation a7309c89-6de1-4361-8b8b-09e82da33cdd · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Training a Large Video Model on a Single Machine in a Day

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.736293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.736293Z digest=sha256:b070be90d9e13144195622b74235da12e9c2ca92454dfd196c8d2d8a486e369f

Observation cbd5f695-acc0-43e6-993a-83801994edfc · outbound

This paper cites Learning video representations from large language mod- els.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning video representations from large language mod- els

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.065933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.740359Z digest=sha256:363cb2b65c1f08a198d720f3712d94b17cfedfdab5fb38d8148348cf95f6b8ff

Observation 872ce457-d2de-463e-b4fc-1e5a2a4f3de8 · outbound

This paper cites Learning video representations from large language models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning video representations from large language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.051092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.744403Z digest=sha256:1ce5d7933fbf27a77e50e1a530ec2df362a205cf6a94ab6aeb4edc4085f443ff

Observation 0fb7ad97-c9ee-46a7-afaa-4dbc2565cbac · outbound

This paper cites Mrsn: Multi-relation support network for video action detec- tion.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Mrsn: Multi-relation support network for video action detec- tion

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.036512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.748483Z digest=sha256:25c235a010fc7d783526414afe1f7b524c7f2201cd26be338c29011a2cf2890a

Pith citing papers

Observation 72f24855-0ec1-4f79-87f8-ae86f6cb6ffc · inbound

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance cites this paper.

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:37.248378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:37.248378Z digest=sha256:33965aba527ac00f9633cfa38bc3b4d9eddc9577e17a8e9cdbdadfc093950690

Observation b8e01093-7009-49ed-a8c0-4e8fe09081b2 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.228004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.228004Z digest=sha256:7bbffcc38462ec5600dbc5c00f5fa8d1f8c174aa889c5df18afaa3a6fad638e5

Observation c9468671-ac66-4673-a0f2-5a325648989c · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:58.390895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:58.390895Z digest=sha256:7b478abecf09459053c864231f2fe8d966617f1a5507c9ae722cc111feac6bfa

Observation ceef77ff-df8a-4e65-a6c3-7beb71ea261f · inbound

Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant cites this paper.

Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:50:51.441624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T10:49:25.312165Z digest=sha256:8a47f547c377492786d6d76eeaff19f891a975b1d2fe32303cfadb6c82d71aff

Observation cbd04bbf-28cc-484d-875c-2bd4d4a42efa · inbound

Human-Inspired Context-Selective Multimodal Memory for Social Robots cites this paper.

Human-Inspired Context-Selective Multimodal Memory for Social Robots Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:05.820629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:37:56.427507Z digest=sha256:cf23b503f77a5609b79ffba7c852930282a40601fdc3a5919487ba63e17b4718

Observation 38d7a5d8-2702-45b7-a673-c5a029976b33 · inbound

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding cites this paper.

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.307729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T00:26:33.306128Z digest=sha256:467c1e09c876cde0412e5783424c4feaffdb0d796d46695229284d1cb58efb01

Observation 8a64f381-c73e-49db-903b-8d37a2beb353 · inbound

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos cites this paper.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T05:01:06.200663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:01:06.200663Z digest=sha256:2be6cd6ce744aada3441711a0826d509bd61d2a66e1ef8bc14f1482f4f2c2946