Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:07:14.748483Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 97 of 97 outbound references and 7 inbound Pith citation observations for arXiv:2412.21080.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:07:14.748483Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:37.248378Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:57.306268Z
97 of 97 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bd6ddc5a-13a4-4b4a-ab2c-4f79252f26e1 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gemini: A Family of Highly Capable Multimodal Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfcf406a-9ecc-4c4b-b4b3-162a9bb23f32 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Swim- master: a wearable assistant for swimmer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0e6e6d-4753-4892-bef7-46e60c23a757 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d301eb7f-de99-49f7-8b55-c29500649ccf · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Analysis of the hands in egocentric vision: A survey
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe67223-53fb-4d19-a7d8-973674cb4640 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8455371b-fdc7-433d-915e-7ec5c1a59c15 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Internlm2 technical report, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2717d3-f7b1-46d3-8f65-7ff8d8ec4e97 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d80c5428-67c9-4371-9864-88b495c6c5d9 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Dcan: improving temporal action detection via dual context aggregation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5eb3f21-fe80-4eee-9c0d-550b576a4a92 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Elan: Enhancing temporal action detection with location awareness
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbaa22a6-33f3-40a5-a81b-6b74f22d5da3 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoLLM: Modeling Video Sequence with Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac52273-6630-49e3-be8e-ddcac21a80af · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6578085-a73b-4727-b924-867d9c0aec1c · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc325425-3ef9-43cf-bd16-228429bf1ae4 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gatehub: Gated history unit with background suppression for online action detection
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966d4664-06ae-4ffa-bcc8-e76b8ceae919 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Videollm-online: Online video large language model for streaming video
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647ee9ad-6757-45c9-95c8-dbc54f48f491 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Seine: Short-to-long video diffusion model for generative transition and prediction
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e0bdab-c45e-44c5-a275-d2b9e7c40fe5 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model gSDF: Geometry-Driven signed distance functions for 3D hand-object reconstruction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4005a2-c553-4284-93bf-4d199e027499 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32aca28f-97e8-4598-b5e7-61f0a41f8529 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20985824-880b-43c4-9a11-69fe87a6f4e5 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Scaling egocentric vision: The epic-kitchens dataset
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb665245-fe7b-4fe7-8099-28367790350a · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 637ce2fc-f4db-4fe4-b052-248e35cac2f5 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Wearable reasoner: towards enhanced human rational- ity through a wearable device with an explainable ai assistant
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07606b05-82c5-4b80-9ec1-0f787131951e · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Summarization of egocentric videos: A com- prehensive survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b030a9-e035-4a6e-85e2-8c3ec4ba9eee · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 578d1f27-e4aa-4a22-a3ec-4589f6a157d3 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning to recognize objects in egocentric activities
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e069c1-68bd-4df0-9b8a-1e6ff126d489 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f462ad46-9f37-4d30-a231-4daebe4d2813 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model What would you expect? anticipating egocentric actions with rolling- unrolling lstms and modality attention
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 89edebc3-d597-43e0-aebc-d6e2b6e32a72 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unsupervised video summarization via relation- aware assignment learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 40f8dee7-3026-4cc8-9e26-4bc195da9efb · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3a10ba8a-273e-4a78-b4f4-5f6d5393a6c9 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego4d: Around the World in 3,000 Hours of Egocentric Video
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9710b1ce-8982-4893-abfb-0f3e58fa7d4f · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04aaf6b1-9e8b-47a9-aef3-457d7cc7bc63 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5f98856a-bfa2-4c60-aa2e-3370688c9bb9 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Pre- dicting gaze in egocentric video by learning task-dependent attention transition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a32170cf-d6d8-4c90-a14e-28271db4de5b · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Mutual context network for jointly estimating egocentric gaze and action
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cb7f878e-9aff-409b-a534-9e67225431f0 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Improving action segmentation via graph-based temporal reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b59425e-f21e-47eb-b220-cfc4b4b14382 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Compound proto- type matching for few-shot action recognition
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 57d405e9-0d19-42da-8818-f7f8fbfd345e · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Weakly supervised temporal sentence grounding with uncertainty-guided self- training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01eab2a6-f1b0-4539-8199-7abbb6f3379d · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5f22b972-d73c-4811-af51-e593c90182bb · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VBench: Com- prehensive benchmark suite for video generative models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f9607eb2-48f9-4602-b167-d3d1df282013 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5043d211-f39b-4b40-a2c1-030ab139a71d · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Towards intelligent wearable assistants
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a9609b7-fd27-44df-89d7-959602087926 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Demonstrating tom: A de- velopment platform for wearable intelligent assistants
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a2c1f624-2e61-4721-85f2-6f3069ba72f2 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 37a921ba-fa10-4e25-899a-45e7e0fa65ab · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7f0aeb99-c015-42ee-b426-06c7b4d2d990 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Time- conditioned action anticipation in one shot
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f55c37f-ffd7-4ebf-bccc-598ff346f20c · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning to discriminate information for online action detection: Anal- ysis and application
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 033e8250-9af6-4c50-bc38-847af0fefbc8 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego-body pose es- timation via ego-head pose estimation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab47944e-5706-4a95-8f8d-40b98bedb0f2 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoChat: Chat-Centric Video Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3818a83-dd62-4b37-8639-268bd91027d6 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Jointly localizing and describing events for dense video captioning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 555b75c5-1459-4a1f-b9ae-1e22cf884672 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egocentric video-language pretraining
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b7c6741-4a86-42e0-aee9-fb177c7aa4c2 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Univtg: Towards unified video-language temporal grounding
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 497a0e72-4612-42d2-9410-fd80b9bbd3c7 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Visual instruction tuning, 2023
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f691b753-7a40-4925-8a99-9d714bcb9a69 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Llava-next: Improved reason- ing, ocr, and world knowledge, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce5f307-e0e7-49f3-8b7b-845bbfca09bc · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Video summariza- tion through reinforcement learning with a 3d spatio-temporal u-net
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 342c998e-cb91-489a-a796-84db9e2d074e · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Latte: Latent Diffusion Transformer for Video Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe3883d7-c724-4e51-9add-9f19b4d31f53 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32dc1b55-0211-4429-8a17-76bb31f0c17b · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6cd379de-9fe6-4d5c-9a60-6585048f0de1 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Integrating human gaze into attention for egocentric activity recognition
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ed6c5aa4-8cde-446b-82d6-450e3f548839 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning affordance landscapes for interaction exploration in 3d environments
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2d62998d-ebd2-4b6e-bd18-f88fb293ecff · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egoenv: Human- centric environment representations from egocentric video
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2e359030-db96-472c-a8fb-ec4463a52fb7 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gpt-4v(ision) system card
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938812b2-fc71-4fcc-9b59-d0733978175e · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Hello gpt-4o
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7b5445b0-eb61-4d24-b845-e537c137950a · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Actionvos: Actions as prompts for video object segmentation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a983fef-1a55-4d00-bcfc-90257592f88d · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Wear- able augmented reality system using gaze interaction
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4afb15f6-053e-426f-b60c-2f45d49edb71 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Deep learning-based smart task assistance in wearable augmented reality
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d4308c20-0da5-4c3f-84bc-367fcd7924b8 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8420b244-53af-4be1-a537-30ec777c7384 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model An Outlook into the Future of Egocentric Vision
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3853078-fcfb-40fc-8cc1-f4888ae96575 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f3df04-9bd8-4ea8-b715-bfa70ff40e5d · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8ac2dbfe-8937-49b8-abff-77075bfa34ab · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Streaming long video understanding with large language models, 2024
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b4a8357-e9e3-4b64-9660-1b69d82fe91a · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Sener, D
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d42fe93-c5a6-4946-b197-a4ad4620c5b4 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Understanding human hands in contact at internet scale
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e27b10e-9437-4c19-b883-ebf2237264d5 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f1364a-4339-4342-9e57-e2ed89ed533c · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Towards diverse paragraph captioning for untrimmed videos
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4e0cec2e-e935-4890-9a0a-3a188af385ec · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego4d goal-step: Toward hierarchical understanding of procedural activities
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb3115e4-9c67-4df4-8146-ae45bb403574 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model H+ o: Unified egocentric recognition of 3d hand-object poses and interactions
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5561dfa5-0ae8-417f-b133-8c3e6663bf47 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Memory-and-anticipation transformer for online action understanding
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6e38406f-8142-4d7d-aebe-d4fd380e192e · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Scene-aware ego- centric 3d human pose estimation
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 48a2e583-2d61-44fd-a5fc-bc6e31e7b6bd · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3e5392-a256-4d5e-b7b3-539d03256832 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c7ba0f18-9f9e-46ca-ba13-7896ae059b56 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Lavie: High-quality video generation with cascaded latent diffusion models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f18991b-1848-4299-9076-7ae89dc2e674 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model In- ternvideo2: Scaling foundation models for multimodal video understanding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1c5e0a52-90b1-47f4-a653-29aa26d6f1ad · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format, 2024
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7f1babe2-6713-49e2-b355-bf6c3774d7d8 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gaze-enabled egocentric video summarization via constrained submodular maximiza- tion
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 60f48183-1dc4-4dc3-afd6-b972d1a6b482 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Retrieval-augmented egocentric video captioning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 17f451e4-2bd5-44c9-b122-156c80b385b0 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 751d0146-664a-4d0c-8e48-52cf0d3f0c64 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f6c3e92b-cf47-4dd9-9c3b-97c06c6c5c89 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Deco: Decomposition and reconstruction for compositional temporal grounding via coarse-to-fine contrastive ranking
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6ad907d8-7b73-404d-8b38-99d84344bd5a · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Basictad: an astounding rgb-only baseline for temporal action detection
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c35436ff-01b6-4ef2-ac97-f289e8add419 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Flash-vstream: Memory- based real-time understanding for long video streams, 2024
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b81ca1b8-ad46-4f13-82ec-552fdac0b5e5 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Fine-grained egocentric hand-object segmentation: Dataset, model, and applications
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b52f9a23-813b-4b13-9d2d-4d2c24a6c2c6 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Masked video and body-worn imu autoencoder for egocentric action recognition
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a38889f8-ece3-444c-9287-1e18ec78ffa5 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Internlm-xcomposer2.5-omnilive: A comprehensive multimodal system for long-term streaming video and audio interactions, 2024
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7a6bd842-ceb9-4c81-9175-3e6b937781aa · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego- body: Human body shape and motion of interacting people from head-mounted devices
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a7309c89-6de1-4361-8b8b-09e82da33cdd · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Training a Large Video Model on a Single Machine in a Day
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd5f695-acc0-43e6-993a-83801994edfc · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning video representations from large language mod- els
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 872ce457-d2de-463e-b4fc-1e5a2a4f3de8 · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning video representations from large language models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0fb7ad97-c9ee-46a7-afaa-4dbc2565cbac · outbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Mrsn: Multi-relation support network for video action detec- tion
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 72f24855-0ec1-4f79-87f8-ae86f6cb6ffc · inbound
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e01093-7009-49ed-a8c0-4e8fe09081b2 · inbound
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9468671-ac66-4673-a0f2-5a325648989c · inbound
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceef77ff-df8a-4e65-a6c3-7beb71ea261f · inbound
Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cbd04bbf-28cc-484d-875c-2bd4d4a42efa · inbound
Human-Inspired Context-Selective Multimodal Memory for Social Robots Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 38d7a5d8-2702-45b7-a673-c5a029976b33 · inbound
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a64f381-c73e-49db-903b-8d37a2beb353 · inbound
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.