Pith. sign in

Paper Citation Record · LEDGER

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

As of 13 August 2026, this Paper Citation Record lists 97 of 97 outbound references and 7 inbound Pith citation observations for arXiv:2412.21080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21080 v1

Coverage vector

measured 97 of 97 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:07:14.748483Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:37.248378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.306268Z

Reference resolution

97 of 97 outbound references displayed

  • verified exact0
  • verified fuzzy56
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd6ddc5a-13a4-4b4a-ab2c-4f79252f26e1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.331971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.331971Z digest=sha256:66298eeb443ced03eec24ffc39de2b90e2c52c948426bab60dcd240a68a9f9cf

Observation cfcf406a-9ecc-4c4b-b4b3-162a9bb23f32 · outbound

This paper cites Swim- master: a wearable assistant for swimmer.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Swim- master: a wearable assistant for swimmer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.337134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.337134Z digest=sha256:dcc0ee0e4d3db996e6b8eb6404ea57b2bc8b7f3d54d852298c933b41c3b5983c

Observation fa0e6e6d-4753-4892-bef7-46e60c23a757 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.341548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.341548Z digest=sha256:1f70b90c2ca86c843152f96bc215a1fcc63f790781f62ddee727d25c273d7ee4

Observation d301eb7f-de99-49f7-8b55-c29500649ccf · outbound

This paper cites Analysis of the hands in egocentric vision: A survey.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Analysis of the hands in egocentric vision: A survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.346187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.346187Z digest=sha256:b47a84843938028b544049d6a6dcd42097d21018f210255f9d0f481590939ff9

Observation 9fe67223-53fb-4d19-a7d8-973674cb4640 · outbound

This paper cites an unresolved cited work.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.350491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.350491Z digest=sha256:b6061d625de5e2118307ea553dbaccef9b7ed7c7a96227918fddc81ac02a5d13

Observation 8455371b-fdc7-433d-915e-7ec5c1a59c15 · outbound

This paper cites Internlm2 technical report, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Internlm2 technical report, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.354875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.354875Z digest=sha256:f34f27e0bb9972f3ef92a7aaca64f0842127184a6c9e274588bfaf5eb533ebe2

Observation 6f2717d3-f7b1-46d3-8f65-7ff8d8ec4e97 · outbound

This paper cites InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.359740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.359740Z digest=sha256:9b93bbd1141ed28b7853cbc7c9f7328ce8f62f3d5e910253ad233c2ae6bc860e

Observation d80c5428-67c9-4371-9864-88b495c6c5d9 · outbound

This paper cites Dcan: improving temporal action detection via dual context aggregation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Dcan: improving temporal action detection via dual context aggregation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.364280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.364280Z digest=sha256:77a66ea157b0d7369d3b74721a5edcd793eec3dfdc803d04bb6a622ad0b9303c

Observation d5eb3f21-fe80-4eee-9c0d-550b576a4a92 · outbound

This paper cites Elan: Enhancing temporal action detection with location awareness.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Elan: Enhancing temporal action detection with location awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.368251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.368251Z digest=sha256:83018c69c603e5fb6ddfa81349098bd5b0787165f321bdf8ed894427921100d9

Observation bbaa22a6-33f3-40a5-a81b-6b74f22d5da3 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoLLM: Modeling Video Sequence with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.372376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.372376Z digest=sha256:bcfca418846ddaf94ea3872bc639b5432950418995d97b7d77d792052cb509ad

Observation cac52273-6630-49e3-be8e-ddcac21a80af · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.376930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.376930Z digest=sha256:56145dd7304be43f23cc6a73650b26d7245ab5bda8f70ee3265cdd3c6312b613

Observation c6578085-a73b-4727-b924-867d9c0aec1c · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.381559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.381559Z digest=sha256:7d31a8b6aa98c6d2436cf403277f0b3fa105ec8b3ad16fbc8834cc88ccdec3e3

Observation dc325425-3ef9-43cf-bd16-228429bf1ae4 · outbound

This paper cites Gatehub: Gated history unit with background suppression for online action detection.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gatehub: Gated history unit with background suppression for online action detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.385876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.385876Z digest=sha256:bcf30c3c386271ee524a33c2057d03e2bccbc762922b9b6d13d0156d99cb032c

Observation 966d4664-06ae-4ffa-bcc8-e76b8ceae919 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Videollm-online: Online video large language model for streaming video

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.389760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.389760Z digest=sha256:95af2c894639efe8793abeec61b83d57bf3cf760638b26ba02651c62b9c08f3a

Observation 647ee9ad-6757-45c9-95c8-dbc54f48f491 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.393831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.393831Z digest=sha256:224863547e6a663593ea3f17d20f35bc09cf15bb6496a9935a19b4f2f41fca6d

Observation f4e0bdab-c45e-44c5-a275-d2b9e7c40fe5 · outbound

This paper cites gSDF: Geometry-Driven signed distance functions for 3D hand-object reconstruction.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model gSDF: Geometry-Driven signed distance functions for 3D hand-object reconstruction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.397896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.397896Z digest=sha256:3b3820baf22c5dd9eb56f167944231b1401356d98af369c3cd05d40bb19a272c

Observation 4a4005a2-c553-4284-93bf-4d199e027499 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.402030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.402030Z digest=sha256:f9986008ef4fbc7c775d67d9809614d94fdb744c5881d958d5f84bd9529f35c8

Observation 32aca28f-97e8-4598-b5e7-61f0a41f8529 · outbound

This paper cites Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.406367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.406367Z digest=sha256:47504d9a3f059a27d82b0f91d53ccd68c99aa46b4703e0bd6f806a40dbdf5bb8

Observation 20985824-880b-43c4-9a11-69fe87a6f4e5 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Scaling egocentric vision: The epic-kitchens dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.410760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.410760Z digest=sha256:a80fbecf33fa250ecfb9e0e9e273d108f6e8d61b076b7d76f4392ed7ca3d0e02

Observation cb665245-fe7b-4fe7-8099-28367790350a · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.415273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.415273Z digest=sha256:8cdf25fac8efcf894eddb5a647feb85b5fc55ee8aa47a9c8dcdbde241c5d2b15

Observation 637ce2fc-f4db-4fe4-b052-248e35cac2f5 · outbound

This paper cites Wearable reasoner: towards enhanced human rational- ity through a wearable device with an explainable ai assistant.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Wearable reasoner: towards enhanced human rational- ity through a wearable device with an explainable ai assistant

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.420252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.420252Z digest=sha256:5b107ccb34ce7e2e33e488ccfbba3e4c76ab0937410063e1cf373af2b465646a

Observation 07606b05-82c5-4b80-9ec1-0f787131951e · outbound

This paper cites Summarization of egocentric videos: A com- prehensive survey.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Summarization of egocentric videos: A com- prehensive survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.424394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.424394Z digest=sha256:98156df0968c5e7e27b84b433dcb9f000382af4862407670421357c4e5e68281

Observation 07b030a9-e035-4a6e-85e2-8c3ec4ba9eee · outbound

This paper cites Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.428452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.428452Z digest=sha256:699f34c3bd10f92c23ba90b0549d6862fcb0b96b12c10f75c5dc454a73bffd4d

Observation 578d1f27-e4aa-4a22-a3ec-4589f6a157d3 · outbound

This paper cites Learning to recognize objects in egocentric activities.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning to recognize objects in egocentric activities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.433221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.433221Z digest=sha256:52ac3cc2fd6a6f0d676d15d541ddbe7cee473b543d1533fe2bd779f778fcb09d

Observation a5e069c1-68bd-4df0-9b8a-1e6ff126d489 · outbound

This paper cites What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.438325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.438325Z digest=sha256:25a1f8758e7598c30618d989272983bff8b671d868090b19938b78e31a9b09f7

Observation f462ad46-9f37-4d30-a231-4daebe4d2813 · outbound

This paper cites What would you expect? anticipating egocentric actions with rolling- unrolling lstms and modality attention.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model What would you expect? anticipating egocentric actions with rolling- unrolling lstms and modality attention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.845089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.442431Z digest=sha256:5b0b56e31e67f2a7b93740d1e15114863c5648258fa50989bb94c3e88e7608ef

Observation 89edebc3-d597-43e0-aebc-d6e2b6e32a72 · outbound

This paper cites Unsupervised video summarization via relation- aware assignment learning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unsupervised video summarization via relation- aware assignment learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.831336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.446939Z digest=sha256:c693c1f531e10e964c292076390366326dab37b9ff1ae4fde6b58859d4b839b4

Observation 40f8dee7-3026-4cc8-9e26-4bc195da9efb · outbound

This paper cites Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.817776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.451258Z digest=sha256:dec49e80bf721f559e41ce548c75a2b7833f87946d2d08c525b801b41a32daf7

Observation 3a10ba8a-273e-4a78-b4f4-5f6d5393a6c9 · outbound

This paper cites Ego4d: Around the World in 3,000 Hours of Egocentric Video.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego4d: Around the World in 3,000 Hours of Egocentric Video

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.804685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.456268Z digest=sha256:ce22ca4143795a5aead43c3e36cd050124500c8e287b97a02f814643cd968721

Observation 9710b1ce-8982-4893-abfb-0f3e58fa7d4f · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.460542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.460542Z digest=sha256:90c0b3bb26db6569aa067e14b13a8ecdab09271d1409672f7b31473d660e7424

Observation 04aaf6b1-9e8b-47a9-aef3-457d7cc7bc63 · outbound

This paper cites an unresolved cited work.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:07:15.791032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.465095Z digest=sha256:7ad724561db70df6d6d052df584572efefb571ccdd53e98271d6fb48ae57b234

Observation 5f98856a-bfa2-4c60-aa2e-3370688c9bb9 · outbound

This paper cites Pre- dicting gaze in egocentric video by learning task-dependent attention transition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Pre- dicting gaze in egocentric video by learning task-dependent attention transition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.777844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.469409Z digest=sha256:61afa05e084a785d8b8ae4f52f5aebb354684a79f02a082e60130727491ca34d

Observation a32170cf-d6d8-4c90-a14e-28271db4de5b · outbound

This paper cites Mutual context network for jointly estimating egocentric gaze and action.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Mutual context network for jointly estimating egocentric gaze and action

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.764659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.473685Z digest=sha256:fbf2b4fe9010f6020899e0863c43e698f6f12292cafddef746b49bdc3b27f93b

Observation cb7f878e-9aff-409b-a534-9e67225431f0 · outbound

This paper cites Improving action segmentation via graph-based temporal reasoning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Improving action segmentation via graph-based temporal reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.751203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.477883Z digest=sha256:2c4867f6dd9b37e8b159aa240cdf1dbb3ae3823dc6b7cc8bb295ded4b550ac29

Observation 5b59425e-f21e-47eb-b220-cfc4b4b14382 · outbound

This paper cites Compound proto- type matching for few-shot action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Compound proto- type matching for few-shot action recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.737841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.482706Z digest=sha256:e5a263a0c9e5b007dff88032217a2db02f2057feddb7bf66d6823809442de397

Observation 57d405e9-0d19-42da-8818-f7f8fbfd345e · outbound

This paper cites Weakly supervised temporal sentence grounding with uncertainty-guided self- training.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Weakly supervised temporal sentence grounding with uncertainty-guided self- training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.724760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.487063Z digest=sha256:a03d53804d3b22b7710c1b6a0c3648c887330288e08616281d675b76cc2680e9

Observation 01eab2a6-f1b0-4539-8199-7abbb6f3379d · outbound

This paper cites Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.710985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.491334Z digest=sha256:5d74d056a2a278cfd5f0179c5c594965d39e5bcc71e9028d011d74400b4d7442

Observation 5f22b972-d73c-4811-af51-e593c90182bb · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VBench: Com- prehensive benchmark suite for video generative models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.697712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.496055Z digest=sha256:5858fa56c66804db1e6e8243089988907a5881a35cb18348bdbe0fc31b653431

Observation f9607eb2-48f9-4602-b167-d3d1df282013 · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.500140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.500140Z digest=sha256:badb1075a46690b5da54d5efd23221d8a14c4b51a17c7ff93abebc80d1031df0

Observation 5043d211-f39b-4b40-a2c1-030ab139a71d · outbound

This paper cites Towards intelligent wearable assistants.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Towards intelligent wearable assistants

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.684460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.504426Z digest=sha256:fedea09b5d3e587ef5f97b5a227e404978b7c7662e8375bf1c4f7a0c7848858e

Observation 8a9609b7-fd27-44df-89d7-959602087926 · outbound

This paper cites Demonstrating tom: A de- velopment platform for wearable intelligent assistants.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Demonstrating tom: A de- velopment platform for wearable intelligent assistants

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.671108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.508558Z digest=sha256:63c0905974d29f6241bb9c3d9abaff6c43c6aa854431af4a7bb5330ae0a354f3

Observation a2c1f624-2e61-4721-85f2-6f3069ba72f2 · outbound

This paper cites Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.657731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.512646Z digest=sha256:f93687547cc7f2ef80a26c12bb26e5a1b6a2f0f84b4cb8e7d7eb91cbfe89c2f3

Observation 37a921ba-fa10-4e25-899a-45e7e0fa65ab · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Epic-fusion: Audio-visual temporal binding for egocentric action recognition

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.643272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.516836Z digest=sha256:ade203c5dbc1ba503136d762e426598d8b465b9e2364b98ebf9fadc2997b2415

Observation 7f0aeb99-c015-42ee-b426-06c7b4d2d990 · outbound

This paper cites Time- conditioned action anticipation in one shot.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Time- conditioned action anticipation in one shot

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.628326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.520886Z digest=sha256:5333a1c1d63180a52f4932112086d5f6f7e5dc7786341af0cd2d5955c27ab5b1

Observation 6f55c37f-ffd7-4ebf-bccc-598ff346f20c · outbound

This paper cites Learning to discriminate information for online action detection: Anal- ysis and application.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning to discriminate information for online action detection: Anal- ysis and application

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.613474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.525005Z digest=sha256:46e841f95984f251d7cc77fce3126cac81394e159aff14ef95fbfb9bb5dd455a

Observation 033e8250-9af6-4c50-bc38-847af0fefbc8 · outbound

This paper cites Ego-body pose es- timation via ego-head pose estimation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego-body pose es- timation via ego-head pose estimation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.599451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.529118Z digest=sha256:89654b400bfa4d3c7731acf7d80640ba90309a4c2c22b16c02cd7a43d45529c8

Observation ab47944e-5706-4a95-8f8d-40b98bedb0f2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model VideoChat: Chat-Centric Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.533205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.533205Z digest=sha256:ef5000a8694b7157fea86fe0544ea815c7580ebd3d0320f7ed6e4c765d500da1

Observation d3818a83-dd62-4b37-8639-268bd91027d6 · outbound

This paper cites Jointly localizing and describing events for dense video captioning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Jointly localizing and describing events for dense video captioning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.585995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.537667Z digest=sha256:b8a73283f1f4488b0522dd7c94b0e44a26cdbbbfa46d44a3abce8f4d174f4bba

Observation 555b75c5-1459-4a1f-b9ae-1e22cf884672 · outbound

This paper cites Egocentric video-language pretraining.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egocentric video-language pretraining

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.571544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.541881Z digest=sha256:88dcdc0ac4c99836de6c1c43873cbf65f7060b9b864d0076fa1597a21a856823

Observation 8b7c6741-4a86-42e0-aee9-fb177c7aa4c2 · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Univtg: Towards unified video-language temporal grounding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.557145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.546735Z digest=sha256:ae9486bf119936503246efd954e52d0e10ae10988c0f1b27a4ee9f42939a146c

Observation 497a0e72-4612-42d2-9410-fd80b9bbd3c7 · outbound

This paper cites Visual instruction tuning, 2023.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Visual instruction tuning, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.551203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.551203Z digest=sha256:ac825817690f1c9eec8b80bca11e15d81058967ab5aa2c2c9c4b92049b6417c8

Observation f691b753-7a40-4925-8a99-9d714bcb9a69 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.555241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.555241Z digest=sha256:c1d971316a21dafa7fa1288794b45d69c20c07b96a269c0b05a4ea4a1da5f6a6

Observation 8ce5f307-e0e7-49f3-8b7b-845bbfca09bc · outbound

This paper cites Video summariza- tion through reinforcement learning with a 3d spatio-temporal u-net.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Video summariza- tion through reinforcement learning with a 3d spatio-temporal u-net

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.524779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.559366Z digest=sha256:cd214bf853d9d2c617d46145f191c9b340dea8333ce7cbef1898dde2694415d1

Observation 342c998e-cb91-489a-a796-84db9e2d074e · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Latte: Latent Diffusion Transformer for Video Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.563844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.563844Z digest=sha256:dd9fce42965154478c6971701e81d11d09af7f5626d9fe856e115c85f802bedb

Observation fe3883d7-c724-4e51-9add-9f19b4d31f53 · outbound

This paper cites Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.568307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.568307Z digest=sha256:e703e39f0d937ca6279357ce3a52c6cc55dae72252d483b5b2dbf3532d794d5b

Observation 32dc1b55-0211-4429-8a17-76bb31f0c17b · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.510814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.572712Z digest=sha256:0b370cf4af778d80007f024ad5303f9c6e07e6012de2259fe5754f5958903b42

Observation 6cd379de-9fe6-4d5c-9a60-6585048f0de1 · outbound

This paper cites Integrating human gaze into attention for egocentric activity recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Integrating human gaze into attention for egocentric activity recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.496469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.576732Z digest=sha256:7df03d4926907b3a190f132d5e8e67b2d1fc60f8a8733b172c13c815d3934bea

Observation ed6c5aa4-8cde-446b-82d6-450e3f548839 · outbound

This paper cites Learning affordance landscapes for interaction exploration in 3d environments.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning affordance landscapes for interaction exploration in 3d environments

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.482815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.581234Z digest=sha256:c6469bfa102006c12dce6f1472d5efa5543f70953015b84831b9e3da7846b877

Observation 2d62998d-ebd2-4b6e-bd18-f88fb293ecff · outbound

This paper cites Egoenv: Human- centric environment representations from egocentric video.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egoenv: Human- centric environment representations from egocentric video

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.469334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.585398Z digest=sha256:180663e57ccbe2ae4c74f158c6216e805cdf6168eeedac1bd2d2cd5b9a6d71df

Observation 2e359030-db96-472c-a8fb-ec4463a52fb7 · outbound

This paper cites Gpt-4v(ision) system card.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gpt-4v(ision) system card

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.589425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.589425Z digest=sha256:fce4d44e20cf80ac30180f2f275c3132c5cbd64af98408ec9f60724f1953f61b

Observation 938812b2-fc71-4fcc-9b59-d0733978175e · outbound

This paper cites Hello gpt-4o.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Hello gpt-4o

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.447088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.593620Z digest=sha256:fbcb7076b275d34ca7efac5ed790630eea5851bf17c0539aa6e3e74340362883

Observation 7b5445b0-eb61-4d24-b845-e537c137950a · outbound

This paper cites Actionvos: Actions as prompts for video object segmentation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Actionvos: Actions as prompts for video object segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.434344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.598767Z digest=sha256:36f4e54f14a1b8ac8f6b8d5e86125158630e9dcbec3ebd37b91c27cfc100a871

Observation 2a983fef-1a55-4d00-bcfc-90257592f88d · outbound

This paper cites Wear- able augmented reality system using gaze interaction.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Wear- able augmented reality system using gaze interaction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.421551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.602833Z digest=sha256:d6fa2a1dd76f5a96b5b787e3db6d1d3f6838d88d8ed16cc4e2b0c1bbbaa5efc4

Observation 4afb15f6-053e-426f-b60c-2f45d49edb71 · outbound

This paper cites Deep learning-based smart task assistance in wearable augmented reality.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Deep learning-based smart task assistance in wearable augmented reality

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.408485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.606907Z digest=sha256:3782c4b92f6cdcad1bc7910114a1cdb8a43190767081db447ac1e23d9f1f0851

Observation d4308c20-0da5-4c3f-84bc-367fcd7924b8 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.610984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.610984Z digest=sha256:17012d63bdff561a923d14aa995852fbd248b5e83b09752be0f4eb3d5bc3ba74

Observation 8420b244-53af-4be1-a537-30ec777c7384 · outbound

This paper cites An Outlook into the Future of Egocentric Vision.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model An Outlook into the Future of Egocentric Vision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.615737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.615737Z digest=sha256:86a4a3782c2c0733a6ad46cb1c34bb1a423cfba1385e1532d256bef0d59f63ee

Observation e3853078-fcfb-40fc-8cc1-f4888ae96575 · outbound

This paper cites Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.620320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.620320Z digest=sha256:edec57a4ca6605e761cfe7dc78f4dc68d61e427530e78cc5f3ee77157b63a7a2

Observation 57f3df04-9bd8-4ea8-b715-bfa70ff40e5d · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.395284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.624620Z digest=sha256:8dc99cd7bb8da3944e154afaee3d5810e95f896f2c125407a927f9842d0d8693

Observation 8ac2dbfe-8937-49b8-abff-77075bfa34ab · outbound

This paper cites Streaming long video understanding with large language models, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Streaming long video understanding with large language models, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.382002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.628691Z digest=sha256:8722fa445f62d7acb8f73933601892a0da9cfc9b86800dc8ba32934487768c0a

Observation 5b4a8357-e9e3-4b64-9660-1b69d82fe91a · outbound

This paper cites Sener, D.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Sener, D

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.368150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.633110Z digest=sha256:f50e66bdb2d3a5dc8359600ecd622508db91b9f995eebd24796b4792cbce3c7d

Observation 3d42fe93-c5a6-4946-b197-a4ad4620c5b4 · outbound

This paper cites Understanding human hands in contact at internet scale.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Understanding human hands in contact at internet scale

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.355037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.637413Z digest=sha256:75b51a4ccde8c694e83d187105b4c4f164eeaf08df0f800ff1b35f1c682a79d5

Observation 0e27b10e-9437-4c19-b883-ebf2237264d5 · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.641633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.641633Z digest=sha256:5eba6d4121719c5bbc3889104d4c1850c79694c6c04c4b45390deb1f0f637f66

Observation 77f1364a-4339-4342-9e57-e2ed89ed533c · outbound

This paper cites Towards diverse paragraph captioning for untrimmed videos.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Towards diverse paragraph captioning for untrimmed videos

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.342012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.645828Z digest=sha256:0277bb79f26f74553713ed8a18ee7e503fd22a5750a8531d9070e463b575c040

Observation 4e0cec2e-e935-4890-9a0a-3a188af385ec · outbound

This paper cites Ego4d goal-step: Toward hierarchical understanding of procedural activities.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego4d goal-step: Toward hierarchical understanding of procedural activities

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.329069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.650645Z digest=sha256:e4868268b974cd15c00a1f6348bd4b670a494e8f7b21308386d1dd05fcca6ce1

Observation fb3115e4-9c67-4df4-8146-ae45bb403574 · outbound

This paper cites H+ o: Unified egocentric recognition of 3d hand-object poses and interactions.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model H+ o: Unified egocentric recognition of 3d hand-object poses and interactions

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.316166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.654852Z digest=sha256:853fd86853b1f0d0d651dea3be63ee1521ecbf372346b3ee09ede623a8752d8a

Observation 5561dfa5-0ae8-417f-b133-8c3e6663bf47 · outbound

This paper cites Memory-and-anticipation transformer for online action understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Memory-and-anticipation transformer for online action understanding

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.302082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.658845Z digest=sha256:71110d4f275ee34ff484c9bebf08d00852d95ed48fafd0f325626c901c5c6951

Observation 6e38406f-8142-4d7d-aebe-d4fd380e192e · outbound

This paper cites Scene-aware ego- centric 3d human pose estimation.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Scene-aware ego- centric 3d human pose estimation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.288716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.663095Z digest=sha256:0d381d9b30ca3de9485e73647b5d422fd51fdf0448c57ec6e290063513abd2a4

Observation 48a2e583-2d61-44fd-a5fc-bc6e31e7b6bd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.668469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.668469Z digest=sha256:1cbebf9cee9025f494abbecf09ec0dfbc803b11c9df5b8d82453792c32011521

Observation fa3e5392-a256-4d5e-b7b3-539d03256832 · outbound

This paper cites an unresolved cited work.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:07:15.275336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.672801Z digest=sha256:00622e491893fa67838e66bc0c5eea5e1128a040d90d15d1e8d2c976d6ccb674

Observation c7ba0f18-9f9e-46ca-ba13-7896ae059b56 · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Lavie: High-quality video generation with cascaded latent diffusion models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.261999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.677185Z digest=sha256:d5398e3172a439d2464895d47d991e1ac6adfa55f7d15330493e8dfdc9702ca5

Observation 6f18991b-1848-4299-9076-7ae89dc2e674 · outbound

This paper cites In- ternvideo2: Scaling foundation models for multimodal video understanding.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model In- ternvideo2: Scaling foundation models for multimodal video understanding

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.248022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.681282Z digest=sha256:2968d922cf2900ec30edbdd272d3604ad3108448c4b83f605463d9266b3774e8

Observation 1c5e0a52-90b1-47f4-a653-29aa26d6f1ad · outbound

This paper cites Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.234972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.685720Z digest=sha256:5a82300a0083f5c2622d34b55f4a4337a9b46614964114d8a94a0dfc4fc6c96b

Observation 7f1babe2-6713-49e2-b355-bf6c3774d7d8 · outbound

This paper cites Gaze-enabled egocentric video summarization via constrained submodular maximiza- tion.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Gaze-enabled egocentric video summarization via constrained submodular maximiza- tion

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.220521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.689868Z digest=sha256:56faf32d50cfe3a51d296075d8d147eae8567bcc893cc0ae2cd7df3de589038f

Observation 60f48183-1dc4-4dc3-afd6-b972d1a6b482 · outbound

This paper cites Retrieval-augmented egocentric video captioning.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Retrieval-augmented egocentric video captioning

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.207403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.693756Z digest=sha256:3d417566139ed58b0153355c401756b7e30b18139f9066b0315eeca373972780

Observation 17f451e4-2bd5-44c9-b122-156c80b385b0 · outbound

This paper cites Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.194360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.697840Z digest=sha256:d4b032c9a10d020ec982d4cc777aa340d0f5e5c93d656553667a812dfc21b352

Observation 751d0146-664a-4d0c-8e48-52cf0d3f0c64 · outbound

This paper cites Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.180905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.702716Z digest=sha256:627eb898b48dce654a98ade5baf51f2a6a53fb976816546ca33b35317c906f59

Observation f6c3e92b-cf47-4dd9-9c3b-97c06c6c5c89 · outbound

This paper cites Deco: Decomposition and reconstruction for compositional temporal grounding via coarse-to-fine contrastive ranking.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Deco: Decomposition and reconstruction for compositional temporal grounding via coarse-to-fine contrastive ranking

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.166916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.707238Z digest=sha256:e30e90f7f9d27e9baeb8793bccf5fbb8b999a551442f24051812b28631596fc6

Observation 6ad907d8-7b73-404d-8b38-99d84344bd5a · outbound

This paper cites Basictad: an astounding rgb-only baseline for temporal action detection.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Basictad: an astounding rgb-only baseline for temporal action detection

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.153204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.711279Z digest=sha256:90d82ac78ba7ce2f1f21b10fda0d875ea992150e8b2ab8a6551dcdf18c3cc9b0

Observation c35436ff-01b6-4ef2-ac97-f289e8add419 · outbound

This paper cites Flash-vstream: Memory- based real-time understanding for long video streams, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Flash-vstream: Memory- based real-time understanding for long video streams, 2024

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.139248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.715438Z digest=sha256:cc9ad0d6c99696824454da33b0253387d65aad3852d0e0ea14e4f0a9622164e1

Observation b81ca1b8-ad46-4f13-82ec-552fdac0b5e5 · outbound

This paper cites Fine-grained egocentric hand-object segmentation: Dataset, model, and applications.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Fine-grained egocentric hand-object segmentation: Dataset, model, and applications

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.125890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.719771Z digest=sha256:12adb87075f36e1c07c8c910eb7ce92169f4c1ef178790d560313697b37f5608

Observation b52f9a23-813b-4b13-9d2d-4d2c24a6c2c6 · outbound

This paper cites Masked video and body-worn imu autoencoder for egocentric action recognition.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Masked video and body-worn imu autoencoder for egocentric action recognition

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.110023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.724292Z digest=sha256:fd978325d338e1681654a4494002da009c0b58072bcd28c30ee770c49655ade8

Observation a38889f8-ece3-444c-9287-1e18ec78ffa5 · outbound

This paper cites Internlm-xcomposer2.5-omnilive: A comprehensive multimodal system for long-term streaming video and audio interactions, 2024.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Internlm-xcomposer2.5-omnilive: A comprehensive multimodal system for long-term streaming video and audio interactions, 2024

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.095032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.728181Z digest=sha256:a112b3a58038dab04d49e83db62d8ec7cf86a4935e55117f41d1f61d2627dc56

Observation 7a6bd842-ceb9-4c81-9175-3e6b937781aa · outbound

This paper cites Ego- body: Human body shape and motion of interacting people from head-mounted devices.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Ego- body: Human body shape and motion of interacting people from head-mounted devices

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.080915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.732377Z digest=sha256:9cd1110ab6535e1df6245bd0b53794da8c92467f37c0ce84ba07e7f9daca2ee8

Observation a7309c89-6de1-4361-8b8b-09e82da33cdd · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Training a Large Video Model on a Single Machine in a Day

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.736293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.736293Z digest=sha256:ddbc871cfd17cb772a1412d6f482c8a6f0cab4a5b20213a602e306393b357a59

Observation cbd5f695-acc0-43e6-993a-83801994edfc · outbound

This paper cites Learning video representations from large language mod- els.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning video representations from large language mod- els

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.065933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.740359Z digest=sha256:e1e12f747aee8c7af9302252ef455cef1a7d0cb112ef84154318990f3c179e88

Observation 872ce457-d2de-463e-b4fc-1e5a2a4f3de8 · outbound

This paper cites Learning video representations from large language models.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Learning video representations from large language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.051092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.744403Z digest=sha256:be6794035466d37bca5da73e762aed90f6cd5436b801fa6850b1731fd128dc14

Observation 0fb7ad97-c9ee-46a7-afaa-4dbc2565cbac · outbound

This paper cites Mrsn: Multi-relation support network for video action detec- tion.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model Mrsn: Multi-relation support network for video action detec- tion

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:07:15.036512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:07:14.748483Z digest=sha256:6720bcfa6dc99e7a62b832a9d8ee25ecb66e4e713b678916f9f32b74ef5d271a

Pith citing papers

Observation 72f24855-0ec1-4f79-87f8-ae86f6cb6ffc · inbound

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance cites this paper.

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:37.248378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:37.248378Z digest=sha256:f1a332b0b41f1d78d1694e1e6eea7eb0b926fc2958164b4a55aad8248a7b7f32

Observation b8e01093-7009-49ed-a8c0-4e8fe09081b2 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.228004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.228004Z digest=sha256:b3bd697ef4d1880a8bfd3120c95bc8a9d83c8fd0a166f1847eb41be3d34a1f76

Observation c9468671-ac66-4673-a0f2-5a325648989c · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:58.390895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:58.390895Z digest=sha256:a325eb7ec477497eabb9b7f08d082c2841cbf84979e3f445d3015339a59ed1f7

Observation ceef77ff-df8a-4e65-a6c3-7beb71ea261f · inbound

Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant cites this paper.

Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:50:51.441624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T10:49:25.312165Z digest=sha256:29d8d024f458f97fc5de222691812167ca8f7fc730bef0a0f10c04600baf8b66

Observation cbd04bbf-28cc-484d-875c-2bd4d4a42efa · inbound

Human-Inspired Context-Selective Multimodal Memory for Social Robots cites this paper.

Human-Inspired Context-Selective Multimodal Memory for Social Robots Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:05.820629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:37:56.427507Z digest=sha256:86c1bc9e8846ce2a01306814d56a7af12903e3c8305746d712dbf0f097acb194

Observation 38d7a5d8-2702-45b7-a673-c5a029976b33 · inbound

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding cites this paper.

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.307729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T00:26:33.306128Z digest=sha256:cd0e56bc70c1d3f3009a6afe75f68faad941d84126a61b83ec43d80a86887fc7

Observation 8a64f381-c73e-49db-903b-8d37a2beb353 · inbound

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos cites this paper.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T05:01:06.200663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:01:06.200663Z digest=sha256:300f7e1549f95d46682c3502e41630a0dc1d9c3a8218fa17759d5280e5e9edfb