Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:34:39.675979Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2411.16156.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:34:39.675979Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:42.424641Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T15:33:51.653731Z
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 082cb338-6733-437d-8440-5807cad345cf · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af876d63-93b8-4127-b704-b78283d9d226 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Spice: Semantic propositional image cap- tion evaluation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation deaab55d-cf4c-4f53-a562-18c9852652f8 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5ccdd3-4d13-4689-9b36-8a0474465e1b · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc23c52c-b1bb-4bbd-a989-988ccd54b296 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 97bff7b1-14ee-44b5-91b8-a57ddaa2f10e · outbound
VideoOrion: Tokenizing Object Dynamics in Videos O’Reilly Media, Inc
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a1d00ac4-016f-4473-9f03-864bc1a8fed2 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cebda45f-0aa7-48a2-8008-19870e9c55de · outbound
VideoOrion: Tokenizing Object Dynamics in Videos End-to- end object detection with transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932aee05-fa87-458e-9788-4c094850c421 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ae6eec-bd68-44f2-ba7a-00b1f27df4a0 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0092740a-55b1-443c-adae-32a5323b2559 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e861eb7f-5a85-464b-bee9-91f7f341c223 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Masked-attention mask transformer for universal image segmentation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b9669a69-5820-4882-bc4f-c4e15f5f35fd · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b9c46bf-5595-42df-9718-aad066aba0a6 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c51a8f-5985-4fb2-9a44-85d4ddf0a131 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Study on density peaks clustering based on k-nearest neighbors and principal component analysis
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7e1ec2af-f5dd-4f42-a043-69003f616162 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07554d6f-ad94-4742-b382-66b35f7d7893 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Lasot: A high-quality benchmark for large-scale single ob- ject tracking
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba229878-cdc1-4eb7-b44e-6e039950c3e3 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f0309b-697f-4c1b-9de5-d049e595a65a · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Ego4d: Around the world in 3,000 hours of egocentric video
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af730cfc-bc5f-4ac0-a6d2-ca3f8a0ada80 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2fee7f5-8897-4d27-9be6-221d9f1a8e6a · outbound
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a837cb10-ca46-4f9d-9b4f-003abde76505 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Long short-term memory
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4048b125-cb13-4ab1-8779-250abed61461 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Open-set image tagging with multi-grained text su- pervision
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 27dda32e-8e56-41b3-9a0e-9e72dc2d10e2 · outbound
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943b5a7f-97fb-4bef-afc4-af9386306b3b · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a52bde1e-85df-45ec-a38c-8608fa117d4c · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9b26ce3e-57e1-479e-b162-949fb9a0ecd2 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Berg, Wan-Yen Lo, Piotr Dollár, and Ross B
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f710ebf-d498-47c5-91e2-489872a8fd1e · outbound
VideoOrion: Tokenizing Object Dynamics in Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07656d73-a75f-4797-a617-fe01cf21ae03 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e8afc645-a8f4-490b-94fe-c1f5f3c39382 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos VideoChat: Chat-Centric Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33e0842-ee4d-4e7c-b7f6-e2896ce0751e · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Unmasked teacher: Towards training-efficient video foundation models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e834733b-c25a-4f95-a7ef-0e177f43659c · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 83738e63-e9d5-4194-95c3-e3931755e1bc · outbound
VideoOrion: Tokenizing Object Dynamics in Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43f120a1-72eb-4099-9685-eef9b93fb36c · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d022d43-6248-497d-9e2b-0896fad80ec2 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Rouge: A package for automatic evaluation of summaries
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6dd8a0b1-f04e-40ac-9649-6e70ebf69249 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Visual instruction tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555b8592-3df8-477c-bbd9-9ef5a4f87786 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Visual instruction tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5f942323-284f-4900-85b3-28d0515c65db · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c47b35c-78e1-4398-81a5-ad030fec65db · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c413464-eddd-4822-99fe-e403e55f749c · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Video-chatgpt: Towards detailed video un- derstanding via large vision and language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f986b07a-4e33-4c13-8e73-80c41666a413 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Trackflow: Multi-object tracking with normalizing flows
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 20f76415-9d50-4e21-b82c-c1e899309713 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7f15f85d-ee84-45c8-a936-81a6ae8bf117 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Trackformer: Multi-object track- ing with transformers
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 630eea15-fd4f-4e25-8309-fb0106e9717e · outbound
VideoOrion: Tokenizing Object Dynamics in Videos OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4cc5ea5-f9a6-4cab-958c-95b4f237fa6a · outbound
VideoOrion: Tokenizing Object Dynamics in Videos GPT-4 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a552a8-a70b-4a0b-82c2-16eb36c79954 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos DINOv2: Learning Robust Visual Features without Supervision
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1011cc5-44bb-4264-86e3-48187b3cb4f8 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Bleu: a method for automatic evaluation of machine translation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a51f1b-ccec-4d26-b9ba-17290cd0476e · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Per- ception test: A diagnostic benchmark for multimodal video models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c6329891-fbc4-4738-94cb-4d3a726e4484 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Learning video object segmentation from static images
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8d531abf-a8fa-4427-83f0-d229a48a7a6f · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Towards generalizable multi-object tracking
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ef256bd6-26c1-47a5-b89f-56f0594d45ed · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Artemis: Towards Referential Understanding in Complex Videos
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ff083d-d8e9-4881-a019-acf7287dc668 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Learning transferable visual models from natural language supervision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b691f4ac-b038-4075-85d5-023e827df575 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Girshick, Kaiming He, and Piotr Dollár
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 41085d68-300e-4ec1-8543-0d927b5a5a7b · outbound
VideoOrion: Tokenizing Object Dynamics in Videos SAM 2: Segment Anything in Images and Videos
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c5a13f-1a39-468d-8cd4-8e9ba56ed472 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ddc4fce0-dfad-44e3-8ba3-5032e811122a · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c6ac71a1-b325-4977-9043-49cd8aaacf96 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Audio-Visual LLM for Video Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9928d5cb-9c40-414b-a00d-4b1ceb7893ec · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Human-centric spatio-temporal video grounding with visual transformers
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ef7c6d51-f861-41bb-8bfa-d15ffe57826f · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Gemini: A Family of Highly Capable Multimodal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e36e0d-49fa-4b6a-af1e-137242892064 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7411bd11-f04d-447e-80ee-fc3989d1c959 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Attention is all you need
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b3d6b349-5e81-4d83-964c-bae5123daf8c · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Cider: Consensus-based image description evalua- tion
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3cb8cb6-8b87-407f-a60d-5feea461e2a3 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Elysium: Exploring object-level perception in videos via mllm
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5fb45a14-08ee-4ea3-8ee4-d0de5f67d7fd · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Fast online object tracking and segmentation: A unifying approach
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8e6574c4-d365-452b-b334-4bf7b58b55fe · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Internvid: A large-scale video-text dataset for multimodal understanding and generation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2e85e2f-10f1-4f0e-b5ab-5d91d0d6dd13 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Slot-VLM: SlowFast Slots for Video-Language Modeling
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b2039b-2828-4a06-a1d8-2a012de8fc73 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Qwen2 Technical Report
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5140dd1-eccf-4722-9257-f3a321c32961 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Associating ob- jects with transformers for video object segmentation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation db39b2ad-338f-496f-a2c0-7588327cd4a3 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a20671a-cbc3-4966-8533-f1dbd5445e5b · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fffe4e55-fa70-4c75-a895-f187e12bac3c · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Merlin: Empowering multimodal llms with foresight minds
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 809c6341-9e23-46dd-8906-80cca33510bc · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7d14f471-a06e-4225-b9e9-e82064e228f9 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Sigmoid loss for language image pre-training
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa19f54-10a2-4e35-94fb-d26a5579b75c · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Ni, and Heung-Yeung Shum
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 62c6b054-d9f2-42b2-81ef-e9de295d7600 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Video-llama: An instruction-tuned audio-visual language model for video un- derstanding
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bdd9b020-43cb-46c0-b9b0-895f15b28c95 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Long Context Transfer from Language to Vision
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3186eee-38ad-4f2a-963b-5e476b1e2ee5 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Bytetrack: Multi-object tracking by associating every detection box
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d127ec19-20b0-4bb5-b547-4e297e93cb3f · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Llava- next: A strong zero-shot video understanding model, 2024
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0168b99f-74c3-484d-b4bc-03fefaf1fb3e · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c030980a-8153-4a00-9e49-f400e8b8c805 · outbound
VideoOrion: Tokenizing Object Dynamics in Videos Tracking Anything in High Quality
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f7feff-207a-4e92-9c5b-053e6c498d97 · inbound
Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI VideoOrion: Tokenizing Object Dynamics in Videos
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c3ae77-95fe-4fe4-a46a-9825521193aa · inbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding VideoOrion: Tokenizing Object Dynamics in Videos
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1b8b98-c779-4bba-ba14-eacaaf7e8583 · inbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos VideoOrion: Tokenizing Object Dynamics in Videos
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.