Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:01:12.829064Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2412.00832.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:01:12.829064Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:01.366536Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:18:57.044501Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f27a5cf4-8386-4842-b18b-319ef0b3c0ea · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cda93b88-d0ba-4f75-8fc4-b88416fabf8d · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models The (r) evolution of multi- modal large language models: A survey
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877ca8af-dd48-45b6-a512-3ebb13ac7acc · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4b315fa-d1df-4ebc-a4fb-60c6b9f3e872 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3f724c32-951e-477f-9df4-0600bbbe027b · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2bc6fd32-242c-49f5-a715-bb0e4893d8a9 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Reproducible scal- ing laws for contrastive language-image learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 05333c3c-b55a-4dcb-91f5-368cdc6e3118 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based vision: A survey
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a8da49b-83ad-49ee-bb36-20e4c81e0fea · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Low-latency auto- motive vision with event cameras
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dddee9e6-87e8-4149-bb7e-7306ef5af691 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Eklt: Asynchronous photometric feature tracking using events and frames
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6c17c68a-1f27-4629-90f8-718b53456217 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Recurrent vision transformers for object detection with event cameras
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe9d2bdf-22a2-4961-a51e-a9d697ac4c09 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Dsec: A stereo event camera dataset for driving scenarios
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9220487-9589-44c5-9e3a-af4cfab38a60 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based Simultaneous Localization and Mapping: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87a4c97-c318-4cce-b567-56c9cff7cff2 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8696cb31-8981-4de7-8ad8-0e9c6db56ac8 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Real-time 3d reconstruction and 6-dof tracking with an event camera
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9a76f3d8-4de0-4523-a974-9f71ad70a25b · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models N-imagenet: Towards robust, fine-grained object recognition with event cameras
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 164a4524-4963-4a2c-ac7e-28b526a83323 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Sodformer: Streaming object detection with transformer using events and frames
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2d517bce-5fa1-43e7-9785-691d9a70c3d4 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a74984-41e2-4f56-931e-d1f364193f75 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8f8b143d-538b-4088-bcd4-f14b415b7773 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Vila: On pre-training for visual language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f159e167-0cad-49fc-9fc7-b59820cb659b · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Improved baselines with visual instruction tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de61ff03-66d6-4813-9971-792a65dd3047 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Visual instruction tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae5221da-37bc-420b-875f-228503fba171 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b9cb82-abac-45af-85a3-abd0e06f9dd0 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e8d8d2-9aad-4e1b-8346-d4062f9c4626 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1285f9-2f14-4f3a-9e42-ac08ff1bbb71 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Data-driven feature tracking for event cameras
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 36c26336-66e5-4419-a6ec-152428026790 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Esl: Event-based structured light
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 159208c6-faaf-4660-80bb-fd0348074fdb · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13b0614-776c-4379-9b42-f765ae47f85d · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968532f8-6901-4f09-a2a3-5980694820d2 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Emvs: Event-based multi-view stereo—3d 9 reconstruction with an event camera in real-time
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 53c8eda1-4a95-47ef-957c-7d272039650d · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Events-to-video: Bringing modern computer vision to event cameras
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6528a202-e7dc-44f5-9e4f-6df6a237dd1f · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf38fdca-0f41-4b83-94e2-aa4adf5b7bc6 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Aligning and prompting everything all at once for univer- sal visual perception
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c169dbb5-d2d0-4f4b-b4bc-83397472b7aa · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models BlinkTrack: Feature Tracking over 80 FPS via Events and Images
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4ef5e9f7-7c74-4902-9b55-dca0f2c8a767 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Flava: A foundational language and vision alignment model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7c054af0-cdea-4ef8-ad61-d2753fc46196 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Cloud-device collaborative learning for multimodal large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 53e4539d-0540-4ab5-8523-50f11d46a391 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a051b0d7-b2bb-4400-8e40-870824ed9ee4 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09726e2-4787-4789-964c-78d9265993ee · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models EventCLIP: Adapting CLIP for Event-based Object Recognition
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 177c1df0-cd5d-4dfd-ac9e-bfeacf6a6b63 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Leod: Label-efficient object detection for event cameras
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e795f0ce-43e0-4479-91b5-ba390758a38f · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models xgen-mm (blip-3): A family of open large multimodal models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce77d6d9-3793-4b57-a355-011819caa0ce · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models A Survey on Multimodal Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f35d5f7-59f1-43ca-8419-a13b448ebc7d · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Eventps: Real-time photometric stereo using an event camera
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8e3c9be-af2f-4cef-8bed-155f1dfaf6bb · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173d49c1-b66d-4209-a7f9-6e1e7aa1f175 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models MM-LLMs: Recent Advances in MultiModal Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b418c2d4-6d9e-4efc-829c-9033f9a779fa · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Spiking transform- ers for event-based single object tracking
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 93adb63c-5752-49c7-90b4-9f28c8c7bfb3 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fd68e8-126b-481c-884d-b0955e5130d9 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86f87bf-1754-44f4-86c0-e7b6f28dd19f · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models E- clip: Towards label-efficient event-based open-world under- standing by clip
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74cf5c2a-b2db-4702-b0a2-d23a41979662 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e43ee78-5e59-493c-99cb-f23f3c94085b · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517844bd-825e-4b7e-bd36-135d6f994670 · outbound
EventGPT: Event Stream Understanding with Multimodal Large Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e341bfd4-a294-4028-a135-048efe39d61c · inbound
Event-Priori-Based Vision-Language Model for Efficient Visual Understanding EventGPT: Event Stream Understanding with Multimodal Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee59036-540e-4c27-8a3f-e3af7a893453 · inbound
EventDrive: Event Cameras for Vision-Language Driving Intelligence EventGPT: Event Stream Understanding with Multimodal Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7ec6b07c-ae39-4d48-b2f3-a86ee3c2bfd6 · inbound
DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments EventGPT: Event Stream Understanding with Multimodal Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.