Pith. sign in

Paper Citation Record · LEDGER

EventGPT: Event Stream Understanding with Multimodal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2412.00832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00832 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:01:12.829064Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:01.366536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:57.044501Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f27a5cf4-8386-4842-b18b-319ef0b3c0ea · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.376332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.146077Z digest=sha256:cff1a3eea4e55f97c10d54759c936dc44c8e31862f74b9ff33f4be7088c0bad6

Observation cda93b88-d0ba-4f75-8fc4-b88416fabf8d · outbound

This paper cites The (r) evolution of multi- modal large language models: A survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models The (r) evolution of multi- modal large language models: A survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.191004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.191004Z digest=sha256:94cd6abed4a4f69f24c38ca64a3ded814044225d979a58e02601fea506ad8428

Observation 877ca8af-dd48-45b6-a512-3ebb13ac7acc · outbound

This paper cites First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment.

EventGPT: Event Stream Understanding with Multimodal Large Language Models First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:01:14.030175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.213248Z digest=sha256:83197065b2dd4a20e2f9ba15895688f5cfb49597e384c5b5b27770b303e98435

Observation d4b315fa-d1df-4ebc-a4fb-60c6b9f3e872 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.345875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.221664Z digest=sha256:06fee35554b1d67758544e2060ba79b6279b254425be07a13347fd319084a498

Observation 3f724c32-951e-477f-9df4-0600bbbe027b · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.311023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.233969Z digest=sha256:6672be5c85262d37d3f79e40f9b624a0bc0e209f581e658eed398726eb853c1d

Observation 2bc6fd32-242c-49f5-a715-bb0e4893d8a9 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.275105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.246718Z digest=sha256:1766b682ea13cd01593d1b00a70ac9f55e2413f7fedbb8b54ee016cf1a1697be

Observation 05333c3c-b55a-4dcb-91f5-368cdc6e3118 · outbound

This paper cites Event-based vision: A survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based vision: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.224190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.264858Z digest=sha256:26d97d47e43da131dcce22434107fdd6350f6bd69a3a8ab80a483af52ca580c1

Observation 3a8da49b-83ad-49ee-bb36-20e4c81e0fea · outbound

This paper cites Low-latency auto- motive vision with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Low-latency auto- motive vision with event cameras

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.272493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.272493Z digest=sha256:dc115ad9f33d52e0cab1b37d6e632161263771c7abea881d5c30202b9002e62e

Observation dddee9e6-87e8-4149-bb7e-7306ef5af691 · outbound

This paper cites Eklt: Asynchronous photometric feature tracking using events and frames.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Eklt: Asynchronous photometric feature tracking using events and frames

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.147226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.299020Z digest=sha256:47520578f2ac153d2b3ef8af6c7283b28698c071d11ddb95ce6cbefba556afa4

Observation 6c17c68a-1f27-4629-90f8-718b53456217 · outbound

This paper cites Recurrent vision transformers for object detection with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Recurrent vision transformers for object detection with event cameras

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.094373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.307967Z digest=sha256:3c84fb0d8b7f8e738184374b7e171e33091f624f818a7ab1b6b5e379764659cc

Observation fe9d2bdf-22a2-4961-a51e-a9d697ac4c09 · outbound

This paper cites Dsec: A stereo event camera dataset for driving scenarios.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Dsec: A stereo event camera dataset for driving scenarios

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.058056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.319807Z digest=sha256:e1eb4bd1dc55904fc0cf8722e2b0d588b5042bc4a2061144d957a751070eb448

Observation d9220487-9589-44c5-9e3a-af4cfab38a60 · outbound

This paper cites Event-based Simultaneous Localization and Mapping: A Comprehensive Survey.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Event-based Simultaneous Localization and Mapping: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.325692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.325692Z digest=sha256:da2cd39b635ce5fdda68a66a7ae05e59b23d3dae73b1549dab31f6117429d2ec

Observation d87a4c97-c318-4cce-b567-56c9cff7cff2 · outbound

This paper cites Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.370094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.370094Z digest=sha256:870d45994c6dde7d4daabc6078e5687c09a8b23c942c9920b04d554b74ccdbef

Observation 8696cb31-8981-4de7-8ad8-0e9c6db56ac8 · outbound

This paper cites Real-time 3d reconstruction and 6-dof tracking with an event camera.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Real-time 3d reconstruction and 6-dof tracking with an event camera

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:15.013695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.380301Z digest=sha256:f6259aa227592b674064ea711027f48c1221d5a15d9ab4366832d1a4ecc733a1

Observation 9a76f3d8-4de0-4523-a974-9f71ad70a25b · outbound

This paper cites N-imagenet: Towards robust, fine-grained object recognition with event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models N-imagenet: Towards robust, fine-grained object recognition with event cameras

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.969940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.389515Z digest=sha256:2a444e8d60009ba638fa241cb80df6832bdbd706cad6fb0ed9b85861e14cec15

Observation 164a4524-4963-4a2c-ac7e-28b526a83323 · outbound

This paper cites Sodformer: Streaming object detection with transformer using events and frames.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Sodformer: Streaming object detection with transformer using events and frames

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.921551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.407322Z digest=sha256:5bbd97bd0d5ca1834a37c7910b085513314e775ca7362fc457a913e7cb5175ed

Observation 2d517bce-5fa1-43e7-9785-691d9a70c3d4 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.419221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.419221Z digest=sha256:bccb527451d6663421e63cc242ce86aedf3524d41290b20c6f2f66d494692a32

Observation 33a74984-41e2-4f56-931e-d1f364193f75 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.861678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.432463Z digest=sha256:0d923ec6d6f7f01a07e9c7b0b1d38993261eee2b1b1ae43e0660363b6e3d84d5

Observation 8f8b143d-538b-4088-bcd4-f14b415b7773 · outbound

This paper cites Vila: On pre-training for visual language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Vila: On pre-training for visual language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.815717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.451299Z digest=sha256:d2886e62b9571ab3df3793f0ac2d35a8d3ea303ac2891ab5fac8e6a31cf59ea4

Observation f159e167-0cad-49fc-9fc7-b59820cb659b · outbound

This paper cites Improved baselines with visual instruction tuning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.463039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.463039Z digest=sha256:6c58e1f3825c78b97c40a10a9e14439380eb6ced4a43aaf68290773ae028ee50

Observation de61ff03-66d6-4813-9971-792a65dd3047 · outbound

This paper cites Visual instruction tuning.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.748627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.471738Z digest=sha256:5b459335215ee26ee3aea5d8affb77834fa3ab8f78ceadfabd75fd52bc9249eb

Observation ae5221da-37bc-420b-875f-228503fba171 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.482716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.482716Z digest=sha256:5c02f0b4ef34a6e05faeec0d2b422174b6b523a6e0da5b5da2c35995d82436c7

Observation 79b9cb82-abac-45af-85a3-abd0e06f9dd0 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

EventGPT: Event Stream Understanding with Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.494304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.494304Z digest=sha256:fe747326951bb8dbecd32b6c9618691b4fab76f3b2db3830f6533c39068749ad

Observation 33e8d8d2-9aad-4e1b-8346-d4062f9c4626 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.504000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.504000Z digest=sha256:cc06f7933b5eace28068ef41335e7d766959c2dd6c0aee60261f077a2dd52e25

Observation 1e1285f9-2f14-4f3a-9e42-ac08ff1bbb71 · outbound

This paper cites Data-driven feature tracking for event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Data-driven feature tracking for event cameras

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.712020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.516781Z digest=sha256:4ef610037e7349eb0094504fe89c4c459d7ca5161cee46996472c7c72f0b0190

Observation 36c26336-66e5-4419-a6ec-152428026790 · outbound

This paper cites Esl: Event-based structured light.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Esl: Event-based structured light

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.660582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.526710Z digest=sha256:0fe74ec17b11d3f5f3ff374896a44a7e1c89c8cd8dd159af6234b26c40edf822

Observation 159208c6-faaf-4660-80bb-fd0348074fdb · outbound

This paper cites Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.536142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.536142Z digest=sha256:2e614fdd3dddab381ecad01e9b1a229bb4a924776a0da0f39fd5ee0dc78f54ad

Observation f13b0614-776c-4379-9b42-f765ae47f85d · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.546821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.546821Z digest=sha256:517af6aa44bec66612831b7dc0abd229a1fd26e66bb8fd93bea3131c44626a2f

Observation 968532f8-6901-4f09-a2a3-5980694820d2 · outbound

This paper cites Emvs: Event-based multi-view stereo—3d 9 reconstruction with an event camera in real-time.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Emvs: Event-based multi-view stereo—3d 9 reconstruction with an event camera in real-time

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.580424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.559931Z digest=sha256:c27ba0045a9d2bc05c423cecd0c14e3f0d813c8c9d04e538646e2a20ce8fcd06

Observation 53c8eda1-4a95-47ef-957c-7d272039650d · outbound

This paper cites Events-to-video: Bringing modern computer vision to event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Events-to-video: Bringing modern computer vision to event cameras

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.535641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.576078Z digest=sha256:3571eb809fe321bb059f9328d2bbcd9509d2924b14e89f7587ef59c37ccbb1b0

Observation 6528a202-e7dc-44f5-9e4f-6df6a237dd1f · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.585257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.585257Z digest=sha256:4f4c268e8aa7183104932115bfabd1a16e0b1313f990fcdebd4caa6c5a665c52

Observation bf38fdca-0f41-4b83-94e2-aa4adf5b7bc6 · outbound

This paper cites Aligning and prompting everything all at once for univer- sal visual perception.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Aligning and prompting everything all at once for univer- sal visual perception

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.504155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.592175Z digest=sha256:b1ebdfad304e58b0ae20073d960fabeebccf1ec9460eba1e19d6f67c1c38c5d1

Observation c169dbb5-d2d0-4f4b-b4bc-83397472b7aa · outbound

This paper cites BlinkTrack: Feature Tracking over 80 FPS via Events and Images.

EventGPT: Event Stream Understanding with Multimodal Large Language Models BlinkTrack: Feature Tracking over 80 FPS via Events and Images

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:01:13.614829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.601049Z digest=sha256:2af3acc894956704767d1b56256e27a9dca2f1760c60acabc9e5c1fabeda51fe

Observation 4ef5e9f7-7c74-4902-9b55-dca0f2c8a767 · outbound

This paper cites Flava: A foundational language and vision alignment model.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Flava: A foundational language and vision alignment model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.470618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.615369Z digest=sha256:9c1bf4ab0bc6e6d1fd628fa900d4f580f28eb2f9138f88b793308153801a4e89

Observation 7c054af0-cdea-4ef8-ad61-d2753fc46196 · outbound

This paper cites Cloud-device collaborative learning for multimodal large language models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Cloud-device collaborative learning for multimodal large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.410556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.629470Z digest=sha256:810e6130be4af7907042634b64e6cb36b4323aed09831dd7762f69f243a046ad

Observation 53e4539d-0540-4ab5-8523-50f11d46a391 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.644835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.644835Z digest=sha256:0e5787ecd40a7b10b06f992f183acd27fcd6580661b0c5e13a1edf6352cc8b53

Observation a051b0d7-b2bb-4400-8e40-870824ed9ee4 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

EventGPT: Event Stream Understanding with Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.652501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.652501Z digest=sha256:8a831e678b248ccb50cdec5dab6d6b424d06526ae95a8268b8c10ebb3d666f61

Observation f09726e2-4787-4789-964c-78d9265993ee · outbound

This paper cites EventCLIP: Adapting CLIP for Event-based Object Recognition.

EventGPT: Event Stream Understanding with Multimodal Large Language Models EventCLIP: Adapting CLIP for Event-based Object Recognition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.664201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.664201Z digest=sha256:822a90d3d9b93d0279278e9b75124d4399057802ff8bd7f0c32ff4d23299ea2b

Observation 177c1df0-cd5d-4dfd-ac9e-bfeacf6a6b63 · outbound

This paper cites Leod: Label-efficient object detection for event cameras.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Leod: Label-efficient object detection for event cameras

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.377574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.675218Z digest=sha256:6a2500d6a337296bd6996fec140e5c0be70242cf2508ba163ddef6825887cf60

Observation e795f0ce-43e0-4479-91b5-ba390758a38f · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models xgen-mm (blip-3): A family of open large multimodal models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.682307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.682307Z digest=sha256:fef6fb7cb21f294d86dd3077d8c73f2facfc8a056696fa1c732c025e79dd5e04

Observation ce77d6d9-3793-4b57-a355-011819caa0ce · outbound

This paper cites A Survey on Multimodal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models A Survey on Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.701293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.701293Z digest=sha256:8265dc7d14f24fd4c27a5507ea3012f666098d211e61986f2c1c94b3ae5ffd35

Observation 1f35d5f7-59f1-43ca-8419-a13b448ebc7d · outbound

This paper cites Eventps: Real-time photometric stereo using an event camera.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Eventps: Real-time photometric stereo using an event camera

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.326105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.711889Z digest=sha256:633b0703cf2dbd6c82fb7883bde2ed0593e699dd38b5eaa37a24c1695d429d20

Observation e8e3c9be-af2f-4cef-8bed-155f1dfaf6bb · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

EventGPT: Event Stream Understanding with Multimodal Large Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.725243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.725243Z digest=sha256:81e98af4ad87229d8bd4c1458d4278c86d66694668fd1983190a43a72667c1bc

Observation 173d49c1-b66d-4209-a7f9-6e1e7aa1f175 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.746090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.746090Z digest=sha256:639623b5bb0a42c83d612db1953d323ab6c11f987ce1ec8ebb7d26e4557615b7

Observation b418c2d4-6d9e-4efc-829c-9033f9a779fa · outbound

This paper cites Spiking transform- ers for event-based single object tracking.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Spiking transform- ers for event-based single object tracking

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.289514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.767968Z digest=sha256:c52bd18797db79cccfab1016597e0b1d9b9ae8ad36e27a008162f1227bd83307

Observation 93adb63c-5752-49c7-90b4-9f28c8c7bfb3 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

EventGPT: Event Stream Understanding with Multimodal Large Language Models OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.778547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.778547Z digest=sha256:9347c948648b0b74017c4fc7aef097ba81151a3a57475952d0a292ecfcf1995d

Observation e8fd68e8-126b-481c-884d-b0955e5130d9 · outbound

This paper cites Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Deep Learning for Event-based Vision: A Comprehensive Survey and Benchmarks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.790227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.790227Z digest=sha256:968283187da88a9ad30482b10d82ba9d5f846425e3fc8b754b12ebb2814bc44f

Observation b86f87bf-1754-44f4-86c0-e7b6f28dd19f · outbound

This paper cites E- clip: Towards label-efficient event-based open-world under- standing by clip.

EventGPT: Event Stream Understanding with Multimodal Large Language Models E- clip: Towards label-efficient event-based open-world under- standing by clip

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.250690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.799667Z digest=sha256:8be0af112dd425e9c2e3808228a4820abc5a991dbc32a230c8bea58be19c4190

Observation 74cf5c2a-b2db-4702-b0a2-d23a41979662 · outbound

This paper cites Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more.

EventGPT: Event Stream Understanding with Multimodal Large Language Models Ex- act: Language-guided conceptual reasoning and uncertainty estimation for event-based action recognition and more

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:01:14.224985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:01:12.811621Z digest=sha256:9ee20081a8b4da756a9a4565cc23ee7bb04e0682cfdd475c6be2b9eb2e4e54bb

Observation 1e43ee78-5e59-493c-99cb-f23f3c94085b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

EventGPT: Event Stream Understanding with Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.822819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.822819Z digest=sha256:366a5e13f6d2dc9147a2672486947b115f0e722900ce98d282bb5f922d9a204b

Observation 517844bd-825e-4b7e-bd36-135d6f994670 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

EventGPT: Event Stream Understanding with Multimodal Large Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.829064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.829064Z digest=sha256:deb0634996e71ca484288eb9f436217f233508e7a8ac21dabf5f93eb9258f51d

Pith citing papers

Observation e341bfd4-a294-4028-a135-048efe39d61c · inbound

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding cites this paper.

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:01.366536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:01.366536Z digest=sha256:8c1bb213d3ce0333e9ddff65457e32a672435d33258129a9976589125ffaf000

Observation cee59036-540e-4c27-8a3f-e3af7a893453 · inbound

EventDrive: Event Cameras for Vision-Language Driving Intelligence cites this paper.

EventDrive: Event Cameras for Vision-Language Driving Intelligence EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.048104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T01:26:43.752335Z digest=sha256:c5ae49c49cb943127d79130b3ef77d9b670d39d34461cb6561fa89492a7cdfda

Observation 7ec6b07c-ae39-4d48-b2f3-a86ee3c2bfd6 · inbound

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments cites this paper.

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments EventGPT: Event Stream Understanding with Multimodal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.246070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T05:35:32.213470Z digest=sha256:13531996583ef9adaee1a2682fc05bf2a4ded8522f1c0979b5b694555f238e66