Pith. sign in

Paper Citation Record · LEDGER

VideoOrion: Tokenizing Object Dynamics in Videos

As of 19 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2411.16156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16156 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:34:39.675979Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:42.424641Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:33:51.653731Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 082cb338-6733-437d-8440-5807cad345cf · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VideoOrion: Tokenizing Object Dynamics in Videos Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.307583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.307583Z digest=sha256:ff32934046591b2add83638dc67193472f318ac057b49e86e0b1f82b3a1b5339

Observation af876d63-93b8-4127-b704-b78283d9d226 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

VideoOrion: Tokenizing Object Dynamics in Videos Spice: Semantic propositional image cap- tion evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.878283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.313027Z digest=sha256:aa0344a42595db07ac91b1f57b7189b73763417d74e411c0f5aad206bad550b4

Observation deaab55d-cf4c-4f53-a562-18c9852652f8 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

VideoOrion: Tokenizing Object Dynamics in Videos MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.318256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.318256Z digest=sha256:7bea50d4a691ef1574ca08dc3ee3b55e1e255400580afd14f5bd437d0ccd23e8

Observation bf5ccdd3-4d13-4689-9b36-8a0474465e1b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VideoOrion: Tokenizing Object Dynamics in Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.323403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.323403Z digest=sha256:5203be0725a32afede341e103efc27d71a1c5f6bc7add58910ab2f52801fb688

Observation cc23c52c-b1bb-4bbd-a989-988ccd54b296 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

VideoOrion: Tokenizing Object Dynamics in Videos Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.863841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.328902Z digest=sha256:f266291de05ba5a081efcbc62fa187d8818ca02f42e28d566a9ebb1e34015ed2

Observation 97bff7b1-14ee-44b5-91b8-a57ddaa2f10e · outbound

This paper cites O’Reilly Media, Inc.

VideoOrion: Tokenizing Object Dynamics in Videos O’Reilly Media, Inc

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.848409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.333670Z digest=sha256:3e9f23114d3fcaf433ff6e32d4ec9b4296a13ad1564f132bf7d08655d07c585e

Observation a1d00ac4-016f-4473-9f03-864bc1a8fed2 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

VideoOrion: Tokenizing Object Dynamics in Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.338126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.338126Z digest=sha256:5462760ae12f417c00eb2fd8615d1428790452a295bc3178df0159425aedf276

Observation cebda45f-0aa7-48a2-8008-19870e9c55de · outbound

This paper cites End-to- end object detection with transformers.

VideoOrion: Tokenizing Object Dynamics in Videos End-to- end object detection with transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.342772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.342772Z digest=sha256:758c0208aa90d33bd2f2d2bd2d5222d5557ae400f8ff770769a35bf1e919ca4f

Observation 932aee05-fa87-458e-9788-4c094850c421 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

VideoOrion: Tokenizing Object Dynamics in Videos MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.348006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.348006Z digest=sha256:87aafad27d44205bde7b9d5980e89a379df0906a2f2a3b654c419959f58b8d5b

Observation d2ae6eec-bd68-44f2-ba7a-00b1f27df4a0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoOrion: Tokenizing Object Dynamics in Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.353218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.353218Z digest=sha256:1672823c34dbd438f0d7bfb28d9cdf7661f932c886c9f5a80e8cea155f9d7459

Observation 0092740a-55b1-443c-adae-32a5323b2559 · outbound

This paper cites Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video.

VideoOrion: Tokenizing Object Dynamics in Videos Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:34:40.052973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.358365Z digest=sha256:d4f7b5cda54b1bafbda65f828a968731aa1608d5ad4bbd42c2a7d769ec601e22

Observation e861eb7f-5a85-464b-bee9-91f7f341c223 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

VideoOrion: Tokenizing Object Dynamics in Videos Masked-attention mask transformer for universal image segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.821627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.363405Z digest=sha256:df9eba0529410bfa4fb87b5bd4c5894bffe5d05480b90000ff58d76d81811e41

Observation b9669a69-5820-4882-bc4f-c4e15f5f35fd · outbound

This paper cites an unresolved cited work.

VideoOrion: Tokenizing Object Dynamics in Videos Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:34:40.806599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.367833Z digest=sha256:0d8e9af65e75c219cf55a0d8f49cfbe4c97133e90a3a4a284c6dad194598af1d

Observation 2b9c46bf-5595-42df-9718-aad066aba0a6 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoOrion: Tokenizing Object Dynamics in Videos VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.372277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.372277Z digest=sha256:760c4641e0a9da943a221c20a641e443a1e3ccb95af27adb5a96db8935ae5fb1

Observation a9c51a8f-5985-4fb2-9a44-85d4ddf0a131 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

VideoOrion: Tokenizing Object Dynamics in Videos Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.791637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.377036Z digest=sha256:00928f3ff0247ece444099c3fd65487f80d6e43cf941686bbdce276b507edc4b

Observation 7e1ec2af-f5dd-4f42-a043-69003f616162 · outbound

This paper cites The Llama 3 Herd of Models.

VideoOrion: Tokenizing Object Dynamics in Videos The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.381626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.381626Z digest=sha256:90921039cfd9532158e016d92989eb101c8ec8140e1791ac649a3d4aab489970

Observation 07554d6f-ad94-4742-b382-66b35f7d7893 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

VideoOrion: Tokenizing Object Dynamics in Videos Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.775737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.386610Z digest=sha256:221937046440af9c0bc3266a0198b2bf3fe3faf9c0cd10c006b64e0cfc9e2281

Observation ba229878-cdc1-4eb7-b44e-6e039950c3e3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoOrion: Tokenizing Object Dynamics in Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.391033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.391033Z digest=sha256:d07c4a97d6dcf50632ed4ba86a0dabb693116299f0116d09e07edee9a9430cc9

Observation 20f0309b-697f-4c1b-9de5-d049e595a65a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VideoOrion: Tokenizing Object Dynamics in Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.396197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.396197Z digest=sha256:7cd4c1918a5451bcf6f1fba3c4ae9eb3a086322cae6186e447c6b2818b59aaa5

Observation af730cfc-bc5f-4ac0-a6d2-ca3f8a0ada80 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

VideoOrion: Tokenizing Object Dynamics in Videos Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.750917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.400683Z digest=sha256:e037dfe15ab9e8ffe3d755f37ac20d022e5944c9187ee7d5e51ae6f88b1e564a

Observation d2fee7f5-8897-4d27-9be6-221d9f1a8e6a · outbound

This paper cites Mask r-cnn.

VideoOrion: Tokenizing Object Dynamics in Videos Mask r-cnn

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.735797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.405042Z digest=sha256:6956db6af7733a05de48a6056697d80250c4ffbe6c33c535d325bd3c1252f677

Observation a837cb10-ca46-4f9d-9b4f-003abde76505 · outbound

This paper cites Long short-term memory.

VideoOrion: Tokenizing Object Dynamics in Videos Long short-term memory

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.721460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.409562Z digest=sha256:57858c5f8d744d318d751291b66e8cc3064ad60378683f81d0366c6cd6c22b91

Observation 4048b125-cb13-4ab1-8779-250abed61461 · outbound

This paper cites Open-set image tagging with multi-grained text su- pervision.

VideoOrion: Tokenizing Object Dynamics in Videos Open-set image tagging with multi-grained text su- pervision

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.706550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.413918Z digest=sha256:764a928e294a1a89cdc5b9dea77bd1329715dffdb24a570e4627a333766fd828

Observation 27dda32e-8e56-41b3-9a0e-9e72dc2d10e2 · outbound

This paper cites Mistral 7B.

VideoOrion: Tokenizing Object Dynamics in Videos Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.418269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.418269Z digest=sha256:758029ace7f33251d2499459ead3aa26271e84ac62f9c30b790d76e0e3ad12d0

Observation 943b5a7f-97fb-4bef-afc4-af9386306b3b · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

VideoOrion: Tokenizing Object Dynamics in Videos Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.691466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.422813Z digest=sha256:088d64bd2316aa8f0168d3c823d5fb05c118e60d727f64a87d69bc0141d568e0

Observation a52bde1e-85df-45ec-a38c-8608fa117d4c · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

VideoOrion: Tokenizing Object Dynamics in Videos Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.675832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.427051Z digest=sha256:77ee35f9459f95c685da154c319fdc338b738289d399d4b44b803261e7754f8b

Observation 9b26ce3e-57e1-479e-b162-949fb9a0ecd2 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollár, and Ross B.

VideoOrion: Tokenizing Object Dynamics in Videos Berg, Wan-Yen Lo, Piotr Dollár, and Ross B

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.658394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.431611Z digest=sha256:d5e4ee84fe37e7ed90e994d14859c822a99f155e041401c5deaeff63059ed3c1

Observation 3f710ebf-d498-47c5-91e2-489872a8fd1e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoOrion: Tokenizing Object Dynamics in Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.436086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.436086Z digest=sha256:2e040be1b3dcf9e4933a365276d7a5807b23d33d596885f908ec0c6fe7a80384

Observation 07656d73-a75f-4797-a617-fe01cf21ae03 · outbound

This paper cites an unresolved cited work.

VideoOrion: Tokenizing Object Dynamics in Videos Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:34:40.643091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.440864Z digest=sha256:34f231771fd147774f3bd1d12204ad18ef3640350cfbffe5a4197effb77fe690

Observation e8afc645-a8f4-490b-94fe-c1f5f3c39382 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoOrion: Tokenizing Object Dynamics in Videos VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.445071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.445071Z digest=sha256:3138bbf980e15c37d3a81e2e79412da8584b8f58c0ca307607c215aa270fd4d0

Observation e33e0842-ee4d-4e7c-b7f6-e2896ce0751e · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

VideoOrion: Tokenizing Object Dynamics in Videos Unmasked teacher: Towards training-efficient video foundation models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.629027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.449828Z digest=sha256:684b9ff005c64a70ac87c14caa479ae069d6392bb91a8263992e242157c4443e

Observation e834733b-c25a-4f95-a7ef-0e177f43659c · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

VideoOrion: Tokenizing Object Dynamics in Videos Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.614415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.454234Z digest=sha256:98e5fb0a4fcb93f1357909de185515737bf119d9bbd15bd7e22150e76e9a1f3d

Observation 83738e63-e9d5-4194-95c3-e3931755e1bc · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

VideoOrion: Tokenizing Object Dynamics in Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.458746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.458746Z digest=sha256:184f032b7bad80c2e46403d1097b56683af12595edfe7d638ee2b6bb19262bbd

Observation 43f120a1-72eb-4099-9685-eef9b93fb36c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VideoOrion: Tokenizing Object Dynamics in Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.463434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.463434Z digest=sha256:01f828665b3cfe94153ebe8a8cf267a783fc5d4fc591639797029a654bf46d88

Observation 5d022d43-6248-497d-9e2b-0896fad80ec2 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

VideoOrion: Tokenizing Object Dynamics in Videos Rouge: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.599544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.468272Z digest=sha256:85abcd9fbfd4a81c31a46164d259ff341e9ea414dffdce640f8ccaf4cbd2faf5

Observation 6dd8a0b1-f04e-40ac-9649-6e70ebf69249 · outbound

This paper cites Visual instruction tuning.

VideoOrion: Tokenizing Object Dynamics in Videos Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.472935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.472935Z digest=sha256:6e5b1560cbf8b8cb1a21f8b7ed11cce0b510621e1ef1341fbc6d6a1bbc37b380

Observation 555b8592-3df8-477c-bbd9-9ef5a4f87786 · outbound

This paper cites Visual instruction tuning.

VideoOrion: Tokenizing Object Dynamics in Videos Visual instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.573691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.477329Z digest=sha256:d8e3a1d998af25fa23ee3bbf67a4ff8159396a658365a9febca3819a829e4836

Observation 5f942323-284f-4900-85b3-28d0515c65db · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

VideoOrion: Tokenizing Object Dynamics in Videos Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.481837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.481837Z digest=sha256:c20fc11f74d84e06f3af3eb19ad9b162750c0e26a46f512e9a14c2fc84ad19df

Observation 3c47b35c-78e1-4398-81a5-ad030fec65db · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

VideoOrion: Tokenizing Object Dynamics in Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.486398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.486398Z digest=sha256:45367d949e73ac8141370a9b045e902891675e85e6d6504bc7c767c12c7a4fe4

Observation 0c413464-eddd-4822-99fe-e403e55f749c · outbound

This paper cites Video-chatgpt: Towards detailed video un- derstanding via large vision and language models.

VideoOrion: Tokenizing Object Dynamics in Videos Video-chatgpt: Towards detailed video un- derstanding via large vision and language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.558291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.491672Z digest=sha256:5fc28ce7246b5bc10d060d806b58c8f87fb757b2c90c6c29e41b7ed09f817d0b

Observation f986b07a-4e33-4c13-8e73-80c41666a413 · outbound

This paper cites Trackflow: Multi-object tracking with normalizing flows.

VideoOrion: Tokenizing Object Dynamics in Videos Trackflow: Multi-object tracking with normalizing flows

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.542588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.496319Z digest=sha256:331e7f73f723da5dd1b9df12b05fb98e13086c80cc83d062528c23884d8f412d

Observation 20f76415-9d50-4e21-b82c-c1e899309713 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

VideoOrion: Tokenizing Object Dynamics in Videos Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.527413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.500910Z digest=sha256:c16d880bf6d9724b864469719a41bbdf72298fa6e74c441f4f662669f1869684

Observation 7f15f85d-ee84-45c8-a936-81a6ae8bf117 · outbound

This paper cites Trackformer: Multi-object track- ing with transformers.

VideoOrion: Tokenizing Object Dynamics in Videos Trackformer: Multi-object track- ing with transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.512988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.506071Z digest=sha256:d5f2dd1ee0783a38de6153b7a0785836cdd195f62cdfbbd8c9e5c8b5d284d38a

Observation 630eea15-fd4f-4e25-8309-fb0106e9717e · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

VideoOrion: Tokenizing Object Dynamics in Videos OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.510390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.510390Z digest=sha256:dd1b0931853a7fa4101e76f5fef29d50bbbb09b58bff686b209b61ab112d65ed

Observation a4cc5ea5-f9a6-4cab-958c-95b4f237fa6a · outbound

This paper cites GPT-4 Technical Report.

VideoOrion: Tokenizing Object Dynamics in Videos GPT-4 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.515137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.515137Z digest=sha256:b89298111ca14c0285b489c30dbf7cf11024d68ad63fee8b1253b74b8c5948d6

Observation 63a552a8-a70b-4a0b-82c2-16eb36c79954 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VideoOrion: Tokenizing Object Dynamics in Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.519690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.519690Z digest=sha256:70f7906e7311035f747ac865a95f40209a2db9c98aae65e812c5809a10e7012c

Observation a1011cc5-44bb-4264-86e3-48187b3cb4f8 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

VideoOrion: Tokenizing Object Dynamics in Videos Bleu: a method for automatic evaluation of machine translation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.523956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.523956Z digest=sha256:8c6ac0cdbbcbc6b6078de1a9445096981f17d42de0dc3f4be6e0ec9e7c7a7006

Observation 22a51f1b-ccec-4d26-b9ba-17290cd0476e · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

VideoOrion: Tokenizing Object Dynamics in Videos Per- ception test: A diagnostic benchmark for multimodal video models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.487829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.528222Z digest=sha256:8da8ade296e8b18d07bce31707708e57bb25907760b4cd2ce3cd83323d07dad9

Observation c6329891-fbc4-4738-94cb-4d3a726e4484 · outbound

This paper cites Learning video object segmentation from static images.

VideoOrion: Tokenizing Object Dynamics in Videos Learning video object segmentation from static images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.473007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.532344Z digest=sha256:99ceb2aa77ac6c393d7b9fc0d1e407eec0ece47470e2cb9fc87481bfa44b77b9

Observation 8d531abf-a8fa-4427-83f0-d229a48a7a6f · outbound

This paper cites Towards generalizable multi-object tracking.

VideoOrion: Tokenizing Object Dynamics in Videos Towards generalizable multi-object tracking

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.458607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.536807Z digest=sha256:fe1c8cb4ac7373325baa6b595dded4a51656c442f060e788b84a25f4648a618f

Observation ef256bd6-26c1-47a5-b89f-56f0594d45ed · outbound

This paper cites Artemis: Towards Referential Understanding in Complex Videos.

VideoOrion: Tokenizing Object Dynamics in Videos Artemis: Towards Referential Understanding in Complex Videos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.541560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.541560Z digest=sha256:975acdb67c44b75aa72726aae68e49ae24ffcf01defccc8fdf24d45329e95791

Observation 72ff083d-d8e9-4881-a019-acf7287dc668 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VideoOrion: Tokenizing Object Dynamics in Videos Learning transferable visual models from natural language supervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.444157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.546238Z digest=sha256:6e94a68b5513149654d6abff77b92377a459a5d7ad8c63cc44ef0bc2cfd67289

Observation b691f4ac-b038-4075-85d5-023e827df575 · outbound

This paper cites Girshick, Kaiming He, and Piotr Dollár.

VideoOrion: Tokenizing Object Dynamics in Videos Girshick, Kaiming He, and Piotr Dollár

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.429322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.550586Z digest=sha256:e25efa17a189f2fe6d31dc17dad7a1a2bf8900257720839a6c9b2bf811a68cf7

Observation 41085d68-300e-4ec1-8543-0d927b5a5a7b · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VideoOrion: Tokenizing Object Dynamics in Videos SAM 2: Segment Anything in Images and Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.554986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.554986Z digest=sha256:e9252bcec23a87e696cd45afe2d77b041f1fa60ec8bf00f9f055b1132fb2e5ed

Observation 75c5a13f-1a39-468d-8cd4-8e9ba56ed472 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

VideoOrion: Tokenizing Object Dynamics in Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.414521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.559784Z digest=sha256:7a8d6af58cdd436b81deb16a64745f2d7ea4853056ba49a2de4eb8e374c99049

Observation ddc4fce0-dfad-44e3-8ba3-5032e811122a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

VideoOrion: Tokenizing Object Dynamics in Videos Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.399619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.564347Z digest=sha256:80d7835c67c74e94136896b4206919d8a87160b3d83f714016b2275ae9d5b04a

Observation c6ac71a1-b325-4977-9043-49cd8aaacf96 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

VideoOrion: Tokenizing Object Dynamics in Videos Audio-Visual LLM for Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.568890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.568890Z digest=sha256:cc948545dd2a38d61737794cfc2171bf340d5afc212e6e7da8b302ec3e57e182

Observation 9928d5cb-9c40-414b-a00d-4b1ceb7893ec · outbound

This paper cites Human-centric spatio-temporal video grounding with visual transformers.

VideoOrion: Tokenizing Object Dynamics in Videos Human-centric spatio-temporal video grounding with visual transformers

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.383909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.573964Z digest=sha256:471ea128c40ceed8ba4dee8ad5d1bfd1f9dabbf833e1f7fb2b07b0b8c3e0190c

Observation ef7c6d51-f861-41bb-8bfa-d15ffe57826f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VideoOrion: Tokenizing Object Dynamics in Videos Gemini: A Family of Highly Capable Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.578928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.578928Z digest=sha256:cf45b7b585ca1273d3b79266be980a3c7b3536a55f36b5aadd42db6105fe8010

Observation 35e36e0d-49fa-4b6a-af1e-137242892064 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

VideoOrion: Tokenizing Object Dynamics in Videos Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.369223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.583883Z digest=sha256:cdc69c7a6646779b8cb2a5f91c57afe07ff0e7812d578055f24d05d40665677a

Observation 7411bd11-f04d-447e-80ee-fc3989d1c959 · outbound

This paper cites Attention is all you need.

VideoOrion: Tokenizing Object Dynamics in Videos Attention is all you need

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.354099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.588317Z digest=sha256:643c7c7d215cbbb9c36bf385c3d9f3c4737e50187ff808cfeb511956e0526f49

Observation b3d6b349-5e81-4d83-964c-bae5123daf8c · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

VideoOrion: Tokenizing Object Dynamics in Videos Cider: Consensus-based image description evalua- tion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.337957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.592868Z digest=sha256:b96fccfbc81c6d776ce4bf100d532d1641599a7304112ad6c23ab30e2cd2bec6

Observation f3cb8cb6-8b87-407f-a60d-5feea461e2a3 · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm.

VideoOrion: Tokenizing Object Dynamics in Videos Elysium: Exploring object-level perception in videos via mllm

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.323665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.597762Z digest=sha256:af5e29b84c9d25382dbbe7bd146ba840083e06df3ff72b1ba6263aafa98e15e1

Observation 5fb45a14-08ee-4ea3-8ee4-d0de5f67d7fd · outbound

This paper cites Fast online object tracking and segmentation: A unifying approach.

VideoOrion: Tokenizing Object Dynamics in Videos Fast online object tracking and segmentation: A unifying approach

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.308679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.602283Z digest=sha256:e276e4ee4e8a645ed1de9f8f3f6271e9007c905ec72aafa646130e7f18b5f17b

Observation 8e6574c4-d365-452b-b334-4bf7b58b55fe · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

VideoOrion: Tokenizing Object Dynamics in Videos Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.293812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.606753Z digest=sha256:56ab74daa6493c2c51f6a137710c4081a69542b21b130b601b2def6ac426e76a

Observation d2e85e2f-10f1-4f0e-b5ab-5d91d0d6dd13 · outbound

This paper cites Slot-VLM: SlowFast Slots for Video-Language Modeling.

VideoOrion: Tokenizing Object Dynamics in Videos Slot-VLM: SlowFast Slots for Video-Language Modeling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.611300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.611300Z digest=sha256:1811ea3b9736630e1646eed7421fec4af696ed2fe0b6e118f878feca89db9255

Observation 53b2039b-2828-4a06-a1d8-2a012de8fc73 · outbound

This paper cites Qwen2 Technical Report.

VideoOrion: Tokenizing Object Dynamics in Videos Qwen2 Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.616540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.616540Z digest=sha256:b013bd6fa0ec54b9e00947502c50db76256d63e401c851b7d117c9a6bd9f673d

Observation a5140dd1-eccf-4722-9257-f3a321c32961 · outbound

This paper cites Associating ob- jects with transformers for video object segmentation.

VideoOrion: Tokenizing Object Dynamics in Videos Associating ob- jects with transformers for video object segmentation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.278454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.621919Z digest=sha256:f84d682c40bcdb87dbf9b9405caa062b3ef970f2723f6f80e9901d6e3d129d7e

Observation db39b2ad-338f-496f-a2c0-7588327cd4a3 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection.

VideoOrion: Tokenizing Object Dynamics in Videos Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.626487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.626487Z digest=sha256:03db00407ff7147f372c7a0ec526a0360ac3d7dbb6d30b1cf88dbaaa4c343cf0

Observation 5a20671a-cbc3-4966-8533-f1dbd5445e5b · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

VideoOrion: Tokenizing Object Dynamics in Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.630943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.630943Z digest=sha256:45a8ec341f19433930332c559226f37ad4bbda57fd5bc1a7e49e2a0b86947edc

Observation fffe4e55-fa70-4c75-a895-f187e12bac3c · outbound

This paper cites Merlin: Empowering multimodal llms with foresight minds.

VideoOrion: Tokenizing Object Dynamics in Videos Merlin: Empowering multimodal llms with foresight minds

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.253003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.635666Z digest=sha256:4a90904023b4256b8edb704bbddd5fcadbd5af50c9c3acace9bcedf48f9a6196

Observation 809c6341-9e23-46dd-8906-80cca33510bc · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VideoOrion: Tokenizing Object Dynamics in Videos Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.238026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.640034Z digest=sha256:07c7ba2ff4d77ab2e72e6d708b6d312eec55f5a8c5af62676d8980ef3be68681

Observation 7d14f471-a06e-4225-b9e9-e82064e228f9 · outbound

This paper cites Sigmoid loss for language image pre-training.

VideoOrion: Tokenizing Object Dynamics in Videos Sigmoid loss for language image pre-training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.644411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.644411Z digest=sha256:a2ec67562178db272a48f59bbde512285ef1efe407aaba546d1ab480642f6ea3

Observation daa19f54-10a2-4e35-94fb-d26a5579b75c · outbound

This paper cites Ni, and Heung-Yeung Shum.

VideoOrion: Tokenizing Object Dynamics in Videos Ni, and Heung-Yeung Shum

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.211398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.648650Z digest=sha256:2e02b44d7d2cb546603bfca302536657407affeeeae27185aa2210516715bf1a

Observation 62c6b054-d9f2-42b2-81ef-e9de295d7600 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

VideoOrion: Tokenizing Object Dynamics in Videos Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.195553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.653519Z digest=sha256:016ea58d31b95ab1be714bf947cddefa364c3edcf2045eff80a2d081695b6104

Observation bdd9b020-43cb-46c0-b9b0-895f15b28c95 · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoOrion: Tokenizing Object Dynamics in Videos Long Context Transfer from Language to Vision

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.658034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.658034Z digest=sha256:9fe272412a2955c46699782276a4a8e351cc7d57e68e5fe930b6d701e67900ef

Observation e3186eee-38ad-4f2a-963b-5e476b1e2ee5 · outbound

This paper cites Bytetrack: Multi-object tracking by associating every detection box.

VideoOrion: Tokenizing Object Dynamics in Videos Bytetrack: Multi-object tracking by associating every detection box

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.178837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.662752Z digest=sha256:d94ffc42d12184296d04720b37f7074ed5cdae0af147dd4205b50d2bb5cd410f

Observation d127ec19-20b0-4bb5-b547-4e297e93cb3f · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

VideoOrion: Tokenizing Object Dynamics in Videos Llava- next: A strong zero-shot video understanding model, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.163011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.667082Z digest=sha256:66897b9a49df32609667a515650318324317dcc8c717b913e0c4be929d7f511a

Observation 0168b99f-74c3-484d-b4bc-03fefaf1fb3e · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

VideoOrion: Tokenizing Object Dynamics in Videos Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.147418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T13:34:39.671530Z digest=sha256:4d17c25fb342d9ecc01aabf1345f188de7fbb9d7eca5179063d4379e5df6e642

Observation c030980a-8153-4a00-9e49-f400e8b8c805 · outbound

This paper cites Tracking Anything in High Quality.

VideoOrion: Tokenizing Object Dynamics in Videos Tracking Anything in High Quality

Reference 80

Resolution
malformed identifier
no resolver link, observed 2026-08-12T13:34:39.675979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.675979Z digest=sha256:e660df016e2290203e1ec48dff8a703c023f2e592e388045d8b7223f99dfa83e

Pith citing papers

Observation 33f7feff-207a-4e92-9c5b-053e6c498d97 · inbound

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI cites this paper.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI VideoOrion: Tokenizing Object Dynamics in Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.424641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.424641Z digest=sha256:87148fae43aaf05e63aba34938eb2b70011667cbb18e4fd59c7319d2fd7097cb

Observation a2c3ae77-95fe-4fe4-a46a-9825521193aa · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding VideoOrion: Tokenizing Object Dynamics in Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.372359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.372359Z digest=sha256:ef2118627f5a1bd41b7190b7a34c6e6f8c5a2df394d6d8ee17a8d3c7ed3ca74a

Observation dc1b8b98-c779-4bba-ba14-eacaaf7e8583 · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos VideoOrion: Tokenizing Object Dynamics in Videos

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:51.694013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T15:33:42.245773Z digest=sha256:56bb24681f979945f355fd55b2d41356de8e173f96b5737ade8dc873a13c0c53