Pith. sign in

Paper Citation Record · LEDGER

VideoOrion: Tokenizing Object Dynamics in Videos

As of 19 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2411.16156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16156 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:34:39.675979Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:42.424641Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:33:51.653731Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 082cb338-6733-437d-8440-5807cad345cf · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VideoOrion: Tokenizing Object Dynamics in Videos Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.307583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.307583Z digest=sha256:ff32934046591b2add83638dc67193472f318ac057b49e86e0b1f82b3a1b5339

Observation af876d63-93b8-4127-b704-b78283d9d226 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

VideoOrion: Tokenizing Object Dynamics in Videos Spice: Semantic propositional image cap- tion evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.878283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.313027Z digest=sha256:8530bb036352261b912bd8ad24a6ace20f1af150ea3f9fee1699052d671c54ff

Observation deaab55d-cf4c-4f53-a562-18c9852652f8 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

VideoOrion: Tokenizing Object Dynamics in Videos MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.318256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.318256Z digest=sha256:7bea50d4a691ef1574ca08dc3ee3b55e1e255400580afd14f5bd437d0ccd23e8

Observation bf5ccdd3-4d13-4689-9b36-8a0474465e1b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VideoOrion: Tokenizing Object Dynamics in Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.323403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.323403Z digest=sha256:5203be0725a32afede341e103efc27d71a1c5f6bc7add58910ab2f52801fb688

Observation cc23c52c-b1bb-4bbd-a989-988ccd54b296 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

VideoOrion: Tokenizing Object Dynamics in Videos Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.863841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.328902Z digest=sha256:3e902c60d431320fc2d9a63c5ea370c0d8ddaa26fba59d356e7ea91502c7f956

Observation 97bff7b1-14ee-44b5-91b8-a57ddaa2f10e · outbound

This paper cites O’Reilly Media, Inc.

VideoOrion: Tokenizing Object Dynamics in Videos O’Reilly Media, Inc

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.848409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.333670Z digest=sha256:61d89ae77d729ed30160e2db2ea35f30e7bafc8c3862284926cf5d31b22e32fc

Observation a1d00ac4-016f-4473-9f03-864bc1a8fed2 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

VideoOrion: Tokenizing Object Dynamics in Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.338126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.338126Z digest=sha256:5462760ae12f417c00eb2fd8615d1428790452a295bc3178df0159425aedf276

Observation cebda45f-0aa7-48a2-8008-19870e9c55de · outbound

This paper cites End-to- end object detection with transformers.

VideoOrion: Tokenizing Object Dynamics in Videos End-to- end object detection with transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.342772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.342772Z digest=sha256:758c0208aa90d33bd2f2d2bd2d5222d5557ae400f8ff770769a35bf1e919ca4f

Observation 932aee05-fa87-458e-9788-4c094850c421 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

VideoOrion: Tokenizing Object Dynamics in Videos MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.348006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.348006Z digest=sha256:87aafad27d44205bde7b9d5980e89a379df0906a2f2a3b654c419959f58b8d5b

Observation d2ae6eec-bd68-44f2-ba7a-00b1f27df4a0 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoOrion: Tokenizing Object Dynamics in Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.353218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.353218Z digest=sha256:1672823c34dbd438f0d7bfb28d9cdf7661f932c886c9f5a80e8cea155f9d7459

Observation 0092740a-55b1-443c-adae-32a5323b2559 · outbound

This paper cites Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video.

VideoOrion: Tokenizing Object Dynamics in Videos Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:34:40.052973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.358365Z digest=sha256:54fb5416858e5109799e99dab3282b14e8eeabc4f46c60eb50f308e9e4804d87

Observation e861eb7f-5a85-464b-bee9-91f7f341c223 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

VideoOrion: Tokenizing Object Dynamics in Videos Masked-attention mask transformer for universal image segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.821627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.363405Z digest=sha256:6c879679ca7cc9ad5268376c814c5582e0f63b99c5d1fe245b4c9f206aa3b0c5

Observation b9669a69-5820-4882-bc4f-c4e15f5f35fd · outbound

This paper cites an unresolved cited work.

VideoOrion: Tokenizing Object Dynamics in Videos Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:34:40.806599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.367833Z digest=sha256:4d77720f62f86f659b973eea72a35d19de0df1214a1339d39d6ee362f0cbe0fe

Observation 2b9c46bf-5595-42df-9718-aad066aba0a6 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoOrion: Tokenizing Object Dynamics in Videos VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.372277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.372277Z digest=sha256:760c4641e0a9da943a221c20a641e443a1e3ccb95af27adb5a96db8935ae5fb1

Observation a9c51a8f-5985-4fb2-9a44-85d4ddf0a131 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

VideoOrion: Tokenizing Object Dynamics in Videos Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.791637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.377036Z digest=sha256:55b202b1ef322f718dabfad04b4bafc9ccb4e33ec1e29051c1b4e890fc3da3c8

Observation 7e1ec2af-f5dd-4f42-a043-69003f616162 · outbound

This paper cites The Llama 3 Herd of Models.

VideoOrion: Tokenizing Object Dynamics in Videos The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.381626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.381626Z digest=sha256:90921039cfd9532158e016d92989eb101c8ec8140e1791ac649a3d4aab489970

Observation 07554d6f-ad94-4742-b382-66b35f7d7893 · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single ob- ject tracking.

VideoOrion: Tokenizing Object Dynamics in Videos Lasot: A high-quality benchmark for large-scale single ob- ject tracking

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.775737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.386610Z digest=sha256:bdec1c64679f8f6d00c9a548e0e0c792d6922d494ede03250208f9c1b580b5e2

Observation ba229878-cdc1-4eb7-b44e-6e039950c3e3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoOrion: Tokenizing Object Dynamics in Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.391033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.391033Z digest=sha256:d07c4a97d6dcf50632ed4ba86a0dabb693116299f0116d09e07edee9a9430cc9

Observation 20f0309b-697f-4c1b-9de5-d049e595a65a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VideoOrion: Tokenizing Object Dynamics in Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.396197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.396197Z digest=sha256:7cd4c1918a5451bcf6f1fba3c4ae9eb3a086322cae6186e447c6b2818b59aaa5

Observation af730cfc-bc5f-4ac0-a6d2-ca3f8a0ada80 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

VideoOrion: Tokenizing Object Dynamics in Videos Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.750917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.400683Z digest=sha256:6bc299fcd5dfc849782ec8182b605b3ee85bff51a6272f6463e427bf8cef6a5d

Observation d2fee7f5-8897-4d27-9be6-221d9f1a8e6a · outbound

This paper cites Mask r-cnn.

VideoOrion: Tokenizing Object Dynamics in Videos Mask r-cnn

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.735797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.405042Z digest=sha256:23929aadc751c1920728a050cf39be8a0a9403b7c4f126c1553f141114eb7e53

Observation a837cb10-ca46-4f9d-9b4f-003abde76505 · outbound

This paper cites Long short-term memory.

VideoOrion: Tokenizing Object Dynamics in Videos Long short-term memory

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.721460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.409562Z digest=sha256:f94baa140f097c88ca9855b482440bbe3aa4191f9e3e6969124febbe667221ab

Observation 4048b125-cb13-4ab1-8779-250abed61461 · outbound

This paper cites Open-set image tagging with multi-grained text su- pervision.

VideoOrion: Tokenizing Object Dynamics in Videos Open-set image tagging with multi-grained text su- pervision

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.706550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.413918Z digest=sha256:429d0801ef7590f55039c57834c3e26880ee504033b049a549749ccaac18430b

Observation 27dda32e-8e56-41b3-9a0e-9e72dc2d10e2 · outbound

This paper cites Mistral 7B.

VideoOrion: Tokenizing Object Dynamics in Videos Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.418269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.418269Z digest=sha256:758029ace7f33251d2499459ead3aa26271e84ac62f9c30b790d76e0e3ad12d0

Observation 943b5a7f-97fb-4bef-afc4-af9386306b3b · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

VideoOrion: Tokenizing Object Dynamics in Videos Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.691466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.422813Z digest=sha256:ee3962993b734b487c2595fb1f9890b3d9b87176cc48665ae118438fb159e20c

Observation a52bde1e-85df-45ec-a38c-8608fa117d4c · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

VideoOrion: Tokenizing Object Dynamics in Videos Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.675832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.427051Z digest=sha256:c7598746ee303d67425a1e891156e98e2ac8af55e1a72c27573a22dabc5ec1ce

Observation 9b26ce3e-57e1-479e-b162-949fb9a0ecd2 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollár, and Ross B.

VideoOrion: Tokenizing Object Dynamics in Videos Berg, Wan-Yen Lo, Piotr Dollár, and Ross B

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.658394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.431611Z digest=sha256:8994dda1753c1b0f761b1d8c5797a1e5d44154b3c17cf5e020eed4a8d29fba89

Observation 3f710ebf-d498-47c5-91e2-489872a8fd1e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoOrion: Tokenizing Object Dynamics in Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.436086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.436086Z digest=sha256:2e040be1b3dcf9e4933a365276d7a5807b23d33d596885f908ec0c6fe7a80384

Observation 07656d73-a75f-4797-a617-fe01cf21ae03 · outbound

This paper cites an unresolved cited work.

VideoOrion: Tokenizing Object Dynamics in Videos Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:34:40.643091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.440864Z digest=sha256:5ff27044c704c7d99d2781c1d4064b0974c0a428227e8098b79d576a19ddc7cc

Observation e8afc645-a8f4-490b-94fe-c1f5f3c39382 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoOrion: Tokenizing Object Dynamics in Videos VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.445071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.445071Z digest=sha256:3138bbf980e15c37d3a81e2e79412da8584b8f58c0ca307607c215aa270fd4d0

Observation e33e0842-ee4d-4e7c-b7f6-e2896ce0751e · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

VideoOrion: Tokenizing Object Dynamics in Videos Unmasked teacher: Towards training-efficient video foundation models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.629027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.449828Z digest=sha256:c05690ec6047843aef5b2fd17a8362b167158d8e6745ff2690f636d8fe75b928

Observation e834733b-c25a-4f95-a7ef-0e177f43659c · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

VideoOrion: Tokenizing Object Dynamics in Videos Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.614415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.454234Z digest=sha256:6da9d4ab6bafa7e6473021a4b42f60ebd6655919f8c02c3feedb3c6c0845b84a

Observation 83738e63-e9d5-4194-95c3-e3931755e1bc · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

VideoOrion: Tokenizing Object Dynamics in Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.458746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.458746Z digest=sha256:184f032b7bad80c2e46403d1097b56683af12595edfe7d638ee2b6bb19262bbd

Observation 43f120a1-72eb-4099-9685-eef9b93fb36c · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VideoOrion: Tokenizing Object Dynamics in Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.463434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.463434Z digest=sha256:01f828665b3cfe94153ebe8a8cf267a783fc5d4fc591639797029a654bf46d88

Observation 5d022d43-6248-497d-9e2b-0896fad80ec2 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

VideoOrion: Tokenizing Object Dynamics in Videos Rouge: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.599544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.468272Z digest=sha256:4d40c9093d6583fb8ffc8e1aa8858aa4c431f5840f3330a3a40726d56278e515

Observation 6dd8a0b1-f04e-40ac-9649-6e70ebf69249 · outbound

This paper cites Visual instruction tuning.

VideoOrion: Tokenizing Object Dynamics in Videos Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.472935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.472935Z digest=sha256:6e5b1560cbf8b8cb1a21f8b7ed11cce0b510621e1ef1341fbc6d6a1bbc37b380

Observation 555b8592-3df8-477c-bbd9-9ef5a4f87786 · outbound

This paper cites Visual instruction tuning.

VideoOrion: Tokenizing Object Dynamics in Videos Visual instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.573691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.477329Z digest=sha256:c4ce49e3e23e787349ca663ef8899d94d0fa68988075c16ab69a2385d9f15e3f

Observation 5f942323-284f-4900-85b3-28d0515c65db · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

VideoOrion: Tokenizing Object Dynamics in Videos Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.481837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.481837Z digest=sha256:c20fc11f74d84e06f3af3eb19ad9b162750c0e26a46f512e9a14c2fc84ad19df

Observation 3c47b35c-78e1-4398-81a5-ad030fec65db · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

VideoOrion: Tokenizing Object Dynamics in Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.486398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.486398Z digest=sha256:45367d949e73ac8141370a9b045e902891675e85e6d6504bc7c767c12c7a4fe4

Observation 0c413464-eddd-4822-99fe-e403e55f749c · outbound

This paper cites Video-chatgpt: Towards detailed video un- derstanding via large vision and language models.

VideoOrion: Tokenizing Object Dynamics in Videos Video-chatgpt: Towards detailed video un- derstanding via large vision and language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.558291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.491672Z digest=sha256:aa7794eb7b6e1fe49de51cbd2db03e440c5d2df4ccf08228a17a7aaa09f1e185

Observation f986b07a-4e33-4c13-8e73-80c41666a413 · outbound

This paper cites Trackflow: Multi-object tracking with normalizing flows.

VideoOrion: Tokenizing Object Dynamics in Videos Trackflow: Multi-object tracking with normalizing flows

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.542588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.496319Z digest=sha256:2184d35fb3407152e82793fba5283bed728739ff65abdf196ed462cfac85e896

Observation 20f76415-9d50-4e21-b82c-c1e899309713 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

VideoOrion: Tokenizing Object Dynamics in Videos Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.527413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.500910Z digest=sha256:1cabd6c6397128aeee553bcab7dffca2d1ebd96014d5f3c2cd4eee2d7819131b

Observation 7f15f85d-ee84-45c8-a936-81a6ae8bf117 · outbound

This paper cites Trackformer: Multi-object track- ing with transformers.

VideoOrion: Tokenizing Object Dynamics in Videos Trackformer: Multi-object track- ing with transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.512988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.506071Z digest=sha256:be5058be724b9d39c77a3d31cae819f29e327b52a2946c0673159378bdbc5e80

Observation 630eea15-fd4f-4e25-8309-fb0106e9717e · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

VideoOrion: Tokenizing Object Dynamics in Videos OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.510390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.510390Z digest=sha256:dd1b0931853a7fa4101e76f5fef29d50bbbb09b58bff686b209b61ab112d65ed

Observation a4cc5ea5-f9a6-4cab-958c-95b4f237fa6a · outbound

This paper cites GPT-4 Technical Report.

VideoOrion: Tokenizing Object Dynamics in Videos GPT-4 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.515137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.515137Z digest=sha256:b89298111ca14c0285b489c30dbf7cf11024d68ad63fee8b1253b74b8c5948d6

Observation 63a552a8-a70b-4a0b-82c2-16eb36c79954 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VideoOrion: Tokenizing Object Dynamics in Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.519690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.519690Z digest=sha256:70f7906e7311035f747ac865a95f40209a2db9c98aae65e812c5809a10e7012c

Observation a1011cc5-44bb-4264-86e3-48187b3cb4f8 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

VideoOrion: Tokenizing Object Dynamics in Videos Bleu: a method for automatic evaluation of machine translation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.523956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.523956Z digest=sha256:8c6ac0cdbbcbc6b6078de1a9445096981f17d42de0dc3f4be6e0ec9e7c7a7006

Observation 22a51f1b-ccec-4d26-b9ba-17290cd0476e · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

VideoOrion: Tokenizing Object Dynamics in Videos Per- ception test: A diagnostic benchmark for multimodal video models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.487829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.528222Z digest=sha256:766a4358535631b162d70430b52a4d9353fc0fa2f6b6e88d791fea1afecaa4bb

Observation c6329891-fbc4-4738-94cb-4d3a726e4484 · outbound

This paper cites Learning video object segmentation from static images.

VideoOrion: Tokenizing Object Dynamics in Videos Learning video object segmentation from static images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.473007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.532344Z digest=sha256:e3144de0ce8d510099756b01dbf78e5476d51302c2168b797a893ef6552113c8

Observation 8d531abf-a8fa-4427-83f0-d229a48a7a6f · outbound

This paper cites Towards generalizable multi-object tracking.

VideoOrion: Tokenizing Object Dynamics in Videos Towards generalizable multi-object tracking

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.458607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.536807Z digest=sha256:9751acae289a05c946ef4b8da841a47f25ebf4237c8bb50ac62cf32e3d1f796d

Observation ef256bd6-26c1-47a5-b89f-56f0594d45ed · outbound

This paper cites Artemis: Towards Referential Understanding in Complex Videos.

VideoOrion: Tokenizing Object Dynamics in Videos Artemis: Towards Referential Understanding in Complex Videos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.541560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.541560Z digest=sha256:975acdb67c44b75aa72726aae68e49ae24ffcf01defccc8fdf24d45329e95791

Observation 72ff083d-d8e9-4881-a019-acf7287dc668 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VideoOrion: Tokenizing Object Dynamics in Videos Learning transferable visual models from natural language supervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.444157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.546238Z digest=sha256:afb79bd8cb8f977e2fb369b7e3647f45d5ebde42d3812578debe325ccbd30be6

Observation b691f4ac-b038-4075-85d5-023e827df575 · outbound

This paper cites Girshick, Kaiming He, and Piotr Dollár.

VideoOrion: Tokenizing Object Dynamics in Videos Girshick, Kaiming He, and Piotr Dollár

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.429322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.550586Z digest=sha256:4a3216d5581922a6915c3d31f65edb5c3fbccb5ce512f206cf4147922694cde2

Observation 41085d68-300e-4ec1-8543-0d927b5a5a7b · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VideoOrion: Tokenizing Object Dynamics in Videos SAM 2: Segment Anything in Images and Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.554986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.554986Z digest=sha256:e9252bcec23a87e696cd45afe2d77b041f1fa60ec8bf00f9f055b1132fb2e5ed

Observation 75c5a13f-1a39-468d-8cd4-8e9ba56ed472 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

VideoOrion: Tokenizing Object Dynamics in Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.414521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.559784Z digest=sha256:4f4f4ac3510ac98d30cee6e90a74591fd262b0a9b87911e4dd2a1fe879c3ae4c

Observation ddc4fce0-dfad-44e3-8ba3-5032e811122a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

VideoOrion: Tokenizing Object Dynamics in Videos Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.399619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.564347Z digest=sha256:811497b59374ad90e83408ddc7d12739bf27c11da757986d961b7f452227c5ea

Observation c6ac71a1-b325-4977-9043-49cd8aaacf96 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

VideoOrion: Tokenizing Object Dynamics in Videos Audio-Visual LLM for Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.568890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.568890Z digest=sha256:cc948545dd2a38d61737794cfc2171bf340d5afc212e6e7da8b302ec3e57e182

Observation 9928d5cb-9c40-414b-a00d-4b1ceb7893ec · outbound

This paper cites Human-centric spatio-temporal video grounding with visual transformers.

VideoOrion: Tokenizing Object Dynamics in Videos Human-centric spatio-temporal video grounding with visual transformers

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.383909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.573964Z digest=sha256:71d161bd74c7e737b19d3c9fb584487fd8c5d527fe77e01af408908cbc93be99

Observation ef7c6d51-f861-41bb-8bfa-d15ffe57826f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VideoOrion: Tokenizing Object Dynamics in Videos Gemini: A Family of Highly Capable Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.578928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.578928Z digest=sha256:cf45b7b585ca1273d3b79266be980a3c7b3536a55f36b5aadd42db6105fe8010

Observation 35e36e0d-49fa-4b6a-af1e-137242892064 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

VideoOrion: Tokenizing Object Dynamics in Videos Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.369223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.583883Z digest=sha256:8c5a559ab61e8e0e057c336816d8ae090e9fb6a60de9157c2604df9f6fea54f9

Observation 7411bd11-f04d-447e-80ee-fc3989d1c959 · outbound

This paper cites Attention is all you need.

VideoOrion: Tokenizing Object Dynamics in Videos Attention is all you need

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.354099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.588317Z digest=sha256:a1d168b55dbc3b5b21dcb6a5d27788ef9f5f103f6e5921bd19432d047125b91f

Observation b3d6b349-5e81-4d83-964c-bae5123daf8c · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

VideoOrion: Tokenizing Object Dynamics in Videos Cider: Consensus-based image description evalua- tion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.337957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.592868Z digest=sha256:da85360e16c6b7ee5737a5cf4429f15f273f48339c97f64e7924d2e381d25939

Observation f3cb8cb6-8b87-407f-a60d-5feea461e2a3 · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm.

VideoOrion: Tokenizing Object Dynamics in Videos Elysium: Exploring object-level perception in videos via mllm

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.323665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.597762Z digest=sha256:89b3ec11ae6fc46a82164e9b46542f32dc7642f17b8aa39fc11e1b989dd0113a

Observation 5fb45a14-08ee-4ea3-8ee4-d0de5f67d7fd · outbound

This paper cites Fast online object tracking and segmentation: A unifying approach.

VideoOrion: Tokenizing Object Dynamics in Videos Fast online object tracking and segmentation: A unifying approach

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.308679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.602283Z digest=sha256:75ccae8369ffa06e81d6688ffcc015d3eb75dc378d6e1ae61f5760be0dc5c365

Observation 8e6574c4-d365-452b-b334-4bf7b58b55fe · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

VideoOrion: Tokenizing Object Dynamics in Videos Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.293812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.606753Z digest=sha256:2551475457d9a25206952195646a183aad55f39a2081c8640083874597ec46f0

Observation d2e85e2f-10f1-4f0e-b5ab-5d91d0d6dd13 · outbound

This paper cites Slot-VLM: SlowFast Slots for Video-Language Modeling.

VideoOrion: Tokenizing Object Dynamics in Videos Slot-VLM: SlowFast Slots for Video-Language Modeling

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.611300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.611300Z digest=sha256:1811ea3b9736630e1646eed7421fec4af696ed2fe0b6e118f878feca89db9255

Observation 53b2039b-2828-4a06-a1d8-2a012de8fc73 · outbound

This paper cites Qwen2 Technical Report.

VideoOrion: Tokenizing Object Dynamics in Videos Qwen2 Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.616540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.616540Z digest=sha256:b013bd6fa0ec54b9e00947502c50db76256d63e401c851b7d117c9a6bd9f673d

Observation a5140dd1-eccf-4722-9257-f3a321c32961 · outbound

This paper cites Associating ob- jects with transformers for video object segmentation.

VideoOrion: Tokenizing Object Dynamics in Videos Associating ob- jects with transformers for video object segmentation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.278454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.621919Z digest=sha256:b7d06950f597ee2a5afd9bf63139e7261bc4913364adb8363c47eb79271f9657

Observation db39b2ad-338f-496f-a2c0-7588327cd4a3 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection.

VideoOrion: Tokenizing Object Dynamics in Videos Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.626487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.626487Z digest=sha256:03db00407ff7147f372c7a0ec526a0360ac3d7dbb6d30b1cf88dbaaa4c343cf0

Observation 5a20671a-cbc3-4966-8533-f1dbd5445e5b · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

VideoOrion: Tokenizing Object Dynamics in Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.630943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.630943Z digest=sha256:45a8ec341f19433930332c559226f37ad4bbda57fd5bc1a7e49e2a0b86947edc

Observation fffe4e55-fa70-4c75-a895-f187e12bac3c · outbound

This paper cites Merlin: Empowering multimodal llms with foresight minds.

VideoOrion: Tokenizing Object Dynamics in Videos Merlin: Empowering multimodal llms with foresight minds

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.253003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.635666Z digest=sha256:a557d8e3e19ec4e05bbb41847daccee34999ff5f25bdb3b7892a9938aa37d522

Observation 809c6341-9e23-46dd-8906-80cca33510bc · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VideoOrion: Tokenizing Object Dynamics in Videos Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.238026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.640034Z digest=sha256:3b1c0c04dc84390cc17adcaab2197ecb43334798eeceef347ec55c3ed0e4e4ec

Observation 7d14f471-a06e-4225-b9e9-e82064e228f9 · outbound

This paper cites Sigmoid loss for language image pre-training.

VideoOrion: Tokenizing Object Dynamics in Videos Sigmoid loss for language image pre-training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.644411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.644411Z digest=sha256:a2ec67562178db272a48f59bbde512285ef1efe407aaba546d1ab480642f6ea3

Observation daa19f54-10a2-4e35-94fb-d26a5579b75c · outbound

This paper cites Ni, and Heung-Yeung Shum.

VideoOrion: Tokenizing Object Dynamics in Videos Ni, and Heung-Yeung Shum

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.211398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.648650Z digest=sha256:c4d77eabc75b306fc1456a29e6453a36e9f2ba1a50ee31c9da61f38a28b68593

Observation 62c6b054-d9f2-42b2-81ef-e9de295d7600 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

VideoOrion: Tokenizing Object Dynamics in Videos Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.195553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.653519Z digest=sha256:639a57fc87fa14d825411b9bc28b2e378f0c700b44caaf88534e34bd8172450f

Observation bdd9b020-43cb-46c0-b9b0-895f15b28c95 · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoOrion: Tokenizing Object Dynamics in Videos Long Context Transfer from Language to Vision

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:39.658034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.658034Z digest=sha256:9fe272412a2955c46699782276a4a8e351cc7d57e68e5fe930b6d701e67900ef

Observation e3186eee-38ad-4f2a-963b-5e476b1e2ee5 · outbound

This paper cites Bytetrack: Multi-object tracking by associating every detection box.

VideoOrion: Tokenizing Object Dynamics in Videos Bytetrack: Multi-object tracking by associating every detection box

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.178837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.662752Z digest=sha256:7248297c146e36c7f586919c7db7e959259fb7cd605d958b9edf3f8e27284e75

Observation d127ec19-20b0-4bb5-b547-4e297e93cb3f · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

VideoOrion: Tokenizing Object Dynamics in Videos Llava- next: A strong zero-shot video understanding model, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.163011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.667082Z digest=sha256:aa0fb066ec4bc492fd08dd2cee5456c840facb6d32e722964f28b64ff19e3784

Observation 0168b99f-74c3-484d-b4bc-03fefaf1fb3e · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

VideoOrion: Tokenizing Object Dynamics in Videos Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:34:40.147418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T13:34:39.671530Z digest=sha256:1067fb2dc616490e6adfb7fea1321e8063be69403fdbbcd23dfd791875aef1d8

Observation c030980a-8153-4a00-9e49-f400e8b8c805 · outbound

This paper cites Tracking Anything in High Quality.

VideoOrion: Tokenizing Object Dynamics in Videos Tracking Anything in High Quality

Reference 80

Resolution
malformed identifier
no resolver link, observed 2026-08-12T13:34:39.675979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:39.675979Z digest=sha256:e660df016e2290203e1ec48dff8a703c023f2e592e388045d8b7223f99dfa83e

Pith citing papers

Observation 33f7feff-207a-4e92-9c5b-053e6c498d97 · inbound

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI cites this paper.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI VideoOrion: Tokenizing Object Dynamics in Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.424641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.424641Z digest=sha256:87148fae43aaf05e63aba34938eb2b70011667cbb18e4fd59c7319d2fd7097cb

Observation a2c3ae77-95fe-4fe4-a46a-9825521193aa · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding VideoOrion: Tokenizing Object Dynamics in Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.372359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.372359Z digest=sha256:ef2118627f5a1bd41b7190b7a34c6e6f8c5a2df394d6d8ee17a8d3c7ed3ca74a

Observation dc1b8b98-c779-4bba-ba14-eacaaf7e8583 · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos VideoOrion: Tokenizing Object Dynamics in Videos

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:33:51.694013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:33:42.245773Z digest=sha256:7ccc09a53a3f9dfecc22dfb4fb0e4dfd85dc3e4dca4952613f302c23319f0a7c