Pith. sign in

Paper Citation Record · LEDGER

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data

As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2509.03501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03501 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:58:20.690107Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T09:20:32.920925Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:21:20.658722Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy44
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f32d129a-eba2-490c-9936-8d5edfae3144 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.439297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.439297Z digest=sha256:8b5c12cac07812cc365b3bff483655b187d1cef172f172117712fbc73e6343a8

Observation b3ccf81e-1b56-427c-9b08-7e20054f5d45 · outbound

This paper cites Vicas: A dataset for combining holistic and pixel-level video un- derstanding using captions with grounded segmentation.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vicas: A dataset for combining holistic and pixel-level video un- derstanding using captions with grounded segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.478984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.478984Z digest=sha256:cc52e9fbdacc1463ff5b6b97334cb6e32bd76e62ba329dc64dcd0ea64025f4ad

Observation 5393e087-47aa-431c-a067-46362548aeee · outbound

This paper cites Qwen2.5-VL Technical Report.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.511934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.511934Z digest=sha256:eb7b071998cc25513fd2e7bc397e6017c421b03a3157e613851fcd3adbae5db8

Observation 0acf6b67-da61-46a4-8d88-8b9b91104798 · outbound

This paper cites PySceneDetect: Video Scene Cut Detection.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PySceneDetect: Video Scene Cut Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.546854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.546854Z digest=sha256:655c9e465d5f86f3922e4337542a18287c052fa659d0c8712e618ebeaeff450f

Observation 0942c527-958a-49d8-a7a2-3e6248af82cb · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sharegpt4video: Improving video understanding and generation with better captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.594910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.594910Z digest=sha256:2798f3d946387b7afa0fbac015c3c0a3eaf7162dfcc1ac09fda768d4776d46c6

Observation 8a097da6-4d02-45a7-9027-91034c18167b · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.625667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.625667Z digest=sha256:81bc23d16c3b694e3b29826bf365d472ee2c57403564a4b063645403ed002e0c

Observation 5b0f2dbe-11db-4cda-bb90-e1f1021855e8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.702318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.702318Z digest=sha256:d1e60e6475a50f46988162e744869c464452473d3e1025a49dc1c2c1557c1c1d

Observation 2d03d88a-5f72-4a3d-bf4a-d778e64fa9fc · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.748440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.748440Z digest=sha256:5de87a277f587f99c2fb42b569d4921f62f924766af07d08068d21ce5db38450

Observation 048c6c3c-bdd1-4fc5-86b3-6ebce32f8450 · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.783270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.783270Z digest=sha256:5a17b546d03675aa1d890cb8a37e1b5de6829efe687dab54fcbf8a33ef48bad2

Observation bd4597e5-0732-488b-8a8c-dc1b7b90ad62 · outbound

This paper cites Unifying Specialized Visual Encoders for Video Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unifying Specialized Visual Encoders for Video Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:58:21.479672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:15.827629Z digest=sha256:0258f9378a8e838e667426dd062807dd37f07f5de0418dd130a3d8d27423b3da

Observation 39097890-3252-438c-801d-864c4e3436f5 · outbound

This paper cites Videorefer benchmark evaluation for general mllms.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videorefer benchmark evaluation for general mllms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.862688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.862688Z digest=sha256:1e8d292846714e2835f366b6a1e9e35cbec22d97d7872242c0f893c4ffb53b27

Observation ea89e910-4024-4328-91d2-40ecbdec82b8 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:36.378923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:15.886354Z digest=sha256:b1ffeafee36418ceb3976be88a85bc93cc9f7457e1ade87e81a52b871aded862

Observation 20a1c4c3-7e96-4250-857e-91bb626e0f3c · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:36.140120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:15.959077Z digest=sha256:197b7e8bb5f1adbdd88656b4a8e56baf9dab2197980eda3b5a42a6ae0040f3c8

Observation c4716635-c7ba-48e6-befe-dd37db008796 · outbound

This paper cites an unresolved cited work.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:58:35.911323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:15.964475Z digest=sha256:4dc3d301e30a0679e53f60864c6a1be9f969be690297af7a97d7542b7164b26d

Observation dc4f1d6e-e92b-4609-96e1-011c9c7cfa23 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:35.677128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:15.997832Z digest=sha256:7ccb8a096510b9e4e845a136d42c6120951514d34991c2154ab4266d00c72f39

Observation 239fdc4d-74dd-4aac-8b13-50ca7b2927b2 · outbound

This paper cites Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.021381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.021381Z digest=sha256:df6f63f0cf34de7080e434126e45b6b0e09a0ad78497f6714f63ae087436b3d4

Observation 6b30dc52-6e55-4609-94cd-8a3165b86594 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.100830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.100830Z digest=sha256:3fe87586b007f31545b0646956ccdc678d5f39106a653e3db8101e32c6a8c0ef

Observation 8d6e30f3-90d9-4439-b413-0fe2d25dcc04 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vtimellm: Empower llm to grasp video moments

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:35.448603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.153565Z digest=sha256:95202eabd3011c17b0ee6d3b30545221a246a54f0153a963f93faf208080d941

Observation ca09e8d2-9d88-44e3-bebf-71617c12196e · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:35.167898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.212317Z digest=sha256:33b6c1437ef9d023de339cf689a52d21cdc2706428d2901a84ecc5b979f51e12

Observation c7f1e1c6-471c-4efa-b6ad-ef27fa137087 · outbound

This paper cites Referring to any person, 2025.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Referring to any person, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:34.871035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.255854Z digest=sha256:7bdbf4dcbc52cb65937009a8c659f95a95e64858ee5bf3405f00fd46f568e228

Observation 8d707d43-44e9-44ae-bad0-275f8ebeb5f5 · outbound

This paper cites Miradata: A large-scale video dataset with long durations and structured captions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Miradata: A large-scale video dataset with long durations and structured captions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:34.590640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.276413Z digest=sha256:d92eb1db5a375b8bc9e7639f303d5a3bb96d74283f2b337e64c43a2c7860c93f

Observation 64164d42-fc27-40eb-8b7d-6a7dd2404fe1 · outbound

This paper cites Large-scale Pre-training for Grounded Video Caption Generation.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Large-scale Pre-training for Grounded Video Caption Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:58:21.238931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.322852Z digest=sha256:3556564dbe44f186c08c9bc1009e9039d6c67eb028512d6631db15ba838f5250

Observation 490269bf-c118-43c1-9d77-91433c9dc0bd · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Detecting mo- ments and highlights in videos via natural language queries

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:34.262660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.360030Z digest=sha256:8ba9731c4d4c643bbc363edb358dc0983dce1927335f1a50cd4b2a7c68b8ae43

Observation bdbad2d6-08e4-4336-8bb6-36cb5609d366 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.394371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.394371Z digest=sha256:f2f640475f73546db175fd56d10a90b6f9a2bb784cd8b5309f5a9f917c0df109

Observation 565fbcb7-361f-4c0f-a99e-47a1c40b960c · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoChat: Chat-Centric Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.444568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.444568Z digest=sha256:ea3b82ad99d4c03b15a991a0d056cab3e0d4a91da446a7812f833f7a6a38fd35

Observation 4e2dcb1b-3744-4395-a5b1-e1c1d46511d2 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.986752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.510722Z digest=sha256:6903a3955d8736b0c1ef92ad34a2e64a77a6c18092ed417e33788f0660423036

Observation 5f901906-4b2c-4118-aa92-62de57bd8c5d · outbound

This paper cites Temporal reasoning transfer from text to video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Temporal reasoning transfer from text to video

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.713478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.553789Z digest=sha256:8016f9650604cb350619b29866c850d01ec119cbbbea17f11bcd2ae87b621be2

Observation b7e4c9cd-d50d-4af7-8df4-24e422889d97 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Llama-vid: An image is worth 2 tokens in large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.367721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.591382Z digest=sha256:990dabea2e01abbe2c6fe61eccb86261a26a9a1ee8ab86305697b6810c6d2c6c

Observation 17a543bd-35a1-449b-abf9-58aa9b4c13ea · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Describe Anything: Detailed Localized Image and Video Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.670856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.670856Z digest=sha256:5548f987ca1ad527ef83979ff961dc7c62e4d1259b1ef50687599d69bdd465a4

Observation 51407507-def0-4822-9ef0-dcab963e35dc · outbound

This paper cites Unleashing hour-scale video train- ing for long video-language understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unleashing hour-scale video train- ing for long video-language understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.708257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.708257Z digest=sha256:cc8c984fedf042c6020045f77a4d1f05e98e1c740c4f4cdfccff5cc7c99f5a48

Observation 2b0f3896-8ba2-43c2-ac5a-aa020adf157d · outbound

This paper cites Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.722899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.722899Z digest=sha256:6dfa806711003fb33298ffc6ab8207bf0af0b4b5e19db30adbed1aa01bba6360

Observation 7d10a3ff-980e-41d7-a53f-13c215e27703 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.779731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.779731Z digest=sha256:42fe1830bc290342509c0379b6a82104ebd1dd6554fc6e1e995218d635ca81d3

Observation 10c054de-5eed-4652-8f0e-67c4dddb1291 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data TempCompass: Do Video LLMs Really Understand Videos?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.815963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.815963Z digest=sha256:606922c30a45b45d376ab70de228055414eafd42d0a03cbf3ed0249e2a7e5ee7

Observation 8baa8687-9803-408e-b636-62050abd02b7 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Groma: Localized visual tokenization for grounding multimodal large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.010944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:16.858999Z digest=sha256:c0ff375dbcca6d7b37173707264e0736d645c75a79e1154437e1add57d2d3fa2

Observation a6ad5f79-7adc-411a-9357-50dd0a011a34 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.903200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.903200Z digest=sha256:dcdbc344ce13bc7cb347c5ae9f7d793c8f432e7f20d66208ebcdc71d424387b1

Observation bd380dbb-488a-4b83-8b6c-16e9af6165b4 · outbound

This paper cites Point and Ask: Incorporating Pointing into Visual Question Answering.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Point and Ask: Incorporating Pointing into Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.959800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.959800Z digest=sha256:62ca330302331bc44484f80b2747a6f1593ec4c38eae16ec5db9e4bd1eb90a66

Observation 94876515-aa3e-476b-a6bb-d9e3c9eeeb79 · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.017701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.017701Z digest=sha256:618587a62a62681a2f137c15dfcb4f0675183464d023a4de37532e5a7c574da7

Observation 716af1ec-e4e1-4635-9627-c571d526b777 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level vi- sual grounding in videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videoglamm: A large multimodal model for pixel-level vi- sual grounding in videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:32.712204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.050270Z digest=sha256:56754315f528e484b0fd78c76dda6f4865d1be9352a50e9025d07b6d61e9fd23

Observation 5ff4291b-5505-4f60-b17f-e45a306f10fb · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.121792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.121792Z digest=sha256:4bd958ed1f002e582fb69378281bb83cc9690eeb20b782327eea4b7ce1374442

Observation 0b2eee42-add2-4e91-92c4-61e8bce2a2ad · outbound

This paper cites Artemis: Towards referential understanding in com- plex videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Artemis: Towards referential understanding in com- plex videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:32.345406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.184642Z digest=sha256:3f682e9f90770eaed621547a009a4d272d14ddb171022117384b3e2ff0bc4713

Observation f1ab99bb-d21e-4920-99d6-f9af787c780e · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Glamm: Pixel grounding large multimodal model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.969203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.240520Z digest=sha256:c60f30e9276920afeac3f6a8e17aa0bbf44783de045a0361117a0208a00cb7d3

Observation cd859a89-9070-41aa-871c-aff310a6b94a · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SAM 2: Segment Anything in Images and Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.383157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.383157Z digest=sha256:f44c6b5967a9e92bb6156d36f943dcde418dc480c8baec4c4a849d8ee1f183c3

Observation a4b7911d-4e68-4380-aa44-c84057af99cf · outbound

This paper cites Grounded sam 2: Ground and track anything in videos with grounding dino, florence-2 and sam 2.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounded sam 2: Ground and track anything in videos with grounding dino, florence-2 and sam 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.616302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.432596Z digest=sha256:ce4306bf8df3a6d1ee05eb790b9b4cd03cf8b2695e1a94f1d6eb07e9928b5392

Observation 2f9abcff-b849-49cd-b4bc-18e75325fbd6 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.487286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.487286Z digest=sha256:32a4a6a89e8194ec61336d364195c67ea4240fd6fe9853d612b3a2aaa45a3e97

Observation 99bcb17b-27a2-48d3-93ce-16d3844b925f · outbound

This paper cites NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.549752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.549752Z digest=sha256:905eba17209d6aff0edfe2fa841fb2f433e3677e1fced4a31b1353455b9d05b7

Observation 1bbee3d4-1b02-4f24-acb0-d7539bf57205 · outbound

This paper cites Sama: Towards multi-turn referen- tial grounded video chat with large language models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sama: Towards multi-turn referen- tial grounded video chat with large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.583118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.583118Z digest=sha256:9c08fc373c9a7419170aa139a1543868fc030a5f7c68110d95eb47bbda4dad59

Observation 0c54bd9b-9a9c-4f0e-9a94-cc1157014dd7 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Qwen2.5: A party of foundation models, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.332165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.612753Z digest=sha256:10cdec19a6b6da520d40e558c6c615d0e10244d685c76485a2d53813d25bd11e

Observation 7214d2b9-9d9e-41f2-8fbf-0bf0ee11b373 · outbound

This paper cites Natural language processing with Python and spaCy: A practical introduction.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Natural language processing with Python and spaCy: A practical introduction

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.040447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.667045Z digest=sha256:60802ba31c4c89d1f32fa1f62bccf91eb32db4eab22b519482403b8605862f6d

Observation 41ca1c2b-a3bf-4f21-bef5-9b4434fbc12a · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.747639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.747639Z digest=sha256:375429927acad786c41def9d9bf3c37be10d45fc7efea877645a65d0e31329ab

Observation 765d890c-9a71-4e79-97ce-21c300b094b2 · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Elysium: Exploring object-level perception in videos via mllm

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:30.665992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.831748Z digest=sha256:68d48a5d70242621050da00b84ff93cc6c1f1fddb319f12006535abb9d4de5c1

Observation ff965332-63fc-4e5b-87ca-42b14e74a0f2 · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:30.357181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.857705Z digest=sha256:1ab341331e42999860f67aa592c52b60ffdc3c711af55a7a04d90d78ae10de14

Observation b15deeb9-3579-48ff-a139-a73346e85ea6 · outbound

This paper cites Language as queries for referring video object segmen- tation.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Language as queries for referring video object segmen- tation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:30.042271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:17.889827Z digest=sha256:fa4466e4cc91746a21fad1b243ef1c504bf820663f814d3c45b800113a7ccbb4

Observation 88627f00-7b73-4104-9607-25201e2aaa50 · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.003545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.003545Z digest=sha256:6ceba393e52c614bf9e8180d940e2064fe296fb856afa277340f03b930d38582

Observation 21cda32f-45a3-4ae8-a22a-8e9a1467a8ae · outbound

This paper cites Number it: Temporal grounding videos like flipping manga.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Number it: Temporal grounding videos like flipping manga

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:29.734566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.090304Z digest=sha256:94f0ba126b550d04e180198afc2510574425c2605e0a473883fafb7c89491e74

Observation db8b1beb-fc4e-4632-af2d-dd2f0be45627 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Next-qa: Next phase of question-answering to explaining temporal actions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:29.392629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.182650Z digest=sha256:b4919d224d4a1655b657d78f8809105085e1520cfe77ba6a5fdf4f3ea416a16d

Observation fca9e09a-1ba7-4a9c-b092-17d5b39bb4d6 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:29.082214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.254568Z digest=sha256:91a4409f4ac6efa7588bd95e2ddb1db75ae481c8977f6b3b404750cb5d5dcdec

Observation 29bf663a-d044-4637-a009-e82cb6143f55 · outbound

This paper cites Pixel- aligned language model.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Pixel- aligned language model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:28.697397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.284425Z digest=sha256:000bb4c0aa4053fe2b58c931f89fbb2b089c7357aeb0f2a99e7a62810e3c865d

Observation fec93ee3-44fe-4d4c-95c9-34e70626e67a · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.338413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.338413Z digest=sha256:c1227a99c070fea5cb223c782cd7cd903bd1fdfa5ea4656ed1780d15e1c75ecf

Observation e49916f7-c529-4d14-8a85-dbf10b8ef374 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.397448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.397448Z digest=sha256:7b2bdcdde73aaf6c1fcb5d2399cb52d4a2c68b5bb2802005be05f062fc701aea

Observation 823bc005-5621-42bd-9c92-70213b7cdcbd · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data xgen-mm (blip-3): A family of open large multimodal models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.453965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.453965Z digest=sha256:9f9f42445010bc9d92940c2c2cadb9e98e25faa647a3ed82ab182dcd7f721115

Observation 77f0dbfa-cd10-4f3b-99b6-b97953a77307 · outbound

This paper cites List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.485504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.485504Z digest=sha256:9c2d398a829081171360e14b035225ac1bc71e0706b4b73c56429842077aab7c

Observation 98767322-0937-462c-b96f-9cabf0103a4f · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.599712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.599712Z digest=sha256:f95ff310c98002520f4b39a7d7a38d3cf177a098a4c33e2a1b7a7e489e7193c6

Observation 93f24843-fe72-45a5-b580-731de430261e · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.651851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.651851Z digest=sha256:302ffaaa2dd2203e3bb4493d641e39cadd8756b16dbb8cd142857b7a008e49d1

Observation ce8cfc46-6dae-45ea-98e2-9f2899656ef4 · outbound

This paper cites Merlin: Empowering multimodal llms with foresight minds.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Merlin: Empowering multimodal llms with foresight minds

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:28.238086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.742480Z digest=sha256:ad185f852faf944e18665435ec54727800087fb9e090382439e1d6f2c2662b01

Observation 2ebddabe-83ff-4b38-bf40-e465dcc7d3d2 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:27.920430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.800425Z digest=sha256:5a47d5bb052f031f9cac5546543c6c079f05045f8c3fc3280beff03108a24c2e

Observation 9aece6be-3583-454e-8cbd-a93e978b42d7 · outbound

This paper cites Osprey: Pixel un- derstanding with visual instruction tuning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Osprey: Pixel un- derstanding with visual instruction tuning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:27.591043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.873688Z digest=sha256:ad86ca0114e5f443d73741f87bd94e0a8e2e28039e00360c2b0e2fcaeff28e70

Observation 3a73b83d-ca66-4222-8e4d-56efb146d259 · outbound

This paper cites Videorefer suite: Advancing spatial- temporal object understanding with video llm.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videorefer suite: Advancing spatial- temporal object understanding with video llm

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:27.131183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:18.929032Z digest=sha256:7b8fe92eecf61382d7685edf83d5ffe50f0fcc2b4db113f7a3546a165bd46404

Observation d5d9aa8a-38c1-4266-bf08-5ad6c73145e3 · outbound

This paper cites Sigmoid loss for language image pre-training.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sigmoid loss for language image pre-training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:26.785238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.027799Z digest=sha256:1d3cab67ab3881f034135209b121c0e23ea7b71cdbc9db543ff61a6165406c8b

Observation 9d81e749-85c4-4ba8-ae83-d53265e10d7d · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.096136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.096136Z digest=sha256:20d4d839ec63e7057858623463338c7b0ea3470794facd54bfead2e55d0938e0

Observation 8918d51d-37d9-45ad-be60-d56aad392477 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Llava-grounding: Grounded visual chat with large multimodal models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:26.325015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.153012Z digest=sha256:4c900940b058f2e22cd64feae74c80481545211ccdbc37a1d950867f3638aed5

Observation 65a71161-201a-4948-826e-c30a6127c282 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.257522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.257522Z digest=sha256:508eaa0f958cc0ef7dc772b6e6d9f6b9821e35d6dfc7d4df5ed4d4803c0281e3

Observation 103af01e-08d4-495a-a86f-d49e1f4b426e · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region- of-interest.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Gpt4roi: Instruction tuning large language model on region- of-interest

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:25.935731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.342543Z digest=sha256:35dd7d2c33a834b9fb5daf883ff41b08dfcf80c062cda3e3cc1c74e99f2d7ad4

Observation 0db6e86d-0d55-44b4-8dbc-7acf2127e4cf · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video instruction tuning with synthetic data, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:25.520312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.436000Z digest=sha256:52ea6f954f1bf22c95b89cf186bca029ade01c613f1348674f506a0ac9533b99

Observation 59a300d9-b8e6-4ef6-a34b-318b14e219a4 · outbound

This paper cites VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.515582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.515582Z digest=sha256:a360d96f20fd1b4761db4b985a69acfc899f0cb6d8b11c55eb15312f72654f32

Observation 29fb73ab-44d5-4544-b861-1cfecb24fc66 · outbound

This paper cites Frames are extracted only from the seg- ment of the video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted only from the seg- ment of the video

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:25.072468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.612113Z digest=sha256:0600392de0e006f64af35982f098818ca9c3fa3559189b0bdfd69fbc03d662cf

Observation 796b7f92-13a3-4030-89cb-11fb6d2ec278 · outbound

This paper cites Sorry, I’m not sure.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sorry, I’m not sure

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:24.720638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.706926Z digest=sha256:bad5b296fb9dabe582b484a9b3ee163ad28bf5fe6301c7cc9b41525644ee01f2

Observation 278c6d5f-1929-483f-8e31-c7a64564ef6a · outbound

This paper cites Frames are extracted only from the seg- ment of the video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted only from the seg- ment of the video

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:24.365748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.711705Z digest=sha256:45db33699c8e9bc4d5f644557bee324e13b52fdb671825a178f45e2e90b419e1

Observation 7ad677b9-c848-4f99-9c64-189edb0320df · outbound

This paper cites Yes” or “No.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Yes” or “No

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:23.998339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:19.893683Z digest=sha256:c415f728d3e2855e61af216c242355ca205ba58dfa0b9c5dcbb3a276deea4934

Observation 7d300569-30cc-4e3d-842d-d40c885d5598 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:23.669493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:20.028670Z digest=sha256:fb96f641e88d68905ab8e4c6e75fc13ff9c08d81b2a51c7ea76dee59faa4f78c

Observation 6bbfbe7e-2d92-4072-90ff-90d6db7ad08a · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:23.298483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:20.152919Z digest=sha256:6a4ba7ded317239bbef5abc6a19a9397127fe1595ddcfb9419b9389b2b13e043

Observation 2f61eb7b-21ac-44fd-b0ec-b3351f8cb620 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.941919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:20.210594Z digest=sha256:2a35652ab98873e32376e5d81c0bbd9e4995b4305310e13c6b052285721a98bf

Observation 5fa46932-f1db-4106-b288-55198e7607f5 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.613750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:20.340722Z digest=sha256:710ded5c0bfb8d653578dedaeeba29eed59cdb3dc9a74b69fc39165d8d81888e

Observation c01a5e71-d7c6-47ad-811f-86fc083866b2 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.332818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:20.543748Z digest=sha256:b9de3ea1bc5d38808a4f563ceea06da147e2ea6e071c519ea5fbfc5143d75a8c

Observation 0c0d1e74-60de-41c2-bae6-e8f3673ee118 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.033684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:20.629580Z digest=sha256:1d433f829a3fe6acce17ef2a3a5dd5afb1dd786bf5a6beb5f961849a61b09690

Observation ac81c5fd-500b-46bf-a0ae-bd3438dac2a7 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:21.791400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:58:20.690107Z digest=sha256:a31cfb37029df473f086893efe518653e42fbdc9f8240c4c2accf385d21eeb53

Pith citing papers

Observation 7d5ed8f9-de26-4309-b4df-0653128618b2 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.660900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:babf077b932c5c21eb90636697be5ba1266fc639318eda63949c334f462d5aa2