Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:58:20.690107Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2509.03501.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:58:20.690107Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-22T09:20:32.920925Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T09:21:20.658722Z
85 of 85 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f32d129a-eba2-490c-9936-8d5edfae3144 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ccf81e-1b56-427c-9b08-7e20054f5d45 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vicas: A dataset for combining holistic and pixel-level video un- derstanding using captions with grounded segmentation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5393e087-47aa-431c-a067-46362548aeee · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0acf6b67-da61-46a4-8d88-8b9b91104798 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PySceneDetect: Video Scene Cut Detection
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0942c527-958a-49d8-a7a2-3e6248af82cb · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sharegpt4video: Improving video understanding and generation with better captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a097da6-4d02-45a7-9027-91034c18167b · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b0f2dbe-11db-4cda-bb90-e1f1021855e8 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d03d88a-5f72-4a3d-bf4a-d778e64fa9fc · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 048c6c3c-bdd1-4fc5-86b3-6ebce32f8450 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4597e5-0732-488b-8a8c-dc1b7b90ad62 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unifying Specialized Visual Encoders for Video Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39097890-3252-438c-801d-864c4e3436f5 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videorefer benchmark evaluation for general mllms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea89e910-4024-4328-91d2-40ecbdec82b8 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Mevis: A large-scale benchmark for video segmentation with motion expressions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20a1c4c3-7e96-4250-857e-91bb626e0f3c · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4716635-c7ba-48e6-befe-dd37db008796 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc4f1d6e-e92b-4609-96e1-011c9c7cfa23 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 239fdc4d-74dd-4aac-8b13-50ca7b2927b2 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b30dc52-6e55-4609-94cd-8a3165b86594 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d6e30f3-90d9-4439-b413-0fe2d25dcc04 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vtimellm: Empower llm to grasp video moments
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca09e8d2-9d88-44e3-bebf-71617c12196e · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7f1e1c6-471c-4efa-b6ad-ef27fa137087 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Referring to any person, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d707d43-44e9-44ae-bad0-275f8ebeb5f5 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Miradata: A large-scale video dataset with long durations and structured captions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64164d42-fc27-40eb-8b7d-6a7dd2404fe1 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Large-scale Pre-training for Grounded Video Caption Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 490269bf-c118-43c1-9d77-91433c9dc0bd · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Detecting mo- ments and highlights in videos via natural language queries
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bdbad2d6-08e4-4336-8bb6-36cb5609d366 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LLaVA-OneVision: Easy Visual Task Transfer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 565fbcb7-361f-4c0f-a99e-47a1c40b960c · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoChat: Chat-Centric Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2dcb1b-3744-4395-a5b1-e1c1d46511d2 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f901906-4b2c-4118-aa92-62de57bd8c5d · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Temporal reasoning transfer from text to video
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7e4c9cd-d50d-4af7-8df4-24e422889d97 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Llama-vid: An image is worth 2 tokens in large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17a543bd-35a1-449b-abf9-58aa9b4c13ea · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Describe Anything: Detailed Localized Image and Video Captioning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51407507-def0-4822-9ef0-dcab963e35dc · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unleashing hour-scale video train- ing for long video-language understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0f3896-8ba2-43c2-ac5a-aa020adf157d · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d10a3ff-980e-41d7-a53f-13c215e27703 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c054de-5eed-4652-8f0e-67c4dddb1291 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data TempCompass: Do Video LLMs Really Understand Videos?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8baa8687-9803-408e-b636-62050abd02b7 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Groma: Localized visual tokenization for grounding multimodal large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6ad5f79-7adc-411a-9357-50dd0a011a34 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd380dbb-488a-4b83-8b6c-16e9af6165b4 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Point and Ask: Incorporating Pointing into Visual Question Answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94876515-aa3e-476b-a6bb-d9e3c9eeeb79 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 716af1ec-e4e1-4635-9627-c571d526b777 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videoglamm: A large multimodal model for pixel-level vi- sual grounding in videos
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ff4291b-5505-4f60-b17f-e45a306f10fb · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2eee42-add2-4e91-92c4-61e8bce2a2ad · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Artemis: Towards referential understanding in com- plex videos
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1ab99bb-d21e-4920-99d6-f9af787c780e · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Glamm: Pixel grounding large multimodal model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd859a89-9070-41aa-871c-aff310a6b94a · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SAM 2: Segment Anything in Images and Videos
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b7911d-4e68-4380-aa44-c84057af99cf · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounded sam 2: Ground and track anything in videos with grounding dino, florence-2 and sam 2
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f9abcff-b849-49cd-b4bc-18e75325fbd6 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99bcb17b-27a2-48d3-93ce-16d3844b925f · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bbee3d4-1b02-4f24-acb0-d7539bf57205 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sama: Towards multi-turn referen- tial grounded video chat with large language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c54bd9b-9a9c-4f0e-9a94-cc1157014dd7 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Qwen2.5: A party of foundation models, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7214d2b9-9d9e-41f2-8fbf-0bf0ee11b373 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Natural language processing with Python and spaCy: A practical introduction
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41ca1c2b-a3bf-4f21-bef5-9b4434fbc12a · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765d890c-9a71-4e79-97ce-21c300b094b2 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Elysium: Exploring object-level perception in videos via mllm
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff965332-63fc-4e5b-87ca-42b14e74a0f2 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Tarsier: Recipes for training and evaluating large video description models, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b15deeb9-3579-48ff-a139-a73346e85ea6 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Language as queries for referring video object segmen- tation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88627f00-7b73-4104-9607-25201e2aaa50 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LongViTU: Instruction Tuning for Long-Form Video Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21cda32f-45a3-4ae8-a22a-8e9a1467a8ae · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Number it: Temporal grounding videos like flipping manga
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db8b1beb-fc4e-4632-af2d-dd2f0be45627 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Next-qa: Next phase of question-answering to explaining temporal actions
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fca9e09a-1ba7-4a9c-b092-17d5b39bb4d6 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29bf663a-d044-4637-a009-e82cb6143f55 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Pixel- aligned language model
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fec93ee3-44fe-4d4c-95c9-34e70626e67a · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49916f7-c529-4d14-8a85-dbf10b8ef374 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823bc005-5621-42bd-9c92-70213b7cdcbd · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data xgen-mm (blip-3): A family of open large multimodal models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f0dbfa-cd10-4f3b-99b6-b97953a77307 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98767322-0937-462c-b96f-9cabf0103a4f · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f24843-fe72-45a5-b580-731de430261e · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8cfc46-6dae-45ea-98e2-9f2899656ef4 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Merlin: Empowering multimodal llms with foresight minds
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ebddabe-83ff-4b38-bf40-e465dcc7d3d2 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9aece6be-3583-454e-8cbd-a93e978b42d7 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Osprey: Pixel un- derstanding with visual instruction tuning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a73b83d-ca66-4222-8e4d-56efb146d259 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videorefer suite: Advancing spatial- temporal object understanding with video llm
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5d9aa8a-38c1-4266-bf08-5ad6c73145e3 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sigmoid loss for language image pre-training
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d81e749-85c4-4ba8-ae83-d53265e10d7d · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8918d51d-37d9-45ad-be60-d56aad392477 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Llava-grounding: Grounded visual chat with large multimodal models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65a71161-201a-4948-826e-c30a6127c282 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 103af01e-08d4-495a-a86f-d49e1f4b426e · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Gpt4roi: Instruction tuning large language model on region- of-interest
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0db6e86d-0d55-44b4-8dbc-7acf2127e4cf · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video instruction tuning with synthetic data, 2024
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59a300d9-b8e6-4ef6-a34b-318b14e219a4 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29fb73ab-44d5-4544-b861-1cfecb24fc66 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted only from the seg- ment of the video
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 796b7f92-13a3-4030-89cb-11fb6d2ec278 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sorry, I’m not sure
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 278c6d5f-1929-483f-8e31-c7a64564ef6a · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted only from the seg- ment of the video
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ad677b9-c848-4f99-9c64-189edb0320df · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Yes” or “No
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d300569-30cc-4e3d-842d-d40c885d5598 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bbfbe7e-2d92-4072-90ff-90d6db7ad08a · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f61eb7b-21ac-44fd-b0ec-b3351f8cb620 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fa46932-f1db-4106-b288-55198e7607f5 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c01a5e71-d7c6-47ad-811f-86fc083866b2 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c0d1e74-60de-41c2-bae6-e8f3673ee118 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac81c5fd-500b-46bf-a0ae-bd3438dac2a7 · outbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d5ed8f9-de26-4309-b4df-0653128618b2 · inbound
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.