Pith. sign in

Paper Citation Record · LEDGER

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 2 inbound Pith citation observations for arXiv:2508.01699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01699 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:30:37.955324Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T19:27:29.843866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:06:08.881592Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 623f6c64-f604-4f32-8938-a8446102c4c8 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:44.003791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:33.412600Z digest=sha256:98a4f2781f0b6cc77e5900736b6c8cd6cba29b8e3bdd70bb5baa6c5490504450

Observation 05c00db8-871c-4775-a3a2-3c3ce4dec1b0 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Activitynet: A large-scale video benchmark for human activity understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.987971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:33.531809Z digest=sha256:96de633445a3b2e520ec792e777196a58723ee3520c75d3189fa128715ed55d6

Observation c78fdb2e-00a1-4b18-bd51-4781f388e678 · outbound

This paper cites Sharegpt4video: Improving video understand- ing and generation with better captions.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Sharegpt4video: Improving video understand- ing and generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.971860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:33.609028Z digest=sha256:ad21d369cde3e2408c3e1466315a14d5dd234b337e7dff289023626c66c19ea7

Observation d594ea80-d3fa-48c4-b294-be2a229abc9a · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.954786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:33.692548Z digest=sha256:8ede789ec01a624f4ed0c71b23597d7edecff87a0f2093360b2de508ea376181

Observation 3ef79b53-06a9-4b21-9f6c-85d5dd158b02 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:33.778727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:33.778727Z digest=sha256:fbc4f4b854175d0b6c8458446b9a2d1625c6930707479a1f29c053133ef79dcf

Observation b9a04bde-8af6-4c7d-b103-ceec727b8e87 · outbound

This paper cites Uni- fied scaling laws for routed language models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Uni- fied scaling laws for routed language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.938058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:33.858259Z digest=sha256:8b408252011c86efdbe42d2efae1ec04fca24a977d77112770a82f62afde67d4

Observation 1a072bc1-35a1-4940-a54e-7d53b065ba84 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.921880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:33.901509Z digest=sha256:4d109cb011b8e01f6605226ce9dc606cfd9c5815d759dc4bb97620781e518ecb

Observation f6ea2427-d8b2-44f8-9eec-6ee7fca202fc · outbound

This paper cites Learning Factored Representations in a Deep Mixture of Experts.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Learning Factored Representations in a Deep Mixture of Experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:33.960108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:33.960108Z digest=sha256:297468f9b43fa899e4520f641430ccb862d5f29510879e423abbb19120a51d1a

Observation c0c05667-94c4-484c-95d8-f322f64d16cc · outbound

This paper cites To- wards an empirical understanding of moe design choices.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding To- wards an empirical understanding of moe design choices

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.905539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.049822Z digest=sha256:38820e903fd3b4dbe90aa96e98e4c9332f7f7db22cecd81352f2b46687f830ab

Observation 157887cb-5000-4b81-8b97-0f9409a752e7 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.890185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.155000Z digest=sha256:b23cf9c7dd882825dc85498748976dd3cf2497492e6404583e01946f7c8a7e15

Observation 1113bebe-bf15-4fd8-894d-e77ba17abd5c · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.872370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.195738Z digest=sha256:dcb42ad54cc4b0e48f30c351f4437b21111e387492492d81f87df1f87adc5fbd

Observation 3df24c47-2fca-46c6-abe5-2bc47edf56da · outbound

This paper cites Soda: Story oriented dense video captioning evaluation framework.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Soda: Story oriented dense video captioning evaluation framework

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.856269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.229575Z digest=sha256:8e5db6f6c7d560d569d183e8b33dabd8b78033189e17a0183c74cbf5f76295e2

Observation 605203f0-9526-4c6f-9f37-717c149fb719 · outbound

This paper cites Tall: Temporal activity localization via language query.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.840107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.278365Z digest=sha256:da9191d232ce3d7f29fc08b31de2a6567cd84eb7f37c2f814a2880a1b2257a67

Observation 093ce128-28cb-4f61-8c89-8d03aad676a0 · outbound

This paper cites Dynamic mixture of experts: An auto- tuning approach for efficient transformer models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Dynamic mixture of experts: An auto- tuning approach for efficient transformer models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.821904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.391428Z digest=sha256:f75b05582bf0e64cb50af9c94d860a8be2cb785d6c99519b7a4c0c89e1a2c1c7

Observation 6889abd9-af84-4ed9-bb06-8b189a4ebf88 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.804342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.460788Z digest=sha256:3294949999fa63d9001d6883b29c5a76a50ab6abe21c7d4aae48ee093e36353a

Observation a1f15d89-1ff6-4c8b-871f-0de6772c6963 · outbound

This paper cites Trace: Temporal grounding video llm via causal event modeling.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Trace: Temporal grounding video llm via causal event modeling

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.787199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.482059Z digest=sha256:5fc4ef799dac6c62230c8873755a01d20bcccfc700cf1ce79be0d9c1ac9ece45

Observation 131fdcac-9151-4a86-b262-05f56902c143 · outbound

This paper cites Creating summaries from user videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Creating summaries from user videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.770016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.561469Z digest=sha256:68f8b3cc98c295f2a103f82cd9d7e9ae0315dc74826b1665efc6223b383c1f74

Observation 17d0bea3-7751-402c-b7f0-41c96bcaaf06 · outbound

This paper cites Unleash the potential of clip for video highlight detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Unleash the potential of clip for video highlight detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.753155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.659566Z digest=sha256:e84031dbbce3ffe7cb97e86a5903fb3a0a7c60e8b792200bbdd09d8a13b884a1

Observation eb98ddcc-bbff-493f-8710-4580b8875508 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vtimellm: Empower llm to grasp video moments

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.735002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.719234Z digest=sha256:fd7578b880ade196f51f314b72fa42926d036e54996ac27ce45c3319d0008ece

Observation 2f14ab98-16ba-440c-ac4c-2eae1cc543bd · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Lita: Language instructed temporal-localization assistant

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.717956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.745891Z digest=sha256:07784c60baebe2c3eac23e21798770214330a5cc0bb289a3ffec6352d98d1897

Observation ea59856b-3465-421e-893d-13195462547c · outbound

This paper cites Harder tasks need more experts: Dynamic routing in moe models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Harder tasks need more experts: Dynamic routing in moe models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.700576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.788888Z digest=sha256:f5146aa4ab6d079ff084c90138a947f7ba009905146458edc9a5edc6db0965e9

Observation 2b6eb25a-64d9-42c0-b027-0f29de08ed76 · outbound

This paper cites Do you remember? dense video captioning with cross-modal memory retrieval.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Do you remember? dense video captioning with cross-modal memory retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.684117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.845670Z digest=sha256:d7164eb897541eb9e0e02ee91163179aa195fa6a10c563afe837858923baaa4c

Observation 087627ca-d238-4b66-b779-31be644de0bd · outbound

This paper cites Dense-captioning events in videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Dense-captioning events in videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.666589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.928179Z digest=sha256:cce3d4e19b278f074bb7b9cf994bb25205208a90e08af792293d730c55f45898

Observation 0d490c18-b50d-4187-83f2-910cd91a7aa5 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Detecting mo- ments and highlights in videos via natural language queries

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.649257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:34.980675Z digest=sha256:06c9ec35fdcf7a33fbf1454a8f41424a3a5ff6e75a9976ce690a51c3992bcafd

Observation 25639214-899c-434e-956e-1357c905573c · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.065340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.065340Z digest=sha256:50f13b9fac364d3a8735291850aaa01a86921735f9b70e22831e4f2729227788

Observation f1aa857f-33c9-40a7-b022-50c49599aa1f · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Unmasked teacher: Towards training-efficient video foundation models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.632430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.126090Z digest=sha256:b7de495a5d5dfa2fbe6bc0f61fd00ff4f511bb11e40fea28f93842ae579e2f53

Observation 17249dd6-7a2b-4b77-bdc5-01035a740a7d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.616051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.192935Z digest=sha256:c8a7b83b4d774d0f7b9dd2d9a12fdfdeb48510d734527e2b1e8835c97dc21a59

Observation 38d5f266-c7c9-4f71-a086-1d79f9b2507a · outbound

This paper cites Uni- moe: Scaling unified multimodal llms with mixture of ex- perts.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Uni- moe: Scaling unified multimodal llms with mixture of ex- perts

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.599390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.268512Z digest=sha256:7bd30d2b0d8bdd16445206dcb131e0d9c380d9840e404cd55f3984fa2b41789b

Observation 0bf732be-fc47-43d5-94dc-a48d507c7e89 · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.583338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.379944Z digest=sha256:3f58fd5c1449246f6fa29a39d7d3532e3278e0fa01a9142cccb49930a76cc4fb

Observation 74721bab-b9d3-4406-83fb-14975db5b8d5 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Univtg: Towards unified video- language temporal grounding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.566814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.452520Z digest=sha256:6208015ddaecf669f6abdd993f1cf5f8168046d33f7a9e27d00261b0a22c0133

Observation 3919a88f-713a-407b-b5e3-5031f14b8e5e · outbound

This paper cites Visual instruction tuning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.270070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.516888Z digest=sha256:3f9a232c1161668ddb8567365490c4ebc799032783f05777912ed1a3014c2f8e

Observation 6fa5e354-bf29-4290-84af-8d37a5a3f3be · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.095980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.562812Z digest=sha256:2e24dc75c2608559e630151602b95b40e4484c3226d97e20f3730a650f6de62f

Observation f21d3b11-8dd8-4f2a-8183-1df5bf4165df · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.652829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.652829Z digest=sha256:0e820ce463001f52285f447177a04e0e9315e62bc779318ad9dfb78de0a78bab

Observation 4f7a096b-b4d3-44bb-8fae-6767e6fa3049 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.962184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.721302Z digest=sha256:a4263b1ce6abdd8db6ffb8a321ae0d325a670dffa90120909a01e7831985a0b1

Observation aad3652a-5708-4e5d-8c17-6efe370f64ca · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.805762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.805762Z digest=sha256:38ef410bc406cf09fcf7fc57f021078201f872a9088d5ca11e1f9a019b56583b

Observation 946c2b1c-897d-422a-a44e-bf2b23a11a99 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.861021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.859524Z digest=sha256:1ea42b7088100ae39a9208009b6cbf87373f11b73e3a898f2732c3f0f01ad9a6

Observation 951b66dd-ded4-46e1-b914-0e7ae09582af · outbound

This paper cites En- coding and controlling global semantics for long-form video question answering.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding En- coding and controlling global semantics for long-form video question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.691002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:35.945738Z digest=sha256:b0087bf8de5ae172d43deca891ba29e02ed077e11b3755916ba205cb7f30c4e5

Observation 5d933afe-c99b-4224-ba99-330958abac2c · outbound

This paper cites Queryd: A video dataset with high-quality text and audio narrations.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Queryd: A video dataset with high-quality text and audio narrations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.342649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.024923Z digest=sha256:bf1eb6ee6452ce4bdc3c62e3daed61ee16b185bb5df8d1359f2f2106d4c932cc

Observation d99fb04b-5f3b-4521-9298-0da608427cec · outbound

This paper cites Momen- tor: Advancing video large language model with fine-grained temporal reasoning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Momen- tor: Advancing video large language model with fine-grained temporal reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.092397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.097163Z digest=sha256:7c28f38863fb5fb305a783afe455c2aba04bc202347aa4f0a592fdea21fc9704

Observation 5147b790-31ef-4d86-aa40-411c7b036680 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.922959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.191091Z digest=sha256:d0a7064399122ed0be4002366de3879175abeb8d13621f8455ccf76bb006f346

Observation d5300c57-1ae8-49f0-8ef3-53426d8d5c33 · outbound

This paper cites Outra- geously large neural networks: The sparsely-gated mixture- of-experts layer.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Outra- geously large neural networks: The sparsely-gated mixture- of-experts layer

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.737786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.236316Z digest=sha256:fc0eec8e7201ecf8302cf07a1adda2282c3756827ca4a4968dd5a50906f585e4

Observation e0ed41b6-1d91-4256-9b4f-85882b0e8c6e · outbound

This paper cites Tvsum: Summarizing web videos using titles.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Tvsum: Summarizing web videos using titles

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.545662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.311506Z digest=sha256:99cc28da74b37283760428a5ec58ae844762d02a3461ea8fb56ff904f4c18700

Observation fe940160-b914-400f-ab09-e401a411b2b2 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.321973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.406795Z digest=sha256:71d1452fba31ae7e226f83df07578cab25bb17eadd9c672ac7a03cc56a74b4da

Observation b25f3dc0-f091-4f80-a22d-fb0b32701f85 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Cider: Consensus-based image description evalua- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.139740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.468685Z digest=sha256:74667988066938193c4005ca4cc59c9c1ca569388ee6823846ecc6071b2bb07f

Observation cb8d8afd-6d9a-4832-a6ef-ffdc5fd7d998 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.525254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.525254Z digest=sha256:5a029acf08c32d92b93004feb741592d17c7bf0dd8eb7a85fd95f4112155774f

Observation 93fa13a4-7f4b-47c4-a171-aad8de3a708e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.592742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.592742Z digest=sha256:4ac3628d1fe994cc23983fa3394c783901a5fccbb5e7b5f6f2b3b24ed0f87605

Observation d2919905-838c-4982-8446-f1457ca57e9c · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding End-to-end dense video captioning with parallel decoding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.943395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.628329Z digest=sha256:6ee62cb9bfa82ff6d5b315e722b9209e6c047616e95779a66ee13a94b32da123

Observation d3969c71-fd3e-4fd9-a942-9c16d6e17c84 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.666749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.666749Z digest=sha256:1afb162bc4dcfdcb9a4c7e73e2a64c67786101b0754f88a7355a6bce12d5a6e4

Observation 87fcc106-ab07-4f0a-8792-1a6c1b708276 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.705649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.705649Z digest=sha256:601314fa9eef6dee61e30a0cc2d18f546a803e630092407e85e43e2078627031

Observation b558d7db-a77f-49b3-bc16-66330f50fb08 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Star: A benchmark for situated reasoning in real-world videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.747026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.758678Z digest=sha256:d119a7b6a61c97bbcb1d15ac54118e5cefca8519ec231344f35c64c58d862708

Observation 3726fe4c-6468-4c6f-bd24-70914bd28998 · outbound

This paper cites A large cross- modal video retrieval dataset with reading comprehension.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding A large cross- modal video retrieval dataset with reading comprehension

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.563938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.832418Z digest=sha256:ce458f1af22a9d7a634397ac3959fe3c286d6974473b203ff5ea44ee76bca21a

Observation 63d58ae2-fa10-454f-98de-3299a0f51c49 · outbound

This paper cites Multi-head mixture-of-experts.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Multi-head mixture-of-experts

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.392301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.911944Z digest=sha256:e33e5b6d406939aec6951c48b1e5b9ebf7a24f8ef286a25f8d5d7fd49ee34d78

Observation 4c366bce-39e4-4cd0-8593-0e30cda2ded0 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.224920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:36.976754Z digest=sha256:9f6d1d204ed4cc78426382ce53e05d7cebb05c061ace67a82ab2d42e3f9e65ca

Observation ba09c676-ddfd-4cd9-a31b-72a8c20d5874 · outbound

This paper cites Videoclip: Contrastive pre-training for zero-shot video-text understanding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Videoclip: Contrastive pre-training for zero-shot video-text understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.028153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.021922Z digest=sha256:3584b4260670d53b2a65ab49e0d33869d805a55e0f99e10745b8aadeec334bcc

Observation 7af50d6b-ec72-4721-b2f1-72a2dbe5712b · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.087279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.087279Z digest=sha256:804483aeaa593b6b04f56aacbe92a8aec3a64dd8ebf66285c5ac824da359ed0c

Observation ea7a8c66-e388-4263-897b-46e884175c77 · outbound

This paper cites M6-T: Exploring Sparse Expert Models and Beyond.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding M6-T: Exploring Sparse Expert Models and Beyond

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.167043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.167043Z digest=sha256:b35ee14887ae1e831acca376a74e8777ea0bd9812dca5d6c0e53f62a01fb7797

Observation 851e325b-d961-44ab-8b75-d640d8e482e4 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.774822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.242864Z digest=sha256:d5d477bf21c081499d3b954335d076b0891d764de08dd7ee06b81ee04cd65076

Observation 5a8dcb2a-b125-4ab5-81b9-631efe2ed276 · outbound

This paper cites Xmoe: Sparse models with fine-grained and adaptive expert selection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Xmoe: Sparse models with fine-grained and adaptive expert selection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.469525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.309509Z digest=sha256:0979dcd95074e381c436f466933a548a461f97dfa0ed8914285291acb3cdbd0e

Observation e3580b97-3d5a-4a1a-8e90-e0f4f2709d2b · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Hierarchical video-moment retrieval and step-captioning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.278032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.434673Z digest=sha256:b3e1679198be191d5c51e9b3928c26d41ddf695c8de8318b144f41f5bb33646b

Observation 4aaf3c91-7ccc-49bc-a41f-3d144cf8a627 · outbound

This paper cites Unimd: Towards unifying moment retrieval and temporal ac- tion detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Unimd: Towards unifying moment retrieval and temporal ac- tion detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.112892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.530612Z digest=sha256:0b585966e314b529254e54c29bbd35d5e8d5e3d36570196fd4d0d55e97279252

Observation 446f505a-9a3c-48ad-b300-2ea55a96d67f · outbound

This paper cites Adamoe: Token-adaptive routing with null ex- perts for mixture-of-experts language models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Adamoe: Token-adaptive routing with null ex- perts for mixture-of-experts language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:38.881163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.591682Z digest=sha256:728a69adc159ebd422fd06181075966e277ea221a3b2f30aaca9e7985bd71df0

Observation 40fedbaa-ed58-438a-ac61-d11118b75316 · outbound

This paper cites Sigmoid loss for language image pre-training.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:38.619902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.662646Z digest=sha256:f00ebb85976feab84876067c75b0db7203eea821fd41784791d80160d0895463

Observation fc61fe64-0088-4c2f-b95f-349400158007 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.791480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.791480Z digest=sha256:181cec8e3723a84809f535b15354742d3b911629542b255f08fd3d21bf2ae586

Observation 01d9109f-96a0-482e-a24e-52e57a29f998 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Towards automatic learning of procedures from web instructional videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:38.418588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:30:37.874247Z digest=sha256:e7e1e2a5c2df3dbc54f50b51b1ace787c1c92cf36c15eea5a5be4214a4d95456

Observation e2d15f56-e33b-49f5-9a5f-b9ea67da2ed7 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.955324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.955324Z digest=sha256:0911e7789a62487f0f221f440a90c750b331e622b09691b56a22fb7aaa0f4be4

Pith citing papers

Observation 8ca6f62c-11c8-46a6-8623-e65dd0c21a46 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Reference 46

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:06:08.883921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:45:50.422679Z digest=sha256:8993a5cb4551889e993afa3fc3ca77445fd0e933391b72b72d1e87fec6823983

Observation a4663677-3956-4d70-991c-cadd2285c595 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:ef8366537c96514ef317fd3960953b6d92d4d3ff0c0e4a00e0131c20a8bf0343