Pith. sign in

Paper Citation Record · LEDGER

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

As of 12 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 2 inbound Pith citation observations for arXiv:2508.01699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01699 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:30:37.955324Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T19:27:29.843866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:06:08.881592Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 623f6c64-f604-4f32-8938-a8446102c4c8 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:44.003791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:33.412600Z digest=sha256:aa29e3efcd4eaf16dfeb06615951d6d8fb7c95144972526a152e362d979cfbdd

Observation 05c00db8-871c-4775-a3a2-3c3ce4dec1b0 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Activitynet: A large-scale video benchmark for human activity understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.987971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:33.531809Z digest=sha256:80eee90486db7ac79cb328c77875ecfdb52f19a1082e28f7145e8fd273c70338

Observation c78fdb2e-00a1-4b18-bd51-4781f388e678 · outbound

This paper cites Sharegpt4video: Improving video understand- ing and generation with better captions.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Sharegpt4video: Improving video understand- ing and generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.971860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:33.609028Z digest=sha256:d7599c1185d90280dee3f5559e6b93feeda541cdc430916147f895f630db7119

Observation d594ea80-d3fa-48c4-b294-be2a229abc9a · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.954786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:33.692548Z digest=sha256:ae6da04c6de12152e7cb105399a10d5b45ca7f0c82816831d87384a8f5987fd2

Observation 3ef79b53-06a9-4b21-9f6c-85d5dd158b02 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:33.778727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:33.778727Z digest=sha256:6fc29d94b28cd1de5628bf291175eea32c5b72987a6e4f0aee206c586855c350

Observation b9a04bde-8af6-4c7d-b103-ceec727b8e87 · outbound

This paper cites Uni- fied scaling laws for routed language models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Uni- fied scaling laws for routed language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.938058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:33.858259Z digest=sha256:0c4968fdad7ddfd226043de04d970d3b55a12b566f45ef96cba4eeb42ff1622b

Observation 1a072bc1-35a1-4940-a54e-7d53b065ba84 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.921880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:33.901509Z digest=sha256:d66ca86c582f192fde4fdc273df4f400ca0652772889b55feb54b6aac44b1488

Observation f6ea2427-d8b2-44f8-9eec-6ee7fca202fc · outbound

This paper cites Learning Factored Representations in a Deep Mixture of Experts.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Learning Factored Representations in a Deep Mixture of Experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:33.960108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:33.960108Z digest=sha256:63ed80580d344cb61f06d84e196f33dde5a56215468ca8bf187b5b3ecc7e809e

Observation c0c05667-94c4-484c-95d8-f322f64d16cc · outbound

This paper cites To- wards an empirical understanding of moe design choices.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding To- wards an empirical understanding of moe design choices

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.905539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.049822Z digest=sha256:6f8a28c3121755725c16ccf35fa6aa7b8e9a35c8a17d875498bef620075c6b36

Observation 157887cb-5000-4b81-8b97-0f9409a752e7 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.890185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.155000Z digest=sha256:933640eaea01c016a7ac6f02ef75234dfff12824bcdcc8d05ace3ea874f0ae4e

Observation 1113bebe-bf15-4fd8-894d-e77ba17abd5c · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.872370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.195738Z digest=sha256:fb71db27499b25dc6587d66aa690509493690945cdb1ef9c57f24e4bfa4e0a45

Observation 3df24c47-2fca-46c6-abe5-2bc47edf56da · outbound

This paper cites Soda: Story oriented dense video captioning evaluation framework.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Soda: Story oriented dense video captioning evaluation framework

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.856269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.229575Z digest=sha256:a5e8f1ae955354437f5447e22cd07d987052fdaa0a6fd2bdc7cddad6691c72d4

Observation 605203f0-9526-4c6f-9f37-717c149fb719 · outbound

This paper cites Tall: Temporal activity localization via language query.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.840107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.278365Z digest=sha256:4cb62db0c2ad2c2b8894c1dc9fee2471a918336e07549bd91056a4c8e19e13dd

Observation 093ce128-28cb-4f61-8c89-8d03aad676a0 · outbound

This paper cites Dynamic mixture of experts: An auto- tuning approach for efficient transformer models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Dynamic mixture of experts: An auto- tuning approach for efficient transformer models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.821904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.391428Z digest=sha256:529ab7c27c99dbdbb06b915debde97f71eb9ea4ac4062bb83a300754e9cdd0ec

Observation 6889abd9-af84-4ed9-bb06-8b189a4ebf88 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.804342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.460788Z digest=sha256:a772a72246cd5415a74ed94d0c02eb2a151079cc609f8f8ccee34eb25c206df6

Observation a1f15d89-1ff6-4c8b-871f-0de6772c6963 · outbound

This paper cites Trace: Temporal grounding video llm via causal event modeling.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Trace: Temporal grounding video llm via causal event modeling

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.787199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.482059Z digest=sha256:ec911ef5510f825fe9fcba1bc74f9da74710a78fad650e85a0597c35432ca3d5

Observation 131fdcac-9151-4a86-b262-05f56902c143 · outbound

This paper cites Creating summaries from user videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Creating summaries from user videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.770016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.561469Z digest=sha256:f9a01029fd3804642d2ab545bc4e64e3384d1e077586b255b7009c77e0334544

Observation 17d0bea3-7751-402c-b7f0-41c96bcaaf06 · outbound

This paper cites Unleash the potential of clip for video highlight detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Unleash the potential of clip for video highlight detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.753155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.659566Z digest=sha256:d7f008eb3d128e1515cf172af45246f86909e827922e6f026064fef570aff6ee

Observation eb98ddcc-bbff-493f-8710-4580b8875508 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vtimellm: Empower llm to grasp video moments

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.735002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.719234Z digest=sha256:1bb4a2579fefaed1cc47f902f2c711fae342901e2a15274ebff60bc76a3e5d3f

Observation 2f14ab98-16ba-440c-ac4c-2eae1cc543bd · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Lita: Language instructed temporal-localization assistant

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.717956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.745891Z digest=sha256:804ed00dfb2630029a82d6a1e56e8e83b1128934232ddbb96e6056435e87572c

Observation ea59856b-3465-421e-893d-13195462547c · outbound

This paper cites Harder tasks need more experts: Dynamic routing in moe models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Harder tasks need more experts: Dynamic routing in moe models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.700576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.788888Z digest=sha256:06926710101183367280309788855d975cdcad2988a6afc0fa69a7128fb55d92

Observation 2b6eb25a-64d9-42c0-b027-0f29de08ed76 · outbound

This paper cites Do you remember? dense video captioning with cross-modal memory retrieval.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Do you remember? dense video captioning with cross-modal memory retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.684117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.845670Z digest=sha256:6ffc52130e3c42b1d3da16d4e79a3cf73b3d5beb5fc18a34cde5f86b761d408b

Observation 087627ca-d238-4b66-b779-31be644de0bd · outbound

This paper cites Dense-captioning events in videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Dense-captioning events in videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.666589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.928179Z digest=sha256:fb68ec922ff8628d4469a247ee15cd6afe4fc02099b6c35c86ea8cee917dc260

Observation 0d490c18-b50d-4187-83f2-910cd91a7aa5 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Detecting mo- ments and highlights in videos via natural language queries

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.649257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:34.980675Z digest=sha256:edc4be487eb497bd734d40c0879cc8bf9219c7ecb2e0911d762d73436fc7bfed

Observation 25639214-899c-434e-956e-1357c905573c · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.065340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.065340Z digest=sha256:997abc3eaaacc2c552952807a2edf84e8fbd88768e2200f329dabcddcb7423af

Observation f1aa857f-33c9-40a7-b022-50c49599aa1f · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Unmasked teacher: Towards training-efficient video foundation models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.632430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.126090Z digest=sha256:fe5cd5a2b134ad9042f33154c79ccbb0f29bc4ff389b8b53a29fa6acd0d31332

Observation 17249dd6-7a2b-4b77-bdc5-01035a740a7d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.616051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.192935Z digest=sha256:308ce1a07823e5814ad336905c6f6133b4cac1c740ce319c26bf6dceca5eade0

Observation 38d5f266-c7c9-4f71-a086-1d79f9b2507a · outbound

This paper cites Uni- moe: Scaling unified multimodal llms with mixture of ex- perts.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Uni- moe: Scaling unified multimodal llms with mixture of ex- perts

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.599390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.268512Z digest=sha256:991c0590c35fde43ad1082f98cb1d0dbbc1838b1a1e3582071a4ce9b9579fa79

Observation 0bf732be-fc47-43d5-94dc-a48d507c7e89 · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.583338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.379944Z digest=sha256:cda2be575db7a931af862bb297d204571993e79031a738607f8ed0e1bfc6ecda

Observation 74721bab-b9d3-4406-83fb-14975db5b8d5 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Univtg: Towards unified video- language temporal grounding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.566814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.452520Z digest=sha256:6517c57e04deec65a2aff193937f333887db1aa20af2fe2cdac2423e93071334

Observation 3919a88f-713a-407b-b5e3-5031f14b8e5e · outbound

This paper cites Visual instruction tuning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.270070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.516888Z digest=sha256:16daf5322fedbf835f4d99c65de85a2cfb077b3417b001fa9690517366b38b9f

Observation 6fa5e354-bf29-4290-84af-8d37a5a3f3be · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:43.095980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.562812Z digest=sha256:5de8ad4ceb4e0fb1ce12bedafa8fd924513e31284e287e313aa6d2aaa8dcde7b

Observation f21d3b11-8dd8-4f2a-8183-1df5bf4165df · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.652829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.652829Z digest=sha256:1bcb8854f778ed251fbf4eb1b186bdb24c6af2fa1d82dfab7ea7542bbf9f6fb4

Observation 4f7a096b-b4d3-44bb-8fae-6767e6fa3049 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.962184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.721302Z digest=sha256:e1176cb06bc36361533aa114cc06244a1fd2b01bb4cae928eef96e6f11191ecf

Observation aad3652a-5708-4e5d-8c17-6efe370f64ca · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.805762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.805762Z digest=sha256:d53a0c893be274b8fed2b04e0887586d9a1704766e4f9f463f0bf9d0792a6fdc

Observation 946c2b1c-897d-422a-a44e-bf2b23a11a99 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.861021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.859524Z digest=sha256:e5a13c3f2d9f7e71d6897a4d6100dc2e3501bd6c257a570971076d76d822cf50

Observation 951b66dd-ded4-46e1-b914-0e7ae09582af · outbound

This paper cites En- coding and controlling global semantics for long-form video question answering.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding En- coding and controlling global semantics for long-form video question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.691002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:35.945738Z digest=sha256:553d44df463cd03dc8d8f7248ba1ef9c80edb15e382b637bc8fe9699fb9e8df8

Observation 5d933afe-c99b-4224-ba99-330958abac2c · outbound

This paper cites Queryd: A video dataset with high-quality text and audio narrations.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Queryd: A video dataset with high-quality text and audio narrations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.342649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.024923Z digest=sha256:792bacd483b3c407ac1e56c28400ac7c3fc9b11de453b644a393ac12cf8b1a73

Observation d99fb04b-5f3b-4521-9298-0da608427cec · outbound

This paper cites Momen- tor: Advancing video large language model with fine-grained temporal reasoning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Momen- tor: Advancing video large language model with fine-grained temporal reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:42.092397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.097163Z digest=sha256:c6dee22b4876dcdb4c31a4eecddd11c2817b61c7dfae6fff4e06df7f3f0b9a98

Observation 5147b790-31ef-4d86-aa40-411c7b036680 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.922959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.191091Z digest=sha256:10bcbebcb62a85dc921cdf0ac825f2f8d3068738576af560684e5b5c2ba8fcfd

Observation d5300c57-1ae8-49f0-8ef3-53426d8d5c33 · outbound

This paper cites Outra- geously large neural networks: The sparsely-gated mixture- of-experts layer.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Outra- geously large neural networks: The sparsely-gated mixture- of-experts layer

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.737786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.236316Z digest=sha256:f10549a4434a793e4fcdfe2b06cff27965386ad690b7aa2f0c1603345cbbf880

Observation e0ed41b6-1d91-4256-9b4f-85882b0e8c6e · outbound

This paper cites Tvsum: Summarizing web videos using titles.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Tvsum: Summarizing web videos using titles

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.545662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.311506Z digest=sha256:11275a7889f50837fa30a4162a29584db9d8f914e0481a53aacee4646f397ca8

Observation fe940160-b914-400f-ab09-e401a411b2b2 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.321973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.406795Z digest=sha256:eb0353d5a94d4b41de2dc5e25a60e5252ab1c9c86aa9a0f5febfb316a9384fcc

Observation b25f3dc0-f091-4f80-a22d-fb0b32701f85 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Cider: Consensus-based image description evalua- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:41.139740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.468685Z digest=sha256:456bc7dcbc040067b9739fbca975e08effe37d52603a7bbf2d15c79ab03063ac

Observation cb8d8afd-6d9a-4832-a6ef-ffdc5fd7d998 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.525254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.525254Z digest=sha256:debf7ce0682988733734c1c8ba3482137bed71a5b8468f231261795c09e67050

Observation 93fa13a4-7f4b-47c4-a171-aad8de3a708e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.592742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.592742Z digest=sha256:8cea11c059cf81f0908a64f4915ad2c2b419439280d42e201f5c0e09fa8bdb04

Observation d2919905-838c-4982-8446-f1457ca57e9c · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding End-to-end dense video captioning with parallel decoding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.943395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.628329Z digest=sha256:523e2311ac0bfe8e5b790396e79e481340de305ae70cc7bff95d6c4a33c3810a

Observation d3969c71-fd3e-4fd9-a942-9c16d6e17c84 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.666749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.666749Z digest=sha256:c718c6b17313d8956f3dc5f08431e92c925a7c616da5b65759dd482a61707485

Observation 87fcc106-ab07-4f0a-8792-1a6c1b708276 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.705649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.705649Z digest=sha256:75fbeae765ca153fb0d1791cde1fda466d91ec850d456c2edd752c4a4d82b34c

Observation b558d7db-a77f-49b3-bc16-66330f50fb08 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Star: A benchmark for situated reasoning in real-world videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.747026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.758678Z digest=sha256:b6ecf8b73d1206fb456d32872833ce92072067339997dfb27758de84649b1bf7

Observation 3726fe4c-6468-4c6f-bd24-70914bd28998 · outbound

This paper cites A large cross- modal video retrieval dataset with reading comprehension.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding A large cross- modal video retrieval dataset with reading comprehension

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.563938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.832418Z digest=sha256:6de4ba55390aca8441284ed4e9afe2265dc2269aacd2d36a9e453ee033c8925b

Observation 63d58ae2-fa10-454f-98de-3299a0f51c49 · outbound

This paper cites Multi-head mixture-of-experts.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Multi-head mixture-of-experts

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.392301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.911944Z digest=sha256:03f2f77c92adf4f70a31683df52322561cedf61eeef01e27233a1680c91d1818

Observation 4c366bce-39e4-4cd0-8593-0e30cda2ded0 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.224920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:36.976754Z digest=sha256:83ea68cb0f96246d42a52883e8212f7e1b7807c507a3361e5d0f82f90d955662

Observation ba09c676-ddfd-4cd9-a31b-72a8c20d5874 · outbound

This paper cites Videoclip: Contrastive pre-training for zero-shot video-text understanding.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Videoclip: Contrastive pre-training for zero-shot video-text understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:40.028153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.021922Z digest=sha256:1c93902f862a370d623db6d5024e83eb25a2173f2c33d03f3a55c22557a2c3de

Observation 7af50d6b-ec72-4721-b2f1-72a2dbe5712b · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.087279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.087279Z digest=sha256:bea4d4689255959ab0dce0862c06932f8a3336c2f4accb5dd6b631105b08e2c2

Observation ea7a8c66-e388-4263-897b-46e884175c77 · outbound

This paper cites M6-T: Exploring Sparse Expert Models and Beyond.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding M6-T: Exploring Sparse Expert Models and Beyond

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.167043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.167043Z digest=sha256:6b4796756f6a2c0a0c8b47d1f8339f0ba2253eb726df04bf83a92b7d1b42396f

Observation 851e325b-d961-44ab-8b75-d640d8e482e4 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.774822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.242864Z digest=sha256:e729ead8e5a76cb94b9ff0c72f77f8bbb523df2092e1a532e8d8d797a30a571a

Observation 5a8dcb2a-b125-4ab5-81b9-631efe2ed276 · outbound

This paper cites Xmoe: Sparse models with fine-grained and adaptive expert selection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Xmoe: Sparse models with fine-grained and adaptive expert selection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.469525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.309509Z digest=sha256:4fbb3ac74df55c255e5d5d234380278c6a8ebb8788a736d90c50f53d64184d95

Observation e3580b97-3d5a-4a1a-8e90-e0f4f2709d2b · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Hierarchical video-moment retrieval and step-captioning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.278032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.434673Z digest=sha256:1e6aafb03c2fa6565b09dc3c2b8b5f17f141c62d3c2d82ee60e83ff0d2cf7934

Observation 4aaf3c91-7ccc-49bc-a41f-3d144cf8a627 · outbound

This paper cites Unimd: Towards unifying moment retrieval and temporal ac- tion detection.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Unimd: Towards unifying moment retrieval and temporal ac- tion detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:39.112892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.530612Z digest=sha256:65a684d50f988fc67f13931bce3643cf34fd270d3541688008049129a6945671

Observation 446f505a-9a3c-48ad-b300-2ea55a96d67f · outbound

This paper cites Adamoe: Token-adaptive routing with null ex- perts for mixture-of-experts language models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Adamoe: Token-adaptive routing with null ex- perts for mixture-of-experts language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:38.881163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.591682Z digest=sha256:ab007562d42c8a42f1998cf2ddb161354d80721980a034ae1ffac11cb8c36117

Observation 40fedbaa-ed58-438a-ac61-d11118b75316 · outbound

This paper cites Sigmoid loss for language image pre-training.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:38.619902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.662646Z digest=sha256:9920543872fa61149141d40bb7a440f8234fa9823b5ad6c867299c3f008c893f

Observation fc61fe64-0088-4c2f-b95f-349400158007 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.791480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.791480Z digest=sha256:3f116dff50665b0bb43bd67d59d584c442b18950d988848a5d98c61c3bb5a7f5

Observation 01d9109f-96a0-482e-a24e-52e57a29f998 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Towards automatic learning of procedures from web instructional videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:30:38.418588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T05:30:37.874247Z digest=sha256:b873e483e387bb977f56b91c2b3c3f7b4b574e53a354c399c5e0004e52fff0f5

Observation e2d15f56-e33b-49f5-9a5f-b9ea67da2ed7 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:37.955324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:37.955324Z digest=sha256:c62d406aca2d79a80288d44efa02461379ddbfb368f8da3473de65287a8a152d

Pith citing papers

Observation 8ca6f62c-11c8-46a6-8623-e65dd0c21a46 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Reference 46

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:06:08.883921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T12:45:50.422679Z digest=sha256:8be8fb593571651ef9a2c4d39cf7728d0c68077b14002ab257567ae6f2cfe8a8

Observation a4663677-3956-4d70-991c-cadd2285c595 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:6abb694be7d589122b879d713cd75a3e1500dc4d107ce48992826f59216c745a