Pith. sign in

Paper Citation Record · LEDGER

HawkEye: Training Video-Text LLMs for Grounding Text in Videos

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2403.10228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.10228 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:35:47.240790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.195732Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 716f8125-2ca2-4693-9d43-eb94caa2de63 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:52:20.788080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:976d6ac862edbd43e921084cba9a2fbaa9ffd422ce1febc7621c7680081b9e9e

Observation 85ce6aae-bad9-42a4-9987-1de562e8821c · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.240790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.240790Z digest=sha256:e96ec9593b4b8de8a6d30d5b711ae325c9ed90e22cb1fc571b71a0663d7ecfd4

Observation b27e939a-30d3-47d2-b4aa-8417e26ffd4b · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:26.972515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:26.972515Z digest=sha256:0cdcaabd8e94a08417e631e2b1a9276f1ffa828f887f22726ece605a9a646acd

Observation 87fcc106-ab07-4f0a-8792-1a6c1b708276 · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.705649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.705649Z digest=sha256:0c3e2788abc6e73236e054e1910c3b0bc66cc333f4f1f1f1bc3bc1d0047954e4

Observation 4a2f41c7-4695-441f-a94a-413cc4806014 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:26.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:26.012600Z digest=sha256:7f7aff728295c3b316be028e0ada08d5f54e3a51bf4bf8287d7c9295bf802141

Observation 989f777a-e959-4819-8ad1-662748ee862c · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.599520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.599520Z digest=sha256:c8b966f710137df256186c8fe4433e85048cfd6d17038e719fe82f0d9f889d20

Observation 8fa2f2a3-cd9b-42d7-ba04-97b5d38e51c1 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.344850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.344850Z digest=sha256:ac8056af9f7b60c6557d5429d5bd8653d5111d2763aa41d2d5c1c3c7026215b5

Observation c2633b38-8492-4fae-b0b6-62c18d264932 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.343793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:29c9f50d036adebf08663aa0ea59711114c8bc37ff04c3183a5ef229edcd949b

Observation 040a74d2-2dee-4802-b188-0c456b63ffee · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.547114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:295e6440225d572440ceafef17621f30280a290868f0a657dbc91e5e58a97dea

Observation a6aa77df-74c7-41c0-87d7-7282fe22140f · inbound

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos cites this paper.

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.514531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T19:43:57.298338Z digest=sha256:9db2cb04cb30095d56bfa9b1a66694fbc8addaa92e4d71765eb2d9c65affbf43

Observation f06b4652-9297-48bf-a0b1-87467b1b6445 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:15:52.667017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:e2af46047c61bb356ba760c2cd468e50f88b2881dc794fcbd56e1cd62f452b42

Observation c35a53c6-e8b0-4b7f-9ab0-a1ec6b784bdb · inbound

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding cites this paper.

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:02.282639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:55:56.385801Z digest=sha256:2573a1a26f0a95fd66e3b8bef73c3f3c9b4841b965c5328201bb8609f0705a13

Observation ba024df3-6600-4bda-8112-4382727cef7b · inbound

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms cites this paper.

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:05:58.951662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:50:48.559703Z digest=sha256:6764dbee7562ebd5189f5099c628f989568f2d5fa2577a5954a8f9ba1c619ccd

Observation b05c671b-1fb7-4c2d-a597-34a0592d090b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:57.115545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:48d383123cb46bf98ba59a4f4fef242588893662dcaa31d62b91afd8dac27992

Observation 08fed6a5-2670-4c49-8ae0-e4892f7e1f78 · inbound

ViLL-E: Video LLM Embeddings for Retrieval cites this paper.

ViLL-E: Video LLM Embeddings for Retrieval HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:01.989634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:00:43.573409Z digest=sha256:90c421b0c26a05ae23f53eb8bb2654c1f82283c755618ca3b24f6ed18a67be14

Observation 2ef16935-95ca-4aef-a608-08692ba7edc3 · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.822999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:5f4b440a8e09610ff46aaf45a71a12af6da3a12efd4e1fb68660875e5147a169

Observation b0ba08a4-59c2-4a30-aaa1-0596f03ff238 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.699623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T14:10:27.416341Z digest=sha256:6b7fdb8c7ff2205521ac77eccb87f897c1a28e1cdd3c1f621b377211ac537200

Observation d46e3d6c-f44f-496b-9a3d-863043f454d6 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.762438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:77d759108d3575e344fcf5f92d2ecc1cfe38e4d19c85f17eefbdc565f62c9f12

Observation ea95495a-a242-433d-a6d2-ca7e8ac05a50 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:32:52.481502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:297f16ab507c53b4a708309ffb741d278fe8988deab26aedb1613acf534cea5e

Observation d38aec51-379d-40fb-a3b8-4ae26b317623 · inbound

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding cites this paper.

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:06:15.420230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-22T08:04:55.680668Z digest=sha256:daad774945d1ce12001e032687ba39f938c090245579c1a2eef66fec383eba7a

Observation 25b9edde-1c18-4302-88f3-8f11480aea48 · inbound

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding cites this paper.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.972853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:33d39d6be8d4d527b6aaa99f47e24052649e27fc8cf2b43dc11f20eb4411b895

Observation b4438bce-39e3-4a81-88d8-8fcb2c7a5e47 · inbound

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models cites this paper.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.672672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:5aff503cb1a22c82886b8b45bd45e9312e903885c53feb73fdc147c953734f4c

Observation 00a19d4c-3a9f-4e3b-a7f2-bfc08c45edc1 · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.834984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:2f576d027b177767f1063fc61224f9341e57e59c6e2e009d3b0eae7f4cd13d2d

Observation 03120690-d260-4787-a437-a7550b942b44 · inbound

Temporal-Aware Reasoning Optimization for Video Temporal Grounding cites this paper.

Temporal-Aware Reasoning Optimization for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.106952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T17:21:30.638275Z digest=sha256:a14f71a89a86403efbddb72d77fb3420a77b6572b0a7c65af301e7410c37aa79

Observation 6046348f-f56f-46c0-93f6-ed10f481f627 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:03.047929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:92acc0e0090059ca375fee5a298ba3b828cc4e0c5f97059081cb764f6e0ff0f3

Observation 0ddc5890-898c-4e74-8d69-05728f1ca9a3 · inbound

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition cites this paper.

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.644785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:01:09.933880Z digest=sha256:772ef7a8df4e74fa670a81c1b9273e356a529358a3136abedeb2ac02b193d4cb

Observation bc60d46a-72b6-452b-a97e-5da54f35f31d · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.197261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:44dbd143e8a39fbb00b54c15422183b4edf8b2fb1349a70ec04c426d10489950

Observation 82ac2554-f4d3-4bf2-977d-3c49f7fa0cba · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:07960f9fcdc0b720c9c171b4168b1dc94b044abc816576823802a6175fda8a38