Pith. sign in

Paper Citation Record · LEDGER

HawkEye: Training Video-Text LLMs for Grounding Text in Videos

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2403.10228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.10228 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:29:26.972515Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.195732Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 716f8125-2ca2-4693-9d43-eb94caa2de63 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:52:20.788080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:9cff3e0f65f54bf7c621846ce4ca3da2a1c4d5adf7d661df10affc8ea9abd354

Observation b27e939a-30d3-47d2-b4aa-8417e26ffd4b · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:26.972515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:26.972515Z digest=sha256:3c9dbb85340acfdd16baa352f51b01cb8c494bcc0070240bcb9d03747ae35b9e

Observation 87fcc106-ab07-4f0a-8792-1a6c1b708276 · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:36.705649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:36.705649Z digest=sha256:601314fa9eef6dee61e30a0cc2d18f546a803e630092407e85e43e2078627031

Observation 4a2f41c7-4695-441f-a94a-413cc4806014 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:26.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:26.012600Z digest=sha256:7f7aff728295c3b316be028e0ada08d5f54e3a51bf4bf8287d7c9295bf802141

Observation 989f777a-e959-4819-8ad1-662748ee862c · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.599520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.599520Z digest=sha256:8d55f1b05c552855a1ca28436e42bf3779952ee14941387ca3bc96e67ac9fc81

Observation 8fa2f2a3-cd9b-42d7-ba04-97b5d38e51c1 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.344850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.344850Z digest=sha256:ac8056af9f7b60c6557d5429d5bd8653d5111d2763aa41d2d5c1c3c7026215b5

Observation c2633b38-8492-4fae-b0b6-62c18d264932 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.343793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ef68af01eeefb4a91e41e863fc24c8da2ed6bf66aabff98f1fd4b3914ea33548

Observation 040a74d2-2dee-4802-b188-0c456b63ffee · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.547114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:606f25fa0cccf46f803224515fb09ad84223a5dbf480c4ceca81c5487c1e54d4

Observation a6aa77df-74c7-41c0-87d7-7282fe22140f · inbound

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos cites this paper.

A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.514531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:43:57.298338Z digest=sha256:283588eced4efd42f792dcd1c7a5f18361d132b6edbfafeb32418bddad06be84

Observation f06b4652-9297-48bf-a0b1-87467b1b6445 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:15:52.667017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:3f1e7832e89461d6ddc308588f71db9b6837b21b9adfc1061610c1d77a070620

Observation c35a53c6-e8b0-4b7f-9ab0-a1ec6b784bdb · inbound

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding cites this paper.

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:02.282639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:55:56.385801Z digest=sha256:9c84abd087591a6bce1c199773fa3f8617614296a25800b7ca46b73b4aa75f64

Observation ba024df3-6600-4bda-8112-4382727cef7b · inbound

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms cites this paper.

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:05:58.951662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:50:48.559703Z digest=sha256:de4d44220e285c6e564f484a8e3565c3322ef6bfc4799976479c08a9c9e95ca7

Observation b05c671b-1fb7-4c2d-a597-34a0592d090b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:57.115545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:a878d29c76b8f176e98f9d1f0a7d7d72c778508ff94b815b27ca7d388a7cbd30

Observation 08fed6a5-2670-4c49-8ae0-e4892f7e1f78 · inbound

ViLL-E: Video LLM Embeddings for Retrieval cites this paper.

ViLL-E: Video LLM Embeddings for Retrieval HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:01.989634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:00:43.573409Z digest=sha256:5faf7b7f2180e90e954836f5ee38f803b14c74d761d8959b715bb093d3c7aabe

Observation 2ef16935-95ca-4aef-a608-08692ba7edc3 · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.822999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:10fcb1bf0bdaec2d28140a615aff54445620d0a4347f68d0dc0d86d30b1a7bf1

Observation b0ba08a4-59c2-4a30-aaa1-0596f03ff238 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.699623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T14:10:27.416341Z digest=sha256:0e4b2134830f3f7a9302d2f221205d8fad5d8d338acbca0b6d9988b858bdf8f2

Observation d46e3d6c-f44f-496b-9a3d-863043f454d6 · inbound

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding cites this paper.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.762438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:df92ae639d81d7dd4bdfedcd7d23960e2296a25355ac52141841d429d5489114

Observation ea95495a-a242-433d-a6d2-ca7e8ac05a50 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:32:52.481502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:7458ff5d95b26947d61e24fbfa3a3f10147d332ea6c9c4dc99620aec2f2c5b81

Observation d38aec51-379d-40fb-a3b8-4ae26b317623 · inbound

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding cites this paper.

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:06:15.420230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-22T08:04:55.680668Z digest=sha256:ca4135ef3bae6a99f9dbb61d1b469a14db8111c727b7a863a8a7d0d16252dbfb

Observation 25b9edde-1c18-4302-88f3-8f11480aea48 · inbound

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding cites this paper.

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.972853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:27:37.631870Z digest=sha256:f70e0ad16047b2ac8c40687c0edbd85dc7db71d06d2a40e22d238a6adb335f8e

Observation b4438bce-39e3-4a81-88d8-8fcb2c7a5e47 · inbound

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models cites this paper.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.672672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:70d10c8f686e7db1cac5c44bdaaa44c024556cedab4b7d94a56478b9643402b6

Observation 00a19d4c-3a9f-4e3b-a7f2-bfc08c45edc1 · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.834984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:e390d7db0c80530bb1a9655e3b231dd6552aa67aa1a50ae5e0b5aff65e9c731d

Observation 03120690-d260-4787-a437-a7550b942b44 · inbound

Temporal-Aware Reasoning Optimization for Video Temporal Grounding cites this paper.

Temporal-Aware Reasoning Optimization for Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.106952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T17:21:30.638275Z digest=sha256:648fc0d9a6245779c7fecc2f4d479e1716714885544c7d70c6fe6e960a422b77

Observation 6046348f-f56f-46c0-93f6-ed10f481f627 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:03.047929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:741c906e69ed04c6183053d0bb70a6e66c8a0023f007d7b2d88e37caf10b2a53

Observation 0ddc5890-898c-4e74-8d69-05728f1ca9a3 · inbound

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition cites this paper.

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.644785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T10:01:09.933880Z digest=sha256:4b2be7a8c1c7ae1838f0b7c8a06de19d85e2a7c8d0baf5543dbf711e91c70d24

Observation bc60d46a-72b6-452b-a97e-5da54f35f31d · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.197261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:c1c702e1b1c0c33414bfe04c208ad69a6ceacdfaee9921f00c5d8dd20d3de6ee

Observation 82ac2554-f4d3-4bf2-977d-3c49f7fa0cba · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:07960f9fcdc0b720c9c171b4168b1dc94b044abc816576823802a6175fda8a38