Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:48.211356Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.10692.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:48.211356Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation db8af9e1-8c39-4713-8d3c-e623fa29db71 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 643381ef-c2b0-4fa5-9f2d-59eb4ecdd82c · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c43ec5cb-6aac-43bb-ba86-362b02fd13de · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection 2 Related work Most previous MR&HD approaches [5, 11, 12] only em- ploy image and text inputs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a2afd7b4-877e-4000-bd7f-f1507ba26e52 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Effectiveness of each module in MRNet on QVHigh- lights val split
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a007fb0a-100a-4d32-971e-d10c7194861e · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The main reason is that Moment-DETR only utilizes RGB, which fails to fully under- Table 5
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2a1b707c-b3d1-49a0-b54c-cb00fc211b4e · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Localizing moments in video with natural lan- guage,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c00b4893-4ba6-4451-b5f7-87942f4596af · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Less is more: Learning highlight detection from video duration,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e4951687-383d-45cd-8e87-f5c8c75176d6 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Detect- ing Moments and Highlights in Videos via Natural Lan- guage Queries,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5c3af061-9e22-46ab-bc1b-32290657fc3e · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection UMT: Unified Multi-modal Transformers for Joint Video Moment Re- trieval and Highlight Detection,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c8c0675-fbc7-4842-ab7a-46be91c2877a · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MomentDiff: Generative Video Moment Retrieval from Random to Real
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c05860d-5739-4c27-9f29-94bb911e5446 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Gmflow: Learning optical flow via global matching,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e1917f8-3afa-4cdc-8c29-bfc7eb45f657 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Depth- cooperated trimodal network for video salient object de- tection,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9fb4ab9b-c1aa-4395-a6fc-23887b50d271 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Pyramid Feature Attention Network for Monocular Depth Pre- diction,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 29f57026-dfaf-4e31-9962-5bfcc611925f · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection How hierarchical is language use?,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1abb676d-9dce-4296-bce8-ef2645e45d9b · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The emergence of hierarchical structure in human language,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb7e6f6a-cb38-494f-8377-fc6072587a61 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection GPTSee: Enhancing moment retrieval and highlight detection via description-based similarity features,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f5b4600-e35f-4a04-85b9-265ac4897fc6 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008f428b-8dcf-4a60-9f1b-351eb1da1fa2 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection An empirical study of end-to-end video-language transformers with masked visual modeling,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation db86f254-899d-4769-902c-65d7b044870c · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection End-to-end object detection with transformers,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d51e2849-68c4-4d12-b15e-08fc9156fa70 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Shifting more attention to visual backbone: Query-modulated re- finement networks for end-to-end visual grounding,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e3fd3cd1-dd33-4563-a411-8c65390926cd · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Re- thinking transformer-based set prediction for object de- tection,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 539371cd-619f-48a9-86e5-08a1ae89813d · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85873ffe-579d-4028-88f7-4c0ea041b52d · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning transferable visual models from natural lan- guage supervision,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef50378e-ec48-4585-a08d-3416706516b6 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Recognizing American Sign Language Manual Signs from RGB-D Videos
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe17b1a6-13ed-47b1-8643-697bde5ccf45 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Layer Normalization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 978ceb9b-bade-4c84-8886-774fdcd1944e · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Temporal Sentence Grounding in Videos: A Survey and Future Directions
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e48aeb59-c1af-459e-8ac0-3ffd49040fc8 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Attention is all you need,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c1646409-cf5d-4aca-9303-6a1b637f9e74 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b48827d7-79ec-4974-b1be-47e0abd404e2 · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Tall: Temporal activity localization via language query,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 324d2ccf-197e-459a-be21-cfa0132c1e6e · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Decoupled Weight Decay Regularization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d7fa82-70c5-4433-a111-b76be3f778ac · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Self-Chained Image-Language Model for Video Localization and Question Answering
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d26cd0a-6015-493c-a3e8-088fb680396d · outbound
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning 2d temporal adjacent networks for moment localization with natural language,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.