Pith. sign in

Paper Citation Record · LEDGER

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

As of 17 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.10692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10692 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:48.211356Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db8af9e1-8c39-4713-8d3c-e623fa29db71 · outbound

This paper cites an unresolved cited work.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:05:49.166365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:47.988030Z digest=sha256:024a690aca8e41d3c1dc67ef7c4f0e7092b869042a32679cf3e196bfcc5919e2

Observation 643381ef-c2b0-4fa5-9f2d-59eb4ecdd82c · outbound

This paper cites an unresolved cited work.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:05:49.136619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:47.994781Z digest=sha256:2cb3bcc1328c2310a571acf9edc3dab0d3c061f9fadda82bed380d36df4a7ed9

Observation c43ec5cb-6aac-43bb-ba86-362b02fd13de · outbound

This paper cites 2 Related work Most previous MR&HD approaches [5, 11, 12] only em- ploy image and text inputs.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection 2 Related work Most previous MR&HD approaches [5, 11, 12] only em- ploy image and text inputs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.115294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.001085Z digest=sha256:0493bd098bd2ae4ef8602db0587efce5115801c12db8ca2922f712272566a4e1

Observation a2afd7b4-877e-4000-bd7f-f1507ba26e52 · outbound

This paper cites Effectiveness of each module in MRNet on QVHigh- lights val split.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Effectiveness of each module in MRNet on QVHigh- lights val split

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.087752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.008082Z digest=sha256:0effcb5aff12e384432f91f41cf6e78056a6086c088ed3035382ea75a915914a

Observation a007fb0a-100a-4d32-971e-d10c7194861e · outbound

This paper cites The main reason is that Moment-DETR only utilizes RGB, which fails to fully under- Table 5.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The main reason is that Moment-DETR only utilizes RGB, which fails to fully under- Table 5

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.054624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.015150Z digest=sha256:e5d6a8ad0768d41930be73739f1bd9a0948c79e12ee31273eaafc76374815c1c

Observation 2a1b707c-b3d1-49a0-b54c-cb00fc211b4e · outbound

This paper cites Localizing moments in video with natural lan- guage,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Localizing moments in video with natural lan- guage,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.034835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.020785Z digest=sha256:badfd00ba03089e6e7ade540e8931df7792cfa131fec3fc724bec8745a572ccb

Observation c00b4893-4ba6-4451-b5f7-87942f4596af · outbound

This paper cites Less is more: Learning highlight detection from video duration,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Less is more: Learning highlight detection from video duration,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.013948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.026677Z digest=sha256:40a16e10a4cbb2784e449dc4666021dcfa46cc6c81ae67bac7434944cba475c7

Observation e4951687-383d-45cd-8e87-f5c8c75176d6 · outbound

This paper cites Detect- ing Moments and Highlights in Videos via Natural Lan- guage Queries,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Detect- ing Moments and Highlights in Videos via Natural Lan- guage Queries,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.990075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.033453Z digest=sha256:7f3c58a61b63d4ed2c8de5a64a8c184857d784a0b9cab9e6f471f9a745e45b1c

Observation 5c3af061-9e22-46ab-bc1b-32290657fc3e · outbound

This paper cites UMT: Unified Multi-modal Transformers for Joint Video Moment Re- trieval and Highlight Detection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection UMT: Unified Multi-modal Transformers for Joint Video Moment Re- trieval and Highlight Detection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.959393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.038646Z digest=sha256:4f42d32cb92fb36d7f2ec9eca4099626c327aba527c58626bc2e39e2b79fbb11

Observation 9c8c0675-fbc7-4842-ab7a-46be91c2877a · outbound

This paper cites MomentDiff: Generative Video Moment Retrieval from Random to Real.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MomentDiff: Generative Video Moment Retrieval from Random to Real

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.545907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.046112Z digest=sha256:af438fe16efde2159e4416892233fd8b986d00ee730996022a71815e2b574013

Observation 8c05860d-5739-4c27-9f29-94bb911e5446 · outbound

This paper cites Gmflow: Learning optical flow via global matching,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Gmflow: Learning optical flow via global matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.934703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.061247Z digest=sha256:1f451e674e67a4549ca11eb1dba367fddf2bb67d26922c0ca3adf5075727deed

Observation 0e1917f8-3afa-4cdc-8c29-bfc7eb45f657 · outbound

This paper cites Depth- cooperated trimodal network for video salient object de- tection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Depth- cooperated trimodal network for video salient object de- tection,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.914079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.071804Z digest=sha256:bde2487273c0f77cfcd1165c6d188ab361c61cfa01f081ea80214e7cd6298a66

Observation 9fb4ab9b-c1aa-4395-a6fc-23887b50d271 · outbound

This paper cites Pyramid Feature Attention Network for Monocular Depth Pre- diction,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Pyramid Feature Attention Network for Monocular Depth Pre- diction,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.888167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.084100Z digest=sha256:c7e50d6e11b0fcf97a483dcc10b27d9e6cf6aa7145a766ebd54fd3e46438a739

Observation 29f57026-dfaf-4e31-9962-5bfcc611925f · outbound

This paper cites How hierarchical is language use?,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection How hierarchical is language use?,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.864654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.091445Z digest=sha256:e1ea10a4e94f1f1aac7f8e82c430cf1f025a8e737ee9bbd2928aecca3dc4714b

Observation 1abb676d-9dce-4296-bce8-ef2645e45d9b · outbound

This paper cites The emergence of hierarchical structure in human language,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The emergence of hierarchical structure in human language,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.843168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.097892Z digest=sha256:ec6ab6722eab5894a81a7f8da955b8aed5f442721eea48e284fe72075bae8daf

Observation eb7e6f6a-cb38-494f-8377-fc6072587a61 · outbound

This paper cites GPTSee: Enhancing moment retrieval and highlight detection via description-based similarity features,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection GPTSee: Enhancing moment retrieval and highlight detection via description-based similarity features,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.812275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.106113Z digest=sha256:dc46b910df6898e3e11817e3ecaab09deb29921600601fb9b7186322717e92b1

Observation 7f5b4600-e35f-4a04-85b9-265ac4897fc6 · outbound

This paper cites MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.113507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.113507Z digest=sha256:90f06021f1af96ca928fbc360bd4b47ce994ae56b7a84c686a3650b8aa39bbc2

Observation 008f428b-8dcf-4a60-9f1b-351eb1da1fa2 · outbound

This paper cites An empirical study of end-to-end video-language transformers with masked visual modeling,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection An empirical study of end-to-end video-language transformers with masked visual modeling,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.781442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.119107Z digest=sha256:0b1deaa56dbcc9ab82643cfd38ff67c4e17b3b398cd58411a62b9caab4203ec1

Observation db86f254-899d-4769-902c-65d7b044870c · outbound

This paper cites End-to-end object detection with transformers,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection End-to-end object detection with transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.758557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.124164Z digest=sha256:c8f7862b9eb6a25db241382bf4af4f3ff71ebc3bfb1278372ee8088d22610a24

Observation d51e2849-68c4-4d12-b15e-08fc9156fa70 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated re- finement networks for end-to-end visual grounding,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Shifting more attention to visual backbone: Query-modulated re- finement networks for end-to-end visual grounding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.732109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.131702Z digest=sha256:77f24d73850871536c01fee21b722dd7a27e603f955d008e895f65506472efa6

Observation e3fd3cd1-dd33-4563-a411-8c65390926cd · outbound

This paper cites Re- thinking transformer-based set prediction for object de- tection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Re- thinking transformer-based set prediction for object de- tection,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.710933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.138847Z digest=sha256:100b6e88e1151e003742e62dddc735aec39332e04e8e3d39a901f6ac370a1ad8

Observation 539371cd-619f-48a9-86e5-08a1ae89813d · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.144753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.144753Z digest=sha256:3ed39cb692b14cd11fcfba7589ce8758857e87f18856ac39d9eb21471dc30bd1

Observation 85873ffe-579d-4028-88f7-4c0ea041b52d · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning transferable visual models from natural lan- guage supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.685286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.150146Z digest=sha256:2348bd5e08e18c476340d7db9c15c56ca69e0137ae23086e1c402ea8748d9062

Observation ef50378e-ec48-4585-a08d-3416706516b6 · outbound

This paper cites Recognizing American Sign Language Manual Signs from RGB-D Videos.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Recognizing American Sign Language Manual Signs from RGB-D Videos

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.435257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.157147Z digest=sha256:6377fcf2b43f7cd177a6b2bfc51d1a7c6f8feb7f5215d9bd37c5169d22366d89

Observation fe17b1a6-13ed-47b1-8643-697bde5ccf45 · outbound

This paper cites Layer Normalization.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Layer Normalization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.162480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.162480Z digest=sha256:159ba366a2f80e27fc4392a6721b97846d9ce44c0347c0ed41f44eaf51071a4a

Observation 978ceb9b-bade-4c84-8886-774fdcd1944e · outbound

This paper cites Temporal Sentence Grounding in Videos: A Survey and Future Directions.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Temporal Sentence Grounding in Videos: A Survey and Future Directions

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.364200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.169189Z digest=sha256:0513bb05304bb4f993c9916d8eaab115de6bd7f6c757a8ca03181a86d1955cff

Observation e48aeb59-c1af-459e-8ac0-3ffd49040fc8 · outbound

This paper cites Attention is all you need,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Attention is all you need,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.654863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.179780Z digest=sha256:8432439ec04178aab2800c95ad22303211f27b5f102cbe5e29f427e2beee34b0

Observation c1646409-cf5d-4aca-9303-6a1b637f9e74 · outbound

This paper cites RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.635334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.186680Z digest=sha256:f866cd39fd2e7c9bba09183c0a914184262219d1d65f20973d64e5caaa351217

Observation b48827d7-79ec-4974-b1be-47e0abd404e2 · outbound

This paper cites Tall: Temporal activity localization via language query,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Tall: Temporal activity localization via language query,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.604469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.192444Z digest=sha256:39668421a69947fa0466d40f5dfd818dfd5378ff676d13d0b04b00b9d6b8c825

Observation 324d2ccf-197e-459a-be21-cfa0132c1e6e · outbound

This paper cites Decoupled Weight Decay Regularization.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.198444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.198444Z digest=sha256:3f93fea48bac94b5c240960e9b6035df741a0e98a5cf28d692db04560df0eb1a

Observation 03d7fa82-70c5-4433-a111-b76be3f778ac · outbound

This paper cites Self-Chained Image-Language Model for Video Localization and Question Answering.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.204282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.204282Z digest=sha256:d157da60facd3c4ff4fdeb271b4bb3dc2468a6d8090d3b8ba0ed9d076753f335

Observation 8d26cd0a-6015-493c-a3e8-088fb680396d · outbound

This paper cites Learning 2d temporal adjacent networks for moment localization with natural language,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning 2d temporal adjacent networks for moment localization with natural language,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.577349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:05:48.211356Z digest=sha256:28f66ee809634987510629837266eda2b4c6e51e2b1e907ea1875cc0cf3dacf4

Pith citing papers

No inbound Pith citation observations are available.