Pith. sign in

Paper Citation Record · LEDGER

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

As of 11 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2501.07972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07972 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:34:12.380805Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T08:04:15.238840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:46:13.641295Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact5
  • verified fuzzy8
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb0021c3-7c2e-4226-b390-b98dcebf3366 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.066631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.066631Z digest=sha256:d954d28d472b93c29b2cd2084f551c0b46e12b395e6148ec70f4e4a313c887b6

Observation e44a10ad-2c5c-464e-aa7b-e04dc8a9ec1c · outbound

This paper cites write newline.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.071863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.071863Z digest=sha256:8effc8b593d9e780443f7578afd3865c7ef9cf4878bfa2fd334ce5b1c882a05c

Observation 256b286e-7b4b-4f0d-bf7f-272b2363632a · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.076756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.076756Z digest=sha256:f4138afe461bd884ff60f3a72e23ff1aab4bdf241e26cdad010c5a4039dca572

Observation f62c59b1-e218-4679-92a6-b9f430b5341e · outbound

This paper cites P.; Barbu, A.; Siddharth, N.; and Siskind, J.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models P.; Barbu, A.; Siddharth, N.; and Siskind, J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.641916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.080865Z digest=sha256:102af60e0b02769f493349742cd53fd14dea8155a06260d8455d3acb4bf5f8c9

Observation e7f76639-f72c-4e54-8acc-9ec4aa645a16 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.627790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.085047Z digest=sha256:d3b75f8ba0441e77ef516b1e37e56385ba4123aea1b71fd300a322dff8e7ed80

Observation 354051e0-70cd-4afa-9cf8-4a188aa44120 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.613611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.089625Z digest=sha256:f2b54482e6a2831dadb1834e00632ac9af2db6246e4b8f75686f5cffb0a0fabe

Observation 475e4609-9f6e-4c7f-be20-e8f98c26067c · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.599636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.093835Z digest=sha256:fdf85fa82159a48e76a2049f7d1e0d4ef545652c12741d7b75ca68e44bf53ceb

Observation 1ee15398-4f56-4c2f-b5b6-df58f1b47963 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.586371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.097724Z digest=sha256:0c95f1025f9227976f2b6a4783d9cf646559af0dcdbcfbdd5e777c574c884ab0

Observation 9200590c-8c4d-4085-b5b2-cdf08cc705ce · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.572804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.102567Z digest=sha256:3879b9e5ccb24202510c0780df0846a71669f708e3ec380b914765bb4bc8cd98

Observation 4216065e-0ec8-4eea-81c2-0698470bd880 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.559501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.106505Z digest=sha256:1e64282ef0c1e8cb229f376da2c27d5f83cbfb8ff2f35e7f8c8923a3822cfa21

Observation 341aa8f8-17dc-46ac-b682-7d2ea043de2b · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.544573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.110455Z digest=sha256:2ef2057b2510d68d39f6c6a8425a25a49a99e0368853e11dff57e41dd8b8aeac

Observation f71b0f06-2866-4329-9828-ed021470ece7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.114574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.114574Z digest=sha256:6760932179745fae86f0575d63de8b1630b71aed870738bb74be53787481380c

Observation e58aa81d-d6a3-4345-8de6-d4d17cd0e4bd · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models VTimeLLM: Empower LLM to Grasp Video Moments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.119139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.119139Z digest=sha256:7af5dfebe69a0547af88376e166df5882bc787fc9c91b2376bad97ca6c1897dc

Observation 2d7acad3-f351-4d19-97b7-78b8d8969e8f · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.530139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.124061Z digest=sha256:ee3148b5092d13c7320b997335b6b47f3742e24d4e8cedf7583cbd09c43c9807

Observation 4001ef24-17d0-4fee-be60-91c001cfca63 · outbound

This paper cites Mistral 7B.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.128101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.128101Z digest=sha256:e65ef20859ef2f8507e005534c46d1dfb5bd034d0e535981035e08d16419affb

Observation 9dde5a58-fa2b-401d-8d59-da131baf827b · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.515433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.132520Z digest=sha256:22a086f36ff500e252a6d05af7ecccb73c6d1434b1a1c3430aad3160fb4e2c72

Observation 6a5e672e-5b5e-4369-96b3-02660bf66d75 · outbound

This paper cites PRewrite: Prompt Rewriting with Reinforcement Learning.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models PRewrite: Prompt Rewriting with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.136449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.136449Z digest=sha256:86d90c657ad0827b4c027176840b39920ea17a84fa0a6cb6396ad6b7fb2967b1

Observation 534b0871-4796-4346-8ee2-a7ddffb4979a · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.501348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.141179Z digest=sha256:bb7845f5d6da53d8edaffe33b706cdf35ec9afaee7b26aaf6d9f067bdf4e4a97

Observation ccedf17e-7969-4c15-a7f9-0cff9f81fec0 · outbound

This paper cites L.; Boucher, A.; Thonnat, M.; and Bremond, F.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models L.; Boucher, A.; Thonnat, M.; and Bremond, F

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.487457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.145626Z digest=sha256:ac0a0a77ffba1ac04107e3fd91ac7e15dc4b8e291341460cb5f1b1162ebf236a

Observation f9e36f95-3e00-483b-a269-701e34f37ab9 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.474511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.149872Z digest=sha256:86187f94b20351f5adcf6eda7ebdf074015c3b0b1818244be7b5b9cbe77e2d7e

Observation fcb0160e-5e32-49a8-a623-d83cd0bf444a · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.153968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.153968Z digest=sha256:62af9f921bb3c7eb6be863c7f2929738bb3beacd9055525484ab9c955eb1561d

Observation 5a6e88d5-6b6b-4b5e-8c8d-597058c38908 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.159521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.159521Z digest=sha256:df6d7a21c2de74b1d5d12c253e755dc541c15325e87a7e2a8ca6cbe09dab21e0

Observation 567dc82d-b37b-4e58-8ad0-209589aac924 · outbound

This paper cites Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:34:12.724034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.164266Z digest=sha256:9a5400f2de39de26db8c9c7a1a31e9ff5a3c8e84a985f6c4da992ce3d71f6d66

Observation d84538eb-e0bf-4167-b794-8e12bcfdf1fe · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.461827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.168888Z digest=sha256:8a4a5a19a890fa62f993051ae96f2d2d5a0222c6cd44315cebd1706e6b44171b

Observation b1617b5a-556e-4091-b76a-d87ffedb67bc · outbound

This paper cites MomentDiff: Generative Video Moment Retrieval from Random to Real.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models MomentDiff: Generative Video Moment Retrieval from Random to Real

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.173111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.173111Z digest=sha256:e0f768342f90b1b980b4891177820719bbd69c6797e04bf6b1b702b52b1aeda9

Observation 62e9dc9a-f0c0-4840-8834-9d26300576d1 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Textbooks Are All You Need II: phi-1.5 technical report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.177877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.177877Z digest=sha256:f819d4854dd36ecf091aaa52117b69e51972176041c20d1aec444bfee1d743be

Observation ad25aa30-b5fc-45f2-95f3-8f47d4e9f884 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.183271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.183271Z digest=sha256:c253263255d7865c80f9f2085c2470d0e3cfa34016499c9f9dcf5a3e3547af31

Observation d1ef7dc2-1c70-416c-b6ba-d3313644d97f · outbound

This paper cites P.; Li, I.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models P.; Li, I

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.448381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.188027Z digest=sha256:cb74c969fc0d1e72fc300f14acee1e9a470b96517aa8792d4cf47fe2ef07f5a3

Observation 6dd683d1-f39d-4b20-99fa-4265d52939cd · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.434061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.192599Z digest=sha256:f969b1c04cc1ee7a4d16c8ff722b890d50df78d891a699dec9f89402f0168638

Observation 4646c320-7b55-4656-8de8-15a77f2df0a6 · outbound

This paper cites Q.; Zhang, P.; Chen, J.; Pramanick, S.; Gao, D.; Wang, A.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Q.; Zhang, P.; Chen, J.; Pramanick, S.; Gao, D.; Wang, A

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.419202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.197423Z digest=sha256:cc7d0d3addf701e49233f1b0a9879d22b90b50eab169b1857df8b2ea1eaf6bf8

Observation 15cc9d1f-c670-492b-946e-75cea26f841f · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.403672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.201320Z digest=sha256:568eb3f9eebe80b4085b5d5913b8de42ffa0d914e2c6a45e582f441997f2aaf8

Observation d4d1c1a7-0dbc-41a4-af50-184eff5b7e51 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.386471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.205355Z digest=sha256:0268fbbd5d78f3515f7897004db19eac68fd4459a93945dc3a541376b2380efb

Observation 78f3c9e6-614a-4baa-bba7-9ce3c7d50168 · outbound

This paper cites Visual Instruction Tuning.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.209489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.209489Z digest=sha256:2c6e6bf32b68ee9d6f5e1880408c165de54d32a30b1552988e61fa4f9f775f62

Observation 26db8e28-8088-4bc1-86a5-b1dbed7fe11d · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.370695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.213492Z digest=sha256:24bfc03395cdfe81dbdcea2ec46b09d834fe1aaefd402cf029299c4458bf7064

Observation c9898231-9276-4cce-ac0d-11391c98d873 · outbound

This paper cites A Decade's Battle on Dataset Bias: Are We There Yet?.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models A Decade's Battle on Dataset Bias: Are We There Yet?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.217303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.217303Z digest=sha256:3b6dfbd2bf5451906b510a6cbf46c5e95d96e3d52bc3abb913b79a6b02e22f5c

Observation 2851c0cd-cf46-4625-a8ec-cce3431106d4 · outbound

This paper cites Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:34:12.632794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.225371Z digest=sha256:d3013b706d5b075ed142ea06711abf5dd3bcb8482728921ca509a72a06e87805

Observation a18c8819-7e3d-466c-8eda-2fe5167fbee2 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.356075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.229853Z digest=sha256:9575b54e2e244e1857c6bd63d1603505ee8c94336ed8ac57ac2fc6f99d9474fa

Observation de6bfcff-74cf-47a4-964d-a951a54c7990 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.340938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.233983Z digest=sha256:1c00fcdbaa89ee22ba2fb4713686d86b4b63ace94b8c3b40611df77b0fccfff2

Observation 23494253-3ef8-48da-8657-7cd744aad1f5 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.238406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.238406Z digest=sha256:af066b3e964ebbccc1f1075c5f98c2846cb49c80ff9fa8f4ef90ffd8451c4fce

Observation 2447d950-2d0a-439d-80cc-938121226d20 · outbound

This paper cites EchoPrompt: Instructing the Model to Rephrase Queries for Improved In-context Learning.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models EchoPrompt: Instructing the Model to Rephrase Queries for Improved In-context Learning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:34:12.598679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.242702Z digest=sha256:d9e07b0cfdaa5a2bda8f267b9e05574ec63a9de52c75ea49156b5cfe3b04d1fa

Observation 16503e9a-1446-43ca-8596-b72eaaeee85d · outbound

This paper cites J.; and Choi, J.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models J.; and Choi, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.326167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.247269Z digest=sha256:6057992dea7aed28a4c5ae046737f9e050a6167ab12556b4c4cfb6976c6b59b0

Observation 2f05673b-b2a0-435c-a483-ddfa0591363e · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.311258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.252417Z digest=sha256:26e9d705ac43f87dcbef25fdd186abcdb6e0a3e390c89075e6561aeb2d2566f3

Observation 4cc89b0c-f4cb-4ddc-9ebd-f9e1be9e3014 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.256532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.256532Z digest=sha256:1719fedbeceeedf0bf7f52dcee678b410b39d80bca44d5f6adff030180d31eef

Observation edcc05da-973e-48c1-878c-7eebc40cc71d · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; and Clark, J.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; and Clark, J

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.286779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.260172Z digest=sha256:15d1feffec1c240b0a528402e9982a48515ec168eb3be533c4461223ed5131cd

Observation 1124415d-0c4b-421a-a124-df65e9d5acbc · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.264313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.264313Z digest=sha256:711d32d979a473a0fa3b4719be921104e3d21dfa5b5f1aeacecef8ebcc5d7a71

Observation 64041b5a-36fc-49bd-ba5d-ddef00691532 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.272366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.268853Z digest=sha256:5d219f942aa4ac719546de20e6044d9bf3f023c93b85f03f3ab937a62a4ca1f3

Observation 6904b105-b3df-446a-a7e6-7c6644003076 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.258861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.273729Z digest=sha256:03b7f053ecc707f75d3d462ce42a1215698316290718cd3d98029423dad7e415

Observation d7b3be0f-b2aa-44df-acbe-4c35fd017854 · outbound

This paper cites A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; and Gupta, A.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; and Gupta, A

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.245005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.277668Z digest=sha256:899ef0bddc007e7effd9bf5c72af691a1e0fdcc998e9d87508c87f13ece928b7

Observation b6227838-ef75-4d96-85c1-c5ab9192aeac · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.230476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.281819Z digest=sha256:00133a7810981fe6452e0bbc51dbe3a32d7cbc37c68f71bf2d90146e67c548cd

Observation 85c11104-1f9a-48c2-99be-a9f8211a461f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.285898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.285898Z digest=sha256:d5faee2bd5bbc5ff7d34b07643f944f905d14626a1e7b515415a93dc0e825320

Observation 77d8f6fa-b86b-4c5d-b16b-eb4f4c8db116 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.215687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.290110Z digest=sha256:5f65346dc7c64ea9fbd29c25ee9e06169196691a4dfcdb64e76c93f33aa9eab2

Observation 3a3acdbf-0840-45e7-9e6a-ed3fd685817b · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.294304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.294304Z digest=sha256:b77c11fe94e320cf480cb41fe679c3003875dce671866f61e134c7d8920d6164

Observation 98dac2c9-f977-490e-9d23-9f728e17f3f6 · outbound

This paper cites I.; Shekhar, S.; Döllner, J.; and Trapp, M.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models I.; Shekhar, S.; Döllner, J.; and Trapp, M

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:34:13.200998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.299217Z digest=sha256:b3e8384023abd0cd922fa1a543ff74d9795471059e9e32ec218a814dacb570aa

Observation b5f4950f-96b1-40fa-a02f-dcfcb035865e · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.187209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.304047Z digest=sha256:1ea54676eed1ab794404085d98dd26fb91c1a179fa73dad0b9bed5c7c086456f

Observation b94c84d3-0ad5-4d41-a48c-53ca696762b6 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.173735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.308515Z digest=sha256:ce3a066e4fee2208de47b7f62be1c4ba4a0387543fe606793e7118cf68e2d550

Observation 24aa0b82-dbdb-4180-acbc-fe1b5b4018c6 · outbound

This paper cites Verbalized Machine Learning: Revisiting Machine Learning with Language Models.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Verbalized Machine Learning: Revisiting Machine Learning with Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.312879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.312879Z digest=sha256:e5949eda0ab645e6f6862196af260ffcaac9ee27be07755cece675ddf32e3542

Observation 518124bb-2ea9-4932-9a99-69a269d6c651 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.317250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.317250Z digest=sha256:1f0504d4f73c6bd48c05f9211ba1ba433c46d9ade0952e2c0a69f25ddf7974a0

Observation eea02911-4dcc-4ddc-9a8c-bd32712ba69f · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.160053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.322375Z digest=sha256:4f7c632b5f7f7595161271e95970910069156f0a8c8c8823d1285f651329af2b

Observation eb513a00-d938-4e95-a1f6-32ce3112111c · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.146154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.326799Z digest=sha256:69654b8d6ab8e80ab1263bb4dfce6681eea577311667724d21ff3e73928f98a4

Observation fb4765d7-7ccf-448e-94e7-52c2f2e3ee16 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.131776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.330949Z digest=sha256:16c51349634c9794813482a2c1fd2774138e10c8c507f28f217376c9c010f3d5

Observation affab108-7f76-47c3-94bf-f5635e06cc81 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.116393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.335292Z digest=sha256:858ed45c68b07f1456c20f5e455a5b5a388e899f6d7ec7bd36346e691986cc5e

Observation 23ca5529-515a-4211-8570-aa374c9c0571 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.101752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.339630Z digest=sha256:8892e5b3268d433f8ff7b5d764dcbb11070669fa810aaca976f20bcb4d2bac3e

Observation 5cd92ac0-2726-47ae-a61a-68860602a4e8 · outbound

This paper cites MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human Motions.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human Motions

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:34:12.505823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.343810Z digest=sha256:1de99e59404578d099b9ce442cc544e4e4c037bef3d4a61d9d04a001a977418f

Observation bbca617f-da0c-41cd-9517-d928f45d922c · outbound

This paper cites Deconfounded Video Moment Retrieval with Causal Intervention.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Deconfounded Video Moment Retrieval with Causal Intervention

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:34:12.482895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.348085Z digest=sha256:7e0ac7a9105b8012945d4e694ae9cc0047ec3d1b2febf576cb2febb8a0565432

Observation d427661b-f2b3-4ea9-86d1-74dc420080bf · outbound

This paper cites SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.352313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.352313Z digest=sha256:95dc9d7031b195a650f17fe957cc9440926591fafefe6a28c1eb59e588da8a9b

Observation 0c9c4cca-bfce-4f33-91ac-e5c0250cd23e · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.356485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.356485Z digest=sha256:c0f0c0ba4245aa642ff6f796a086b90440e4d08d34f7db748ee118ec97bde009

Observation 13f1873e-5f20-41ee-8448-5442652b193d · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.087847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.360568Z digest=sha256:34b7bf1c76effb7a3534ac7a9193189c7232a6080793301d9b635f386a31ddd1

Observation b5009508-9f94-4d48-9ad9-a154b051ce08 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.364724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.364724Z digest=sha256:24114374a2565ce6abdae4546a56bd77264bd7a7971bf147e9aead01ae32fce9

Observation dc36f000-4dae-40c2-828d-14639f1d93f6 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.073445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.368702Z digest=sha256:4981695742b9fff6977b6be3dde4511dc94b8a945f9a864cb7ecece430837c66

Observation 19ace2d6-af39-4b42-8e56-d321ea147a70 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.059507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.372736Z digest=sha256:2f74117c023aa8dccca5b29693fda4985b942776fcb7bbf7726c1702ccca130d

Observation e03d85be-bf0b-49e7-820a-1643cae88297 · outbound

This paper cites an unresolved cited work.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:34:13.044474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:34:12.376826Z digest=sha256:d12de59a8593f8b1a224119cba066fec89f66443893478095f64df1968a5194c

Observation f29df06d-4659-48c8-8685-279ab39d136b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:12.380805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:12.380805Z digest=sha256:3337e27184bd8533f2057722aa52d1203517e6a2fb3dde074c4cffc55720abb2

Pith citing papers

Observation 5a7677e6-bf89-4d80-83a9-4868eded77a7 · inbound

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning cites this paper.

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.644297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-08T08:04:15.238840Z digest=sha256:f89ae62bd9359d1fe22f7867fb107e0a604b8726c84ac263c064adfbcdb25c77