Pith. sign in

Paper Citation Record · LEDGER

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2508.17442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17442 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:04.858493Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f1d62b3-e4ff-44e3-88cc-74fe63ac909f · outbound

This paper cites Human action recognition and prediction: A survey,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Human action recognition and prediction: A survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.377306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:02.710999Z digest=sha256:a432a1d67d389a583ff6ff4c2392a5cf2bca7ef43707c6eb1b5f3101976602b1

Observation ca917620-26f2-4f5a-a9ef-15b7c3fb264e · outbound

This paper cites How would autonomous vehicles behave in real-world crash scenarios?.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition How would autonomous vehicles behave in real-world crash scenarios?

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.365851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:02.768294Z digest=sha256:e88ad502ad06ffb9b69d7abfbc166a6dcc3f4fdaa1d090c98740fa9722af58dc

Observation f23be0e3-7c48-4726-b026-9ec4087861db · outbound

This paper cites Crash-based safety testing of autonomous vehicles: Insights from generating safety- critical scenarios based on in-depth crash data,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Crash-based safety testing of autonomous vehicles: Insights from generating safety- critical scenarios based on in-depth crash data,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.355542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:02.849541Z digest=sha256:60755e738a71f2fb059e65465dcba1980ef8898335b9c4a2bbe62b0cf3651891

Observation 3c4d22bc-9847-4c6b-be31-85d70a0292ae · outbound

This paper cites Diffcrash: Leveraging denoising diffusion probabilistic models to expand high- risk testing scenarios using in-depth crash data,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Diffcrash: Leveraging denoising diffusion probabilistic models to expand high- risk testing scenarios using in-depth crash data,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.181785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:02.930306Z digest=sha256:de61b4db5eb94e7b4b4ac71eb5189adaedab573e1d569daf9dc48ff9e6bbe94a

Observation 6bf11373-c719-4f82-8435-1b6698d8172d · outbound

This paper cites Search-to-crash: Generating safety-critical scenarios from in-depth crash data for testing autonomous vehicles,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Search-to-crash: Generating safety-critical scenarios from in-depth crash data for testing autonomous vehicles,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.901226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.044456Z digest=sha256:7249b165f931448af62659f3f32fa585b5deafc133b981c89abfe51415485c43

Observation b35bfb34-040c-4576-8ba5-94b3c409ed6b · outbound

This paper cites Visual features of interme- diate complexity and their use in classification,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Visual features of interme- diate complexity and their use in classification,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.705485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.140644Z digest=sha256:76790c2c482f10eb043dc38fb5e328edfe5d3ab98671b3d20c5d99cd16fc5fe9

Observation 0bf434c0-2c9b-4dc3-bd96-29e13993419b · outbound

This paper cites Visual in-context learning for large vision-language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Visual in-context learning for large vision-language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:03.247944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:03.247944Z digest=sha256:ea1d2a8f992e5fa1ea554870e642c04e62feab255fedaaaebee5bee0b7c6b707

Observation 9211e064-77c1-4010-bfdf-c90ba7e07438 · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Weak to strong generalization for large language models with multi-capabilities,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.527606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.335362Z digest=sha256:b9530e9cb1902dbfaa0adf506390a66148ef88b2b06bfca85ec3a025d70bcb35

Observation 224aaed7-67a0-4e16-a210-5ce0c1e3997f · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:03.423729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:03.423729Z digest=sha256:11de086988e923364dc37643dd749e32b86cf01e43869b773daf9a9a55592e79

Observation d4a7da62-e621-486a-9fa3-16dc0b86d2cd · outbound

This paper cites Cbr-net: Cascade boundary refinement network for action detection: Submission to activitynet challenge 2020 (task 1),.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cbr-net: Cascade boundary refinement network for action detection: Submission to activitynet challenge 2020 (task 1),

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.342940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.530428Z digest=sha256:aff73764fcf9b5f11d5cfcc872d2024e67da5dcc6e000f203321ff570a3e8589

Observation 03899956-4933-490a-a4c7-629d929cd2a9 · outbound

This paper cites Fineaction: A fine- grained video dataset for temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Fineaction: A fine- grained video dataset for temporal action localization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.146218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.602521Z digest=sha256:a9fa2d60c9a63cd55cd1d5e722a3c49d2c303a90d9fd430ee275e27538133a21

Observation ef41a983-1680-4f33-b553-8d6fb8d8fe5b · outbound

This paper cites The THUMOS challenge on action recognition for videos.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition The THUMOS challenge on action recognition for videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.940537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.718104Z digest=sha256:a751c2de8a63aaf1f9d4533cf4844c5f15c8a81d1f792a32f227d098f8ccb065

Observation a7b9a5fe-129e-4185-9f99-899b1726cfc1 · outbound

This paper cites A survey on temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A survey on temporal action localization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.725629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.778064Z digest=sha256:4ac02670f63cd54aed534449d0a9e59d5386d12a9c331198bea825315d502c48

Observation 9382a4f2-34cd-4395-8309-5206f9e61901 · outbound

This paper cites Tallformer: Temporal action localization with a long-memory transformer,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Tallformer: Temporal action localization with a long-memory transformer,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.524109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.844680Z digest=sha256:e0cc226502ffed7a51682f2dc0dfb08900bfab68e592b708850a7ed785ad90ee

Observation f3f02e0a-06f0-4cdf-b677-c39fa59aa05a · outbound

This paper cites Cross-fiber spatial-temporal co-enhanced networks for video action recognition,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cross-fiber spatial-temporal co-enhanced networks for video action recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.294764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.899195Z digest=sha256:d3a7c1b37268bc60437a47e2b5ae6582434733bb440b0144dd5875d608d122d2

Observation ac0b2a75-226b-45d7-b09d-38dcca55f0db · outbound

This paper cites Mutually reinforced spatio-temporal convolutional tube for human action recognition.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Mutually reinforced spatio-temporal convolutional tube for human action recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.192111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:03.980769Z digest=sha256:b90d06cb71550c3fe96274f071c178044a7b661b21bfc20d5baea9dc8e3cbe8a

Observation 02d844f6-21ee-476a-8b3b-42c23c94aa4b · outbound

This paper cites Multi-scale spatial- temporal integration convolutional tube for human action recognition,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Multi-scale spatial- temporal integration convolutional tube for human action recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.050514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.066995Z digest=sha256:fda421d1ae7a400fffdacadce9e78bb7a363ba9b2243cf9a7d49c0b7a48ed876

Observation 567dc6b1-2132-407c-9605-a08dc09a600b · outbound

This paper cites Cross-video contextual knowledge exploration and exploitation for ambiguity reduction in weakly supervised temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cross-video contextual knowledge exploration and exploitation for ambiguity reduction in weakly supervised temporal action localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.901413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.094400Z digest=sha256:52863dba53ff4d99ca917367ce7048fc9bb9ac4ab1af6a613d8fbd9dc8adfa83

Observation 0f799125-56a3-4d42-bd85-2a1197b630a1 · outbound

This paper cites Trigger is not sufficient: Exploiting frame-aware knowledge for implicit event argument extraction,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Trigger is not sufficient: Exploiting frame-aware knowledge for implicit event argument extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.793739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.157516Z digest=sha256:1488e13412bf094da91bb07006d098f9626f34441f238ed17728ec9b3ff3ef52

Observation d4beb4ae-e0ac-4bd4-80e5-35d32144f7e8 · outbound

This paper cites Guide the many-to-one assignment: Open informa- tion extraction via iou-aware optimal transport,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Guide the many-to-one assignment: Open informa- tion extraction via iou-aware optimal transport,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.618967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.215955Z digest=sha256:867045dcf6819472652c335d49a69ab835d2303e7591a8a53f8d2f211754eace

Observation 53ef3b59-ce1f-4abe-b2b6-42c2904fb1cb · outbound

This paper cites Video activity localisation with uncertainties in temporal boundary,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Video activity localisation with uncertainties in temporal boundary,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.438970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.236450Z digest=sha256:b174747a512c22bc33cff96db4870aa08b895d88f830d075f255454a34b5e54d

Observation ea17fdfa-a670-4471-85df-f46710cca323 · outbound

This paper cites Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.310617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.290223Z digest=sha256:0f3ba81f31eaa242132f53bdfbf7979f316ea20734ffeb44c7eabe49666f3741

Observation 551a893b-fd60-4b3e-a6f2-1f062e4f7392 · outbound

This paper cites How vision-language tasks benefit from large pre-trained models: A survey,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition How vision-language tasks benefit from large pre-trained models: A survey,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.167646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.314652Z digest=sha256:7479ab4b3cd75e63eac44a5a07b7a74fcd02776514fdebbceb4548cbd22217bd

Observation ad1ced01-72c5-4aaf-8cad-4b5b40f9fa1e · outbound

This paper cites From linguistic giants to sensory maestros: A survey on cross-modal reasoning with large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition From linguistic giants to sensory maestros: A survey on cross-modal reasoning with large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.993465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.373024Z digest=sha256:53c48ede8b756ef59888d46d231efc87a8ca6ba558527db5a21ef6efd0c70ab1

Observation f94019dc-ab73-426b-a526-9fd2d4d125db · outbound

This paper cites A systematic survey of prompt engineering on vision-language foundation models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A systematic survey of prompt engineering on vision-language foundation models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.845189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.457379Z digest=sha256:3d8184950433cafb69f9c6453ffd053506009838a2fae7e4605018326bc663da

Observation 950d90b7-9a6d-4b03-923f-f2ae9d0f4cc6 · outbound

This paper cites Evaluating multimodal vision- language model prompting strategies for visual question answering in road scene understanding,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Evaluating multimodal vision- language model prompting strategies for visual question answering in road scene understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.669528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.548139Z digest=sha256:acec94344df1c70c27b88a66b39632b39e0a32298230d5753804feb8e31ec632

Observation 4b5364e7-7739-4df5-9ba8-15d560a207f0 · outbound

This paper cites Towards grounded visual spatial reasoning in multi-modal vision language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Towards grounded visual spatial reasoning in multi-modal vision language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.509382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.608059Z digest=sha256:8e94149f90632311af85cfdf56c5d4b38c629052a1ada992a40141fd7ded0de9

Observation 51bb8a7c-3808-432c-b3e3-95078126c0b1 · outbound

This paper cites Enhancing advanced visual reasoning ability of large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Enhancing advanced visual reasoning ability of large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.335787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.641263Z digest=sha256:2b9e362a47874cb86bc4706d8c0a9257cd751dbc94e50094d8b1847249b465be

Observation 3357097d-1759-4009-8866-21e72b58ae99 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:04.691360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:04.691360Z digest=sha256:47cae16086fbb6656d9734fbe5b61e7f4693f9f9f1999e0eb7cc854d730a469d

Observation ce6576ed-c6cb-49b3-8ca9-b10ef9db3a17 · outbound

This paper cites Chain-of-specificity: Enhancing task-specific constraint adherence in large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Chain-of-specificity: Enhancing task-specific constraint adherence in large language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.221014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.745657Z digest=sha256:e0de01bfc0d7d8fcccb6df3e2482b3cfb5645097283a51bb09e19d534727fcfb

Observation f4f25b48-e5c4-4bb2-bae9-b8d309399ca6 · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.054392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T16:55:04.858493Z digest=sha256:f20c7f18c79c3f2def933619b9b82837a9d74021c615837eecf08441a8252003

Pith citing papers

No inbound Pith citation observations are available.