Pith. sign in

Paper Citation Record · LEDGER

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2508.17442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17442 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:04.858493Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f1d62b3-e4ff-44e3-88cc-74fe63ac909f · outbound

This paper cites Human action recognition and prediction: A survey,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Human action recognition and prediction: A survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.377306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:02.710999Z digest=sha256:247d12eb533a0506a031461b30d2968d6cf52c15990bff84ec4cd49eba86a89f

Observation ca917620-26f2-4f5a-a9ef-15b7c3fb264e · outbound

This paper cites How would autonomous vehicles behave in real-world crash scenarios?.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition How would autonomous vehicles behave in real-world crash scenarios?

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.365851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:02.768294Z digest=sha256:edde259fe60ac29efed546bf56338009955f54981ba3462ef83af2fed5cec8f2

Observation f23be0e3-7c48-4726-b026-9ec4087861db · outbound

This paper cites Crash-based safety testing of autonomous vehicles: Insights from generating safety- critical scenarios based on in-depth crash data,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Crash-based safety testing of autonomous vehicles: Insights from generating safety- critical scenarios based on in-depth crash data,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.355542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:02.849541Z digest=sha256:71a7c129680fc59d39beb1b401ac739eb24b407959056fbf5a96be290e1cd6a8

Observation 3c4d22bc-9847-4c6b-be31-85d70a0292ae · outbound

This paper cites Diffcrash: Leveraging denoising diffusion probabilistic models to expand high- risk testing scenarios using in-depth crash data,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Diffcrash: Leveraging denoising diffusion probabilistic models to expand high- risk testing scenarios using in-depth crash data,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.181785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:02.930306Z digest=sha256:58b04ba75369e0cec57c5424038c35dc3a049e9fc69186a10a54ccb5b8c0c087

Observation 6bf11373-c719-4f82-8435-1b6698d8172d · outbound

This paper cites Search-to-crash: Generating safety-critical scenarios from in-depth crash data for testing autonomous vehicles,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Search-to-crash: Generating safety-critical scenarios from in-depth crash data for testing autonomous vehicles,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.901226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.044456Z digest=sha256:9ff2997ac95093e4a84c877361a69920482256dbd1579997d535e50884cfb868

Observation b35bfb34-040c-4576-8ba5-94b3c409ed6b · outbound

This paper cites Visual features of interme- diate complexity and their use in classification,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Visual features of interme- diate complexity and their use in classification,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.705485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.140644Z digest=sha256:e220edea664e018fa0aef78a0acdde1e2644a6824e206dc159d95b2fe09001f8

Observation 0bf434c0-2c9b-4dc3-bd96-29e13993419b · outbound

This paper cites Visual in-context learning for large vision-language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Visual in-context learning for large vision-language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:03.247944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:03.247944Z digest=sha256:2b3f2411b77cb29d2f41b81a73868e85ab234fdb18bf7bc0d90fea0e0afcf119

Observation 9211e064-77c1-4010-bfdf-c90ba7e07438 · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Weak to strong generalization for large language models with multi-capabilities,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.527606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.335362Z digest=sha256:c6eb7d4c5facf6c20038b46204a0316bffff9f43379118e6525397224a05b632

Observation 224aaed7-67a0-4e16-a210-5ce0c1e3997f · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:03.423729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:03.423729Z digest=sha256:6eb71476e88e9c6493f8d09e76ed903bcdc1fbe5797b4bac342f5797b99e14f3

Observation d4a7da62-e621-486a-9fa3-16dc0b86d2cd · outbound

This paper cites Cbr-net: Cascade boundary refinement network for action detection: Submission to activitynet challenge 2020 (task 1),.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cbr-net: Cascade boundary refinement network for action detection: Submission to activitynet challenge 2020 (task 1),

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.342940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.530428Z digest=sha256:fb349518e5af8c0812272d7715c3c39d625ab23e0d8088b43ce404b3858b7a4e

Observation 03899956-4933-490a-a4c7-629d929cd2a9 · outbound

This paper cites Fineaction: A fine- grained video dataset for temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Fineaction: A fine- grained video dataset for temporal action localization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.146218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.602521Z digest=sha256:b3f9e5a47d8ec2ef36a0c4356cbf821fe172a51eb92d1a1bb04336e054e60546

Observation ef41a983-1680-4f33-b553-8d6fb8d8fe5b · outbound

This paper cites The THUMOS challenge on action recognition for videos.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition The THUMOS challenge on action recognition for videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.940537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.718104Z digest=sha256:2d2c822242512b32ab4833530ae665ac285c58d6265bec58582358ce3b27df55

Observation a7b9a5fe-129e-4185-9f99-899b1726cfc1 · outbound

This paper cites A survey on temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A survey on temporal action localization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.725629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.778064Z digest=sha256:755092a22a83e1a797e31fb3490d2a20e23e7ce5a78df4998e337f45523f4ef1

Observation 9382a4f2-34cd-4395-8309-5206f9e61901 · outbound

This paper cites Tallformer: Temporal action localization with a long-memory transformer,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Tallformer: Temporal action localization with a long-memory transformer,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.524109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.844680Z digest=sha256:b3b6cd56e7116e3e91535e7a8eca7d538ea18d336bef50ec5411087c27d0832a

Observation f3f02e0a-06f0-4cdf-b677-c39fa59aa05a · outbound

This paper cites Cross-fiber spatial-temporal co-enhanced networks for video action recognition,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cross-fiber spatial-temporal co-enhanced networks for video action recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.294764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.899195Z digest=sha256:3126a23fead5540df752bb98729441d53a01aad2403e4918f88126d4df178a47

Observation ac0b2a75-226b-45d7-b09d-38dcca55f0db · outbound

This paper cites Mutually reinforced spatio-temporal convolutional tube for human action recognition.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Mutually reinforced spatio-temporal convolutional tube for human action recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.192111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:03.980769Z digest=sha256:a6417a58a7c8c4cba961e51aa7927294d1cda922c725d9f9b92d92e4114219aa

Observation 02d844f6-21ee-476a-8b3b-42c23c94aa4b · outbound

This paper cites Multi-scale spatial- temporal integration convolutional tube for human action recognition,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Multi-scale spatial- temporal integration convolutional tube for human action recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.050514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.066995Z digest=sha256:aa41a1566c385919fdb9f797fd3bb3932d7bb948fba7be8a63abb2cfbde39f0f

Observation 567dc6b1-2132-407c-9605-a08dc09a600b · outbound

This paper cites Cross-video contextual knowledge exploration and exploitation for ambiguity reduction in weakly supervised temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cross-video contextual knowledge exploration and exploitation for ambiguity reduction in weakly supervised temporal action localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.901413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.094400Z digest=sha256:5fb9300bd2c9ed4987dc4f4b29499635e6e148e5b5fc4451f25fa31ac0385654

Observation 0f799125-56a3-4d42-bd85-2a1197b630a1 · outbound

This paper cites Trigger is not sufficient: Exploiting frame-aware knowledge for implicit event argument extraction,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Trigger is not sufficient: Exploiting frame-aware knowledge for implicit event argument extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.793739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.157516Z digest=sha256:84f5596b65c0addbacab83d0006e2db382da461d4c1166dc1f154ee39f3c885b

Observation d4beb4ae-e0ac-4bd4-80e5-35d32144f7e8 · outbound

This paper cites Guide the many-to-one assignment: Open informa- tion extraction via iou-aware optimal transport,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Guide the many-to-one assignment: Open informa- tion extraction via iou-aware optimal transport,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.618967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.215955Z digest=sha256:97c16db65f137bfac6037f502b7d1621ee223ee34b087ebb0dae3e68804a6981

Observation 53ef3b59-ce1f-4abe-b2b6-42c2904fb1cb · outbound

This paper cites Video activity localisation with uncertainties in temporal boundary,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Video activity localisation with uncertainties in temporal boundary,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.438970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.236450Z digest=sha256:1aac0e19caf83ec7c953613e6bf7f0fb3918ac23fb9f1506e6cfe50b4e9b48a0

Observation ea17fdfa-a670-4471-85df-f46710cca323 · outbound

This paper cites Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.310617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.290223Z digest=sha256:15d1e75df796c48bb4cfe30faa78039f3509f0e2be4ff056953ea42736377a7a

Observation 551a893b-fd60-4b3e-a6f2-1f062e4f7392 · outbound

This paper cites How vision-language tasks benefit from large pre-trained models: A survey,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition How vision-language tasks benefit from large pre-trained models: A survey,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.167646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.314652Z digest=sha256:d92556f6307481768b79ec95538110d2f9165fc73220707fbd164111e17cb96d

Observation ad1ced01-72c5-4aaf-8cad-4b5b40f9fa1e · outbound

This paper cites From linguistic giants to sensory maestros: A survey on cross-modal reasoning with large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition From linguistic giants to sensory maestros: A survey on cross-modal reasoning with large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.993465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.373024Z digest=sha256:54ade1fecce5058ed535c735c4a56567532fca97f21a495694507c74c793487b

Observation f94019dc-ab73-426b-a526-9fd2d4d125db · outbound

This paper cites A systematic survey of prompt engineering on vision-language foundation models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A systematic survey of prompt engineering on vision-language foundation models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.845189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.457379Z digest=sha256:fbceec1036122ef12b771cb61bbeea882bf88c1cacb521a468b789a33da8c2e2

Observation 950d90b7-9a6d-4b03-923f-f2ae9d0f4cc6 · outbound

This paper cites Evaluating multimodal vision- language model prompting strategies for visual question answering in road scene understanding,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Evaluating multimodal vision- language model prompting strategies for visual question answering in road scene understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.669528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.548139Z digest=sha256:9d9f53f9a1bf2fe2c00bb1e470e9cde0b338a2393445b7d1061581fe6eee5880

Observation 4b5364e7-7739-4df5-9ba8-15d560a207f0 · outbound

This paper cites Towards grounded visual spatial reasoning in multi-modal vision language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Towards grounded visual spatial reasoning in multi-modal vision language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.509382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.608059Z digest=sha256:da106cd425a31e1385f0eae56dac656ce4894a78aa3a95bd5e02df4030419359

Observation 51bb8a7c-3808-432c-b3e3-95078126c0b1 · outbound

This paper cites Enhancing advanced visual reasoning ability of large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Enhancing advanced visual reasoning ability of large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.335787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.641263Z digest=sha256:983ed31a8d9fc7702ebeaaa8141257bf1252414e6fb1a9ead53f540fd1c1096c

Observation 3357097d-1759-4009-8866-21e72b58ae99 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:04.691360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:04.691360Z digest=sha256:efe0c647163b6f50162994c8d436b97feec9da5799ffaf748e316b26598d000d

Observation ce6576ed-c6cb-49b3-8ca9-b10ef9db3a17 · outbound

This paper cites Chain-of-specificity: Enhancing task-specific constraint adherence in large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Chain-of-specificity: Enhancing task-specific constraint adherence in large language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.221014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.745657Z digest=sha256:f578d495433aaa938dcceb447f870455c9638243f0d6f2028845fb158e480977

Observation f4f25b48-e5c4-4bb2-bae9-b8d309399ca6 · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.054392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:55:04.858493Z digest=sha256:92dec415001f3ceb2ecb67510c58aa0e3997ff474d58853ae7f3001ab871c105

Pith citing papers

No inbound Pith citation observations are available.