Pith. sign in

Paper Citation Record · LEDGER

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos

As of 18 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2508.15903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15903 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:45:25.588833Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30546734-549b-421e-9c22-975274f512e2 · outbound

This paper cites Human action recognition from various data modalities: A review,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Human action recognition from various data modalities: A review,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:29.123621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:23.060592Z digest=sha256:9f40645c81939e201361875c7ac7f7efb628e230e8310fa5622fa04ab054b7e3

Observation d252faa4-4aab-465e-b185-618752afe075 · outbound

This paper cites 2d versus 3d convolutional spiking neural networks trained with unsupervised STDP for human ac- tion recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos 2d versus 3d convolutional spiking neural networks trained with unsupervised STDP for human ac- tion recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.949894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:23.146665Z digest=sha256:4f5615804c6b4d8d5ee6945ff591a82a4e8affb5b7ee4dddbaae2d79a0dd4fc2

Observation e50b3ee1-68e1-40bf-8e19-38eda428f0c7 · outbound

This paper cites Overview of the transformer-based models for NLP tasks,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Overview of the transformer-based models for NLP tasks,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.789963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:23.228078Z digest=sha256:d98df1b9cdc256d5f29074489a8acea971f47923dc68249f5d71c2cf98d624cb

Observation 75379562-ca53-4ae5-9323-5bcc340240f1 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:23.362858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:23.362858Z digest=sha256:bf3ffacb0b6db86ec8e2bbaa83c42b4625bf487e69b9079bbad0b9a463bd0626

Observation 393296a3-aa62-407b-9c96-27bfb523ed99 · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.625634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:23.471682Z digest=sha256:800bf11557070483bf3455c40e7e052c16b2b4f77d14f4f42e0f8145a9050ce3

Observation b93cb2b3-2c7e-46dc-a229-4e7af91113e1 · outbound

This paper cites Visual in-context learning for large vision-language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Visual in-context learning for large vision-language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:23.600178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:23.600178Z digest=sha256:5af47c635b0470d6fc38ffff3c9110bebbedc13a3c3fcc67a6db1b616ec40809

Observation 9f7e4b3b-4da5-4477-ab48-b442500aef50 · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Weak to strong generalization for large language models with multi-capabilities,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:23.702577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:23.702577Z digest=sha256:489afaf1c552fc15ec90134c84ec3e6d478add2a130454395957ccc08b253d95

Observation c2c5cd88-7f4b-4abf-a9a9-6be46a81d246 · outbound

This paper cites Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.462300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:23.829216Z digest=sha256:92f8017a7e53c93de6f365af40d22311e40946f7b6c96a9279ef4cd38fec64c3

Observation beb377fc-8139-40d5-8afe-405dedf492f2 · outbound

This paper cites Adaptive prompt: Unlocking the power of visual prompt tuning,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Adaptive prompt: Unlocking the power of visual prompt tuning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.243612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:23.958129Z digest=sha256:2587edcdb11b9478fcc7f73a8181da5418a2514e29c215c2c5abd39cda85a1ec

Observation 4bdffa8b-4257-4b3c-b07d-b184751a1123 · outbound

This paper cites NTU RGB+D: A large scale dataset for 3d human activity analysis,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos NTU RGB+D: A large scale dataset for 3d human activity analysis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.046819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:24.084139Z digest=sha256:075dd17a9156d71d23df3380728a61d4ab964f8c2a6d8db1c6516502338105f0

Observation 4e6dc61f-30be-4c19-9a88-f3b716458e0a · outbound

This paper cites Human action recognition and prediction: A survey,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Human action recognition and prediction: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.168894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.168894Z digest=sha256:149640870c830383d92d36b414c42091350b2e2d876fe5309534dbd172c08ff3

Observation 2361002e-23bb-49df-8111-06a8d05a507e · outbound

This paper cites A comprehensive study of deep video action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A comprehensive study of deep video action recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.884575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:24.257880Z digest=sha256:74ad9676856c268068eddf1b5216de29c15c7f8da5c5fcb5910ebbed10c6fee9

Observation 79808f44-f637-4276-a7fe-b7fa2f194032 · outbound

This paper cites Action recognition based on efficient deep feature learning in the spatio-temporal domain,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Action recognition based on efficient deep feature learning in the spatio-temporal domain,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.720535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:24.356694Z digest=sha256:2b56ba6edf27e22cd2f9574bbd1179ff0c5ce5f1ea7763b62b3eb2e55b16a36e

Observation e7ebe177-7957-4a59-a98a-6e0c5f04b272 · outbound

This paper cites Cross-fiber spatial-temporal co-enhanced networks for video action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Cross-fiber spatial-temporal co-enhanced networks for video action recognition,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.437028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.437028Z digest=sha256:5c138a7bb9ded62fa3d0d6eb8d75f7e574f13f5860005ca987db2604c3606694

Observation d38b2284-ef73-454c-8d50-24c3e7f36bde · outbound

This paper cites Mutually reinforced spatio-temporal convolutional tube for human action recognition.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Mutually reinforced spatio-temporal convolutional tube for human action recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.563094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.563094Z digest=sha256:ce8f4ba3abb94aa6ff1fba4319bb8a4a1a771b8de3e3f8da839e2342411e1b5c

Observation 1b2ad58e-545b-4a36-8881-17ec5cec0e07 · outbound

This paper cites Multi-scale spatial- temporal integration convolutional tube for human action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Multi-scale spatial- temporal integration convolutional tube for human action recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.625203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.625203Z digest=sha256:4ddea4e82da02f34535b51ff86adbfb035e330787089dc143ec01ce67b364b61

Observation 5a6e302f-a6d6-43c0-a56b-c4ee621fe105 · outbound

This paper cites Long-term temporal convolutions for action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Long-term temporal convolutions for action recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.583506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:24.707195Z digest=sha256:b7b435d332f5bab2ec9113638f124a27ca06e34cc730c9a3515b9cb320891c47

Observation ec2b4758-7c2f-4a5c-9952-0cd638e80b82 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Finegym: A hierarchical video dataset for fine-grained action understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.375395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:24.810896Z digest=sha256:d7f3f90bbd511c9b34fddcc1ec214215dc5a0cbce5e5b5d9d84675c3ee3ed666

Observation 6cc55d62-1891-4051-8279-ad907b4c8366 · outbound

This paper cites End-to-end video-level representation learning for action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos End-to-end video-level representation learning for action recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.153608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:24.875343Z digest=sha256:f7bf1355d82a9c7ba1506a01711bae5e45cfa200406b1830875a7746ef0e8d6a

Observation 9ba7246f-d83e-48b7-9661-d24dd9c51378 · outbound

This paper cites Stnet: Local and global spatial-temporal modeling for action recogni- tion,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Stnet: Local and global spatial-temporal modeling for action recogni- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.955176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:24.943847Z digest=sha256:55892f6b10d750ec6f4a8aa515ede15197584576f6b9e5db088447dea2d09eba

Observation f36d86e0-9f0b-4436-b838-d6922f39f59e · outbound

This paper cites A survey on efficient vision-language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A survey on efficient vision-language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.739960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:25.030702Z digest=sha256:982f269043befa18f83cc924f0d70418065b913bba41690186fe5540114fc5ef

Observation 60923db0-71b2-4be5-a736-e99c2b677598 · outbound

This paper cites Cheap and quick: Efficient vision-language instruction tuning for large language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Cheap and quick: Efficient vision-language instruction tuning for large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.578959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:25.141410Z digest=sha256:0f8964d039414d4c8d63fe41737268be7cde50bbb6ff0e1b46b75fadf2eef360

Observation 5ea32719-ee71-46ea-b40b-b0638e01085c · outbound

This paper cites PEFT A2Z: parameter-efficient fine-tuning survey for large language and vision models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos PEFT A2Z: parameter-efficient fine-tuning survey for large language and vision models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.404939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:25.221123Z digest=sha256:5628868f29ff50ef2a14f6a885d35f8c69a4abc90fc10e0e3d3c95ca88a694f2

Observation fa2bd55c-8085-4894-a159-46f8e5e017c8 · outbound

This paper cites Visualgpt: Data- efficient adaptation of pretrained language models for image captioning,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Visualgpt: Data- efficient adaptation of pretrained language models for image captioning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.204200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:25.317266Z digest=sha256:8796cc3d63a3bf54ed0e5f64658146efa1e8f3c5b1fa643df95d36146a31589d

Observation ed8435e0-1800-4210-b87d-c91cf362d305 · outbound

This paper cites Multi-modal large language models are effective vision learners,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Multi-modal large language models are effective vision learners,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:25.997782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:25.381519Z digest=sha256:c6ca447e30063ebd968f96dde23c067af618ee2acbed394d469bd38f44e20ed6

Observation 1280096c-bfe2-4c67-938d-9ca493b7f1bb · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:25.477587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:25.477587Z digest=sha256:a8461e791cc5e374a8ee2c132ea26e19a5a9810d2bc78ecba9ac86c5b92b1d97

Observation 0e97a3f1-20f4-44d3-a72d-793938ec7988 · outbound

This paper cites Aligngpt: Multi-modal large language models with adaptive alignment capability,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Aligngpt: Multi-modal large language models with adaptive alignment capability,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:25.836109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T17:45:25.588833Z digest=sha256:19664fff9034d3258d5bf5bf7d496edb2128b8dc25aaefbe491f58b38d526a72

Pith citing papers

No inbound Pith citation observations are available.