Pith. sign in

Paper Citation Record · LEDGER

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering

As of 19 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2502.09573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09573 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:03:13.329200Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de2ca75f-24bb-4753-b803-a9624a78f9df · outbound

This paper cites GPT-4 Technical Report.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.196965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.196965Z digest=sha256:12bd4b8435d2c1259fd74ee3e1ebdb7fa431915a413d3b75a73f034df6b66b58

Observation 1624cbe1-9f77-4ade-bdfd-9f68932708a3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.203362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.203362Z digest=sha256:ed907e6925dbff5664ac51ad82404c7bd54ba4aea9bb8fc04381326225030e7c

Observation 8ddb77f4-17e0-45ba-b278-4f49f0224a65 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.209644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.209644Z digest=sha256:8241cecd3091c50bdee30ed901fa03863663f5b7b6405ff2230028fa352939fa

Observation b3f95510-baa5-443b-9ed1-d0c4c216c43a · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.216250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.216250Z digest=sha256:11d7b443c38779ea28f8bf8d85ecfae543109dfad11445cd7e5eb6d318921665

Observation 67d1ec17-df35-401b-844b-1cc8ae526129 · outbound

This paper cites A Multimodal CNN-based Tool to Censure Inappropriate Video Scenes.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering A Multimodal CNN-based Tool to Censure Inappropriate Video Scenes

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T21:03:17.669138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.222137Z digest=sha256:c84dfb1c3d01f1f809cf622538d2339f9d36c245fe371e5fc3dd61e79e019837

Observation 17e20f24-31b8-4b5c-8da7-44d4284838d7 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.228207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.228207Z digest=sha256:528500bf340fe2cdd6ad2836b30bc46e81283ade93e024fac4be7837a261439a

Observation db718e0c-6c29-4e3f-9505-7265c13a891c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.234882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.234882Z digest=sha256:475dd7aed6424fd08d8628df010714a442166822bd9f7f72370492e7bb14c8f4

Observation 789aac87-32e9-4c96-934a-497f7fa53db5 · outbound

This paper cites Deep residual learning for image recognition.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Deep residual learning for image recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.241015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.241015Z digest=sha256:2afba250346955c8c675b3415e6e388991a6cd103576f38131e1d17a1b77114c

Observation 090a0efb-7a03-4b0e-89fe-23ca011b8ccd · outbound

This paper cites H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.246481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.246481Z digest=sha256:665cd067f8544628fc1188d868e0c428cfd22270bc37dccd89b253400f00aaa7

Observation face6fb8-7606-4124-98d4-77ae618e30fa · outbound

This paper cites an unresolved cited work.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.251781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.251781Z digest=sha256:d1b40e5d4ebb9a2d4c461f4ecaa1fa8357bd47dccd1e0c6cb389fc6b523561d8

Observation a0bc587c-2b28-4d79-ad1a-0651955a31b0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.257128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.257128Z digest=sha256:5f7b1e2a881b8aff425ded593e96035de78302cbc1e0745dcab34994539f7d50

Observation c85518da-c75f-43c9-ab7e-5cb99c1ad88f · outbound

This paper cites Openai official api documentation, 2025.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Openai official api documentation, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.873073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.263148Z digest=sha256:41a97f316378a03bcbc8bf0893f2119ade34657bd29180700afef294a9bb6b65

Observation f802c8c8-efde-455f-a883-d7b3fb7cf566 · outbound

This paper cites J., Mahmud, A., Sobuj, M.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering J., Mahmud, A., Sobuj, M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.852117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.268323Z digest=sha256:7cd81a4402769e46538ffa9707a639cef681851dc56c6226de2f08156332fc3d

Observation 45ade3e0-9fd1-444f-aeac-847b27b071e8 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.273301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.273301Z digest=sha256:a37119a142b769c9dfd67de0ba76e02b5fd4f3a73cd77f5d0248ce5c4883b762

Observation 8c7641b8-503c-4d55-ab5d-42edf16a08e7 · outbound

This paper cites Will the \ 1 trillion of generative ai investment pay off?, 2024.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Will the \ 1 trillion of generative ai investment pay off?, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.798891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.278477Z digest=sha256:08db94e70ea4a16fc54c3a1c8587a4d5ad47388138993f176177a1d19b05c538

Observation 631c7893-09fa-4324-a0d3-12c3c8e99927 · outbound

This paper cites Video understanding with large language models: A survey.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Video understanding with large language models: A survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.283982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.283982Z digest=sha256:85180cd247e942c298d5fd720323541272f21676a4c4cfd52932570aa6d24787

Observation c6fdfcdf-b3b3-4dc0-9a83-71a1ed37e48a · outbound

This paper cites PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.289380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.289380Z digest=sha256:1b26d51f2129d5a92e197dd513d9619b0399303391e2e10fc0dc531f644197f8

Observation be6885a3-3740-4ae4-bd0e-13f3421c0db6 · outbound

This paper cites Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T21:03:13.409766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.295499Z digest=sha256:41af8cf6008e1bad6c3a3bd51db890302b8d8ef1fcfba213e663b7315963c5fa

Observation ff8f4093-4c2c-4c11-8553-c46bc9b15223 · outbound

This paper cites and Nawaz, T.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering and Nawaz, T

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.780330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.300703Z digest=sha256:9db9d6bea646d9098b734228723108c535a85230c01eb57d91a5711822429d8b

Observation 43e1c9d8-a0fa-4e48-878b-a258b5620dc3 · outbound

This paper cites Fine-tuning Large Language Models for Domain-specific Machine Translation.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Fine-tuning Large Language Models for Domain-specific Machine Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.306086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.306086Z digest=sha256:31da5987c9e4c8f3dbc6fe8d19563c2594bfa395edb13cd547bb45e98eff9cbb

Observation 64a935d5-5239-436a-aaaa-507bab5120fd · outbound

This paper cites write newline.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering write newline

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.311663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.311663Z digest=sha256:3f520e0c8aecb89404a5bf7e6270dd14964a10210cdd64db54cbdb658f104bb3

Observation 56d84102-f72a-4a7e-b68f-f0a97fef3345 · outbound

This paper cites @esa (Ref.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering @esa (Ref

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.318053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.318053Z digest=sha256:b85d890d0b560b3367c5f34aaf39907519026fd16e55347c52f072d11adf6e97

Observation 955c383f-d17d-4541-a83a-83e4be449b2a · outbound

This paper cites an unresolved cited work.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.323649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.323649Z digest=sha256:41bea1fdb5abc92ca9ad2ac4089a986a21d3df218fbafbd19ffdc10d08384bd0

Observation 2f342397-317a-445a-b690-75ab64be3803 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering DriveLM: Driving with Graph Visual Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.329200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.329200Z digest=sha256:0c446060580c8836a76c9433c10b1ebfa3f7bb26862b95daba277938b14741d5

Pith citing papers

No inbound Pith citation observations are available.