Pith. sign in

Paper Citation Record · LEDGER

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

As of 11 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.16701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16701 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:21.811848Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b72c9351-9ef0-465e-b48e-f1ce3130d5c8 · outbound

This paper cites Action genome: Actions as compositions of spatio- temporal scene graphs.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Action genome: Actions as compositions of spatio- temporal scene graphs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.520530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:19.034906Z digest=sha256:da201f85ae643e59fdb400cddf59a33e0952e7d42b1596f8fbcc3d2f52a8b0d3

Observation 33a8dfb9-c426-4331-85c2-628f14f2c224 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition OPT: Open Pre-trained Transformer Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:19.144627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:19.144627Z digest=sha256:01b428fb51db2f829326037e11bdde73014cd8629c4c021e88f8186891aee8b5

Observation f8c50662-bf37-4081-9d4d-80586bc74098 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:19.256942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:19.256942Z digest=sha256:6df39e7cd289545cd831d7198550ce5bd21290efb774187b663c62c5eff7d266

Observation 78e5ac41-9a02-42fd-9a94-117b65d18d08 · outbound

This paper cites Chatgpt-4o, 2024.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Chatgpt-4o, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.508520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:19.370286Z digest=sha256:8b087b08da208049fd054dfc08f88439f6ba0481b8e66d46137fa76c08b7453a

Observation 20c30f0a-6199-46e8-b882-0cbeac9c7b31 · outbound

This paper cites Gemini 2.0 flash, 2024.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Gemini 2.0 flash, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.497191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:19.505819Z digest=sha256:d774419947a708743d9e60eeb00bafc14fbb455d911f4f8596c448dc2fcc976d

Observation a01dcf5b-733c-4f47-ad8c-b86ba066d818 · outbound

This paper cites Qwen2.5- vl technical report, 2025.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Qwen2.5- vl technical report, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.484668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:19.665627Z digest=sha256:d265ef060f91ccccfe3047192033e9bb394c268bdeae1d271756a7ae2eb0aea7

Observation 373aa4ee-a0f3-4b4f-af5b-ad4fc7eb3c53 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompting visual-language models for efficient video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.465022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:19.780660Z digest=sha256:c84eb59f6fc5a6f9dbe0b02bc73c4cf7513b2f41f39f6fdd4b9d5fe49822410b

Observation 239205f9-1874-448e-a453-9ab9238b2366 · outbound

This paper cites Llms are good action recognizers.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Llms are good action recognizers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.451449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:19.851409Z digest=sha256:b1cc54a69c54cefcffa165951ae72cd57aed21e91e3ce9e041998faadbb7505e

Observation 04db58d5-6d67-424b-bbdd-7b8eff7218e1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.430934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.001362Z digest=sha256:92b56e63a55d872a177e38f61b85769f0d9019ccfcb9e228f3de2dba051ededb

Observation d9c25835-2ced-4e4d-8a4c-7fafab3cb869 · outbound

This paper cites M-llm based video frame selec- tion for efficient video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition M-llm based video frame selec- tion for efficient video understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.418534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.132309Z digest=sha256:8ca479ea1c9876c706e002165e39eacf7e60bee19e4b7e412cd309b58b2ca420

Observation 0db42020-bc5d-4b84-bfe1-195443b6a256 · outbound

This paper cites HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:20.273161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:20.273161Z digest=sha256:c247ed04625333327de2f14ff99aa1929ddc32a9966727efcf53cf17ac2abf00

Observation cb149e37-2dad-424f-9e59-0b8fafce6447 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Two-stream con- volutional networks for action recognition in videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.385943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.354728Z digest=sha256:22562761c53a5eeed902602bda2ed6bedaca75944a8374359da3b453cd1039b2

Observation 259051f0-9365-45ca-b031-1b99d5ba3968 · outbound

This paper cites Temporal segment networks for action recognition in videos.IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740– 2755, 2019.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Temporal segment networks for action recognition in videos.IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740– 2755, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.374140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.427751Z digest=sha256:5e40b0cabe5b84a5550c04ef1c8b3807bd4e70b6fddb511d78bf223e7b4460d2

Observation 284c16f2-3140-4f03-9820-0cd072e8dd07 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning spatiotemporal features with 3d convolutional networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.360639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.486930Z digest=sha256:7532b73436ba059cbb63bcd1e195a13b0e2d1562dc86f2ebf469057ab74d2d66

Observation b1b3f256-2266-4cea-bdc4-10175834fbc6 · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Quo vadis, action recognition? A new model and the kinetics dataset

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.347775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.570357Z digest=sha256:9787067496c3e0e61ae2b348bf69210d4ce1a019c73d41ed2ba64542199f9742

Observation 786b4f0b-98ff-44c7-b9f0-b03b7d76232e · outbound

This paper cites Smeulders.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Smeulders

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.335524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.634188Z digest=sha256:8f226bac5be1ea57c789b3e0c5336028c04dd77dd15f868d5353c605002dcd74

Observation 53831e79-1ec1-4d5f-b53d-b1ce24f67bbb · outbound

This paper cites Slowfast networks for video recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Slowfast networks for video recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.322239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.732330Z digest=sha256:581feb15adac00ee04669f53578625c7117ca527786e9146009943b1fd41f218

Observation 64d244a2-55d8-420c-8cd6-4fdb5da3b246 · outbound

This paper cites an unresolved cited work.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:42:22.308568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.811147Z digest=sha256:b090c378bb0c376c1588d7c72301a66a5470ba141d23158291d2ebf855dc4fc6

Observation 44aab4ff-dcef-4ac7-9517-6380c2b044cc · outbound

This paper cites Tokenlearner: Adaptive space- time tokenization for videos.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Tokenlearner: Adaptive space- time tokenization for videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.293076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:20.821125Z digest=sha256:5b1e370bf46cc1850e2dbe22a76321a1d57379e2d3cc080035c03403d78bf849

Observation c3ec7672-fe3d-4eb1-97aa-ee2ba973dbd7 · outbound

This paper cites ConceptNet 5.5: An Open Multilingual Graph of General Knowledge.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ConceptNet 5.5: An Open Multilingual Graph of General Knowledge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:20.983666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:20.983666Z digest=sha256:b32b737004f06528a6821280e8afa540f979cb58c5cb1d314b2968f2a2f33578

Observation 0a4e7caf-2a1f-4f04-87ff-69fc0c041152 · outbound

This paper cites ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.078067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.078067Z digest=sha256:fd7b4444dee6a9538b3ec71b757429a00fd4576e82be1268a1fef5d2436647b8

Observation 6e20bce3-48c2-4f9a-a53b-ae7be4ed1c08 · outbound

This paper cites We- bChild 2.0 : Fine-grained commonsense knowledge distilla- tion.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition We- bChild 2.0 : Fine-grained commonsense knowledge distilla- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.271807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.192655Z digest=sha256:67659437b5c50c8c00bde9f3f6f74a11213bea7599a51a6e8864f4de9b7163a9

Observation 990f1d3b-7972-4649-bcdb-6f1381f3ae86 · outbound

This paper cites COMET: Com- monsense transformers for automatic knowledge graph con- struction.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition COMET: Com- monsense transformers for automatic knowledge graph con- struction

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.233158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.287096Z digest=sha256:0b08c02a45db5bc9ab706d1fd0cd96458994bcda80676005bcce123d6131ffee

Observation 2c46a1d6-4e5e-4fe1-b9bc-faeb99f209f8 · outbound

This paper cites Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.213346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.306892Z digest=sha256:241a1852b6b68aa3168b0b12cb4c2d0bf19ca7f4f1cb7b3735805fc28cbaf46f

Observation eb5cccf0-a671-4502-a497-e36f674092a3 · outbound

This paper cites Conditional prompt learning for vision-language models.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Conditional prompt learning for vision-language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.194019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.444572Z digest=sha256:7829f68c6d166c6d16ba0b46be207b1746d11320afe8caa9040bfad3f2714d1e

Observation 6f5971ea-0443-4a1d-95db-f1e2e582866b · outbound

This paper cites Prompt distribution learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Prompt distribution learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.173041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.544753Z digest=sha256:d6f4657f64e8e83c6eb1ba823978d0c0c2addd4ea41b97f680e31cdec4331816

Observation 7b6802c2-d0dc-4f3b-a310-e6ab3adf4e31 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning transferable visual models from natural language supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.574755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.574755Z digest=sha256:9f6077e865cf919775cae182027e2aaa2212ed769a8bc9eee8b06b07ca73f84a

Observation b9fe3d8f-922b-4af2-a320-e0060a24cc92 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Denseclip: Language-guided dense prediction with context- aware prompting

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.151063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.633977Z digest=sha256:69eaf78a36af2ae67502b15c144b819233615906d92be417b5ac72d72e57c720

Observation 58ea54c6-2266-46b1-8fb9-d96decd84584 · outbound

This paper cites Expanding language-image pretrained models for general video recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Expanding language-image pretrained models for general video recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.137255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.742005Z digest=sha256:45578a2f444ee59fd1dbd368ce68bdf75a7c364081fbfea304c50b73c417ed19

Observation e16178ab-908b-460c-8c71-61bebae55ae0 · outbound

This paper cites Learning to prompt for continual learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning to prompt for continual learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.122393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.745915Z digest=sha256:0e5f0241185de8c401e69f5c32801eefae106c587fad5afa8c54215980d36170

Observation e60e9bb0-1d7c-461e-a67e-361fec39616a · outbound

This paper cites Visual prompt tuning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual prompt tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.109828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.749643Z digest=sha256:a3cacdcec08cb1e9c1e1777b8768a9bf70ff0cbaf400ff483885c9c6103f11eb

Observation 626eb1b9-c04f-45cf-84c3-95f5616aca36 · outbound

This paper cites Videobert: A joint model for video and language representation learning.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Videobert: A joint model for video and language representation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.096230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.756955Z digest=sha256:1feb2b0f9667de034f1fbca5c11d98f74bd9e7d431fd44e2c1276845a4cbbd48

Observation f7615d3c-ec66-474b-b69c-708751fcd5ac · outbound

This paper cites Language models are few- shot learners.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Language models are few- shot learners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.075878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.763188Z digest=sha256:4d6525094eb477a0e4be4fa3014ffcda0575f2d14b7e90504b0d395e644258ba

Observation edc83934-8674-4252-8dd8-09a0bd6a8aa6 · outbound

This paper cites Le, and Christo- pher D.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Le, and Christo- pher D

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.054885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.768055Z digest=sha256:bc6c1e104c82b8157e9f6881492fb771969f8aa3b138aab55c69b5c748baeb2c

Observation ad77e2ab-4f84-44c4-9a81-51ecbd5c80a6 · outbound

This paper cites Multimodal few-shot learn- ing with frozen language models.Proc.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Multimodal few-shot learn- ing with frozen language models.Proc

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.042786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.772487Z digest=sha256:596ed3414814e16278b6ad1435cde33930470024951a7e4339d031d0441694f1

Observation 79406379-5f3b-4823-aa3b-eb5e72351b8e · outbound

This paper cites UNIFIEDQA: Crossing format boundaries with a single QA system.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition UNIFIEDQA: Crossing format boundaries with a single QA system

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.019630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.777249Z digest=sha256:17c4efec6d80d8edb9c96a6b4fecbafdc6b78d00d4acbaf4f0ef2700394d760a

Observation e08c08b5-1488-4201-afdf-32b8fcefbaaf · outbound

This paper cites Few-shot text generation with natural language instructions.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Few-shot text generation with natural language instructions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:22.005297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.781863Z digest=sha256:0f852c76025e2e1a0a467c0cb9b88cdd9fe4dd8bffadc3b9ed6b4e79dfbbf46f

Observation 67d8d725-f251-4fa0-acce-d060c304a619 · outbound

This paper cites Generating action-conditioned prompts for open-vocabulary video action recognition.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Generating action-conditioned prompts for open-vocabulary video action recognition

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.987818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.786252Z digest=sha256:e4e6a4e5673ffe1f9744b2a168601eb115baebec89ee913833b33f4cf8d3e739

Observation 2e59d9d7-4bd7-43b6-9229-22b0bd086f9c · outbound

This paper cites Kronecker mask and interpretive prompts are language-action video learners.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Kronecker mask and interpretive prompts are language-action video learners

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.973922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.790912Z digest=sha256:7fe78622bf36d42bf8313b5a3d6dc88364bad3744f4d384d5f1e6f7968aa1969

Observation ad2c47ed-e007-499d-bad7-4a94f141ceae · outbound

This paper cites Visual semantic role labeling for video understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Visual semantic role labeling for video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.960619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.794614Z digest=sha256:53371c2f916beb9b324d1b6b5a3941b27310edd92491af6e6c47c46e4a7561f1

Observation 8d9974a3-f7ec-40c7-b23f-fb6a181a81b6 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Learning Transferable Visual Models From Natural Language Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.799058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.799058Z digest=sha256:784094a5abb7715e9ab2bf9a42a8f95acb8528cfb93489688f8c6fea59453f6a

Observation cb55a17d-2ff0-4de0-9d46-7fc584a684b9 · outbound

This paper cites Flava: A foundational language and vision alignment model.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Flava: A foundational language and vision alignment model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:42:21.947281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:42:21.807905Z digest=sha256:198ad71de64a5e59d6cb321a8a5479c10062833365b99b6b682cd932c32b50b6

Observation 83e83f57-78f2-4534-b8fb-0af7dffdfdea · outbound

This paper cites Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding.

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:21.811848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:21.811848Z digest=sha256:26f4b187713cc9e40387d98752fc53baac0e2d5a6aa16150d092a411086f2fdf

Pith citing papers

No inbound Pith citation observations are available.