Pith. sign in

Paper Citation Record · LEDGER

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.18675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18675 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:16:15.771212Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c906201-ce95-4b06-8361-e5bf22231b03 · outbound

This paper cites Leveraging Vision -Language Models for Improving Domain Generalization in Image Classification.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Leveraging Vision -Language Models for Improving Domain Generalization in Image Classification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.363661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.602955Z digest=sha256:6364840f035f70335f9d21251194133fb419cc8abbb4e39fd734b9457ca55068

Observation b42e8a9c-025b-42ac-a052-645e77104a56 · outbound

This paper cites Are Visual-Language Models Effective Action Recognition? A Comparative Study.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Are Visual-Language Models Effective Action Recognition? A Comparative Study

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.348253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.608943Z digest=sha256:25217c8693df5ae2142e863f1b00e75eab612aa669965101436be0af6a35fb9f

Observation 5091d515-2eca-4f6b-b38b-2a0d3f6841d2 · outbound

This paper cites Vision–language model for visual question answering in medical imagery.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision–language model for visual question answering in medical imagery

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:16:16.333516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.613602Z digest=sha256:4b8a7d11acc5f51fc9d6b2eca4a185d09649babd4bf7b238a4a3f3d876b616ba

Observation c64017ab-cd13-4b9a-a5f9-3f23b1f9ff1f · outbound

This paper cites When deep learners change their mind: Learning dynamics for active learning.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks When deep learners change their mind: Learning dynamics for active learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.317730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.618379Z digest=sha256:8896da66515fda322b3437b74567f80cbffde11baf90b1f939c16725a1a4c35f

Observation 7116ff25-bad7-432e-913f-b73209cc6692 · outbound

This paper cites An Introduction to Vision-Language Modeling.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks An Introduction to Vision-Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.623253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.623253Z digest=sha256:57d23fc0bc264c1135817f5b108510e20371bfbecd379b0364af4b8a3616de96

Observation 5a18d4b3-1509-478e-b4d8-5816046489a7 · outbound

This paper cites PracticalDG: Perturbation Distillation on Vision- Language Models for Hybrid Domain Generalization.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PracticalDG: Perturbation Distillation on Vision- Language Models for Hybrid Domain Generalization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.301352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.628962Z digest=sha256:dd0f2d036dfa6eb07d641ccdc8b8c796ecda77a0097bc5d511a2fd8f89dab248

Observation 4cb39969-b1da-4a9e-a20f-d0374c62e8fe · outbound

This paper cites Pub - medclip: How much does clip benefit visual question answering in the medical domain?.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Pub - medclip: How much does clip benefit visual question answering in the medical domain?

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.284660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.634590Z digest=sha256:9adb4a21e4a428e7c80c412d9ff4e45863dd112dc5df9da3e41a30a779becb3d

Observation a01ce183-c5e8-4d54-9b26-c00771e8b294 · outbound

This paper cites Clipsyntel: clip and llm synergy for multi - modal question summarization in healthcare.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Clipsyntel: clip and llm synergy for multi - modal question summarization in healthcare

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.268728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.639482Z digest=sha256:ecc1d538c502d1d5994c8860fd4c65f56bd0594db8b8ee5b8b48a8af8bcd6c25

Observation 8599834e-4962-45b3-afce-5fc894760d96 · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in action recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learn2augment: learning to composite videos for data augmentation in action recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.253854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.643907Z digest=sha256:a15ac98e2327a86b85746a1e96273cb851251609048587a56becac6c3b5a6aec

Observation 99c5f218-63d8-4483-a43b-e727a90ef1ea · outbound

This paper cites Class-Specific Noise Injection for Improved Road Segmentation.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Class-Specific Noise Injection for Improved Road Segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.239299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.648059Z digest=sha256:d307b3bd88e8d84f26bce3f00abe8d1ae0f3c71d77276c04c4b36909fbb8208c

Observation 9c393a4b-7fef-4d1f-aaf6-e1540c4185f7 · outbound

This paper cites Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.223687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.652509Z digest=sha256:ee38654d5b2b7bfacc5390d7c8f40035072ec5b9e2dbe908260c234a415f499b

Observation e05f7981-d6fa-474f-97f9-3828fb363c00 · outbound

This paper cites Perturbation- based methods for explaining deep neural networks: A survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Perturbation- based methods for explaining deep neural networks: A survey

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.207039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.657146Z digest=sha256:a5fed40d179887b433bb521ba6c668c6e871266751302534f7a3cacf0fdbaee9

Observation a1a36eec-d4d2-4eaf-b7f6-39708a0b1ad7 · outbound

This paper cites A dversarial attack on yolov5 for traffic and road sign detection.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks A dversarial attack on yolov5 for traffic and road sign detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.192083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.661596Z digest=sha256:933f64ee28eefb16cf50c61fcaf4966181bd337cbed917e538cf8f63f6febaf3

Observation 2814c903-1eea-4b62-8ec2-a39b84838e12 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.176666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.666027Z digest=sha256:b653685fe012e387ae3223e95d6e1c688f32ab4c5585e2e65b59f42f8b51674e

Observation e99755b9-7c87-4a02-8f9c-c7eebdc838d3 · outbound

This paper cites Language augmentation in clip for improved anatomy detection on multi-modal medical images.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Language augmentation in clip for improved anatomy detection on multi-modal medical images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.161183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.670475Z digest=sha256:c2c595e994b8cb70362ecfb816005bcce272dd05694e099e4e9489705e09f9ec

Observation c0a9211b-c188-4843-a68b-8cbbf12409c7 · outbound

This paper cites PALM: Predicting Actions through Language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PALM: Predicting Actions through Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:16:15.971825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.675732Z digest=sha256:0bfdc7c183070f3ae68792b08839b16c18e974ef1efb9d5299e92638b1149ad6

Observation 2a182daa-20e5-425d-a840-eb420f81bad6 · outbound

This paper cites Segment Anything.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Segment Anything

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.680772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.680772Z digest=sha256:18abee5ab9c4fa10a52f6142883ef1ae4e4ed5d064ca9f10b1c4754ce0d8f842

Observation 6e691084-371f-4eee-9828-7132922ec243 · outbound

This paper cites Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.685735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.685735Z digest=sha256:42c8085e94279d3ac4a69d7da993c6b234082837d05d31f37f7e78a31a94f55b

Observation 57251939-f85f-4b06-908c-fd9499ee11db · outbound

This paper cites Enhancing clip with gpt-4: Harness- ing visual descriptions as prompts.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Enhancing clip with gpt-4: Harness- ing visual descriptions as prompts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.145623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.690697Z digest=sha256:0e990a1911319e9155f9ff3b86260dc8cc38d106f30924679b2d3b822960573c

Observation da2b2fb5-34e3-45a8-afb0-7628970b5995 · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learning transferable visual models from nat- ural language supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.130369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.695366Z digest=sha256:ddf7ca36a7f2a8ead09f49f541e9ef62b0e8e024a92497a6d7e764b409d0ed57

Observation 7f484ea5-da3b-4100-8470-099fa43f2cb0 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.699872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.699872Z digest=sha256:c03f316d71ca9f9452a142d9d1a1a234920f75d945f9f262f6f7cddeee212d48

Observation 331a88c7-9971-4d18-8149-35d44671dac7 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.704561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.704561Z digest=sha256:12fc7126ca7bfe85542899e15a7320f595b06037b3dbc6c5d0da915da382c820

Observation ddf7199f-94ae-4eed-86a3-d5d6845f63a9 · outbound

This paper cites Test-time prompt tuning for zero-shot general- ization in vision-language models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Test-time prompt tuning for zero-shot general- ization in vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.115214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.709612Z digest=sha256:ce1adfb817c8d771965a5eba27f709e0952091bf2d9e170e2c9085b36770f96c

Observation bf7884b7-6150-40d6-b059-befcade4fa8c · outbound

This paper cites Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.713836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.713836Z digest=sha256:aff64ca551efdf5613d1a597b62bcc8f84d26798aa3f2b7e0c34a074cf3d8ed3

Observation b738c268-606b-4352-b927-1ba254d6912f · outbound

This paper cites Motionclip: Exposing human motion genera - tion to clip space.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Motionclip: Exposing human motion genera - tion to clip space

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.100219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.718612Z digest=sha256:1ba1760fb2efe6b92537d5ef2c81907f31810325cffb62c46bb44ae8765135a4

Observation 0b0f2520-2a16-4ca4-b37d-d030f0fcd319 · outbound

This paper cites XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.723157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.723157Z digest=sha256:834601751562a62a6a3d39c7e8491fa074f5c5f28b8be5e22e565dbce61630af

Observation e3929c2c-f003-455f-b096-1c14574783e1 · outbound

This paper cites CLIP with Quality Captions: A Strong Pretraining for Vision Tasks.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP with Quality Captions: A Strong Pretraining for Vision Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.728021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.728021Z digest=sha256:b4b75f73f9902951a4a6f30fbc01535dfa1ff6d75f1c475e98a0de36cc8a2ddf

Observation 315d1def-45df-424e-974c-f6743b83bf3c · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks ActionCLIP: A New Paradigm for Video Action Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.732959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.732959Z digest=sha256:a047c6ad64e813678379da13b30ba4f42c293ba04c573493afbd25b5f05c4946

Observation 8dbdee37-4579-415f-b08d-0d3b788052dc · outbound

This paper cites Actionclip: Adapting language-image pre- trained models for video action recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Actionclip: Adapting language-image pre- trained models for video action recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.084863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.737740Z digest=sha256:7b04070e90d2e7474f5e4653fa83e14e8bd3876b3727609e1d7962e0f8536373

Observation d57a36ad-9822-466a-87d4-44028a193620 · outbound

This paper cites Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.067933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.742143Z digest=sha256:a493d1729ecf8921e9bb4ec6ed1abf2fe37a49be8d2c0714d9a17ed4577f90da

Observation 088b7585-9ecb-44ad-bdeb-4998e8802493 · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.746661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.746661Z digest=sha256:4828d38f422a3f19460046bb64d93379eb3e0a92424003d39412ecc384801727

Observation ba183e20-0401-4f89-9bb3-aaceb3f15e20 · outbound

This paper cites Revisiting classi - fier: Transferring vision-language models for video recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Revisiting classi - fier: Transferring vision-language models for video recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.051759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.751782Z digest=sha256:5b8a4bc95f536049ceca5edb07df8df42c38b51603496a36de683c494b400dfa

Observation e8df1d71-a006-40a7-aed9-07546bda48d5 · outbound

This paper cites Investigating Compositional Challenges in Vision- Language Models for Visual Grounding.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Investigating Compositional Challenges in Vision- Language Models for Visual Grounding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.036514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.756446Z digest=sha256:a09683d3bc3e125bd1641bd7f543c69564a7c0b6438854b3f7227a17b1a4036c

Observation 9bd7a2d6-847d-41af-9280-a3edbb1c6d6b · outbound

This paper cites PeVL : Pose -Enhanced Vision-Language Model for Fine-Grained Human Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PeVL : Pose -Enhanced Vision-Language Model for Fine-Grained Human Action Recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.020935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.760727Z digest=sha256:6c87171b00d342d864cd5766a616608f2393e90e39b701040817368ef9645572

Observation 88089a6f-6aff-4ae9-978b-d96fc226bd36 · outbound

This paper cites Vision -language models for vision tasks: A survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision -language models for vision tasks: A survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.004907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.765144Z digest=sha256:55428c81682ed8f0ccd602972912eafd107963330b5790ca0b5c90f8326382e7

Observation c828b911-0368-4390-9f3a-ef08f5229124 · outbound

This paper cites CLIP in Medical Imaging: A Survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP in Medical Imaging: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.771212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.771212Z digest=sha256:f6d1436c3d50ec517b0acb95f9edbcbc8d092db1f375dfa087f6d87199b1e1ca

Pith citing papers

No inbound Pith citation observations are available.