Pith. sign in

Paper Citation Record · LEDGER

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks

As of 19 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.18675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18675 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:16:15.771212Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c906201-ce95-4b06-8361-e5bf22231b03 · outbound

This paper cites Leveraging Vision -Language Models for Improving Domain Generalization in Image Classification.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Leveraging Vision -Language Models for Improving Domain Generalization in Image Classification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.363661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.602955Z digest=sha256:332925d7a18eef417a208e8ad67fe986246e14a85db168d6a3eca2980a0dacdb

Observation b42e8a9c-025b-42ac-a052-645e77104a56 · outbound

This paper cites Are Visual-Language Models Effective Action Recognition? A Comparative Study.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Are Visual-Language Models Effective Action Recognition? A Comparative Study

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.348253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.608943Z digest=sha256:1db67e23c8e3d033b8dd87b4c81f34458db57b9d808736e19c76e8a725d6392e

Observation 5091d515-2eca-4f6b-b38b-2a0d3f6841d2 · outbound

This paper cites Vision–language model for visual question answering in medical imagery.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision–language model for visual question answering in medical imagery

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:16:16.333516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.613602Z digest=sha256:28305a5bebab9bb9b30fb6305dbb1808813bb9872e5e56fd319ff5e29afbbd36

Observation c64017ab-cd13-4b9a-a5f9-3f23b1f9ff1f · outbound

This paper cites When deep learners change their mind: Learning dynamics for active learning.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks When deep learners change their mind: Learning dynamics for active learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.317730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.618379Z digest=sha256:53a2260833e1badacb48597b431fa5a8b059dae30103ca8f5ae32ec75334b6c2

Observation 7116ff25-bad7-432e-913f-b73209cc6692 · outbound

This paper cites An Introduction to Vision-Language Modeling.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks An Introduction to Vision-Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.623253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.623253Z digest=sha256:7a1af999f8ecae39085523d65d4a66f88a3abda50f9ac5e53d44a99fd311387a

Observation 5a18d4b3-1509-478e-b4d8-5816046489a7 · outbound

This paper cites PracticalDG: Perturbation Distillation on Vision- Language Models for Hybrid Domain Generalization.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PracticalDG: Perturbation Distillation on Vision- Language Models for Hybrid Domain Generalization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.301352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.628962Z digest=sha256:01ff0d0a6feb5628ca37c5c1e9932d39dc2b31f239bb39250db4450ac4fd7f04

Observation 4cb39969-b1da-4a9e-a20f-d0374c62e8fe · outbound

This paper cites Pub - medclip: How much does clip benefit visual question answering in the medical domain?.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Pub - medclip: How much does clip benefit visual question answering in the medical domain?

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.284660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.634590Z digest=sha256:88c3c865711d21546f7231ad69905ff0ccad4bcdfc5ce0315facda9a2f0ad4b0

Observation a01ce183-c5e8-4d54-9b26-c00771e8b294 · outbound

This paper cites Clipsyntel: clip and llm synergy for multi - modal question summarization in healthcare.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Clipsyntel: clip and llm synergy for multi - modal question summarization in healthcare

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.268728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.639482Z digest=sha256:3b1a44d3e640bfc80480f630a5c11fcbaf9ff81eac4b79ebf9776866f048e392

Observation 8599834e-4962-45b3-afce-5fc894760d96 · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in action recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learn2augment: learning to composite videos for data augmentation in action recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.253854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.643907Z digest=sha256:b9c536d498935b8defbea672c2e8a28d85f11aff8f17606f3545a5183a737139

Observation 99c5f218-63d8-4483-a43b-e727a90ef1ea · outbound

This paper cites Class-Specific Noise Injection for Improved Road Segmentation.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Class-Specific Noise Injection for Improved Road Segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.239299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.648059Z digest=sha256:85757d105c9c2635a681bd2a16788fef214407da93cfa51b13a0e9498f22949d

Observation 9c393a4b-7fef-4d1f-aaf6-e1540c4185f7 · outbound

This paper cites Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.223687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.652509Z digest=sha256:a7cf7b3e9070c61f5934bcd1844ab6776d93549d39c49d9437189a572a577f47

Observation e05f7981-d6fa-474f-97f9-3828fb363c00 · outbound

This paper cites Perturbation- based methods for explaining deep neural networks: A survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Perturbation- based methods for explaining deep neural networks: A survey

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.207039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.657146Z digest=sha256:37efe50e576651d043bdf45cb82ec11c241dc39fdb413dcc9cf35ffd369098cb

Observation a1a36eec-d4d2-4eaf-b7f6-39708a0b1ad7 · outbound

This paper cites A dversarial attack on yolov5 for traffic and road sign detection.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks A dversarial attack on yolov5 for traffic and road sign detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.192083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.661596Z digest=sha256:f584f49fa9678fc478a2023a8f879155a437b8d45bbc74a100aeaf03bb43b4e0

Observation 2814c903-1eea-4b62-8ec2-a39b84838e12 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.176666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.666027Z digest=sha256:514993d98968543f8cbe18f74858460452d1d3237d3d1530e147b8c2f18e23bb

Observation e99755b9-7c87-4a02-8f9c-c7eebdc838d3 · outbound

This paper cites Language augmentation in clip for improved anatomy detection on multi-modal medical images.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Language augmentation in clip for improved anatomy detection on multi-modal medical images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.161183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.670475Z digest=sha256:aba95a49690ba4f97e23f7671f0a7fb9b945e4ee187e19ebdcd84ddea6c6ccd4

Observation c0a9211b-c188-4843-a68b-8cbbf12409c7 · outbound

This paper cites PALM: Predicting Actions through Language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PALM: Predicting Actions through Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:16:15.971825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.675732Z digest=sha256:55a701340f491093099326d18e9e807a7f99647505312783e6d9dc52ed07a3e1

Observation 2a182daa-20e5-425d-a840-eb420f81bad6 · outbound

This paper cites Segment Anything.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Segment Anything

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.680772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.680772Z digest=sha256:15267964374a754600a89ee6ef5bc10edf1402fc323bd29b9ec14057de4e525c

Observation 6e691084-371f-4eee-9828-7132922ec243 · outbound

This paper cites Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.685735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.685735Z digest=sha256:f6ff71e981573bd9affec6fcd60ce655f2e12b7070718cfba3512451eb2977a8

Observation 57251939-f85f-4b06-908c-fd9499ee11db · outbound

This paper cites Enhancing clip with gpt-4: Harness- ing visual descriptions as prompts.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Enhancing clip with gpt-4: Harness- ing visual descriptions as prompts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.145623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.690697Z digest=sha256:af3e97ef04f4a53215eef341dc038f6bb8e8dc7963b676cce6763edb8db6e7e5

Observation da2b2fb5-34e3-45a8-afb0-7628970b5995 · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learning transferable visual models from nat- ural language supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.130369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.695366Z digest=sha256:63dbd01de78c3507290c66fe38cb8409b44073792b7a56ad63f6b1f513d6c8ff

Observation 7f484ea5-da3b-4100-8470-099fa43f2cb0 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.699872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.699872Z digest=sha256:8b82f2069a1b8d55eabcb006f4c9803166945ee6e03a4cbbbc622848ab6b6d0c

Observation 331a88c7-9971-4d18-8149-35d44671dac7 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.704561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.704561Z digest=sha256:3c91b1a65e25251b0dc9ccc74d05d62f253061ef26ad2a838b76cd00db02af77

Observation ddf7199f-94ae-4eed-86a3-d5d6845f63a9 · outbound

This paper cites Test-time prompt tuning for zero-shot general- ization in vision-language models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Test-time prompt tuning for zero-shot general- ization in vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.115214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.709612Z digest=sha256:c3deee2b62ea198123528872439e234a625a3572d78af39ed5913227ee40746f

Observation bf7884b7-6150-40d6-b059-befcade4fa8c · outbound

This paper cites Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.713836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.713836Z digest=sha256:28c8e96c508685806491f8c058468dcab9510e2eca8832313bb94df54cd4aaa2

Observation b738c268-606b-4352-b927-1ba254d6912f · outbound

This paper cites Motionclip: Exposing human motion genera - tion to clip space.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Motionclip: Exposing human motion genera - tion to clip space

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.100219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.718612Z digest=sha256:7fa9ba48c04dad0c6904a24d9fd1713f6be23c3130a6b06b62ed9d0790ee0462

Observation 0b0f2520-2a16-4ca4-b37d-d030f0fcd319 · outbound

This paper cites XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.723157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.723157Z digest=sha256:90505b99e561eec4b66a5c323c869be131c43b2acc8c99b1adedb34d0ba74e18

Observation e3929c2c-f003-455f-b096-1c14574783e1 · outbound

This paper cites CLIP with Quality Captions: A Strong Pretraining for Vision Tasks.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP with Quality Captions: A Strong Pretraining for Vision Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.728021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.728021Z digest=sha256:90e3dc0aa30ec26b49706045cdf4a3508ccc8211828a7c2f6d328a73eabcf13e

Observation 315d1def-45df-424e-974c-f6743b83bf3c · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks ActionCLIP: A New Paradigm for Video Action Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.732959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.732959Z digest=sha256:d9e0e6ad583d0be334dceebe764a7574a33b9058781399dc592b6c079f30d58a

Observation 8dbdee37-4579-415f-b08d-0d3b788052dc · outbound

This paper cites Actionclip: Adapting language-image pre- trained models for video action recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Actionclip: Adapting language-image pre- trained models for video action recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.084863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.737740Z digest=sha256:308d72574b2101084b02d9c3b2a69fbb3c94f6fcf02713d8d86bfd2effbf18cd

Observation d57a36ad-9822-466a-87d4-44028a193620 · outbound

This paper cites Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.067933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.742143Z digest=sha256:f120890bd4b92d239de447b61440f354d0909e36325014d1f70c5901d1205b80

Observation 088b7585-9ecb-44ad-bdeb-4998e8802493 · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.746661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.746661Z digest=sha256:1a48615587898f7cc700099bc83f95369f2b11c3247949c25e3960d8eacdd05a

Observation ba183e20-0401-4f89-9bb3-aaceb3f15e20 · outbound

This paper cites Revisiting classi - fier: Transferring vision-language models for video recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Revisiting classi - fier: Transferring vision-language models for video recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.051759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.751782Z digest=sha256:69d82530c219719eb85e7019fe971d1c3387f0bb04723fecf0b26321b58e9bf4

Observation e8df1d71-a006-40a7-aed9-07546bda48d5 · outbound

This paper cites Investigating Compositional Challenges in Vision- Language Models for Visual Grounding.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Investigating Compositional Challenges in Vision- Language Models for Visual Grounding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.036514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.756446Z digest=sha256:e1500bd376a44c9e877d817814c2e62c48fb7dcc88e46d924273e4344eb8605d

Observation 9bd7a2d6-847d-41af-9280-a3edbb1c6d6b · outbound

This paper cites PeVL : Pose -Enhanced Vision-Language Model for Fine-Grained Human Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PeVL : Pose -Enhanced Vision-Language Model for Fine-Grained Human Action Recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.020935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.760727Z digest=sha256:d98d68a9dc026ff296c1ec9f56544a496fbba5c1aa203c4eb5d76167a6f8786a

Observation 88089a6f-6aff-4ae9-978b-d96fc226bd36 · outbound

This paper cites Vision -language models for vision tasks: A survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision -language models for vision tasks: A survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.004907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T18:16:15.765144Z digest=sha256:d9764d53de4a3881f5011f723e0f3a9624a4352ed588d39f597430d20f2680c8

Observation c828b911-0368-4390-9f3a-ef08f5229124 · outbound

This paper cites CLIP in Medical Imaging: A Survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP in Medical Imaging: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.771212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.771212Z digest=sha256:652d9f7e12eb8005bd62eb64dffe2c16a458f3462a3b2bcb66b85e13194ee638

Pith citing papers

No inbound Pith citation observations are available.