Pith. sign in

Paper Citation Record · LEDGER

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP

As of 17 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2603.05962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.05962 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-15T14:07:39.892672Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b75b2e5-1c3a-41ef-b654-34161a5a6193 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:00b9db4a2c8f70257c40bbfd092452c3cc6c0de02813d2854dfce59f0de5d5a0

Observation 16f48e47-3aa3-4e18-b672-e4c201c07bc8 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked autoencoders are scalable vision learners,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:d194d46c5e5c295eacf88d7d0a90d12c0cb7f88f4f90c066ea06b0b3784709c6

Observation 92c9d89a-73d8-4c1b-ac04-65fb376299ef · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:042eca92875a03e0dab0f1424d425d4e6f8fb801e219a593f4ef4261152c1e6c

Observation 611b93f7-047a-4f4e-93b2-bb6f6c8d3918 · outbound

This paper cites Flava: A foundational language and vision alignment model,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Flava: A foundational language and vision alignment model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:ef3091455cacd95ce867833ee6928deaa6c6ed632bcb7ef98dc2c8eff8e0ca2f

Observation 0fa5598e-b729-4188-9b37-d89559577fa9 · outbound

This paper cites Large Language Models Can Understanding Depth from Monocular Images.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Large Language Models Can Understanding Depth from Monocular Images

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:8afc4b59456d028a51d859d9d5539d7ed5ceb72241be1acedc0ef9506eec882a

Observation b723cfb6-611a-40cd-8b77-87c161635084 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Depth anything: Unleashing the power of large-scale unlabeled data,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e401eb0cd93caef63f1ecbe74ee7a7f0d1636f9e42e3ec08ea2afb18f7cfd9ec

Observation 111a5061-44b3-4a7b-b6f5-f05989653757 · outbound

This paper cites Pad: Self-supervised pre-training with patchwise-scale adapter for infrared images,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Pad: Self-supervised pre-training with patchwise-scale adapter for infrared images,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:d9b097bac34020aa414c88ab100199081b01e14ac1a91ace868748dd21e6eaa1

Observation 4d44a9db-978a-4e93-a2e7-8b5f2b33831d · outbound

This paper cites F-ViTA: Foundation Model Guided Visible to Thermal Translation.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP F-ViTA: Foundation Model Guided Visible to Thermal Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:bc5680c566d687138ba780dad32025405b915d834ea765c5fc40531386fd1661

Observation cb6d7a1b-7ea7-4cf0-8a8e-490e188954be · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:eeb8eaea786b6b08677721a1a64beb75ba38321fde649600afbabb85a5df614e

Observation a1a9f6db-c3d4-4f08-a9f3-795658efce97 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Videomae v2: Scaling video masked autoencoders with dual masking,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:cd701a95809f7ab8d862b7be89cb962a31bd6903bc5fcfa976e164cee73769e5

Observation 5cd75477-a027-4bc3-a3bb-406a26e6666f · outbound

This paper cites Omnivl: One foundation model for image-language and video- language tasks,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Omnivl: One foundation model for image-language and video- language tasks,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:5fd4a8a7d51db78edc4bbca98ee8cf74fe717f75b1e8cc7eabe25b54181f8aad

Observation 2488037c-a727-417a-86f3-ba826fa6122d · outbound

This paper cites P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:57390ca1ecedd2e8fb79388ce39d0a464943af174b1b6707681f9583a20ce368

Observation 8ff34b17-245a-4985-b3c0-d6a3df311583 · outbound

This paper cites Pointclip: Point cloud understanding by clip,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Pointclip: Point cloud understanding by clip,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:252bb211f6843f6deaff57bbf8f216038e448241bf06ef8881e46851ba28b863

Observation e8ed976e-f73d-4c0e-a362-3c5a0ec23c39 · outbound

This paper cites Diffusion models as masked autoencoders,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Diffusion models as masked autoencoders,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:27bd54364f70d276048f21033070e9fc5b7a7f5799a9d929640e10a0d9e4f69a

Observation e2978196-0f8b-4c2b-8286-b08a02d76dab · outbound

This paper cites Hierarchical recurrent neural network for skeleton based action recogni- tion,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hierarchical recurrent neural network for skeleton based action recogni- tion,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:f3a44a625b31fbbb77991fc26b35a51c9994fdd45b37b9c8b92506e427390798

Observation 5523ccd9-7c6d-4732-8217-9691868bda82 · outbound

This paper cites Skeleton-based action recognition using spatio-temporal lstm network with trust gates,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton-based action recognition using spatio-temporal lstm network with trust gates,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:a6fce24618afc0df97528b957ac72dc7bb6a32138e898122905cd79797a35144

Observation 796176f6-e6ad-4eae-a3b5-6aef00f31692 · outbound

This paper cites View adaptive recurrent neural networks for high performance human action recognition from skeleton data,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP View adaptive recurrent neural networks for high performance human action recognition from skeleton data,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:fcd2dc98ed985d0d41ea4a305e262422d7bb4bcf6efb4103ce97f3be86fb5a13

Observation 20f3bb99-eea7-42ce-890b-d1798d643c43 · outbound

This paper cites Skeleton based action recognition with convolutional neural network,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton based action recognition with convolutional neural network,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:97b04a421fa9adad9dc78082e211453de4a25ee86b53e78f7755ed816d07c791

Observation 15bc7c7d-9e61-43b4-955d-1681751aecb2 · outbound

This paper cites A new representation of skeleton sequences for 3d action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A new representation of skeleton sequences for 3d action recognition,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:8af82c3cb67c796470546daaf384e7f0917cd50623e73c422e7939f103bc85af

Observation 2e481554-1aec-4c89-a9af-94710de46594 · outbound

This paper cites Co-occurrence feature learning from skeleton data for action recog- nition and detection with hierarchical aggregation,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Co-occurrence feature learning from skeleton data for action recog- nition and detection with hierarchical aggregation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:0195006b163adf2a73e7207d8285ae539483f62448091942bd0275b796c49287

Observation 35da2927-a66f-4c57-8013-ed70c2d900e5 · outbound

This paper cites Spatial temporal graph convolutional networks for skeleton-based ac- tion recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Spatial temporal graph convolutional networks for skeleton-based ac- tion recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:5c4261e71cc798385f7208039a8442e6ee5493ee217d7695bc370f72f60e0922

Observation 3b0ed4d5-9bcd-4331-8754-fc6dbfa591b0 · outbound

This paper cites Two- stream adaptive graph convolutional networks for skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Two- stream adaptive graph convolutional networks for skeleton-based action recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:74b67a2e469b104622b12f75fb312b5fe0f37dbdde530d1e7bad941083036f39

Observation 685c5728-0bbd-42f3-a779-efdebabb2314 · outbound

This paper cites Channel-wise topology refinement graph convolution for skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Channel-wise topology refinement graph convolution for skeleton-based action recognition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:4135f371f4a1c83ac5c6da39f53bf896a05e913e72f7906ea40069604904987e

Observation 0dd03cef-bd49-4e90-8736-0dc1cf81cf84 · outbound

This paper cites Stst: Spatial-temporal specialized transformer for skeleton- based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Stst: Spatial-temporal specialized transformer for skeleton- based action recognition,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e8842393c9c7297fae58a0dcd1cb8b04324d077a6371c2cc1085b3d1a245a470

Observation 40655a0a-789d-48f3-841e-4af85e150dc8 · outbound

This paper cites Hypergraph transformer for skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hypergraph transformer for skeleton-based action recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:2ee497211f2aa3fa7306f09eb5163c6406fbcc8de52feee31d5afdf9ac941127

Observation 0e4e9a83-972a-4ea9-bee0-aafbbdb67098 · outbound

This paper cites 3d human action representation learning via cross-view consistency pursuit,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP 3d human action representation learning via cross-view consistency pursuit,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:5be443d35cafadb853a14089eb8de0300a2f7fe925144d694ea131ba172c6b48

Observation ff3fe65c-ac7b-4946-866c-0cf3d22890f4 · outbound

This paper cites Contrastive learning from extremely aug- mented skeleton sequences for self-supervised action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Contrastive learning from extremely aug- mented skeleton sequences for self-supervised action recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:3e2e1980c7c60d6fe5ef2a9f8add5876b81ecdb073d4870a70fe694622cb8b2c

Observation 36cbc4a1-ebc3-4c8d-9e15-b43af9bbe8e0 · outbound

This paper cites Contrastive positive mining for unsupervised 3d action represen- tation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Contrastive positive mining for unsupervised 3d action represen- tation learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:1fd17d51e16e7bd1de343e1f46a6981a558daff838dd0027cc79d58ebd44ffe5

Observation d7610019-3dee-4acd-ae89-32d6669e20f4 · outbound

This paper cites Masked motion predictors are strong 3d action representation learners,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked motion predictors are strong 3d action representation learners,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:448b9a23b87d7efb2723e4cc1eaafb60e45b67613563f367c070627aa77f512a

Observation e1b9ff96-6b20-470d-83d8-866757908ec9 · outbound

This paper cites Macdiff: Unified skeleton modeling with masked conditional diffusion,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Macdiff: Unified skeleton modeling with masked conditional diffusion,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:da887a3a7b0b0935178b2f054839ca3a3852ccf2058a91c3c3542a56466fc042

Observation 8e3d93a3-9fd4-44d4-ad14-3bc87e06cc2c · outbound

This paper cites Skeletonmae: Spatial-temporal masked au- toencoders for self-supervised skeleton action recog- nition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeletonmae: Spatial-temporal masked au- toencoders for self-supervised skeleton action recog- nition,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:a38f8d05cb0569c57ebbc1f1aaf30d9c1e05da64b2047309500ac73015d852eb

Observation c1d1644d-adc5-4f5d-9dae-0c4a96b9cc1a · outbound

This paper cites Momen- tum contrast for unsupervised visual representation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Momen- tum contrast for unsupervised visual representation learning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:41402f64dd12dd9d4dfe0a7a802c9881a06bb4736efc8f3c98cced73f9ee845a

Observation bc34bf90-a585-42aa-b5e5-4d1c071e875b · outbound

This paper cites A simple framework for contrastive learning of visual representations,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A simple framework for contrastive learning of visual representations,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:88f9d8c251fae0612b048300286ff560db6554c910dae07f0472cf70cc3b9e95

Observation 2d1cac2a-ec23-459a-ac93-88679ac68309 · outbound

This paper cites Spatiotemporal contrastive video representation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Spatiotemporal contrastive video representation learning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:6e584dd90cb91ebdc12bca4580f4fb559a176958a43de2091962c0a166f4cb1b

Observation a1e20b50-8770-416c-9a95-5c0ebfe960be · outbound

This paper cites BEit: BERT pre-training of image transformers,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP BEit: BERT pre-training of image transformers,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:ba536a8990901277990a9aa3c21a0485eec64776bf249c9249b4955c4161399d

Observation 963da4a3-adcc-4625-9d6d-06b317d49c95 · outbound

This paper cites Masked feature prediction for self- supervised visual pre-training,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked feature prediction for self- supervised visual pre-training,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:2ad14bebe40711d0010864d8cd2be51e431256f8e0765f23a385f7fb1cad7502

Observation d72f2d0e-9b95-4758-a0ca-4fc926622d97 · outbound

This paper cites Dynamic multiscale graph neural networks for 3d skeleton based human motion prediction,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Dynamic multiscale graph neural networks for 3d skeleton based human motion prediction,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:acb8ff8e2aa2fbef90d1d330a064314452720fca0f33fa7660068609c353484e

Observation 45bf7f52-8a73-4808-89d0-3ae3b0957aa4 · outbound

This paper cites Exploiting spatial-temporal relationships for 3d pose estimation via graph con- volutional networks,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Exploiting spatial-temporal relationships for 3d pose estimation via graph con- volutional networks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:ff6a602c0d4f6314e975e078bfd85755d5300ee14ead06fab938e843c076e143

Observation e7fc133c-520b-4a72-8b5e-26dd6b02075d · outbound

This paper cites Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e5f68fecb3e421d5c8c32f349c0a0e648f3fa8df370885ba769525892259993b

Observation d86fc034-19a8-4dd6-919a-b288fbf868cd · outbound

This paper cites Skele- ton cloud colorization for unsupervised 3d action rep- resentation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skele- ton cloud colorization for unsupervised 3d action rep- resentation learning,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:5d42cb12fd264fc159823ef0ba372bfac96770d6e58b508e0eb8de0d99a7b225

Observation 954ff6b5-b7a8-4f83-9fd0-fb72b4f4c071 · outbound

This paper cites Collaborating domain-shared and target-specific feature clustering for cross-domain 3d action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Collaborating domain-shared and target-specific feature clustering for cross-domain 3d action recognition,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:8dbd69ba0f97e1d07fa6bd8ab1d5f4c81d7f97d76a3528ee885f5cb36af2853a

Observation ad617174-7455-431d-a686-caf1be3f010a · outbound

This paper cites Ntu rgb+d: A large scale dataset for 3d human activity analysis,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ntu rgb+d: A large scale dataset for 3d human activity analysis,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:9ce766d7da34fc2eeff6cc5e6f02adbc72bbfe640e6c2ec4ce6f1c5b1dc5c811

Observation 740dac90-8092-4db9-8de0-d5d43b59df45 · outbound

This paper cites Ntu rgb+d 120: A large-scale bench- mark for 3d human activity understanding,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ntu rgb+d 120: A large-scale bench- mark for 3d human activity understanding,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:95e3f346bc61574e4f8fb6b717e218edcf88548c9aee92beea62d1f3e579ee15

Observation ff22536a-2572-415b-b8c0-3906f6120615 · outbound

This paper cites A bench- mark dataset and comparison study for multi-modal human action analytics,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A bench- mark dataset and comparison study for multi-modal human action analytics,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:545f67c199998928f2b7f9bd220564c69c3c46d71f63360d4a2b4b76c615c4a6

Observation f5abd89c-e6a2-404d-9f42-56e41e5d601f · outbound

This paper cites Cross- view action modeling, learning and recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Cross- view action modeling, learning and recognition,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e80b40e120bd53cdd01b9f378d2a33eb8459af46766915f9fca49f917cbba82d

Observation d2685dd8-9338-4a53-9c5f-5a8340194899 · outbound

This paper cites Toyota smarthome: Real-world activities of daily living,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Toyota smarthome: Real-world activities of daily living,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:f3c611493b85509ff34fd3ea2b98892a4dac676e3b34f55c87d12465b5218384

Observation e3c34f84-7f86-4f64-bbe4-10bc643a5541 · outbound

This paper cites Semantics-guided neural networks for efficient skeleton-based human action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Semantics-guided neural networks for efficient skeleton-based human action recognition,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:bc4b885e002e1cea73b5bae0930e78ce4638059fcfd4b16c4591554201ed2792

Observation 07840b59-f146-4267-82a8-24f6ea7549ab · outbound

This paper cites Skeleton-based action recognition with shift graph convolutional network,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton-based action recognition with shift graph convolutional network,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:a628f8b43129fcc465b46880ac7ba4ca00d528e295708e3828fcd8ff47bb068e

Observation b3831f3d-0de6-44a6-a7fa-d9684c4d738e · outbound

This paper cites Unsupervised representation learning with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 long-term dynamics for skeleton based action recogni- tion,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Unsupervised representation learning with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 long-term dynamics for skeleton based action recogni- tion,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:9a7eead89ffa2591fc6259989878d71f92b60ef454f6ed0ebad7c1da518c8a36

Observation a4dc1a56-1fcd-4595-93c2-7a8bcba78426 · outbound

This paper cites Predict & cluster: Unsupervised skeleton based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Predict & cluster: Unsupervised skeleton based action recognition,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:61cf1a96918e27df2392498a2e575adfe26132501bb6446452089409a6410874

Observation 366592a3-c611-4679-9308-2319924da831 · outbound

This paper cites Ms2l: Multi- task self-supervised learning for skeleton based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ms2l: Multi- task self-supervised learning for skeleton based action recognition,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:6a8351bbfe77b882a83e3358b6baa7c1ce8e992e0cbe81ec820193c6fc5dfe7c

Observation a19d53fe-c254-490d-9dd5-be4d7fee82b3 · outbound

This paper cites Skeleton- contrastive 3d action representation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton- contrastive 3d action representation learning,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:450338a64ddf16851872b030ca9f298f1820670449cc9776fb8d67861c7c0fbb

Observation 802284f4-345b-49f0-ac4c-5ccd63fcca5c · outbound

This paper cites Global- local motion transformer for unsupervised skeleton- based action learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Global- local motion transformer for unsupervised skeleton- based action learning,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:05052dec158d6ae2cf08e3f10144a9bbbc707a6f7fe0622e9de371f113d087c2

Observation 7e277e30-1c62-4f1d-9927-ac0c2ce9a2f2 · outbound

This paper cites Cmd: Self-supervised 3d action representation learning with cross-modal mutual distillation,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Cmd: Self-supervised 3d action representation learning with cross-modal mutual distillation,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:fb957a8f60661f08848c3b0ad2de31594e33aeb79cc716a9ff7aca0f16c688bc

Observation c8c519ac-761c-4d41-afd3-011be7da6b74 · outbound

This paper cites Actionlet-dependent contrastive learning for unsupervised skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Actionlet-dependent contrastive learning for unsupervised skeleton-based action recognition,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:0c94f79ed70b04c0e318edd100df832bc353858f8ad3fd5221d76149d780c0bd

Observation 90aecfd8-ccd5-491b-92e3-9ecb43fed72d · outbound

This paper cites Self-supervised 3d skeleton action representation learning with motion consis- tency and continuity,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Self-supervised 3d skeleton action representation learning with motion consis- tency and continuity,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:562b6d2fa64b24e5f18b04167c7b054bff04323c79a7c8b6ff9bca364907ecbf

Observation 6db2b46a-de0f-4d0c-ba3b-f61c51bc18c4 · outbound

This paper cites View-invariant skele- ton action representation learning via motion retarget- ing,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP View-invariant skele- ton action representation learning via motion retarget- ing,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e3667bc2aa43b9a530a5242c9bccf995b10fb50519aa9cd35d2718470e658868

Observation d3d46f76-0065-48ac-8b8a-6ff8312fb83e · outbound

This paper cites Hierarchically self- supervised transformer for human skeleton represen- tation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hierarchically self- supervised transformer for human skeleton represen- tation learning,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:a7e0b63b5f326b8ab188acc292188c68d85e75dd98557e6c0ae0cbbf670b7766

Observation c66f9d00-dc82-41f9-aac8-616cf56e6707 · outbound

This paper cites Adversarial self-supervised learning for semi-supervised 3d action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Adversarial self-supervised learning for semi-supervised 3d action recognition,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:2a3115a9db3e43f0f95310dc39bdf6903dbd64b8def996049989a34f9f784023

Observation 3d8de597-38e6-4b43-a71f-51f897f56250 · outbound

This paper cites UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:70d2d470f77966fe02cbd286b9153143e176b3ee2a57850dec48f694601d339c

Observation 0d29d8a5-a399-450c-a9ac-72ec1b85f4ca · outbound

This paper cites Attention is all you need,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Attention is all you need,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:249c624fc15ead5aac3e6ec2f34d9a6c7d09d1a69ba79a7d4c92d47dcb42c7da

Observation 074e8d82-25c1-415f-b0b8-120a18f04c28 · outbound

This paper cites Layer Normalization.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Layer Normalization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:fda467152230c88b645089a584efd02c64f41f5ebbd2b4618c1fbbb49b51eada

Observation 1649a5a0-0926-4997-bedf-1ebce901ada5 · outbound

This paper cites Denoising diffusion probabilistic models,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Denoising diffusion probabilistic models,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:20692d0ba4a9a69eee0ed3cba62f9f7f69269e4c91110f6f18a8b4c13aebd40e

Observation 9cb9a11b-b15b-47ae-9fe5-27f1c93b65f1 · outbound

This paper cites Lcr-net++: Multi-person 2d and 3d pose detection in natural im- ages,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Lcr-net++: Multi-person 2d and 3d pose detection in natural im- ages,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:cea8d0cb92dae9b0eb25b0470ba3ee7257193b5e14f80d0a63a8e249b5767ca1

Pith citing papers

No inbound Pith citation observations are available.