Pith. sign in

Paper Citation Record · LEDGER

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

As of 17 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2412.09895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09895 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:41:35.796320Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:55:34.614403Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T22:55:35.017324Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved62
  • parse uncertain5
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1e4052c-37f7-49c0-8ff1-b206c51fa4a5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:37.008579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.460295Z digest=sha256:0284aa1dc794dd1cfdcc90c90811bc69a332ce14c84d8e1602a46d9d0abf97c2

Observation 8d83d00c-8508-4542-be35-1b33c1091153 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.993580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.464893Z digest=sha256:b2ff259d88472ea2c4fb1f3b1ade0a704064906d58083a2d074b42a562cf70f6

Observation cd1035d4-28eb-48c7-9093-06919a7e5916 · outbound

This paper cites A Short Note about Kinetics-600.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP A Short Note about Kinetics-600

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.425233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.425233Z digest=sha256:6631ce187875a30eca4be4a48c6e86faec08da63a33c11ede0c20e915f785056

Observation b5628ccf-e913-441f-bc6f-25fd9d9b3ff3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.962123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.474613Z digest=sha256:2f1fa06b9c669b5c1018d5ad9b804b98ccbef6130819b7dc5d7e2577239fb0b3

Observation 0748889f-d01e-4c92-8b6d-569f533d136c · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.932708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.479269Z digest=sha256:cc9374840bb1b6b2d6d155f6591d58eeda757cbed6f52e2cf0eb8c9a83f779ef

Observation cece37e7-1654-4128-9149-d896670024fe · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.918232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.484238Z digest=sha256:272a252349e8edb8b7312b5371fc52664a80df667bd8acb76068b32f29875ba2

Observation 6e3c96fc-4ad9-4dd4-895c-14abf033f7c0 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.444683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.444683Z digest=sha256:54cb1fe6a6c1f23a936bd368fee794c76cd79d2acedd4ad74e941ca531c61713

Observation abcbb0da-2040-4bd1-90e5-bde471121f85 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP ActionCLIP: A New Paradigm for Video Action Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.449384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.449384Z digest=sha256:a19906f2b67cb24bc773f3000e77d8cf162c4353c78ffc42e1f22025086581cb

Observation 7d050f96-786a-4bde-8973-7ca2b622a978 · outbound

This paper cites AIM: Adapting Image Models for Efficient Video Action Recognition.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.455213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.455213Z digest=sha256:11539a08359c5ddaef45338a00c3de4a65b160029513c0ef55d699b9a3d7f839

Observation 2d5d870c-9dd9-4c77-96f1-7ef6de7e99a9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.977628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.469970Z digest=sha256:68cf63a8bf8b2ccc43a7514153485a6eb8d8d477510c3472e51d48820caebf8c

Observation 24df1153-402e-4455-a5ab-8c5887c7c98c · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.903915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.489068Z digest=sha256:4f0ce955f18cec118791d09db087c3474bed2b5a936cb7bb98d314a962765f24

Observation 2c3238b8-7ef8-4f67-b250-f54f529a39ba · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 17

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.888249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.493873Z digest=sha256:adba225dfe8ecb519374de5465cec36b29d342a16d7accaa87914cdf1f092598

Observation 658f3d45-73e9-426b-a703-3bca23e6cac8 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 18

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.873477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.498702Z digest=sha256:cee4a988ce6987e9c4a896fe39d332cce9cde11272afe8cf3b8e503e61d04370

Observation c29aa97d-2e3b-4258-b881-a5d3b385ac85 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.857743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.504067Z digest=sha256:08ab01cabf351a275c16bf6ab3c274b9a047897a951730d7d6207681982bea01

Observation 000d4a80-9583-490e-b2a8-012a0159c461 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.842178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.509306Z digest=sha256:07ad776f1347df2f657642b6908307a549e1040f2a43c78dbdba8ee8e085a86a

Observation d0408fee-b2f2-463d-b198-de4a8d6559a5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.827080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.514358Z digest=sha256:5c1523ea4bf8e104397b345844dfa0cb6aa4cbbac69c6f8a8f852ed896951051

Observation 4f86c36b-5afa-45ce-a71d-9884aa93d8f2 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.812538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.519169Z digest=sha256:f157b31a3e97c31030d1dfe06d22590f876c085d58e7beb541a80be9dfbd8a1d

Observation 2e37585c-0929-45ce-8feb-356115eb4bfc · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.796896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.524367Z digest=sha256:64132b88326fab9b4ef3b8bcf2cc9eebfc353404d299910440e732d5c46960cf

Observation 472d5c52-89ce-4ddc-8f95-977771c9997c · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.781860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.529457Z digest=sha256:3698b03316c5978448a72f61a647bb9d2b24d63166aa5e4a17d155fd9b2a2b28

Observation 907bf1e0-99c7-4291-a594-569dad3c5d78 · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.765399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.534387Z digest=sha256:4a846b062eb6981e6761b6cc7da2cdfeac79ced6b539b07f281fa54d44d40c3e

Observation d464e9af-b9d3-41a4-bf4c-b103b0615b69 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.750950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.538823Z digest=sha256:06eba61ffc42a47ba4ff9e7218283bdf5eee491023c86cf6ff1995525aeb2ba4

Observation 2905054b-43ef-4078-b1a9-f87f9efb41c3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.737523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.543406Z digest=sha256:f142312657b0d5635ea5fa94541a5371b08174fa968c73db4350e6248a2f44b4

Observation d43fb80d-a3ac-4d7d-b238-37da53c63b3a · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.723612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.548859Z digest=sha256:5b1155deb7e15d3ad2dd291f273aea1e449c1ffa1aeb78bfdbc5ed2ec33d833c

Observation 520871b4-8713-4c6b-97aa-4e9dd1b395b0 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.708199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.554641Z digest=sha256:bb59b2b3b57907413f087190ee246305ba0d944fd7868bca66f13649924215e2

Observation c24c487f-8871-495a-9cf6-939cce21ac34 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.689492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.560349Z digest=sha256:ebb915cdc17bcd6dab187fbe6217efc854be79e220ae5080d69ea9bbbf163d75

Observation 34bc97d2-29ae-42d0-bb6c-6090aaa87c95 · outbound

This paper cites Responses for action category “Surfing”: Object list for spatial text prompts [S]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Responses for action category “Surfing”: Object list for spatial text prompts [S]:

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.675106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.565565Z digest=sha256:2c663427a3712ba45e97a1360a38a3808e3b55d11af3b53a5d917ba9550cee7a

Observation 3508cd0c-47ef-4aea-890d-89a117075176 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.660917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.570823Z digest=sha256:9d32caa5ba002dfaff6656861e2ba06d5a4a3bc436ffddb7ac3809a6f3e49d87

Observation 4b833654-c6df-455d-8124-7cb60090ae95 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.646147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.574896Z digest=sha256:aed65716f9756b4f5831f1c49aca48563911196fecbc07858bb2a1d1991348f1

Observation c76308dd-8a8f-4355-9f6b-9d027017bb75 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.631409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.579164Z digest=sha256:3ac905dae302172b50372403734f48cd573ec608bcc0ef82837dfc0633920bf6

Observation ed8cb8bb-fcfa-4798-a8b5-2c31f6b40fc6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.616920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.583420Z digest=sha256:296fe45467d52ef32b3f4fb2655cd8e1746a1b9e5df9642ae537baed30b4e188

Observation b407b063-6e62-4508-bfe9-6f13d09b2999 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.600275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.587957Z digest=sha256:4f3778a3817ed83806d93c818354ede6ea409dea09ba8e1615d889c26860cd63

Observation 0ff8ec5c-6c18-4ec7-9a34-f818a585fc4d · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.585007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.592642Z digest=sha256:ec1a4e4fd9277adbaf6c15b69e7531313beed9546bc81b2b02b001a9f6348e6e

Observation 3310ed52-22ca-41e7-9fd2-dafb21b4346c · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.568910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.596857Z digest=sha256:ff2283f1d2e675568c5668791f10e50eb35e544445960d2adccb52d1ff3cd0d5

Observation 1e8ce5c8-232c-4fa1-a67f-f0418efaf330 · outbound

This paper cites given action name.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP given action name

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.554491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.602066Z digest=sha256:4fc1a852fe9b02474a02b7c56add96a008e41ecf39c21f6b8cf535585dd92e16

Observation 715ba0c4-9293-4c8c-9823-1f53f9e321c3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.539984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.607394Z digest=sha256:23094e8538f3079b71cd7b9d47b446856499ea81bbec794262229f2c9505f5c1

Observation cb3a4fe5-4b16-4954-8e1f-3e346d6c43fa · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.525766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.612239Z digest=sha256:af39e0879a2861e64542aadb2deef1430d555d450efd59c84bac844c8a7f12fb

Observation e6896d1f-d830-4e01-bb3b-e727d0669cdb · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.512304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.617075Z digest=sha256:82710f69a1a90dc5c113893864635723212e8a2348b2ce354f48dc73cc264828

Observation 4604b74a-3dc0-4a02-810e-741808f2f607 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.498650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.621966Z digest=sha256:f339051fce89fdc81176359aafd8478b892167cbc1c51022e24f94c903048969

Observation 7eef8db0-8f5c-4155-beaa-c48496b87f3f · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.483352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.627226Z digest=sha256:e15e385eb956a00ee7ba4673ca0448990e0645aae53b246d0c634deb1341b887

Observation c327faa7-29f0-4752-a44f-e59524ef7e60 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.468702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.632248Z digest=sha256:544e5362768aa82cc22a7b5d94970d56befeb44078c544e46a88b3504ce31f2a

Observation e4b7d01f-ae90-45de-903b-3630190e1f80 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.453446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.636966Z digest=sha256:4dc212161a93213d62f74b3c591507b8bbc5578b68d91b67a363d54f0db9fe22

Observation 72cbe206-ccde-4d35-9645-5921daf7cccf · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.438626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.641426Z digest=sha256:51632fd275a4fc91fe6f6ccc580a6f76d03fd0362a34008b42cbc52def9370d3

Observation c69e605d-6355-4977-a6d3-7ab9f20977a8 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.423754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.645907Z digest=sha256:069c14865e0dbe20c794403a8bb9534d3d7fda8cb30d1807c110367ec2907a49

Observation ace0df3f-cfcd-4585-9a75-981c6360c71f · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.408609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.650859Z digest=sha256:4fb7402f86811c7c2866011cc8292b8c166faac8dfcb76cd542b63b4e192fdb5

Observation 63a5e824-3a9a-4b6e-bfef-1b7d990d91a9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.394323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.655428Z digest=sha256:a4b9e5821dd2ec2e25c0d82b46b6c7a9af60d8cd27a7657de8636877401f5d69

Observation 054c740e-499b-4282-a1e6-3e626481ffe5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.379530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.659914Z digest=sha256:43e55c8d08d380a2ecc4f24e287cfda4639a84aa9ffd81e32dc5d2afdcc5fcc9

Observation 1be595b1-c095-4013-8b28-87e8dce80b7b · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.365014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.664511Z digest=sha256:1792855e5ecfdb38ec77e7070cff2e13557c04c059bab6506b1b483d068555cf

Observation e5002ea4-d56b-4fc6-8726-2f32aae58328 · outbound

This paper cites Clean and jerk.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Clean and jerk

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.351397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.668719Z digest=sha256:3456057bcddc09281634b2a4679f41e337f8b366e806dcfb11a1258e5d839c0f

Observation 92631da5-4954-48f6-ac02-b668ce0ea1e9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.337236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.673142Z digest=sha256:c3fc5caa8079d0d0c0c12bc2173f1e6963c98ed0f84dd4f1fd8858aec358c19c

Observation 2b06b85c-74b6-4174-90fb-c06b0342bb72 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.322517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.677555Z digest=sha256:8db20632e1db02c0fdf28a44bc02204c45321fb27efd60ac5d81fe853d960851

Observation 76a85673-baba-4a35-ae23-0d898b89b1ed · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.307906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.682303Z digest=sha256:f69c5304922d109bef76ddbf3c79cc9043028d54d92e43ed5c33c6fcd879169c

Observation d26f89f6-a4c9-4616-820a-bb6ed704e2b6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.292915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.687056Z digest=sha256:b97219a0a889ad53558919142491eb9fa1d5fe8af07ec1df3732f4fd74c6d424

Observation 9bbbd63d-cf7b-4899-92c6-cc284b8693bb · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.278314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.692402Z digest=sha256:cb8528ec375dae1a67762bafdf9410cb56914281fefdd88472d7ca4643df234a

Observation 97282e13-6b49-4ac9-890d-c36f39908512 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.264381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.696907Z digest=sha256:acfe9fb31915e728e11bdc0fdc1de0a51541d590eb759412c13903b6e81829ba

Observation aa8e3779-b8cb-4e43-9620-c1c47bef0f37 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.249948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.702217Z digest=sha256:97401065406e7cf01924a748963a2160224f6738806611e597bc6a86618b5c7d

Observation c19b2ae5-1b73-46c3-b185-870babd944b3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.236238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.706824Z digest=sha256:1f9227077c2bd42bda59645a1e6f2d744fabbfb5e243cb8a2b4e8b78044011ed

Observation 6bccac1b-34ee-4e92-bd23-0208af35b7d6 · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.221966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.711922Z digest=sha256:dadfea8ffd84429e5bf7f0e3601f9b01a9599e5e19f8e6481eb389313eaf08b8

Observation 2ac87c45-565a-470e-bc54-2ec053dfc9bf · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 63

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.207027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.716236Z digest=sha256:0f4e701e06a76b198ffb638f393fd3ca61c0aa613ae0761c74158c32cc8dba2b

Observation 9dbecae1-f80c-41ec-8178-5857411271f3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 64

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.191558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.720612Z digest=sha256:912c14c4e7e18e64f2874288a958e8e09858ab905d692d979302b0b03b2df2a1

Observation 88f8cd3e-9d23-44be-b1b4-61b21f20c193 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 65

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.177192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.724883Z digest=sha256:c5f535787e1b089933a586713bf72c84e6486b8ef3270424aec507c82f165e71

Observation 994ac44b-c81e-47e8-8238-0f61e20ab453 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.164074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.729353Z digest=sha256:2c743fe16968a8590f160167446c15913d8d5d1131aff242bcd7ef22eee7a2e1

Observation 5383e4d2-5546-4226-bd41-b6efc58cadb6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.150134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.734000Z digest=sha256:da5b9e442d0294cc3f140d8120e6fa7479127d244bdf0afb31beb9a78178fc77

Observation 56de4781-f5b4-4ce5-9357-2f6d7cd7ac92 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.136455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.740225Z digest=sha256:a35a0d75d195fe88b21b9cf11252339eaed898736b0f72679e4087bf800c71c4

Observation a976d8ea-9993-4140-93ab-85344d7a0622 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.121130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.744329Z digest=sha256:19f9ae64b49429c99a4a4e3736d119a313455e7a57a834ea03f27ca79ba047c5

Observation 8ff0ee90-0967-4aa0-a5bb-ad98b5030455 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.106517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.748776Z digest=sha256:857be4bcc802afdb31056303f53f24caaa2c10189456c3f62102f9dd47995af3

Observation 600e6fb0-18a8-43d1-a4f0-ff57bbd78436 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.090831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.753751Z digest=sha256:2a44cb4b36c3305bc56978a8fa65d56c114731d434c7087e87e50a387d01f0c0

Observation 1947000e-f008-4154-9278-f7db4966acf2 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.074870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.757928Z digest=sha256:1e4b7f5c3fb4c54f0a256faa4898ff4f552ac29cb9ae672abe40213fddbe88a4

Observation d4832963-bdfc-4ba9-a2ca-c9703a2967af · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.058938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.761853Z digest=sha256:29f1f2efc6eef71239e9580b3247f3fd9b097a0e4f5bea14510d4698cd15b8dd

Observation 0b614176-2cc3-4a8a-86d9-09896f602c06 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.043397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.765894Z digest=sha256:dd67c05a0708dc086e0d51359aba9f94d39723fe91622b8fd0ac2201f4b315c2

Observation 0f47195c-4b35-4bdf-b4be-2eeb045cc8ea · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.027022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.769936Z digest=sha256:5efeceba44815fb10b9bc786390c166acdc664ed9d6461c47793ce049a949c10

Observation 62da1bdf-4bef-463b-9509-9bc11828d69d · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.009686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.774046Z digest=sha256:de8dadf1a4a759ec0eb204f9698e03564fd25d91ec68d38bdb858f6c6daf8766

Observation 918abf29-90e5-486d-a3d9-c4483d0ac1c6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:35.994903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.777983Z digest=sha256:ec7b1daa4e539926010e4afcc6d093603e7ef4e474c92198243ef0c6771210f4

Observation 16f78237-19b5-48a5-8385-a8228b15cbf3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:35.980916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.782355Z digest=sha256:96c741c53ba54f0dab7666009e97954b2b901ad9b2b12cf718287498f9d84ccb

Observation 2927d39a-e0fd-471f-83c9-e617c65621fd · outbound

This paper cites B Details of Datasets and Evaluation Protocols Datasets We conduct the training process on Kinetics-400 (Kay et al.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP B Details of Datasets and Evaluation Protocols Datasets We conduct the training process on Kinetics-400 (Kay et al

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.965754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.786825Z digest=sha256:cfa61b7ed697d89308d83d50a14d9f428d640cad4bccce4d1679800ea9d20bf2

Observation 20acf465-aae6-4172-8886-cee4e5fbbbc9 · outbound

This paper cites jump”, “kiss.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP jump”, “kiss

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.950023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.791432Z digest=sha256:4581e4e5b970894d106930e70b14380f0bfb9516a989fd633940fc0d2246ba5d

Observation 3af675e4-82c9-4254-8a9b-b2d059eaab1d · outbound

This paper cites This can be explained by the fact that larger temporal scales result in sparser interactions for boundary frames during channel mixing.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP This can be explained by the fact that larger temporal scales result in sparser interactions for boundary frames during channel mixing

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.932844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.796320Z digest=sha256:528eaff575e6c0f1c9f705f21be04b6fec2f33a01f74cfdddf6b12d6fa693efd

Observation 8cae5fb0-5c0b-4c5a-8374-82e7c0d27291 · outbound

This paper cites In 2011 International conference on computer vision, 2556–.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP In 2011 International conference on computer vision, 2556–

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:37.042145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.435691Z digest=sha256:440ad4a57103ed7a2125506b8b78ef5fe397caeb9661cdd924b1c2b0f92aa592

Observation 3649627a-c22f-4a9b-9bf1-6911b23e172c · outbound

This paper cites The Kinetics Human Action Video Dataset.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP The Kinetics Human Action Video Dataset

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.431226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.431226Z digest=sha256:339d24e7fdfd354019a1d9c9b885de4249e259aa33eb9002ed5ede0f84ed15a7

Observation f095a1d2-2ee5-49c7-926c-d9e245a1484c · outbound

This paper cites InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 4613–4623.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 4613–4623

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.420099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.420099Z digest=sha256:b614ca3d4fb9a960241a28bc756225b4db7b197e733a17f53eec6aa641fc32f3

Observation 37347fda-c884-4304-88e1-dc77d1cbd936 · outbound

This paper cites GPT-4 Technical Report.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.414650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.414650Z digest=sha256:1cabe47ce061b5341b730899f8da0f3439d5a3aceb020af8de5857cb8094d391

Observation 55190b9c-e992-4a7c-9057-f09fae1fba50 · outbound

This paper cites Lee, D.; Lee, J.; and Choi, J.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Lee, D.; Lee, J.; and Choi, J

Reference 2563

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:37.025080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:41:35.439885Z digest=sha256:9a01166e44ba467a22d134f6cf51b2d91517c4f34b68a31da4d42b79d2cd3fc2

Pith citing papers

Observation a789da5b-9901-4a69-8fca-e9115415c18e · inbound

Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition cites this paper.

Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:55:35.022671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:55:34.614403Z digest=sha256:98de91d90281ae473e53c8c1ff191531f6f16167ac719bd1a5e249b127116a4f