Pith. sign in

Paper Citation Record · LEDGER

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

As of 16 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2412.09895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09895 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:41:35.796320Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:55:34.614403Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T22:55:35.017324Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved62
  • parse uncertain5
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1e4052c-37f7-49c0-8ff1-b206c51fa4a5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:37.008579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.460295Z digest=sha256:eac3f0f34e3d80ece1afdaee6983323cea7557489c31215ffcb34e09e5c77c73

Observation 8d83d00c-8508-4542-be35-1b33c1091153 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.993580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.464893Z digest=sha256:8d13ff8ed9fadd6b55b719b657da6ae509da873df1b3e6f5630bbd39c5716ebf

Observation cd1035d4-28eb-48c7-9093-06919a7e5916 · outbound

This paper cites A Short Note about Kinetics-600.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP A Short Note about Kinetics-600

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.425233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.425233Z digest=sha256:6631ce187875a30eca4be4a48c6e86faec08da63a33c11ede0c20e915f785056

Observation b5628ccf-e913-441f-bc6f-25fd9d9b3ff3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.962123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.474613Z digest=sha256:4e8be4b971705c3e78fc304aac6d3006e054d66a15bbc54f7f8181029fb6350b

Observation 0748889f-d01e-4c92-8b6d-569f533d136c · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.932708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.479269Z digest=sha256:a4d0a25cff2b7fff266a1db92109b0cc6a55d3e457502aeb563b09c6b2a85b2a

Observation cece37e7-1654-4128-9149-d896670024fe · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.918232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.484238Z digest=sha256:c58c81fea4be82f679a0f8030540a8f2d277840019f13747dc4a46fbc486438e

Observation 6e3c96fc-4ad9-4dd4-895c-14abf033f7c0 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.444683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.444683Z digest=sha256:54cb1fe6a6c1f23a936bd368fee794c76cd79d2acedd4ad74e941ca531c61713

Observation abcbb0da-2040-4bd1-90e5-bde471121f85 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP ActionCLIP: A New Paradigm for Video Action Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.449384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.449384Z digest=sha256:a19906f2b67cb24bc773f3000e77d8cf162c4353c78ffc42e1f22025086581cb

Observation 7d050f96-786a-4bde-8973-7ca2b622a978 · outbound

This paper cites AIM: Adapting Image Models for Efficient Video Action Recognition.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.455213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.455213Z digest=sha256:11539a08359c5ddaef45338a00c3de4a65b160029513c0ef55d699b9a3d7f839

Observation 2d5d870c-9dd9-4c77-96f1-7ef6de7e99a9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.977628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.469970Z digest=sha256:a39241aa48bc2067b92de13303a42e8dd603f07c0a7bd9db39442bbe05cd48df

Observation 24df1153-402e-4455-a5ab-8c5887c7c98c · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.903915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.489068Z digest=sha256:2c58e89e20054968ff0b8ab30e3083799dd31dcc56fda7d1f97d10b156f16413

Observation 2c3238b8-7ef8-4f67-b250-f54f529a39ba · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 17

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.888249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.493873Z digest=sha256:4af19a137c9a47d9c34cc021cb85d99c7dcd536790c1c1350d167923b356474b

Observation 658f3d45-73e9-426b-a703-3bca23e6cac8 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 18

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.873477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.498702Z digest=sha256:2385e9eebde1952456271b07d1272698524f7e0aca47693b4ebc249b71050554

Observation c29aa97d-2e3b-4258-b881-a5d3b385ac85 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.857743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.504067Z digest=sha256:4b83580b4b934316adeade4cfc597c0e8a8c3acfd4553e7b7b4bde15a9af797c

Observation 000d4a80-9583-490e-b2a8-012a0159c461 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.842178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.509306Z digest=sha256:d6a141fec7f2ee00ab042640aef7e79bd59fe408a6df88a5d9c084a23dffdf90

Observation d0408fee-b2f2-463d-b198-de4a8d6559a5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.827080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.514358Z digest=sha256:54a58ff7e315289a256ed65d1390755498712c045b1903c1690bcf0142506314

Observation 4f86c36b-5afa-45ce-a71d-9884aa93d8f2 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.812538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.519169Z digest=sha256:25227c63880045c8e1a93dc737447fe813782eef34cb0a19ec417432415d142c

Observation 2e37585c-0929-45ce-8feb-356115eb4bfc · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.796896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.524367Z digest=sha256:ead976b03b8d80cb8716a9ceccaa1947388bb5c56fbed5f0343d6105abbafed5

Observation 472d5c52-89ce-4ddc-8f95-977771c9997c · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.781860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.529457Z digest=sha256:bd54f942cc7792c8ee437753e2166d52922682b805913a92b7ef817743727f65

Observation 907bf1e0-99c7-4291-a594-569dad3c5d78 · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.765399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.534387Z digest=sha256:56dfbc62b5477c75fe30cada594677bdebd157d22c6ba629e3fabf2c2bedceb0

Observation d464e9af-b9d3-41a4-bf4c-b103b0615b69 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.750950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.538823Z digest=sha256:541a8af4df4f546bbe5381387e63c356d36550b49f7e1d91a0d4e1874185b469

Observation 2905054b-43ef-4078-b1a9-f87f9efb41c3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.737523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.543406Z digest=sha256:59cdcdf4e042ea1f630a4d7fdfcee3c77c60ceea1734bacfd06d32c5deaa2412

Observation d43fb80d-a3ac-4d7d-b238-37da53c63b3a · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.723612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.548859Z digest=sha256:178195c4132e454082e56cbbd3cbe54a201c38e0d398c199dd9fb15e71dbca82

Observation 520871b4-8713-4c6b-97aa-4e9dd1b395b0 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.708199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.554641Z digest=sha256:f5402d0d9647de86a918f3e08db3e5f33a294e306d37d56d6594126a1a7fb2d6

Observation c24c487f-8871-495a-9cf6-939cce21ac34 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.689492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.560349Z digest=sha256:52a40544dff1d5f4186f5ac17d857cf597fe4fe3d78fd3896df122c287a45965

Observation 34bc97d2-29ae-42d0-bb6c-6090aaa87c95 · outbound

This paper cites Responses for action category “Surfing”: Object list for spatial text prompts [S]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Responses for action category “Surfing”: Object list for spatial text prompts [S]:

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.675106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.565565Z digest=sha256:7ecdba7d11954e3a78e0836d8292fac6d9e779df68e0824046482126cd4d0f45

Observation 3508cd0c-47ef-4aea-890d-89a117075176 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.660917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.570823Z digest=sha256:4d8f29480afdc516853cfc49134c5702482977c6e44957b0ae67d0b0809d2776

Observation 4b833654-c6df-455d-8124-7cb60090ae95 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.646147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.574896Z digest=sha256:97d87ee7f92175ee312d28e1ffe3f9a4195e6da70010d53ec788f5cd684b0adb

Observation c76308dd-8a8f-4355-9f6b-9d027017bb75 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.631409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.579164Z digest=sha256:a2075408559adda8165899d1bdb5835e834bb4c10708343e6b278cd4d61f1623

Observation ed8cb8bb-fcfa-4798-a8b5-2c31f6b40fc6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.616920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.583420Z digest=sha256:c8aba586fac0fb5ff6b03457ae0946e57ece220e2cfd5fe5ea1597bf193db61f

Observation b407b063-6e62-4508-bfe9-6f13d09b2999 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.600275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.587957Z digest=sha256:397881def3adc24449ec1c7190b5b6746a1c47c352b916a64bce48c33ee05bc2

Observation 0ff8ec5c-6c18-4ec7-9a34-f818a585fc4d · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.585007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.592642Z digest=sha256:3c245873f1c17d91cd48ad7314fb8dd19fe4f8f25bac6c778e715f2f9073515c

Observation 3310ed52-22ca-41e7-9fd2-dafb21b4346c · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.568910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.596857Z digest=sha256:b7647d5c0a1d6857812b5b459cf51ef695bbbc4f03bdf23a929c9b406d802de6

Observation 1e8ce5c8-232c-4fa1-a67f-f0418efaf330 · outbound

This paper cites given action name.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP given action name

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.554491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.602066Z digest=sha256:317724474261a637f369a32cd40f0d3f970814905b46c62c829b88ede50878f4

Observation 715ba0c4-9293-4c8c-9823-1f53f9e321c3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.539984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.607394Z digest=sha256:57bd2e990f68faf31d4702e3d3e37b9b587f59fbfb892a39f41da8bf5660ca36

Observation cb3a4fe5-4b16-4954-8e1f-3e346d6c43fa · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.525766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.612239Z digest=sha256:f1015e343c3ecb711415e86d206f4e6fac41701d42bf7b8befee2e3eae08eb61

Observation e6896d1f-d830-4e01-bb3b-e727d0669cdb · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.512304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.617075Z digest=sha256:5bd710c82adfe828e67b6deb0937d35a3cc3647c96c67ad7f842d8ea08ceb940

Observation 4604b74a-3dc0-4a02-810e-741808f2f607 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.498650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.621966Z digest=sha256:3d2033306cfa723295317ebfbada27ca762ab2485db7771b31fc5f8e6a613c37

Observation 7eef8db0-8f5c-4155-beaa-c48496b87f3f · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.483352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.627226Z digest=sha256:85c5a6ffa01e5a6462f0d4b979aff3ad892526414f9c2f2dfe9ac87ac1fb17fa

Observation c327faa7-29f0-4752-a44f-e59524ef7e60 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.468702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.632248Z digest=sha256:f08555de117a29fe28a7acb64e96cf30ba695613e27080ef93161593fe34564d

Observation e4b7d01f-ae90-45de-903b-3630190e1f80 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.453446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.636966Z digest=sha256:3efe6724261be757d777f3bf6283265540b02da1ea578c3d1851e398f6d178be

Observation 72cbe206-ccde-4d35-9645-5921daf7cccf · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.438626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.641426Z digest=sha256:e2caa12b8a663626da424fbb869b319474c2431f1cec36ca3eef1a71c9371963

Observation c69e605d-6355-4977-a6d3-7ab9f20977a8 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.423754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.645907Z digest=sha256:15b06c3dfa3e85e2d21a749045bcca511d578a297412241b588ba14cd541bb5b

Observation ace0df3f-cfcd-4585-9a75-981c6360c71f · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.408609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.650859Z digest=sha256:f3470e6cdadeb3b49c53aaf789a215ab8c2882c3a027b111712466949561327b

Observation 63a5e824-3a9a-4b6e-bfef-1b7d990d91a9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.394323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.655428Z digest=sha256:61f8a72dc6cdb13a51be3c116e8ba6f74723b2fd5c36ac8881445c0675316daa

Observation 054c740e-499b-4282-a1e6-3e626481ffe5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.379530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.659914Z digest=sha256:d69183693365d8b62bdae2b248668ae29543efe0c103ca0dd4c9c7b328facc59

Observation 1be595b1-c095-4013-8b28-87e8dce80b7b · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.365014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.664511Z digest=sha256:fb32300c46195fe01667262a87e1e09d0f92198447b304bd7ddeb3395893e3d7

Observation e5002ea4-d56b-4fc6-8726-2f32aae58328 · outbound

This paper cites Clean and jerk.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Clean and jerk

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.351397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.668719Z digest=sha256:9b5ae9a4b6de546605248ab5cd7aec4fcc2f15caa4f571e3ca967c050b3641f1

Observation 92631da5-4954-48f6-ac02-b668ce0ea1e9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.337236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.673142Z digest=sha256:1ca9ea180ab5c795b681ce11cbd3ea279614ad03658ab55c02442425817f26c1

Observation 2b06b85c-74b6-4174-90fb-c06b0342bb72 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.322517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.677555Z digest=sha256:aa4702a369de8e7865fe6da47d5bb781c6e4f2aa8a6f5eece6c4166897610a15

Observation 76a85673-baba-4a35-ae23-0d898b89b1ed · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.307906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.682303Z digest=sha256:b3d2af7fee0ce66e4e3c324df701cfa3d5ac261ac98167f7bbd40178622e2323

Observation d26f89f6-a4c9-4616-820a-bb6ed704e2b6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.292915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.687056Z digest=sha256:7fc8cdeaf768ef0a52b37bbc73af41a9c5b7d286bebdfbaa48edc89b016f890b

Observation 9bbbd63d-cf7b-4899-92c6-cc284b8693bb · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.278314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.692402Z digest=sha256:6ecfa51c74047c483cd3ac2d4898b6a12f55005c7dbac1c24cb0c0fb3c361e13

Observation 97282e13-6b49-4ac9-890d-c36f39908512 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.264381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.696907Z digest=sha256:309105065a8db3feead5be363a48b7a283f311f148171466c0fd12146c77955d

Observation aa8e3779-b8cb-4e43-9620-c1c47bef0f37 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.249948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.702217Z digest=sha256:79de22157df9b6e68593896d90a644c36aea6231199914835fabd32b5a11dbaa

Observation c19b2ae5-1b73-46c3-b185-870babd944b3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.236238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.706824Z digest=sha256:f59ca8160e61e66a9d7479c20613b41819d5e43944a926808591b93e60c278f4

Observation 6bccac1b-34ee-4e92-bd23-0208af35b7d6 · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.221966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.711922Z digest=sha256:7ead921b79f67c1d918bf0994c06f59965caeb200cb8ee58a093a53caae9cb94

Observation 2ac87c45-565a-470e-bc54-2ec053dfc9bf · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 63

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.207027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.716236Z digest=sha256:35d4c944a8f2b5b56b3d29b21ac69b2fc3852fe7a49dc80c393aa58a906373cf

Observation 9dbecae1-f80c-41ec-8178-5857411271f3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 64

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.191558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.720612Z digest=sha256:151a1ca0d18d0b1de75a09ec52a4767cafb8d27dcfd94bcc39ffa4062237ee50

Observation 88f8cd3e-9d23-44be-b1b4-61b21f20c193 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 65

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.177192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.724883Z digest=sha256:fa059ed1f503c276e7fa97687d6ed1a5131abab8a148f2832871840b43e34eec

Observation 994ac44b-c81e-47e8-8238-0f61e20ab453 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.164074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.729353Z digest=sha256:229d3506c838d3d4986c112339277f1fe7b2899728dca300f76cf4959007a439

Observation 5383e4d2-5546-4226-bd41-b6efc58cadb6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.150134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.734000Z digest=sha256:3ded377e9c9d1a703e59978e0d3d336fd65564579bdcf089bd64b86b09b52e08

Observation 56de4781-f5b4-4ce5-9357-2f6d7cd7ac92 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.136455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.740225Z digest=sha256:6a410e62e610045e08a2bdc7b38bd91a436f8f519834982b8e7dff034df45737

Observation a976d8ea-9993-4140-93ab-85344d7a0622 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.121130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.744329Z digest=sha256:b893380f03d1811648bc9ba1f003b53b24061b3471afe6912a0ee43730dc4246

Observation 8ff0ee90-0967-4aa0-a5bb-ad98b5030455 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.106517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.748776Z digest=sha256:f614926b5b038c65fa3d0a29d6df89f944c34fedb057326fbc9253d5730cce89

Observation 600e6fb0-18a8-43d1-a4f0-ff57bbd78436 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.090831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.753751Z digest=sha256:57429b8587353a88eed7511784769331c628c470aabaa04009b4ea693a5318b9

Observation 1947000e-f008-4154-9278-f7db4966acf2 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.074870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.757928Z digest=sha256:b8592c6fd4edc4a9821196e734ddcab5261cff3aff13f311e73e9f41129995f1

Observation d4832963-bdfc-4ba9-a2ca-c9703a2967af · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.058938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.761853Z digest=sha256:754ef67822a0e51766ae1590e5f47f6f1c3f656f283705a92d92b594d68c8c1e

Observation 0b614176-2cc3-4a8a-86d9-09896f602c06 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.043397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.765894Z digest=sha256:8b7df9fc6cd0432991d06575c4acf7d176e3319d2d79a7fae02d88e78c8c4716

Observation 0f47195c-4b35-4bdf-b4be-2eeb045cc8ea · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.027022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.769936Z digest=sha256:7e5893807b20eb7bb391890411cc66604b710948493f9298e655d582f0a94f9c

Observation 62da1bdf-4bef-463b-9509-9bc11828d69d · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.009686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.774046Z digest=sha256:3775c4bcd0021cae6fc3361d470d8b53d8e99bc12bb488f9e7f4d34d74234b8b

Observation 918abf29-90e5-486d-a3d9-c4483d0ac1c6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:35.994903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.777983Z digest=sha256:b90fe443cf52e1c5b523e1399358f93d83cf9e9d842ceb98fa9066e10b05b484

Observation 16f78237-19b5-48a5-8385-a8228b15cbf3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:35.980916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.782355Z digest=sha256:2e9f6c7e3ec733d0d4f3f690e2cce89ccb521bb53402ff8555260d48f2a95ed6

Observation 2927d39a-e0fd-471f-83c9-e617c65621fd · outbound

This paper cites B Details of Datasets and Evaluation Protocols Datasets We conduct the training process on Kinetics-400 (Kay et al.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP B Details of Datasets and Evaluation Protocols Datasets We conduct the training process on Kinetics-400 (Kay et al

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.965754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.786825Z digest=sha256:372e9a8031a537d99106a4fa55b4152785f0202be48bc411f800bcffd93391b6

Observation 20acf465-aae6-4172-8886-cee4e5fbbbc9 · outbound

This paper cites jump”, “kiss.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP jump”, “kiss

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.950023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.791432Z digest=sha256:e45522ad4a30fd2806e497cd33b5be9e4ffbe243f6df4112e8a6009b5ef0d068

Observation 3af675e4-82c9-4254-8a9b-b2d059eaab1d · outbound

This paper cites This can be explained by the fact that larger temporal scales result in sparser interactions for boundary frames during channel mixing.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP This can be explained by the fact that larger temporal scales result in sparser interactions for boundary frames during channel mixing

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.932844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.796320Z digest=sha256:5d493f270a7333c5ff0d1fabcf07f671e7614c80226905728cbdda9b4289155d

Observation 8cae5fb0-5c0b-4c5a-8374-82e7c0d27291 · outbound

This paper cites In 2011 International conference on computer vision, 2556–.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP In 2011 International conference on computer vision, 2556–

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:37.042145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.435691Z digest=sha256:8605b517bb97ac8e60c814559c538d6c76bf63f87c7d1546db63d9df4580f3b8

Observation 3649627a-c22f-4a9b-9bf1-6911b23e172c · outbound

This paper cites The Kinetics Human Action Video Dataset.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP The Kinetics Human Action Video Dataset

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.431226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.431226Z digest=sha256:339d24e7fdfd354019a1d9c9b885de4249e259aa33eb9002ed5ede0f84ed15a7

Observation f095a1d2-2ee5-49c7-926c-d9e245a1484c · outbound

This paper cites InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 4613–4623.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 4613–4623

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.420099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.420099Z digest=sha256:b614ca3d4fb9a960241a28bc756225b4db7b197e733a17f53eec6aa641fc32f3

Observation 37347fda-c884-4304-88e1-dc77d1cbd936 · outbound

This paper cites GPT-4 Technical Report.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.414650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.414650Z digest=sha256:02e00e439fcc1bc3dcc355b7ead65f11fb58b5518b636cfe2830e916627ac45a

Observation 55190b9c-e992-4a7c-9057-f09fae1fba50 · outbound

This paper cites Lee, D.; Lee, J.; and Choi, J.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Lee, D.; Lee, J.; and Choi, J

Reference 2563

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:37.025080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T16:41:35.439885Z digest=sha256:c80203bc38445ae38bd9346d652dadc1dce3db7efd62f521a23efa749908fa1a

Pith citing papers

Observation a789da5b-9901-4a69-8fca-e9115415c18e · inbound

Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition cites this paper.

Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:55:35.022671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:55:34.614403Z digest=sha256:f2a14bcac18be6bfae2584a39b4f220d7b8fa1a2c88f83879e271288e8fdf397