Pith. sign in

Paper Citation Record · LEDGER

Conformal Predictions for Human Action Recognition with Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2502.06631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06631 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:57:12.636325Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:57:12.479732Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T14:57:12.816902Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0696ad96-3d03-4398-baa9-416714377ac5 · outbound

This paper cites Conformal Predictions for Human Action Recognition with Vision-Language Models.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions for Human Action Recognition with Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:57:12.824486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.479732Z digest=sha256:df8982b0ffbfea6c6d707c8c38eb0b33387522534c9e3d2d5485c9be7f4c06d1

Observation 14ddd180-af64-41a0-8915-b6c19d18bcd8 · outbound

This paper cites Conformal Predictions Providing reliable confidence estimates for predictions made by deep learning models is essential in many applications.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions Providing reliable confidence estimates for predictions made by deep learning models is essential in many applications

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.269711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.486097Z digest=sha256:970bdf7af467b2d32ab9f4ce397306d74890c5f1df5ace3814bad7d87a5a3282

Observation 937c3d2d-80a0-4e38-ba04-34cc503621e2 · outbound

This paper cites Experimental settings We conduct our experiments using three video clip datasets: HMDB51 (51 classes) [23], UCF101 (101 classes) [24] and Kinetics400 (400 classes) [20].

Conformal Predictions for Human Action Recognition with Vision-Language Models Experimental settings We conduct our experiments using three video clip datasets: HMDB51 (51 classes) [23], UCF101 (101 classes) [24] and Kinetics400 (400 classes) [20]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.253420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.491485Z digest=sha256:8fcd698f33738604035b66f5a73295abbbe34dfe54657e1116166a9081f0e905

Observation a4dc4bf8-b50e-4c10-974f-14b4c4860548 · outbound

This paper cites Red dots mark 1/τ∗, our estimate for minimizing tail size, while green dots indicate the true optimum 1/τopt when it differs.

Conformal Predictions for Human Action Recognition with Vision-Language Models Red dots mark 1/τ∗, our estimate for minimizing tail size, while green dots indicate the true optimum 1/τopt when it differs

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T14:57:13.222306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.502192Z digest=sha256:8a5d851ac6a568ea22104067ea825c3c6ba3a7ef5b8fa06435fefcce40302c54

Observation d271ec38-87e0-4102-b171-7f46843b75e1 · outbound

This paper cites an unresolved cited work.

Conformal Predictions for Human Action Recognition with Vision-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:57:13.205996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.507668Z digest=sha256:4910df7e2505b7429db6602a80dfd31400d741334a2510f32d7a3bb6dbf7958e

Observation 26d32a1e-c5df-4c3a-9808-2337a5e9288e · outbound

This paper cites Behavior recognition via sparse spatio-temporal features,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Behavior recognition via sparse spatio-temporal features,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.108366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.538451Z digest=sha256:51dafa58c19ef303b01ead8c7448c2e7d1b07c69ff95ad2a2f283769e35ee925

Observation e53a7350-a19a-46af-bf66-01c5b696be03 · outbound

This paper cites Fast user-guided video object segmen- tation by interaction-and-propagation networks,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Fast user-guided video object segmen- tation by interaction-and-propagation networks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.190000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.512849Z digest=sha256:04c77d6c03c234eb1a1b00ef8e44544c369a671aa8c744abf34a5658bfb8ce19

Observation d623cd3e-30ec-4323-b602-6198080154fa · outbound

This paper cites Human-in-the-loop vehicle reid,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Human-in-the-loop vehicle reid,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.173741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.517839Z digest=sha256:1ad4b45101d105ed0fa5ae712a9de413e07bd959bd99626c9293746f1987d7c4

Observation db002e33-79c5-412c-be50-1a43bcc1e030 · outbound

This paper cites Surveillance video querying with a human-in-the-loop,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Surveillance video querying with a human-in-the-loop,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.157968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.523018Z digest=sha256:e39b84e9846bdb5ffaf9adeddaed48fd625098672f2d71ce14059ab91993ee45

Observation 2bf64a99-9bb0-400a-8ffc-ef1fb4a94a89 · outbound

This paper cites De- signing decision support systems using counterfactual prediction sets,.

Conformal Predictions for Human Action Recognition with Vision-Language Models De- signing decision support systems using counterfactual prediction sets,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.140776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.528008Z digest=sha256:d035ba094b76ba12bb399863540ccfa2e03bc6a5a9ecff972329d1927fdab63f

Observation 983ee99c-06c2-4e4e-a30d-dc7291c60a6e · outbound

This paper cites Conformal prediction sets improve human de- cision making,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal prediction sets improve human de- cision making,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.124735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.532984Z digest=sha256:a1d2de78099c62d5c358ec8dcd00295e96c3087bfe92455ae18f53c72d90a6b0

Observation b0c089b6-f903-40ce-836e-ad89ca43de0b · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Learning transferable visual models from natural lan- guage supervision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.027326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.568095Z digest=sha256:6a1a15f9911000167da07fe38e2b9d93829773257b6da53cf85dde10c7898f16

Observation e0520cae-d70e-437d-9a80-e678ddd82761 · outbound

This paper cites Temporal segment networks: Towards good practices for deep ac- tion recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Temporal segment networks: Towards good practices for deep ac- tion recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.092733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.543548Z digest=sha256:2821a1f7f56e2c082e94ad219cddc5e0c91191144957b19fdb6e71dc8c760759

Observation 40a9001d-cfa5-414b-892c-87b68362b7ef · outbound

This paper cites Slowfast networks for video recog- nition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Slowfast networks for video recog- nition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.076522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.548509Z digest=sha256:a165e92860827b4d19f1bd07ee3b430848709c5035f2a8cc53801a502cdcfaef

Observation f5c9e6ea-c327-487b-b214-588895d3b2a0 · outbound

This paper cites Stm: Spatiotemporal and motion en- coding for action recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Stm: Spatiotemporal and motion en- coding for action recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.060036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.553338Z digest=sha256:7cb87eef721865c91a33971349c0dbe7909c273d61066cf1215daf35cc6c0b3c

Observation 1c500f2c-4430-4e4f-8aa2-aea9dcb298ac · outbound

This paper cites Vivit: A video vision transformer,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Vivit: A video vision transformer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.043979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.558219Z digest=sha256:7c1fa8f7ab182046a14f9c4d3b2a202e48ae6e9ba3a815f4a7d12d0cb6fed18d

Observation db75da33-556d-40dd-8625-aa56000599e1 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Conformal Predictions for Human Action Recognition with Vision-Language Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.562855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.562855Z digest=sha256:dd51260b9dd2be7aabba69f9192e3fe64486cedf4fbffeb7e3c5eb6b6bffb9cb

Observation 7f5472a1-8231-4ad2-ba2e-fc17ed0bcccf · outbound

This paper cites An electronic nose-based assistive diagnostic prototype for lung cancer detection with conformal prediction,.

Conformal Predictions for Human Action Recognition with Vision-Language Models An electronic nose-based assistive diagnostic prototype for lung cancer detection with conformal prediction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.944526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.597946Z digest=sha256:1a393b7db0e3d934b21f2bdfb83a66a5b88d926138589283ed00de2125367edf

Observation b923805b-bfda-4892-b4ee-f3ecc20ef08e · outbound

This paper cites Are foundation models for computer vision good conformal predictors?,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Are foundation models for computer vision good conformal predictors?,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.572903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.572903Z digest=sha256:908832baa41bee9fcc7c89a7d163b87a040f46237c5b4559577722c771a36f54

Observation 1ed4e0fd-0643-484b-98cf-27143aa82821 · outbound

This paper cites On the rate of gain of information,.

Conformal Predictions for Human Action Recognition with Vision-Language Models On the rate of gain of information,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.010707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.577660Z digest=sha256:a7bd45fd45098d78c82b96184e54118eb42945477dbdb72a160c68650a727ce0

Observation a24f7d33-6532-45e0-b602-fb6ae52e3661 · outbound

This paper cites Selection from alphabetic and numeric menu trees using a touch screen: breadth, depth, and width,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Selection from alphabetic and numeric menu trees using a touch screen: breadth, depth, and width,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.994010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.583346Z digest=sha256:97799faa1494eeb31ae5e2598a1046175da3a969716e54e42148c675d919a92f

Observation da4e27eb-11d8-4d45-bc02-038e32f85a63 · outbound

This paper cites 29, Springer, 2005.

Conformal Predictions for Human Action Recognition with Vision-Language Models 29, Springer, 2005

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.977618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.588442Z digest=sha256:48d03d258bc88c90f89b45f542b9d53f18f8636ba00381a4e9fc57cea1e8136e

Observation 721a98cc-f580-44f3-ad99-72371ade35c7 · outbound

This paper cites Least ambiguous set-valued classifiers with bounded error levels,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Least ambiguous set-valued classifiers with bounded error levels,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.961553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.593351Z digest=sha256:6987a76e7a7883dd63ecf27033b86b071a8ed37871c17ea313595e517c8cbb2a

Observation 6646e6a3-2c74-42a6-8ec3-0cdf63832d99 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Conformal Predictions for Human Action Recognition with Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.626354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.626354Z digest=sha256:13d63e3135d7cfb895eb6bf1409962ab0f27eaee32af9e665ff0ff04977e59e4

Observation e8f59025-c1f8-483f-9807-cf3da5364aad · outbound

This paper cites an unresolved cited work.

Conformal Predictions for Human Action Recognition with Vision-Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:57:13.237883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.496797Z digest=sha256:1522c9488b5aa549ea896e5a0b548ced9a390dd48dbb4984de2611bddd32e29f

Observation b09704fd-e233-4b1d-a05d-12750ded83d7 · outbound

This paper cites Dense trajectories and motion bound- ary descriptors for action recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Dense trajectories and motion bound- ary descriptors for action recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.928675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.602719Z digest=sha256:9acb51f27975947bdcd86ce95bb8257ebcacbb1e7d8f072a2265f0bcc69ba70e

Observation bf0d9c6c-9e8d-4b62-9611-04c033e18da2 · outbound

This paper cites Quo vadis, ac- tion recognition? a new model and the kinetics dataset,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Quo vadis, ac- tion recognition? a new model and the kinetics dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.910729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.607464Z digest=sha256:75f1698ff7eb2c21694d607a40db767bf29a4ef6c410ddd7f99dc318f83759d1

Observation 0eea96b6-72fb-453c-84f4-41be3bb39538 · outbound

This paper cites Expanding language-image pre- trained models for general video recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Expanding language-image pre- trained models for general video recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.894085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.612212Z digest=sha256:2c983a84f98609b52e6d1c7613f32f5691b34e953b7808bd4002ade5851a2d66

Observation da2774e8-bb9f-449c-a0f4-21fb6629e4d9 · outbound

This paper cites Fine-tuned clip models are efficient video learners,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Fine-tuned clip models are efficient video learners,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.876493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.616854Z digest=sha256:01c85de7f0d7f67a55a856dca3244a9ea00305b13ec87cbfea2b25e9043fa2ac

Observation 68c96306-e637-4d2b-8e1e-a2fb307bdab1 · outbound

This paper cites Hmdb: A large video database for human motion recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Hmdb: A large video database for human motion recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.858847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.621764Z digest=sha256:31ec61d959ea21feba8ec056860d477bb64e919003d22c5d4b23ef325744bbbf

Observation 3f56ef6a-da66-46bb-ae9a-b5593bf1c613 · outbound

This paper cites On sequence learning models: Open-loop control not strictly guided by hick’s law,.

Conformal Predictions for Human Action Recognition with Vision-Language Models On sequence learning models: Open-loop control not strictly guided by hick’s law,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.842209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.631485Z digest=sha256:527d3ae1eb70dde1ad25fd074966e1f796afb1b9da4a0960fbb4547416df7350

Observation 6bf53bbe-ea34-459b-95a4-8b2c145abedf · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

Conformal Predictions for Human Action Recognition with Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.636325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.636325Z digest=sha256:7b4557f3f99df59966dd874a178494bf8fa69c68cab6ba837a47aa9caa17452f

Pith citing papers

Observation 0696ad96-3d03-4398-baa9-416714377ac5 · inbound

Conformal Predictions for Human Action Recognition with Vision-Language Models cites this paper.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions for Human Action Recognition with Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:57:12.824486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:57:12.479732Z digest=sha256:df8982b0ffbfea6c6d707c8c38eb0b33387522534c9e3d2d5485c9be7f4c06d1