Pith. sign in

Paper Citation Record · LEDGER

Conformal Predictions for Human Action Recognition with Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2502.06631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06631 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:57:12.636325Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:57:12.479732Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T14:57:12.816902Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0696ad96-3d03-4398-baa9-416714377ac5 · outbound

This paper cites Conformal Predictions for Human Action Recognition with Vision-Language Models.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions for Human Action Recognition with Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:57:12.824486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.479732Z digest=sha256:58e3b0d3b427efa4fda738117d5cc3a3e17a8ffa833b87b6186d7fdada156b99

Observation 14ddd180-af64-41a0-8915-b6c19d18bcd8 · outbound

This paper cites Conformal Predictions Providing reliable confidence estimates for predictions made by deep learning models is essential in many applications.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions Providing reliable confidence estimates for predictions made by deep learning models is essential in many applications

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.269711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.486097Z digest=sha256:f44353b63c0fb1d3154453b37e5d8ee16bae119606c51ab77aad5d1c4c247499

Observation 937c3d2d-80a0-4e38-ba04-34cc503621e2 · outbound

This paper cites Experimental settings We conduct our experiments using three video clip datasets: HMDB51 (51 classes) [23], UCF101 (101 classes) [24] and Kinetics400 (400 classes) [20].

Conformal Predictions for Human Action Recognition with Vision-Language Models Experimental settings We conduct our experiments using three video clip datasets: HMDB51 (51 classes) [23], UCF101 (101 classes) [24] and Kinetics400 (400 classes) [20]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.253420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.491485Z digest=sha256:ccfe35c3d5695d4003ef95b0ba519857ab4b692e029c608714c9107d05e0a722

Observation a4dc4bf8-b50e-4c10-974f-14b4c4860548 · outbound

This paper cites Red dots mark 1/τ∗, our estimate for minimizing tail size, while green dots indicate the true optimum 1/τopt when it differs.

Conformal Predictions for Human Action Recognition with Vision-Language Models Red dots mark 1/τ∗, our estimate for minimizing tail size, while green dots indicate the true optimum 1/τopt when it differs

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T14:57:13.222306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.502192Z digest=sha256:a8e4804742a1916a1f3f3f0733877fcd3d60244d6e596908ed80f467004c23e8

Observation d271ec38-87e0-4102-b171-7f46843b75e1 · outbound

This paper cites an unresolved cited work.

Conformal Predictions for Human Action Recognition with Vision-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:57:13.205996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.507668Z digest=sha256:6f0f28fdb45eb33c6d62506ff2077a01f84042f5d885f84bf3cf1b8593a3146b

Observation 26d32a1e-c5df-4c3a-9808-2337a5e9288e · outbound

This paper cites Behavior recognition via sparse spatio-temporal features,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Behavior recognition via sparse spatio-temporal features,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.108366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.538451Z digest=sha256:6c52846d631b5ace8b1aa18e3ee3ba31eff3ecef0f1478f6c4ee1fc2d8bc53ba

Observation e53a7350-a19a-46af-bf66-01c5b696be03 · outbound

This paper cites Fast user-guided video object segmen- tation by interaction-and-propagation networks,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Fast user-guided video object segmen- tation by interaction-and-propagation networks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.190000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.512849Z digest=sha256:d09631d6212ad87db87cad6b487beeca74dd9a93b61ad16ae033bfb230e9fc5c

Observation d623cd3e-30ec-4323-b602-6198080154fa · outbound

This paper cites Human-in-the-loop vehicle reid,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Human-in-the-loop vehicle reid,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.173741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.517839Z digest=sha256:09cfaf89ecdcbaedd80f10899424ba387ebde4c2d9009fef126efa08049da12c

Observation db002e33-79c5-412c-be50-1a43bcc1e030 · outbound

This paper cites Surveillance video querying with a human-in-the-loop,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Surveillance video querying with a human-in-the-loop,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.157968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.523018Z digest=sha256:6a90bc05411dda79b7b2b501fd18b7a5382f38b6d50a321be5d4368a18349d2c

Observation 2bf64a99-9bb0-400a-8ffc-ef1fb4a94a89 · outbound

This paper cites De- signing decision support systems using counterfactual prediction sets,.

Conformal Predictions for Human Action Recognition with Vision-Language Models De- signing decision support systems using counterfactual prediction sets,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.140776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.528008Z digest=sha256:e71232ec8937e09ced4882e723145ee941570488231b1868afb3bc08b176f077

Observation 983ee99c-06c2-4e4e-a30d-dc7291c60a6e · outbound

This paper cites Conformal prediction sets improve human de- cision making,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal prediction sets improve human de- cision making,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.124735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.532984Z digest=sha256:9f3aba05bda46497950791e53ce1262643c88c72f0e7a5620616b548624e55ef

Observation b0c089b6-f903-40ce-836e-ad89ca43de0b · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Learning transferable visual models from natural lan- guage supervision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.027326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.568095Z digest=sha256:514ec59aa1ae8c6e7b553ade8e26d4c5802d7de8a2137e8a270a1ee4a02b1ca1

Observation e0520cae-d70e-437d-9a80-e678ddd82761 · outbound

This paper cites Temporal segment networks: Towards good practices for deep ac- tion recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Temporal segment networks: Towards good practices for deep ac- tion recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.092733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.543548Z digest=sha256:338c05a79698767bb9ef1b712e3c5602f345be8fa7606c007db9836a00dd4a76

Observation 40a9001d-cfa5-414b-892c-87b68362b7ef · outbound

This paper cites Slowfast networks for video recog- nition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Slowfast networks for video recog- nition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.076522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.548509Z digest=sha256:c48fd89085788fea50896ecf12d428f70725646e8a705c534e663cb53f417fe1

Observation f5c9e6ea-c327-487b-b214-588895d3b2a0 · outbound

This paper cites Stm: Spatiotemporal and motion en- coding for action recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Stm: Spatiotemporal and motion en- coding for action recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.060036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.553338Z digest=sha256:a87201a34381442cc194dffefce3c11f059f00e64643bd012e3540d0c8dc732a

Observation 1c500f2c-4430-4e4f-8aa2-aea9dcb298ac · outbound

This paper cites Vivit: A video vision transformer,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Vivit: A video vision transformer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.043979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.558219Z digest=sha256:c306718a6fbaeda72603f382f686e9bf69cbdacc405378c1c9674c33585c8730

Observation db75da33-556d-40dd-8625-aa56000599e1 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Conformal Predictions for Human Action Recognition with Vision-Language Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.562855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.562855Z digest=sha256:5e7c29d70dbdec4fae7be54508e8fa8d1ca7ae7f0f524ce00131b99043f38fc8

Observation 7f5472a1-8231-4ad2-ba2e-fc17ed0bcccf · outbound

This paper cites An electronic nose-based assistive diagnostic prototype for lung cancer detection with conformal prediction,.

Conformal Predictions for Human Action Recognition with Vision-Language Models An electronic nose-based assistive diagnostic prototype for lung cancer detection with conformal prediction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.944526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.597946Z digest=sha256:ca6b715b8b3c7f38258c2f340350fa6ff600e7bea4bef42f70964c25ae74b96a

Observation b923805b-bfda-4892-b4ee-f3ecc20ef08e · outbound

This paper cites Are foundation models for computer vision good conformal predictors?,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Are foundation models for computer vision good conformal predictors?,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.572903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.572903Z digest=sha256:dc316b4a5d359fcf312014a2ab1540e0e0bd6d65a9ba69bb8a5c90a129483208

Observation 1ed4e0fd-0643-484b-98cf-27143aa82821 · outbound

This paper cites On the rate of gain of information,.

Conformal Predictions for Human Action Recognition with Vision-Language Models On the rate of gain of information,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.010707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.577660Z digest=sha256:6897c89394f2ba320431a042232412fa22a0780d35b47167756b96738436ca05

Observation a24f7d33-6532-45e0-b602-fb6ae52e3661 · outbound

This paper cites Selection from alphabetic and numeric menu trees using a touch screen: breadth, depth, and width,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Selection from alphabetic and numeric menu trees using a touch screen: breadth, depth, and width,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.994010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.583346Z digest=sha256:35fc95e3e7035b16d654728159d3a235b84bb5b28e1b11fe82f3649097a1525d

Observation da4e27eb-11d8-4d45-bc02-038e32f85a63 · outbound

This paper cites 29, Springer, 2005.

Conformal Predictions for Human Action Recognition with Vision-Language Models 29, Springer, 2005

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.977618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.588442Z digest=sha256:fcab07d5034996f9bcbe98453de1ec7d47484d4810606562b2807c4718150ef3

Observation 721a98cc-f580-44f3-ad99-72371ade35c7 · outbound

This paper cites Least ambiguous set-valued classifiers with bounded error levels,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Least ambiguous set-valued classifiers with bounded error levels,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.961553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.593351Z digest=sha256:de305cf5300a84e9196386d7cdfa09123097269b1155582fecb19b7966a52640

Observation 6646e6a3-2c74-42a6-8ec3-0cdf63832d99 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Conformal Predictions for Human Action Recognition with Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.626354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.626354Z digest=sha256:547d16ce8309bd1ae707d5d748b7be4ea47c4592200f39c8244b33be4e4caae3

Observation e8f59025-c1f8-483f-9807-cf3da5364aad · outbound

This paper cites an unresolved cited work.

Conformal Predictions for Human Action Recognition with Vision-Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:57:13.237883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.496797Z digest=sha256:c9a48f970493732099c11987d3bcf5c2d14460edebef76a4586f1b0f54e95fed

Observation b09704fd-e233-4b1d-a05d-12750ded83d7 · outbound

This paper cites Dense trajectories and motion bound- ary descriptors for action recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Dense trajectories and motion bound- ary descriptors for action recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.928675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.602719Z digest=sha256:d9a0404e12f5a228b5cc90598b936125471dc58ea85bfcce4fccb933dce49984

Observation bf0d9c6c-9e8d-4b62-9611-04c033e18da2 · outbound

This paper cites Quo vadis, ac- tion recognition? a new model and the kinetics dataset,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Quo vadis, ac- tion recognition? a new model and the kinetics dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.910729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.607464Z digest=sha256:576684cf761a7b24b473fff1a316ed256c697b017c2c93f52813ed07455eb884

Observation 0eea96b6-72fb-453c-84f4-41be3bb39538 · outbound

This paper cites Expanding language-image pre- trained models for general video recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Expanding language-image pre- trained models for general video recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.894085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.612212Z digest=sha256:89bae3133ff8f0e102fef61682c438d1bffc5ef8160f5fd86b6055f6beff0fc3

Observation da2774e8-bb9f-449c-a0f4-21fb6629e4d9 · outbound

This paper cites Fine-tuned clip models are efficient video learners,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Fine-tuned clip models are efficient video learners,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.876493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.616854Z digest=sha256:504b8ce5933821444c3e9101c40e7503e98b5c3ba59e676f5c8297def71b4ece

Observation 68c96306-e637-4d2b-8e1e-a2fb307bdab1 · outbound

This paper cites Hmdb: A large video database for human motion recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Hmdb: A large video database for human motion recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.858847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.621764Z digest=sha256:031ed4097688c7f983b2c270cfdc6342dc7ba16ea8691b04bb6eb707eb4dcc6c

Observation 3f56ef6a-da66-46bb-ae9a-b5593bf1c613 · outbound

This paper cites On sequence learning models: Open-loop control not strictly guided by hick’s law,.

Conformal Predictions for Human Action Recognition with Vision-Language Models On sequence learning models: Open-loop control not strictly guided by hick’s law,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.842209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.631485Z digest=sha256:b8e3da9461f42e70180f49df73e072b0f3a3a050b7f59296458d6d19b3de2a42

Observation 6bf53bbe-ea34-459b-95a4-8b2c145abedf · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

Conformal Predictions for Human Action Recognition with Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.636325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.636325Z digest=sha256:cb75e79087078dbf9bb86f03dc1e1d275f31a7f7d4e6321eff48479e2f0f1aa2

Pith citing papers

Observation 0696ad96-3d03-4398-baa9-416714377ac5 · inbound

Conformal Predictions for Human Action Recognition with Vision-Language Models cites this paper.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions for Human Action Recognition with Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:57:12.824486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T14:57:12.479732Z digest=sha256:58e3b0d3b427efa4fda738117d5cc3a3e17a8ffa833b87b6186d7fdada156b99