Pith. sign in

Paper Citation Record · LEDGER

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos

As of 18 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2411.15628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15628 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:08:42.308404Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy54
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bc9bc2a2-bd41-45fd-a3e5-6572c2935db5 · outbound

This paper cites GPT-4 Technical Report.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:40.821947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:40.821947Z digest=sha256:f16461462ba5576e6b0175c7f36d281fbe473c66d76ca833dfc4531403d94adf

Observation eb663771-06c0-4920-b6a0-434d580c9460 · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.540893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:40.888612Z digest=sha256:26dfb608fdb2ac6f9f3aff1a3c7a3beeea8298ecf0d4b573b4052709a21223b5

Observation 96065e79-07ec-478b-a35d-a020b140eec8 · outbound

This paper cites Exploring synonyms as context in zero-shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Exploring synonyms as context in zero-shot action recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.472400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:40.995780Z digest=sha256:25dc472e754db4783aa1882632434c6f64a9d2037f417409869afa92df8e56e1

Observation 84fa500b-d883-473c-b707-e4b8e2c44fa2 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Hiervl: Learning hierarchical video- language embeddings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.425059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.016803Z digest=sha256:368845386cdeea11cc0fd12267b79ed05e7192bb4b2ccec142802c29322167bc

Observation 732405b6-5ed4-4b7a-9f92-fae80af0857d · outbound

This paper cites The ikea asm dataset: Understanding people assem- bling furniture through actions, objects and pose.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos The ikea asm dataset: Understanding people assem- bling furniture through actions, objects and pose

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.409979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.022207Z digest=sha256:81b1296c34360c7a78e13fd5c784fb6fd1181cb0ec68d0cf284cbeb08f81b6d2

Observation e7551e14-bab1-4879-b945-d862870784d2 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.396127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.027101Z digest=sha256:ecb768a440a3119bb53fe50cc88a84f238f428c7be2120f4bc02b3d3de78ad93

Observation 80e0cd83-12ec-4d51-a6d1-3e5312a09955 · outbound

This paper cites Rethinking zero-shot video classification: End-to-end training for realistic applications.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Rethinking zero-shot video classification: End-to-end training for realistic applications

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.381534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.083324Z digest=sha256:7032c9b706683033dfe63a370f26fe29c4991cc5eb0a919f1ec65b8966921ba9

Observation ff4d3c15-0b5a-44a9-8445-fed95ac1fa51 · outbound

This paper cites Regen: A good generative zero-shot video classifier should be rewarded.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Regen: A good generative zero-shot video classifier should be rewarded

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.299158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.090381Z digest=sha256:1e4526a3dc3821ca19291021d432d595126e65b3a6a3942c65edcc12d8b99d4c

Observation 9a249bd3-7f50-4a03-b916-3db3601cda56 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Quo vadis, action recognition? a new model and the kinetics dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.285381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.094991Z digest=sha256:9d4b3b2bb7dacb3c4a2402f161360dc5c3bcee17248a140048a5da04f0885a6d

Observation 74109dde-6953-4ac5-bf01-dbd2129eacc2 · outbound

This paper cites Elaborative rehearsal for zero-shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Elaborative rehearsal for zero-shot action recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.268691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.101071Z digest=sha256:989de9804c743f8f7e5b102c811703c7b9ab0738a870776e627936326eea61b7

Observation b606cf03-a89a-4827-a2cf-a92ae546cf73 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.148077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.148077Z digest=sha256:8ece7be7fe1414ab9f05b1d837aa2be6694b2638a69e738a0a2d1087bef1639c

Observation 11326423-3c91-49f8-92de-fb0fedaa099c · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Teaching structured vision & language concepts to vision & language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.157402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.215173Z digest=sha256:75840f63f8b66ab30a3ffb90dc4ba34d18660b3247305d0db9835bd28ef1bc79

Observation a75dd326-4fb5-44bc-8847-7ade7db59dda · outbound

This paper cites Step- former: Self-supervised step discovery and localization in instructional videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Step- former: Self-supervised step discovery and localization in instructional videos

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.142507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.240189Z digest=sha256:e3e5989f42d0e8288de230b2bb97a5c7ee0f7f701e0ffeaba9c66063423a7bfb

Observation 2fec3f59-f902-4123-84b1-945a3c7abf3e · outbound

This paper cites Zero-shot action recognition in videos: A survey.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot action recognition in videos: A survey

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.071841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.244645Z digest=sha256:27172bccd79d1dcb5474b5d3c5af24437b28ca1e1e16ccdd53e7473d8c27f413

Observation 7d0628ee-9900-45d8-868a-4c71a8f9d486 · outbound

This paper cites Ms-tcn: Multi-stage tem- poral convolutional network for action segmentation.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ms-tcn: Multi-stage tem- poral convolutional network for action segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.968430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.329232Z digest=sha256:030829f845d288da2bf28c896579c65af9b05cea0b3a2c5ccf216e4176b19772

Observation bea9a6b6-a870-40a8-bd16-a779afd8df73 · outbound

This paper cites Learning to recognize objects in egocentric activities.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to recognize objects in egocentric activities

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.953527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.421730Z digest=sha256:66471d7ef3c434683c8ba5c703a3f00681795f7b9144bbc864e4abbb2a007c88

Observation d9e4b6be-e75c-4f34-9bd5-bfb91ebfc51d · outbound

This paper cites Slowfast networks for video recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Slowfast networks for video recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.936844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.502547Z digest=sha256:65431ae34091a1833dd8e3ca51369fdeea0a54a12b7b87d56bf9c648b149245e

Observation 9e8b1f8f-6ab6-4707-8485-c0dbb80803cd · outbound

This paper cites Prego: online mistake detection in procedural ego- centric videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Prego: online mistake detection in procedural ego- centric videos

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.921960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.535699Z digest=sha256:f8aa334813270bd0602870fbc0e69749e27235734212e247388a1ee68e074b76

Observation 1d271bf2-ada1-4cb9-81a9-db6c660651a9 · outbound

This paper cites Improving zero-shot gen- eralization and robustness of multi-modal models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Improving zero-shot gen- eralization and robustness of multi-modal models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.821681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.558624Z digest=sha256:cd8ad0d83b6386ad43ce038f39c1610614e9d5d130e65ac8ba9b95ae038971e0

Observation 0d3393ef-0247-49d9-9225-eb6b77159975 · outbound

This paper cites Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.769725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.563967Z digest=sha256:154a496c5ee22bccd09984c4f71f6120b72b1f51f687aa1c49c6a85f4e665876

Observation 7263b4e0-2589-4854-b751-0c54a2799b8f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.754663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.569698Z digest=sha256:9b5ae890787566929f4d050de2d93f3b9684871c763faec2bbca01024f38c0d6

Observation 90c21025-a295-4590-91b0-ee49897d6703 · outbound

This paper cites Temporal alignment networks for long-term video.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Temporal alignment networks for long-term video

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.620096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.620096Z digest=sha256:df5d640cb67a718b7abd47cfae2235ffc88ad7ac908fb7efcc50f09532dad98b

Observation fa393e83-200d-46a6-825e-39e6971c9c23 · outbound

This paper cites Probing Image-Language Transformers for Verb Understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Probing Image-Language Transformers for Verb Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.680083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.680083Z digest=sha256:b658d43deec1b367a3666a229a049c15c84e72626126905b1a8d3b4836ed4965

Observation ac89540c-787a-4938-a1f6-9275e2aa9601 · outbound

This paper cites Clover: Towards a unified video-language alignment and fusion model.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Clover: Towards a unified video-language alignment and fusion model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.722893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.709050Z digest=sha256:a4a27db7080bac6e9bb874040de3d3dcade1373105be59a86dc6636fbada2e8a

Observation be5709eb-168f-4b2c-940d-cbf8a686689b · outbound

This paper cites Fine-grained generalized zero-shot learning via dense attribute-based attention.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Fine-grained generalized zero-shot learning via dense attribute-based attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.639994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.714001Z digest=sha256:0a41ab358aa44e8e6fa703199e21b7b1dcc91b4850b1e0925f85d2171e664093

Observation 74b88be9-9871-4bae-b5b4-d24e6670ea2b · outbound

This paper cites Objects2action: Classifying and localiz- ing actions without any video example.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Objects2action: Classifying and localiz- ing actions without any video example

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.624457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.718983Z digest=sha256:a2e9f2ba16029b2a1ee0e1ef6a84f1d06988429effecc30dbed8a56b88581f42

Observation e9921a13-86a5-40c0-af2e-6a837086a824 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Prompting visual-language models for efficient video understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.606588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.723234Z digest=sha256:72f042b16a1cd9aab614e34b37ffd9b3d50d40e728a6a329d53d8ddd86efb3ed

Observation df167b3c-640b-476e-92fe-196eaf923582 · outbound

This paper cites Error detection in egocentric procedural task videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Error detection in egocentric procedural task videos

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.535690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.783002Z digest=sha256:867ef2d70c53ab5b00d5fb54cedf2d413827d2d70fc60d673cfb9058689661e3

Observation e749117b-00d5-444c-995e-e6de40abf633 · outbound

This paper cites Align and prompt: Video-and-language pre-training with entity prompts.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Align and prompt: Video-and-language pre-training with entity prompts

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.468170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.825070Z digest=sha256:5ed7bc1a63921eec49002c6a92a41b3934acf588adcca99a52ca44bff2fcb92d

Observation 2a32aef5-e3fe-42f0-bfeb-c374ee904f9f · outbound

This paper cites Cross-modal representation learning for zero- shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Cross-modal representation learning for zero- shot action recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.453055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.829748Z digest=sha256:88eab71576bc4721396cd62a74f6dbc16d855cdd6f8fce2188afc209565f4894

Observation 8b00805e-9d19-4f0d-93f4-716827babb06 · outbound

This paper cites Match, expand and im- prove: Unsupervised finetuning for zero-shot action recog- nition with language knowledge.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Match, expand and im- prove: Unsupervised finetuning for zero-shot action recog- nition with language knowledge

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.354680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.834114Z digest=sha256:e058b4d5cf2c3ee0a896dd4b761fd221703aee382491e850497a2084193f301c

Observation bb8f3bb8-0369-4a79-ba98-800b5391f2a3 · outbound

This paper cites Learning to recognize procedural activities with distant supervision.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to recognize procedural activities with distant supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.838632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.838632Z digest=sha256:bccc36326d35af723fd1435405f5cd4fcfa9609045dd78b8c3262027f0ab0781

Observation 56a4e60a-3ba5-486c-b277-0223da873a3d · outbound

This paper cites Out-of-distribution detection for gener- alized zero-shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Out-of-distribution detection for gener- alized zero-shot action recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.260818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.917429Z digest=sha256:06bb85dc551b0d7a4c018d04bc0e78129f055bf60757210a159a2b37557373b7

Observation f768d3f8-f83e-4dc0-8c91-00a35ec1b7e3 · outbound

This paper cites Learning to ground instructional articles in videos through narrations.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to ground instructional articles in videos through narrations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.922633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.922633Z digest=sha256:d38063d293cb7395d59cbe05b622688cfefab56cd26f709de6c59adb8455e930

Observation 392843a1-d2b0-42bc-a566-d817eda44526 · outbound

This paper cites Object priors for classifying and localizing unseen actions.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Object priors for classifying and localizing unseen actions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.233566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.927079Z digest=sha256:bdac085bd5a799edd0977c9f9ce72fc540811bac9811b275e26c1eae11543875

Observation f7caf8e7-e8af-42fc-8c45-92f4ced659d3 · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.205632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.931494Z digest=sha256:15eb0d10418f45b6a4445ce045e5a705926a304babde6ece4705dc2e7706583d

Observation 5439a55e-bf2d-4057-9040-60008bbf1ecc · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips, 2019.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips, 2019

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.151592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.936116Z digest=sha256:44c341a290ff487df3be9d58b18305f196b667dd6d32cfaedf9098ee44c667f4

Observation d6fba638-7398-40f5-8d77-219f0861028d · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Efficient Estimation of Word Representations in Vector Space

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.941389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.941389Z digest=sha256:ffb12f70817faa98fa85917989f65771317db5971fd2e97ed74bcf169f574667

Observation f35a470c-b545-4034-a694-4679b5b08965 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Verbs in action: Improving verb understanding in video-language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.123711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.946220Z digest=sha256:c656d8f7e061ea7d10f20d15ee98bf8fdf55ef047d5f76fbd94236abb8ef00ba

Observation 366828e0-c3b9-4c9c-a70b-1df0eefb606c · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Spoken moments: Learning joint audio-visual representations from video descriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.028356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.951307Z digest=sha256:5e3ccf46b9235258d7270171fbb4640e10e90a59d6ac8f40d198b9c20a90924c

Observation c20dc994-7087-4b41-9e42-0a02bfd86572 · outbound

This paper cites Zero-shot temporal action detection via vision-language prompting.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot temporal action detection via vision-language prompting

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.007255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:41.996161Z digest=sha256:7f2797a852adefc526b537c9d3cb5dce16fb8c52debcd77c4b22c77ac35ff872

Observation 15acba4f-6f76-4565-bc7e-44d4020b9676 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Expanding language-image pretrained models for gen- eral video recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.896043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.021559Z digest=sha256:ca40226309390cafd23d8651e5b0c2e5373cee884f5df42216dc83d4d0925205

Observation f86fbf1e-492b-422a-8aa7-857260a243b9 · outbound

This paper cites Learning multimodal representations for unseen activities.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning multimodal representations for unseen activities

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.867587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.027762Z digest=sha256:f13ef0bf49118f5b571a59ebb5a97d7e0d4150c7f9026a57f64237b019dfb56a

Observation 4cc7ed0a-b02c-45a6-9cab-ef19423003a3 · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.761687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.032409Z digest=sha256:beda30533f0f0fc3163099202d64c84f78005dcd657cd1505d7ad46d11bdb4bc

Observation 5efc7eac-01e5-4a80-a6f9-02e7aae9ae5e · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classification.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos What does a platypus look like? generating customized prompts for zero-shot image classification

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.043056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.043056Z digest=sha256:6387716c017bf47dd497922addf5cc9bdf54008d706bec3a4abe24994513d73a

Observation 68d4deaf-c8e6-4719-bcb3-44f0b725a8b2 · outbound

This paper cites Alignment-uniformity aware representation learning for zero-shot video classifica- tion.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Alignment-uniformity aware representation learning for zero-shot video classifica- tion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.736529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.090035Z digest=sha256:3790980e80147aced7f9aa9df8b06c0caf3273b555a7a86f3d9ebdbc5bbc4ea1

Observation 08ddf227-795d-407f-a990-67c7f1f1ef5e · outbound

This paper cites Rethinking zero-shot action recognition: Learning from latent atomic actions.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Rethinking zero-shot action recognition: Learning from latent atomic actions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.648754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.101490Z digest=sha256:f1eb067d13b968fb78b70a4d34dfcc32d5af786e230f336eec433ba52b01ba45

Observation 28f404fb-d0db-4909-bc7a-7117967cc0c3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.105973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.105973Z digest=sha256:592b120f31ad79ffd3c3d733a45269652fe29df0d380fd14d474f7db86288a84

Observation d556de65-a245-475d-94c7-ea542c650a66 · outbound

This paper cites Language-based action concept spaces improve video self-supervised learn- ing.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Language-based action concept spaces improve video self-supervised learn- ing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.568529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.110115Z digest=sha256:fec3062a972980e5b15c5761f351c223135be28589e550570f9eb23707225a09

Observation 60c15b2d-c2ce-4b4c-abcd-b2b204a7c148 · outbound

This paper cites Fine-tuned clip models are efficient video learners.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Fine-tuned clip models are efficient video learners

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.480138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.114311Z digest=sha256:06ab00dbffd53f56083a5a9884b740f3189f73d3d9af3087e4e2ea0fa744c0b5

Observation 91ad1310-2332-43cf-933a-e037b138d7c1 · outbound

This paper cites Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.119921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.119921Z digest=sha256:fce0516bfb969655208f9bdecb2ed37e59feb04b27f080cb6f52f4be9103b6c7

Observation 2aaf919a-3816-4afe-bbb6-cd3f35a5bd35 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.414332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.126043Z digest=sha256:30ea59076fc8f6f3f481077d04f313b4d7d40b1622f685573c78fb805fc78b4c

Observation 230c78fd-ddb1-42ee-9656-e57fd4294125 · outbound

This paper cites Mpnet: Masked and permuted pre-training for language understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Mpnet: Masked and permuted pre-training for language understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.398085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.157706Z digest=sha256:cee8b3a81c6a9964ec75f21b5c55151fdfb2312f1b026c77f30de3a7ce468342

Observation 43391558-a256-44fc-9939-dd713ce37841 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.295076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.184990Z digest=sha256:8adb1dc66d57be2db0d28eaa96c85f98c0c8e428789867aa2a9b6a511cdffe43

Observation 04979466-8c37-4949-a8df-3dead8f8a0aa · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos ActionCLIP: A New Paradigm for Video Action Recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.189536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.189536Z digest=sha256:0c640115b14b6e590c1c3965e993ecad46f1c57c05b0ab23572caf1a13c35583

Observation b9aa2de1-0f43-4a3e-9a2b-0f8408ebdcb6 · outbound

This paper cites Vilta: Enhancing vision-language pre-training through textual augmentation.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Vilta: Enhancing vision-language pre-training through textual augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.235796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.193948Z digest=sha256:48448035bc960021f0ab1b722145ad69887eca06c9a5bbb72b97921f6bbec6f9

Observation c871377a-b35e-4725-98b1-44ce9e340e4f · outbound

This paper cites Pax- ion: Patching action knowledge in video-language founda- tion models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Pax- ion: Patching action knowledge in video-language founda- tion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.152810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.198093Z digest=sha256:7c5290a6ab9c2a4e14167904b654cd9a9e84458b7bcb26ceb5e0b040527e7f7a

Observation 002a0166-01aa-4425-8f9b-4ebf5272c2d3 · outbound

This paper cites Zero-shot event detection using multi-modal fusion of weakly supervised concepts.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot event detection using multi-modal fusion of weakly supervised concepts

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.059961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.202643Z digest=sha256:5031a11065b4fe0ce3f00ba606c46a9fbef01abf7b1e14852cf7f4b63f53d997

Observation fa7edda3-578f-463d-942c-58c3ed6cb094 · outbound

This paper cites Cap4video: What can auxiliary captions do for text-video retrieval? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10704–10713, 2023.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Cap4video: What can auxiliary captions do for text-video retrieval? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10704–10713, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.945797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.207694Z digest=sha256:7c84bdc5d3318616c0c59fb0918eaade5ce2c88a642d4a04bbad690f3df82900

Observation d8a3b65a-cb63-4822-9401-d98af018659b · outbound

This paper cites Revisiting clas- sifier: Transferring vision-language models for video recog- nition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Revisiting clas- sifier: Transferring vision-language models for video recog- nition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.879421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.213370Z digest=sha256:3aea070a8b0594f4754018a59b483fd7b596732b54881bdb0b6999f66a48df85

Observation d61984bc-1b77-4ae9-894d-e0c97895b061 · outbound

This paper cites Bidirectional cross- modal knowledge exploration for video recognition with pre-trained vision-language models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Bidirectional cross- modal knowledge exploration for video recognition with pre-trained vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.835387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.227643Z digest=sha256:300d0a852dc2b10a25e00154a327db5b201fd0da515da94adef14050f63dc24b

Observation cb7ad621-edbc-4d41-8701-0aad67f7b9c2 · outbound

This paper cites Generative action description prompts for skeleton-based action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Generative action description prompts for skeleton-based action recognition

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.651926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.259861Z digest=sha256:c7aba69b99dd1cf66f6e2aef335815f5ca832ee81ed79ca7f08d9be1b244b54b

Observation 31beffb7-5a7b-4c4d-b56c-48ca0150bb0e · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.293896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.293896Z digest=sha256:c0e72abb9e35d3bd1f28b73ea943b989108dede32c214e246fb60916462d15a5

Observation cb277bc7-0b42-4946-9107-b6beed78ed43 · outbound

This paper cites Movie genre classification by language augmentation and shot sampling.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Movie genre classification by language augmentation and shot sampling

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.622826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.298672Z digest=sha256:b11e813fadfff584bd25b413859467d04b00bdcd8155e15a43e539f07edcdd4e

Observation 5cbe71b3-35a1-47ab-b07f-abd064bcdac0 · outbound

This paper cites Learning video representations from large lan- guage models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning video representations from large lan- guage models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.566586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.303134Z digest=sha256:509e51188e85d10cbf1f15d7080fb0225b1c54393040a2f660c2bc04dd9c6eda

Observation 79a20398-21bc-4633-ad0f-c8d514f8455c · outbound

This paper cites Learning procedure-aware video represen- tation from instructional videos and their narrations.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning procedure-aware video represen- tation from instructional videos and their narrations

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.457418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:08:42.308404Z digest=sha256:9fac2af2d251c5962bed34761bf3038efda735b552a1ae0be34c8992c7236b94

Pith citing papers

No inbound Pith citation observations are available.