Pith. sign in

Paper Citation Record · LEDGER

HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:1906.03327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1906.03327 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:13.520730Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:29.697481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ebee176c-4b2a-4cb3-9dba-6b5cdd17d903 · inbound

VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos cites this paper.

VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.520730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.520730Z digest=sha256:29fc1870f3a10a9afbc471949fa335d5a0cdcb2be6fee5329b10746a41828afb

Observation aac3bf2f-26c9-4bde-ba18-30b1539858c3 · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.100081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:fb68e67e893bf1638593c70d5f8f614ebe1b6fe079c5604d32e3e756c302c3a0

Observation b83efce2-60f6-4720-99e4-c3695c01f03a · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.921649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:1eb13aec88ccbb1becb9d875a56552104e09a2a73f708f3b854ebfd56ae8858e

Observation 22c02f2b-bdb6-488b-9561-1ca4a1e8c67f · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:39.975072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:b962edd018610bc9b746e86dd34bf11ec9400a0db25c2f5c91cbe977ef9f8aa1

Observation 17e3486a-d4a0-40ef-b324-1f48e5ea6594 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:334493c63789d30ae05eaec3fd593e8d1c64c4c6861ef2081740027418a63626

Observation 6047af69-0c9f-4351-98f7-3c54a5e26d1f · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:50.886671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:50.886671Z digest=sha256:e63cd70ed53d7f279197ff309cd3303a36bbfb36d43a128828532c9ef771b183

Observation 8487369e-e5c6-4ade-a681-4925339743b7 · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.654352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:3b2da4328d717a3f9a53030a88654a1d31e2d0278898738ae876814470b1e109

Observation d80fbd2a-e196-4a80-aacb-ec0f662604df · inbound

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining cites this paper.

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.699812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T17:53:24.431287Z digest=sha256:d5f8d71e64205e865514ede0f7009ef20631f36dd783ca11228afd56a1b057e1

Observation 4d429d72-093f-45c4-b88b-ad0ad2c03cbb · inbound

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living cites this paper.

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 182

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.244239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T18:22:13.147215Z digest=sha256:d000359b71d98078023dda42db4cb3c8149e54df9e77e452914026911f954f8e

Observation 1e0810f7-9469-4c59-894d-39418a7e82a0 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.772882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.772882Z digest=sha256:899ad85c4c234daca9e025e1ddd1a9910bd1991f8f94e28d7f5494a0f5b3ae0e