Pith. sign in

Paper Citation Record · LEDGER

SV3.3B: A Sports Video Understanding Model for Action Recognition

As of 10 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.17844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17844 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:47:28.818411Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3859547-cc78-47d7-a1a8-e3fb88d0406d · outbound

This paper cites Review on wearable technology in sports: Concepts, challenges and opportunities,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Review on wearable technology in sports: Concepts, challenges and opportunities,

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T14:47:28.838947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.735332Z digest=sha256:f9067303c24ad11368e48ba1271b27bda965004d0795e169735a3c3012ab062b

Observation 338bca84-3529-4456-bff5-dc53d909da87 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.737949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.737949Z digest=sha256:58cb4eabe4018904724fb53f35732cb4e5fdfeee8172ce730b2da0cf72241de3

Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.740370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.740370Z digest=sha256:0e2217f834fa633ed986e6121a8a4f21e40e62ec8013d1f63db8f5891da99519

Observation 74989adb-e10d-4565-aa3a-e6a220763406 · outbound

This paper cites A path towards autonomous machine intelligence,.

SV3.3B: A Sports Video Understanding Model for Action Recognition A path towards autonomous machine intelligence,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.152382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.742825Z digest=sha256:de78fb71e7c2262309c80f83d0a46890bf7c2aef0d746f04d46eadd8e99d5a98

Observation 41494e1c-70a4-4022-8f07-fb841d1a0c80 · outbound

This paper cites Self-supervised learning from images with a joint- embedding predictive architecture,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Self-supervised learning from images with a joint- embedding predictive architecture,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.146394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.745119Z digest=sha256:4942144f76926f4409b65b5a18954c42f30d7922e691bacf2fa19c09d71d66dd

Observation a7bd5611-59ef-4565-afb7-7d959e6fd077 · outbound

This paper cites V -JEPA: Latent video prediction for visual representation learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition V -JEPA: Latent video prediction for visual representation learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.140706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.747272Z digest=sha256:af5a9dce261bfdaf236038e88528591cf8a8cd35c3d55d041e59113084814196

Observation 643eb1ac-91a6-4c1e-bbc1-b16805fde231 · outbound

This paper cites UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity.

SV3.3B: A Sports Video Understanding Model for Action Recognition UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.940562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.749711Z digest=sha256:28a26a18dea83bacedcce41b9df0275d6129edb009a0e5cd394f496eae5a06e4

Observation 029376f8-78cd-460f-9024-e915cb0e2d86 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

SV3.3B: A Sports Video Understanding Model for Action Recognition V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.752004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.752004Z digest=sha256:2e38b8e278f90ec98ad0efa43cea76290821428bcad9b5ba5d6ff24588a0999f

Observation 354d65ae-8848-46b7-af8c-44abcdf96f4f · outbound

This paper cites Computer vision for sports: Current applications and research topics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Computer vision for sports: Current applications and research topics,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.134924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.754962Z digest=sha256:f37b6930f04ea38950b70f1a7864243bfaa486287e01bf8ca3e2195551298590

Observation dc57ca70-1e9c-48d0-8886-435839cb8a2a · outbound

This paper cites Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.129217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.757216Z digest=sha256:a088cbdcada12bbc1f2c8ba7e946648a516d729b5e7be8283218f609aaf05f9c

Observation 79293dc3-e8d5-4646-81d1-c6af34e48e68 · outbound

This paper cites Soccernet: A scalable dataset for action spotting in soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet: A scalable dataset for action spotting in soccer videos,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.123627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.759138Z digest=sha256:9447f05c7a8ea816c702000d12911298071d7ec1b8a97c19f6f17ec0cea0bd60

Observation fac1c96e-771e-4fe3-bb4b-d68df025fad1 · outbound

This paper cites Fine-grained action recognition on a novel basketball dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained action recognition on a novel basketball dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.118061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.761422Z digest=sha256:d5fb3a5a8af0e77eb87bc9d40dbcdddb331620ca7e61204110ccbdcd82b58a3c

Observation 8fd68349-1c4d-498b-8c8e-be5fefeb4b81 · outbound

This paper cites Soccernet caption: Dense video captioning for soccer broadcasts commentaries,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet caption: Dense video captioning for soccer broadcasts commentaries,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.112178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.763299Z digest=sha256:79c96b94ed659a9d00d4b23e021fdd1f5ec10099e46db155014ce1c36d41282a

Observation f03400f5-4020-4fe7-bf01-b888895bf08e · outbound

This paper cites Sports video captioning via attentive motion representation and group relationship modeling,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports video captioning via attentive motion representation and group relationship modeling,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.106353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.765170Z digest=sha256:8d6e64baa793cdde63db8823be5b9798cbb8c74f733508e1dfe7da67beda8048

Observation c99070ce-cc63-4f03-b9a5-68d95982b5fe · outbound

This paper cites Matchtime: Towards automatic soccer game commentary generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Matchtime: Towards automatic soccer game commentary generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.100713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.767048Z digest=sha256:f6de1cf15ebe5d8c16686bbac22c51c09f65b7404a93f8e679e9be774cc1036b

Observation 1709e277-9f4e-4a28-9110-8115693abf3d · outbound

This paper cites Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark.

SV3.3B: A Sports Video Understanding Model for Action Recognition Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.925700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.768978Z digest=sha256:1ab87b691b6eada5202f8ab4e30dc7402aed37b9ee8fe1ae00e44a299e528d91

Observation 7c6771fa-87fd-482a-9d0d-3a0ade95999c · outbound

This paper cites Fine-grained video captioning for sports narrative,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained video captioning for sports narrative,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.095279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.771064Z digest=sha256:0565e75c2da17dbaaf50c0ef046fba83707e512b1a362fef80488f85087d1cbb

Observation 21d0f4e1-c2ea-4a6d-8575-36f75976bcd0 · outbound

This paper cites Finegym: A hierarchical video dataset for fine -grained action understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finegym: A hierarchical video dataset for fine -grained action understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.089788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.772938Z digest=sha256:062a93b101ae0033b814cd29976b7d41767b213f11896fb379f27482f20336ca

Observation 35700ec6-dfc7-473b-8984-aab3d5f17bf9 · outbound

This paper cites Finediving: A fine - grained dataset for procedure-aware action quality assessment,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finediving: A fine - grained dataset for procedure-aware action quality assessment,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.083763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.775051Z digest=sha256:8df5fff143798868609523c6d9c9e728c71181ead54c0ac1cfdb98456c9e2c3c

Observation 3d7bb799-2d82-4796-b9fd-f7259c4d6e1c · outbound

This paper cites Tacticai: An AI assistant for football tactics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Tacticai: An AI assistant for football tactics,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.077698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.776961Z digest=sha256:a7ebc2c2e36349f7f1356d9cf1e75c0fb746cb650498c6f1cb8368f48eb069d0

Observation 96745f7f-0097-49c8-af54-2e418a36fc36 · outbound

This paper cites VARS: Video assistant referee system for automated soccer decision making from multiple views,.

SV3.3B: A Sports Video Understanding Model for Action Recognition VARS: Video assistant referee system for automated soccer decision making from multiple views,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.071515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.778846Z digest=sha256:46d4a075e7ddc87d6c03e59b744673edc620959dd92f730927ed3d34cbf264a4

Observation 18204d28-eecd-429a-89a7-660da1fa7d85 · outbound

This paper cites X -VARS: Introducing explainability in football refereeing with multimodal large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition X -VARS: Introducing explainability in football refereeing with multimodal large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.065929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.780649Z digest=sha256:d20636f91ffce95282f064c2e6bd7e15f024daa5b517436691a0811bca8decb6

Observation 19ac599a-1c98-46cc-b946-2036d5ea9ff0 · outbound

This paper cites Sports-QA: A large -scale video question answering benchmark for complex and professional sports,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports-QA: A large -scale video question answering benchmark for complex and professional sports,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.782472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.782472Z digest=sha256:234763a6a0db209429f1e3c4d60aa4cc94a55034b48f090e4de964c2c4e21761

Observation 1e851dbc-fdf2-424a-a7fb-05e4c106ecd6 · outbound

This paper cites SportQA: A benchmark for sports understanding in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition SportQA: A benchmark for sports understanding in large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.060399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.784339Z digest=sha256:cdfa296d9995a13b671581dda0c723de7a0f19b537fde144d5c4b1f436a91350

Observation 5c80b69f-ef7f-4e47-a145-9fe7c26de692 · outbound

This paper cites SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.786186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.786186Z digest=sha256:7277cb1c917dd7be15faae147014231ddae7c5972620847b8f08324aaf9a8195

Observation 05827de8-aa86-48d4-b09d-4a6f57139e35 · outbound

This paper cites Flamingo: a visual language model for few -shot learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Flamingo: a visual language model for few -shot learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.054842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.788390Z digest=sha256:12725a72f77ec90463ef81b688b6a37190cd97d542a9d030c29cb77baab91154

Observation 6326f665-380b-4582-b3d0-ee35e9d3a584 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.049117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.790361Z digest=sha256:98eec8e0ee52c4ccdbc877d80dd1c1e3608522ae44350bd3c1c671e544816507

Observation 27132ff3-fe73-447b-85dc-23a36b03cca7 · outbound

This paper cites BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.043418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.792353Z digest=sha256:35c99a7f91ab02aa43c2d3b7189278b46a7fc5a50d4815cea9fa7669a07767fc

Observation 1e314612-64b6-4c4f-bfae-a239a2f1a818 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Learning transferable visual models from natural language supervision,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.037574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.794410Z digest=sha256:2d134a2f3383d170cb9879d89ab4a6d24cbbcbb7f63ae19efd472c8586bb983d

Observation c3e888dd-8ada-490a-b844-0a4a727806c1 · outbound

This paper cites Sigmoid loss for language image pre -training,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sigmoid loss for language image pre -training,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.031708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.796296Z digest=sha256:c16db3b41b01c60ae67863c83a49d5b3b31131ad3019aca01e4990ee2970532a

Observation 8cd9abe6-d824-4158-9990-45f7127f4079 · outbound

This paper cites MVBench: A comprehensive multi-modal video understanding benchmark,.

SV3.3B: A Sports Video Understanding Model for Action Recognition MVBench: A comprehensive multi-modal video understanding benchmark,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.026208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.798156Z digest=sha256:b1a91c3310be288643ee7c23a6ec98d12353ded456c385fddc54b2ac5166d06c

Observation 09140373-a5f7-4ad0-adf8-01dfc9870cf4 · outbound

This paper cites Llama -vid: An image is worth 2 tokens in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Llama -vid: An image is worth 2 tokens in large language models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.020360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.799820Z digest=sha256:46e9b48b2b71d7dc2835622650cc4e33f25631879e6e93d7ab443a9427e4f7ef

Observation 8613031e-91c7-49b0-9caf-e3f67a6c1907 · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.014390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.801681Z digest=sha256:5ac4d7e1b84079a20ec552550f2855f24c698dae03bfa70d18133655ec19106f

Observation b5588f34-ed5d-4267-aedc-2f891d43403c · outbound

This paper cites Temporal alignment networks for long-term video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Temporal alignment networks for long-term video,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.008523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.803498Z digest=sha256:c9e421f181254796d76eaf6e3cd66be150f2edcc265949c716d51a5f7e517096

Observation 94b8b072-21f6-433a-b120-af04d79c59e9 · outbound

This paper cites Multi-sentence grounding for long -term instructional video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Multi-sentence grounding for long -term instructional video,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.002751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.805361Z digest=sha256:6f36e748320549c4caf3419828a43a3feda9766583a299964cb38d20f8fbeec2

Observation b3295252-d1f3-4c72-8202-08f2a77f17de · outbound

This paper cites Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.996492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.807354Z digest=sha256:4f5599917035a75f2c905476a1615eea53e610ba3aea24ec9df8f23888da06a7

Observation a1e043d5-b327-4f58-97d0-c7c0245eea61 · outbound

This paper cites Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.990368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.809252Z digest=sha256:3b016aba7cf6f521b4e8e70ecf042b897736a663958bbd18e0e9cc2afa901eda

Observation 542eae72-54e7-47eb-a817-15cb45c4713e · outbound

This paper cites Streaming dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Streaming dense video captioning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.984046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.811016Z digest=sha256:6df5927bacc997673d6f00779474a76c744d8fc6009b2364ab8e406f8225b566

Observation 8c1bb864-2069-46bd-ad3c-82db58b13626 · outbound

This paper cites Autoad: Movie description in context,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad: Movie description in context,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.978020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.812820Z digest=sha256:278238b4aac5738cc0d3515cd25e3457e83323fe82eec769c07cf6a64a8e04e7

Observation 2287eebb-0f85-4024-bd33-826f84e35ab4 · outbound

This paper cites Autoad II: The sequel —who, when, and what in movie audio description,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad II: The sequel —who, when, and what in movie audio description,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.971953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.814682Z digest=sha256:5f47200daf3a35fdd7c8ea1874835e509dc838c9c55ce0629dbe3ace32867d07

Observation b3b75a8e-7dfc-4aaa-9db7-271e0fde46ba · outbound

This paper cites Autoad III: The prequel —back to the pixels,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad III: The prequel —back to the pixels,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.965659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.816566Z digest=sha256:994aee8af93a4fabeb634b9c553c80e3750e3e9d8aec8da6068ad6fbd2af9fd2

Observation 225b382c-0702-4d58-aeff-48fbc85917e1 · outbound

This paper cites NSVA Subset: Basketball Video -Text Dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition NSVA Subset: Basketball Video -Text Dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.959004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:47:28.818411Z digest=sha256:15657361e967b4c8a0c52c632242f2ddc0140e4810719dc2e04ae142110d4cfa

Pith citing papers

No inbound Pith citation observations are available.