Pith. sign in

Paper Citation Record · LEDGER

SV3.3B: A Sports Video Understanding Model for Action Recognition

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.17844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17844 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:47:28.818411Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3859547-cc78-47d7-a1a8-e3fb88d0406d · outbound

This paper cites Review on wearable technology in sports: Concepts, challenges and opportunities,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Review on wearable technology in sports: Concepts, challenges and opportunities,

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T14:47:28.838947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.735332Z digest=sha256:79a24f6c407bb4b767a95984a28d05f6a38ea62f1cb85b75c46d7de38cea7d39

Observation 338bca84-3529-4456-bff5-dc53d909da87 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.737949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.737949Z digest=sha256:4649f1385a1cdc5757cef6d69b0a46c4c38ee36aa69875e83ba84101d94719db

Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.740370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.740370Z digest=sha256:356258d0f6e8e5cbd3e0ad08bfc9b7888772e6b2b85de14a6b2f76b9c7a50312

Observation 74989adb-e10d-4565-aa3a-e6a220763406 · outbound

This paper cites A path towards autonomous machine intelligence,.

SV3.3B: A Sports Video Understanding Model for Action Recognition A path towards autonomous machine intelligence,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.152382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.742825Z digest=sha256:3a694a08c2c29113073cec2c70d8f1dd546ef345309fd5717af8485323c2d3ae

Observation 41494e1c-70a4-4022-8f07-fb841d1a0c80 · outbound

This paper cites Self-supervised learning from images with a joint- embedding predictive architecture,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Self-supervised learning from images with a joint- embedding predictive architecture,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.146394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.745119Z digest=sha256:253cb445041f5d589df2a05b33741fe02f94c45c3c7d2d5c64596a71477b9871

Observation a7bd5611-59ef-4565-afb7-7d959e6fd077 · outbound

This paper cites V -JEPA: Latent video prediction for visual representation learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition V -JEPA: Latent video prediction for visual representation learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.140706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.747272Z digest=sha256:6ad3fdfb57778d9879ccba86e8e0ea8afe1173496f81e15781cd6f3a35b6ec7a

Observation 643eb1ac-91a6-4c1e-bbc1-b16805fde231 · outbound

This paper cites UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity.

SV3.3B: A Sports Video Understanding Model for Action Recognition UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.940562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.749711Z digest=sha256:d975766746c5b1d4c6d37b8e68add2d400088d6db22eeaff022af72bdc82edd5

Observation 029376f8-78cd-460f-9024-e915cb0e2d86 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

SV3.3B: A Sports Video Understanding Model for Action Recognition V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.752004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.752004Z digest=sha256:dfed6468eef24d5bf007cd754753f2d3fc977c9f60b0e8a8166bddfe30fa1e9d

Observation 354d65ae-8848-46b7-af8c-44abcdf96f4f · outbound

This paper cites Computer vision for sports: Current applications and research topics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Computer vision for sports: Current applications and research topics,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.134924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.754962Z digest=sha256:4e6b998a3435ff530938977f2ab780abeb37fcfb0579a761afc6d77b99ca86b3

Observation dc57ca70-1e9c-48d0-8886-435839cb8a2a · outbound

This paper cites Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.129217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.757216Z digest=sha256:17922ee6deb69af7005ceac6eff3f0a881844bcc87f52bfc9f9e914d9457dade

Observation 79293dc3-e8d5-4646-81d1-c6af34e48e68 · outbound

This paper cites Soccernet: A scalable dataset for action spotting in soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet: A scalable dataset for action spotting in soccer videos,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.123627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.759138Z digest=sha256:04c47797ce0bc1a1c83d8069331fa242527770c5b061ab51ff2cc400a14a8b9c

Observation fac1c96e-771e-4fe3-bb4b-d68df025fad1 · outbound

This paper cites Fine-grained action recognition on a novel basketball dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained action recognition on a novel basketball dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.118061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.761422Z digest=sha256:b8c0137c5ff9709be2b3daf4c744b5864e581f396610a6bd8139e1f33dafc9b8

Observation 8fd68349-1c4d-498b-8c8e-be5fefeb4b81 · outbound

This paper cites Soccernet caption: Dense video captioning for soccer broadcasts commentaries,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet caption: Dense video captioning for soccer broadcasts commentaries,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.112178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.763299Z digest=sha256:5751a79d93984705d1bc4be53b22acd7f2b6c08f3016692a12ad715f5e1df5db

Observation f03400f5-4020-4fe7-bf01-b888895bf08e · outbound

This paper cites Sports video captioning via attentive motion representation and group relationship modeling,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports video captioning via attentive motion representation and group relationship modeling,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.106353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.765170Z digest=sha256:f586ffcbad76b82c3ccfa465fee668060c632ff7eab7c5947066b6c006839821

Observation c99070ce-cc63-4f03-b9a5-68d95982b5fe · outbound

This paper cites Matchtime: Towards automatic soccer game commentary generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Matchtime: Towards automatic soccer game commentary generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.100713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.767048Z digest=sha256:65ec5b67ca5a7115972590d74a37fb23de09ba20e365172c609632b5609694a1

Observation 1709e277-9f4e-4a28-9110-8115693abf3d · outbound

This paper cites Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark.

SV3.3B: A Sports Video Understanding Model for Action Recognition Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.925700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.768978Z digest=sha256:335cb70d205f797b2b7937819de845d4212603073517da740220cc8c0c479cb4

Observation 7c6771fa-87fd-482a-9d0d-3a0ade95999c · outbound

This paper cites Fine-grained video captioning for sports narrative,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained video captioning for sports narrative,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.095279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.771064Z digest=sha256:1ae3a8ef3e21d389d92f95aff80e74b30947040ad506dc5951e2643cfc9532b1

Observation 21d0f4e1-c2ea-4a6d-8575-36f75976bcd0 · outbound

This paper cites Finegym: A hierarchical video dataset for fine -grained action understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finegym: A hierarchical video dataset for fine -grained action understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.089788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.772938Z digest=sha256:86539a00502e926c9d9b6f1df21403584d43a40e2b47afd4b5023f1d946c1591

Observation 35700ec6-dfc7-473b-8984-aab3d5f17bf9 · outbound

This paper cites Finediving: A fine - grained dataset for procedure-aware action quality assessment,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finediving: A fine - grained dataset for procedure-aware action quality assessment,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.083763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.775051Z digest=sha256:ccb6e57e4812e7ccb7e8f0b6fca93e761e369692a662a6c13606fa86c6b025c2

Observation 3d7bb799-2d82-4796-b9fd-f7259c4d6e1c · outbound

This paper cites Tacticai: An AI assistant for football tactics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Tacticai: An AI assistant for football tactics,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.077698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.776961Z digest=sha256:daefcbf5b3654fc74956130e9bb55df39ccdd41109c5c631c8726d5b210e7187

Observation 96745f7f-0097-49c8-af54-2e418a36fc36 · outbound

This paper cites VARS: Video assistant referee system for automated soccer decision making from multiple views,.

SV3.3B: A Sports Video Understanding Model for Action Recognition VARS: Video assistant referee system for automated soccer decision making from multiple views,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.071515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.778846Z digest=sha256:3c599ce78a515c46c05e7c24731d238302426ba824beb3b0b36f00dc6fdb7155

Observation 18204d28-eecd-429a-89a7-660da1fa7d85 · outbound

This paper cites X -VARS: Introducing explainability in football refereeing with multimodal large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition X -VARS: Introducing explainability in football refereeing with multimodal large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.065929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.780649Z digest=sha256:93a1f9b7ecbebf8888a81d45a04c2f18e009dfc94014c1d764f80bc90cfc9e3b

Observation 19ac599a-1c98-46cc-b946-2036d5ea9ff0 · outbound

This paper cites Sports-QA: A large -scale video question answering benchmark for complex and professional sports,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports-QA: A large -scale video question answering benchmark for complex and professional sports,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.782472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.782472Z digest=sha256:9c1003bb6d55dd72c18dd41923145466828e99a5ae2c5ac886078c1d05c07add

Observation 1e851dbc-fdf2-424a-a7fb-05e4c106ecd6 · outbound

This paper cites SportQA: A benchmark for sports understanding in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition SportQA: A benchmark for sports understanding in large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.060399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.784339Z digest=sha256:0340efc7716134331393664e29136cc75713e820072e23d5bed9156e4816a94b

Observation 5c80b69f-ef7f-4e47-a145-9fe7c26de692 · outbound

This paper cites SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.786186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.786186Z digest=sha256:bab2255f52d58a06c3b4d025e57699f452db3acd3ef28a9ca22556f75aeaf7a6

Observation 05827de8-aa86-48d4-b09d-4a6f57139e35 · outbound

This paper cites Flamingo: a visual language model for few -shot learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Flamingo: a visual language model for few -shot learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.054842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.788390Z digest=sha256:d29b40036e9c15eb121dc702f6dabeb88140e05e584563994f4b16fa436db77e

Observation 6326f665-380b-4582-b3d0-ee35e9d3a584 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.049117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.790361Z digest=sha256:b5072bd6386913f68e662b78b1f1b9052539546c0ff596e658f448b0885dab5b

Observation 27132ff3-fe73-447b-85dc-23a36b03cca7 · outbound

This paper cites BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.043418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.792353Z digest=sha256:4d2b87a8c747629cb608e893a334c4539dd594da3c9b1e7b3a965b2921ce87e7

Observation 1e314612-64b6-4c4f-bfae-a239a2f1a818 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Learning transferable visual models from natural language supervision,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.037574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.794410Z digest=sha256:ab8481c97ccca2bfa02231f40dd2415c7c2b31e8ee7b652ee913b6471d961674

Observation c3e888dd-8ada-490a-b844-0a4a727806c1 · outbound

This paper cites Sigmoid loss for language image pre -training,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sigmoid loss for language image pre -training,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.031708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.796296Z digest=sha256:439f2a163c9780f55825f1d62fb823ab88229754e6ce0dbcef9468ee18d5edac

Observation 8cd9abe6-d824-4158-9990-45f7127f4079 · outbound

This paper cites MVBench: A comprehensive multi-modal video understanding benchmark,.

SV3.3B: A Sports Video Understanding Model for Action Recognition MVBench: A comprehensive multi-modal video understanding benchmark,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.026208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.798156Z digest=sha256:4887fe20395d9eb8a2257b33ce5174d08921d80a7856d8a8fc0af26640eb3532

Observation 09140373-a5f7-4ad0-adf8-01dfc9870cf4 · outbound

This paper cites Llama -vid: An image is worth 2 tokens in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Llama -vid: An image is worth 2 tokens in large language models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.020360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.799820Z digest=sha256:1ce7f0101de4845c3dc5b14f81f60f8896906b71910b62b467e4ea3eaff82f42

Observation 8613031e-91c7-49b0-9caf-e3f67a6c1907 · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.014390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.801681Z digest=sha256:5fc591277c3a5cacddb09e8422602bb1260651988dc4ada74820f5e52c8e096a

Observation b5588f34-ed5d-4267-aedc-2f891d43403c · outbound

This paper cites Temporal alignment networks for long-term video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Temporal alignment networks for long-term video,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.008523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.803498Z digest=sha256:fe5f8dcf45346bf570f7d7998a39cf9dab267923a6f5edbbbaf3981154bd0249

Observation 94b8b072-21f6-433a-b120-af04d79c59e9 · outbound

This paper cites Multi-sentence grounding for long -term instructional video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Multi-sentence grounding for long -term instructional video,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.002751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.805361Z digest=sha256:be36ce2643148951c1613a3f5748d2af12032cde6a3eca79bb3623a92a3c2946

Observation b3295252-d1f3-4c72-8202-08f2a77f17de · outbound

This paper cites Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.996492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.807354Z digest=sha256:ea7d628480710abf081423da106af2cadf05585462ecf9e2ffe918d74e61ab57

Observation a1e043d5-b327-4f58-97d0-c7c0245eea61 · outbound

This paper cites Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.990368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.809252Z digest=sha256:152adda6319dc115da32b6384400f3fd5dcaf327b8509a4656fb13275cd377bf

Observation 542eae72-54e7-47eb-a817-15cb45c4713e · outbound

This paper cites Streaming dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Streaming dense video captioning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.984046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.811016Z digest=sha256:68408cd89267d3241623df6f897a476f3410e491d068d5fd5aba66c0665b10e1

Observation 8c1bb864-2069-46bd-ad3c-82db58b13626 · outbound

This paper cites Autoad: Movie description in context,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad: Movie description in context,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.978020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.812820Z digest=sha256:856cd57d50f4263490ca6a6ab7822b624e7d079156fe9720e5968b54382ddfc7

Observation 2287eebb-0f85-4024-bd33-826f84e35ab4 · outbound

This paper cites Autoad II: The sequel —who, when, and what in movie audio description,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad II: The sequel —who, when, and what in movie audio description,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.971953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.814682Z digest=sha256:4cd6ad847895b4965492cf84952f7f1429a7e28adbae19884ca367099faaac5d

Observation b3b75a8e-7dfc-4aaa-9db7-271e0fde46ba · outbound

This paper cites Autoad III: The prequel —back to the pixels,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad III: The prequel —back to the pixels,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.965659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.816566Z digest=sha256:7f77732a2106f69124fbacd260721f397d9ca873b2d3d78391b5d4d08937b9c5

Observation 225b382c-0702-4d58-aeff-48fbc85917e1 · outbound

This paper cites NSVA Subset: Basketball Video -Text Dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition NSVA Subset: Basketball Video -Text Dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.959004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:47:28.818411Z digest=sha256:943cc55574a9a05089c384946477fcd6c82fa43d16234ebc2e48b0a1f2ad0721

Pith citing papers

No inbound Pith citation observations are available.