Pith. sign in

Paper Citation Record · LEDGER

Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.06189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06189 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T06:02:34.065885Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.790278Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 950c5de4-7ac2-446d-abc0-7c45252c75f1 · inbound

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training cites this paper.

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T06:02:34.065885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:02:34.065885Z digest=sha256:bacd8c1c718ac6db485a6353eff7320225500ade96b9560148fd56bc0afb4f6e

Observation 7bbb85ce-bd01-4a79-94a0-db5e973566a4 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.831110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.831110Z digest=sha256:9188b8d0e81595129211f5b20b42147d3b5703fe7f39b90829a3053e374817f3

Observation 3c592c7c-ab04-44c7-8713-b9e59e2c950a · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.330222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:bcddc12d57db3ac6e6a4054c3189aeab71073525aa8993df11746d5874354cae

Observation cbd74da5-b010-4ae6-9fc1-70794fdbb2eb · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.335661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.335661Z digest=sha256:1e49b15b704d616b02e3a8401e6ad13b2f7abcb3e9fda98148aaac768c05f1da

Observation e637b7cf-0310-4c32-940f-16401aff3cc1 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.741475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:34d2ad24472193da01189f29c05c64bc334abf4760dafc77965d3a55da98fae2

Observation de93c1eb-8727-46ec-a8d0-58a2198a708d · inbound

See, Think, Learn: A Self-Taught Multimodal Reasoner cites this paper.

See, Think, Learn: A Self-Taught Multimodal Reasoner Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-08-03T19:04:29.148776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:04:29.148776Z digest=sha256:7b479128fc65f6958abd5e2ca3d3b1223d0202c9616f06482e23c70aff26c0ef

Observation 27b97f05-089c-498c-b445-df9a8fa132ef · inbound

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding cites this paper.

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.690880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T11:52:26.828633Z digest=sha256:5ec59f2f6d509f1164f37f6a88ef253fcd5c1c8865bb859953e88e1e25aaa03a

Observation 082d1ef7-ff9d-4be5-88c6-a5803e6dd4ec · inbound

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding cites this paper.

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:45.455092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T07:00:21.192082Z digest=sha256:03795c382736b1442acd0415c226ecad1eb7c7566bdd17ee3af2eb1d3850cccc

Observation 8f2e2cf0-ce10-4037-a833-ae03c78aef61 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:40.791756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T12:14:35.109298Z digest=sha256:d3dde2c21a4587238d1018e5d3287fb7235f940e5a795970dd789894f3aa12e2

Observation ffa5e41b-1a77-46a7-8fe8-57b3a5625bdb · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:39.637541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T06:29:51.635039Z digest=sha256:a312b44f9665b6c819e9442f3d0d8c194d213af01186ffe1b9791a6f79e60de7

Observation 0e6c1eb6-3d0a-4e8f-8554-fd73f8a1d567 · inbound

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models cites this paper.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:3541e2157f5727d69085b9494255065ad1c8eafa71bb40729da203a6190a4aaf