Pith. sign in

Paper Citation Record · LEDGER

Fostering Video Reasoning via Next-Event Prediction

As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 5 inbound Pith citation observations for arXiv:2505.22457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22457 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:50.329904Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:30:02.227226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:08.052620Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8b243e2-10bb-45c5-9874-d119d25b607b · outbound

This paper cites GPT-4 Technical Report.

Fostering Video Reasoning via Next-Event Prediction GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:39.795809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:39.795809Z digest=sha256:5395688760fe8e14ee30f7effd7bbdc6e5c081b47219a8b8eac5bf5cf15a862f

Observation e1344f05-fac6-4083-b108-4d82a83f50ae · outbound

This paper cites Qwen Technical Report.

Fostering Video Reasoning via Next-Event Prediction Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:39.920205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:39.920205Z digest=sha256:76e03a0ce64f9dab14fbd1b7b9fda050912c5d8e4e2cd37754b9298973c0f686

Observation f7b28bc2-524b-4bd9-828b-d89f18750afa · outbound

This paper cites Qwen2.5-VL Technical Report.

Fostering Video Reasoning via Next-Event Prediction Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.123547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.123547Z digest=sha256:ff801e614428a2623fa62fe0bd687343108e5328045d92489a3badff1f3bb576

Observation b97c9119-9168-41fa-acfe-81ee98f6fb38 · outbound

This paper cites Perspective—discovery within validation logic: Deliberately surfacing, complementing, and substituting abductive reasoning in hypothetico- deductive inquiry.

Fostering Video Reasoning via Next-Event Prediction Perspective—discovery within validation logic: Deliberately surfacing, complementing, and substituting abductive reasoning in hypothetico- deductive inquiry

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.806676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:40.267143Z digest=sha256:fc334c98cc0984da98a7896a2b3477ca39c054307a490239024bbb0ace062804

Observation 8c97c53f-09c4-49cc-a2be-eebe6f273efc · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Fostering Video Reasoning via Next-Event Prediction TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.431117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.431117Z digest=sha256:933f3880c1a24b1d7d1d090421d2931917479c6befe479e68aa1e46107bfd94f

Observation 1ef549fe-3383-4fb6-8f03-86eaa348bf98 · outbound

This paper cites Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1.

Fostering Video Reasoning via Next-Event Prediction Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.647005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.647005Z digest=sha256:3629a56f191a523c8096c5b2bf3754da0fb65ec5e54978bc88c478526b612843

Observation 0791a0bd-f9d3-47f2-86ad-e93cedb888e3 · outbound

This paper cites Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs.

Fostering Video Reasoning via Next-Event Prediction Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.763426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.763426Z digest=sha256:a14e9d4e500a2b0babd27cae04ba26f684e189a27e5c85d0d4b6334ba1091fb4

Observation ec8e0ee3-de3c-4682-8d6c-becc925384ba · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

Fostering Video Reasoning via Next-Event Prediction Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.921513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.921513Z digest=sha256:3b3ecafebedaccf5c1be3c97c76c1649955e1666d8305bf5e3de1fa0df8b4350

Observation fbbad63a-714e-4059-9099-f57c7e3879de · outbound

This paper cites Abduction.

Fostering Video Reasoning via Next-Event Prediction Abduction

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.638631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:41.070272Z digest=sha256:295276f40c1a3e2b6e3d6640992370cc69accba61c003830d674b511a7a8817b

Observation 1b26eba1-f86e-44d0-a8c9-12b66e821277 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Fostering Video Reasoning via Next-Event Prediction Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:41.225273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:41.225273Z digest=sha256:9af5791ace14216414cf0cae346f61aa914f91dccbfaf356b7baa569e32e6757

Observation 2e760cb1-9cb8-4792-91bb-0b81ab3a281b · outbound

This paper cites Predicting the future: A jointly learnt model for action anticipation.

Fostering Video Reasoning via Next-Event Prediction Predicting the future: A jointly learnt model for action anticipation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.366839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:41.432388Z digest=sha256:8e11e086d72df6c3d6ec2ef38101a1fdf8580fb79cd95596777e463fde2c67d8

Observation 9b41be81-a1e1-4170-869d-90f1cb2964ec · outbound

This paper cites Inductive and deductive reasoning.

Fostering Video Reasoning via Next-Event Prediction Inductive and deductive reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.139151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:41.566747Z digest=sha256:dd1ecbcf0f87e8caba48214141da0a77c9a42dd0f97aefedbc57acdeafcf9b72

Observation da3b778a-9a18-401d-8f84-3217ffe933f1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Fostering Video Reasoning via Next-Event Prediction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:41.740834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:41.740834Z digest=sha256:583bbf20a60a0a27187594fab07735cefc021f17da5de7be2cc673fe11fae306

Observation 96b992c6-8922-45e8-99df-b3e8f37d511e · outbound

This paper cites Large Language Models Are Reasoning Teachers.

Fostering Video Reasoning via Next-Event Prediction Large Language Models Are Reasoning Teachers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:41.915486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:41.915486Z digest=sha256:aa0447e1916e0e8b1baf30ff910c9cb2254552e8d39cfe2681675f3e9b2e7d9b

Observation 3b1b8a55-698f-4d8f-be06-2c98a3185948 · outbound

This paper cites GPT-4o System Card.

Fostering Video Reasoning via Next-Event Prediction GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.048397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.048397Z digest=sha256:eb619cd796c641492fbf01cf478b28057ac3a05f55f726b6ebd722cac51b2770

Observation 14476e3b-fe1c-4b0a-b790-b5bb495e1be7 · outbound

This paper cites A hierarchical representation for future action prediction.

Fostering Video Reasoning via Next-Event Prediction A hierarchical representation for future action prediction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:57.895421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:42.221938Z digest=sha256:77acfc16d43554d9509ef322a7f95dee08e8bc5794e4bc8d71d287cac179d3e0

Observation 59de5145-1f66-4698-94f4-6ffbe66328ce · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Fostering Video Reasoning via Next-Event Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.300383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.300383Z digest=sha256:65d0559e8ee6722f6c21ed192df7076550740258b6a7e742ab08d7b067f87583

Observation 481d3352-d8e1-4c05-97e5-2b773533b625 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Fostering Video Reasoning via Next-Event Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.429393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.429393Z digest=sha256:bced780421283c06290d593331d0752a050dd42a3955fe68814258f3cf34255c

Observation a60049e2-c685-48d0-a8f6-1cad7c51724f · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Fostering Video Reasoning via Next-Event Prediction Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.566011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.566011Z digest=sha256:249567710c761e524efbff3da00d695c771259cdb2dc261a43ac557e6a8f9e93

Observation 450bb94b-80e9-4bc3-ae25-9225c7105aa6 · outbound

This paper cites A survey of multimodel large language models.

Fostering Video Reasoning via Next-Event Prediction A survey of multimodel large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.721721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.721721Z digest=sha256:aa78f848e198445e30cb1ab87cd40ef6e6e8cb5ce5835d964f7d2d672666651e

Observation 84733f08-cbc8-4014-8a9b-a32583d0674b · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Fostering Video Reasoning via Next-Event Prediction Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.860422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.860422Z digest=sha256:7d8bcf104504cb43049d950fa891b36116f6132827b9e00a0f6070894dd2421d

Observation 0160fd17-3872-42e9-8f2a-753a0538dea6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Fostering Video Reasoning via Next-Event Prediction Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.994162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.994162Z digest=sha256:64c3b5ebb1b0e276071f5ada34d9ed865fd781e64f57394181baedbf509631d0

Observation fc1f4c3c-6c0f-41b0-a710-4a3506c49205 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Fostering Video Reasoning via Next-Event Prediction TempCompass: Do Video LLMs Really Understand Videos?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.177211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.177211Z digest=sha256:61f656eb6851e62dfed48217c4c17d90938585cd0dd6cdc6d8525b2fa096f5cf

Observation afe341c0-7073-4c1d-882c-34ab70420a31 · outbound

This paper cites Training language models to follow instructions with human feedback.

Fostering Video Reasoning via Next-Event Prediction Training language models to follow instructions with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.311301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.311301Z digest=sha256:a777855f39e6a06706f201b53d11fbb023853b688dcc55657e3e36dc8e878a04

Observation a9610f74-06b6-42ca-8980-6a344a6c5767 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Fostering Video Reasoning via Next-Event Prediction Learning transferable visual models from natural language supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.456195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.456195Z digest=sha256:62f1a35a3ef43401a9871b46fa7f448f0c235e0a2f6a71696977dea010b38cbb

Observation 454627e6-d7b8-46a8-8f4f-5d3637e20f86 · outbound

This paper cites Video (language) modeling: a baseline for generative models of natural videos.

Fostering Video Reasoning via Next-Event Prediction Video (language) modeling: a baseline for generative models of natural videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.584292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.584292Z digest=sha256:00e4ff97f35c15779cde1c9f4636697f7bb896b1e515496fd92f8b46e82d6a46

Observation 633a548a-9fec-4e11-aa6c-8bd1ebd3fa99 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Fostering Video Reasoning via Next-Event Prediction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.689968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.689968Z digest=sha256:4e1e780c79d3930ca7048d28300627b1e2610bae245587a8fa79038b93b17221

Observation 9bf018a1-6fb0-448d-b95e-efa4b250609d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Fostering Video Reasoning via Next-Event Prediction HybridFlow: A Flexible and Efficient RLHF Framework

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.851289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.851289Z digest=sha256:1910deff4783f706449945789b9b600f240ccbbf1f1bd118546f8e543f75689e

Observation f42e9495-07d5-43ff-a659-14d3113954da · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Fostering Video Reasoning via Next-Event Prediction To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.988173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.988173Z digest=sha256:52611ff16d563746ee7e52e7c4a9fc55c7aed8da5ef9d3f5acdad085698840cb

Observation e931fd75-e0b7-4087-af4f-9c8819e9dbff · outbound

This paper cites The wisdom of crowds: Temporal progressive attention for early action prediction.

Fostering Video Reasoning via Next-Event Prediction The wisdom of crowds: Temporal progressive attention for early action prediction

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:57.587779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:44.194165Z digest=sha256:ccd6dd4d4909a7b2223d264c65275f6d9dad37edbd9acc7cdbab2bff20a69679

Observation f8ba82e5-fe4e-461e-8d5e-52c2b8254de2 · outbound

This paper cites Video understanding with large language models: A survey.

Fostering Video Reasoning via Next-Event Prediction Video understanding with large language models: A survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.360273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.360273Z digest=sha256:3076e599fa27b376e96f614a5eb6ffc894d8bc25c175014c9b2fb48353cff318

Observation 5465fe51-8e45-4d3b-86f2-106cb8700161 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Fostering Video Reasoning via Next-Event Prediction Gemini: A Family of Highly Capable Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.672213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.672213Z digest=sha256:0468457f812e737284aa0bd1d4e701987763e91cd5c1ede8b1f9b59452449d97

Observation 02023953-c1a5-41ac-8ba0-e57c4663e488 · outbound

This paper cites Generating videos with scene dynamics.

Fostering Video Reasoning via Next-Event Prediction Generating videos with scene dynamics

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.830136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.830136Z digest=sha256:0fdf888e76ed27e2c147d9aa6cb5b61ea63a6c5f5fa05b886302554a04047096

Observation 761182a1-f92a-4d9d-b1f1-15faf9eef1d0 · outbound

This paper cites Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate.

Fostering Video Reasoning via Next-Event Prediction Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.921265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.921265Z digest=sha256:c5d85bad7b16e95379b3baaf1a89b9580e2ce602f061f2a64d1ea67476c608aa

Observation 2aa31ecd-1b36-4c79-a23a-2fd80ed5d43f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Fostering Video Reasoning via Next-Event Prediction Chain-of-thought prompting elicits reasoning in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.073595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.073595Z digest=sha256:3fdb1c44e18ab3c9b2d1f0be1445dd578bda4b06821d99fee3ff6919a26ed73c

Observation d84633e9-bfe1-4859-8de3-ad1ba266b0c6 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.

Fostering Video Reasoning via Next-Event Prediction Longvideobench: A benchmark for long- context interleaved video-language understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:57.285209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:45.200203Z digest=sha256:d49df1ddf831fe12095c8af4db63ed94c1c731803060ec6bf83645186630f82b

Observation 7786aaa2-42f6-4983-8025-236172866b37 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Fostering Video Reasoning via Next-Event Prediction Next-qa: Next phase of question- answering to explaining temporal actions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.325130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.325130Z digest=sha256:2d95520e0f45a11e15652c9eda1d3475a12a9df9e4f03d1f1a31f1923257a3b0

Observation c4bdd91e-916d-48b7-8ae6-ba21a585a8d9 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Fostering Video Reasoning via Next-Event Prediction Tree of thoughts: Deliberate problem solving with large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.470309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.470309Z digest=sha256:12eb05c902e182f7bb42da1ba16994a30788f8274be82117cb1c49861ef340f8

Observation 8f62cfd8-7bdb-4bc1-b40d-e42e352344dc · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Fostering Video Reasoning via Next-Event Prediction ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.635718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.635718Z digest=sha256:642608b8bff1dd5ca11d18da5d213cb5300b599529cecf4423339f8851b50ffa

Observation 554e8628-ed79-4325-8914-01eb210ef96d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Fostering Video Reasoning via Next-Event Prediction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.780544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.780544Z digest=sha256:79f45d1276ad6353d68c938b84b84f3aabc3db4909ddbed13971644874ba31a9

Observation 154b2e67-dfd9-49d3-8247-306637013fb2 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Fostering Video Reasoning via Next-Event Prediction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.965589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.965589Z digest=sha256:9e52d327e4e5e4222db1a45ab7d331cfea047e968dfed11afa24a8cbcdbd234a

Observation 5065b583-6453-4089-8ddf-b0e1f36b6f9d · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Fostering Video Reasoning via Next-Event Prediction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:46.144299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:46.144299Z digest=sha256:aead5256df60042a89d93fe383dc99722c07b60b0265671a7db574cc371d2be3

Observation 3f835619-fbea-477d-a8f5-aa583d881a46 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Fostering Video Reasoning via Next-Event Prediction LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:46.365557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:46.365557Z digest=sha256:619dc1e643b0d53dac294766736d5eb063d12db4cacb30ef0ee61c9eade4c415

Observation b9a95e2c-c00f-4580-9b5d-6551957c4539 · outbound

This paper cites Achild…”},{Scene2:“Legsare….

Fostering Video Reasoning via Next-Event Prediction Achild…”},{Scene2:“Legsare…

Reference 46

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:10:56.948195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:46.541780Z digest=sha256:149fd3b5c258a3a6793fdf7a2d6ffc4db14f82a6211e6704ebac6bf201f3f68f

Observation 6208c882-a1e4-4b0e-bd89-8c07fdcb3b5c · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:56.637066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:46.784745Z digest=sha256:05f660cd0d7e6878b31f7e3bc7dcb6698701105de3206bbbcef476697c7221b3

Observation d76ef48b-4255-4107-a1a6-693015df19f4 · outbound

This paper cites events.

Fostering Video Reasoning via Next-Event Prediction events

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:56.298105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:46.969026Z digest=sha256:229ab684e9c70e3b920a57f04c92dc98193f37ccc90070288f75b4bd556e4cc5

Observation db8aad78-2cee-4390-97f6-b7b146a25760 · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:56.014666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:47.160274Z digest=sha256:a716a111a7c27a72967b640e876d956820e4b24df4ef8a60ea0ea957ad604ccb

Observation 9ab6bacb-9e18-42c4-8962-c6b42e505ed6 · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:55.716992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:47.341401Z digest=sha256:9f3a252159df98d78288543186e38f18b4c814b034ebbdbd844b981faea99b69

Observation f597f086-9003-4464-ac15-afa20cba306b · outbound

This paper cites suitable.

Fostering Video Reasoning via Next-Event Prediction suitable

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:55.431653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:47.485560Z digest=sha256:24bb5e2050318c398f26330acf69cd53037624b4eb99d8c983c2ca11d5b2d47c

Observation 92fcad3b-2201-482f-a665-5f0f263ba338 · outbound

This paper cites d e s c r i p t i o n.

Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:55.086036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:47.692629Z digest=sha256:fc875341b185bd55d1fd99176e1b206db9db1bd0d7f8011f339f941388dab8e6

Observation 65ad7b78-c82f-496e-b895-ac0d179512ba · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:54.826050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:47.902944Z digest=sha256:b7c4a33da9140cdd6462f83b587a5116136c1f6a2000ae7ad24971a945e9075f

Observation 3dc27cc6-0011-43e3-9455-e1c67fec9fd4 · outbound

This paper cites d e s c r i p t i o n.

Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:54.512731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:48.042558Z digest=sha256:8873110d90de378cad4a27928d821d6b12946de6bc6de3fa4fda5e3d5ba227fd

Observation 84b3ed2f-467c-4790-8f1d-79d50c88cff6 · outbound

This paper cites d e s c r i p t i o n.

Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:54.188702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:48.174609Z digest=sha256:8047a8809552f3bb67e71634ca52aaba4764c63b31a6f72342c2ad7d6e6ab65e

Observation eccaab00-adda-4cac-ae6f-5c2cb4d935ad · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:53.697297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:48.281096Z digest=sha256:1216c786144e04d974e6d29e21dc94538c2447ec7cd5c9b1526e44efac65795b

Observation 1f1cdcd4-1d27-43a9-b28a-efcbd22a09db · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:53.211950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:48.488636Z digest=sha256:2d48e30f0f76a959d3937be37195d7f0e5d2c3b4f5a793184bdcd14c01fc7f81

Observation b511f957-3141-4348-bff4-5b1eab709b5e · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:52.783190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:48.685909Z digest=sha256:680a5529e8f8dfb297817429e4bccb5e742ea9c0075f1f3c43cd1b0be440445f

Observation 22328854-7e91-457c-ab7c-b0555fe4e041 · outbound

This paper cites C o n c l u s i o n : right.

Fostering Video Reasoning via Next-Event Prediction C o n c l u s i o n : right

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:52.398983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:48.873561Z digest=sha256:07c17057962d1ab110166b55557a4e424740ec663cbb457d42bf8510d12eba3a

Observation 627e4afa-4a32-46f5-b949-8c9971d99f08 · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:52.182869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:49.053040Z digest=sha256:8a44b7d0cdc259d9b5e30859278ca36d34e6ee0e64f0ebccee620e351f782d94

Observation ff0e6b85-545b-419d-ab95-d3c6aa46106c · outbound

This paper cites - Ensure only one correct answer and that the r em ai nin g three options are wrong.

Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:52.002985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:49.269402Z digest=sha256:b63f5f4a94a8a90c2bd35545976c24a566efc798ed2e5ce4f411fbee517afb7f

Observation 13c7557d-d7c2-4d67-b2ab-c3048e22dcae · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.837463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:49.416429Z digest=sha256:5dd145cb46095767d4183680b0f8cdd44dde8a7b9c1ae5dc1887b8e835d89ed5

Observation 0a22f794-bfb4-45da-bb6b-cda04bf24d4c · outbound

This paper cites - Ensure only one correct answer and that the r em ai nin g three options are wrong.

Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.690049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:49.612554Z digest=sha256:f53c5f671b4636266a6ea3dc4453be5a43c7263fc1e466548dd6e75e92aeeff2

Observation 3e9e1026-edaa-4cd2-b433-e990e86b4565 · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.549366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:49.785664Z digest=sha256:2285d48e7a418e1659bef677fc5e560ae6a38a2811544dd1ced9048fc44534bb

Observation 02fa0c0c-327e-4877-86c2-325c0418a7f0 · outbound

This paper cites - Ensure only one correct answer and that the r em ai nin g three options are wrong.

Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.367016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:49.909121Z digest=sha256:513a5587904f81541ceaa5759a31c519b0472def2947f34de844f4e762f647d8

Observation aaf58cc8-58b2-43d2-b779-18a2023b4d3d · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.201914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:50.061912Z digest=sha256:8c625072100f2eb1b79cfd7b3570947b7358a831b03a5ef69ee8374029299c46

Observation 26d5b6dc-02ae-4016-a9b7-aec1ff6e3ed3 · outbound

This paper cites - Answer options should be built upon the scenes after the observed scenes and before the last scene.

Fostering Video Reasoning via Next-Event Prediction - Answer options should be built upon the scenes after the observed scenes and before the last scene

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:50.917351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:10:50.329904Z digest=sha256:d5af67e7ceffa102e9b6e30ec58a8faaa9c55f703f3560dfdd73b9b7a77802f5

Pith citing papers

Observation 999cd9e9-7450-41b6-81d1-bffa35dbba45 · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Fostering Video Reasoning via Next-Event Prediction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.227226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.227226Z digest=sha256:e1cbd5a9dbb6142c748c776c0cdaf13832dc8f8d44bc475048b7dacd05f2e218

Observation 9a9dbaae-8460-4fa6-a39e-5f07afa41cf8 · inbound

EgoSelf: From Memory to Personalized Egocentric Assistant cites this paper.

EgoSelf: From Memory to Personalized Egocentric Assistant Fostering Video Reasoning via Next-Event Prediction

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:03.692100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:23:21.119521Z digest=sha256:3d00f3f1f765d1612f59a0377acab747c3380e5cac896b22b1b4f3192a0e6366

Observation d5dbfd85-61ed-4384-aae0-8589d7c3ab8a · inbound

EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs cites this paper.

EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs Fostering Video Reasoning via Next-Event Prediction

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:15.231278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T08:34:14.297331Z digest=sha256:409cd5d6a79492354cac516668412cfbb1bd17d5685dc4afd28ff48e50cef1ca

Observation b5ace437-a666-4879-885c-3a35be18a4dc · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Fostering Video Reasoning via Next-Event Prediction

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.449447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:acabf43374afbb6939455156e8bb9008e5ddc1e22e6e364b15fdf025dac21e18

Observation 416437f4-13df-4c83-977f-687485dc6d1b · inbound

FeVOS: Foresight Expression Video Object Segmentation cites this paper.

FeVOS: Foresight Expression Video Object Segmentation Fostering Video Reasoning via Next-Event Prediction

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:08.053934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T21:13:21.500674Z digest=sha256:cacaff8809c51d7e8900cb6e4e9817c153409c9f8e3256fb47c177000c729a13