Pith. sign in

Paper Citation Record · LEDGER

Fostering Video Reasoning via Next-Event Prediction

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 5 inbound Pith citation observations for arXiv:2505.22457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22457 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:50.329904Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:30:02.227226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:08.052620Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8b243e2-10bb-45c5-9874-d119d25b607b · outbound

This paper cites GPT-4 Technical Report.

Fostering Video Reasoning via Next-Event Prediction GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:39.795809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:39.795809Z digest=sha256:3f3717547df7c321cd238522996e7679d498ba9bf9361fb51fb7437f1ac8ab54

Observation e1344f05-fac6-4083-b108-4d82a83f50ae · outbound

This paper cites Qwen Technical Report.

Fostering Video Reasoning via Next-Event Prediction Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:39.920205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:39.920205Z digest=sha256:4611dc19c3603ad29ce1a3baa66741d165046a43dfb3a08ebd4a9a16221a0e55

Observation f7b28bc2-524b-4bd9-828b-d89f18750afa · outbound

This paper cites Qwen2.5-VL Technical Report.

Fostering Video Reasoning via Next-Event Prediction Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.123547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.123547Z digest=sha256:7fbb379691cf924cf8a0a4f65cb9afdab87e91fd841ae85e2054173f0b5e3535

Observation b97c9119-9168-41fa-acfe-81ee98f6fb38 · outbound

This paper cites Perspective—discovery within validation logic: Deliberately surfacing, complementing, and substituting abductive reasoning in hypothetico- deductive inquiry.

Fostering Video Reasoning via Next-Event Prediction Perspective—discovery within validation logic: Deliberately surfacing, complementing, and substituting abductive reasoning in hypothetico- deductive inquiry

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.806676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:40.267143Z digest=sha256:838c174ebbfded5c03e39dcc947d8b163961b0680df5789fadec03fddf8ce8bc

Observation 8c97c53f-09c4-49cc-a2be-eebe6f273efc · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Fostering Video Reasoning via Next-Event Prediction TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.431117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.431117Z digest=sha256:0e7ff06564237f6aa751ddbf696078d6dda08a6887985b1a52f2de54d8069f8f

Observation 1ef549fe-3383-4fb6-8f03-86eaa348bf98 · outbound

This paper cites Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1.

Fostering Video Reasoning via Next-Event Prediction Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.647005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.647005Z digest=sha256:3d61d2db8e93ca728b27c9196d52b662dde8c85f47a219889af9580e5367aa1d

Observation 0791a0bd-f9d3-47f2-86ad-e93cedb888e3 · outbound

This paper cites Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs.

Fostering Video Reasoning via Next-Event Prediction Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.763426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.763426Z digest=sha256:cca8cd231cab91b1f5f55e405e01004a5705323ee026837a4a6c57b25127069f

Observation ec8e0ee3-de3c-4682-8d6c-becc925384ba · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

Fostering Video Reasoning via Next-Event Prediction Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:40.921513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:40.921513Z digest=sha256:1ba28635b2d155e6e2c3fd87038becfc1cab34eda5f3612c57eee405727ce695

Observation fbbad63a-714e-4059-9099-f57c7e3879de · outbound

This paper cites Abduction.

Fostering Video Reasoning via Next-Event Prediction Abduction

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.638631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:41.070272Z digest=sha256:425f4b10d24afff902a71e060e4a14577348c138eecd06c3ecd30f3f22ba7940

Observation 1b26eba1-f86e-44d0-a8c9-12b66e821277 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Fostering Video Reasoning via Next-Event Prediction Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:41.225273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:41.225273Z digest=sha256:5e0e485cb1338e547a65619b830aeff47e267e5c0158b3ec0b24ed57d67e83ac

Observation 2e760cb1-9cb8-4792-91bb-0b81ab3a281b · outbound

This paper cites Predicting the future: A jointly learnt model for action anticipation.

Fostering Video Reasoning via Next-Event Prediction Predicting the future: A jointly learnt model for action anticipation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.366839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:41.432388Z digest=sha256:d7b36f4d73393fb018760823e54b0d31c64f1e258ea4b5c8f0371021111a0602

Observation 9b41be81-a1e1-4170-869d-90f1cb2964ec · outbound

This paper cites Inductive and deductive reasoning.

Fostering Video Reasoning via Next-Event Prediction Inductive and deductive reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:58.139151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:41.566747Z digest=sha256:3ea4a888f2a0d3b733f2c42faecdb118cc96cbac6d1bd68113a7d20d038bbfc4

Observation da3b778a-9a18-401d-8f84-3217ffe933f1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Fostering Video Reasoning via Next-Event Prediction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:41.740834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:41.740834Z digest=sha256:985a611a61831da911f21e779b42bf14fc8803c776537b923dc7f70037b9d38d

Observation 96b992c6-8922-45e8-99df-b3e8f37d511e · outbound

This paper cites Large Language Models Are Reasoning Teachers.

Fostering Video Reasoning via Next-Event Prediction Large Language Models Are Reasoning Teachers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:41.915486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:41.915486Z digest=sha256:5dd0090692a96e0a904251c2d62d299e5ebf14379cafb9f3c94ce92925697714

Observation 3b1b8a55-698f-4d8f-be06-2c98a3185948 · outbound

This paper cites GPT-4o System Card.

Fostering Video Reasoning via Next-Event Prediction GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.048397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.048397Z digest=sha256:1667fdb92986772f6b7247bb3358278a4aac2fe90ab9e81647b3e00e3e8d2f57

Observation 14476e3b-fe1c-4b0a-b790-b5bb495e1be7 · outbound

This paper cites A hierarchical representation for future action prediction.

Fostering Video Reasoning via Next-Event Prediction A hierarchical representation for future action prediction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:57.895421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:42.221938Z digest=sha256:11423d311dc983afa7b26805c8343da4ed943bb741c9dcd4fb3b9c8be87d8f9a

Observation 59de5145-1f66-4698-94f4-6ffbe66328ce · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Fostering Video Reasoning via Next-Event Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.300383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.300383Z digest=sha256:466bf666530b35883a38857d5f0b89ab5f6c029e009c2694d98b7b6ff2260eae

Observation 481d3352-d8e1-4c05-97e5-2b773533b625 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Fostering Video Reasoning via Next-Event Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.429393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.429393Z digest=sha256:d18c93c6f351403e8279877c408e6e9630776c08961689378d5d7297cceec6e7

Observation a60049e2-c685-48d0-a8f6-1cad7c51724f · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Fostering Video Reasoning via Next-Event Prediction Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.566011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.566011Z digest=sha256:489b4b60f93028f5be02c3db2181d708250e6681cbdd739dad8cb80f11ec8ef9

Observation 450bb94b-80e9-4bc3-ae25-9225c7105aa6 · outbound

This paper cites A survey of multimodel large language models.

Fostering Video Reasoning via Next-Event Prediction A survey of multimodel large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.721721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.721721Z digest=sha256:4fc69ea539e42b2d638726fb0bd8f3513442cdecd3b193d2f967ca902e068c5a

Observation 84733f08-cbc8-4014-8a9b-a32583d0674b · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Fostering Video Reasoning via Next-Event Prediction Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.860422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.860422Z digest=sha256:6a968acb4e9b6101c928c55f5bd014f2dc5dd211eaeffb7d5962098bc8547f31

Observation 0160fd17-3872-42e9-8f2a-753a0538dea6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Fostering Video Reasoning via Next-Event Prediction Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:42.994162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:42.994162Z digest=sha256:2fc8244d42c4ceebfa4e430807512a450efdf5cf4453803fe0e1042f1be77766

Observation fc1f4c3c-6c0f-41b0-a710-4a3506c49205 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Fostering Video Reasoning via Next-Event Prediction TempCompass: Do Video LLMs Really Understand Videos?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.177211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.177211Z digest=sha256:94d3bcba1b9515e00ac496d1fef9d7428228a8214e64c1008c8b9c64ae8f6640

Observation afe341c0-7073-4c1d-882c-34ab70420a31 · outbound

This paper cites Training language models to follow instructions with human feedback.

Fostering Video Reasoning via Next-Event Prediction Training language models to follow instructions with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.311301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.311301Z digest=sha256:f32bb886355f9704642f59d17f238693f602a121808c80f414e4192e869d239e

Observation a9610f74-06b6-42ca-8980-6a344a6c5767 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Fostering Video Reasoning via Next-Event Prediction Learning transferable visual models from natural language supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.456195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.456195Z digest=sha256:060b5502777d624d9688b6a3403b0d15bc15938508e2c1e512e0d1bdabdca287

Observation 454627e6-d7b8-46a8-8f4f-5d3637e20f86 · outbound

This paper cites Video (language) modeling: a baseline for generative models of natural videos.

Fostering Video Reasoning via Next-Event Prediction Video (language) modeling: a baseline for generative models of natural videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.584292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.584292Z digest=sha256:a6e206e4fe47223752761665369fed3f70d5afc50ae65010d5301ae907575d26

Observation 633a548a-9fec-4e11-aa6c-8bd1ebd3fa99 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Fostering Video Reasoning via Next-Event Prediction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.689968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.689968Z digest=sha256:dc7545610adfb4c25294afda9997bcf592a1892c0936a47de5080f92e0c1a416

Observation 9bf018a1-6fb0-448d-b95e-efa4b250609d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Fostering Video Reasoning via Next-Event Prediction HybridFlow: A Flexible and Efficient RLHF Framework

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.851289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.851289Z digest=sha256:129bd8e11ec1e3027d1891cce0c8ddd347b1d8f932f7de7a702cde5b43d04856

Observation f42e9495-07d5-43ff-a659-14d3113954da · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Fostering Video Reasoning via Next-Event Prediction To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:43.988173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:43.988173Z digest=sha256:e6368a0b6ced0bcd0a6ec2d55176f1abab766070aa369908df57b408a59de6a6

Observation e931fd75-e0b7-4087-af4f-9c8819e9dbff · outbound

This paper cites The wisdom of crowds: Temporal progressive attention for early action prediction.

Fostering Video Reasoning via Next-Event Prediction The wisdom of crowds: Temporal progressive attention for early action prediction

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:57.587779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:44.194165Z digest=sha256:753798e9a255338849eb93ad72d0f5ee367ccd1df6fe43e7b2fd68c869e73370

Observation f8ba82e5-fe4e-461e-8d5e-52c2b8254de2 · outbound

This paper cites Video understanding with large language models: A survey.

Fostering Video Reasoning via Next-Event Prediction Video understanding with large language models: A survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.360273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.360273Z digest=sha256:20a488a5dff30826add8d31a12b03deff597ed215820560a4e1e11af89be1ea3

Observation 5465fe51-8e45-4d3b-86f2-106cb8700161 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Fostering Video Reasoning via Next-Event Prediction Gemini: A Family of Highly Capable Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.672213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.672213Z digest=sha256:817e85c74add012afbb73d954bfc559f4420faabf5cd6805d81f3cb65c1ad1de

Observation 02023953-c1a5-41ac-8ba0-e57c4663e488 · outbound

This paper cites Generating videos with scene dynamics.

Fostering Video Reasoning via Next-Event Prediction Generating videos with scene dynamics

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.830136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.830136Z digest=sha256:ae6f049fd86f97895b240b45dc2cda5e107ded915a1bfa5d3f51cae67fc30b6f

Observation 761182a1-f92a-4d9d-b1f1-15faf9eef1d0 · outbound

This paper cites Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate.

Fostering Video Reasoning via Next-Event Prediction Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:44.921265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:44.921265Z digest=sha256:7e06951ebc3cc12b878235fb9cec085af8163cb5d688b881df1afc991fc6ea99

Observation 2aa31ecd-1b36-4c79-a23a-2fd80ed5d43f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Fostering Video Reasoning via Next-Event Prediction Chain-of-thought prompting elicits reasoning in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.073595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.073595Z digest=sha256:a8e482dcd2766c09d0932381e342248eb6d674ea0cb1192ffa469616800f22d2

Observation d84633e9-bfe1-4859-8de3-ad1ba266b0c6 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.

Fostering Video Reasoning via Next-Event Prediction Longvideobench: A benchmark for long- context interleaved video-language understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:57.285209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:45.200203Z digest=sha256:2d9c629e3708304cc6905820d9d22763e23c52132003bb0645b2a3169c6bbeb6

Observation 7786aaa2-42f6-4983-8025-236172866b37 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Fostering Video Reasoning via Next-Event Prediction Next-qa: Next phase of question- answering to explaining temporal actions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.325130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.325130Z digest=sha256:8593c4883d99f7feb6058b65c47c64d4b6255dbf6f022ba7675c2c0d09867f4c

Observation c4bdd91e-916d-48b7-8ae6-ba21a585a8d9 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Fostering Video Reasoning via Next-Event Prediction Tree of thoughts: Deliberate problem solving with large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.470309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.470309Z digest=sha256:e3c96d18778026af945109a9f00e4ec2227c8ed597ea7419456b72a2abc057cd

Observation 8f62cfd8-7bdb-4bc1-b40d-e42e352344dc · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Fostering Video Reasoning via Next-Event Prediction ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.635718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.635718Z digest=sha256:5b8a5e83f3f61070d00b7e8a9169fe43422e38f360c00fd9e13c71d7ed50b1e8

Observation 554e8628-ed79-4325-8914-01eb210ef96d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Fostering Video Reasoning via Next-Event Prediction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.780544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.780544Z digest=sha256:97ad81ae26af4a977fde85060d19cb6aa801f97ea7feb4126055613861e65cc1

Observation 154b2e67-dfd9-49d3-8247-306637013fb2 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Fostering Video Reasoning via Next-Event Prediction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.965589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.965589Z digest=sha256:7225c14de85eb3f45977cd9ae98896439295e5369d3a60682cd70b1be0ef194e

Observation 5065b583-6453-4089-8ddf-b0e1f36b6f9d · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Fostering Video Reasoning via Next-Event Prediction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:46.144299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:46.144299Z digest=sha256:e58c7e38bd17fda62fba82d276f6de6ebd8ae60b0bb4e7364c2d822bbcff3573

Observation 3f835619-fbea-477d-a8f5-aa583d881a46 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Fostering Video Reasoning via Next-Event Prediction LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:46.365557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:46.365557Z digest=sha256:befc1b4c2070d8707b87fe994c6da2d1df2b32c8d4bba24734b43b2db3ad0a2f

Observation b9a95e2c-c00f-4580-9b5d-6551957c4539 · outbound

This paper cites Achild…”},{Scene2:“Legsare….

Fostering Video Reasoning via Next-Event Prediction Achild…”},{Scene2:“Legsare…

Reference 46

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:10:56.948195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:46.541780Z digest=sha256:26ecc6648bd0e0de34c65bcef1e15c656171f882c07b9753708951107281a2ce

Observation 6208c882-a1e4-4b0e-bd89-8c07fdcb3b5c · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:56.637066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:46.784745Z digest=sha256:f6cf9d82dcc2fb085ec04c31cff3151e1cef520ae8ac873394717ce1cdc2c39d

Observation d76ef48b-4255-4107-a1a6-693015df19f4 · outbound

This paper cites events.

Fostering Video Reasoning via Next-Event Prediction events

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:56.298105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:46.969026Z digest=sha256:d8ab175e145d380893b977dbc4bfee05e25c67a4131978d32732d7d7aaa09811

Observation db8aad78-2cee-4390-97f6-b7b146a25760 · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:56.014666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:47.160274Z digest=sha256:c424e95de67514fe52c26578d29b36aeeb52ea75caa37655591b6a5a9aa29d79

Observation 9ab6bacb-9e18-42c4-8962-c6b42e505ed6 · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:55.716992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:47.341401Z digest=sha256:a0154ef6704db50cb501848705c7af647e1f2f8593b20985161b68a9aec6c61d

Observation f597f086-9003-4464-ac15-afa20cba306b · outbound

This paper cites suitable.

Fostering Video Reasoning via Next-Event Prediction suitable

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:55.431653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:47.485560Z digest=sha256:467f7f200558af70d38541a371e2506515f5aa6c1f3c730267a72f80393a2328

Observation 92fcad3b-2201-482f-a665-5f0f263ba338 · outbound

This paper cites d e s c r i p t i o n.

Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:55.086036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:47.692629Z digest=sha256:03e7b2079474587c5a4ee1918bcb5382e3842fff3ba1a3de414ab95e50ad237c

Observation 65ad7b78-c82f-496e-b895-ac0d179512ba · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:54.826050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:47.902944Z digest=sha256:1b75c0aa90986c02736ffd7babece37fbcbb1781194c5e16651d9b9bbba73ad5

Observation 3dc27cc6-0011-43e3-9455-e1c67fec9fd4 · outbound

This paper cites d e s c r i p t i o n.

Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:54.512731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:48.042558Z digest=sha256:e9df67d4b9a2e5a2bf3303eddb0da5732210958d64cf8f8cff372366a8b4f95b

Observation 84b3ed2f-467c-4790-8f1d-79d50c88cff6 · outbound

This paper cites d e s c r i p t i o n.

Fostering Video Reasoning via Next-Event Prediction d e s c r i p t i o n

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:54.188702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:48.174609Z digest=sha256:45793e0c64b17673d234ee01df436a9982264044b5952e1a4298e81c5d40476f

Observation eccaab00-adda-4cac-ae6f-5c2cb4d935ad · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:53.697297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:48.281096Z digest=sha256:bbc227292a0f0fe32a5c8eb2922a4c594b5c5a30bf4bb0b0870e1b2fa774fe38

Observation 1f1cdcd4-1d27-43a9-b28a-efcbd22a09db · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:53.211950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:48.488636Z digest=sha256:b0bbecd03ce27774b0762deb4e4b80a71b64891b60a447cf3140687144fbf4a1

Observation b511f957-3141-4348-bff4-5b1eab709b5e · outbound

This paper cites an unresolved cited work.

Fostering Video Reasoning via Next-Event Prediction Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:52.783190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:48.685909Z digest=sha256:60941c06e695fd3f2f3910ecbe1ec8e1b9f659b80ffc4a98af70d12e0ae7b598

Observation 22328854-7e91-457c-ab7c-b0555fe4e041 · outbound

This paper cites C o n c l u s i o n : right.

Fostering Video Reasoning via Next-Event Prediction C o n c l u s i o n : right

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:52.398983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:48.873561Z digest=sha256:84263d1f497cefdb3c829c38501b00d01070efa36557217d1bd4ed6fe691ffc1

Observation 627e4afa-4a32-46f5-b949-8c9971d99f08 · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:52.182869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:49.053040Z digest=sha256:742fc7654443c30db5d2762ff717db4ed77af1279d655f1550a3b01dc44bf21b

Observation ff0e6b85-545b-419d-ab95-d3c6aa46106c · outbound

This paper cites - Ensure only one correct answer and that the r em ai nin g three options are wrong.

Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:52.002985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:49.269402Z digest=sha256:2cff628518c8cb9c0d216b2cc8b5bf254b9741271d74e281610902af490eb98c

Observation 13c7557d-d7c2-4d67-b2ab-c3048e22dcae · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.837463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:49.416429Z digest=sha256:87e22602b6addb63b75747424f2b675982ace6752bd850100917b4f2b358b8e7

Observation 0a22f794-bfb4-45da-bb6b-cda04bf24d4c · outbound

This paper cites - Ensure only one correct answer and that the r em ai nin g three options are wrong.

Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.690049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:49.612554Z digest=sha256:652381a3bea96ddd58d6b0d085ca59559be37f1217753f7b15652f2f51ae4392

Observation 3e9e1026-edaa-4cd2-b433-e990e86b4565 · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.549366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:49.785664Z digest=sha256:7fb3b159f39d896a8b1f5e8b29612869826bf58ebcf08c2b3424240ce9309df3

Observation 02fa0c0c-327e-4877-86c2-325c0418a7f0 · outbound

This paper cites - Ensure only one correct answer and that the r em ai nin g three options are wrong.

Fostering Video Reasoning via Next-Event Prediction - Ensure only one correct answer and that the r em ai nin g three options are wrong

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.367016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:49.909121Z digest=sha256:535a4c561741b86d0065a204493d70055e5492bccf4fcaabb373457d7b780bfa

Observation aaf58cc8-58b2-43d2-b779-18a2023b4d3d · outbound

This paper cites Question.

Fostering Video Reasoning via Next-Event Prediction Question

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:51.201914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:50.061912Z digest=sha256:0b51b8bbc3397adf9913e7d943c3b912695274955e5bd1963cf55b85c7cd65fb

Observation 26d5b6dc-02ae-4016-a9b7-aec1ff6e3ed3 · outbound

This paper cites - Answer options should be built upon the scenes after the observed scenes and before the last scene.

Fostering Video Reasoning via Next-Event Prediction - Answer options should be built upon the scenes after the observed scenes and before the last scene

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:50.917351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:50.329904Z digest=sha256:efbfc6f13620e351927c955202f07e9f45b5b034ebc23dc59d7e477ed830ceff

Pith citing papers

Observation 999cd9e9-7450-41b6-81d1-bffa35dbba45 · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Fostering Video Reasoning via Next-Event Prediction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.227226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.227226Z digest=sha256:d37d9a5b470fcabb7bfc438fc6b086e1e1543eb72f21d12206a2fc1dcf7ade1f

Observation 9a9dbaae-8460-4fa6-a39e-5f07afa41cf8 · inbound

EgoSelf: From Memory to Personalized Egocentric Assistant cites this paper.

EgoSelf: From Memory to Personalized Egocentric Assistant Fostering Video Reasoning via Next-Event Prediction

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:03.692100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:23:21.119521Z digest=sha256:d55dd1629476ce2d8e4484748f80aa578a56eab374eec01712be5c3a4d6627f9

Observation d5dbfd85-61ed-4384-aae0-8589d7c3ab8a · inbound

EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs cites this paper.

EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs Fostering Video Reasoning via Next-Event Prediction

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:15.231278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:34:14.297331Z digest=sha256:b9d96abf27cd9c11849be8ca5aa6ecba418600229a012962bc5027bd43e42a95

Observation b5ace437-a666-4879-885c-3a35be18a4dc · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Fostering Video Reasoning via Next-Event Prediction

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.449447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:c16190f695c847a2c74435a7fd3721a52b754927a0ccb00da8878be94cb088cf

Observation 416437f4-13df-4c83-977f-687485dc6d1b · inbound

FeVOS: Foresight Expression Video Object Segmentation cites this paper.

FeVOS: Foresight Expression Video Object Segmentation Fostering Video Reasoning via Next-Event Prediction

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:08.053934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T21:13:21.500674Z digest=sha256:5cbabf5cb15e8ecb80b4df79643e131786de4634d20c292f7b22d2127b3200a5