Pith. sign in

Paper Citation Record · LEDGER

Thinking in Video: Can Video Generators Really Reason About the Real World?

As of 19 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2607.17523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17523 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:13.750039Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3b7b776-f939-4895-a031-237815ada913 · outbound

This paper cites Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:04.983265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:04.983265Z digest=sha256:14800d831ae932307f844a238b086c9f3b473d6c4e78892e0ead81b59a60b8a2

Observation b268883f-6059-47e9-814b-78ba70662e6f · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Thinking in Video: Can Video Generators Really Reason About the Real World? Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.062133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.062133Z digest=sha256:da075aafc0b64eb4b312f814a9523e3bd5ba4a01b634c153bc5185a4467642d9

Observation 3362a094-430d-475b-b4a9-73d0ce640b1d · outbound

This paper cites Sora 2.https://openai.com/index/sora-2, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Sora 2.https://openai.com/index/sora-2, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.183713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.183713Z digest=sha256:e44c8920f1b058ac9d9f6f126b70651af623edf69572bed790e9e300942bbd11

Observation 92675ba5-d455-41dd-a856-07ad00e3b75d · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? HunyuanVideo 1.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.356951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.356951Z digest=sha256:bdbd2d52c3cf55cf68ee848dfe3a427b3c24c3b2c44d685b710f3c74c9714a03

Observation 6c2d970b-d489-4ddd-ba65-e5a47f9ce6df · outbound

This paper cites Veo 3.https://aistudio.google.com/models/veo-3, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Veo 3.https://aistudio.google.com/models/veo-3, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.530332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.530332Z digest=sha256:b09895d8482765ccb361adff8c942bf3149d9d477e34df94bd77e794281a3dcc

Observation ea11f366-8280-42aa-bc23-0f5cda6c1acc · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Wan: Open and Advanced Large-Scale Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.639163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.639163Z digest=sha256:ee517413218a7eae5c9c85aaf8eeb67aed70232f38c5e01827298d374de8f567

Observation 41acdf30-f207-439f-a381-4e6b87a1b8b2 · outbound

This paper cites From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.751020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.751020Z digest=sha256:40c6e893bf3f912c4dabf6b5f2c46412ab1efc982454156783832bec0fc97c49

Observation 6f71d0e0-f3e8-46bc-a99b-beb811933f59 · outbound

This paper cites World Simulation with Video Foundation Models for Physical AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.927225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.927225Z digest=sha256:5656cf0ad3911e9839b5ae4c6f93aeaeac19c5e49757eab36920997bd8a09660

Observation 183adc3c-406c-406f-9150-098284b56556 · outbound

This paper cites Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.134736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.134736Z digest=sha256:96f6fc61d0e3fd7ec46ce78837f0bf42e37ab2d5557ba6c3e2df5c64cfdd8a37

Observation 5920ead6-c2f8-48fb-9f7d-3462801e86f6 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.272600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.272600Z digest=sha256:ba0e3e6ff21801c7718e28f128c04ae66304f8476f1ce7311086e20b6b229ec4

Observation 28b0cf78-ac9b-4bb2-842d-17a99a13261c · outbound

This paper cites The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking.

Thinking in Video: Can Video Generators Really Reason About the Real World? The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.472741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.472741Z digest=sha256:29ef66453c4f354cbff51d47f1971d9d2d5f456ac12fb4fb79b7e4369fda71f0

Observation bcf34f7e-8d07-4f11-b174-56cac699311d · outbound

This paper cites Think- ing with video: Video generation as a promising multimodal reasoning paradigm.

Thinking in Video: Can Video Generators Really Reason About the Real World? Think- ing with video: Video generation as a promising multimodal reasoning paradigm

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.688624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.688624Z digest=sha256:2ec42ecc0b3dea3cb523a0da82a5a60fb09cfaa40bc1389733a03854b162a5cb

Observation 79b302f7-46fa-476c-9608-8aaf018b30bb · outbound

This paper cites Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks.arXiv preprint arXiv:2511.15065, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks.arXiv preprint arXiv:2511.15065, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.902542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.902542Z digest=sha256:71cc9468c8d1b5b594ab50d9a0d56f8e14c0e62bd8102aa4b1c9a10e613c2425

Observation 2849c63f-b004-45be-a2a4-aea3c427aa1b · outbound

This paper cites Weave: Unleashing and benchmarking the in-context interleaved comprehension and generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Weave: Unleashing and benchmarking the in-context interleaved comprehension and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.063535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.063535Z digest=sha256:cd9086f522be99304dae4f5c8d94eea872412de4d30ee5e2ef567d1f4f0e60e0

Observation eb068e54-0837-4871-aaad-0c0c9c255731 · outbound

This paper cites Tivibench: Benchmarking think-in-video reasoning for video generative models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Tivibench: Benchmarking think-in-video reasoning for video generative models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.188868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.188868Z digest=sha256:6b7cbf20b4af7d8e0fbd3ea6cfc3ed751945c25299adc6f5c66e9654d01c0dd0

Observation df8b54b6-e33e-4890-b847-28dae3b270fa · outbound

This paper cites Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.292420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.292420Z digest=sha256:3a2356e38c3138fbba1dc18cbc01f3f10f7192c49d0cb5c7af3eb648fe804a82

Observation 340e3464-76b5-4ebb-8910-eeda2658bb8a · outbound

This paper cites Ggbench: A geometric generative reasoning benchmark for unified multimodal models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ggbench: A geometric generative reasoning benchmark for unified multimodal models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.399742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.399742Z digest=sha256:3fb456670324e1082a1a46f8475961612f3ce67f94ee960fc2848968ab87a51d

Observation c588afa7-58f2-49db-bfbf-e07be89a9d08 · outbound

This paper cites Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.

Thinking in Video: Can Video Generators Really Reason About the Real World? Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.550254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.550254Z digest=sha256:67b83cc61152287165d6f8019848a53e56665e656b83ee8c8963f289273c06b8

Observation 7c54f164-3993-4179-9b48-475a6b2ee650 · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Physgen: Rigid-body physics-grounded image-to-video generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.733004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.733004Z digest=sha256:f7f6a670d1594268551e52ac12ec21f81bfe378dfaa7beb6d4c139841a29f8b3

Observation 4e7da424-a68d-4c27-8ee9-6419c9091ffb · outbound

This paper cites Exploring the Evolution of Physics Cognition in Video Generation: A Survey.

Thinking in Video: Can Video Generators Really Reason About the Real World? Exploring the Evolution of Physics Cognition in Video Generation: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.906970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.906970Z digest=sha256:3cf61c3c0127b6e66ea746baaf6d3a8976f53c7b5219be186dee9c948d35d3a4

Observation 2d7942be-06a5-4a8f-8624-74c3ee6f3621 · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29, 2016.

Thinking in Video: Can Video Generators Really Reason About the Real World? Improved techniques for training gans.Advances in neural information processing systems, 29, 2016

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.045401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.045401Z digest=sha256:c77df298b291c4a5e8fb0bec15c69cb1842760126208763c35f573eff6818487

Observation 4384fc51-e1a3-477b-8302-9fe4f393075b · outbound

This paper cites FVD: A new metric for video generation,.

Thinking in Video: Can Video Generators Really Reason About the Real World? FVD: A new metric for video generation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.215955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.215955Z digest=sha256:4a7aa79de9dbb46ed48f2ad66aa93bfada5b1e3ea1cc604d9f11deb09962f165

Observation 4bbd32dc-1b98-4c46-bbe5-53e1ca6f2694 · outbound

This paper cites On the content bias in fréchet video distance.

Thinking in Video: Can Video Generators Really Reason About the Real World? On the content bias in fréchet video distance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.494735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.494735Z digest=sha256:f85a81309f367281caf09465e81ccba7bf23a88f69a7a64600840099feaae558

Observation 0993dd37-f5d0-4ae7-a738-2c0c41d2a414 · outbound

This paper cites The Essential Role of Causality in Foundation World Models for Embodied AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? The Essential Role of Causality in Foundation World Models for Embodied AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.632064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.632064Z digest=sha256:159bb822ef9f4fac39f9252c67054fc5b1e75400bd8a5f4ba397d633ec1af92b

Observation f4d139b7-71ad-4e7d-a290-2d3d1a41e8fd · outbound

This paper cites Diffusion art or digital forgery? investigating data replication in diffusion models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Diffusion art or digital forgery? investigating data replication in diffusion models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.796963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.796963Z digest=sha256:777118385aefe1f147111ca755162bb3a8e78b6ca2d6cfae8eb72cc5531161ea

Observation 8d3fbe95-ac2c-4864-9de3-afdfa3d18455 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.955102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.955102Z digest=sha256:214961381dc79fdbdd3291297fb75d0db8a71a1fd60c09cffd291458640d30a0

Observation 346eedb5-5d52-47ae-8da8-8821aa4496aa · outbound

This paper cites Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.107864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.107864Z digest=sha256:eabe1cfb197a93f9f146b0069b8ae3235e7084b3c57c2b0dba5043aee86f8742

Observation 681bab52-da7f-4f80-8208-920d355eba84 · outbound

This paper cites RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.268892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.268892Z digest=sha256:3bf115c93a471773a64eadea3796f8af1b30deed7c65eaea3094da1b178c4f76

Observation 7cfdd990-47cb-419e-9141-1f8aba064fdb · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vbench: Comprehensive benchmark suite for video generative models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.423659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.423659Z digest=sha256:80c0c77e30e48411567c560491f001e1cb1b74ecc39a43d3e2219a3357b36d2c

Observation b03da153-1e28-4dde-80d4-d80e60feac85 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Evalcrafter: Benchmarking and evaluating large video generation models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.568545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.568545Z digest=sha256:1c4358bee284a43757ef5ad332480ef2e7772c58d4ed238da5daaaabdb4f4282

Observation d16d4994-edc7-486f-8ca0-c70b18120e6d · outbound

This paper cites Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.717609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.717609Z digest=sha256:683cf0b0517f437ec4fc25db6a343ad404ca89f78ccce391ccfbcbb0f61b46e6

Observation cfbbbf81-aae1-4bb4-8f0c-b48111c36c11 · outbound

This paper cites What about gravity in video generation? post-training newton’s laws with verifiable rewards.arXiv preprint arXiv:2512.00425, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? What about gravity in video generation? post-training newton’s laws with verifiable rewards.arXiv preprint arXiv:2512.00425, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.883305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.883305Z digest=sha256:3427142a9684f3a8237650fb7f71daacea67f4451c3f84466eea358175696319

Observation f1606849-161b-4dde-92ac-bcd5521a617c · outbound

This paper cites T2v-compbench: A comprehensive benchmark for compositional text-to- video generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? T2v-compbench: A comprehensive benchmark for compositional text-to- video generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.028648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.028648Z digest=sha256:dd4f180c294b4f7360f34c767b3cc6db70e787788560be6c57d4b4b0d3321376

Observation 81b672e5-e006-454e-9c39-3baadfdbd46b · outbound

This paper cites Vibe: A text-to-video benchmark for evaluating hal- lucination in large multimodal models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vibe: A text-to-video benchmark for evaluating hal- lucination in large multimodal models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.192580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.192580Z digest=sha256:f4503f9155071e6905b3192905a75173c5c245412d94e314c63594633caba6ef

Observation 6005acd2-3838-4655-b5b2-f94ef9c5749b · outbound

This paper cites World Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? World Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.369130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.369130Z digest=sha256:a79dfc13e7a68ef9565c03e960e5a48487dd6306c835c51b28fae6f312c32f44

Observation 865b148a-0639-4b83-bed5-70275e86a37e · outbound

This paper cites A path towards autonomous machine intelligence version 0.9.

Thinking in Video: Can Video Generators Really Reason About the Real World? A path towards autonomous machine intelligence version 0.9

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.488113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.488113Z digest=sha256:ac513a18336ea3dbd108204c3a82de37aa9ac4d893c63281b9426bfdecee345e

Observation 1fe3a9cc-1e28-4c26-bd81-91db74b9caa6 · outbound

This paper cites Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability.

Thinking in Video: Can Video Generators Really Reason About the Real World? Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.650083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.650083Z digest=sha256:69947ce53c61c54b57d766915b46b1b66610f0875a14ec5e9e4a5af6c3968991

Observation bdc30d75-403e-4271-bb63-5728bbb293b0 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.778104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.778104Z digest=sha256:54cd3de86d96f9292b555e201190cbf477dbc7584f90ab0aa04ac8906d52a912

Observation 6d7ff8a3-ca4f-4be4-a1bd-f9fb1fdcd440 · outbound

This paper cites The sound of water: Inferring physical properties from pouring liquids.

Thinking in Video: Can Video Generators Really Reason About the Real World? The sound of water: Inferring physical properties from pouring liquids

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.959513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.959513Z digest=sha256:b887f78c74084a74656cde3744cdccb5ca344cc05854f145038fd71c2002d5a1

Observation e1672087-e3a2-4a30-8032-89efe9002863 · outbound

This paper cites Physion: Evaluating Physical Prediction from Vision in Humans and Machines.

Thinking in Video: Can Video Generators Really Reason About the Real World? Physion: Evaluating Physical Prediction from Vision in Humans and Machines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.071407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.071407Z digest=sha256:5cfc71dd3f8e6b1ad466f5d8af86eed7b4891159c1cece3b64a9f0361ac70265

Observation 56746adb-4ee7-424c-b99f-fe235f2ce35d · outbound

This paper cites TLD: A Vehicle Tail Light signal Dataset and Benchmark.

Thinking in Video: Can Video Generators Really Reason About the Real World? TLD: A Vehicle Tail Light signal Dataset and Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.177637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.177637Z digest=sha256:979b5c1701d73d319138036a15356344b325a2f59f6baab702099e5ae705c2e2

Observation 3119fe02-6076-421d-8623-8fb47675dd61 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Thinking in Video: Can Video Generators Really Reason About the Real World? The Kinetics Human Action Video Dataset

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.299311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.299311Z digest=sha256:f98537799da9e3f513a67dacaa0785b29ff7d3c0c780e3f62f84e856cdc514f8

Observation 0f149436-6454-4854-ba04-5fbd4a2ec778 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.456060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.456060Z digest=sha256:647bcd832a8b46b33aa572b98347bfedea57cb7b1a8adb713b85c7ec1d52b28b

Observation 0c291af9-b8e3-4bc6-b573-09e9e32aa427 · outbound

This paper cites Gemma 3 Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? Gemma 3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.624753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.624753Z digest=sha256:2badb910072ab49b2e479cdb9caa7b79df13145c6beb275845031d453673b47a

Observation 7c66c50e-c190-47bc-94e1-ae62a2776045 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Thinking in Video: Can Video Generators Really Reason About the Real World? Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.737381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.737381Z digest=sha256:edad9d8fd6870be90280dd18426b0880250b559201e53490bf8f205b77b91038

Observation bb04f007-3916-4d36-9401-c8711314876a · outbound

This paper cites Qwen3-VL Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? Qwen3-VL Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.848598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.848598Z digest=sha256:8f6051ae78a55eed70fc04d6ec261e32bbf38960791308b1020421e632386362

Observation 289c7a03-9513-4594-ae01-e88413b45fc5 · outbound

This paper cites OpenAI GPT-5 System Card.

Thinking in Video: Can Video Generators Really Reason About the Real World? OpenAI GPT-5 System Card

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.966954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.966954Z digest=sha256:fda29290c07633a994143ae2b73e9e3537849353b2183e7698225f1d10825cf5

Observation a6f56fdd-2510-4055-8df2-90eceae09a42 · outbound

This paper cites Scaling zero-shot reference-to-video generation.arXiv preprint, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Scaling zero-shot reference-to-video generation.arXiv preprint, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.088675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.088675Z digest=sha256:10c301b478bd71f8c47a24f1d35197155011be5c936d63290303e3de1c8845bc

Observation 4917b50f-6f35-4b29-a0d0-58d904216787 · outbound

This paper cites Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.237328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.237328Z digest=sha256:610e8d556d059d967610cda9eaeb1218b797b4c6e1e70c84f652987ae9089cd0

Observation f2523d10-3c97-4cc1-bb65-31af0cb879ce · outbound

This paper cites Stable video infinity: Infinite-length video generation with error recycling.arXiv preprint arXiv:2510.09212, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Stable video infinity: Infinite-length video generation with error recycling.arXiv preprint arXiv:2510.09212, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.350191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.350191Z digest=sha256:f2fe0353553a931a5f5bc4efd1f18c7a306629b828707fd06222c19db13611b0

Observation 5f57e935-e817-4ec8-8ad7-79753866fcaa · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

Thinking in Video: Can Video Generators Really Reason About the Real World? LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.458101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.458101Z digest=sha256:5ae631c4cb4eb131f3c9545595d0517a335d4a6438f895c9b73294a26f8ef62a

Observation 349359b0-731f-4220-8f57-5b05ee3c02d4 · outbound

This paper cites V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.611214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.611214Z digest=sha256:723d57c28bf6b8596c21c70774bfc044ed5e03b3fb839a105b516496652d5f78

Observation 27ab5293-97ed-488f-9cb7-7bcc0a468d3a · outbound

This paper cites Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.825883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.825883Z digest=sha256:c5689891007fc9248cf0d10d6189b5d8797e22b346e6e8a2fe832ca509267cd7

Observation 2fc1c363-0dc2-4bfa-a372-cdecc44e767f · outbound

This paper cites Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.956508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.956508Z digest=sha256:3f20528357c7e10cb9e3b89b1685b45d88ac22b1b1bf102466c68c7f2de9f7e2

Observation 542dec33-ec6a-4828-b0a5-1c969dd8ed89 · outbound

This paper cites Can world simulators reason? gen-vire: A generative visual reasoning benchmark,.

Thinking in Video: Can Video Generators Really Reason About the Real World? Can world simulators reason? gen-vire: A generative visual reasoning benchmark,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.071487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.071487Z digest=sha256:7687a110212620f1c1cf96be9848e8c6c351221c1db394553b522d9d5544d25a

Observation 8d8fd854-c9e1-49ed-a6c3-6e3595abf3dc · outbound

This paper cites CCHall: A novel benchmark for joint cross-lingual and cross- modal hallucinations detection in large language models.

Thinking in Video: Can Video Generators Really Reason About the Real World? CCHall: A novel benchmark for joint cross-lingual and cross- modal hallucinations detection in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.292641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.292641Z digest=sha256:37d71e7cd634545b5fc7daccddd3287befe59b9f9f697be2919579181b3d5356

Observation 6d0f5da2-5c50-4887-a9a1-c7fd7163c42a · outbound

This paper cites Large language models meet nlp: A survey.Frontiers of Computer Science, 20(11):2011361, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Large language models meet nlp: A survey.Frontiers of Computer Science, 20(11):2011361, 2026

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.421131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.421131Z digest=sha256:fdbf0a5111f20597bfeca95dd06d1dc02bc67872663f4dbe0cd574d45b006163

Observation dcc1aacd-a901-45d8-a871-ea3b0a1777c7 · outbound

This paper cites Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.arXiv preprint arXiv:2510.26802, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.arXiv preprint arXiv:2510.26802, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.478604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.478604Z digest=sha256:205a9a6562d2515cb5f4fc355917f19f8bf7a9c0c572ee7a39f0fa82f285a134

Observation ba003690-3d5e-4339-81a0-d116a936937c · outbound

This paper cites ISBN 979-8-89176-251-0.

Thinking in Video: Can Video Generators Really Reason About the Real World? ISBN 979-8-89176-251-0

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.365314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.365314Z digest=sha256:e23f22924d3a983fa97b1938b25cd2224d91ca764b6e26164f316f3eb4f923ae

Observation 30cf12d7-beed-4d12-9caa-a1ba8ba98eb2 · outbound

This paper cites Latent Visual Cache for Video Reasoning.

Thinking in Video: Can Video Generators Really Reason About the Real World? Latent Visual Cache for Video Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.597841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.597841Z digest=sha256:934baf5c2947dd09da3e2f28b055678801359d260a74b3cb549a0ec17b9667bf

Observation e985d63c-cca1-4f9b-9f41-16b1dcd9e458 · outbound

This paper cites Mmgr: Multi-modal generative reasoning.arXiv preprint arXiv:2512.14691, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Mmgr: Multi-modal generative reasoning.arXiv preprint arXiv:2512.14691, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.651779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.651779Z digest=sha256:29522a7dfea46bcd6999b99cc136172e702b4ced576d2a381c6f81a72005f05d

Observation 38e2166c-e519-4414-a023-a8ccac217909 · outbound

This paper cites Video models start to solve chess, maze, sudoku, mental rotation, and raven’matrices.arXiv preprint arXiv:2512.05969, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video models start to solve chess, maze, sudoku, mental rotation, and raven’matrices.arXiv preprint arXiv:2512.05969, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.540223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.540223Z digest=sha256:2d91b0fe664c2a3136a6431b73e3bd977db72393e5ba1ac5b78933e6b76e5179

Observation 307fd135-8114-41b9-a092-782e9b338d94 · outbound

This paper cites Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.750039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.750039Z digest=sha256:89aec65a55ee9910843315b69b2ebb47ba96b7c9dd2bb1ceff0a564795638518

Observation 605f4c84-aa7a-48b6-9c5d-73b0fd3f4a94 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.331772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.331772Z digest=sha256:264b94ad61fbc1eea3494c33f6c884264acd228c34ec5822bf0cbca93c9778ec

Observation 179ec37e-e83d-4301-9e90-1933f2fb7440 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.188710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.188710Z digest=sha256:e2faf5644e088aa113abcb360cd797aa037f71f324a4795bd192ad4c4077c25c

Pith citing papers

No inbound Pith citation observations are available.