Pith. sign in

Paper Citation Record · LEDGER

Thinking in Video: Can Video Generators Really Reason About the Real World?

As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2607.17523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17523 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:47:13.750039Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3b7b776-f939-4895-a031-237815ada913 · outbound

This paper cites Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Youtu-llm: Unlocking the native agentic potential for lightweight large language models, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:04.983265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:04.983265Z digest=sha256:a3f7e27eefcdf43cadf9a91ae0aebaf425ef925e23c5bf4bfa857c1467be1dac

Observation b268883f-6059-47e9-814b-78ba70662e6f · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Thinking in Video: Can Video Generators Really Reason About the Real World? Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.062133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.062133Z digest=sha256:ad67e5fa72d3f4d85ed81a1b1e158ab4ffebf99a52972576fc61034d6f054294

Observation 3362a094-430d-475b-b4a9-73d0ce640b1d · outbound

This paper cites Sora 2.https://openai.com/index/sora-2, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Sora 2.https://openai.com/index/sora-2, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.183713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.183713Z digest=sha256:230f77bcfa53b1a7dcb9f4d421b176404b5f67c931f933854764b0b70c108ebb

Observation 92675ba5-d455-41dd-a856-07ad00e3b75d · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? HunyuanVideo 1.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.356951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.356951Z digest=sha256:fe9c7db7e67cbce2a79a583945e2bcae9d8623423a27489760002ed715bef342

Observation 6c2d970b-d489-4ddd-ba65-e5a47f9ce6df · outbound

This paper cites Veo 3.https://aistudio.google.com/models/veo-3, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Veo 3.https://aistudio.google.com/models/veo-3, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.530332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.530332Z digest=sha256:c80cf25759f7b89a94f0fc4d30dd25db3f26b5fee96ec1f3a9fdc16e5486d976

Observation ea11f366-8280-42aa-bc23-0f5cda6c1acc · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Wan: Open and Advanced Large-Scale Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.639163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.639163Z digest=sha256:b2ae0b25034353d5e5b8a3a622c8da96aec0589fc0f5a3b410a13ac6abcbf66e

Observation 41acdf30-f207-439f-a381-4e6b87a1b8b2 · outbound

This paper cites From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.751020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.751020Z digest=sha256:5990b241e9829bf14048df67809ee0dfef7bca2126b16425357b9d9096003283

Observation 6f71d0e0-f3e8-46bc-a99b-beb811933f59 · outbound

This paper cites World Simulation with Video Foundation Models for Physical AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? World Simulation with Video Foundation Models for Physical AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:05.927225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:05.927225Z digest=sha256:76b09c2303b30df7d70037864d2b525eaaa93a2df9d924323c7a1beb4b5aca1c

Observation 183adc3c-406c-406f-9150-098284b56556 · outbound

This paper cites Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.134736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.134736Z digest=sha256:ec312dac2e9b52aec80574af289d8e716b4e8f1d618420c84d6c65c44a4d330a

Observation 5920ead6-c2f8-48fb-9f7d-3462801e86f6 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.272600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.272600Z digest=sha256:f7439c8c1fd45054af4b69f048757a0c3520df60030f051de58da971dfab0207

Observation 28b0cf78-ac9b-4bb2-842d-17a99a13261c · outbound

This paper cites The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking.

Thinking in Video: Can Video Generators Really Reason About the Real World? The past mistake is the future wisdom: Error-driven contrastive probability optimization for chinese spell checking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.472741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.472741Z digest=sha256:b747bd33ef8414b5a5c59243ed1a6a949d360e950c0d8cc5af47a94ada8f15f7

Observation bcf34f7e-8d07-4f11-b174-56cac699311d · outbound

This paper cites Think- ing with video: Video generation as a promising multimodal reasoning paradigm.

Thinking in Video: Can Video Generators Really Reason About the Real World? Think- ing with video: Video generation as a promising multimodal reasoning paradigm

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.688624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.688624Z digest=sha256:225e9256e42ae5314871bb684c6899f25a7756361d6efe5f6776a7b64e9c7cf1

Observation 79b302f7-46fa-476c-9608-8aaf018b30bb · outbound

This paper cites Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks.arXiv preprint arXiv:2511.15065, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Reasoning via video: The first evaluation of video models’ reasoning abilities through maze-solving tasks.arXiv preprint arXiv:2511.15065, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:06.902542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:06.902542Z digest=sha256:313ed3b959b9f56f22c23bbc5b90b0b713c953fc1f5f705dbebd3d5a1021a21c

Observation 2849c63f-b004-45be-a2a4-aea3c427aa1b · outbound

This paper cites Weave: Unleashing and benchmarking the in-context interleaved comprehension and generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Weave: Unleashing and benchmarking the in-context interleaved comprehension and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.063535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.063535Z digest=sha256:6fd0b56fcd34e13ea7d8777c4704997f84f4ff8b50e3b06100081eec4b0e1b61

Observation eb068e54-0837-4871-aaad-0c0c9c255731 · outbound

This paper cites Tivibench: Benchmarking think-in-video reasoning for video generative models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Tivibench: Benchmarking think-in-video reasoning for video generative models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.188868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.188868Z digest=sha256:6e3b402a4cc2c347f36d4fa34dbd538fae3c2fcf1e99819ec743eb123c135148

Observation df8b54b6-e33e-4890-b847-28dae3b270fa · outbound

This paper cites Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Evaluating gemini robotics policies in a veo world simulator.arXiv preprint arXiv:2512.10675, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.292420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.292420Z digest=sha256:03d85fb73fe5346f6334f915531444f726ba172ffd6fd4c2eaf6553d5a68a257

Observation 340e3464-76b5-4ebb-8910-eeda2658bb8a · outbound

This paper cites Ggbench: A geometric generative reasoning benchmark for unified multimodal models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ggbench: A geometric generative reasoning benchmark for unified multimodal models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.399742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.399742Z digest=sha256:414cb8f1cd2ada108f643f1f7a596a9585a0322c873f1d4e0097de5eeaa95ce9

Observation c588afa7-58f2-49db-bfbf-e07be89a9d08 · outbound

This paper cites Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.

Thinking in Video: Can Video Generators Really Reason About the Real World? Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.550254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.550254Z digest=sha256:9d1c6dd545c5c3baa81920ce7101e35cb2c9e3874395d5e67d09ba4e94fb0928

Observation 7c54f164-3993-4179-9b48-475a6b2ee650 · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Physgen: Rigid-body physics-grounded image-to-video generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.733004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.733004Z digest=sha256:590976d7f943914ad2cc0602104d2d5dee000917ecde7a0597a40034ae472872

Observation 4e7da424-a68d-4c27-8ee9-6419c9091ffb · outbound

This paper cites Exploring the Evolution of Physics Cognition in Video Generation: A Survey.

Thinking in Video: Can Video Generators Really Reason About the Real World? Exploring the Evolution of Physics Cognition in Video Generation: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:07.906970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:07.906970Z digest=sha256:c925b3a02c9092299cb4d1e9ce63be94405ed889f600b960f91d587b832cb020

Observation 2d7942be-06a5-4a8f-8624-74c3ee6f3621 · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29, 2016.

Thinking in Video: Can Video Generators Really Reason About the Real World? Improved techniques for training gans.Advances in neural information processing systems, 29, 2016

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.045401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.045401Z digest=sha256:1c7523f772711744c0b94752ecade0dd8dcfa7aec9baaf342bbe9c523a153228

Observation 4384fc51-e1a3-477b-8302-9fe4f393075b · outbound

This paper cites FVD: A new metric for video generation,.

Thinking in Video: Can Video Generators Really Reason About the Real World? FVD: A new metric for video generation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.215955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.215955Z digest=sha256:590fa5a6c48296565539b6044f25de4aa3971ef188bdd153059ae58262a87c93

Observation 4bbd32dc-1b98-4c46-bbe5-53e1ca6f2694 · outbound

This paper cites On the content bias in fréchet video distance.

Thinking in Video: Can Video Generators Really Reason About the Real World? On the content bias in fréchet video distance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.494735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.494735Z digest=sha256:955ec2f588b7754f5d44b246b920c5977d744610cd27a5f6174ca6646881e145

Observation 0993dd37-f5d0-4ae7-a738-2c0c41d2a414 · outbound

This paper cites The Essential Role of Causality in Foundation World Models for Embodied AI.

Thinking in Video: Can Video Generators Really Reason About the Real World? The Essential Role of Causality in Foundation World Models for Embodied AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.632064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.632064Z digest=sha256:093a555e9fb533b7e8f4bfcb139e0f586e8f71c6060846073234df11e644aae3

Observation f4d139b7-71ad-4e7d-a290-2d3d1a41e8fd · outbound

This paper cites Diffusion art or digital forgery? investigating data replication in diffusion models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Diffusion art or digital forgery? investigating data replication in diffusion models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.796963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.796963Z digest=sha256:376b04cb4d9c45d947752507695c28b3e3333fde4ec2d88f55955689602790c3

Observation 8d3fbe95-ac2c-4864-9de3-afdfa3d18455 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.955102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.955102Z digest=sha256:e52b021dda8d1c4a3bfeee13545d502c16629cffa0a62cdeeb4efc853e0d2b8a

Observation 346eedb5-5d52-47ae-8da8-8821aa4496aa · outbound

This paper cites Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.107864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.107864Z digest=sha256:07bb2aac11fe3211398456320b88edd78524ff2ab50548947fc80338dbfbbe06

Observation 681bab52-da7f-4f80-8208-920d355eba84 · outbound

This paper cites RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.268892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.268892Z digest=sha256:c3ce48b4ba34b73cc5cabaff096e4785be59545825ad894594e5036aab54c0ac

Observation 7cfdd990-47cb-419e-9141-1f8aba064fdb · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vbench: Comprehensive benchmark suite for video generative models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.423659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.423659Z digest=sha256:1b452bd042c1ae1c5a663bfaa5f1ecff3a6ca0acdab6bd8dcae381c44f857675

Observation b03da153-1e28-4dde-80d4-d80e60feac85 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Evalcrafter: Benchmarking and evaluating large video generation models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.568545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.568545Z digest=sha256:d1aa62946b3d1754ef5537804b78389525b3e34fb8173f17141f0226047c6fc7

Observation d16d4994-edc7-486f-8ca0-c70b18120e6d · outbound

This paper cites Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.717609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.717609Z digest=sha256:2ba6e32aa8843d00b4649227f5ae40139606b124d9cb91c3be172e386fc5fa5b

Observation cfbbbf81-aae1-4bb4-8f0c-b48111c36c11 · outbound

This paper cites What about gravity in video generation? post-training newton’s laws with verifiable rewards.arXiv preprint arXiv:2512.00425, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? What about gravity in video generation? post-training newton’s laws with verifiable rewards.arXiv preprint arXiv:2512.00425, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:09.883305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:09.883305Z digest=sha256:b2684568d754b651c654b5f6c80bd2b573e7e41b23b759b3863c3a77a0e31efc

Observation f1606849-161b-4dde-92ac-bcd5521a617c · outbound

This paper cites T2v-compbench: A comprehensive benchmark for compositional text-to- video generation.

Thinking in Video: Can Video Generators Really Reason About the Real World? T2v-compbench: A comprehensive benchmark for compositional text-to- video generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.028648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.028648Z digest=sha256:39babfb7e5d979c36fbe0a071193d988cad5a0742106b735c8fe2a0ae64b2150

Observation 81b672e5-e006-454e-9c39-3baadfdbd46b · outbound

This paper cites Vibe: A text-to-video benchmark for evaluating hal- lucination in large multimodal models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vibe: A text-to-video benchmark for evaluating hal- lucination in large multimodal models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.192580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.192580Z digest=sha256:6bc577fcfc404d2fc9101a05b307c787f128a4823357c176688fdecb6bb92eee

Observation 6005acd2-3838-4655-b5b2-f94ef9c5749b · outbound

This paper cites World Models.

Thinking in Video: Can Video Generators Really Reason About the Real World? World Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.369130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.369130Z digest=sha256:af1a39eedaac5121753a566f71b12d13a77c61810638018ddcdf0128bdb0436c

Observation 865b148a-0639-4b83-bed5-70275e86a37e · outbound

This paper cites A path towards autonomous machine intelligence version 0.9.

Thinking in Video: Can Video Generators Really Reason About the Real World? A path towards autonomous machine intelligence version 0.9

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.488113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.488113Z digest=sha256:28fe03d9fe4383240a5bc50664a672159964dbf898e69ceefdbfba5249700ec7

Observation 1fe3a9cc-1e28-4c26-bd81-91db74b9caa6 · outbound

This paper cites Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability.

Thinking in Video: Can Video Generators Really Reason About the Real World? Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.650083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.650083Z digest=sha256:ab54cd8f3a86496a3495040af2fa178d183d0cb8fbdb42b34ac0d8b962d0f1de

Observation bdc30d75-403e-4271-bb63-5728bbb293b0 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.778104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.778104Z digest=sha256:82b33b0cfec2d934d075b720cf18f6c20e8b332eac3847202449644be80e9e4b

Observation 6d7ff8a3-ca4f-4be4-a1bd-f9fb1fdcd440 · outbound

This paper cites The sound of water: Inferring physical properties from pouring liquids.

Thinking in Video: Can Video Generators Really Reason About the Real World? The sound of water: Inferring physical properties from pouring liquids

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:10.959513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:10.959513Z digest=sha256:f5d01d56d713b3a2c80c82b4c92487a07584b0bdcd84fd7cb8c6dc9b53f3ba66

Observation e1672087-e3a2-4a30-8032-89efe9002863 · outbound

This paper cites Physion: Evaluating Physical Prediction from Vision in Humans and Machines.

Thinking in Video: Can Video Generators Really Reason About the Real World? Physion: Evaluating Physical Prediction from Vision in Humans and Machines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.071407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.071407Z digest=sha256:73a140eb10e7758284a913f16e1927c1a1f4c98fbfcaea26d6ef0d68918964e4

Observation 56746adb-4ee7-424c-b99f-fe235f2ce35d · outbound

This paper cites TLD: A Vehicle Tail Light signal Dataset and Benchmark.

Thinking in Video: Can Video Generators Really Reason About the Real World? TLD: A Vehicle Tail Light signal Dataset and Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.177637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.177637Z digest=sha256:c4d7383e1737bbb3b8174945f78a6479ee1f8d3b9ab5f1c4fa1ff40d1474bd84

Observation 3119fe02-6076-421d-8623-8fb47675dd61 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Thinking in Video: Can Video Generators Really Reason About the Real World? The Kinetics Human Action Video Dataset

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.299311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.299311Z digest=sha256:47840ad1487244101f525e7b53601b91d0374d3c20c69556c9d444fbeb434847

Observation 0f149436-6454-4854-ba04-5fbd4a2ec778 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ego4d: Around the world in 3,000 hours of egocentric video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.456060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.456060Z digest=sha256:966a7049c8e67ca3bb5b4166d2e30245b009dbece006e247e840893cd514e0ee

Observation 0c291af9-b8e3-4bc6-b573-09e9e32aa427 · outbound

This paper cites Gemma 3 Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? Gemma 3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.624753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.624753Z digest=sha256:9dbc465c9fe6a501f9bfdd04a1cd53b9377e1413f1e33ed76cae5906e6a45524

Observation 7c66c50e-c190-47bc-94e1-ae62a2776045 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Thinking in Video: Can Video Generators Really Reason About the Real World? Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.737381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.737381Z digest=sha256:2fd7824a9259294ea028ef73412fe230045806b15ab0bac8dcfc29d10a08dca4

Observation bb04f007-3916-4d36-9401-c8711314876a · outbound

This paper cites Qwen3-VL Technical Report.

Thinking in Video: Can Video Generators Really Reason About the Real World? Qwen3-VL Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.848598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.848598Z digest=sha256:54edd3964c7f9f0d69fe476b87d39e3c21980dc444ad3a134fc171c7ccfbd9f6

Observation 289c7a03-9513-4594-ae01-e88413b45fc5 · outbound

This paper cites OpenAI GPT-5 System Card.

Thinking in Video: Can Video Generators Really Reason About the Real World? OpenAI GPT-5 System Card

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:11.966954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:11.966954Z digest=sha256:ed46a5352f7cdd835063690ea7d466f111242ccaf965bb8f520ea226e3b01bb1

Observation a6f56fdd-2510-4055-8df2-90eceae09a42 · outbound

This paper cites Scaling zero-shot reference-to-video generation.arXiv preprint, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Scaling zero-shot reference-to-video generation.arXiv preprint, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.088675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.088675Z digest=sha256:f2bf547de9d8bd17a29f355425afa78ec6263160129fdc272b0b1954c9c1e5c0

Observation 4917b50f-6f35-4b29-a0d0-58d904216787 · outbound

This paper cites Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation.

Thinking in Video: Can Video Generators Really Reason About the Real World? Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.237328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.237328Z digest=sha256:94136fa538747c0ecab7c36ad0f5eecf1b8ec566d4944441ef740b4362218946

Observation f2523d10-3c97-4cc1-bb65-31af0cb879ce · outbound

This paper cites Stable video infinity: Infinite-length video generation with error recycling.arXiv preprint arXiv:2510.09212, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Stable video infinity: Infinite-length video generation with error recycling.arXiv preprint arXiv:2510.09212, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.350191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.350191Z digest=sha256:85b043cebb056d1ef524c4b10a44125ebd431a567db82085653015f3d843473c

Observation 5f57e935-e817-4ec8-8ad7-79753866fcaa · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

Thinking in Video: Can Video Generators Really Reason About the Real World? LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.458101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.458101Z digest=sha256:6fc38a94c93edef511853ca9ef6df0dc8fc2be7e5585d975357bd0a1e8888dd6

Observation 349359b0-731f-4220-8f57-5b05ee3c02d4 · outbound

This paper cites V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.611214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.611214Z digest=sha256:8edeae0eb11c7c0c4dc8a0b4c4b7e6ac35dcb2076edc5bac022ace4a50270fcf

Observation 27ab5293-97ed-488f-9cb7-7bcc0a468d3a · outbound

This paper cites Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models.

Thinking in Video: Can Video Generators Really Reason About the Real World? Vitcot: Video-text interleaved chain-of-thought for boosting video understanding in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.825883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.825883Z digest=sha256:69a2e3fc8cd884abdc04049859050d21136c11977f92b8f8afeadf47a6b77023

Observation 2fc1c363-0dc2-4bfa-a372-cdecc44e767f · outbound

This paper cites Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Ruler-bench: Probing rule-based reasoning abilities of next-level video generation models for vision foundation intelligence, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:12.956508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:12.956508Z digest=sha256:8f05c1de759cd824ec675936c883efcb985441553df5f4ab0c41b108b97ade30

Observation 542dec33-ec6a-4828-b0a5-1c969dd8ed89 · outbound

This paper cites Can world simulators reason? gen-vire: A generative visual reasoning benchmark,.

Thinking in Video: Can Video Generators Really Reason About the Real World? Can world simulators reason? gen-vire: A generative visual reasoning benchmark,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.071487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.071487Z digest=sha256:71bc88053b90adde3a62218a9697761bf3f39ff770b3275e35568a66616bf746

Observation 8d8fd854-c9e1-49ed-a6c3-6e3595abf3dc · outbound

This paper cites CCHall: A novel benchmark for joint cross-lingual and cross- modal hallucinations detection in large language models.

Thinking in Video: Can Video Generators Really Reason About the Real World? CCHall: A novel benchmark for joint cross-lingual and cross- modal hallucinations detection in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.292641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.292641Z digest=sha256:c978fb1cc3d260196d2999417b56eb16b119e14b79762fe2989e411ef5b0e825

Observation 6d0f5da2-5c50-4887-a9a1-c7fd7163c42a · outbound

This paper cites Large language models meet nlp: A survey.Frontiers of Computer Science, 20(11):2011361, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Large language models meet nlp: A survey.Frontiers of Computer Science, 20(11):2011361, 2026

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.421131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.421131Z digest=sha256:02d124d51b0c3d1c6cc2a1a15f14d97d83cf69e5b2f668ddeeb9cd5366783307

Observation dcc1aacd-a901-45d8-a871-ea3b0a1777c7 · outbound

This paper cites Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.arXiv preprint arXiv:2510.26802, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark.arXiv preprint arXiv:2510.26802, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.478604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.478604Z digest=sha256:486cbe5130b8f7c926362619b65c2eb6f499e4ec6a792b72dd586ab53444a064

Observation ba003690-3d5e-4339-81a0-d116a936937c · outbound

This paper cites ISBN 979-8-89176-251-0.

Thinking in Video: Can Video Generators Really Reason About the Real World? ISBN 979-8-89176-251-0

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.365314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.365314Z digest=sha256:dcdd0fda32c51361034b38a3c6ce7f16fffd01305fd1c275b9516490b1bf1f96

Observation 30cf12d7-beed-4d12-9caa-a1ba8ba98eb2 · outbound

This paper cites Latent Visual Cache for Video Reasoning.

Thinking in Video: Can Video Generators Really Reason About the Real World? Latent Visual Cache for Video Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.597841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.597841Z digest=sha256:ef13e9eb8b3b22686fe510044e7d10ef2ab52f73240aad6255730b3da1287fb1

Observation e985d63c-cca1-4f9b-9f41-16b1dcd9e458 · outbound

This paper cites Mmgr: Multi-modal generative reasoning.arXiv preprint arXiv:2512.14691, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Mmgr: Multi-modal generative reasoning.arXiv preprint arXiv:2512.14691, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.651779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.651779Z digest=sha256:27beb242f57e46afc5b9ef44a557b9d80ac3d03e594e3f2670430299801b793b

Observation 38e2166c-e519-4414-a023-a8ccac217909 · outbound

This paper cites Video models start to solve chess, maze, sudoku, mental rotation, and raven’matrices.arXiv preprint arXiv:2512.05969, 2025.

Thinking in Video: Can Video Generators Really Reason About the Real World? Video models start to solve chess, maze, sudoku, mental rotation, and raven’matrices.arXiv preprint arXiv:2512.05969, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.540223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.540223Z digest=sha256:9b06f984899728b815de5b17b9894141feb13faf5de14ee3552998b5fbf88c96

Observation 307fd135-8114-41b9-a092-782e9b338d94 · outbound

This paper cites Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026.

Thinking in Video: Can Video Generators Really Reason About the Real World? Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.750039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.750039Z digest=sha256:2e93ed3ad6e781196176dc05618e5aa72304931a0cafbccf129d28f46493d1c9

Observation 605f4c84-aa7a-48b6-9c5d-73b0fd3f4a94 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:08.331772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:08.331772Z digest=sha256:dfe9bc8f7824c4d987970a8f4928392aa3cfbde9255d86a7ac4c60a8fd94a1fd

Observation 179ec37e-e83d-4301-9e90-1933f2fb7440 · outbound

This paper cites an unresolved cited work.

Thinking in Video: Can Video Generators Really Reason About the Real World? Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:13.188710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:13.188710Z digest=sha256:8616446423f0efe9638315cc3d975047f3fce06aa8b735ee36318c006743e0db

Pith citing papers

No inbound Pith citation observations are available.