Pith. sign in

Paper Citation Record · LEDGER

Taming Teacher Forcing for Masked Autoregressive Video Generation

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 8 inbound Pith citation observations for arXiv:2501.12389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12389 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:16:47.640589Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:38:30.597797Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:33:15.180082Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8217c3c-1ec2-4900-92fd-e6926197cce8 · outbound

This paper cites Video generation models as world simulators.

Taming Teacher Forcing for Masked Autoregressive Video Generation Video generation models as world simulators

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.399672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.399672Z digest=sha256:40602f06b5081ac56e4a262c1acc3a4ae95a6f7b7dc11bb93a3315a75de96448

Observation ffde491b-2e00-4ec5-8920-a7b15e856a4b · outbound

This paper cites Ge- nie: Generative interactive environments.

Taming Teacher Forcing for Masked Autoregressive Video Generation Ge- nie: Generative interactive environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.220054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.404985Z digest=sha256:088ce1eb288226f017a1e479e9d4b2e7951ffbfc4d11719e892a1da1dcef530a

Observation a0ad7479-a4f6-4d31-92cd-85adb495846b · outbound

This paper cites A short note about kinetics-.

Taming Teacher Forcing for Masked Autoregressive Video Generation A short note about kinetics-

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.410911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.410911Z digest=sha256:def2afd51a84243f2021956fe1116fac955bda7e71e6ed2678c08d3652941c2a

Observation 0f6a07e0-4a77-47f0-90b2-a47632a79e8b · outbound

This paper cites Maskgit: Masked generative image transformer.

Taming Teacher Forcing for Masked Autoregressive Video Generation Maskgit: Masked generative image transformer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.201480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.423531Z digest=sha256:9716f2ba24d9ee102cf8548d02170d1c5d8878013380741d0a4509aea0ce3a23

Observation 732eac7f-53c7-4c1b-a645-30a724bc8542 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Taming Teacher Forcing for Masked Autoregressive Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.428417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.428417Z digest=sha256:6ac9b386a8f9b57a8a0febb0a475601ba831a6a8175290c19bff20f053827a51

Observation c65c365c-b92e-45bf-b664-3df86884fa53 · outbound

This paper cites Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion.

Taming Teacher Forcing for Masked Autoregressive Video Generation Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.433670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.433670Z digest=sha256:1f272d5d0e7f658e3e60b74f3d8ffb79d88f678f6733caa515afececd1a5754e

Observation 4938ebbd-83b1-41e1-94ed-0f664cc0ba32 · outbound

This paper cites Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,.

Taming Teacher Forcing for Masked Autoregressive Video Generation Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.438930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.438930Z digest=sha256:c6fdd94e7c98971783842fe070cf0ac84abbc1551a7d27c8e47cff0299b45999

Observation cbcca705-c57c-42c6-801d-1a7716c2711d · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

Taming Teacher Forcing for Masked Autoregressive Video Generation Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.443669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.443669Z digest=sha256:6779779ca1afcb18846550c62ae05603f3c70245eeeefcc810eb5045836ddee3

Observation 1445b4f2-170d-4b06-8734-2ab442a3606c · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for llms.

Taming Teacher Forcing for Masked Autoregressive Video Generation Model tells you what to discard: Adaptive KV cache compression for llms

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.183696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.448666Z digest=sha256:2945ba9d2058a218b8ec65bd27f41d28cbc0bdc0883f6961b3bf129a4217ec70

Observation 3747f35a-5d14-43f0-91d7-ffaa29e449cc · outbound

This paper cites Courville, and Yoshua Bengio.

Taming Teacher Forcing for Masked Autoregressive Video Generation Courville, and Yoshua Bengio

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.173588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.453799Z digest=sha256:3104e28da479b8d722877eeee520cc24d013e4616081c7560072697d2890c870

Observation 5c2b762b-3225-4277-8b34-04e9748f73ae · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Taming Teacher Forcing for Masked Autoregressive Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.459726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.459726Z digest=sha256:0bfafd1c2116b67ccbdf2707fbb09135afc8e577c2cfff8a355289b55700be65

Observation 7d6aa9f6-43dd-4825-9437-05471e6be217 · outbound

This paper cites Girshick.

Taming Teacher Forcing for Masked Autoregressive Video Generation Girshick

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.162173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.466231Z digest=sha256:d5148b89a194f445d874cff61b56ad2e7f02c7379ae98084685ec8b93da626df

Observation 9f9aade9-01a3-42b4-b197-10715474df14 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.471762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.471762Z digest=sha256:d1e271726c684682674ad2514cd5bfee0b841907dd9f4c21aaf9074af22e378e

Observation 1fef39fe-5590-46f6-86dd-06d1877d2ddd · outbound

This paper cites Denoising dif- fusion probabilistic models.

Taming Teacher Forcing for Masked Autoregressive Video Generation Denoising dif- fusion probabilistic models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.150461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.476470Z digest=sha256:53eedec3803351f8869c51e36074ab77ec6d0c40ec813b5f029fdf05aeddcb11

Observation 20159b2a-dec4-49b3-a901-7709bb2635ef · outbound

This paper cites Video dif- fusion models.

Taming Teacher Forcing for Masked Autoregressive Video Generation Video dif- fusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.138967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.480978Z digest=sha256:a3e527f60ed451fd78add684250b0ca9b95fa808c4df8a5f68965406f66d3273

Observation b9eed922-6652-4e54-b265-d2e2379095d5 · outbound

This paper cites Scalable Adaptive Computation for Iterative Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Scalable Adaptive Computation for Iterative Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.485858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.485858Z digest=sha256:e1785798f594b86ef57e4874fbfccc4c1237e26e2f7134aa19f37571e513433e

Observation 360b747a-18fe-4dfd-b74e-0848a829e321 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.490673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.490673Z digest=sha256:f3c1272560d9a1a92e73dd967cd42ec664f6998fe89d0e962e2447446e753437

Observation 81b6d41c-4923-40d1-8397-0282b6c91ce6 · outbound

This paper cites Dart: Noise injection for robust imitation learning.

Taming Teacher Forcing for Masked Autoregressive Video Generation Dart: Noise injection for robust imitation learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.125866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.495655Z digest=sha256:a43e4475d9eda9f730454378c7b7cbd72eb97f7928f4694d402d6eaab85cd3e0

Observation fcb101ee-7e9c-4071-bc4e-82a394b57b61 · outbound

This paper cites Autoregressive image generation without vec- tor quantization.

Taming Teacher Forcing for Masked Autoregressive Video Generation Autoregressive image generation without vec- tor quantization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.112751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.500178Z digest=sha256:9a3df4676a1ab0ca7a78c4992d404387f2392cc97898879014fdaa2eb3fc0673

Observation 4bc11b39-4bc1-4e78-b94c-91c8ebb49620 · outbound

This paper cites Keep the cost down: A review on methods to opti- mize LLM’s KV-cache consumption.

Taming Teacher Forcing for Masked Autoregressive Video Generation Keep the cost down: A review on methods to opti- mize LLM’s KV-cache consumption

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.098726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.505391Z digest=sha256:f4dd7c68c8b2b96db7dd249f504d8cfbd69a03b02db48194a77c72e8495e4298

Observation c565ebda-6104-4ab8-8f56-29cb9a5cc815 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.509903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.509903Z digest=sha256:3289d86206fc422246cf417e74eb63280d906f373e4affa18d20994e1518b6ea

Observation c9bd6572-827e-4f82-9b6a-c9a64e551df3 · outbound

This paper cites Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.515030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.515030Z digest=sha256:14b5a05e6d619746e6993776d8342f783479e331ba2caa9075c71f297aded916

Observation 8ca12290-9639-49e7-931b-b5d69692e94d · outbound

This paper cites Cosmos tokenizer.

Taming Teacher Forcing for Masked Autoregressive Video Generation Cosmos tokenizer

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.085171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.519984Z digest=sha256:a7b24b83ed0994ab3b2dd26c16177b6fc27eaa5426e224555c4e4bd3140956c9

Observation 5f30d9b2-3407-4071-95f8-236f20103467 · outbound

This paper cites Improving language understanding by gener- ative pre-training.

Taming Teacher Forcing for Masked Autoregressive Video Generation Improving language understanding by gener- ative pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.071303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.524356Z digest=sha256:40a5a43dc412345295df7fdf9abd47de8113c4721504340555f910ab891e63ce

Observation 4841542e-b71a-45a4-ba93-583127cb2c90 · outbound

This paper cites Language models are unsu- pervised multitask learners.

Taming Teacher Forcing for Masked Autoregressive Video Generation Language models are unsu- pervised multitask learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.057648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.528509Z digest=sha256:8c004ba2f425776e2494aa850f6bb9fb48e7f276b1091fa184be39417a0fb3dc

Observation 0fd74360-c00e-47ac-959d-3f07d4277060 · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Zero-Shot Text-to-Image Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.533348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.533348Z digest=sha256:1b567a4539af4f73ac27fcf3e8eee306c279976470ef279f1970333f32bafbab

Observation 4cfb776c-15b3-4afa-b194-77d3ae117dce · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Taming Teacher Forcing for Masked Autoregressive Video Generation High-resolution image synthesis with latent diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.044619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.538415Z digest=sha256:f7bcfdf4542f982143bc5f833e1341ee8a1e2eb7ef6a28f9cd353d39d3495de1

Observation 1b4086cd-b949-400e-bb3e-e2f776919639 · outbound

This paper cites Generalization in generation: A closer look at exposure bias.

Taming Teacher Forcing for Masked Autoregressive Video Generation Generalization in generation: A closer look at exposure bias

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.031417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.542539Z digest=sha256:0307c6aa70c33ed2c1ec86370e351551fcb340a2d4cbf7c0eaf640cbab87643c

Observation 2e1a87e8-e7f9-4084-a594-9a27adf35bb0 · outbound

This paper cites Mostgan-v: Video generation with temporal motion styles.

Taming Teacher Forcing for Masked Autoregressive Video Generation Mostgan-v: Video generation with temporal motion styles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.016832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.546816Z digest=sha256:64598345343676db2063b58ea1cef52da933f96032312635f7404f8c109e1b5e

Observation ab58ef13-103c-403c-a0e1-709477ce37cd · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Taming Teacher Forcing for Masked Autoregressive Video Generation Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.003201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.550916Z digest=sha256:fd72f63acc99dc5552a82dc521d9b72b2653da1294a8217bbd9c6a8519bafc83

Observation 712eb7b0-13ba-4a16-9873-7280b52418ca · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Taming Teacher Forcing for Masked Autoregressive Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.555283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.555283Z digest=sha256:997b350b77d1094e3e08e2cbe0acfe88268e867be9622443bd28f5885580b4a8

Observation 781189e5-eab2-4e27-972f-02b01933f344 · outbound

This paper cites DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis.

Taming Teacher Forcing for Masked Autoregressive Video Generation DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.559474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.559474Z digest=sha256:6c178d84087bca8ed3237811d67e672311617dc0b81b4a769d88507ffd9bf665

Observation b7f2e466-cb66-47fe-b237-27131a2c7756 · outbound

This paper cites A Good Image Generator Is What You Need for High-Resolution Video Synthesis.

Taming Teacher Forcing for Masked Autoregressive Video Generation A Good Image Generator Is What You Need for High-Resolution Video Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.563491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.563491Z digest=sha256:1bb49d67dc44f9b30ac64bc61f7d1025d619f6deb2bb03ecdc0312796cafc58d

Observation 620a12e9-8525-42e0-833a-bdddd904b986 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Mocogan: Decomposing motion and content for video generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.568171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.568171Z digest=sha256:2b3a1fcc83b9de06ec3bd9ba6796413d97aa49c1196280dcc54a6b14cabc744a

Observation eb982e26-a697-4d87-bdb8-d9b38f1bf0e4 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Taming Teacher Forcing for Masked Autoregressive Video Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.572241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.572241Z digest=sha256:263e96e1463cec2b862742e5881cbe10306f911389f47a07b94bdcbe547a1038

Observation 55f9b719-ed24-401f-88fb-51c87abc9ecd · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Taming Teacher Forcing for Masked Autoregressive Video Generation Diffusion Models Are Real-Time Game Engines

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.576958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.576958Z digest=sha256:5ff5ac4ff1f6400f7b621e219a4512ff4c466113307f0a102fe3f1a9cbebcb3f

Observation 6ccbc150-f911-4b07-8925-2c523cbc6894 · outbound

This paper cites Neural discrete representation learning.

Taming Teacher Forcing for Masked Autoregressive Video Generation Neural discrete representation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.983020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.581515Z digest=sha256:6a92e15b1dc65a37b3ef415993328007ee14ae3703578ad5c0e8f32b826a4102

Observation d09973aa-30d3-4d5e-a283-40586f223a62 · outbound

This paper cites Attention is all you need.

Taming Teacher Forcing for Masked Autoregressive Video Generation Attention is all you need

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.969438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.585602Z digest=sha256:83381894b2dfc40ff41ef26cbe074a1656f57b1838cc939c03767fa0506e9d3e

Observation 05191d2d-92ee-447f-bcbf-bb57554b3d3a · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Taming Teacher Forcing for Masked Autoregressive Video Generation Phenaki: Variable length video generation from open domain textual descriptions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.957586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.590142Z digest=sha256:78e0fa164071a5c7fda78f0ee5fd484c0d6f382a3a9ec43f5e112996963fdfc2

Observation 80260792-2f5c-477d-a704-a08243a0995d · outbound

This paper cites Predicting Video with VQVAE.

Taming Teacher Forcing for Masked Autoregressive Video Generation Predicting Video with VQVAE

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.594537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.594537Z digest=sha256:086ca009913f57f2043368a19da45e756171c9aa448c91c5207c1a8323337e9f

Observation df6172df-8e15-4382-aa81-bcc7abf63b89 · outbound

This paper cites OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.599371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.599371Z digest=sha256:17f1f2b71b0e06b84f85f72c6cdf2a515975f82cb23d4954029886a991eb488d

Observation 7f39279b-c28e-4617-bd95-c2194542e5bf · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Taming Teacher Forcing for Masked Autoregressive Video Generation Emu3: Next-Token Prediction is All You Need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.604165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.604165Z digest=sha256:3064fa2b6711a85f6c187bc71950b023a948eb7da92e4131e796955e9529bcab

Observation a85941ed-ffd9-4442-a632-1245d13c93db · outbound

This paper cites Williams and David Zipser.

Taming Teacher Forcing for Masked Autoregressive Video Generation Williams and David Zipser

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.945800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.608815Z digest=sha256:e908c8a6ef69b05dd57267f2732905e3846b05bf2830cbb9ac04dcd6057478d0

Observation 8e5f15cc-13c9-4714-9b9c-a9ba42b1bce1 · outbound

This paper cites N ¨uwa: Visual synthesis pre- training for neural visual world creation.

Taming Teacher Forcing for Masked Autoregressive Video Generation N ¨uwa: Visual synthesis pre- training for neural visual world creation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.932679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.613408Z digest=sha256:3cfb635e9740a891f0b9d9fa9e6172200548fed69c6a3c381ddf9cfeb414b159

Observation 1a778c2c-eb55-42ff-a069-2a61d8e47074 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Taming Teacher Forcing for Masked Autoregressive Video Generation VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.617853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.617853Z digest=sha256:e621bca7e0b4112fa740d2fe91d8067ae89e3a33af52aeb4292157458d8f992d

Observation f9606f3f-391e-49cf-a136-982f5ac0cdef · outbound

This paper cites Magvit: Masked generative video transformer.

Taming Teacher Forcing for Masked Autoregressive Video Generation Magvit: Masked generative video transformer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.919331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T17:16:47.622561Z digest=sha256:2f4c217abaaf1b3073a65671ae1db3f636469a732ec5bf16e7c76e46af443f52

Observation 02f907dd-8d63-440c-8fb2-970dcbbf1b5e · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.626847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.626847Z digest=sha256:d87106377a9b160f32ad3f117cef9fca2b7e26e4d8344f0e7f2e2bb528a1e2cb

Observation 38ab387c-7f8e-4f81-901f-9c640634f9eb · outbound

This paper cites Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks.

Taming Teacher Forcing for Masked Autoregressive Video Generation Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.631553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.631553Z digest=sha256:df5ee74c3fa30b148018f35b703b40775b8a89d016158931ded4993bec680270

Observation adb9fa19-0050-4e3e-9ea8-a9e00091d17f · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Taming Teacher Forcing for Masked Autoregressive Video Generation Video probabilistic diffusion models in projected latent space

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.636085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.636085Z digest=sha256:03e618292d1647dc04a29dd4df2eab24b7f588fe65313e70a312754d2b904486

Observation b6517ac2-4092-4ea4-b88b-0c4730de4ddd · outbound

This paper cites Bridging the Gap between Training and Inference for Neural Machine Translation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Bridging the Gap between Training and Inference for Neural Machine Translation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.640589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.640589Z digest=sha256:24babc5a5d943ec8187ecafc98048532dcae2d85306492d46dc873d7d1ec1a37

Observation e6e7aa42-9bc0-4694-a878-c8052c99a8b8 · outbound

This paper cites A Short Note about Kinetics-600.

Taming Teacher Forcing for Masked Autoregressive Video Generation A Short Note about Kinetics-600

Reference 600

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.415805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.415805Z digest=sha256:7761e7a6fee2b58cd8820ca4f4efa0d9981aa38210fb1dcfeb030044ff4dbf01

Pith citing papers

Observation 345c7c81-7ca0-4f60-9dcd-030355ebec85 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.524350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6805d6d2a8c27819ec6ce602ea75e17eb6f25bcc5e9d26ee8c48b53acdaefe9c

Observation 55028210-12d4-41f0-a85b-e1e6bffa0950 · inbound

Long-Context Autoregressive Video Modeling with Next-Frame Prediction cites this paper.

Long-Context Autoregressive Video Modeling with Next-Frame Prediction Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:05:17.429070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:05:17.201790Z digest=sha256:c3e66277c00cffcc63e0d43014f0f24a766e0513e8993941944524fef514f57a

Observation 3c4f6c91-da64-4f9a-bd15-5b1934e03dc2 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.751003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:d07a8ceb1d00c81d13a8686a5bc7ab67e12b893d75b8ad7fac1805021d5670e5

Observation cd409e12-1029-48c4-809c-a2fcf0b85d56 · inbound

Video-GPT via Next Clip Diffusion cites this paper.

Video-GPT via Next Clip Diffusion Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:30.597797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:30.597797Z digest=sha256:dac3068640ebf4dd39b85222f9789421d2ef0be94b58e06fc4db0a4b485af733

Observation b302fce3-bb57-4116-9db3-a28e75632476 · inbound

VideoMAR: Autoregressive Video Generatio with Continuous Tokens cites this paper.

VideoMAR: Autoregressive Video Generatio with Continuous Tokens Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:30:29.656888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:29.656888Z digest=sha256:63affe52ef1038b25d4ecc6186bba6af0b028b72633459ba1f9135cd1810554a

Observation 6fa80ef5-e219-4eeb-9890-eb40db8a9af9 · inbound

LongLive: Real-time Interactive Long Video Generation cites this paper.

LongLive: Real-time Interactive Long Video Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:52:59.443581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T03:52:59.287555Z digest=sha256:8b9d4bdbb2c774ffffa8785148f730858a0672b19985dee023e0c06a625de31c

Observation 802ff53e-027c-462f-90df-0204cb5e436c · inbound

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation cites this paper.

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:15.181604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T08:30:00.438334Z digest=sha256:74b8e6c103dc0f7995addc3a06e987560a09659214957ba63c0e4147be44b327

Observation ea3bcdc1-ed91-4ee5-aa03-278c26125255 · inbound

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation cites this paper.

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:58.686040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:58.686040Z digest=sha256:f4df802facd4941a2600d550c5bc922b81f5cadee19833c1f0a9df6abba9ea44