Pith. sign in

Paper Citation Record · LEDGER

Towards Understanding Camera Motions in Any Video

As of 24 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 18 inbound Pith citation observations for arXiv:2504.15376.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15376 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:31:56.629967Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:01:32.165361Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.919006Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved54
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c54bb016-e98f-4b83-8775-f0af14e55556 · outbound

This paper cites The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing.

Towards Understanding Camera Motions in Any Video The anatomy of video editing: A dataset and benchmark suite for ai-assisted video editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.889056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.252069Z digest=sha256:9bdfe78e7cae3cc2a54dad470e5c7ee8aa3f2bfb4e9d5e60ec4fcc981df06b95

Observation f81330bf-6c70-45e6-862b-517b081a15ad · outbound

This paper cites AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers.

Towards Understanding Camera Motions in Any Video AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.258440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.258440Z digest=sha256:f3f05dde4f27d05a5edee4e0983a10241e951bd7d09c04bd76930b1f3023f928

Observation a7717c8b-0f33-455f-912e-4372ebbf1651 · outbound

This paper cites VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control.

Towards Understanding Camera Motions in Any Video VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.264019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.264019Z digest=sha256:b8089d2caef4eb2004f8ecd2f35171cffd77e2628f5322e1269e7196a8b745ae

Observation fdb7c8e7-ee6c-4f90-9ab7-1f46c709a41d · outbound

This paper cites Qwen2.5-VL Technical Report.

Towards Understanding Camera Motions in Any Video Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.269418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.269418Z digest=sha256:8a3dadd4b2a43d3efbe455b813920f639c21a35d464b483581dd26467d779919

Observation accaba5b-d68a-491b-9d1f-baec2974b67c · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Towards Understanding Camera Motions in Any Video AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.274698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.274698Z digest=sha256:7349da51fdea9b7b095c60ef8c038bda5e23e462226e5dbd2f008dc36ae59c56

Observation c2004e56-598b-46ed-974f-df4c2b89c741 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

Towards Understanding Camera Motions in Any Video SkyReels-V2: Infinite-length Film Generative Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.279937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.279937Z digest=sha256:8280b2cbb07b4b809e9ea48fac41f7b3665ae53b53fc9fd7e116099ccb1a9321

Observation 8e191f2c-bb0b-42af-894f-7019a1eb7067 · outbound

This paper cites Goku: Flow Based Video Generative Foundation Models.

Towards Understanding Camera Motions in Any Video Goku: Flow Based Video Generative Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.285275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.285275Z digest=sha256:51dc46b967272e78407434c558770b5ae78b1df806d2aa25b77ac8ea26696e0c

Observation 194de423-8b6b-4acf-94fb-c7fa89165f39 · outbound

This paper cites Boosting Camera Motion Control for Video Diffusion Transformers.

Towards Understanding Camera Motions in Any Video Boosting Camera Motion Control for Video Diffusion Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.290615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.290615Z digest=sha256:d1ea8f92b0a6a238036852b139aff1e9f80f421b0b4d47d4d648253d2fa197c6

Observation 0cdbb948-919d-40b4-acd8-d76b40151c35 · outbound

This paper cites Optimal structure from motion: Local ambiguities and global estimates.

Towards Understanding Camera Motions in Any Video Optimal structure from motion: Local ambiguities and global estimates

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.874327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.295078Z digest=sha256:ba07fab99773612f0dbe453311a56c8bba7bdadba1f92a88e3f87f5306c827f9

Observation bc8485f0-ba4a-4194-ba59-721496c639f7 · outbound

This paper cites Et the exceptional trajectories: Text-to-camera-trajectory generation with character awareness.

Towards Understanding Camera Motions in Any Video Et the exceptional trajectories: Text-to-camera-trajectory generation with character awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.859637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.299791Z digest=sha256:d6b8e95e08f846f0fc6d6c5d7be7352a2acf5457bad0fb0413aca046c841ff8c

Observation 05e04c7e-d9af-4f9a-8802-59a1ab559042 · outbound

This paper cites Understanding noise sensitivity in structure from motion.

Towards Understanding Camera Motions in Any Video Understanding noise sensitivity in structure from motion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.844861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.304159Z digest=sha256:b77274f0b3cceaaa3251e7b523c944b453eba75dfbdfda8591f9b75070d46144

Observation a44c2f2d-ed5b-41e9-82f1-53c6640bc6a3 · outbound

This paper cites Monoslam: Real-time single camera slam.

Towards Understanding Camera Motions in Any Video Monoslam: Real-time single camera slam

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.308530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.308530Z digest=sha256:5d3e50fff108ab4dbe334ca6b19ec7f3217d587363b862c7228c5c984aa5a041

Observation beed10a5-b494-403d-909a-37bcdae994f4 · outbound

This paper cites Types of camera movements in film explained: Definitive guide, 2020.

Towards Understanding Camera Motions in Any Video Types of camera movements in film explained: Definitive guide, 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.821017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.313096Z digest=sha256:a28852d421f74abbc8772df1d86d9c1e2ffddb3daa085a2d10cfa1b34847c85b

Observation 545c363a-3ed7-4e3f-8229-cec04c572bdd · outbound

This paper cites MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion.

Towards Understanding Camera Motions in Any Video MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.317723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.317723Z digest=sha256:cb0a427d28d873279d79cbcd182881e156440113f9d5b91605c7411fed3ea083

Observation 39d37a5e-1ef6-42e8-a58c-bc8ffcdc7009 · outbound

This paper cites Cinematic journeys: Film and movement.

Towards Understanding Camera Motions in Any Video Cinematic journeys: Film and movement

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.805875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.322432Z digest=sha256:9131d5bbf02f3c45893d98b50d96b7ac6fdf29a12eb07344d52c995628583245

Observation b2e3959d-9749-4199-a70c-7150ab5a1047 · outbound

This paper cites Lsd-slam: Large-scale direct monocular slam.

Towards Understanding Camera Motions in Any Video Lsd-slam: Large-scale direct monocular slam

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.326706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.326706Z digest=sha256:5707ecc9ddc8df94c36d799270aa5ca472476f19f8ec86c0d7f49c7604c0e911

Observation 0ac2ce22-0e9d-4171-b651-a5333cfa160a · outbound

This paper cites Ambiguity in structure from motion: Sphere versus plane.

Towards Understanding Camera Motions in Any Video Ambiguity in structure from motion: Sphere versus plane

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.782580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.331335Z digest=sha256:b011eda2db285bd252d8d727ea70f7ea80cd7ef7c648aaf0f50968866cb28e57

Observation 2ddb2f6d-1772-402d-b384-4ca8c33c6df4 · outbound

This paper cites Motion parallax and absolute distance.

Towards Understanding Camera Motions in Any Video Motion parallax and absolute distance

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.767325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.335953Z digest=sha256:24844cbb15ab38bd11a6bea9de0868cb64a6b4aba4a75f3026d606bf89d252d0

Observation a540214e-40af-4634-8f5c-3869f94607cb · outbound

This paper cites The ecological approach to visual perception.

Towards Understanding Camera Motions in Any Video The ecological approach to visual perception

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.751836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.340442Z digest=sha256:c8dcbee8ce3bb26191337a4bf4afaf12a2f91d1f048321493db15885c557ac83

Observation 430c5e9d-989a-4db9-aeb5-e003df86707e · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Towards Understanding Camera Motions in Any Video Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.345091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.345091Z digest=sha256:615da30730a43a5e285433193d5b8fac3ae4cb6630c073c3bd9ca17f26a06e28

Observation 2a24dbd7-63d1-4f24-9b47-f851edfb3b64 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

Towards Understanding Camera Motions in Any Video Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.349754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.349754Z digest=sha256:0bbb49fc85edbf372970e3268ca41e190f4c291b50cbba6f7d9e519195c54e21

Observation 2f9c43d9-b117-4825-96a8-3cf6a98ca19a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards Understanding Camera Motions in Any Video DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.354373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.354373Z digest=sha256:91f437ef275feba2ad68e87bdfe66c73e6aa1e13d617ca3a47a885c472d47e95

Observation 98a39f6a-1f83-4ca4-b286-686804300c6b · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Towards Understanding Camera Motions in Any Video CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.359376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.359376Z digest=sha256:b3be53e02726f000611c1464014b7b5a6a09a7a18123a8bb95fd288b3e5bc5eb

Observation d463f054-8252-4d31-8bdf-5061e46b5dde · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Towards Understanding Camera Motions in Any Video CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.364098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.364098Z digest=sha256:3f5c4d94b9db7e072a440102e124de8c8d54f705b22ba0f912d0174469b1c154

Observation d8f146aa-1697-494e-8a8c-d57fcf873e06 · outbound

This paper cites MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models.

Towards Understanding Camera Motions in Any Video MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.369178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.369178Z digest=sha256:8d3312883f8f03f551a49623b050122a4c76d65297da2a705c80f6144bfaa098

Observation b88b164a-1527-47b1-bd16-db773aba8bca · outbound

This paper cites Learning Camera Movement Control from Real-World Drone Videos.

Towards Understanding Camera Motions in Any Video Learning Camera Movement Control from Real-World Drone Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.373996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.373996Z digest=sha256:0329b86ded0dde008312b6feb75ec9a66421f1a770e097756587f14b06d73f4b

Observation 5691f187-d59f-4c7f-8bb0-ce620a8a1016 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

Towards Understanding Camera Motions in Any Video Movienet: A holistic dataset for movie understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.716321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.379122Z digest=sha256:8488fa44d00e1dd643e30e7465ec0f6c52ae22adf8c8c37fbf19df3ff85e715d

Observation 1b56008b-6a2b-4696-9f28-ed128c074d76 · outbound

This paper cites Cinematographic camera diffusion model.

Towards Understanding Camera Motions in Any Video Cinematographic camera diffusion model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.700062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.383477Z digest=sha256:402549ae016b59b25df4770f21367917bdaaa78b65883dd9bcf4bd1c608e10bb

Observation 6dc3cf78-ab73-43a5-a23b-714df86e6745 · outbound

This paper cites Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos.

Towards Understanding Camera Motions in Any Video Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.387923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.387923Z digest=sha256:5d213fc1bbdf69920f4da96fa5fcac4c2d2895210930cba723b0764fac11cea7

Observation 92eed851-c879-4518-a321-f88129306833 · outbound

This paper cites AnimateAnything: Consistent and Controllable Animation for Video Generation.

Towards Understanding Camera Motions in Any Video AnimateAnything: Consistent and Controllable Animation for Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.392654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.392654Z digest=sha256:960e8a6b0c15977c587bdb157db9a7ad61e26f4dc02191ca389ca274317ff6c7

Observation c2db6ad2-4d91-42cb-b98e-2f293c0bf4a8 · outbound

This paper cites Evaluating and improving compositional text-to-visual generation.

Towards Understanding Camera Motions in Any Video Evaluating and improving compositional text-to-visual generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.685508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.397771Z digest=sha256:cb6b76469c94e32566a5c504e2e5a7bf442b177541a9ea9d03db4928843d3eec

Observation 713c1d77-b5a0-4ca7-96e9-2a1700735e90 · outbound

This paper cites Naturalbench: Evaluating vision-language models on natural adversarial samples.

Towards Understanding Camera Motions in Any Video Naturalbench: Evaluating vision-language models on natural adversarial samples

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.670981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.402293Z digest=sha256:19a9d6b5f08c6b0e75effe2d390c744be2c48caa4be979286efb45fccbe59b49

Observation d723d971-3cd1-4517-bcec-fbb382390815 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Towards Understanding Camera Motions in Any Video LLaVA-OneVision: Easy Visual Task Transfer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.406857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.406857Z digest=sha256:f968dbe0f56c609de269991d67631b0582cf664dd641c563dbbf708386cbad8f

Observation 49e4444d-50fc-4f7b-bab0-a6694352bbe1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Towards Understanding Camera Motions in Any Video BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.411838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.411838Z digest=sha256:a8d9cc4128e1e37ee4a99042538f3a390ca5fc61689b823e83394ad978e867f3

Observation 40f5637a-9e6d-45a8-83a0-0d1725b13853 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

Towards Understanding Camera Motions in Any Video Unmasked teacher: Towards training-efficient video foundation models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.655657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.416716Z digest=sha256:d367ded66179b2eded121c98db5511beee5c6a2d15bdb32e756c4f913dab915c

Observation 0d0b4cfa-6c03-4f3d-b303-e9884970246f · outbound

This paper cites RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control.

Towards Understanding Camera Motions in Any Video RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.421465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.421465Z digest=sha256:9505bc81d9c4666d518b5b811a9be462d62a4b7a9de1db32b9668f543b83bf7e

Observation da734167-211c-4d1c-a33b-8ed90ab449ab · outbound

This paper cites Can video generation replace cinematographers? Research on the cinematic language of generated video.

Towards Understanding Camera Motions in Any Video Can video generation replace cinematographers? Research on the cinematic language of generated video

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.426125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.426125Z digest=sha256:b05373afc4d7bec46444e0dafe1d956ae6a796ee8200176d1ec2f19069464afa

Observation 6c57dced-b6bd-4e26-8913-07409ef61ca6 · outbound

This paper cites MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos.

Towards Understanding Camera Motions in Any Video MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.430839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.430839Z digest=sha256:042a1cae45a1aca93091e07a887073862cdbabdca50e7345d563d7df3e6bb812

Observation cd01364b-fee4-4072-929c-885a395e1243 · outbound

This paper cites Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models, 2023.

Towards Understanding Camera Motions in Any Video Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.640371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.435524Z digest=sha256:1e380f58d1583a7a2a5e570388a375d310b281486a458f12d5c906f15792b874

Observation 4e3b89c1-6997-4504-8905-d1072be5caae · outbound

This paper cites Revisiting the Role of Language Priors in Vision-Language Models.

Towards Understanding Camera Motions in Any Video Revisiting the Role of Language Priors in Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.440129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.440129Z digest=sha256:55f6695e37ec42995a1e77007fb6bfde6bccec52476a3418a2c8ba8144455b1f

Observation 958bfd97-f328-49b7-b1d3-4b08c20d8ad2 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Towards Understanding Camera Motions in Any Video Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.444840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.444840Z digest=sha256:3ff30ac30678473331f24272f18af5d5072dbc67a552a61f3d744db13f9b7499

Observation 9a46f69e-d210-48ed-b606-30f628382720 · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision.

Towards Understanding Camera Motions in Any Video Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.450223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.450223Z digest=sha256:471dbed00f456246380a1b8c4091439b4772fa8b491c1084d7aa2871d9082b79

Observation 12ce1517-8066-402d-b816-8a313ccab6c1 · outbound

This paper cites Language Models as Black-Box Optimizers for Vision-Language Models.

Towards Understanding Camera Motions in Any Video Language Models as Black-Box Optimizers for Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.455342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.455342Z digest=sha256:fe4d44c782950014f80a913f9f2a782e4aae4462d3589ea89bdf4e4b237df292

Observation 34a5909f-7418-4156-bcd5-4dc071bb190a · outbound

This paper cites Chatcam: Empowering camera control through conversational ai.

Towards Understanding Camera Motions in Any Video Chatcam: Empowering camera control through conversational ai

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.615581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.460415Z digest=sha256:7e5d6388bddaa53bfa7c8ce1666c392394c0e330586a7f21c252687ad4ed8d5c

Observation 34a8e7d2-b4b6-4a8c-a794-a11837d588d5 · outbound

This paper cites GPT-4 Technical Report.

Towards Understanding Camera Motions in Any Video GPT-4 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.464994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.464994Z digest=sha256:7d14aa2459f7694fe1387aa5aaf50aae500e0f7663c336b095459320b06f6eb2

Observation 51305558-5bdb-4ba2-bd34-d1f27e75d8c4 · outbound

This paper cites The Neglected Tails in Vision-Language Models.

Towards Understanding Camera Motions in Any Video The Neglected Tails in Vision-Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.469528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.469528Z digest=sha256:d8415badfb31299b38e7ca88a08b614c1d56b382f779d818c0153adc4073feca

Observation 8df05936-99b6-4286-9707-52b9972a6bdb · outbound

This paper cites Movie gen: A cast of media foundation models.

Towards Understanding Camera Motions in Any Video Movie gen: A cast of media foundation models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.600678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.474305Z digest=sha256:a2a11d43ae21820aa708956b2d17ebf66a70ff831d1b533570ae25e2b89b01a6

Observation eb53926f-b00b-4e43-8155-de4a81340e2f · outbound

This paper cites Learning transferable visual models from natural language supervision.

Towards Understanding Camera Motions in Any Video Learning transferable visual models from natural language supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.478859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.478859Z digest=sha256:1e6ee75a0de5c7de360bbaeed1ef3427059fb04b022da3458b5aa686d2f5d583

Observation c66a896e-68ef-4b6d-9378-acced6cda7f1 · outbound

This paper cites A unified framework for shot type classification based on subject centric lens.

Towards Understanding Camera Motions in Any Video A unified framework for shot type classification based on subject centric lens

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.576429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.483209Z digest=sha256:5c4fbfdf0ae57580fc25068cdb1ee42868b65237ad72cac87c3b25a5daa72e79

Observation 2c6430ab-ca37-4373-ac90-9db3b852d0ec · outbound

This paper cites Motion parallax as an independent cue for depth perception.Perception, 8(2):125–134, 1979.

Towards Understanding Camera Motions in Any Video Motion parallax as an independent cue for depth perception.Perception, 8(2):125–134, 1979

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.561587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.487993Z digest=sha256:77fefc7ad126847d6552adc2be1444b6709b666b23067cb59c1b4fd0a5f41d71

Observation a6427f87-5a08-45f1-af59-1fb2215e5e05 · outbound

This paper cites Structure-from-motion revisited.

Towards Understanding Camera Motions in Any Video Structure-from-motion revisited

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.492394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.492394Z digest=sha256:b96b5a4b33a1b8bcf3a6419ceb53ba26c4d3ba48a9e9a293a7e570b31f510a22

Observation e51cdf39-5066-457a-9ed1-d478544a2e20 · outbound

This paper cites Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation.

Towards Understanding Camera Motions in Any Video Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.496850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.496850Z digest=sha256:9d8f033fb869a01c0ece68f7c6d941c383bc28a7e27384f0c0808f28d18a3ea1

Observation cd2172a3-50e6-4b28-887b-943632cd2997 · outbound

This paper cites Transnet v2: An effective deep network architecture for fast shot transition detection.

Towards Understanding Camera Motions in Any Video Transnet v2: An effective deep network architecture for fast shot transition detection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.501833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.501833Z digest=sha256:f8b8509925ef9a499729fac38d02f48e2f65e99219ebc76fe23042ee72b7315d

Observation 8d5323e1-d71d-40c3-aed7-1c47f60bd1b3 · outbound

This paper cites A grammar of the film: An analysis of film technique.

Towards Understanding Camera Motions in Any Video A grammar of the film: An analysis of film technique

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.526846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.506107Z digest=sha256:7e87ff23a9f055808ea8b8c3852df4996d4dc0ba473818e283f06ec79827a861

Observation 57b43a73-8850-4c63-90b3-08bc0986677a · outbound

This paper cites Visual slam algorithms: A survey from 2010 to.

Towards Understanding Camera Motions in Any Video Visual slam algorithms: A survey from 2010 to

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.512189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.510614Z digest=sha256:de070c8d4a3fcac5f67ff9c5dc1c1feadf3a27d588280baf48eeadc75f69c29d

Observation a7990c5c-a314-421d-a5ed-711bb1e5d876 · outbound

This paper cites Vidcomposition: Can mllms analyze compositions in compiled videos? arXiv preprint arXiv:2411.10979, 2024.

Towards Understanding Camera Motions in Any Video Vidcomposition: Can mllms analyze compositions in compiled videos? arXiv preprint arXiv:2411.10979, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.519524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.519524Z digest=sha256:0ec87dadd1824116d06ad0799386eae411d594b90eb0627e5f9fb38240f31831

Observation cd7f0e3a-35c3-4d29-b77c-aa4b4bb11f04 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Towards Understanding Camera Motions in Any Video Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.524009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.524009Z digest=sha256:f0d16e41ab1b6d917a89c717e237141ec96e0c67ede38fbe4e480847ac0d6e7f

Observation 69b43521-4bf4-40af-97dd-08715950dc79 · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Towards Understanding Camera Motions in Any Video Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.481654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.528838Z digest=sha256:bd2578a97cf76052b7c17a87d772be1a8be8853698c515805afb8d34a0c9138d

Observation e1e1073e-887e-47af-90f5-ee74b3a70efe · outbound

This paper cites Vggsfm: Visual geometry grounded deep structure from motion.

Towards Understanding Camera Motions in Any Video Vggsfm: Visual geometry grounded deep structure from motion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.533385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.533385Z digest=sha256:78e633df6992274d077e49207089acf93cec15e758628c580b2cdba80f8ca929

Observation 5c05d351-778a-496b-9157-98155d2fb5f4 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Towards Understanding Camera Motions in Any Video Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.538046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.538046Z digest=sha256:0d6868b9a5c6dae6375e20e0c7b43f68c39e5f9086effb5c7c40bef4bb33ac17

Observation 49154c57-10c3-4d2d-acff-39452055c994 · outbound

This paper cites Continuous 3D Perception Model with Persistent State.

Towards Understanding Camera Motions in Any Video Continuous 3D Perception Model with Persistent State

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.543345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.543345Z digest=sha256:0338338a4f55683a22511730fa12c7e6fac8b03ba7b5654185480475dfa0f2b8

Observation fe7c90a6-0914-4366-93df-b08451e1e27e · outbound

This paper cites Dust3r: Geometric 3d vision made easy.

Towards Understanding Camera Motions in Any Video Dust3r: Geometric 3d vision made easy

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.547973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.547973Z digest=sha256:44d488ac9230a4fb4155345e8ffe6c053f0b881d7b0ac7942e494f4cf9b595e5

Observation 7c84f8ba-0b31-43f5-afcd-dd55dc26c0ed · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Towards Understanding Camera Motions in Any Video Internvideo2: Scaling foundation models for multimodal video understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.447191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.552469Z digest=sha256:6afe6c8c535d9d5cff88f9af820df54d8f0f9cb931518b49879c64b3fe463b67

Observation c13e7384-7680-4cde-bc54-3664d9db2ddb · outbound

This paper cites CPA: Camera-pose-awareness Diffusion Transformer for Video Generation.

Towards Understanding Camera Motions in Any Video CPA: Camera-pose-awareness Diffusion Transformer for Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.557164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.557164Z digest=sha256:da2dcbc6d13b2d5a04ec593638a6f61dada33f60be22ea71ed133db13d4dec3d

Observation 6c1241ba-f910-4fea-bafa-7d1c940d5128 · outbound

This paper cites Motionbooth: Motion-aware customized text-to-video generation.

Towards Understanding Camera Motions in Any Video Motionbooth: Motion-aware customized text-to-video generation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.431615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.562638Z digest=sha256:92869569795aedbc0d2b4593fc0ce5aca68c710dccf3f90744d9d9cee8b49ecc

Observation e751bc5a-ee03-48c2-bae1-bf5a391a1f79 · outbound

This paper cites Trajectory Attention for Fine-grained Video Motion Control.

Towards Understanding Camera Motions in Any Video Trajectory Attention for Fine-grained Video Motion Control

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.567142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.567142Z digest=sha256:8b8f64a476fe3bf760fd3848ef3a4699719f23172e4c552bf799dc948539ad24

Observation 2605e6ca-d8fa-49ca-bfc0-35ed89f33818 · outbound

This paper cites MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation.

Towards Understanding Camera Motions in Any Video MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.571810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.571810Z digest=sha256:0c008b786312fe9a7c9b236f7349b73cfe51496feda7a3fe6406e79b2ee65929

Observation 316e1eff-a610-46e2-8250-f6d02fc2e671 · outbound

This paper cites Direct-a-video: Customized video generation with user-directed camera movement and object motion.

Towards Understanding Camera Motions in Any Video Direct-a-video: Customized video generation with user-directed camera movement and object motion

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.576677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.576677Z digest=sha256:a3eabec74328fdfc5cee3f5461ec25a9c862fb903919a26e0fdf29a0e7acb7e2

Observation db84b5e2-f00e-4e45-9445-43f40a683aa4 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Towards Understanding Camera Motions in Any Video CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.581101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.581101Z digest=sha256:68fb5ff3d3f65fe7463c886c23b2f35c3a61ebc449417a4fdb56644e2fd5ce39

Observation 85cdf9c8-dcd6-41b6-b758-a480572fb9ad · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

Towards Understanding Camera Motions in Any Video Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.585601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.585601Z digest=sha256:020c7d6b6a5b704dbd079df61782f9cf51e5337350f463451f96fe6b5cec8348

Observation 61a438b0-bfb6-49f1-9949-94f0d8549a88 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

Towards Understanding Camera Motions in Any Video InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.590340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.590340Z digest=sha256:714ff4fcaeb1059db4501581416f9e087e0ae41ad4bded8fb10acb9540e69fd1

Observation 3dad9ab5-fc2e-420b-bb21-ccd8b7691102 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Towards Understanding Camera Motions in Any Video LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.595009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.595009Z digest=sha256:2c748835e0ea092e64bb92b548254d259c1426142f6586c4e57063c68d44ac01

Observation 35aec037-5ca6-40b3-8106-f2f4822493df · outbound

This paper cites Structure and motion from casual videos.

Towards Understanding Camera Motions in Any Video Structure and motion from casual videos

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.599970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.599970Z digest=sha256:bd131dc24098321b647f60af8a2956c48a0f1868a3e0f4cc7aa427f5a4034af2

Observation f7b9e279-574e-42f4-b805-606a3f4650c4 · outbound

This paper cites VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation.

Towards Understanding Camera Motions in Any Video VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.604627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.604627Z digest=sha256:7534a8f83db3162dc0a86ac536d8708d89f546315ebc0a4db2cc56a12f4390ca

Observation 5bd11483-4e0e-48b5-9d22-6632337f1a79 · outbound

This paper cites Stereo Magnification: Learning View Synthesis using Multiplane Images.

Towards Understanding Camera Motions in Any Video Stereo Magnification: Learning View Synthesis using Multiplane Images

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.609266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.609266Z digest=sha256:6810f12fe43b5710bc304d521414537d6bbd5e0ec2eeb8f4b63eaadaa371e1be

Observation 822936e3-7891-4abc-93ed-32d1f31dd543 · outbound

This paper cites Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training.

Towards Understanding Camera Motions in Any Video Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T11:31:56.614294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:31:56.614294Z digest=sha256:445cad039e9897da7d750e06c42bca4090599b0d479e5c87cfb3a1c2067785d7

Observation 2cf50800-f62e-48ad-ad3d-a838b8a13be2 · outbound

This paper cites the camera zooms in for a push shot.

Towards Understanding Camera Motions in Any Video the camera zooms in for a push shot

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:31:57.397945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.619693Z digest=sha256:9b04d601ed744b2f07188038669134cee413d85f953975be694eabaeaf7c5244

Observation c2a284e2-9185-48f2-a719-ae44accdb98f · outbound

This paper cites an unresolved cited work.

Towards Understanding Camera Motions in Any Video Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:31:57.381270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.625203Z digest=sha256:8cce6670ce7df2bfd28fdf7eaaa8b76cafab43cfaebd2beb94c646ec99e463b5

Observation a181f2ec-b58e-4747-9d98-538559842eef · outbound

This paper cites north” or toward the top of the frame, and backward motion ( dolly-out) as moving “south.

Towards Understanding Camera Motions in Any Video north” or toward the top of the frame, and backward motion ( dolly-out) as moving “south

Reference 80

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T11:31:57.364891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.629967Z digest=sha256:11dc73a111141411f986c5458e0f956ed8fd4298073751281c49376cf74a8630

Observation 414dbcd9-68f1-447e-bc8a-ed9fa28adea6 · outbound

This paper cites an unresolved cited work.

Towards Understanding Camera Motions in Any Video Unresolved cited work

Reference 2016

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:31:57.496846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:31:56.515070Z digest=sha256:c20f7de85bdb1950ae161c76b062328f528e197d586bb6853d5d075be240430f

Pith citing papers

Observation 127364a5-59da-42e0-910c-65fda40ffce6 · inbound

MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual Triplets cites this paper.

MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual Triplets Towards Understanding Camera Motions in Any Video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:32.165361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:32.165361Z digest=sha256:3ae86b449c6c742aee272a8ca0e568b2e9daa22d5a800f0748f32708337174cd

Observation 6e78c1f2-76d4-4f05-9d49-5a51a5a27c98 · inbound

Leveraging Optimal Transport for Distributed Two-Sample Testing: An Integrated Transportation Distance-based Framework cites this paper.

Leveraging Optimal Transport for Distributed Two-Sample Testing: An Integrated Transportation Distance-based Framework Towards Understanding Camera Motions in Any Video

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:49.976023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:49.976023Z digest=sha256:c5b0a32849ea1e868ebad4dade63cf359a375b0ae7cca6f0db80a07ec6a4954f

Observation 580e7ed0-d1f7-4435-968c-c454c2334b5e · inbound

ViPE: Video Pose Engine for 3D Geometric Perception cites this paper.

ViPE: Video Pose Engine for 3D Geometric Perception Towards Understanding Camera Motions in Any Video

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:08.740824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T16:41:08.620285Z digest=sha256:58b0787faea745741bc6564739c2f8b35d74544ab389a5b53a0089baa10ec832

Observation 25a997bb-2a43-4045-b099-9f931125154c · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding Towards Understanding Camera Motions in Any Video

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.685627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:244f6849acf0bb78edc9980a96e66645c70b23bff33339321ecf62941e502d03

Observation 711ea883-a30c-4f23-9442-4316bc0f85f9 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Towards Understanding Camera Motions in Any Video

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.471377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:e3cc18edb86811f99a3358f316db9ebf7dcfd910b483be4ff14ed8618cb8c1b6

Observation 220a4362-b9a8-4aaf-9cc7-53c8e9864681 · inbound

HumanScore: Benchmarking Human Motions in Generated Videos cites this paper.

HumanScore: Benchmarking Human Motions in Generated Videos Towards Understanding Camera Motions in Any Video

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:04.336350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:17:16.513512Z digest=sha256:888a77d94969a82ddfa1d9202857b38e53399433ab306650fcc55064c3d83449

Observation fc14d044-fdaf-4d48-9efb-2c47221a3f00 · inbound

OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer cites this paper.

OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:49.149459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T04:15:20.045060Z digest=sha256:1e7bc13430de21e790a7955794bc10b2a46e80bf5c8489c75306f812f91f39e8

Observation 245941f5-f87c-4279-a090-4c45357dd2de · inbound

DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models cites this paper.

DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models Towards Understanding Camera Motions in Any Video

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:30.177543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T17:42:00.634333Z digest=sha256:e9a4d290845a26ac76a04d45b572cbf5c65591cf4703764368c5aa7b326ef021

Observation 1d7bdf41-e112-4a18-aecf-9514926b0aa0 · inbound

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs cites this paper.

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs Towards Understanding Camera Motions in Any Video

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.446378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:10:27.595446Z digest=sha256:3e3bbc9862b1214395dccc014c32e8be22e10fa6c397b163aff4fcc780ecae03

Observation f2102e95-3756-42a1-8188-7d646072807c · inbound

Probing into Camera Control of Video Models cites this paper.

Probing into Camera Control of Video Models Towards Understanding Camera Motions in Any Video

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:15:04.352310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T21:11:19.441408Z digest=sha256:a820f1b39290a8c933eedb72a2bbb823e5c940ac45084518092fdc6c8e2b228c

Observation b6df5c30-0758-4b1f-a8b9-e4a030b805ea · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Towards Understanding Camera Motions in Any Video

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:04.580227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:42ff21b9c62b7b6c2c67e790f3e55cd81c0133448bbada32fbf22eae0d0cde1a

Observation 309e9f1b-23d3-45b4-adf1-cce801da9589 · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Towards Understanding Camera Motions in Any Video

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.920422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:9110c28be2ba5b728aa558661eb09d3f77cf2a87f39e24138f987f32e961b6c1

Observation 4f0f49a4-5bcf-435d-8d7f-30686c822220 · inbound

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds cites this paper.

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:49:52.340383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T04:51:38.355857Z digest=sha256:34c9e3bab3bea958273bb52bb338d4add484504ed293f69729dd91c3a701d436

Observation ed033391-79e8-45d1-a3d5-dfdef1130468 · inbound

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds cites this paper.

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:13:52.839672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T04:55:32.982338Z digest=sha256:9c954a3503e40df74146a11b74b700125abd3f5743a4383e5e5c052c3de4e7d4

Observation 91c0c318-7fc0-4dc1-a37a-11ee04506787 · inbound

Natural Language Camera Movement Understanding cites this paper.

Natural Language Camera Movement Understanding Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T05:16:19.750976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:16:19.750976Z digest=sha256:3547b186a451aee352a5825f30db6929ff0c3190a3fea3e5febdb21b8ac9f77a

Observation 18e482bd-c9b6-4f99-8506-30cf03e0c49c · inbound

GS-Agent: Creating 4D Physical Worlds With Generative Simulation cites this paper.

GS-Agent: Creating 4D Physical Worlds With Generative Simulation Towards Understanding Camera Motions in Any Video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T07:14:42.446756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:14:42.446756Z digest=sha256:d51ecbebf67e8964cf546164b5a3be0fba011f1992dead109f5739ab2f02c883

Observation 28d0b606-de80-41da-bdf2-c07efc56d5bc · inbound

Self-Supervised Learning of Structured Dynamics from Videos cites this paper.

Self-Supervised Learning of Structured Dynamics from Videos Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.948987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.948987Z digest=sha256:1d73382ae8b5680d27687c19e4aa1fa1532f6d092c82ef053e8d1fed1325454a

Observation a29f6672-bca4-4dc4-97d7-1e24f41ee6f1 · inbound

CameraAnything: Refilming Videos with Arbitrary Camera Control cites this paper.

CameraAnything: Refilming Videos with Arbitrary Camera Control Towards Understanding Camera Motions in Any Video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T11:10:41.548774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:10:41.548774Z digest=sha256:f811fc0cc9c429e665f250bd4173422820f75ca8c4ae3338fff579656bb3b20b