Pith. sign in

Paper Citation Record · LEDGER

CoT-Edit: Let CoT Guide Instruction Video Editing

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.01113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01113 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:14:36.146697Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 657b10a4-a0bb-42b0-8f78-5823b143d706 · outbound

This paper cites Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.975532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.975532Z digest=sha256:1da335b0478f7706d26b181fb067c739f75dd203b281845f7acbcc1cb90eb6dd

Observation 1856cfa5-fccf-4d0a-825a-d496f3c5a391 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.987406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:35.979280Z digest=sha256:3a7404d0a2cae5c3351e6e77e3d43a248f9d1be7851c24573e2c173dc0d1bd82

Observation 408e0bc6-58b0-4827-b1be-96613547c5e7 · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.982238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.982238Z digest=sha256:dbc30262ebbe6f41be239eddb6351ed8748ef4e0df73495ade27909ed4087a43

Observation 18f9807a-5d72-4114-a3c3-70770a3eedaf · outbound

This paper cites Consistent Video-to-Video Transfer Using Synthetic Dataset.

CoT-Edit: Let CoT Guide Instruction Video Editing Consistent Video-to-Video Transfer Using Synthetic Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.986744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.986744Z digest=sha256:589971785e62fad5f7f3295ddcb7546e3cf74a9d01020d02145a991f44eb8bea

Observation b75ff8c6-70e0-4f0d-9f67-4e751ab27f24 · outbound

This paper cites Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler.

CoT-Edit: Let CoT Guide Instruction Video Editing Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.976630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:35.990411Z digest=sha256:6effb6d215e1d3b7105de88113f0be8eab3e5cb687d6afb1be25189b96b62cc5

Observation db7eded1-6605-4037-9244-54308783348a · outbound

This paper cites Why compress what you can generate? when gpt-4o generation ushers in image compression fields.

CoT-Edit: Let CoT Guide Instruction Video Editing Why compress what you can generate? when gpt-4o generation ushers in image compression fields

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.966130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:35.994108Z digest=sha256:1c0bd9d78f7d40dde4a50afe2b6597e1ae6e9cf6f1bd267ba775cb1a07a33f0a

Observation c7e8d454-32fa-46ec-9c31-56a81201e8ce · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

CoT-Edit: Let CoT Guide Instruction Video Editing SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.997634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.997634Z digest=sha256:64a01ee2e607f1f7e11e685dd91f1e2294d0ab02121ab2a4618a80b10a14e4d9

Observation eb512c56-8fb9-4959-b4af-05742885a328 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

CoT-Edit: Let CoT Guide Instruction Video Editing Clipscore: A reference-free evaluation met- ric for image captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.953342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.001663Z digest=sha256:05983cded37098748511b94d55526dd160627f16e9677999e528079145214295

Observation 48d98d9a-6c21-44d9-86c0-995bf19d8587 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Imagen Video: High Definition Video Generation with Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.005696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.005696Z digest=sha256:22d3a739a2c68b29a1d5107f17433e140899c1b03c72a050cd8e4c8b28aebd8e

Observation f8cba0fc-3fa7-4c3d-93bf-6345952e11fa · outbound

This paper cites Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.009394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.009394Z digest=sha256:3f69eba948a9b915489b8b3074a61c74e2c5ae249f5589a20ac3cafae58e7264

Observation 25ef2ba6-7d42-469c-831d-7876d2fdfc32 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

CoT-Edit: Let CoT Guide Instruction Video Editing CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.012707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.012707Z digest=sha256:381cf7edfd6077e21f3a8f82b9cfa2c88e92457e9a48453a92328f1b91323634

Observation c1b5049f-49ca-4103-864e-8cf6af35b0b1 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

CoT-Edit: Let CoT Guide Instruction Video Editing Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.934410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.016317Z digest=sha256:13b95699ac8fc6973febd75c651b1d221df8c4cd83a0a9865803c04a83987bc8

Observation 49884281-ad6b-42ba-a8a2-3f56dd32d611 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.019799Z digest=sha256:3fb3cf4217a80776dbcc973e4537852f5a821639fbd37ceaa39094cff1b2c35c

Observation 7cde7cc7-1d24-4b9c-aac6-7a87e4f26e5f · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

CoT-Edit: Let CoT Guide Instruction Video Editing Vbench: Comprehensive bench- mark suite for video generative models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.023233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.023233Z digest=sha256:f665e24cf8a18254ac3b8dda67e4c5cce38e0afc80a5615a414d23db17b76510

Observation 4dd9d16c-bfc1-46db-894b-b7216878607e · outbound

This paper cites Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.026401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.026401Z digest=sha256:40a3f7731c1f013141650c3a92e2f9d1c95ad6b3214fbf62571f3d56262ba60e

Observation c4b2e429-8f5b-4a6f-b212-61d7d3bcaf88 · outbound

This paper cites Vace: All-in-one video creation and editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Vace: All-in-one video creation and editing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.029525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.029525Z digest=sha256:83959f668b89521eb03123202a09b9f43772755c947efaa021a8d17f61ace5cc

Observation e2c88d85-f4ca-499b-becf-0f87bb95e566 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

CoT-Edit: Let CoT Guide Instruction Video Editing Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.032244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.032244Z digest=sha256:3ad908c7d05bb364bb37a86d31444f11ef32e2b672c0e654b3190474dd08fece

Observation 0242e969-b0d2-454e-8224-5ea3437cee65 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.035188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.035188Z digest=sha256:35878da6546567688b482d48ccc90308725f819d29e8c6d5c9b7d0afa701ed3a

Observation 05eb5eef-d43d-4443-912a-ad663df13610 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

CoT-Edit: Let CoT Guide Instruction Video Editing AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.039061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.039061Z digest=sha256:348515fe142b62c914a1f2cc5a0e6a17a15080ea2b2624e9b9871c55fe804d50

Observation e3f22532-7f1b-41d9-85ae-d90033f338a7 · outbound

This paper cites OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation.

CoT-Edit: Let CoT Guide Instruction Video Editing OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.042340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.042340Z digest=sha256:727f2d7aebb7b29bcd6d3d04db0035d0c5c07d2908751427956ee13759d0f069

Observation 4c7772c1-8450-4621-a2fe-7362429c370a · outbound

This paper cites Grounding 3D Scene Affordance From Egocentric Interactions.

CoT-Edit: Let CoT Guide Instruction Video Editing Grounding 3D Scene Affordance From Egocentric Interactions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.046347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.046347Z digest=sha256:cded11b88de5e9f83e1d568c3f254f9a482410a65b6e1ed0ea7a82917a49712a

Observation b1fa8719-768e-4bd2-bd97-7155ad412c6f · outbound

This paper cites Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.904108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.049901Z digest=sha256:b2887bf894323489abd13ad03e2bb432761d23808c762f70cebec1399c4223b1

Observation 409ef355-6b9c-41e2-945a-6a7dd9d41699 · outbound

This paper cites The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density.

CoT-Edit: Let CoT Guide Instruction Video Editing The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.893023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.052062Z digest=sha256:fe4b0bb0e8d2c6b3a456bb3da9803d03de4539ee2e6a6e32cb1822bb5370a3e6

Observation 4f40e9e3-6641-4c18-a16c-9f04c1aa84a9 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

CoT-Edit: Let CoT Guide Instruction Video Editing Video-p2p: Video editing with cross-attention control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.054760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.054760Z digest=sha256:3ce45cef36bf7334df6eef9a7ae3837b244a2ac85fc9ce127183909c6d7df686

Observation 34839ddd-5a94-46ac-b8ce-f8a987f2cca6 · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

CoT-Edit: Let CoT Guide Instruction Video Editing Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.057103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.057103Z digest=sha256:175ddc86280c7bd7b8f1414dbafd01a3fd8a320405ab597ec38e0051084e22a1

Observation 5be0a624-c720-4b35-ab97-43354a0f784d · outbound

This paper cites Magic- stick: Controllable video editing via control handle transfor- mations.

CoT-Edit: Let CoT Guide Instruction Video Editing Magic- stick: Controllable video editing via control handle transfor- mations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.868611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.059802Z digest=sha256:3d01963f1533c28615d4a893f04ba6deec27e29b34c322e33d5a1fe5e1e43955

Observation 6d48c170-3a29-4bf3-98e0-67d67d71e691 · outbound

This paper cites In- structx: Towards unified visual editing with mllm guidance.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structx: Towards unified visual editing with mllm guidance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.062312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.062312Z digest=sha256:e8c3d5cf8b0bfbce91a49d45a666785818256ac119057fbfdcebd077f6a5dade

Observation d2585ad4-dec7-4090-9764-6c2391d8a21d · outbound

This paper cites Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.855427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.064649Z digest=sha256:0c4019e684890978701e467f4b951952afee37c307edfab5134251024ef85808

Observation 98cde1d6-e308-4324-ab13-0f0a5d78d10a · outbound

This paper cites Instructvid2vid: Controllable video editing with natural language instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing Instructvid2vid: Controllable video editing with natural language instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.843912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.067106Z digest=sha256:9b933a4070d6a98c58ba0ccd61c525388f83c2383a372a9a0bdc89b3bafe9764

Observation 64bf5aa5-8938-4b18-8474-807bb8e824c3 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

CoT-Edit: Let CoT Guide Instruction Video Editing Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.832657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.069273Z digest=sha256:679ab6e2944b0e518e641001a3b3cdce6e580faa406e5bb0fcf5f7dc79b94633

Observation 2141106f-c39d-43ae-bfa7-3cc3fee87fb2 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

CoT-Edit: Let CoT Guide Instruction Video Editing Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.071477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.071477Z digest=sha256:04ec2f27d9d517c8c1a257878c0bc74fe8311a1acd7b76b879f9e2b7e7dec354

Observation 3fcd91b3-7f33-49e9-8c2f-baff1f40219b · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

CoT-Edit: Let CoT Guide Instruction Video Editing Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.821389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.075387Z digest=sha256:63dc167892feaa6ac7871227b9755c70b6fdc1e0f5d8b7ee0dc1e94199a45ce7

Observation 83fff1b4-c3a0-4f43-bf40-4a6491434b31 · outbound

This paper cites Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.078800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.078800Z digest=sha256:4c406515402aea9bd61839b6c274e246930111fa164276f865b6075290b17ea3

Observation af04fc58-2e23-4c2a-bbe7-cc285382811b · outbound

This paper cites Lucy edit: Open-weight text-guided video editing, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Lucy edit: Open-weight text-guided video editing, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.811263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.081791Z digest=sha256:95c806ba7c55cd3a914fa0dfb2ad19c26fc1057c60bd08620ec3ed6a89efc025

Observation d0dc72c7-8972-4b1e-bc7e-45e9454f15ff · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.085011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.085011Z digest=sha256:23b5335f0c55ec79888ba3fa5d0db3dfa7edde55f53f7fd90b2f330107b992d5

Observation 57023dca-7e67-4dd7-8704-606aa4574292 · outbound

This paper cites Fvd: A new metric for video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Fvd: A new metric for video generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.800032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.088655Z digest=sha256:423c881b9398b1f251e37bc1b51144ba49ab5af2c13bb0150151a4b2b9299e12

Observation 9ab29126-2f77-4800-ae0e-3ff4fbd845f6 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

CoT-Edit: Let CoT Guide Instruction Video Editing Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.091525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.091525Z digest=sha256:d5b10bc8a574df6ef7b52be1db18fb5deb087a947ff28b5c14d432a30a73a6b3

Observation 3cb4d851-d6ff-425f-9239-27bf6f9efad8 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

CoT-Edit: Let CoT Guide Instruction Video Editing ModelScope Text-to-Video Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.096079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.096079Z digest=sha256:43bf09b08e8c630f680df98a0e0854d6a7c9d1d7159e6df0654c272b0ce381b9

Observation e663e69f-bf22-4cbb-9f2b-755e9db47ae8 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

CoT-Edit: Let CoT Guide Instruction Video Editing Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.788680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.099979Z digest=sha256:25212b128d06cac6ef890eadb8ed2611933646f4a28ea5459d7fa5a30f72c045

Observation 8fb1b249-ffac-4e93-a955-c98c5a7a2aa0 · outbound

This paper cites Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration.

CoT-Edit: Let CoT Guide Instruction Video Editing Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.778051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.103092Z digest=sha256:c6fea643ba15f5994849c4510b453e88296476d5e1d13e78824454e1b9d237cf

Observation f013a9ff-17b8-40ca-9547-f6c71f7c7fc6 · outbound

This paper cites Re-attentional con- trollable video diffusion editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Re-attentional con- trollable video diffusion editing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.765446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.106568Z digest=sha256:dafffeba279e967bbcc3b1de23454dd9f3cb90955270e8c6e04a72f46752cee1

Observation 43379ae1-bdc5-4091-8111-ffe361a8a947 · outbound

This paper cites Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026.

CoT-Edit: Let CoT Guide Instruction Video Editing Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.755803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.109921Z digest=sha256:2af595f4fde5d3d8880efe22f25c0a188ed9b21e60efe338683edabf7dec5e47

Observation d7a7aec6-83e1-4638-adb3-63faecad4f77 · outbound

This paper cites Phrasecut: Language-based image segmen- tation in the wild.

CoT-Edit: Let CoT Guide Instruction Video Editing Phrasecut: Language-based image segmen- tation in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.745577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.112795Z digest=sha256:e0fcc7f867820386b41dd1fa8ae35131c79591e52728ef350d1abd2485bce36c

Observation c685c6db-e0dd-4fca-b8ee-7b27984456af · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.733642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.115970Z digest=sha256:f8372c2c558ddf001261173e4279972576099bffdf437328434d379d86b2c923

Observation 2f70e7ea-a7e9-4628-b819-5d9dc8fb09d1 · outbound

This paper cites Insvie-1m: Effective instruction-based video editing with elaborate dataset construction.

CoT-Edit: Let CoT Guide Instruction Video Editing Insvie-1m: Effective instruction-based video editing with elaborate dataset construction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.721670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.119760Z digest=sha256:120e5e24b224a25045b82382edfca2b7f20335ab76328d22cc08664395c974b3

Observation 67b4273c-b3c0-4edc-8f4b-aa0bce3e75e6 · outbound

This paper cites Veg- gie: Instructional editing and reasoning video concepts with grounded generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Veg- gie: Instructional editing and reasoning video concepts with grounded generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.122798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.122798Z digest=sha256:32fafa901c756a65d64fdd05114e9da967c8a458676f36e45952463c17f9cef2

Observation 222f1b9e-0d29-495a-98a4-845715a0059e · outbound

This paper cites Editworld: Simulating world dynamics for instruction- following image editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Editworld: Simulating world dynamics for instruction- following image editing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.704257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.125635Z digest=sha256:831121e6a7d4e5d791aff9b370bc2218d886d826f40226cf90a652c8580732df

Observation 8ba8b5ea-f2e3-400f-bd20-5da9728d9eab · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.128527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.128527Z digest=sha256:c3162c335f2abe21c6d77ae28b88c47974604fc2970c1c7d9be44341829f82df

Observation 89c0ee69-c1c6-411c-8e1c-6e5af7d15788 · outbound

This paper cites EffiVED:Efficient Video Editing via Text-instruction Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing EffiVED:Efficient Video Editing via Text-instruction Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.132218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.132218Z digest=sha256:cf91d7e0957bc1a79aa1676fc8e0d441b241f27def4d6ee959919e5aab4c1214

Observation 106458a7-e319-4faf-9c0d-7a77c611b3f4 · outbound

This paper cites Motionpro: A precise mo- tion controller for image-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Motionpro: A precise mo- tion controller for image-to-video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.693440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.135815Z digest=sha256:36611d83298960f2f3931ba3636e68fb58ef67cca893dd38bc6b42ed0c7be317

Observation 892ca003-1808-4665-9ca2-76f4c6b06c01 · outbound

This paper cites Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024.

CoT-Edit: Let CoT Guide Instruction Video Editing Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.681969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.139111Z digest=sha256:565f13fa03ab58f0e89d99f5c6e71283e41a9cd922c509287e1b219782f974be

Observation 6937549a-7058-4ae0-a562-31ce580c7c73 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019.

CoT-Edit: Let CoT Guide Instruction Video Editing Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.670892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:14:36.142912Z digest=sha256:4e42c8ba5ce21249953ac1128bc401881af3a5b10cf26ccf6e0e92dbeec9bdad

Observation f2bbfc9a-9c3a-44b6-b89a-7f1cfbfb6813 · outbound

This paper cites Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists.

CoT-Edit: Let CoT Guide Instruction Video Editing Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.146697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.146697Z digest=sha256:0bb376b116a02b9c950cbac79b5472760e645dec17dbd6fc97383cd997f37032

Pith citing papers

No inbound Pith citation observations are available.