Pith. sign in

Paper Citation Record · LEDGER

CoT-Edit: Let CoT Guide Instruction Video Editing

As of 19 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2608.01113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01113 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:14:36.146697Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 657b10a4-a0bb-42b0-8f78-5823b143d706 · outbound

This paper cites Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Scaling instruction-based video editing with a high-quality synthetic dataset.arXiv preprint arXiv:2510.15742, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.975532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.975532Z digest=sha256:1da335b0478f7706d26b181fb067c739f75dd203b281845f7acbcc1cb90eb6dd

Observation 1856cfa5-fccf-4d0a-825a-d496f3c5a391 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.987406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:35.979280Z digest=sha256:2a52174b5352c31d49b0ba9b38d57169a76f33527abda1ab379c400b5b843422

Observation 408e0bc6-58b0-4827-b1be-96613547c5e7 · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.982238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.982238Z digest=sha256:dbc30262ebbe6f41be239eddb6351ed8748ef4e0df73495ade27909ed4087a43

Observation 18f9807a-5d72-4114-a3c3-70770a3eedaf · outbound

This paper cites Consistent Video-to-Video Transfer Using Synthetic Dataset.

CoT-Edit: Let CoT Guide Instruction Video Editing Consistent Video-to-Video Transfer Using Synthetic Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.986744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.986744Z digest=sha256:589971785e62fad5f7f3295ddcb7546e3cf74a9d01020d02145a991f44eb8bea

Observation b75ff8c6-70e0-4f0d-9f67-4e751ab27f24 · outbound

This paper cites Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler.

CoT-Edit: Let CoT Guide Instruction Video Editing Rass: Improving denoising diffusion sam- plers with reinforced active sampling scheduler

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.976630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:35.990411Z digest=sha256:2528356b88e89c2d939bba1be326d7599724f4df1bf7c91ee0202d8262984026

Observation db7eded1-6605-4037-9244-54308783348a · outbound

This paper cites Why compress what you can generate? when gpt-4o generation ushers in image compression fields.

CoT-Edit: Let CoT Guide Instruction Video Editing Why compress what you can generate? when gpt-4o generation ushers in image compression fields

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.966130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:35.994108Z digest=sha256:16f5e633fc6364965f378224f2078a1f2342d6fe4390bddebaf33acb95905725

Observation c7e8d454-32fa-46ec-9c31-56a81201e8ce · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

CoT-Edit: Let CoT Guide Instruction Video Editing SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:35.997634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:35.997634Z digest=sha256:867a1ec579271eff36c7695210e4d98afa5b01ffa6fc62662821058579bd6a51

Observation eb512c56-8fb9-4959-b4af-05742885a328 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

CoT-Edit: Let CoT Guide Instruction Video Editing Clipscore: A reference-free evaluation met- ric for image captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.953342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.001663Z digest=sha256:95ad2aeef6dea5e9f20fd8b10e2722aa795784eb8f8f9a5f30945bd09baf34e9

Observation 48d98d9a-6c21-44d9-86c0-995bf19d8587 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Imagen Video: High Definition Video Generation with Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.005696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.005696Z digest=sha256:22d3a739a2c68b29a1d5107f17433e140899c1b03c72a050cd8e4c8b28aebd8e

Observation f8cba0fc-3fa7-4c3d-93bf-6345952e11fa · outbound

This paper cites Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.009394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.009394Z digest=sha256:3f69eba948a9b915489b8b3074a61c74e2c5ae249f5589a20ac3cafae58e7264

Observation 25ef2ba6-7d42-469c-831d-7876d2fdfc32 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

CoT-Edit: Let CoT Guide Instruction Video Editing CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.012707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.012707Z digest=sha256:381cf7edfd6077e21f3a8f82b9cfa2c88e92457e9a48453a92328f1b91323634

Observation c1b5049f-49ca-4103-864e-8cf6af35b0b1 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

CoT-Edit: Let CoT Guide Instruction Video Editing Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.934410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.016317Z digest=sha256:3af42cf671c33f38afe496945acfc634a423bddcf40f0533a7a0e7e12173cddb

Observation 49884281-ad6b-42ba-a8a2-3f56dd32d611 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.019799Z digest=sha256:3fb3cf4217a80776dbcc973e4537852f5a821639fbd37ceaa39094cff1b2c35c

Observation 7cde7cc7-1d24-4b9c-aac6-7a87e4f26e5f · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

CoT-Edit: Let CoT Guide Instruction Video Editing Vbench: Comprehensive bench- mark suite for video generative models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.023233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.023233Z digest=sha256:f665e24cf8a18254ac3b8dda67e4c5cce38e0afc80a5615a414d23db17b76510

Observation 4dd9d16c-bfc1-46db-894b-b7216878607e · outbound

This paper cites Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Anyedit: Edit any knowledge encoded in language models.arXiv preprint arXiv:2502.05628, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.026401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.026401Z digest=sha256:40a3f7731c1f013141650c3a92e2f9d1c95ad6b3214fbf62571f3d56262ba60e

Observation c4b2e429-8f5b-4a6f-b212-61d7d3bcaf88 · outbound

This paper cites Vace: All-in-one video creation and editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Vace: All-in-one video creation and editing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.029525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.029525Z digest=sha256:83959f668b89521eb03123202a09b9f43772755c947efaa021a8d17f61ace5cc

Observation e2c88d85-f4ca-499b-becf-0f87bb95e566 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

CoT-Edit: Let CoT Guide Instruction Video Editing Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.032244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.032244Z digest=sha256:3ad908c7d05bb364bb37a86d31444f11ef32e2b672c0e654b3190474dd08fece

Observation 0242e969-b0d2-454e-8224-5ea3437cee65 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

CoT-Edit: Let CoT Guide Instruction Video Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.035188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.035188Z digest=sha256:35878da6546567688b482d48ccc90308725f819d29e8c6d5c9b7d0afa701ed3a

Observation 05eb5eef-d43d-4443-912a-ad663df13610 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

CoT-Edit: Let CoT Guide Instruction Video Editing AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.039061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.039061Z digest=sha256:348515fe142b62c914a1f2cc5a0e6a17a15080ea2b2624e9b9871c55fe804d50

Observation e3f22532-7f1b-41d9-85ae-d90033f338a7 · outbound

This paper cites OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation.

CoT-Edit: Let CoT Guide Instruction Video Editing OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.042340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.042340Z digest=sha256:727f2d7aebb7b29bcd6d3d04db0035d0c5c07d2908751427956ee13759d0f069

Observation 4c7772c1-8450-4621-a2fe-7362429c370a · outbound

This paper cites Grounding 3D Scene Affordance From Egocentric Interactions.

CoT-Edit: Let CoT Guide Instruction Video Editing Grounding 3D Scene Affordance From Egocentric Interactions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.046347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.046347Z digest=sha256:cded11b88de5e9f83e1d568c3f254f9a482410a65b6e1ed0ea7a82917a49712a

Observation b1fa8719-768e-4bd2-bd97-7155ad412c6f · outbound

This paper cites Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Stablev2v: Stabilizing shape consistency in video-to- video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.904108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.049901Z digest=sha256:b0c1332a0f3523e4605bd7c73a47de6b1496afe03604917c085c74a5c88bb450

Observation 409ef355-6b9c-41e2-945a-6a7dd9d41699 · outbound

This paper cites The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density.

CoT-Edit: Let CoT Guide Instruction Video Editing The health- wealth gradient in labor markets: Integrating health, in- surance, and social metrics to predict employment density

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.893023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.052062Z digest=sha256:2316fee01cc3dc68b356aa891db48876529c19aa39beeca87f222ced780abe0c

Observation 4f40e9e3-6641-4c18-a16c-9f04c1aa84a9 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

CoT-Edit: Let CoT Guide Instruction Video Editing Video-p2p: Video editing with cross-attention control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.054760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.054760Z digest=sha256:3ce45cef36bf7334df6eef9a7ae3837b244a2ac85fc9ce127183909c6d7df686

Observation 34839ddd-5a94-46ac-b8ce-f8a987f2cca6 · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

CoT-Edit: Let CoT Guide Instruction Video Editing Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.057103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.057103Z digest=sha256:175ddc86280c7bd7b8f1414dbafd01a3fd8a320405ab597ec38e0051084e22a1

Observation 5be0a624-c720-4b35-ab97-43354a0f784d · outbound

This paper cites Magic- stick: Controllable video editing via control handle transfor- mations.

CoT-Edit: Let CoT Guide Instruction Video Editing Magic- stick: Controllable video editing via control handle transfor- mations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.868611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.059802Z digest=sha256:0987f8fb36b10062191bd94becfe19a89f4853f48af5ee51aa49b7f62d9c910f

Observation 6d48c170-3a29-4bf3-98e0-67d67d71e691 · outbound

This paper cites In- structx: Towards unified visual editing with mllm guidance.

CoT-Edit: Let CoT Guide Instruction Video Editing In- structx: Towards unified visual editing with mllm guidance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.062312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.062312Z digest=sha256:e8c3d5cf8b0bfbce91a49d45a666785818256ac119057fbfdcebd077f6a5dade

Observation d2585ad4-dec7-4090-9764-6c2391d8a21d · outbound

This paper cites Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022.

CoT-Edit: Let CoT Guide Instruction Video Editing Occluded video instance segmentation: A bench- mark.International Journal of Computer Vision, 130(8): 2022–2039, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.855427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.064649Z digest=sha256:fb2cf933c9963e307405dc33897eda69d29a412967135346ba57627b0310150c

Observation 98cde1d6-e308-4324-ab13-0f0a5d78d10a · outbound

This paper cites Instructvid2vid: Controllable video editing with natural language instructions.

CoT-Edit: Let CoT Guide Instruction Video Editing Instructvid2vid: Controllable video editing with natural language instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.843912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.067106Z digest=sha256:97dfbf701c45a63c07eff25ebac070d23d684feccd2fdb3944655f86277122e7

Observation 64bf5aa5-8938-4b18-8474-807bb8e824c3 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

CoT-Edit: Let CoT Guide Instruction Video Editing Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.832657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.069273Z digest=sha256:94454f285ebfdf3c5f3ad4b9727d896a791de2080ca9a631606dea3ffcb14146

Observation 2141106f-c39d-43ae-bfa7-3cc3fee87fb2 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

CoT-Edit: Let CoT Guide Instruction Video Editing Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.071477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.071477Z digest=sha256:04ec2f27d9d517c8c1a257878c0bc74fe8311a1acd7b76b879f9e2b7e7dec354

Observation 3fcd91b3-7f33-49e9-8c2f-baff1f40219b · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

CoT-Edit: Let CoT Guide Instruction Video Editing Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.821389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.075387Z digest=sha256:7a69f07303a07549d8bb7abb7573765711c3cd5594fc70103f2e1ee6540d39c4

Observation 83fff1b4-c3a0-4f43-bf40-4a6491434b31 · outbound

This paper cites Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Omni-video: Democratizing uni- fied video understanding and generation.arXiv preprint arXiv:2507.06119, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.078800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.078800Z digest=sha256:4c406515402aea9bd61839b6c274e246930111fa164276f865b6075290b17ea3

Observation af04fc58-2e23-4c2a-bbe7-cc285382811b · outbound

This paper cites Lucy edit: Open-weight text-guided video editing, 2025.

CoT-Edit: Let CoT Guide Instruction Video Editing Lucy edit: Open-weight text-guided video editing, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.811263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.081791Z digest=sha256:66880bdbff66df011d2e48d170aef3adbaafa97f4238c8ba75bc55270575d6aa

Observation d0dc72c7-8972-4b1e-bc7e-45e9454f15ff · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CoT-Edit: Let CoT Guide Instruction Video Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.085011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.085011Z digest=sha256:23b5335f0c55ec79888ba3fa5d0db3dfa7edde55f53f7fd90b2f330107b992d5

Observation 57023dca-7e67-4dd7-8704-606aa4574292 · outbound

This paper cites Fvd: A new metric for video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Fvd: A new metric for video generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.800032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.088655Z digest=sha256:116f990d6ab23e50a60310b806526b4edb8c0febf9bc983327b1f4a889b15d59

Observation 9ab29126-2f77-4800-ae0e-3ff4fbd845f6 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

CoT-Edit: Let CoT Guide Instruction Video Editing Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.091525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.091525Z digest=sha256:73a2ac1c2d3315f1699c81f947fddcea8d065b914783e78302be7d227d36d0ce

Observation 3cb4d851-d6ff-425f-9239-27bf6f9efad8 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

CoT-Edit: Let CoT Guide Instruction Video Editing ModelScope Text-to-Video Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.096079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.096079Z digest=sha256:43bf09b08e8c630f680df98a0e0854d6a7c9d1d7159e6df0654c272b0ce381b9

Observation e663e69f-bf22-4cbb-9f2b-755e9db47ae8 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

CoT-Edit: Let CoT Guide Instruction Video Editing Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.788680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.099979Z digest=sha256:e613f62357d1e7230e42a20173080a5541ddaff6897217a7d28b6e5992c5c90a

Observation 8fb1b249-ffac-4e93-a955-c98c5a7a2aa0 · outbound

This paper cites Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration.

CoT-Edit: Let CoT Guide Instruction Video Editing Tiv-diffusion: Towards object-centric movement for text-driven image to video gen- eration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.778051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.103092Z digest=sha256:d09d4591bfb719765cce45ac5e0b7198ecfd76463f30e8530ade915a05eb8af7

Observation f013a9ff-17b8-40ca-9547-f6c71f7c7fc6 · outbound

This paper cites Re-attentional con- trollable video diffusion editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Re-attentional con- trollable video diffusion editing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.765446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.106568Z digest=sha256:f23577c2ba7c061fd570a72dd410b001211132c8f14ac46b255b0be111d182d8

Observation 43379ae1-bdc5-4091-8111-ffe361a8a947 · outbound

This paper cites Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026.

CoT-Edit: Let CoT Guide Instruction Video Editing Training-free controllable text-guided video editing.IEEE Transactions on Circuits and Systems for Video Technology, 2026

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.755803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.109921Z digest=sha256:2d7f72a8c88a4f6e74987b3865a36a1a8ac88c234fde5d304718d1de9f283308

Observation d7a7aec6-83e1-4638-adb3-63faecad4f77 · outbound

This paper cites Phrasecut: Language-based image segmen- tation in the wild.

CoT-Edit: Let CoT Guide Instruction Video Editing Phrasecut: Language-based image segmen- tation in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.745577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.112795Z digest=sha256:61a4a2089202cd49134276ca4b8631534edbc055fbab0719a68b73250135b0b9

Observation c685c6db-e0dd-4fca-b8ee-7b27984456af · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.733642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.115970Z digest=sha256:38d092681b5b74a72ffef2e7040fadbc19123b0f58e52355cc84a8b590d42df0

Observation 2f70e7ea-a7e9-4628-b819-5d9dc8fb09d1 · outbound

This paper cites Insvie-1m: Effective instruction-based video editing with elaborate dataset construction.

CoT-Edit: Let CoT Guide Instruction Video Editing Insvie-1m: Effective instruction-based video editing with elaborate dataset construction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.721670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.119760Z digest=sha256:caf7ac1c9fda33f8a7e1a59f91b599ba7069bbd24c0147b234b9b07f6c751ead

Observation 67b4273c-b3c0-4edc-8f4b-aa0bce3e75e6 · outbound

This paper cites Veg- gie: Instructional editing and reasoning video concepts with grounded generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Veg- gie: Instructional editing and reasoning video concepts with grounded generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.122798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.122798Z digest=sha256:32fafa901c756a65d64fdd05114e9da967c8a458676f36e45952463c17f9cef2

Observation 222f1b9e-0d29-495a-98a4-845715a0059e · outbound

This paper cites Editworld: Simulating world dynamics for instruction- following image editing.

CoT-Edit: Let CoT Guide Instruction Video Editing Editworld: Simulating world dynamics for instruction- following image editing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.704257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.125635Z digest=sha256:df4079d7b022e6de00b05995d3d92551344530122c1b46c12363e9a6da4fdc0b

Observation 8ba8b5ea-f2e3-400f-bd20-5da9728d9eab · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.128527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.128527Z digest=sha256:c3162c335f2abe21c6d77ae28b88c47974604fc2970c1c7d9be44341829f82df

Observation 89c0ee69-c1c6-411c-8e1c-6e5af7d15788 · outbound

This paper cites EffiVED:Efficient Video Editing via Text-instruction Diffusion Models.

CoT-Edit: Let CoT Guide Instruction Video Editing EffiVED:Efficient Video Editing via Text-instruction Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.132218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.132218Z digest=sha256:cf91d7e0957bc1a79aa1676fc8e0d441b241f27def4d6ee959919e5aab4c1214

Observation 106458a7-e319-4faf-9c0d-7a77c611b3f4 · outbound

This paper cites Motionpro: A precise mo- tion controller for image-to-video generation.

CoT-Edit: Let CoT Guide Instruction Video Editing Motionpro: A precise mo- tion controller for image-to-video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.693440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.135815Z digest=sha256:6013fbd63ca389623a86b6e0b5a17b54c9a20faca796c059de07f806423b8566

Observation 892ca003-1808-4665-9ca2-76f4c6b06c01 · outbound

This paper cites Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024.

CoT-Edit: Let CoT Guide Instruction Video Editing Ultraedit: Instruction-based fine-grained im- age editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.681969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.139111Z digest=sha256:8880ea5b23d27c86cab8549dc0d5bb7fc931142c8d91b7f5cb21adcd4502fec0

Observation 6937549a-7058-4ae0-a562-31ce580c7c73 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019.

CoT-Edit: Let CoT Guide Instruction Video Editing Semantic under- standing of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:14:36.670892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:14:36.142912Z digest=sha256:182fcad0c8cd49f1219095378e0d5abd47991785590443026c658dc0eb535ae9

Observation f2bbfc9a-9c3a-44b6-b89a-7f1cfbfb6813 · outbound

This paper cites Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists.

CoT-Edit: Let CoT Guide Instruction Video Editing Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:14:36.146697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:14:36.146697Z digest=sha256:0bb376b116a02b9c950cbac79b5472760e645dec17dbd6fc97383cd997f37032

Pith citing papers

No inbound Pith citation observations are available.