Pith. sign in

Paper Citation Record · LEDGER

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

As of 17 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2606.23254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.23254 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:52:23.330736Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact36
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01a54191-87f2-47e2-81d0-9c26a4ff3c39 · outbound

This paper cites Tongyi wan 2.7 video generation, 2026.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Tongyi wan 2.7 video generation, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:f5ab44f1c5970ae1fef337cf534a6b83aff644c595ed87b2dc3f8fffe26c1487

Observation f5a5541a-0e81-4b0a-b632-0aa8a524679e · outbound

This paper cites Qwen3-VL Technical Report.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.529172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:17c603e894ee05ed67adf015eb0c300935cd8c5f2045df173756df5c005393ff

Observation a8e581ee-5a31-44db-a591-0439240b8e92 · outbound

This paper cites Videopainter: Any-length video inpainting and editing with plug-and-play context control.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Videopainter: Any-length video inpainting and editing with plug-and-play context control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:4718799aaf1e22dc845cea991d5b65316b2925866eb61e8c1c02c65b775a379f

Observation e5d4207f-e998-4989-bdde-93c969933e4c · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Instructpix2pix: Learning to follow image editing instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:6b0dc72f2254187e76e1584b900e4ee87df7464629ab9847a0543d940abb07b1

Observation 82b66182-6fbf-4582-a3a4-8ca3f67792c9 · outbound

This paper cites Diffute: Universal text editing diffusion model.Advances in Neural Information Processing Systems, 36:63062–63074, 2023.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Diffute: Universal text editing diffusion model.Advances in Neural Information Processing Systems, 36:63062–63074, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:b8c1aed6c5401d13896aac8c5d6d02f28a2749c4b452619ea40e4badddeed41e

Observation c2f3816c-bbbf-4401-9a69-fe2164c75568 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36:9353–9387, 2023.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Textdiffuser: Diffusion models as text painters.Advances in Neural Information Processing Systems, 36:9353–9387, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:cabaaf74e73b1eb6ab2674b9f6cb06ca1dd32982ca45e8c6c2ea9eededbd9a0e

Observation 5a3f1490-c515-4cde-983d-12e21b079ab6 · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:e3483190fe3337334bea669bc0cdef9db6a096ace88760cbbfe23ab7c2fc2db6

Observation a370b689-cc2a-4335-adfc-9adb30f92714 · outbound

This paper cites Consistent Video-to-Video Transfer Using Synthetic Dataset.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Consistent Video-to-Video Transfer Using Synthetic Dataset

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.591005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:84b341c5883b2ae2d1ceaa4844cfaeae1624b99badc1a53db6ff11ea82e58bb1

Observation b9a5b256-51e6-46db-8dd6-724c0ae8b81b · outbound

This paper cites Viva: Vlm-guided instruction-based video editing with reward optimization.arXiv preprint arXiv:2512.16906, 2025.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Viva: Vlm-guided instruction-based video editing with reward optimization.arXiv preprint arXiv:2512.16906, 2025

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.571018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:36f36b39f9adb1df918ea379cbd104eacd388a2ac260399b8ca7445faff62e2e

Observation 3d47ce53-766d-4b8f-9f9d-9269ded40262 · outbound

This paper cites FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.582861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:43cf4852be0af856a10f8b7b2fb54ae8b6bc17898c6db5fc9768a48b19ea46ed

Observation 6caf3afb-c236-4023-b898-9f1f0f0739d5 · outbound

This paper cites Peak signal-to-noise ratio, 2026.https://en.wikipedia.org/w/index.php?title=Peak_ signal-to-noise_ratio&oldid=1210897995.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Peak signal-to-noise ratio, 2026.https://en.wikipedia.org/w/index.php?title=Peak_ signal-to-noise_ratio&oldid=1210897995

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:9c662fe15102e25ac7e64eb3d24842ae25c8311145ecae409b700a882aaa67d0

Observation 45e57e04-289b-4fac-a70c-cfd1cf28c8be · outbound

This paper cites PaddleOCR 3.0 Technical Report.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control PaddleOCR 3.0 Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.514082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:963ecee06a93acd71d75db9ea097a500405d2f8590c519a700002dc0f977bfeb

Observation ca51d50a-9086-430e-942b-7a7ee93c3a91 · outbound

This paper cites Diffusion models beat gans on image synthesis.Advances in neural information processing systems, 34:8780–8794, 2021.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Diffusion models beat gans on image synthesis.Advances in neural information processing systems, 34:8780–8794, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:4bf3c32b3f2d4612b6e1dec942ed7f76c7da08db338316338851efafc386fb3f

Observation 02321dd6-8e9f-4c76-8cda-fe9205f194eb · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Scaling rectified flow transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:7d14011db438fe754503499ff57684bae9186040fec4e7d457181e39974f268e

Observation 39976ede-3408-4917-bff5-aac230d806a9 · outbound

This paper cites Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.508863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:ff5d2bc890d90bf574c9516afa0db4bce17bef10f8beff1acd33f0cc9d657c86

Observation a4d07c76-0f05-4796-9be8-13b9875d467c · outbound

This paper cites google-10000-english, 2026.https://github.com/first20hours/google-10000-english.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control google-10000-english, 2026.https://github.com/first20hours/google-10000-english

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:a189d09a320221b2fc9ef683dd9cdd14f727a05884b66dce4efb9eb0050a486c

Observation 80c54c28-66af-4bd3-8eaa-12ea81f79a36 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.550528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:94f3fee3789a9eba61e9c6850471d022e52234481a046956370545501efd40c6

Observation db5d6a78-ceb2-4b5c-82b7-202c0b33c627 · outbound

This paper cites Google fonts, 2026.https://fonts.google.com/.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Google fonts, 2026.https://fonts.google.com/

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:b204d480c7c100d6ab0c4b410b9bbff778be7f5f43c2582f88f323140af1b281

Observation 94b77fb0-d500-4aaa-a4de-81ea45ae305c · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:efe738d96ccb91a21468abe4a157eeb8ce744baf5c6ce5e5092735fa286a65c8

Observation 163fca16-bd3a-477a-9883-eef26dc9e62a · outbound

This paper cites Videoswap: Customized video subject swapping with interactive semantic point correspondence.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Videoswap: Customized video subject swapping with interactive semantic point correspondence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:2e696229e3ac4fcae43fdbe741a5d69480f542b089ce9cb2ccb9819a47f77ec2

Observation 99feef22-cde9-4aa7-a86e-3b2fdc8fffa0 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control LTX-Video: Realtime Video Latent Diffusion

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.562273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:8b869cab95b17838e408e60fcb2375faa6e127c112907feecc947bc40818ec1f

Observation ab0ee45a-30a5-4170-b466-36597b38f5b4 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:e3c22ff5e9d041daeceaf94918c40481f21ef9f390d6c9840c854e917f2566b7

Observation fea129e5-139b-43c7-a561-a2221eecddd0 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:64f905fb9e48c1f1505d2fb45d49e6262e52e1dfc2e82ed6d3d91a4305b00d65

Observation 4c66eb26-f31d-41c2-8a29-273ca8836406 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:22bd0be6fae0cf80a71a132d697d11456f1feed69b19940835114ea8f31fadff

Observation de744b06-fdfd-4b1f-a811-e9516906f7d2 · outbound

This paper cites VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.520621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:1a5c428c320c6b8dad2e0d7578cb4c6b70aa19fff442ebb947f1012487621750

Observation 31992bb2-30c1-4363-9861-da234462f905 · outbound

This paper cites Vace: All-in-one video creation and editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Vace: All-in-one video creation and editing

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:803956bfcf1630a8942921901c3434b6d12204383ab1c46dd91474925bef2675

Observation 716a80d8-ade9-4c87-bf29-4d1ecf04c5f2 · outbound

This paper cites Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Rave: Randomized noise shuffling for fast and consistent video editing with diffusion models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:0ec7549b089a7c36e26f6c60c9d36177b2d8470f5a8f4b5b1423c08a1f751da6

Observation 970cdce0-79e0-4a2f-a309-b764aa4db3f2 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.576283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:170e9dbe44ca5cc37dbc8d20b03a35fdcbe3b88b06beb10f34691963fe61c66e

Observation e60b96ff-866e-42f0-916a-f4d686669f78 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.543311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:e29055fe5b1d22e71f40717146597053dfaebf90063e61c75381366a84eff6fb

Observation 7b3a8a59-ee5f-4978-a0db-42ba31de23fc · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.559581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:4defe6f18b0e4428d79f11cfa04a0298ef9fe816c5460ab9be9e643e5d07069b

Observation 0e6450ec-c234-46db-8d4e-270342e6a960 · outbound

This paper cites Flux-text: A simple and advanced diffusion transformer baseline for scene text editing.arXiv preprint arXiv:2505.03329.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Flux-text: A simple and advanced diffusion transformer baseline for scene text editing.arXiv preprint arXiv:2505.03329

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.568287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:75162b17a54b3e9a969480f21047ef03f85adfed77d3fb21341f1915fadf5e58

Observation a8b7dad8-42ba-41a3-be69-491d225708ec · outbound

This paper cites Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.573737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:56f715c547843ef8a8fc58e900ac8299f8da69369a1f64640da6e26cd1e7cc62

Observation 162c3efd-3b47-4d16-96c2-bc468a422520 · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Vidtome: Video token merging for zero-shot video editing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:b6fa16c27649ac483c829fbf7b6c0878f986a2eb82b4064d467642a59b56947f

Observation 4075c8c0-c139-4129-abc9-ffa784b4d233 · outbound

This paper cites arXiv preprint arXiv:2510.14648 , year=.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control arXiv preprint arXiv:2510.14648 , year=

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.537726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:543d611c488764aa6c40e931055dd5181273e64a2ddb47a9641e11d2c70c40fe

Observation fd2dd144-f5ed-48d7-bd30-c705d67c59cc · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Open-Sora Plan: Open-Source Large Video Generation Model

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.579365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:e66cd7d4be7b5ee2ebb6e4a50b0920fcad01063fe8035033adffc71b4a051564

Observation e343b75a-e0c2-49da-a78a-8ea2aaf92f89 · outbound

This paper cites Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.551320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:c818e2ce5123f9ad37dbd91cf9dda85b8ba7b392b1bf449322858f28a8a4aa85

Observation 5cb9e119-0ec8-44d8-9d43-50f9c9741cda · outbound

This paper cites Flow Matching for Generative Modeling.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Flow Matching for Generative Modeling

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.544793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:4d82176c2b283cec29b4b60376007f39b1673554add1737f548fb856d8276ed8

Observation 73f69ce3-4f74-4156-976c-604ca9272cee · outbound

This paper cites Generative Video Propagation.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Generative Video Propagation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.517233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:97b3ba9b9f107e0dc2bf09f784285dc87472aaa3a6f74c4c5b550f8abf802c68

Observation 1594abdc-5ce4-4deb-be71-76803918dd70 · outbound

This paper cites Glyph-byt5: A customized text encoder for accurate visual text rendering.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Glyph-byt5: A customized text encoder for accurate visual text rendering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:d5830ecc8ff2910bf503592c248d68173c7d6aa95225a6784cfc67e7758f595f

Observation 0849a903-e6b9-4a6e-97c2-802de07a1f7f · outbound

This paper cites GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.565365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:015bd09c52c9ce0edf5177a212c393969e89e7de117315f24d5f3482c72ff945

Observation 1ae076a6-6e3e-450f-b706-8be6613d2342 · outbound

This paper cites UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.552995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:ac71ce276ceb123675e2d8a8f1aea0def9a27b71cfaa9b9a13b5935587227b9a

Observation a22746e4-c98e-4943-bd37-6bc5e83c4dd0 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control DINOv2: Learning Robust Visual Features without Supervision

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.562508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:ddde3d1b329b2a5403cc5fcf6bcc9b45cdfe8676c3604c5b1b1eb9ee20bc5da6

Observation fd7a2065-072e-4816-9843-95b573d6226c · outbound

This paper cites Codef: Content deformation fields for temporally consistent video processing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Codef: Content deformation fields for temporally consistent video processing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:fa9673e50da3b6ef62d60fe993a430db95484034e97198fc2dbb48bcaf693f6b

Observation b022a04b-53e8-4314-88b8-90b5af90a866 · outbound

This paper cites I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.564974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:4ca1ba23769101bc713177edfc8158afb24e0f90ca05a51a1865d9c6774f128a

Observation 131c6acd-c94f-48d5-9604-c33d349abd47 · outbound

This paper cites Scalable diffusion models with transformers.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Scalable diffusion models with transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:909410bd65e8731a4200852173d034b0c4806433a8374dc4d9544641b5523709

Observation 6c217845-113e-4cb8-9587-02cd7f631cef · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Fatezero: Fusing attentions for zero-shot text-based video editing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:3f2aa412ad55615b688831d96e0d74b4506fe5620ab46206ae6eadc2b76041b7

Observation 039162f6-3535-407c-8bcc-656a6148cd0e · outbound

This paper cites Learning transferable visual models from natural language supervision.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Learning transferable visual models from natural language supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:fcbe8427f916e854a192c7a6cabddc9ca0a5ceefb732e379d95aa71b3e52512c

Observation 3fee72c8-64aa-44de-ac2f-4b3dc5c6edac · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control High-resolution image synthesis with latent diffusion models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:629946a487f38d9b782f9f5d94c48e1882badce1e783fcd73635326317a803a0

Observation 1a4526fa-525f-4115-94b2-ca259bb71f8b · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Seedance 2.0: Advancing Video Generation for World Complexity

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.587961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:6c960bb2d58e7cc1dc336f4f1684108c17914551cd33e3e0e3ea092ee58afe56

Observation 4f39e394-0a6d-4586-881d-283783523626 · outbound

This paper cites Stellar: Scene text editor for low-resource languages and real-world data.arXiv preprint arXiv:2511.09977, 2025.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Stellar: Scene text editor for low-resource languages and real-world data.arXiv preprint arXiv:2511.09977, 2025

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.573176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:88ec0265047943ec050d1554fdee8dd550d8ca9d7286b636990a9fda93d17b92

Observation 2e0780c9-28ef-4057-8ba1-7b10bf1bd489 · outbound

This paper cites Fonts: Text rendering with typography and style controls.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Fonts: Text rendering with typography and style controls

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:046ca536158e3ded2d704bcb59d455fa5ff90c238346f1007a86de81df4f0c9d

Observation 46f3b731-7b71-4bdc-ae8c-38b3072f35ca · outbound

This paper cites Denoising Diffusion Implicit Models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Denoising Diffusion Implicit Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.511476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:598322fabb8a54cf957322fca811c3d281bddc345204bd873c1ae290b6d866b4

Observation 849f1ebf-63aa-41a4-8f95-888c08d12ac4 · outbound

This paper cites DoPI: Doctor-like Proactive Interrogation LLM for Traditional Chinese Medicine.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control DoPI: Doctor-like Proactive Interrogation LLM for Traditional Chinese Medicine

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.542349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:839d0718467d4c05220697e8f79f54c0ebf1bf4d99f1ebf8043c409d5865be57

Observation 6284fcb0-9c0c-422d-8eb8-26dfa3c47c5c · outbound

This paper cites Omni-video: Democra- tizing unified video understanding and generation.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Omni-video: Democra- tizing unified video understanding and generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.548965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:bf9150bf1ce4e47f01c4f760891d12aace68033c33bf6f10ca04f779d3e14ba5

Observation 4cee4815-1fb5-4fb6-b3de-2dc13e548267 · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control AnyText: Multilingual Visual Text Generation And Editing

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.520381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:9353c95149c643c5c23884cf82532f4480a62e577b67ef2d944898f0d9f68ddf

Observation 1ab31c13-c9fa-461f-9c30-e80c086922b3 · outbound

This paper cites AnyText2: Visual Text Generation and Editing With Customizable Attributes.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control AnyText2: Visual Text Generation and Editing With Customizable Attributes

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.559435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:7ec26142fa1a45cc9a30a23798834a69ddcce37286b6e845d86fba019d19f6c3

Observation 38f2a444-445c-4b0b-96c6-7d8f2e29de11 · outbound

This paper cites Fvd: A new metric for video generation.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Fvd: A new metric for video generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:d2b650fc7b14bf8607c115d082e015033f45ac02ed76a6765d602a25c0e2ddc9

Observation 4d5063b6-a8da-498c-9713-a811437e9612 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Wan: Open and Advanced Large-Scale Video Generative Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.536395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:ce87e21f43c9d0c93ff986f7dd7556a871fae1824fd1a5368a00c65a5ad447e7

Observation 34ca7572-60dd-4cad-93dc-509e429f7a61 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:102b658682b7232df131787e2b8ba9cd7124ff63ce4dd45e11b96444138d309c

Observation 6a1d2a9e-063f-4382-b861-e68e9b452ec6 · outbound

This paper cites UniVideo: Unified Understanding, Generation, and Editing for Videos.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:d92291982e4ba2c759a4cc2f1a7d2d92cb96f42b9e10ae021ebb048609ae07f6

Observation e57e41c7-7de8-4d33-8edc-026a552eeb0c · outbound

This paper cites Qwen-Image Technical Report.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Qwen-Image Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.585396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:1a7b1617114d6869ddf6aadd322ffa572060d68ff9f98c600481c79623c1556f

Observation 8c156b7a-7c66-4bc6-86a2-d349d631273b · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:5179ea4476aac89387f2ab0b20644d0b0eee6f13f60122b4605ddbc2185ad223

Observation c3ac5148-4af1-40b7-9969-ed7c7e6b0484 · outbound

This paper cites A Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control A Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.535026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:9699f411d79539d4d14b3d1424163dfb7889972220c0a7727d81bc23ca769047

Observation b4ff0b57-e9ac-4aac-8189-fd97169f3b9d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:29:44.556552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:58ad2fa29ea83887404ef05d2f7e6574d5049504032ceed3f46e9787b7396c43

Observation f3df1a79-2162-49c6-b5d4-cee49bec0530 · outbound

This paper cites Textctrl: Diffusion-based scene text editing with prior guidance control.Advances in Neural Information Processing Systems, 37:138569–138594, 2024.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Textctrl: Diffusion-based scene text editing with prior guidance control.Advances in Neural Information Processing Systems, 37:138569–138594, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:1b9f5e62edfe38e426217d816dc44962fd1a1653202767d7e49e328bc5dcad0f

Observation 9f357f31-8b22-4086-aa85-ed00d312a9c2 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Adding conditional control to text-to-image diffusion models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:d410254e0840ea53981153e7baa59f43d3d1023771ec9ad7523b524d7a533d45

Observation cdd7e922-f6c2-4ff2-a9e4-da55d3f08ec6 · outbound

This paper cites EffiVED:Efficient Video Editing via Text-instruction Diffusion Models.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control EffiVED:Efficient Video Editing via Text-instruction Diffusion Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.570350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:f7d5bc5e919bc0f110b6205de3287aebec7c2739d97be98e8d2da67150188cac

Observation ffe46721-f12e-41a3-9d14-18c55bbd7460 · outbound

This paper cites Utdesign: A unified framework for stylized text editing and generation in graphic design images.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Utdesign: A unified framework for stylized text editing and generation in graphic design images

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:905467c54fd8fd0c33beab003fd3da16e251e49f5c9b61106b6dfc4bb786cfa4

Observation bc22fe16-ddbd-48d5-ba36-3cdbac3751a8 · outbound

This paper cites Describe in detail the typography, color, style, text material, and rendering effects of the text regions {source text} in this image.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control Describe in detail the typography, color, style, text material, and rendering effects of the text regions {source text} in this image

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-26T08:52:23.330736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:33373866b464a3cec073e8993da07092a820c84f6df5c0c8725021347b4a34c3

Pith citing papers

No inbound Pith citation observations are available.