Pith. sign in

Paper Citation Record · LEDGER

Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2305.10874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10874 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:35.266101Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:34:37.702024Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 532c46f3-6ed1-43f3-8103-8712c7bd60d1 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.704069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:7183e99766104d68406e316404f21b8f9700bf066c0c6d5a73febdaaf400dc04

Observation c5788791-343a-4a42-b610-91b72e515481 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.266101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.266101Z digest=sha256:40407b570fa3d915d8704952cd7bde0183e4693cb648e10bde7e4baf650d3bef

Observation d06e90c2-55dd-4d97-b791-7ce4f5559953 · inbound

TextMesh4D: Zero-shot Text-to-4D Mesh Generation cites this paper.

TextMesh4D: Zero-shot Text-to-4D Mesh Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:50.065412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:50.065412Z digest=sha256:65f85447018dc580acf3271766895b8274a696d0df5c8fa14587155cc1262f83

Observation d299757f-a3c3-43ad-86ee-fdca42af45cb · inbound

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion cites this paper.

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:18.331975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:18.331975Z digest=sha256:f7c212932507ffff2f683065b90753958ab078fbcc037f9275a09e4163e3b25c

Observation cfc808d4-8e26-44b4-83a7-5c2aba9e5c35 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.322880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.322880Z digest=sha256:8b4848989447e2a6cddbc22a6b33e964336b07169a53330d3f891c555429d91a

Observation c1512bc3-fb6e-4207-aee3-43338d15919a · inbound

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences cites this paper.

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:40.408942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:20:40.408942Z digest=sha256:14a26b0d8b053bcdb813c14128bdeddc91bcc5f971ccc0d9e01e8864d5bdb35c

Observation 4b8f6296-d361-4a00-84b4-4ba3751f0150 · inbound

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency cites this paper.

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:30.901204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:30.901204Z digest=sha256:57cd352892a2fa4662f8a3ce229cde93380851222ded8c02bf70e2562c7414e2

Observation 0c99a30f-fd9e-415b-ab35-3235bac5faa3 · inbound

UNICA: A Unified Neural Framework for Controllable 3D Avatars cites this paper.

UNICA: A Unified Neural Framework for Controllable 3D Avatars Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:48:11.255188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:47:13.052204Z digest=sha256:4771f7179ba12faff4e7799380bd96e078757d42b473a26ad8227f91ab484a7c

Observation 4a9695ad-c41e-40c6-b1a4-01a26f35a35b · inbound

Detecting AI-Generated Videos with Spiking Neural Networks cites this paper.

Detecting AI-Generated Videos with Spiking Neural Networks Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:12.328250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T16:15:42.113951Z digest=sha256:25b0176bd97aeef461a72a7c60bfcbb6b1a90cac944be3c25573ad04396d2fae

Observation 824d2951-fae7-4812-972e-8320785a452d · inbound

Detecting AI-Generated Videos with Spiking Neural Networks cites this paper.

Detecting AI-Generated Videos with Spiking Neural Networks Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T02:25:01.702234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:25:01.702234Z digest=sha256:a6c9a25ce314cbbbaf35130765e8f6297c48c50e0a9161298ca253e01a288d51

Observation 1f389f1e-0f8a-4008-899c-9103a5bf502d · inbound

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation cites this paper.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:48.054091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:23:48.054091Z digest=sha256:4ceaa2b7f123d06f87ce3878f54248ac893988a39ea64586abe139743593d314

Observation e7c544ed-a4d6-4341-b448-2587a11b1bdc · inbound

Retrieval-Driven Training-Free AI-Generated Video Attribution cites this paper.

Retrieval-Driven Training-Free AI-Generated Video Attribution Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:36:48.967450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:36:48.967450Z digest=sha256:7fb5c0ce35188c08ec14b6a9817d2bbe83937d570d63498113a667657ad49e10

Observation 8c187cd9-4c89-44ba-8d5c-d7695cb93257 · inbound

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images cites this paper.

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:20:04.584508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:20:04.584508Z digest=sha256:508dcac840625e552f8e54f2522e1773d6042b166ae9e2f95546f74334de0e02

Observation c23c8670-7780-4b06-a1a1-8747df9e688f · inbound

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection cites this paper.

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:24:41.080064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:24:41.080064Z digest=sha256:88a22103eacc92f7393a5ae82e1da1f2a7f3b4b974f869276be69934371a7742