Pith. sign in

Paper Citation Record · LEDGER

Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2305.10874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10874 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:35.266101Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:34:37.702024Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 532c46f3-6ed1-43f3-8103-8712c7bd60d1 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.704069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:5c7bb33d6be534674eb7aa05db3af6c5b64ed29bfaf5d2a1cf38a1a75bf56c06

Observation c5788791-343a-4a42-b610-91b72e515481 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.266101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.266101Z digest=sha256:b3a123694456967e1e4edaf1d761fd315567f52bd5f3b7272f22b60df0fc236e

Observation d06e90c2-55dd-4d97-b791-7ce4f5559953 · inbound

TextMesh4D: Zero-shot Text-to-4D Mesh Generation cites this paper.

TextMesh4D: Zero-shot Text-to-4D Mesh Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:50.065412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:50.065412Z digest=sha256:ce7328601b6b2895931ab3b50f3daf457d4ef24b37054aebaab06c9ee91c3ba1

Observation d299757f-a3c3-43ad-86ee-fdca42af45cb · inbound

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion cites this paper.

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:18.331975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:18.331975Z digest=sha256:e7de10144a45bd093736d0557d797a93e2ae017987e3ccf2edd550d113f3ead3

Observation cfc808d4-8e26-44b4-83a7-5c2aba9e5c35 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.322880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.322880Z digest=sha256:02e96a9eefdd0a3ec11846f7632287b11f80f412b2e514d06ee1ae7a5b085fbe

Observation c1512bc3-fb6e-4207-aee3-43338d15919a · inbound

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences cites this paper.

AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:40.408942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:20:40.408942Z digest=sha256:e4c4dd8338c5d873cb71d5083a6ac2c46e9ac935b0f8647c3599ad2054b60f04

Observation 4b8f6296-d361-4a00-84b4-4ba3751f0150 · inbound

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency cites this paper.

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:30.901204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:30.901204Z digest=sha256:0b9d24eafa833d191fcb59e873a60aabe379ec38c689c4e7167a9148a4700fb5

Observation 0c99a30f-fd9e-415b-ab35-3235bac5faa3 · inbound

UNICA: A Unified Neural Framework for Controllable 3D Avatars cites this paper.

UNICA: A Unified Neural Framework for Controllable 3D Avatars Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:48:11.255188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T19:47:13.052204Z digest=sha256:25e302e14d58f51b57922591fe6aec4d622e85c7a9ae696ae9f7ba1d3b4af7f6

Observation 4a9695ad-c41e-40c6-b1a4-01a26f35a35b · inbound

Detecting AI-Generated Videos with Spiking Neural Networks cites this paper.

Detecting AI-Generated Videos with Spiking Neural Networks Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:12.328250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T16:15:42.113951Z digest=sha256:4f74775c0b08edcc689ef16342c511353a00801dfe3a7f111f7189372cdfe5b8

Observation 824d2951-fae7-4812-972e-8320785a452d · inbound

Detecting AI-Generated Videos with Spiking Neural Networks cites this paper.

Detecting AI-Generated Videos with Spiking Neural Networks Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T02:25:01.702234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:25:01.702234Z digest=sha256:738441937fb659ecbd88b163efb038e0877b2c181ae324b8c63772f59722b469

Observation 1f389f1e-0f8a-4008-899c-9103a5bf502d · inbound

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation cites this paper.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:48.054091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:23:48.054091Z digest=sha256:a0b4116c102375644be0230a90ec9af3a904bd544901d52efde6c4b1749778d1

Observation e7c544ed-a4d6-4341-b448-2587a11b1bdc · inbound

Retrieval-Driven Training-Free AI-Generated Video Attribution cites this paper.

Retrieval-Driven Training-Free AI-Generated Video Attribution Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:36:48.967450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:36:48.967450Z digest=sha256:2b719c515ea476407e8ecb2f208631eddeff22d4222d8e5ea605ce33989d3475

Observation 8c187cd9-4c89-44ba-8d5c-d7695cb93257 · inbound

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images cites this paper.

RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:20:04.584508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:20:04.584508Z digest=sha256:7ee043502d5b8005f26f839a7a3594d363ce01c4032af97bf5f611cd9108a6e0

Observation c23c8670-7780-4b06-a1a1-8747df9e688f · inbound

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection cites this paper.

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:24:41.080064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:24:41.080064Z digest=sha256:93add0f9e174a4bdfc2f6a6d07e9d904af03c5772f0280954f9f8d6622a15aa8