Pith. sign in

Paper Citation Record · LEDGER

CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.10391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10391 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:29.343261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.682483Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7344fd96-2a37-4a56-9bff-dc864b12df8a · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.343261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.343261Z digest=sha256:faf790b2ea6948d1701d60901d5ebae52a5593095ddd8cf63938536a769189ef

Observation 10bbdf14-7a25-4c48-bb27-3789b2a20813 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:08.227600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:08.227600Z digest=sha256:ec74314f1f7f681c36ffa20492cee9ffbc12c4b3acfac62f9c418624cc5b8d74

Observation aa4e6714-16c5-4dc3-9c33-4edf724954cd · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:36.010543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:36.010543Z digest=sha256:5f6c9ca6d34207d19b7cb9e70319857b1ee6af6f4d62be496981b8405b06ad6c

Observation 3e6fad62-7218-485c-ae16-187812f6a501 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:08.256451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:08.256451Z digest=sha256:f4b1e9baef3756344fd7037d2b7246afa2e23d034c73433e86a93d969e6fa6c4

Observation 6a475883-7016-4dfc-a2ec-e9afb0e3bc01 · inbound

HOComp: Interaction-Aware Human-Object Composition cites this paper.

HOComp: Interaction-Aware Human-Object Composition CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:09:51.852399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:09:51.852399Z digest=sha256:e6612dac0748bbdc9d805d9e3d888e128facc466c33b49ce80aa89e7292b1d99

Observation 6b52e863-a64e-45b6-8a0b-47dda76c78a3 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:b495934add35c96991d066b84d2610060b2244f1c954b112b84e90cc643a0554

Observation a6452147-1bd9-46b1-b879-ee5b15d83f5b · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.912925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:226b9915ee79bbd8ae0a3939b78516594a72f412ebf8d55771b7e30250e61b87

Observation 54963091-81c4-4cea-8fc1-ebc352845893 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.977015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:6006c855eb694b91acbfec46f4fde86a676a29ca21aa720fd41e6d83f6994019

Observation b80fc3bc-d5c4-4034-9335-58cd8047c016 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.343949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:354080054bd8ea4c5615bc39abd0f950e6c04d297918be4300712ce57c5514cb

Observation 38fef8f3-95d7-4a33-bf2c-af5c51c038ba · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.976977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:0b351ab9603e3a7582154dd048e853f3788ca369c68eabd428385ad7fc9dd723

Observation 3d29057a-593a-41ea-925b-47a9414aae75 · inbound

Customizing Video Portraits via Identity-ActionDecoupling cites this paper.

Customizing Video Portraits via Identity-ActionDecoupling CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.380535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T10:56:11.553110Z digest=sha256:90e7cac5cf766b67aa461fbf3a02d7603dff2bab3bd940c1d457e21aa02f221d

Observation dcce61f5-bf18-4c2d-9ad4-326d741f59bf · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.684393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:b0f6340d5f0a4103dd74ac7c736d57fc01843f64b3da38dfd0ae3f39c5fc58a4

Observation 75e11631-dd5b-46f4-bfa7-df8687893a99 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:42c5d470af6bd2842383e3c61fe0d7c05c23c03d66f318f724b8e74f8a8c14f3

Observation 4a50d647-fe85-45be-b497-2b84e4849794 · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:19.025878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:19.025878Z digest=sha256:f633252a0f9eb5f6d68652e06b77f69b1acbf75387e19acc68ac2f4a6432895e

Observation c2ea64d7-54ec-4967-92ac-74862def1255 · inbound

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry cites this paper.

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:04.839943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:04.839943Z digest=sha256:a877d1c4b8cdcc11f911226b49fd612cc120c727efdb7309b7f75aff26c278ae