Pith. sign in

Paper Citation Record · LEDGER

CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.10391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10391 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:29.343261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.682483Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7344fd96-2a37-4a56-9bff-dc864b12df8a · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.343261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.343261Z digest=sha256:acf66465529e0ca9112331bec15593755ba2c64611ab47990f2bfb0d4264ad07

Observation 10bbdf14-7a25-4c48-bb27-3789b2a20813 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:08.227600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:08.227600Z digest=sha256:61ce74c0de7ba6b76d15cd5ce470e07891e0a90eba77ebaf24148799bfcddb77

Observation aa4e6714-16c5-4dc3-9c33-4edf724954cd · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:36.010543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:36.010543Z digest=sha256:eb542c86180c081427da24ae8e69493515bff9558a63fe2cb2bf87c5d8d6c56a

Observation 3e6fad62-7218-485c-ae16-187812f6a501 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:08.256451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:08.256451Z digest=sha256:cefb818fbfdac994467155840ec0de0c3665d930c7623edd543f0d71a0aa7159

Observation 6a475883-7016-4dfc-a2ec-e9afb0e3bc01 · inbound

HOComp: Interaction-Aware Human-Object Composition cites this paper.

HOComp: Interaction-Aware Human-Object Composition CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:09:51.852399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:09:51.852399Z digest=sha256:ee14d4e50191af17d89c01fa1c12f9e213d232ae00c971a98746fb3e5e28e2ea

Observation 6b52e863-a64e-45b6-8a0b-47dda76c78a3 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:0ee97f9f21d5689e198247fd05ac9f0bbe35e267e2ba9580a0628e80913ee617

Observation a6452147-1bd9-46b1-b879-ee5b15d83f5b · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.912925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:aade5bc3eb5f3fc270b1fc561fec90e64cca5ea13ddb75811faead69921c7ae8

Observation 54963091-81c4-4cea-8fc1-ebc352845893 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.977015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:e93d68c7710220fb37b884481db61813f8c1eb3fa42b928faa1abd381e2b396f

Observation b80fc3bc-d5c4-4034-9335-58cd8047c016 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.343949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:1d9de33c652d56bea2dae519eae36eea0f3353d4215d3330683cf1d40a7a1744

Observation 38fef8f3-95d7-4a33-bf2c-af5c51c038ba · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.976977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:d09ec29109ed8179aba9dc04824c28744bfc5d34bf92c675151ee9c701a9fefa

Observation 3d29057a-593a-41ea-925b-47a9414aae75 · inbound

Customizing Video Portraits via Identity-ActionDecoupling cites this paper.

Customizing Video Portraits via Identity-ActionDecoupling CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.380535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T10:56:11.553110Z digest=sha256:aac432768491e6e0b7e8d6ecdb609bd45c62a77e10b348d277a0106e03386ab2

Observation dcce61f5-bf18-4c2d-9ad4-326d741f59bf · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.684393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:6b190335bdf7ffff53be2dad9d795d60bbe740973a4b659433ac1c099b15f28a

Observation 75e11631-dd5b-46f4-bfa7-df8687893a99 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:dea8c15bd4a066a074f2f6089efdffef57b659c859efdbb56bab44da49e5d25f

Observation 4a50d647-fe85-45be-b497-2b84e4849794 · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:19.025878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:19.025878Z digest=sha256:4e80f27bf8514e7c8c0b2430f860be6f230c510fa3b6726479c310a52769969d

Observation c2ea64d7-54ec-4967-92ac-74862def1255 · inbound

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry cites this paper.

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:04.839943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:04.839943Z digest=sha256:e3b60d9f784bf88ec86018ab464588aac2b9914d099d4877535d76131e92ed9d