Pith. sign in

Paper Citation Record · LEDGER

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

As of 14 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 2 inbound Pith citation observations for arXiv:2412.04440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04440 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:28:21.301097Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T05:27:16.731285Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:28:52.699573Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact3
  • verified fuzzy38
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aeab1852-7b01-4467-9d91-286f901f7c31 · outbound

This paper cites https : / / huggingface.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration https : / / huggingface

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:24.048844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.831324Z digest=sha256:fa921fbf774eef61721775d6becc6d0031135bcbc95931714fbdeeccd117d598

Observation 1e572bb3-01f3-4908-bbc1-bb97e7630d3f · outbound

This paper cites https://pika.art/, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration https://pika.art/, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.993015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.837758Z digest=sha256:3c4f610b3082a2eb0e88006285457a9c010957839418e05b3ca1ea015a835984

Observation 00949303-f71d-41d2-ab55-2af4b7c42bd2 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.845667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.845667Z digest=sha256:26183c3adcaa9ac46a514e7b8fdb66093d3182fef038c8d9398122e65523002b

Observation 9f78e4b8-de1c-4845-b63f-431a7f652e93 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.853812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.853812Z digest=sha256:135a5683084719f15c2f8a396a36dd1eda4824370739d0df2a8f7d68c74530bb

Observation e99f32f2-07b0-4a7f-a178-e43bb1e08591 · outbound

This paper cites Maskgit: Masked generative image transformer.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Maskgit: Masked generative image transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.860177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.860177Z digest=sha256:8e941439b0f4b2a2decbcd82dbc31d2c07b942f9f98a9c130a0c0996bf7a6e2e

Observation 29928a2a-1cda-4304-9885-ba0406f0924a · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.865826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.865826Z digest=sha256:28eb2244404a9abf28232205b1185726ed85ebc9b98f0bd2449f73c8cd846cdb

Observation 6138a6df-d913-445a-8d82-ac3d5e9ba11a · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.806787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.875964Z digest=sha256:f9333361861b630b98f50835704c78f80100101e842ba1ea43e374dfd8041521

Observation 8e0ca905-fd28-49c0-bb89-10eb207807d1 · outbound

This paper cites VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.884384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.884384Z digest=sha256:9900d1b53d3415980d34620817dde272aa01a832d416ba0ef30515172454d3c6

Observation 066dbf94-aa40-489c-b320-9e17afd41d7c · outbound

This paper cites Training- free layout control with cross-attention guidance.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Training- free layout control with cross-attention guidance

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.771600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.890486Z digest=sha256:8258ba6dd626d2a6af663dbe699f98a8d2222f47d0ced96f0d3e0f93a06d4240

Observation 312f5fcb-2635-4e50-a77f-746ae2dc346d · outbound

This paper cites Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.896387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.896387Z digest=sha256:7ac7255756a41fa56c15ff88f41d32701c323689b6bbb62828f66a714235d6b4

Observation 669d1303-adaa-4124-9bc6-408fcf5d7652 · outbound

This paper cites Mind2web: Towards a generalist agent for the web, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Mind2web: Towards a generalist agent for the web, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.747093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.902905Z digest=sha256:1febb034ad39d9422d74faa83f9b561b252390d0a12b4a3f19955496d4804f7b

Observation 329da210-fafc-444c-8bf0-1b790c8f685b · outbound

This paper cites an unresolved cited work.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:28:23.724429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.909328Z digest=sha256:243e3341e0b43cac11dadaac5f009e8c70c0632b5bfbddaf49f207793f136c99

Observation 5652a914-6f6f-41ae-a5a1-5ace08a38c5d · outbound

This paper cites Training-free structured diffu- sion guidance for compositional text-to-image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Training-free structured diffu- sion guidance for compositional text-to-image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.690044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.916649Z digest=sha256:31d752c48a974966776994a7b4e7bb03e5221165fc43bc759f7ae255641ecae5

Observation 26858460-4a3b-4e46-ae38-82b138f5d128 · outbound

This paper cites LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.922439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.922439Z digest=sha256:443b0b3420807b29d5d3f42b16caa9ff93b52696fc8adef0deabf10723afc7e8

Observation 9b2ac166-c2b6-40db-acfd-a90c8f860000 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.928645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.928645Z digest=sha256:bd067a4e7bbee8be94afebe52185eda3cd744140b1817f6248604d6de399f204

Observation a0ab5767-075c-456e-a619-15566b85d415 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.934639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.934639Z digest=sha256:9bbb7e4ac774292f04a6a90095f513c902dabbeab656782c773f7f9d604769af

Observation 04a428a9-d5fa-4965-ade5-0c35fd06b96b · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Denoising diffu- sion probabilistic models, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.941001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.941001Z digest=sha256:68f2f105bb7425a448c63719187787ee496d17478396d10e7c3e4a6bd741b286

Observation 3e93b76c-5144-466c-a344-6fe5aae94940 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Imagen Video: High Definition Video Generation with Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.946653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.946653Z digest=sha256:dec5816288bde4aa94637d32e75485f5ad4eb6f8636fc16b94e783ec12298912

Observation 71828357-ae88-4d65-9c55-caf33b423133 · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Metagpt: Meta programming for a multi-agent collaborative framework, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.636186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.953139Z digest=sha256:6ef11e321345daeb0202475b41316324ad46f6c6b15353d4f9aa06a482d51c51

Observation cf084e4a-06a1-4d8f-8d7d-3607b8e36167 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.960029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.960029Z digest=sha256:46514c41b0a0e70ce67cfc7afbb59f12d04829e227d6c43c2fc078a859726896

Observation 42b64119-d7f0-43f5-a8ec-c9bb6bbda202 · outbound

This paper cites Open-sora: Democratizing efficient video pro- duction for all, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Open-sora: Democratizing efficient video pro- duction for all, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.610098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.965574Z digest=sha256:9cc7025d45dfbc2c0a9464f343b590df16283f017f922590a2e5c74e7478f574

Observation 3aa475d1-7bf5-41d1-ba44-0ea23c3a5db5 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.576063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.972538Z digest=sha256:41c30dc6c99f9558541ef78005aa1d06e0437a94e6fce578484bd75974031625

Observation e612d41b-dd4a-4e01-9bdb-366cb2138164 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.978399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.978399Z digest=sha256:91c33d5f180de58f00095f62e87d4264297b668615dfc9a6d48c95757377a92e

Observation 576d222d-0cec-4343-8aa9-8dc60fcb3a59 · outbound

This paper cites Dense text-to-image generation with attention modulation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Dense text-to-image generation with attention modulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.521624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.984086Z digest=sha256:89984a393381f623095e9d8adaf07212bef6211670264e9fb69139eb603809e0

Observation 72995c45-1d99-47d1-98ca-5be64aecafe3 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.989614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.989614Z digest=sha256:9329d54daf438bf408a2adadb1730498f92a710c786be9dd2e56ae69abf8cabb

Observation 6439a76a-ddde-400c-932f-87ea3dd8f53a · outbound

This paper cites Open-sora-plan, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Open-sora-plan, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.498187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:20.995116Z digest=sha256:c577ea01cb750a220c2dd6193d9078d33c7021d93ec7b3fc071559a4376b3a2b

Observation 136af769-2592-4035-bf3f-b883f43945ac · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Gligen: Open-set grounded text-to-image generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.475878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.000489Z digest=sha256:704bf94bbc5c00a8724c79ac28093a84503212f23a3946f993299a3d9d780e30

Observation 94a1fafa-c26a-42a6-95dc-c6aba4f41d32 · outbound

This paper cites Stylet2i: Toward compositional and high-fidelity text- to-image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Stylet2i: Toward compositional and high-fidelity text- to-image synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.453249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.006516Z digest=sha256:4e7557cba188833ef3b1074a6477e7f7597468fd75e673c453c74a5a48998363

Observation 6521fa0a-33c7-4ef2-8f1b-009026c3036c · outbound

This paper cites LLM-grounded Video Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration LLM-grounded Video Diffusion Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.012192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.012192Z digest=sha256:a6c3c1b202bb2469f075509d52a950789433e84699bc62644332b77992a99913

Observation cce2e134-ac28-47d5-a1cf-1bb3d2984ef9 · outbound

This paper cites Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.433787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.017820Z digest=sha256:1e038de00f1a8061ce4b64c9ace47e92469add7b709e54d83c189be230a15995

Observation 6371f809-fffd-491b-a83b-5c2469aa593b · outbound

This paper cites Compositional visual generation with composable diffusion models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Compositional visual generation with composable diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.413400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.024562Z digest=sha256:29ffe0200fa4e126a16fe970ddacbbd57e0d2daffd644a6699f4c17205a0918d

Observation 8feb338f-f5b1-478a-a0d4-22e1450bbe3f · outbound

This paper cites Referee Can Play: An Alternative Approach to Conditional Generation via Model Inversion.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Referee Can Play: An Alternative Approach to Conditional Generation via Model Inversion

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:28:21.972406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.030570Z digest=sha256:a17f58c2085ac94088a58418d8de43d104b92619e88dd4ada234f1726a3c2fcf

Observation 157268f6-b8e6-4b87-9ac8-364bd77b8a3d · outbound

This paper cites Videofusion: Decomposed diffusion mod- els for high-quality video generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Videofusion: Decomposed diffusion mod- els for high-quality video generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.037497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.037497Z digest=sha256:f344987eb2ea4b6d098426e1d6177a6473c2bd9656ea1c699259662cf21d4d03

Observation b54aa1d8-d318-476c-b764-ec1dc3bd9910 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Latte: Latent Diffusion Transformer for Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.043729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.043729Z digest=sha256:9e47fd0af41da96f915501d1aba5229e002fd29ffb9a94971d8346e45ab3a6eb

Observation a15a99a0-4c63-44b2-92ef-2367495773b8 · outbound

This paper cites CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:28:21.841376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.050467Z digest=sha256:5ce23452d0f02bca70f4f8a2257a2b6aa78f9e93a4d16ccc1eb01207ee564cf8

Observation c22d51a9-b364-430d-ba1c-ec13120a1db0 · outbound

This paper cites Hello GPT-4o.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Hello GPT-4o

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.373456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.060129Z digest=sha256:14be2f27fc3c54202403b82af692d4a10a3461125e2f1572fabdba65cf9ce5f7

Observation c44c8540-211f-4c8c-bd29-13b59bfbf17c · outbound

This paper cites Benchmark for compositional text-to- image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Benchmark for compositional text-to- image synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.346098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.065423Z digest=sha256:6024be0ec159582e29a0efac9f37edc86c41a23d13831ade7a306c902e1dbaa1

Observation 41d8b418-01cd-4c26-933d-4d201263626c · outbound

This paper cites O’Brien, Carrie J.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration O’Brien, Carrie J

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.070725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.070725Z digest=sha256:8e54bf71f10cca191e6b9df8ce79e97f98c4a098484feb5ec1ccccd8deded4b7

Observation 2834ab87-5a69-4fa6-859f-e324e5f24fcb · outbound

This paper cites ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.077475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.077475Z digest=sha256:b2b655dd86025430057b9f1ef833fe01de7c1f4f3e45cfe7d8443307d1c1d754

Observation 82e7c30b-7ceb-4d65-85dc-7e5b0a23679d · outbound

This paper cites Chatdev: Communicative agents for software development,.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Chatdev: Communicative agents for software development,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.296452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.084722Z digest=sha256:0a6c14ac07bbbb2b507b510e4571aee8f82e30b404f1963a1cf09b7a3288f469

Observation ff19bb7a-faec-4850-84eb-9a81bfa2940c · outbound

This paper cites Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.275394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.090503Z digest=sha256:9ded0078d86febe1d697b79f30ef0815b32d79186c08f254bb407bf8137618c1

Observation 670c746c-b792-41e2-9aa9-6dbc1dae81cd · outbound

This paper cites an unresolved cited work.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:28:23.253803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.095761Z digest=sha256:474eedea6a09f0bf3766d7f7428e9954443c52901991875728a56bccedcb96b7

Observation e4d34dee-ca87-434b-b2ae-72f80196fc43 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.101571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.101571Z digest=sha256:ff3f97efb0bf8d09603c5eb95eb36b1d946a37b15b8de0ac5ce5f45958dfc34a

Observation 10d5fbc1-b7aa-4ff8-84a3-ddbfea11f1d4 · outbound

This paper cites Weiss, Niru Mah- eswaranathan, and Surya Ganguli.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Weiss, Niru Mah- eswaranathan, and Surya Ganguli

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.108351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.108351Z digest=sha256:6cd7a46ab75e0937ae3ab60696d6a486a3acdd7eb4aeb1622d8dae9cf82798f2

Observation 95fcbebf-e845-4075-825d-4d21b10b961f · outbound

This paper cites Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.223987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.113904Z digest=sha256:9e5c987e14f1b081f49e2e2c67e3d5e8ff327186ec117a87e74cd3ad09daca94

Observation 0d6b7f9e-7fda-4fda-88f6-e6b44df02246 · outbound

This paper cites T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation,.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.118962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.118962Z digest=sha256:3a2dcab3fbe7d558475fd5e6acb47ef16b9a97c91855a8a366a5614915a3da5d

Observation 6573c5c3-ecf7-4add-9040-9f3fc96951d8 · outbound

This paper cites A survey of neural code intelligence: Paradigms, advances and beyond, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration A survey of neural code intelligence: Paradigms, advances and beyond, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.190825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.125094Z digest=sha256:d924e020f86f61ea223a8b3983846e4877b98077aefa102ab64141c10edccb5f

Observation 5a05c921-02da-40ff-8e2f-1df422505941 · outbound

This paper cites Corex: Pushing the boundaries of complex reasoning through multi-model collaboration, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Corex: Pushing the boundaries of complex reasoning through multi-model collaboration, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.165481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.130009Z digest=sha256:fa244177a0239b2015ec58a69e1834fbf9df60b5059edc40b77f2b0a92f5fce2

Observation a4455696-ccec-4398-8a8a-56f3e3a6c2c7 · outbound

This paper cites Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.135761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.135761Z digest=sha256:fbeadeae9cb34abef808ddef093ee119d7cbc0f37507c1f11ba30c503d17b9f2

Observation 4d632eb3-f1da-4d43-b1c4-4a221893d24c · outbound

This paper cites Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.142452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.141345Z digest=sha256:c996b31d74e8dd4504a82b396138ba8d45c1ae813a348b1963e119f66dd8e838

Observation 5149d81d-e953-4d74-ae43-56418f823262 · outbound

This paper cites Videotetris: Towards compo- sitional text-to-video generation, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Videotetris: Towards compo- sitional text-to-video generation, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.025872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.146892Z digest=sha256:535b61ffd4acc4d517eed1f6905da90a8168180d2112d7e2650f733589bf4d0c

Observation b59720f1-3ece-412e-911a-05fcf03c7e8b · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Phenaki: Variable length video generation from open domain textual descriptions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.152506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.152506Z digest=sha256:60d978d009cbc974f6398ee36b5c785ae2c867849cd38c2931a078d76a870731

Observation 1efbd637-6b74-493d-90ed-0314e0ebe016 · outbound

This paper cites V oyager: An open-ended embodied agent with large language models, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration V oyager: An open-ended embodied agent with large language models, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.979056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.157710Z digest=sha256:cb90ce108b867b64166f371c07fcb83943ba383a18ffd296d627faf67ce72ce0

Observation ff78e883-b859-4042-89d8-9d4c4f69b305 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration ModelScope Text-to-Video Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.164023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.164023Z digest=sha256:756baebf1428e8b42aeb2096ec01ee0285d58c61a0a29641f2e940ccdaa72872

Observation a3fb2505-7ccb-4402-ba3e-77c4e901c229 · outbound

This paper cites Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:28:21.624631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.169358Z digest=sha256:26d00be40eec771300fbbcd9b64d6e6a0f96edfdc4bfacf82bb2b215ca96fbd6

Observation 7316e80a-c952-4674-9b79-96f0a490c7e9 · outbound

This paper cites Xu, Xian- gru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Xu, Xian- gru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.962499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.175569Z digest=sha256:2ac8f1b0e91a1cda17cb0ad444ae2e4484a62362bdf7439a1918d2de9fdbeff8

Observation fcf40c8d-af20-4f0f-ae94-866d3074409d · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image gen- eration and editing, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Genartist: Multimodal llm as an agent for unified image gen- eration and editing, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.936409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.181372Z digest=sha256:28e7cec8c413746523af9ce05bbf053605007819a58da941d2b579ec97fbd5e3

Observation a6f86732-2de4-403c-b4d3-cb37a9736e51 · outbound

This paper cites Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.187840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.187840Z digest=sha256:d30a04fd5ab106a70d2eace090dc6810301b4216077d2e95b87a65f0e2bfb997

Observation 27ad0aa7-92dd-4dc5-ab06-0db6b0e011fe · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.193411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.193411Z digest=sha256:3002b88592e200d86bc10bc698ee87cf6a193aba0d084362b69a6b3af0cdc109

Observation b63825ae-12ee-481d-8340-099df8e54477 · outbound

This paper cites N ¨uwa: Visual synthesis pre- training for neural visual world creation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration N ¨uwa: Visual synthesis pre- training for neural visual world creation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.913597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.199173Z digest=sha256:b9b63edda686acf835818f4145fd7acdd8a423e4e68e0c687f0db7ac26505272

Observation f5df9352-2121-47bb-a119-2d62c8d5d7fc · outbound

This paper cites Autogen: Enabling next-gen llm ap- plications via multi-agent conversation, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Autogen: Enabling next-gen llm ap- plications via multi-agent conversation, 2023

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.885930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.204183Z digest=sha256:4f8656b01478108762c35569b5e8197ba7caaf693a40192828fd1df643089700

Observation c0fcfc26-6ac0-4cff-9a7a-25c7ceb18449 · outbound

This paper cites Harnessing the spatial- temporal attention of diffusion models for high-fidelity text- to-image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Harnessing the spatial- temporal attention of diffusion models for high-fidelity text- to-image synthesis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.861717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.209518Z digest=sha256:1d1f7ddbd5f3b5c4d7f8d1a0debe34c671ac93ef1848be9956876928c4a9c8a4

Observation 51b5997b-13ba-4828-b530-76bf3d409050 · outbound

This paper cites Gonzalez, Boyi Li, and Trevor Darrell.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Gonzalez, Boyi Li, and Trevor Darrell

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.835815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.215085Z digest=sha256:be0689fad6be4b3cfefb820a8f7041c1d4fe99335bcc96d5203333bcc5880186

Observation d18b211b-9d12-46c5-9a1e-15169c08dc39 · outbound

This paper cites Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.219961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.219961Z digest=sha256:dc004710685ccc0b0880c81d72aae6c16053b6c4c98d18f8ed3ce856529cb28c

Observation 3179a8cc-d168-4553-ad50-504e9936edf5 · outbound

This paper cites Compositional video gen- eration as flow equalization, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Compositional video gen- eration as flow equalization, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.789679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.225607Z digest=sha256:eea907a9c46a9c6ec36371ce84ecac9302552d71356a6e31dd034cd63cf44d55

Observation 6fe3f2b2-8432-45fa-a21b-54d7148d909d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.230637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.230637Z digest=sha256:cd21a0e3604e2147452ff4fb94979213c3fb9ad0ef91f231ab35876a6eb140fd

Observation 4b1cb226-ee5e-472f-bd39-fe4fcbdc1f87 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Webshop: Towards scalable real-world web interaction with grounded language agents, 2023

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.735156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.236451Z digest=sha256:57e2e9a2612aecdb7fd2a1b912a4e9d9c6dff7b5273435a90a80704a8af94c3e

Observation fd5e7ede-a858-4c9e-ae27-41ba85e031f3 · outbound

This paper cites Magvit: Masked generative video transformer.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Magvit: Masked generative video transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.242478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.242478Z digest=sha256:4b743a211d87d1ab46c608adb24079e1a1deb9e890cbd1e8a13d25bebd5c38e6

Observation 0f53b9e2-e613-46b4-bb31-e3be780a5bdd · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.247742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.247742Z digest=sha256:5d0a84ba659c2d0ee2393069c4b0215594f06853061a3f1a8c40c26ad888f98b

Observation e2a25fc4-6444-4931-a668-21f21abf41ae · outbound

This paper cites MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.252713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.252713Z digest=sha256:3d9a7a06e5afb17d7164bd61031bdb15fc2621b47580b2c28c0f871543475bb1

Observation 05aaf634-d709-4dca-a0cd-73623d514bd3 · outbound

This paper cites Mora: En- abling generalist video generation via a multi-agent frame- work, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Mora: En- abling generalist video generation via a multi-agent frame- work, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.646011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.259547Z digest=sha256:780346a8bbb4cc694a3ee4382bcef3a0f929f5d8cf22407fa80afc707901af02

Observation 4000ac06-1fad-46c9-b138-cd51eea9b84c · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.264413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.264413Z digest=sha256:5557529e47b5cca93fb1a5209bdd6fe7162c05f4d2209d7a4794cab14b384c4e

Observation 9f010a1f-46a8-4d8c-80e8-b16658ce10fc · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.269581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.269581Z digest=sha256:6119077600f791b7e98d1d156f2cfdfbb7ba0603f9b47b77633aa6bbe16c35e1

Observation 727b02a2-57de-4ab2-901e-9bbf57dbf13f · outbound

This paper cites a glass sculpture.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration a glass sculpture

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.623594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.274783Z digest=sha256:187d54571653e4227bcc4223a68db20d6966a7f22cab3bbc44272aa1372f75ca

Observation 942799bb-1375-47e3-ba4f-7d28aa06eac9 · outbound

This paper cites porcelain rabbit.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration porcelain rabbit

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.525704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.281714Z digest=sha256:218864a5ebe26e1ede31c7704d40ce3e46a7676b307ea40c0a25058bf6787b67

Observation edca6bb7-095b-4d6c-82c5-2b347e971ccf · outbound

This paper cites Rabbit police officer directs traffic.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Rabbit police officer directs traffic

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.431932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.293365Z digest=sha256:67883ef3f0c53a67c78bddb71a464b2f3cea8c1a1702b9a646e798536b645e1d

Observation cec2947f-d5c4-424a-81a0-70f46b6003bf · outbound

This paper cites This could involve positioning the rabbit with an arm raised or using a gesture to indicate traffic direction.2.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration This could involve positioning the rabbit with an arm raised or using a gesture to indicate traffic direction.2

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.330466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:28:21.301097Z digest=sha256:f67bf3dec1cf7d9e1faf3a05d264a0582d83dadb73b76549038540e9076804f9

Pith citing papers

Observation c7a59eb6-89c6-4ad0-868b-3e2c394d4aab · inbound

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation cites this paper.

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:52.701228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T01:46:14.430539Z digest=sha256:02735617b681d2b569ad3d5002a93e9874910bbdbb514b9633c949716b86e0ba

Observation e5024161-ddd4-46bd-b301-a60898369938 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:16.731285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:16.731285Z digest=sha256:1e4d92104e27a9dfe1d557a11196eaea0783aa7ad606c588f8f6cefdfcff8ffa