Pith. sign in

Paper Citation Record · LEDGER

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

As of 14 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 2 inbound Pith citation observations for arXiv:2412.04440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04440 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:28:21.301097Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T05:27:16.731285Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:28:52.699573Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact3
  • verified fuzzy38
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aeab1852-7b01-4467-9d91-286f901f7c31 · outbound

This paper cites https : / / huggingface.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration https : / / huggingface

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:24.048844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.831324Z digest=sha256:66cef52f9419891a2bdaeb9d4723063c0054494e712e049ef408f66cb833ab81

Observation 1e572bb3-01f3-4908-bbc1-bb97e7630d3f · outbound

This paper cites https://pika.art/, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration https://pika.art/, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.993015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.837758Z digest=sha256:c92293b9fc0927562ed5b0345fdf3c95482aaf205ed49a2dc0a35e7c8f172395

Observation 00949303-f71d-41d2-ab55-2af4b7c42bd2 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.845667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.845667Z digest=sha256:26183c3adcaa9ac46a514e7b8fdb66093d3182fef038c8d9398122e65523002b

Observation 9f78e4b8-de1c-4845-b63f-431a7f652e93 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.853812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.853812Z digest=sha256:135a5683084719f15c2f8a396a36dd1eda4824370739d0df2a8f7d68c74530bb

Observation e99f32f2-07b0-4a7f-a178-e43bb1e08591 · outbound

This paper cites Maskgit: Masked generative image transformer.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Maskgit: Masked generative image transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.860177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.860177Z digest=sha256:8e941439b0f4b2a2decbcd82dbc31d2c07b942f9f98a9c130a0c0996bf7a6e2e

Observation 29928a2a-1cda-4304-9885-ba0406f0924a · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.865826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.865826Z digest=sha256:28eb2244404a9abf28232205b1185726ed85ebc9b98f0bd2449f73c8cd846cdb

Observation 6138a6df-d913-445a-8d82-ac3d5e9ba11a · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.806787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.875964Z digest=sha256:8dc39860e8cffc452c283d4ea313851ca9828a0c02d70eb87bcc6a5626c0b52e

Observation 8e0ca905-fd28-49c0-bb89-10eb207807d1 · outbound

This paper cites VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.884384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.884384Z digest=sha256:9900d1b53d3415980d34620817dde272aa01a832d416ba0ef30515172454d3c6

Observation 066dbf94-aa40-489c-b320-9e17afd41d7c · outbound

This paper cites Training- free layout control with cross-attention guidance.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Training- free layout control with cross-attention guidance

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.771600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.890486Z digest=sha256:3d34045fa54c302fac33d4aabe4947a060d2b2d5241296acdf3b753961a49f49

Observation 312f5fcb-2635-4e50-a77f-746ae2dc346d · outbound

This paper cites Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.896387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.896387Z digest=sha256:8e85d9cbc73c80d17d4d38b90a4bed4ccbf5b573a4bd680fa1e17edb34018109

Observation 669d1303-adaa-4124-9bc6-408fcf5d7652 · outbound

This paper cites Mind2web: Towards a generalist agent for the web, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Mind2web: Towards a generalist agent for the web, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.747093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.902905Z digest=sha256:39138087cd65f29205e8782eb99617cb79b2ba3b316f6751e329425c1e204a02

Observation 329da210-fafc-444c-8bf0-1b790c8f685b · outbound

This paper cites an unresolved cited work.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:28:23.724429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.909328Z digest=sha256:ff1c818c04cd3f4d1c2a1ef1c53b63b0af6fa3d697674471720ac846b8520429

Observation 5652a914-6f6f-41ae-a5a1-5ace08a38c5d · outbound

This paper cites Training-free structured diffu- sion guidance for compositional text-to-image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Training-free structured diffu- sion guidance for compositional text-to-image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.690044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.916649Z digest=sha256:1fa23c031e3dbb2b4e783439c7d8a5c00c22d2c56b6af878463160a5f8a3c1fa

Observation 26858460-4a3b-4e46-ae38-82b138f5d128 · outbound

This paper cites LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.922439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.922439Z digest=sha256:443b0b3420807b29d5d3f42b16caa9ff93b52696fc8adef0deabf10723afc7e8

Observation 9b2ac166-c2b6-40db-acfd-a90c8f860000 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.928645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.928645Z digest=sha256:bd067a4e7bbee8be94afebe52185eda3cd744140b1817f6248604d6de399f204

Observation a0ab5767-075c-456e-a619-15566b85d415 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.934639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.934639Z digest=sha256:9bbb7e4ac774292f04a6a90095f513c902dabbeab656782c773f7f9d604769af

Observation 04a428a9-d5fa-4965-ade5-0c35fd06b96b · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Denoising diffu- sion probabilistic models, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.941001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.941001Z digest=sha256:68f2f105bb7425a448c63719187787ee496d17478396d10e7c3e4a6bd741b286

Observation 3e93b76c-5144-466c-a344-6fe5aae94940 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Imagen Video: High Definition Video Generation with Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.946653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.946653Z digest=sha256:dec5816288bde4aa94637d32e75485f5ad4eb6f8636fc16b94e783ec12298912

Observation 71828357-ae88-4d65-9c55-caf33b423133 · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Metagpt: Meta programming for a multi-agent collaborative framework, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.636186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.953139Z digest=sha256:891959c381308427dec5493488e8c3715d49d20a0244a45b444a9ced2a880c38

Observation cf084e4a-06a1-4d8f-8d7d-3607b8e36167 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.960029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.960029Z digest=sha256:46514c41b0a0e70ce67cfc7afbb59f12d04829e227d6c43c2fc078a859726896

Observation 42b64119-d7f0-43f5-a8ec-c9bb6bbda202 · outbound

This paper cites Open-sora: Democratizing efficient video pro- duction for all, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Open-sora: Democratizing efficient video pro- duction for all, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.610098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.965574Z digest=sha256:84060ebf55ab06d2c883f28124221eb3d5df5004d16bdfd994a5e29ff951c2a2

Observation 3aa475d1-7bf5-41d1-ba44-0ea23c3a5db5 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.576063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.972538Z digest=sha256:ce5480f60ab6c99705658262b7ebfa455f7188648fb402b9fee7911428708c92

Observation e612d41b-dd4a-4e01-9bdb-366cb2138164 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.978399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.978399Z digest=sha256:91c33d5f180de58f00095f62e87d4264297b668615dfc9a6d48c95757377a92e

Observation 576d222d-0cec-4343-8aa9-8dc60fcb3a59 · outbound

This paper cites Dense text-to-image generation with attention modulation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Dense text-to-image generation with attention modulation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.521624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.984086Z digest=sha256:1d2d52b6c4cc8ff015c8041f59ee394f40bf17b0fbb76ebc1e0479aa600a2f7c

Observation 72995c45-1d99-47d1-98ca-5be64aecafe3 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.989614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.989614Z digest=sha256:9329d54daf438bf408a2adadb1730498f92a710c786be9dd2e56ae69abf8cabb

Observation 6439a76a-ddde-400c-932f-87ea3dd8f53a · outbound

This paper cites Open-sora-plan, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Open-sora-plan, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.498187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:20.995116Z digest=sha256:08fd706a0605d7432f802a0d56d54a05d534798efbb8cf4c6a3c1ccfd50e863a

Observation 136af769-2592-4035-bf3f-b883f43945ac · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Gligen: Open-set grounded text-to-image generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.475878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.000489Z digest=sha256:0ab3fe325bb9b9e553852ae9351b5c57b09faece7182bb0faf2d3c6c60a9a950

Observation 94a1fafa-c26a-42a6-95dc-c6aba4f41d32 · outbound

This paper cites Stylet2i: Toward compositional and high-fidelity text- to-image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Stylet2i: Toward compositional and high-fidelity text- to-image synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.453249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.006516Z digest=sha256:f9c2c25d652922329a4ac6feb208b46ec80224772b6b30dfad4271c244bbc5ff

Observation 6521fa0a-33c7-4ef2-8f1b-009026c3036c · outbound

This paper cites LLM-grounded Video Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration LLM-grounded Video Diffusion Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.012192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.012192Z digest=sha256:a6c3c1b202bb2469f075509d52a950789433e84699bc62644332b77992a99913

Observation cce2e134-ac28-47d5-a1cf-1bb3d2984ef9 · outbound

This paper cites Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.433787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.017820Z digest=sha256:50beea549e1b35da5c4453cfd51eef04005df31200d74691662475bd8c1eed35

Observation 6371f809-fffd-491b-a83b-5c2469aa593b · outbound

This paper cites Compositional visual generation with composable diffusion models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Compositional visual generation with composable diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.413400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.024562Z digest=sha256:f21f126e211d7eedba16e5895a12489b0d429fc50f9700fc64db99e8c27e8ddd

Observation 8feb338f-f5b1-478a-a0d4-22e1450bbe3f · outbound

This paper cites Referee Can Play: An Alternative Approach to Conditional Generation via Model Inversion.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Referee Can Play: An Alternative Approach to Conditional Generation via Model Inversion

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:28:21.972406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.030570Z digest=sha256:208ac7304b0b6e8c077991245ed48de1e9e297855f144d1f4bc024ccd66056cc

Observation 157268f6-b8e6-4b87-9ac8-364bd77b8a3d · outbound

This paper cites Videofusion: Decomposed diffusion mod- els for high-quality video generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Videofusion: Decomposed diffusion mod- els for high-quality video generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.037497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.037497Z digest=sha256:f344987eb2ea4b6d098426e1d6177a6473c2bd9656ea1c699259662cf21d4d03

Observation b54aa1d8-d318-476c-b764-ec1dc3bd9910 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Latte: Latent Diffusion Transformer for Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.043729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.043729Z digest=sha256:9e47fd0af41da96f915501d1aba5229e002fd29ffb9a94971d8346e45ab3a6eb

Observation a15a99a0-4c63-44b2-92ef-2367495773b8 · outbound

This paper cites CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:28:21.841376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.050467Z digest=sha256:cb08f61b91bc43b382ce82eea5661a09a0e88f0b8cb6c32a4ea6fa65135e1a6f

Observation c22d51a9-b364-430d-ba1c-ec13120a1db0 · outbound

This paper cites Hello GPT-4o.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Hello GPT-4o

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.373456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.060129Z digest=sha256:54aab8782d6f3d13ae3e54d745c76b7b12a36b0f4f2c217d7d12a21c4a2e8603

Observation c44c8540-211f-4c8c-bd29-13b59bfbf17c · outbound

This paper cites Benchmark for compositional text-to- image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Benchmark for compositional text-to- image synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.346098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.065423Z digest=sha256:86b83b4e89f410af82acb656c31e6d812a1dd2722048b2a08d6c9a4a88dedc51

Observation 41d8b418-01cd-4c26-933d-4d201263626c · outbound

This paper cites O’Brien, Carrie J.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration O’Brien, Carrie J

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.070725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.070725Z digest=sha256:8e54bf71f10cca191e6b9df8ce79e97f98c4a098484feb5ec1ccccd8deded4b7

Observation 2834ab87-5a69-4fa6-859f-e324e5f24fcb · outbound

This paper cites ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.077475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.077475Z digest=sha256:b2b655dd86025430057b9f1ef833fe01de7c1f4f3e45cfe7d8443307d1c1d754

Observation 82e7c30b-7ceb-4d65-85dc-7e5b0a23679d · outbound

This paper cites Chatdev: Communicative agents for software development,.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Chatdev: Communicative agents for software development,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.296452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.084722Z digest=sha256:4ff6ae44efa862a45e00f47a2038de90021d18adfcc911c23671ef1057cce0b1

Observation ff19bb7a-faec-4850-84eb-9a81bfa2940c · outbound

This paper cites Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.275394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.090503Z digest=sha256:60b77dccb76390bb92e352f17eabf154aa21f7c50c4578aec27bd9656df946a2

Observation 670c746c-b792-41e2-9aa9-6dbc1dae81cd · outbound

This paper cites an unresolved cited work.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:28:23.253803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.095761Z digest=sha256:5680dfca861289011505f664246f8dc4fc7931c983a41dfac3835c1f66e91287

Observation e4d34dee-ca87-434b-b2ae-72f80196fc43 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.101571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.101571Z digest=sha256:ff3f97efb0bf8d09603c5eb95eb36b1d946a37b15b8de0ac5ce5f45958dfc34a

Observation 10d5fbc1-b7aa-4ff8-84a3-ddbfea11f1d4 · outbound

This paper cites Weiss, Niru Mah- eswaranathan, and Surya Ganguli.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Weiss, Niru Mah- eswaranathan, and Surya Ganguli

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.108351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.108351Z digest=sha256:6cd7a46ab75e0937ae3ab60696d6a486a3acdd7eb4aeb1622d8dae9cf82798f2

Observation 95fcbebf-e845-4075-825d-4d21b10b961f · outbound

This paper cites Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.223987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.113904Z digest=sha256:4cf377a0c4a66abc717f0171aa0a41a73789b46e968d7a9214c50327f3b50ec4

Observation 0d6b7f9e-7fda-4fda-88f6-e6b44df02246 · outbound

This paper cites T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation,.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.118962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.118962Z digest=sha256:3a2dcab3fbe7d558475fd5e6acb47ef16b9a97c91855a8a366a5614915a3da5d

Observation 6573c5c3-ecf7-4add-9040-9f3fc96951d8 · outbound

This paper cites A survey of neural code intelligence: Paradigms, advances and beyond, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration A survey of neural code intelligence: Paradigms, advances and beyond, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.190825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.125094Z digest=sha256:ccb9a9fbd6600d1007bf98e708cca7cfe5eb4a4d832f5f2d3f21cda63bb50bb0

Observation 5a05c921-02da-40ff-8e2f-1df422505941 · outbound

This paper cites Corex: Pushing the boundaries of complex reasoning through multi-model collaboration, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Corex: Pushing the boundaries of complex reasoning through multi-model collaboration, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.165481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.130009Z digest=sha256:63efb05124e8732504bd0e7d490a27361deb76f53c85caaa7642b0b8bdf25bcd

Observation a4455696-ccec-4398-8a8a-56f3e3a6c2c7 · outbound

This paper cites Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.135761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.135761Z digest=sha256:fbeadeae9cb34abef808ddef093ee119d7cbc0f37507c1f11ba30c503d17b9f2

Observation 4d632eb3-f1da-4d43-b1c4-4a221893d24c · outbound

This paper cites Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.142452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.141345Z digest=sha256:6dfe6aa05b56a19bf08dceb16918fa74a69f95423af1ae1cada29c499688cd3f

Observation 5149d81d-e953-4d74-ae43-56418f823262 · outbound

This paper cites Videotetris: Towards compo- sitional text-to-video generation, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Videotetris: Towards compo- sitional text-to-video generation, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:23.025872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.146892Z digest=sha256:012cfe1cd135a1994034a6d8935690a7892741fda8e09a5f4dd4431b1ed32f0d

Observation b59720f1-3ece-412e-911a-05fcf03c7e8b · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Phenaki: Variable length video generation from open domain textual descriptions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.152506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.152506Z digest=sha256:60d978d009cbc974f6398ee36b5c785ae2c867849cd38c2931a078d76a870731

Observation 1efbd637-6b74-493d-90ed-0314e0ebe016 · outbound

This paper cites V oyager: An open-ended embodied agent with large language models, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration V oyager: An open-ended embodied agent with large language models, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.979056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.157710Z digest=sha256:6f222b4055e5222d90874e6df21163f695dd432cf6723ec6e1395bed6c318299

Observation ff78e883-b859-4042-89d8-9d4c4f69b305 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration ModelScope Text-to-Video Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.164023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.164023Z digest=sha256:756baebf1428e8b42aeb2096ec01ee0285d58c61a0a29641f2e940ccdaa72872

Observation a3fb2505-7ccb-4402-ba3e-77c4e901c229 · outbound

This paper cites Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:28:21.624631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.169358Z digest=sha256:7504935bbef2fcb157861e3d32919de1617b347f5a431582ec0998ecc3adcccb

Observation 7316e80a-c952-4674-9b79-96f0a490c7e9 · outbound

This paper cites Xu, Xian- gru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Xu, Xian- gru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.962499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.175569Z digest=sha256:e7c5ccb889d01b791ae8db2eefc2fe09122c1ac8757552cd9c56d4d3372db19f

Observation fcf40c8d-af20-4f0f-ae94-866d3074409d · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image gen- eration and editing, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Genartist: Multimodal llm as an agent for unified image gen- eration and editing, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.936409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.181372Z digest=sha256:960ee75068a75acdbada439b9487f423899a3d99636471bb14f49309fbf3baca

Observation a6f86732-2de4-403c-b4d3-cb37a9736e51 · outbound

This paper cites Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.187840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.187840Z digest=sha256:d30a04fd5ab106a70d2eace090dc6810301b4216077d2e95b87a65f0e2bfb997

Observation 27ad0aa7-92dd-4dc5-ab06-0db6b0e011fe · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.193411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.193411Z digest=sha256:3002b88592e200d86bc10bc698ee87cf6a193aba0d084362b69a6b3af0cdc109

Observation b63825ae-12ee-481d-8340-099df8e54477 · outbound

This paper cites N ¨uwa: Visual synthesis pre- training for neural visual world creation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration N ¨uwa: Visual synthesis pre- training for neural visual world creation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.913597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.199173Z digest=sha256:29dead0a3ae44f4ad64564e2b4fd46a8bb132bda46c730ca10750bfc35d9efb0

Observation f5df9352-2121-47bb-a119-2d62c8d5d7fc · outbound

This paper cites Autogen: Enabling next-gen llm ap- plications via multi-agent conversation, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Autogen: Enabling next-gen llm ap- plications via multi-agent conversation, 2023

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.885930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.204183Z digest=sha256:eaf7fe694202cad32de4d6464c9a0d528c7312ea4682ad84559bdfdbd79ed2cd

Observation c0fcfc26-6ac0-4cff-9a7a-25c7ceb18449 · outbound

This paper cites Harnessing the spatial- temporal attention of diffusion models for high-fidelity text- to-image synthesis.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Harnessing the spatial- temporal attention of diffusion models for high-fidelity text- to-image synthesis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.861717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.209518Z digest=sha256:f7f319088386aa66d61ec3aeb7d5c603ee3f16236ec5be66555ea9980eecc9bb

Observation 51b5997b-13ba-4828-b530-76bf3d409050 · outbound

This paper cites Gonzalez, Boyi Li, and Trevor Darrell.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Gonzalez, Boyi Li, and Trevor Darrell

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.835815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.215085Z digest=sha256:7c71b4f964e4d65d6a66ab1bce1a1bf157d389fe1b97ecbea0cbbf1577e4a4bd

Observation d18b211b-9d12-46c5-9a1e-15169c08dc39 · outbound

This paper cites Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.219961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.219961Z digest=sha256:dc004710685ccc0b0880c81d72aae6c16053b6c4c98d18f8ed3ce856529cb28c

Observation 3179a8cc-d168-4553-ad50-504e9936edf5 · outbound

This paper cites Compositional video gen- eration as flow equalization, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Compositional video gen- eration as flow equalization, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.789679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.225607Z digest=sha256:d0673d37dc3926b78bb514181746c5e06d2d173f6199ffb58293151c498ff01f

Observation 6fe3f2b2-8432-45fa-a21b-54d7148d909d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.230637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.230637Z digest=sha256:cd21a0e3604e2147452ff4fb94979213c3fb9ad0ef91f231ab35876a6eb140fd

Observation 4b1cb226-ee5e-472f-bd39-fe4fcbdc1f87 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents, 2023.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Webshop: Towards scalable real-world web interaction with grounded language agents, 2023

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.735156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.236451Z digest=sha256:0fd235aeccdc843a7feb97bf07c605106387d4f2ce181bd9cd6a7ba1c5eedfb1

Observation fd5e7ede-a858-4c9e-ae27-41ba85e031f3 · outbound

This paper cites Magvit: Masked generative video transformer.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Magvit: Masked generative video transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.242478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.242478Z digest=sha256:4b743a211d87d1ab46c608adb24079e1a1deb9e890cbd1e8a13d25bebd5c38e6

Observation 0f53b9e2-e613-46b4-bb31-e3be780a5bdd · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.247742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.247742Z digest=sha256:5d0a84ba659c2d0ee2393069c4b0215594f06853061a3f1a8c40c26ad888f98b

Observation e2a25fc4-6444-4931-a668-21f21abf41ae · outbound

This paper cites MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.252713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.252713Z digest=sha256:3d9a7a06e5afb17d7164bd61031bdb15fc2621b47580b2c28c0f871543475bb1

Observation 05aaf634-d709-4dca-a0cd-73623d514bd3 · outbound

This paper cites Mora: En- abling generalist video generation via a multi-agent frame- work, 2024.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Mora: En- abling generalist video generation via a multi-agent frame- work, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.646011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.259547Z digest=sha256:cdcdeba1b3a304770eda6b5a9f8a3a5298c550cff2d7b5c129761810bff452f9

Observation 4000ac06-1fad-46c9-b138-cd51eea9b84c · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.264413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.264413Z digest=sha256:5557529e47b5cca93fb1a5209bdd6fe7162c05f4d2209d7a4794cab14b384c4e

Observation 9f010a1f-46a8-4d8c-80e8-b16658ce10fc · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:21.269581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:21.269581Z digest=sha256:6119077600f791b7e98d1d156f2cfdfbb7ba0603f9b47b77633aa6bbe16c35e1

Observation 727b02a2-57de-4ab2-901e-9bbf57dbf13f · outbound

This paper cites a glass sculpture.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration a glass sculpture

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.623594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.274783Z digest=sha256:1741eb004c6fe41396badcb96c66a19736a4d5410ec77d998adb326a66ea4c08

Observation 942799bb-1375-47e3-ba4f-7d28aa06eac9 · outbound

This paper cites porcelain rabbit.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration porcelain rabbit

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.525704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.281714Z digest=sha256:780fb9e8cb0bdc89f11bd236878d7658334f0d4256f0ef356371aafc7331b4ea

Observation edca6bb7-095b-4d6c-82c5-2b347e971ccf · outbound

This paper cites Rabbit police officer directs traffic.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Rabbit police officer directs traffic

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.431932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.293365Z digest=sha256:79060dfafd2f41c0297e316f6e7912e95bc5f2d89b9b3ffc827ae5cf33270929

Observation cec2947f-d5c4-424a-81a0-70f46b6003bf · outbound

This paper cites This could involve positioning the rabbit with an arm raised or using a gesture to indicate traffic direction.2.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration This could involve positioning the rabbit with an arm raised or using a gesture to indicate traffic direction.2

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:22.330466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:28:21.301097Z digest=sha256:ebcbfc5837a21e95fbe8eacd3032b22c89a95235546140301fdc42bb50d40a83

Pith citing papers

Observation c7a59eb6-89c6-4ad0-868b-3e2c394d4aab · inbound

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation cites this paper.

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:52.701228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:46:14.430539Z digest=sha256:31f541930dd776537cfd7acea5ff5b5966351f31d70807b8bb5bbfca63ba7e7f

Observation e5024161-ddd4-46bd-b301-a60898369938 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:16.731285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:16.731285Z digest=sha256:1e4d92104e27a9dfe1d557a11196eaea0783aa7ad606c588f8f6cefdfcff8ffa