Pith. sign in

Paper Citation Record · LEDGER

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

As of 19 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2412.09283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09283 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:10:22.228795Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f098c3e9-9f2f-4ffc-ae90-4dc2769f8607 · outbound

This paper cites Chen and William B.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Chen and William B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.322697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.885108Z digest=sha256:c0027dfb8c281284da2cd1142dbd78306129f37dbe752955ac57271586833d6b

Observation d57e239f-5270-4cfb-9c6a-ac55c55d4ef2 · outbound

This paper cites VideoCrafter2: Overcoming data limitations for high-quality video diffusion models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption VideoCrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.306471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.891074Z digest=sha256:1c244062cec7529302afe3045b5b3456ef9055185738348422a65ab31b49549c

Observation 4ebfefe2-7ddf-48c5-b74f-a80380af84ad · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.897243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.897243Z digest=sha256:89d88790f5eec50a8261f815b55170753eeb299aa4388f07a2605d4f4516b3e0

Observation 85085582-55ce-400c-b4df-279fb124da4a · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.903106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.903106Z digest=sha256:8f12ae3bb2f9f65aa544d73485b51a0cd46390a240dd3190b4ff3b8f5eb00cce

Observation d23350ec-f60d-4ba0-b2ce-ae3afb25ebc2 · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.909455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.909455Z digest=sha256:2634975b0b9fd514b7cd87fe6f38384b73d034dab671f30438a192766238db2a

Observation 1ec20df2-4806-43e7-802c-3ede21157525 · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption VBench: Com- prehensive benchmark suite for video generative models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.290914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.915566Z digest=sha256:c0dea58bd8a781229bae4ecb0bf0183a454e2b0f35243ae916130b8c9ab892ba

Observation 8edaebb5-2364-49c7-81e8-10cffce4725a · outbound

This paper cites Survey of hallucination in natural language generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Survey of hallucination in natural language generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.274241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.920916Z digest=sha256:2e7ac6c42048ff443c104d262761a1c2a4325bed68a38a1a05ff95c454f1c919

Observation 4c7fbfbb-a11c-4a9e-9d90-b56edcc76973 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Pyramidal flow matching for efficient video generative modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.927418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.927418Z digest=sha256:ee3721c0954f4361b76e19601b0e9b1ad8d7536ecf096f17bffa49d4c0ac8b1b

Observation 5af02d9c-2cc1-411a-8e94-3e020a7ddb98 · outbound

This paper cites MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.932432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.932432Z digest=sha256:414a796a050897a62738bade0baa84104ad277261a289ec98916332a5f03f709

Observation 6118f070-b677-4b94-bfa0-8843cfd55cb1 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:53.257171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.937711Z digest=sha256:7e1a188624f7c27fd6fe23729978354acd5405f79882c0aebdc2ceb4b0dd6042

Observation 77183dd7-47d2-40e3-b7b2-6b85611a20a7 · outbound

This paper cites Open-sora-plan, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-sora-plan, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.241808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.942609Z digest=sha256:8888ba8bf5c8be7748885d5b4b2010f480545ca519a6ea4da93f1120f106c150

Observation af50f70e-e17f-423e-ae7f-d685af56199d · outbound

This paper cites T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.947570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.947570Z digest=sha256:244e6667a7ad3022e2188dc2e564fbb1f051b623df1d23b7a032c3a331cc3477

Observation 4b43d96d-9d72-4105-866b-a4156803ea0e · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Evaluating text-to-visual generation with image-to-text gen- eration, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.223906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.952218Z digest=sha256:afac9c9041d4c2b3f9ffe235f9030eb5e2a1b42bf7553dfa78492144e9eaf031

Observation 73ec0bcf-72e9-4f34-a786-a13eef6bfc78 · outbound

This paper cites EvalCrafter: Benchmarking and Evaluating Large Video Generation Models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.957349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.957349Z digest=sha256:c656d8dcf8b3d5f7c07b47799bb3b9670bf28201cde7a93de153d985f9ee7eb0

Observation fd21542b-0112-4ad1-99bd-b9d3eb2f66ab · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Latte: Latent Diffusion Transformer for Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.963009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.963009Z digest=sha256:72e30eab4fed0601f310344cb46a273d6835fa41fd217a563a445597fac4f0cb

Observation 2c2e2bc9-c3de-4886-87df-a7dc1680c0ce · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.969388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.969388Z digest=sha256:4fd903f5bb2adc598dff5e3aef0fc28c9fa8f033f561e9c024d4c29d397426a2

Observation fc576829-b7bc-4388-b783-e820dfee3991 · outbound

This paper cites Animal kingdom: A large and diverse dataset for animal behavior understanding.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Animal kingdom: A large and diverse dataset for animal behavior understanding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.206937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.975014Z digest=sha256:5ac9dcfaa359e922df7b2955e7102cad013f873f7e38e3f35538708ca0b8748d

Observation 0054ac26-a500-4ce8-910c-64dd8e702de0 · outbound

This paper cites Pika 1.0.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Pika 1.0

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.189391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.980801Z digest=sha256:e33573a80cd4bdbbc6a8a1dbd3114aecbaab8b11e5329af480299962601a3fc5

Observation 0ec2a843-92e2-4d53-a88b-bc84f7fd44a4 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Learning transferable visual models from natural language supervision, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.172796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.986290Z digest=sha256:1a29ef4f76ac3988db8e5624656456984120c61b64f4cabe70194be41d33b2b9

Observation 9bea47ba-645f-4e3f-b1da-5748cc258da8 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption SAM 2: Segment Anything in Images and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.991371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.991371Z digest=sha256:8ae33c332940490534a9cba0747cc4d79bf693653d42afbe751a4b1e81cba60d

Observation ed6640e2-e24f-4d5d-aa8e-584e8ddaf2be · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:53.156086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:21.996655Z digest=sha256:e3a0627ae980ad52124003cf3871bc2042310351ff105749aff40a97d319ce9b

Observation d9812e93-3ff8-40bc-bfca-468994f453ef · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.137911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.003117Z digest=sha256:ce945030bdb48a4eefac63975f75724671e8244a8a46fe868da0e2f859cc272f

Observation fee51e57-7e0b-4dac-aff5-61ac3bdbaa7a · outbound

This paper cites ModelScope Text-to-Video Technical Report.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ModelScope Text-to-Video Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.008174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.008174Z digest=sha256:6da547c1730c1ae770fc9a3b5c54e56e0f5195a0545d38fb04e92be1366f4482

Observation 9e21e2c1-3a3e-4bb0-bc4c-0e12308fd022 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research,.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Vatex: A large-scale, high- quality multilingual dataset for video-and-language research,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.114885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.014362Z digest=sha256:32e77d122d68087207e4a082737b558864d04c1d352f71e6266c49301a7d8b51

Observation aa616b46-75ad-4d33-9d72-c21d88d7e1c5 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.020323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.020323Z digest=sha256:3c25108318c0ce42d4df5fa9ee4d8e1d3919ed311254ea87eb3b8cb457c9dc48

Observation 7c4a398b-520a-4cc2-b9ba-31dfd64e4ff7 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.097058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.030087Z digest=sha256:7fd4cd973c541425a9a26c66fa727286023b6f749253addc34874c23c8a7b4c2

Observation 95c1768c-6195-40ed-9b97-b37311e9020d · outbound

This paper cites Ifadapter: Instance feature control for grounded text-to-image generation, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Ifadapter: Instance feature control for grounded text-to-image generation, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.079417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.037006Z digest=sha256:cd5ab50e2264b3d9553c49c90f27587916af3f6a2e76405eb8f3785b18bd3823

Observation 6ac07ef2-95a5-4b98-a847-e3c1166cad5c · outbound

This paper cites Vript: A video is worth thousands of words, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Vript: A video is worth thousands of words, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.061606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.042805Z digest=sha256:52d3d695021375092458d3ef04086694c9972c69024c3c7b3f22a534b7d8e124

Observation a450c8be-4544-4f05-b717-c9bd27bf6505 · outbound

This paper cites Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.045703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.055946Z digest=sha256:8bbf4a70f1a8eb8b673365e3c6811e051f2e699a7ac2536a22cab327ed94efe6

Observation 68f197f2-abba-4d39-9807-5a81253164fe · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.062619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.062619Z digest=sha256:383483df5201b9b0b3f51a4121ccd9160d2e931db5fa39132221cfd0ff102155

Observation 031a6f72-fe6e-42f0-ade5-027b6ab50fe2 · outbound

This paper cites Cpt: Colorful prompt tuning for pre-trained vision-language models.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Cpt: Colorful prompt tuning for pre-trained vision-language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:22.069047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:22.069047Z digest=sha256:c819bb04d4ef1ae57ff74cf641e0d2fa1e6109a33b9c5fc7937a06f349cd7b0c

Observation 5fa98523-e966-4445-9cde-15042c4c94a0 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.018846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.074359Z digest=sha256:84ec60a8df7a47ee81379dad23fd1be1e08950ed85753fd60ee44caaa8a0d5ef

Observation d19740f5-efad-4cb3-a74c-17647e40ea17 · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Efros, Eli Shecht- man, and Oliver Wang

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:53.002158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.081845Z digest=sha256:215429f314633f36a842e30e20c3d37e9d9841a50108946cd0e91c77aad00cb2

Observation ca537b78-a784-43d9-9adb-6e20807cc484 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Video instruction tuning with synthetic data, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.984990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.087507Z digest=sha256:600fb1f8bdfabc662dab808f49fc7c954230524588b16d095059e0f5512e17af

Observation f809e469-f5a1-44f6-a958-4e226d7f06e2 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-sora: Democratizing efficient video production for all, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.968053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.095179Z digest=sha256:8a5afdf2fb15fcc2722bde09fe775d34b86cfbc1b31514e6579345bb68d742fc

Observation cc17b6b7-170b-4017-870e-873af743e6aa · outbound

This paper cites Open-Sora: Democratizing efficient video production for all.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Open-Sora: Democratizing efficient video production for all

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.951776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.101440Z digest=sha256:9a8f1c8c0030f0c30fd351fe7bb531c9ff0932aaf6666c1ad0e002829d9f195d

Observation 95ddd32b-5419-4f6e-bcbd-1f2343d8fd40 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.935738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.107772Z digest=sha256:af4bc6380426b3832d0311ec7546b8939eccc6cff46dd803e26aea0aeb0d1e79

Observation 270d6517-3512-4130-80b8-5de4d34bb86d · outbound

This paper cites Detrs with col- laborative hybrid assignments training.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Detrs with col- laborative hybrid assignments training

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.917077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.113084Z digest=sha256:cad092f6cb30e8054a33cccc9d4f8d4ee4ee316c1daeb280815d9c3c248aaae0

Observation 318e6561-5430-4add-b673-a6b1154ef611 · outbound

This paper cites Conversely, we manually constructed a Negative Lexicon, which was further enriched using the powerful LLM, GPT-4o.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Conversely, we manually constructed a Negative Lexicon, which was further enriched using the powerful LLM, GPT-4o

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.899020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.118425Z digest=sha256:f62823340f9f19f99fcec440f10789fddfc19d2164887d8cf045a330f7e0f050

Observation 73d602a7-17e8-4818-a477-2f4a5bf17aef · outbound

This paper cites Please describe the car by its color, make, model, condition, license plate (if visible), and any distinguishing features such as stickers, dents, or modifications.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Please describe the car by its color, make, model, condition, license plate (if visible), and any distinguishing features such as stickers, dents, or modifications

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.881436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.125141Z digest=sha256:bc8fee25341a58265e52decc0062d8c4bc58eaa25d48fccbe7d023d1a561e269

Observation 231bd2a2-cb13-47c6-85f5-b7a4a6953f5b · outbound

This paper cites Please describe this video in one sentence, no more than 20 words.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Please describe this video in one sentence, no more than 20 words

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.866461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.130340Z digest=sha256:35f8bb9c9b26c184d1eb50f1f3c51be7884492929cf71fe4b9c52f2ab66abff2

Observation 57eef00a-8192-48ac-9c89-8e9d1f2d0758 · outbound

This paper cites To pro- vide more precise instructions to LLMs, we meticulously designed multiple examples as part of the CoT, which are fed into the LLMs.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption To pro- vide more precise instructions to LLMs, we meticulously designed multiple examples as part of the CoT, which are fed into the LLMs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.851374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.135176Z digest=sha256:34083c64a2137fc7675d21ce18922809959fa0c2776462a08f8ca6a939e390bb

Observation 05cacb38-087a-4f84-9b17-3324de6d7a87 · outbound

This paper cites subject" +.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption subject" +

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.835383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.141222Z digest=sha256:2792fe6cfe79748893e661db13846ff7ddfaad3ae369f84c619eabdc07b68df2

Observation 88254254-5fef-4b45-a053-24043bc1c34b · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.818327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.146262Z digest=sha256:d2361b37301b28f742e63fa576fa70fc33cda397f97bbed84e71440336de1e83

Observation 37a2c3dc-17c1-47f4-ac61-a21a0412bb15 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.802974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.151056Z digest=sha256:bc9b207668d89f6f6a01f640984c1125488c0c0597b92c570e019b48cdef99ae

Observation 86159046-675d-4300-944e-ae7f004f9195 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.786731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.156056Z digest=sha256:b9b141c4cf79dc66a8ec763841b3d18ddea240b6aab547ea72f2ee51dd997b16

Observation 4b89d5ee-6326-4aa5-ad58-fce3b1848ceb · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.767594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.160851Z digest=sha256:b7648f566c1e774097ba67f591f6da8db3bbf52877b1adead973cf5efabcc4d0

Observation 09805f3b-88cf-4c9d-ad31-12813f0b514a · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.753427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.167400Z digest=sha256:ab7ca68813eb4b494aa7d036015c45d19ec60c56d70a8181d57d134f24259aba

Observation 68c54bb7-34e2-47d1-9e7b-eb28801a1685 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.738613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.172740Z digest=sha256:3c1dca8a960bf5de17cf7b1253d5d9bd2296cef9635c970bf4ae28d7afcaccc9

Observation 2733460a-5410-4e53-a018-d61641052d41 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.723950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.177978Z digest=sha256:33b3434e5f24bfa6003d6b1b739752e2fcaf9aeecc72bada07bc36629db87299

Observation 97ad5e53-6765-42f1-98b7-882556d71f13 · outbound

This paper cites an unresolved cited work.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:52.709918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.183120Z digest=sha256:2ab59cab0e79d92933a0141fda368cbe49c7b1a99a4f82096b090360f6e293c0

Observation ea38dc8d-3864-4df4-9681-ab08e95dbe5e · outbound

This paper cites Global Description.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Global Description

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.695354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.188382Z digest=sha256:00faab495760b1d230b5a57f8c89dfed429c3776fcaef9e6a99e9623ca905e8a

Observation 6b3ff845-bbdb-4d94-a251-0079f1e3e7a1 · outbound

This paper cites 3) Ex- trinsic Hallucination: Evaluate whether the text introduces content that is not present in the video.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption 3) Ex- trinsic Hallucination: Evaluate whether the text introduces content that is not present in the video

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.680131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.193343Z digest=sha256:f707ac4b2ba9f0ee20c62b91a9f29b4f053fca6a2699752bc424b6c1d5e16ec8

Observation c46daef7-2472-4254-905a-553e26f6af1c · outbound

This paper cites counter-intuitive.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption counter-intuitive

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.665365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.198272Z digest=sha256:6b29f7010e04f45084e38a7cb7ef284d4998a694b4479f06da7fda27ec05409b

Observation 02fd437d-b1f0-4537-bef6-fd4705d237b6 · outbound

This paper cites ,".join([f.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption ,".join([f

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.648749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.202993Z digest=sha256:7e432c96a755f9e0b4f63f01c33ea6478c3cd7a7fefb1ef49015491750b790dd

Observation f2da4476-454b-4771-bfa6-fb07cec18eee · outbound

This paper cites subject" +.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption subject" +

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.631507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.208917Z digest=sha256:707090f2aecb7625bbc8dafc9d555cf88114df76e65e82fa2330f0c1141ae8a9

Observation 90cdd62c-75c5-439f-9dd4-c3f4f6cc0e8f · outbound

This paper cites The video shows.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption The video shows

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.614918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.213643Z digest=sha256:9f8c14257ac0c18f4f98977041c37dd315a291de04335a01351605b29cf1db43

Observation b486091b-22ac-4c79-92e5-72752dcde0a6 · outbound

This paper cites A man... and a woman..., and a man.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption A man... and a woman..., and a man

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.596761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.218481Z digest=sha256:27e3aafc8c4f2b6f6e78b02f451c9fb15b7d258271091d52429c93fcc7bd2f73

Observation e2d51b3a-c5fb-4e9b-ba52-394fe430e6eb · outbound

This paper cites The scene is.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption The scene is

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.581473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.223004Z digest=sha256:954b836b05b6d04117e6ee887bd7d0f0c30a4388c17afc3224a492ed83d3d2b2

Observation a1f40f2a-eb5d-465d-ae52-8fc043bcb8c4 · outbound

This paper cites Aligning prompt used during alignment with the open source model.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Aligning prompt used during alignment with the open source model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:52.564812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:22.228795Z digest=sha256:327249bc6aa7a63476de6a72ac0100bb631472cd13e818fa884b25ddfa3ed36d

Pith citing papers

No inbound Pith citation observations are available.