Pith. sign in

Paper Citation Record · LEDGER

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective

As of 19 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 1 inbound Pith citation observation for arXiv:2411.14062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14062 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:39:32.833572Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:50:43.512312Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:50:46.335683Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6c7de02-b8f7-4689-8b48-9fcb1a5c0988 · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.349613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.349613Z digest=sha256:7c2bdfcbd1959a8ea3fc08a1cc9c30b94a51e74937ac6759409fec92ccb5ec0f

Observation b48fb7d6-b12b-4d2b-a6e1-c079d9f10e26 · outbound

This paper cites Lawrence Zitnick, Dhruv Batra, and Devi Parikh.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Lawrence Zitnick, Dhruv Batra, and Devi Parikh

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.355907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.355907Z digest=sha256:674473a5d3665387759b596beb93374410b13e0ff54c9ce8772ee2da44b4d8b1

Observation 6b13b67a-2826-4d42-a277-939c243e4e08 · outbound

This paper cites Pixtral 12b, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Pixtral 12b, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.361140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.361140Z digest=sha256:56653b45168e1ec1b39460ce1a0b5850f25d36c4f62423325521f7fd57b1b929

Observation 88503e0e-eb56-4065-834c-333684f824a2 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.365700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.365700Z digest=sha256:64db2addda5f300c4dac3a7893bed82e39af025ee603f121da97b3ab835be8bc

Observation 620f42f1-e921-4387-b5af-0bea0aa64520 · outbound

This paper cites Unicom: Universal and compact representation learning for image re- trieval, 2023.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unicom: Universal and compact representation learning for image re- trieval, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.370491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.370491Z digest=sha256:f6a09a002fc9b59ae70871e509a0ad27ceeca1d7e5ba173b60662791e058cddb

Observation 392d676f-d717-43b3-a963-6a4e784f0ad4 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.375751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.375751Z digest=sha256:6101a32b963271b68993e50aa43091537eddbf011957a21a446c40e87b52dbd3

Observation 5a3cb3c3-c7b5-4b63-8759-2e2c9483f6ca · outbound

This paper cites Benchmarking foundation models with language- model-as-an-examiner.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Benchmarking foundation models with language- model-as-an-examiner

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.382279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.382279Z digest=sha256:9dae68097372e326b7158a4cf19dccd5e278e01599a9b5d313d0b57629858fd0

Observation 3a569c53-eca4-41cc-aed1-e0e13b528f6a · outbound

This paper cites AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.387555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.387555Z digest=sha256:03e8c1c251f6d7aef2514267a3933a8726fc2727c3860a5266972b713cf57655

Observation 34af97cc-c653-4fc4-9f3e-53580ba164c6 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.392716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.392716Z digest=sha256:a86878729369b3cc624dfeefa53f36d809402d3e02046e5b4730767674011c0a

Observation 4d426660-b006-4629-9d50-508e9e74dc1b · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.397796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.397796Z digest=sha256:55cd351768dd7265f4933a5570ba27e61fff16bb03eb86b17057513a3055f8bb

Observation 01337695-0e8e-4add-bfb0-1d19d00479f6 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.402784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.402784Z digest=sha256:23f028bfc21be57aa5421ce42bd8712290e5e82f78ac4cec979fe73ea4f6ef0a

Observation e25963b0-5c21-4047-95e0-4209caa29209 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.408242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.408242Z digest=sha256:617f710b834e9447789a9e3e27a49db15adeaadc9fcc689dae87f4613f9557ec

Observation 89cf1586-6f20-4a4f-a2c2-8eafcca0e571 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Opencompass: A universal evaluation platform for foundation models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.413729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.413729Z digest=sha256:41bed2d531ad82b562242993542f78e1118d4770201824d09693779de45fbf42

Observation 1761d266-6cdf-430a-90a2-27df1ace2ead · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.418920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.418920Z digest=sha256:25cd0b68758bb65ac5dc04822cd09974a29c92d2220f4e9a64515c7c377c7d9d

Observation d8c43b71-31f3-4880-a221-e0fbd37bd55e · outbound

This paper cites Vlmevalkit: An open- source toolkit for evaluating large multi-modality models,.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Vlmevalkit: An open- source toolkit for evaluating large multi-modality models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.424755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.424755Z digest=sha256:d9d6ab872a4d6e08969f42ae505b88519e1f1632bcbb4a5b28f5ac19a46cf56e

Observation 1b4b3530-0bfb-4f77-ae21-051394a1d927 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.429872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.429872Z digest=sha256:a0b7e85e933c34ddcaf888fe120648864fe74a0279185541e1173c0030729efd

Observation e2144318-a87c-4cf8-af47-a01a5c975885 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.436371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.436371Z digest=sha256:b7db85eff1db0f87203378a6789f77a903733bf905edcaaea5cf6cef17522816

Observation 043b6139-149a-47eb-acba-cb8576dc294f · outbound

This paper cites Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.441638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.441638Z digest=sha256:4aa51403ffb706908be8ca155f4712ea0d4c791e74e4202ba9aa4f98fa057570

Observation 4a7cc438-a291-4348-a1f4-bfaaa58aa2c2 · outbound

This paper cites Smith, Wei-Chiu Ma, and Ranjay Krishna.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Smith, Wei-Chiu Ma, and Ranjay Krishna

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.447691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.447691Z digest=sha256:324553ff7016864e77c05406c171e9cdbe04809bdff3d9ea6c216a68f818e07a

Observation bf3e3257-aaa7-48c2-a71c-89d72ecb73bd · outbound

This paper cites Lumina-t2x: Transforming text into any modality, resolution, and dura- tion via flow-based large diffusion transformers, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Lumina-t2x: Transforming text into any modality, resolution, and dura- tion via flow-based large diffusion transformers, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.453204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.453204Z digest=sha256:4ca45686672b073670c9bae115d50f6ae734121bda9b674b54fb45d57f79626c

Observation c5ece020-f293-4eae-85a2-43dcce0cec50 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.457788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.457788Z digest=sha256:64c16918105671273c6c4c394376704e7b386bb8f58861f1e5671c18c6870c38

Observation c78fc9af-9c35-49fb-9379-e274a6031276 · outbound

This paper cites Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.462845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.462845Z digest=sha256:de35d5202d33cf0c65dd8895a3877369b9c7dfbc2e9964faf59ac8f8df350319

Observation 24d9c418-a3c1-4770-8901-12ff2e6cc96d · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.467716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.467716Z digest=sha256:3db22e4c52a3dc4334eae49bc61e46863b8383636a435226c5963dca4fcdfa20

Observation e87e963a-4527-4e4c-807f-9faec4e66514 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.473356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.473356Z digest=sha256:627a1875333d2655b7df8acd9ce16622cf744959174c2210d4d280b4e3b77912

Observation c52fba38-7d99-4a5c-a8d5-d72edc5df881 · outbound

This paper cites Denoising dif- fusion probabilistic models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Denoising dif- fusion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.478293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.478293Z digest=sha256:197bb8c2f793a35a6d21133f6087d1ff2325e8cdfa5d285c8a34a10c5df5e69e

Observation 65865f0e-ac7c-437a-ad06-8ba98191d52c · outbound

This paper cites Cogvlm2: Visual language models for image and video understanding, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Cogvlm2: Visual language models for image and video understanding, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.276710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.482907Z digest=sha256:e5b983231b3a6a49c79201085145b44c29b759d8704c2a0ca039cb65f99a8237

Observation 928cb6f8-4a69-42c3-beb4-245f80d879fc · outbound

This paper cites Chatgpt for shaping the future of 9 dentistry: the potential of multi-modal large language model.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Chatgpt for shaping the future of 9 dentistry: the potential of multi-modal large language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.260258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.487705Z digest=sha256:1e79126d4e269870ee4b3dfed44931935765e5654a9fb311b686695b650d7b94

Observation b98100ee-c7ff-44be-b245-1c10baff7f79 · outbound

This paper cites Genmac: Compositional text-to-video generation with multi-agent collaboration, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Genmac: Compositional text-to-video generation with multi-agent collaboration, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.242519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.492378Z digest=sha256:f4f102dc40c83e42cdfc25bac9edfec05df5c85158854c3cb088ac8376f3c484

Observation 998299a5-2518-461b-bb68-19aa7336bed3 · outbound

This paper cites Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight mllms via complementary image pyramid, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight mllms via complementary image pyramid, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.225390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.497759Z digest=sha256:7233ec2c7cc4c1c18a7ee5089d14979860598309c31c160cc388a475d34e59f3

Observation 994807de-2d76-408d-b19b-83b065a3f60f · outbound

This paper cites Hudson and Christopher D.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Hudson and Christopher D

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.502293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.502293Z digest=sha256:e5b815b93dee14065fc8c328ee576fc90f14b7417452e6596f3f0847716ec961

Observation ded9f1a8-8f81-4d6a-955b-a09e68a7221a · outbound

This paper cites Ku, Qian Liu, and Wenhu Chen.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Ku, Qian Liu, and Wenhu Chen

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.198095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.507414Z digest=sha256:8d8c90dbc17258aca23add9901a33e0b65cd4012bfb49b4e1ec8f11c193055bc

Observation 499efcef-457d-4012-b8ca-bba28fb504b1 · outbound

This paper cites Chatgpt for good? on opportuni- ties and challenges of large language models for education.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Chatgpt for good? on opportuni- ties and challenges of large language models for education

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.182486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.512552Z digest=sha256:7cd47e15ae0344bef15c1306f0f0d3ee704c917e542bb31209bd71b712a17cff

Observation 4b9286f8-31cc-40c9-842e-31262017dad3 · outbound

This paper cites Reflective decoding network for image captioning.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Reflective decoding network for image captioning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.517611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.517611Z digest=sha256:15a696a4e53071a987770a726e6764621b7ecdeddec5abd725cb06c2d745a692

Observation 93d3472c-f64a-45aa-9faf-dba61c362afb · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:34.156297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.522155Z digest=sha256:96a96d0fbcd8c7cd36153b068fb63b3d82e287c3d05a2d09388ba5571aae0c4d

Observation d4e2799c-b6c2-4394-915b-400f3e2f62d7 · outbound

This paper cites Building and better understanding vision- language models: insights and future directions., 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Building and better understanding vision- language models: insights and future directions., 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.140216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.527662Z digest=sha256:68d290fe76ee920d3824f84d38e520d41bd77b2a0e7afb46a5137f2452d90d46

Observation b745b877-e7ee-44cc-9cd5-20d6df8ef4ca · outbound

This paper cites What matters when building vision-language models?,.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective What matters when building vision-language models?,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.533626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.533626Z digest=sha256:27fc813af8bfd90cbce8fd8fca5f244a5d650752508bb74db7c816c24a4c0c10

Observation abfeeb7c-77c2-49e9-9b15-1d47c90208bb · outbound

This paper cites Llava-onevision: Easy visual task transfer, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Llava-onevision: Easy visual task transfer, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:34.113810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.538387Z digest=sha256:00032f99badab8e482f754df3e3f5c50fdd8cc78a08ea0755a24a3d224aaac14

Observation e4660870-ddae-4088-a8f4-b2bcb5c4e527 · outbound

This paper cites AutoBencher: Towards Declarative Benchmark Construction.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective AutoBencher: Towards Declarative Benchmark Construction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.543675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.543675Z digest=sha256:bdf2c6762350bf9299d7f7beeab8941404952e3c2d3a40d76bc39b6877c6ba5e

Observation 603b06ca-39c5-4251-94fd-94203ff1d0cb · outbound

This paper cites Llm-grounded video diffusion models, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Llm-grounded video diffusion models, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.975313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.549335Z digest=sha256:0ea78da31f35af7404fde44282d2af787313675fbaec380da1a5ae4aad16a3b3

Observation 894a6aa1-2230-4cdf-a99e-f594e9c768f1 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Vila: On pre-training for visual language models, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.959405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.554553Z digest=sha256:25f462ac8047564b6f19ded8d31759b9f984d74efc40b49c8ad8a9ab55d251ee

Observation 97984e33-8afc-4b0f-b019-ec6c0f30fff3 · outbound

This paper cites Microsoft coco: Common objects in context.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Microsoft coco: Common objects in context

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.944293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.559117Z digest=sha256:d69facea79bc512b3c234e192d1e7d0842182d4cbd57c344194531cc71d34483

Observation d8f55a43-6f18-48cd-87bc-5e13ab9587a1 · outbound

This paper cites Improved baselines with visual instruction tuning.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Improved baselines with visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.563553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.563553Z digest=sha256:59455a4f2eb01020e33cd9ae2e6e4078031088265cd60dd3681407cdc1071046

Observation efbb7fb3-2d82-47db-beac-e355f76519c8 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.916778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.569554Z digest=sha256:bce56792b5bed976e81692b831d49f276cd33e584ef93a877d5ef196617ba78a

Observation 6ec08b7a-98d7-4311-a264-2f5311eaa607 · outbound

This paper cites Tempcom- pass: Do video llms really understand videos?, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Tempcom- pass: Do video llms really understand videos?, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.899473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.575050Z digest=sha256:8bfe0bf6eab1c82138e5c7b815107efc55667c0ed20da037d126aebdf0160a99

Observation 72df9ccf-b402-4686-831d-2b7fa24dd0f9 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.579767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.579767Z digest=sha256:201912194e57f8a153ab5748432eb8704c52024ee6a4efaf2149b86c9d97f3e6

Observation a0b82f83-739c-4368-befa-a19ba47778dc · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.872275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.585731Z digest=sha256:982a4f7025d5e264741e692b195533c9f32192c09213abd12b692918765556f2

Observation de84a044-9868-4e8b-85c4-d245ff68e623 · outbound

This paper cites Mmdu: A multi-turn multi-image dia- log understanding benchmark and instruction-tuning dataset for lvlms, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mmdu: A multi-turn multi-image dia- log understanding benchmark and instruction-tuning dataset for lvlms, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.855859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.591250Z digest=sha256:9ed9cf22256a7daeaebb89807f330bb4d869125bdd2d409e9cff36605f6605a1

Observation 6b2a2f7e-e0dc-4e9b-830f-9fc7de4a394c · outbound

This paper cites Mmalaya2.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mmalaya2

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.839643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.595832Z digest=sha256:76c061bba29149add5d5fb693f4d15a36002692491436d8c65626c0fa894435f

Observation 19f4c7ed-d0e8-4f7f-a8f7-e97484ad952e · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.822483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.600831Z digest=sha256:d4004a074fecc088c832acdc63779deae778779f83c0c56770a423553ae4f7fc

Observation 1fefffe4-841f-4476-ad1b-aab72a973087 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.605366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.605366Z digest=sha256:33be63ee5bf393d8fae8a3ff9b5fa253b9e5efaa4658082304ab87cdb16beaac

Observation dad6b72c-ab48-4ee2-b150-6c4992ebac98 · outbound

This paper cites Mmlongbench-doc: Bench- marking long-context document understanding with visual- izations, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mmlongbench-doc: Bench- marking long-context document understanding with visual- izations, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.806687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.610774Z digest=sha256:e1f3ab3ed1f9a32270211a3166cdd329befa88a13f8baa0b6c942e37941817a5

Observation 99adf939-6ebf-4b71-abd2-739a3ea8e914 · outbound

This paper cites The llama 3 herd of models, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective The llama 3 herd of models, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.791517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.615244Z digest=sha256:3565ffb1f0c4b65748f448df094817646ccd15a35164e8d18b7b22f95634f830

Observation dc7f92d1-c0ea-4920-86d2-78b33d333f96 · outbound

This paper cites Gpt-4o system card, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Gpt-4o system card, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.775496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.620381Z digest=sha256:faa417710fe11cb66bdcf8256a3cb55d53623bfa84ee06f7f0ea435d790693d2

Observation f6c63cb7-7aef-443b-a138-e5014b289879 · outbound

This paper cites GPT-4 Technical Report.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective GPT-4 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.625018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.625018Z digest=sha256:d7ed2c45346ac922a39bc023821b116d6b0975e54c7df60882de6985a045b476

Observation eec231ab-3931-4a8f-8ee7-b49b1d93bb09 · outbound

This paper cites Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annota- tions, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annota- tions, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.760173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.629879Z digest=sha256:e46cfb10939bf814e3c969c43836f18fdfb137544353e2cc018ccfca9decaa17

Observation 451e32c6-30bd-46f2-af7d-313d7bfb8822 · outbound

This paper cites Sowing information: Cultivating con- textual coherence with mllms in image generation, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Sowing information: Cultivating con- textual coherence with mllms in image generation, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.743326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.634453Z digest=sha256:a5f4f4c95773390683632d4be014b317ac64d4939562851c129e660cefd2e295

Observation af72decb-b80c-483c-93b6-e71f39f199fd · outbound

This paper cites Rbdash-v1.2-72b.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Rbdash-v1.2-72b

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.726582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.639736Z digest=sha256:bdb16a88e1fd70afa69b3d4ca85de939fe07cccab66fca69728e5c603a206b49

Observation 0e3ad64d-18be-49fe-8188-374e0702de0b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective High-resolution image synthesis with latent diffusion models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.711065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.644263Z digest=sha256:2530451eab1c4d32629ab94fb33a43087a1828a1c78e619f5ad71c4e3141e264

Observation 2e14a600-424d-46e1-ae8d-9f53737c730d · outbound

This paper cites Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.695399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.649249Z digest=sha256:46e32793d00ce42601bc564df06ea64edeb3639e3f7f72e58ed334f15a5cace7

Observation c4f332b9-96ce-415e-8db0-7f0afaeb2b0d · outbound

This paper cites Journeydb: A benchmark for generative im- age understanding.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Journeydb: A benchmark for generative im- age understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.678629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.653934Z digest=sha256:a429a9fd533f9a1c8d4d972c834523c1b9ae845a693473fe658495b1d0d4c4f3

Observation b7d4203c-012c-4d6f-a4ce-96b627f69fb4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Gemini: A Family of Highly Capable Multimodal Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.658812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.658812Z digest=sha256:9ee5f52b2f8032544b69bc5de0c9ccf25f6cc30ed1200af595fe886e1ba5bf29

Observation 7912cc58-07eb-4614-aa15-2a5fd20f5c43 · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.663550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.663550Z digest=sha256:d0bf1b9d703a0b2bb0e711e0f5455aaa4f86611ad190422b9fa6aa5864d868e5

Observation ca0b0a84-17aa-422e-9259-f029bb20aeaa · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.668528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.668528Z digest=sha256:1608a7370b89a3c2b367e38701c07b286c1932501f7a5247aec5ce3e0a10ac7c

Observation e0680d9d-1c65-4540-8eff-861fabde1c63 · outbound

This paper cites Llama-3-mixsensev1 1.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Llama-3-mixsensev1 1

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.652759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.673733Z digest=sha256:39e2634e9ceba6374d62e10a7b71041ec1cbd677fcf77edbb47390de524e2ac6

Observation b9e32ecc-4031-4d02-b48e-45a77f86cbea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.678916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.678916Z digest=sha256:c016811c52d2e9b78392b6dbeb86823461e22fccda09b97da998f6e40971cc9f

Observation 77e1f016-8d2a-4f80-bf98-81ea93970981 · outbound

This paper cites Large-scale multi-modal pre-trained models: A comprehensive survey.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Large-scale multi-modal pre-trained models: A comprehensive survey

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.638570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.683799Z digest=sha256:e04298068c3faff6163378ecaadb6308e91ae326dab957289270f7ed430f9bbd

Observation d911fd21-c8b9-428b-b48d-0bd46701a678 · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.623022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.688773Z digest=sha256:aa6d3800c0babe288075f761d045e8ebc982e879eb96cb9db73450ebdb63bdd0

Observation 7530f209-eea4-4070-a986-39e837a3f2e5 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective BloombergGPT: A Large Language Model for Finance

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.693315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.693315Z digest=sha256:a6e608d09d70e34c99e77f273e791bd821f4e54d03f5517228835d036c9910d8

Observation 10ba1df2-2978-4223-94fa-ad2afe866761 · outbound

This paper cites Unigen: A unified framework for textual dataset generation using large language models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unigen: A unified framework for textual dataset generation using large language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.698376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.698376Z digest=sha256:847784592bea6d03b2958410efc043de8c35f602d8a0aeb49ea5d7ad9fee1435

Observation ec0b6ef1-fb73-4fd3-8e0b-71d15839f9ec · outbound

This paper cites Self-correcting llm-controlled diffu- sion models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Self-correcting llm-controlled diffu- sion models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.607735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.702934Z digest=sha256:263f75aace21ffeac697c194d4de555a65c4aa33458deeca894f464259f1aac0

Observation 90e9bae8-6ace-4529-8971-259e9c382a40 · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.592508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.707386Z digest=sha256:e85c3a34302a750587169df6bc0debdbbde8e411068838663589baf998e73890

Observation b007837a-dc2f-44b1-b02d-f0988bbd5bd7 · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.711740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.711740Z digest=sha256:bb083f90ce8a28069469e2a4f11f0cbb9410fdda0638391fc2e4572882dcc91c

Observation 1017213b-7fe3-4803-b0f4-1117cc6fb4bb · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective xgen-mm (blip-3): A family of open large multimodal models, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.576802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.716481Z digest=sha256:cd7335d2ac2ad84671694d0b1fd80b652e164696823cb498eb2504d31e9afb85

Observation 2eff9fe7-f24f-4fac-a20e-ae515ea03c3c · outbound

This paper cites Cc-ocr: A comprehensive and challenging ocr benchmark for evalu- ating large multimodal models in literacy, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Cc-ocr: A comprehensive and challenging ocr benchmark for evalu- ating large multimodal models in literacy, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.561138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.721684Z digest=sha256:cb9f28b13720277d27d400120bb3984d402fb342b3c1aebe75157204d93ee5e4

Observation 5a146a39-7414-4757-a6de-acab5baa365d · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.726534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.726534Z digest=sha256:caa0143d4afea949a2e31859b755e9cac69d8b02d4d7912ead72978c0a9fa5bd

Observation 7df529b8-8387-446e-9b09-be541edef8e3 · outbound

This paper cites Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.545432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.731371Z digest=sha256:d3c0682a9e1df0a3af1cb1d154de52e5828593964c5eefa8b101291b6e300b6b

Observation 6aa2bdd9-6a97-435c-b870-3ba4ce4dd882 · outbound

This paper cites Benchmarking chinese text recognition: Datasets, baselines, and an empirical study, 2022.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Benchmarking chinese text recognition: Datasets, baselines, and an empirical study, 2022

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.529052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.735935Z digest=sha256:dd6d5fdc7b581e57482ddc279a00fbbd79507e48386612d40478e806b2e5833b

Observation 695ee012-2c3b-46ae-bbba-e70926fb0b97 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.740916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.740916Z digest=sha256:d51f82be77b57de88da4e894ce40969373826baeb31cf798668a20cd01f89854

Observation ec4128b2-cdd5-4f5d-b7a5-01df68b31685 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.512768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.745577Z digest=sha256:7a37832a60d7db7e22fe484a93fcf3ffe98a58d4251010f20af2a67d608b7547

Observation d00654b8-e5bc-4326-a09e-6b903bbc1853 · outbound

This paper cites Task Me Anything.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Task Me Anything

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.750706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.750706Z digest=sha256:3ccda67a063bf665375c690d1a0f9e25c702456ae71c0ba3da0378dca302bf53

Observation 3e43a91c-b0d8-44d8-b76c-198276ae641d · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T15:39:32.756471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:39:32.756471Z digest=sha256:4bc11d77a9ca3d769d3b859ff566433d1ff3210ce826fbf35ef0eff4b29e3fb4

Observation b12178ac-ae59-4a29-a9c0-45ea07ced1ab · outbound

This paper cites Omchat: A recipe to train multimodal language models with strong long context and video under- standing, 2024.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Omchat: A recipe to train multimodal language models with strong long context and video under- standing, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.496825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.761214Z digest=sha256:c4d13b9e1850765632d2bf1e99f96b451e54103f6b47da11864e2dc7f9e9f4b2

Observation 254e35a6-d337-45af-8368-ab56ccc548e9 · outbound

This paper cites Dyval: Dynamic evalua- tion of large language models for reasoning tasks.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Dyval: Dynamic evalua- tion of large language models for reasoning tasks

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.480755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.765818Z digest=sha256:57da872ba5dba138a7b61782f2fac4f55b7f4aa453867399dc07da428340316e

Observation 5a8b3be0-124b-4f00-980f-eceea2f9ad56 · outbound

This paper cites role”, “definition.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective role”, “definition

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.463792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.770057Z digest=sha256:0be1de826fd41329648f375a2fd060f69291dbdc291c908b0d45c721706dec28

Observation 2a295e12-c4d6-4c4a-aec9-35beabeab30d · outbound

This paper cites image pattern.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective image pattern

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.430819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.780399Z digest=sha256:381e896c2deae66c3fee2a83a376498cc77838db3ad7c29f3ce8de5920c9a6b6

Observation 5f766ab4-efdc-4edb-b508-1d5d5eea9f56 · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.415152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.784721Z digest=sha256:032be2a2980473107d50a51a0cbb864720332a358b645457629682e328f46c20

Observation 9fe9aa89-99dd-4282-9ab5-ad91dd43825c · outbound

This paper cites Surreal”: 2262, “Lighting.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Surreal”: 2262, “Lighting

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.397689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.789967Z digest=sha256:3a32d2a87fa6bb7d4c3118351a732e489b97482ddd833db4d588cc7ac3aad4c1

Observation 165296a6-f543-4b96-9ec8-fe604d315a75 · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.447604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.794248Z digest=sha256:28967b4ba2d06e27fb8cbc38e85722e0e232bf3736613b05c47f022db714d5de

Observation e18433df-7137-4cd0-9fe9-44f9097ecb03 · outbound

This paper cites # Key Points.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective # Key Points

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.382471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.798628Z digest=sha256:63449eafcca23beb09876e6de3e9e97fd2e5cca3261af93a574d0a2e21e7ed23

Observation 7d7f9315-0a39-49e0-91fc-20d079ce331d · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.367650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.804235Z digest=sha256:3e45528ff483c3bfd94a1ce815e80714b260f1dfa3b932867cfe77d46c28f8cc

Observation ff20a192-012f-46d2-9461-884fdf8d3c61 · outbound

This paper cites You may annotate multiple patterns as appropriate.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective You may annotate multiple patterns as appropriate

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.351417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.809587Z digest=sha256:cb53d07eb9e72d09c4c46e9a6311afc6b82358450e3eb4a99cd554e721b7b0f2

Observation cb493ce5-771e-4a9d-a2a8-cd789f8efacb · outbound

This paper cites Surreal”: “This pattern is characterized by its prevalence in depicting scenes that mix elements of fantasy with reality, often creating imaginative or dream-like visuals.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Surreal”: “This pattern is characterized by its prevalence in depicting scenes that mix elements of fantasy with reality, often creating imaginative or dream-like visuals

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.334412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.814633Z digest=sha256:383430a155a14d5da4c4a9712f5ea8dad111a1538055c60abecb349d7d039e05

Observation 8a05b0ac-0152-48d7-9186-ce600dc46680 · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.317571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.819720Z digest=sha256:b93f4783c7363fb4270df094eb4cfe3ae1843a7560756f0217127ae61b59d660

Observation 35228db3-7b17-4e64-868b-e644dd2c0fb4 · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.301575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.824272Z digest=sha256:9d5d84b450388afdfe6eb313f720ef43b599b024c95726c396e33140d7ad3eb9

Observation 5da42579-5ce4-4599-b205-0e82eb4f5245 · outbound

This paper cites an unresolved cited work.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:39:33.286138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.829093Z digest=sha256:ef01d21af88336621f588b087a0f5a0b36358f27b3434449fb4076ad12f262ed

Observation a416d5f0-2a35-4b0e-946e-751778cb5ece · outbound

This paper cites BHNORAK TOP.

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective BHNORAK TOP

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:39:33.269335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T15:39:32.833572Z digest=sha256:7e7efd36ec62b73cf5819caaaf3f8c9761c8b8786d462ac7f8800c71b36add9b

Pith citing papers

Observation 17068afa-0c71-4cec-9951-a0b0225303ec · inbound

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models cites this paper.

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:50:46.486732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:50:43.512312Z digest=sha256:57e13f72552f94fe62e2ef0fd8e7bbbb2f79a80459c8941699ff64cc9fb2d7a0