Pith. sign in

Paper Citation Record · LEDGER

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark

As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 2 inbound Pith citation observations for arXiv:2501.16688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16688 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:26:29.251213Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:40:53.903009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:00:50.182088Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56dee464-c6dc-4e45-94ae-38a77f8ab241 · outbound

This paper cites GPT-4 Technical Report.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.095678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.095678Z digest=sha256:9441175f8ac6935756c38350b3f4aace46ceb6906ac5a86f88fc2c675cb2935f

Observation 0ee1f056-df51-40de-a25d-1334ec1bf855 · outbound

This paper cites Qwen Technical Report.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.106417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.106417Z digest=sha256:4818a98763206e2f5f3884a4db6e005701b76d52bdcdf9ba60be4ee94df2636c

Observation 7541bcbd-747b-4a1e-949a-f766a4eb5ec7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.111354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.111354Z digest=sha256:d4905676d9ec9c4c1b97575c43577dd17b95f51c5eec79dde61fff0eabf6ad61

Observation 6a476e83-9698-48be-bf23-0ffde00e9e27 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.116371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.116371Z digest=sha256:190472954e623042c4a7222c75453038d6d43258e453156147dd4e36ce065c20

Observation 082f44b9-beeb-4f8b-adc0-9a808a21cd22 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.130502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.130502Z digest=sha256:fbf9c0fa2f671c6c088532dda6e6e1bb1dcd9a79258c8745b9498da6171f9a88

Observation 7b7b239c-115f-466d-b19a-0da33d7cbb05 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.135777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.135777Z digest=sha256:7762978222ad0483f0eec9efc4a6f00ae6f9a4d14398f2affe6932b8e57cabcb

Observation e58c7726-ab28-4d01-867d-e240faf4511c · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.746437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.140299Z digest=sha256:7c4d1909f0c98cfb17a9028f70ce1d61d1d4745ab35c55b3d756a0fc9d499e45

Observation f59c7468-ccd1-43b9-b979-c4bd9fb569cb · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.144676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.144676Z digest=sha256:78a2d5a2395cc44e55ef611955f54cf17818be4e37e0d7bf8147693956297f44

Observation 71da6fb4-77e6-47e2-9fa3-fe8291ac4b4e · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.149135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.149135Z digest=sha256:56ebfb09d4f890728846214844fc111ba2a7305b42460218bbfb8124c20d15df

Observation 82b174a3-5d82-4411-a330-5a0c2363df82 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.153567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.153567Z digest=sha256:c9083f1562a1ec27a2bec738ad7eaf5a5617ce99b425706a4b046768ff3bcf3c

Observation da99e05f-54cd-4157-bac8-f596a5559318 · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.717602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.161982Z digest=sha256:9cf9858d52e2e3ebeb2b610c082c2b2734b43a0c0859f503bc1bfb3d13e3ec3d

Observation a9ae5d96-87d5-4fa8-915e-3fcf61f6979f · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Monkey: Image resolution and text label are important things for large multi-modal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.704491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.165933Z digest=sha256:6f775e51c430d51621cb9775cf0934d5dabc253af6e4e45d1bfb4e655567f25c

Observation 945af6d2-102f-47e0-84a2-4a0a090dafa8 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark MMBench: Is Your Multi-modal Model an All-around Player?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.169977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.169977Z digest=sha256:779516699cb1a0778bc862eb6a164cb18bee1f99487a1ccfd4b439818df308a9

Observation e83c63da-d2a5-43fd-98e7-9449af5fd5e3 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.174488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.174488Z digest=sha256:378d24572ef4e6551f5e95bcd691e95b15fd5d023e2ef0d9ba73b56b417de964

Observation 7b675fb5-e226-4e93-853a-ee05c37b0b8b · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.178954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.178954Z digest=sha256:7d90f2522441948b653e7057cf66ab963b8d6aa8fe9bd7ef3e50bdb15f44fb60

Observation 07d84a8d-8c82-4c1e-9f4e-114d04f26442 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.183284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.183284Z digest=sha256:b5671b705dbe513a3284accbb5ef8c4af29976b38ad200606c61ae4ab2532dca

Observation 40f43e90-910f-4e14-881c-eb6eff3996f9 · outbound

This paper cites Training language models to follow instruc- tions with human feedback.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Training language models to follow instruc- tions with human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.190299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.190299Z digest=sha256:7f13229bb4224c36316fdbc11b4fffa27948490cc0456206bf094f47fe49a4ad

Observation 6b632a4b-1418-4641-be56-30d1e43b18a5 · outbound

This paper cites Multilayer perceptron (mlp).

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Multilayer perceptron (mlp)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.672920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.193733Z digest=sha256:c4898c65fffb02d1bb31dff75988826280a8d9505da6c5e6dcaad1f12136a843

Observation 519b8997-851e-41ea-8111-42621960c03b · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities,.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Internlm: A multilingual language model with progressively enhanced capabilities,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.659920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.201842Z digest=sha256:726805ed17bf8224a11fc0391a07a09bc42366c15fcb19767c748f58e0ae4897

Observation 38df5ea7-3fb6-450b-ad8c-f086bdfc7e49 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.205860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.205860Z digest=sha256:4dcff6df746a208029e6d22b3148e0b29b3f3970bb32dab27b325470f5878598

Observation e5d65e6f-ce52-4940-894f-d073327309c7 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.210443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.210443Z digest=sha256:e51a344ab1ebf218392d5df40d2a10c74abad4debd698bee9286d887306d244c

Observation 50448d07-38c6-45dc-997b-dde94c40c5b2 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark CogVLM: Visual Expert for Pretrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.215489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.215489Z digest=sha256:acc506802ea50226e43f28653f240daf63027df573fbb542a20b27add87cae22

Observation b88beee7-c818-4108-82b2-687df36d5a08 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Baichuan 2: Open Large-scale Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.219967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.219967Z digest=sha256:419646c7159023f806705881ee3b8d5c67c0da58c3bfb6f29cda206f812c1e93

Observation 37d07d4e-13e8-45f8-ab6c-0509483a77fb · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.224231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.224231Z digest=sha256:72150553179e8e3cccf9533d8ab5ad906c4b4375d65f981912c9481d7f86fc0c

Observation 3093fe24-26ce-4dd0-8863-13feed24e6dc · outbound

This paper cites mplug-owl3: Towards long image-sequence under- standing in multi-modal large language models,.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark mplug-owl3: Towards long image-sequence under- standing in multi-modal large language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.646159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.228675Z digest=sha256:96db758be0e19bfd677c97f0d3347bec61722daa74cfc52f589e80d71b36b27a

Observation 0941ea37-996e-46b0-a446-d7a3510215e6 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.232724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.232724Z digest=sha256:caa448d3fd14c7236ff9a7bdbff0b1c7ee7f3647e078ffd36e119e1035c533fa

Observation 05f96119-e74e-4678-ae06-f8e452a0beba · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.630621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.237069Z digest=sha256:234a768ae928e3ac76a1fdb6f93facd5588143bb7fb4502b59106a38c4fb98bc

Observation 8e1a6271-5d9e-4084-b77c-f2a4a9a756f8 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.241357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.241357Z digest=sha256:c640ed173458ff5941f842a27a9922c629fd2c289952954bd4c4323d89bdfdbe

Observation 4b5ce515-2087-4987-9dc4-0acbe47afee9 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.246647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.246647Z digest=sha256:472a1a9e73b3df33469f8a55680c6a78e4c8a03997bff7ac55b24c4cc44a4a11

Observation e6d42da3-5a4f-4899-80b3-07d14b2c7270 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.251213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.251213Z digest=sha256:6d0fc45b638cbefa3f88174a1b281463d8376d9f6a16689620dc25b1695cbedd

Observation 361c4051-056e-45f1-919e-07e7e8cbebaf · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Gemini: A Family of Highly Capable Multimodal Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.197255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.197255Z digest=sha256:07aa8b61d4b7d420e526155aac6acf0ac72b81708fe3d35f6d2ea97f2a4ef820

Observation de929faf-3875-4ebd-bf05-fa60ce8b9334 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text su- pervision.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Scaling up visual and vision-language representation learning with noisy text su- pervision

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.732271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.157837Z digest=sha256:600d1385f35fe3cc5828c9b7002d54c4dca033347ec8b13b4e381849921b42f6

Observation 4b5406b8-2c45-4456-89e2-507d466cc627 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Docvqa: A dataset for vqa on document images

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.187027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.187027Z digest=sha256:a6088e6914878567521c6536ee2ab0363ad2a63640e8651f41c7f48f2aeac454

Observation 9064c011-934f-4443-b298-8b06eea20d22 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.101250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.101250Z digest=sha256:66d29b44df20bd4c49ea3446b7a8530d5d46cee1eea15ce0f6d7c36308c8fb41

Observation bfdcad41-56da-4654-9e62-6125effff869 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Sharegpt4v: Improving large multi-modal models with better captions

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.776399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.121335Z digest=sha256:046812a576fe5ee17c5610498c85962bd8445cc1e62c01d519c948a9f0a6b40a

Observation 1a8370a7-1b44-4ba9-80bb-363b55e9d0bd · outbound

This paper cites Palm: Scaling lan- guage modeling with pathways.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark Palm: Scaling lan- guage modeling with pathways

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:26:29.761004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T11:26:29.125959Z digest=sha256:3f61c62966c75bc35c0429d7639e9c29256c9f026c73ee356901cbe3c302d223

Pith citing papers

Observation 54de9421-cea9-4749-b7f6-ae489d8dfa0a · inbound

CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model cites this paper.

CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:53.903009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:53.903009Z digest=sha256:58860b541b9594faa990b2ab25fffa8e29940ee190f1ec7c2f694fb3f0515fe5

Observation e28b1841-9f53-4e79-9ad8-b828921e8db5 · inbound

FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios cites this paper.

FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:50.187662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:24:23.367540Z digest=sha256:cd463e6f8b4d111a264d4df620187e143e8208756a3f116a76f4da89ce1e8c77