Pith. sign in

Paper Citation Record · LEDGER

SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2404.16790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.16790 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:09:15.070161Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T04:30:40.718071Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb1b67c9-cb8f-4cf3-9bed-c91a0b6ad9eb · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.782957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:85c369004b68d71427d4f4a3668cd1cc84a68ef3f86ba4e48196aae72d624e3b

Observation b7911f6b-4b0f-4c65-abcf-bbfa683b9ea9 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:36.852356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:36.852356Z digest=sha256:721ee9ff80fd20d62bb159a902d22e8fc0f3d663f69de4a19de9b1d9c154413b

Observation b7e5b5d7-16fe-4807-a4fc-faf20d7c2657 · inbound

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey cites this paper.

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 230

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:31.476504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:02:31.476504Z digest=sha256:0c931af7069aeefdc270d063f7300ac1738fdf0ea85d0db59ae2582b98feac37

Observation 18e991f0-6e74-440d-ac84-0659174b7709 · inbound

Detailed Object Description with Controllable Dimensions cites this paper.

Detailed Object Description with Controllable Dimensions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.966757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.966757Z digest=sha256:49cc504ea4f88febd31059f0d3c385e705d43ae543beddfb95f1a8b08d594943

Observation 99cd16a7-4900-44ee-8ea3-b69c39b6ea7c · inbound

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks cites this paper.

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:21:22.513649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:21:22.513649Z digest=sha256:b3442b09c5e3bfa0112c209fec6c8ac004473fa1da3b437faee2002e6b43af50

Observation 47953182-0762-4380-aaaf-0871ed770f91 · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:23.227643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:23.227643Z digest=sha256:1c72913f495168bf1d8b4acb96f72d6042ecb84b56cf9eeab02a92b9fae142df

Observation d48a5e20-08fe-4e3b-8a18-2b2305f3b4ce · inbound

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? cites this paper.

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:19:11.525972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:19:11.525972Z digest=sha256:27308d0855d3d04fbe7e12789354af32749d5249117b2ab8abd9605f16b2c5e0

Observation 71bc3799-1b16-4aa9-9020-0ae1e1c36bc5 · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.031765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.031765Z digest=sha256:9db53b38253f338d202e38b477722adb69189f9001385444490e64e39a05c2f5

Observation bf601f4f-cc32-4b4f-884f-34104f92a460 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.074856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:921a3ed4423f43b42da21f79171093fc2de40248532b4d41a34b0915d7643ae5

Observation 7fed4a12-5eaf-40e0-bb7f-58a14df4529b · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.706356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:10152741491390ecf678660800d876f416670276d30fc6a5505ced60236d7dd7

Observation 611c7e4c-0ec4-4ffb-99de-ecdc32dfe686 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.937654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.937654Z digest=sha256:936dd3a5cd018e33822b0c59c7605f8ecb62acf2412f1486e946515d4b18521d

Observation 6ffa1d7f-aaee-47f1-9efa-8d2f643abf52 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.337965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:061c70fb1b5e3d24de7f6cc6f4658d1a67e9428e70cc538bc7db86e5eff071f3

Observation 334ef63f-17c2-494d-8997-515465fa4dae · inbound

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding cites this paper.

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:09:15.070161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:09:15.070161Z digest=sha256:1775951b9bdc12fc3f160976f4cfd19194ed239131620340329e032a672f3b57

Observation 1a1bcee3-8c27-4888-bb25-b0af90329a1c · inbound

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling cites this paper.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.477616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.477616Z digest=sha256:f36c2ee898092bb7f877357cf966d81c01ddeddd3d4f5549c402992c55fec0c5

Observation 31e98ba9-e124-4b41-b4d7-cd1e9217cb1a · inbound

SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning cites this paper.

SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:53:21.037307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:53:21.037307Z digest=sha256:8b30bc0205b868299c5d24ec3b97cf3ece2e1326433cad6ede4398f03b895cfe

Observation bb312b7c-9db3-4237-b007-03eaee6fbdc8 · inbound

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation cites this paper.

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:50.351887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:50.351887Z digest=sha256:6f95f1c8b9cc13892331bc12bb9c24c2e7bfb41421996abf45164399545152ad

Observation 0e6110f5-c63c-467d-9361-79217ce2c16e · inbound

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? cites this paper.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.506036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.506036Z digest=sha256:e4de22b8bb1801ef1a4b51c4eb63b41a1fe63140547a4cde4cf6beec6e16590c

Observation a5bc9a1d-66e4-4615-8dc2-80065ca6c88b · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:18.053002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:18.053002Z digest=sha256:86c6b626dd17530c65e78723f1b9aa0f0eb73fc8f77e5cf898bf1ec419ff8c9d

Observation 8ebc8c25-c5e4-4cd8-9ae5-9f2d8400e533 · inbound

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models cites this paper.

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.226798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:55.226798Z digest=sha256:6f37649f4462374515107e10453a4989a39f4499e732ac5c52fa18a1bd03ec46

Observation a6a1a634-b303-4b6e-b6c7-e8b03cf126b3 · inbound

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation cites this paper.

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.130801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.130801Z digest=sha256:ee32556498dda59fddda0f0957bf4029d7a871c8eb7fa745787d507f9b50f11b

Observation a87c0ae9-e60d-4481-af81-960fba18264f · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:11.985639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:11.985639Z digest=sha256:024ea419b1cfbce0b20a76b82f2189e1d2bb1a91872ec27a5a354a9cd881ef31

Observation ab25a3d4-ae8e-46a7-9ad8-26a6507a2431 · inbound

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review cites this paper.

Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 203

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:48.672990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:48.672990Z digest=sha256:ef38225a537a6a81f3ea6c54bb599d1f7d64b0d84376cd90c45038088cd86daf

Observation f24a4638-6649-4dad-9272-65e048d80726 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.116722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.116722Z digest=sha256:41714aa3973d32b92f166fa12d6cb8b396171ee375d5c7128736c4810521d79e

Observation d0599d4f-1320-4427-957d-93ceb82ad4d3 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.185030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.185030Z digest=sha256:47b15c433f6ebc06186b6f5f97f997a9db4aaedd4d74d77e68dc4acf560782d0

Observation 3e46d437-d4fa-4afc-800a-916d564b420e · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:23.943515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:23.943515Z digest=sha256:321176a20734690a08a19e57d79c8783f4b7403440a886eb68aff24c4dd4feef

Observation 0f0d8dc1-58fc-44ab-9453-12b949975ae9 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.998784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.998784Z digest=sha256:923a41ba8215967f01695116bfeaa84df4c21aed857aab4b52997371f615293f

Observation e446c60f-acc8-4f68-9a7d-af9aa250ad29 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:08.173192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:08.173192Z digest=sha256:21a4f6c5802b143e0f15785968246ffd133272abe0b7d4bfee9ffc4f15cb608f

Observation 26a51221-4347-4ac5-88fe-8b51366e9d8b · inbound

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents cites this paper.

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:18:54.074131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:18:54.074131Z digest=sha256:64afbe77cfa7b14f7297c8b5daa160bbc312764458cbf92cb7f648b2944d9231

Observation 92bbbf6b-91b7-40b3-8fd0-8423d9c3780a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.159584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:568535b3ee421e39f440be952224981124936d92d40e3e99bc117cfb373e3ffa

Observation 93a3a951-c912-49f2-8512-585b1fc62165 · inbound

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning cites this paper.

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:29:24.990122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:29:24.990122Z digest=sha256:d93caa9930854085ffd5bc6b991c5ee943ce5bc82553256ff1d903b1c610e120

Observation a1b56a09-3792-4a48-af5c-26eeedbf0746 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.335668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:19de42abb110350fa04d0e4f7f23e5acac2fa80cbd389b6312f2b7eca50aad62

Observation 8b799e30-2925-4d39-a890-2e38cc6736f3 · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.514676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:48972baba5eb4562bc4e027ea32d7b05e702fd2ae737c102305113e9eb89faeb

Observation d2f8dcbc-db99-45ae-9962-7e9141e36ba9 · inbound

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning cites this paper.

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:01:05.178856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T15:58:03.654150Z digest=sha256:6e7f974751a179e078053c6798f85ad09592729cb3ebd678741c265ad7d0ed9a

Observation 9ed6f8f7-af64-4e22-8ed1-c90afa19f5ab · inbound

Learning More from Less: Unlocking Internal Representations for Benchmark Compression cites this paper.

Learning More from Less: Unlocking Internal Representations for Benchmark Compression SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:57.751567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:57.751567Z digest=sha256:baac86239dd801e78df832d58ac72e1019a55f4b15ddf846027c8ba5ea4cebf2

Observation 87b8e29e-dfbe-4f72-81af-26fe4762df7f · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:c2ba9ba53282c90b13dcd946d3bb6d37a2aa441559b7d34b45b8e979f73694a0

Observation e40648cf-11bb-41eb-af3e-203b45d85a7d · inbound

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models cites this paper.

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:53.805058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:35:21.514502Z digest=sha256:a61d8823aea03c8da28ebf1f73d6b7f48cc59cc0e4e9ec435bbd7edd35d9d28f

Observation 3e929891-5d7f-4ccf-9fc2-81d28e7affe8 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.284210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:88215a1eb289e6879618fb0261b65bd89f32519f471cbc62fa0e695b76dd712c

Observation 4adad9f9-fad9-43a7-8d8e-15396f9d3750 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:26.854067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:2e44c2c2c050490207b4e5ee72536ea2bdefd419c4fb61f17eaf5e8d300b37ec

Observation 4b66fcbb-5d88-4ff7-abdd-cff4dad2ba57 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.383248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:44cc11495e55e24916ea9322e86df41f4e0cced539db7e2cd9be41ab1cc464d0

Observation f451cf29-e6b8-4dd3-a0e5-1b08896683eb · inbound

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow cites this paper.

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:24.877582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T09:02:09.075097Z digest=sha256:639110e88d7ba50cfc92c7734517f17b0729dca6e2e9a81b48e1327c3c945d02

Observation 75727995-6101-4240-89b7-3e9c522d0f1c · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:04.248345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T01:17:33.668451Z digest=sha256:e7ab6865a8b720f4aa739fba36fcb49574532f4356d6495b8af7591934a90f4c

Observation b0fe877f-5bf5-449c-9e76-3a59ace3047f · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-05T04:30:40.720043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-05T04:29:52.480873Z digest=sha256:ea876581b70f98fda9cb84631db371d18b6dd042adfdb7837053230ad4de8e89

Observation abb05a70-1fbc-41c8-a357-413b306d3a7e · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.909922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T20:24:30.679219Z digest=sha256:2096a4eb1409c592186fb32c7b51ca53d50c7170a307d5e5e2b4b0496f7e3c2b

Observation 16914475-6daa-40bb-ad3e-c59290015fe3 · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.730394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T09:12:22.712240Z digest=sha256:5c3cd889d1f9376dc4f66b9afb1d7011d13d5a40f25d327ae968fabb00e33775

Observation 579364dd-983f-4c76-bfd3-ebb0c2944b74 · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:27:39.577374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:7186b6cb4cf241248cd36ca2aa157b1f84e7df20fedcf38bb50cd9eb38e3434f

Observation efcc01c1-daba-4837-90d9-1bdce7d19ce4 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.878107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:8668c12f2f02008b93ad2deb57eb55a53e96271135c93f96078088c44c9fdb7a

Observation 46758ddf-3d42-469b-872d-32c4c5cb52da · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.447465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:dbdb255874d5caccb6e6622040d7b402e252e7d05b5c6517aac3f3537c3added

Observation 0ecaf041-da59-4b17-bf3e-18d4eb31c305 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.039683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:a0a9873baf184c2b7261eb8a1061ac64d8ba3c3e0d06d4a5a393428bb52b13f6

Observation 929228b3-17f2-4f93-a933-9619e559e750 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.641668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:a13d0fce6ffe78cf20660afc1fe2634b6964f683223d4339a844a70c5772f5a4

Observation 7c4026ce-6994-4e23-8df0-612519f719b0 · inbound

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering cites this paper.

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.614558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T00:24:20.208132Z digest=sha256:74392fa0ef209d58631a5c78594a1d70e7c8514302b415708ddfa4b6070e9361

Observation 238a0bfb-d25a-41b8-90ea-c3601466b7c3 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.731258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:697a87999e3a6ed72560cf36b539cd1917d10dab4bb721b33ee565134192ef30

Observation 8e6216d3-e66c-4f9a-9e92-8eb022bc2b9a · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.162441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:c921b5fbd135b8b9ea840a11f5bcca7ac1fef7c4850956a950f3083fd33014c5

Observation 73ee3721-3cd8-460a-a73f-6e71170c2f86 · inbound

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning cites this paper.

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:07:03.920797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T15:04:18.880016Z digest=sha256:a9a58eb4392dba6d346a02382f638665509555a39c6abca21a1f64ff4e91a79b

Observation 6c79f30b-a7c0-4f88-904a-6c8d6a29cdf5 · inbound

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions cites this paper.

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:01.879688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:01.879688Z digest=sha256:32e222851b9ee307a5653922bff81ad0d17ba7dfda8a4f514278357f489d09d2

Observation b8adfbc4-0697-4892-8d1e-9bce6b39468f · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 143

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.200481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.200481Z digest=sha256:4ba8d7012675c8d963535363d6e2055640f6257ea9909c8f62ca457fe085bfc4

Observation d6899fd5-23ac-494e-848d-f0574a043eac · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.312878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.312878Z digest=sha256:532652d2f73ea96f125fdd25e90d726671df4e53536e93d1d42bde0f56f3ab2d