Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:40:25.866299Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2505.21389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:40:25.866299Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
98 of 98 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 859e0860-1a10-44ea-97f4-4cd191c0163c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0729bc2a-b115-45d9-acc8-ac2492a13ee9 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079b0580-3e41-495c-9245-cd3bdebbd3e7 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Pixtral 12B
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e11e2f29-ff11-441f-916e-ae9f96db639d · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Item response theory: What it is and how you can use the irt procedure to apply it
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f164e09-d6e1-49a0-936e-b035439e018c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7927f561-cbe5-4108-927e-2b125ae26319 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Qwen2.5-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbdff70e-2be6-4d25-9704-bddeebb5f246 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs PaliGemma: A versatile 3B VLM for transfer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3d10c7-ce79-499f-8a62-2560758f21e8 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs A Conceptual Introduction to Hamiltonian Monte Carlo
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7256ea90-91fe-4931-adab-9785d9e88e4d · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Item response theory
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a3f6e8-ed5f-428f-ae38-e30f8c9e180a · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8d5d8f-a9db-45a7-b901-2184cca4158c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60714eb7-06d6-49be-8b88-0d9642f4ba7d · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ce60de2-ff36-4248-b395-deaec309f805 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d5e796-3d5d-45fe-92d6-accde172e660 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Guiding the Growth: Difficulty-Controllable Question Generation through Step-by-Step Rewriting
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d7fa42ba-1ca5-4c80-aaa8-c3ec8efc44e9 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs https://github.com/jiutiancv/JT-VL-Chat , 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7b1f43-7a6b-431e-9147-1e3b5b05c94c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs NVLM: Open Frontier-Class Multimodal LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa5e6d45-5701-4755-b6de-8fc3dbb449fd · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41571838-bf52-4676-847b-b06030067006 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f96212fa-8499-41f4-a7f8-0867af10a55c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Easy2hard-bench: Standardized difficulty labels for profiling llm performance and generalization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9320d1f5-1fce-4579-85aa-230e6bdb3984 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e074296d-34d8-4ca8-97c5-a1dff4d401d4 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51196814-becb-4580-811f-3854ec73b33e · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98834668-9201-4eae-8e30-e3bf8c7b63cb · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305abdde-2cf4-4737-95c8-7fd3ec942c1c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Blink: Multimodal large language models can see but not perceive
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25bf317-5f2d-40ca-955e-33badcf75361 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs H2OVL-Mississippi Vision Language Models Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7790e8-d9bc-4313-af73-73e7ae890635 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Difficulty Controllable Generation of Reading Comprehension Questions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d85e1e-a194-4f42-a0bb-7503eae4c33d · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 66d1d3e2-7d90-4b33-b486-4af3e58fe14e · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Example of the glicko-2 system
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9a8c5b9-0124-4481-8765-c5224935fc06 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18be7ea-ff13-41a2-981c-a9343786aeaa · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ae5757-6f11-422c-a62d-ccb59f4426d7 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Efficient Multimodal Learning from Data-centric Perspective
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a4f9f4-96cc-4e33-9af1-591cddb241f5 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb5378d-d934-481d-b76c-bc1ad7a36ae3 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs GPT-4o System Card
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d85cb312-c88b-420a-ac44-aa7a40c2c00f · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 503d17eb-a3c7-42d3-861e-26b2633339aa · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Ku, Qian Liu, and Wenhu Chen
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f61ed3-621c-4886-b07e-099b92ef3a73 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Automatic educational question generation with difficulty level controls
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b155af45-9c20-464f-b9ec-a9994238aa44 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs A diagram is worth a dozen images
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a03eb3-3c59-43d6-b8ef-1a0cfccaeb51 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Adam: A Method for Stochastic Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f0ecd7-0d61-4caa-b63c-9f5112eca615 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Auto-Encoding Variational Bayes
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d13fbe-517a-4dc7-968d-fb2e65c111ca · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Moondream2: A vision-language model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a52ced-8ff9-455f-9d0c-2d6e37ac6711 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs What matters when building vision- language models? Advances in Neural Information Processing Systems, 37:87874–87907, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0c972b-48de-4849-a504-0cbdf2699521 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Difficulty-Focused Contrastive Learning for Knowledge Tracing with a Large Language Model-Based Difficulty Prediction
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ea43b17b-bf85-4e6e-819d-8a48bff9a40f · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13917c5a-bb02-4254-bbd5-83c110a4c83f · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50813fc-6d66-46ed-98d9-610c938ae99b · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Seed-bench: Benchmarking multimodal large language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df7fb95-3801-4a8d-83c5-60b5fff0a581 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Reform-eval: Evaluating large vision language models via unified re-formulation of task-oriented benchmarks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa363727-9181-443f-a5be-3fd6cbe7a698 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Continuous or discrete, that is the question: A survey on large multi-modal models from the perspective of input-output space extension
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 810e2824-ad71-43d2-b5fb-4cbea895bf9b · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c8a659-4aa0-42aa-ba8f-6b73c7f96a09 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs A survey of state of the art large vision language models: Alignment, benchmark, evaluations and challenges
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2060fb34-e3cf-4ee8-b7c6-167d41a4c4c5 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs DeepSeek-V3 Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e75d31a-f5b7-4617-aa5f-3ec23d9331f3 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b734376c-1a68-4c36-be08-0f4910be9971 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Improved Baselines with Visual Instruction Tuning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2037cf67-c670-4b76-95c5-4406e91a5acf · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf79efd-d0cc-404b-9aad-a36ac0a2091a · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7b6a80-6563-476d-bf84-ec4f198ca983 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs POINTS: Improving Your Vision-language Model with Affordable Strategies
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03b9e430-240a-41a5-b3ea-b30911b9beff · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf27b58-1937-4909-9ec2-6de1a6df3046 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Applications of item response theory to practical testing problems
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7ef08d00-959d-4ee2-95ec-540db9247e00 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Mmalaya2
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 297b0366-7d48-44e1-8630-d5f8542fc998 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Deepseek- vl: Towards real-world vision-language understanding, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ea0cfe11-27b0-466a-b325-f0517896acb3 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Deepseek- vl: Towards real-world vision-language understanding, 2024
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 609752e8-c675-46bc-b10a-040ce47a9637 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e772b1-2662-4301-b53f-c44867191a32 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66051f3d-87e3-4bc9-be8c-dd4563c51b30 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94969d3f-a8bd-4ac6-a698-06814836e0d4 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520597f7-df58-4bc6-a649-8e2f9dcf1183 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasks
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d0ade989-03d2-4893-9686-1a9ad66289db · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Gpt-4v-system-card
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation da47228a-1241-4eea-81b9-7a3371231ebe · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Large language models are students at various levels: Zero-shot question difficulty estimation
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 244d940c-0c30-4aca-99a0-f4a451543ee4 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Automatic differentiation in pytorch
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0892b2d4-f7cb-48cc-b7e2-ac4e0151c6cb · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Transcore-m: Multimodal foundation model for transportation research
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b256a703-e9a1-4633-883a-ef3724999d41 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d52ffe-f59f-4b81-baae-5b941819db4e · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Efficient Benchmarking of Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91455da6-36ac-4a4e-a460-a335005e4c24 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs tinyBenchmarks: evaluating LLMs with fewer examples
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f84bff-13ca-43a7-8ba9-0b07af0b27ad · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs qihoo360/360vl-70b: An open-source large vision-language model based on llama3-70b
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 295d6d78-4cb0-4b82-9c92-e5588ff9ede2 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Learning transferable visual models from natural language supervision
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483bd89e-23a3-4733-a06c-d82a11abeb39 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Probabilistic models for some intelligence and attainment tests
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d47bf71d-1318-4c93-a5c0-783ff35a655a · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef07471-8c90-4b3e-a5f4-abe0866ca726 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 35fe3493-ed07-4dd7-b2ab-7eef11a78a2d · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15471aa2-dd38-4627-b45e-f1a6f99d2d4b · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs tiiuae/falcon-11b-vlm: A vision-language model based on falcon- 11b
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b7e95455-61d5-4a24-a70e-853cd13385c7 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Evaluating Cost-Accuracy Trade-offs in Multimodal Search Relevance Judgements
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8d7e483c-391c-4725-994d-a8d24ac5dadc · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs LLaMA: Open and Efficient Foundation Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343cb7bb-915e-49a8-9ab6-0fe700ab8fd4 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Comparing Test Sets with Item Response Theory
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486e7a53-c6a3-466e-a9d4-4b5e2ffcbfdd · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Moondream1
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5b4da071-ba41-44d2-ada8-c4d333ccc169 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Anchor Points: Benchmarking Models with Much Fewer Examples
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b942e48d-4611-4c30-b882-f65be3a79527 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2590573f-3035-4269-b202-ff079e461bb5 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba1d4ed-d1df-43e7-a8bf-19aa739846be · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Chain-of-thought prompting elicits reasoning in large language models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac17db1a-9f78-4d02-9572-ad99f23879b9 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe6123b-dd9f-4f2b-ab07-a73f2d78e955 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Adaption-of-thought: Learning question difficulty improves large language models for reasoning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 292c4461-a251-49bc-afa8-8774d69dfb8f · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c8dcad-48dd-43e4-b100-9a4a16e77231 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Collageprompt: A benchmark for budget- friendly visual recognition with gpt-4v
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8890e8ef-0ae4-4903-85c3-f15493e5c72c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Qwen2.5 Technical Report
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5336192-b0dd-44e1-8dea-3ac387e47985 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd1247fd-ca59-4c6e-8d84-ab8666766be7 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa99504-6013-429a-97a8-9a2d03f49ce8 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f2e1f6-1317-4010-b328-cee0148f0e27 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7b04f8e-7b3d-4e43-96a6-b8aaf36a9168 · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Efficiently measuring the cognitive ability of llms: An adaptive testing perspective
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 05ed1da9-f499-4840-9920-0ff28eff1f3c · outbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Position: AI Evaluation Should Learn from How We Test Humans
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.