Pith. sign in

Paper Citation Record · LEDGER

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

As of 23 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2507.19525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19525 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:51:17.233907Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:05:29.093803Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T10:07:43.266061Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact4
  • verified fuzzy17
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a419aeb3-04f8-4616-8ac4-741362428b3c · outbound

This paper cites ChipNeMo: Domain-Adapted LLMs for Chip Design.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs ChipNeMo: Domain-Adapted LLMs for Chip Design

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:16.912733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:16.912733Z digest=sha256:3278523046a250dc2f9265691c05d208003ceab5959fffa5c364f923f13eb82a

Observation 109b7cfc-d319-40ae-8bf0-d11fa5b1fcaa · outbound

This paper cites SemiKong: Curating, Training, and Evaluating A Semiconductor Industry-Specific Large Language Model.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs SemiKong: Curating, Training, and Evaluating A Semiconductor Industry-Specific Large Language Model

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:51:17.895784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.921758Z digest=sha256:b6b84309e84b89d7a8acabbac437da58a91797e64386b78363855115bebb9f3b

Observation 08ade2e4-3b06-4b60-9559-4002daf21789 · outbound

This paper cites Openllm-rtl: Open dataset and benchmark for llm-aided design rtl generation,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Openllm-rtl: Open dataset and benchmark for llm-aided design rtl generation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.221774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.929211Z digest=sha256:d12011a5aed24e360083eae10d60a59cf1f410ef5c52e51df19f1796c5596314

Observation 3cac6b90-f9e0-49a8-b66e-d70dcd92e8d5 · outbound

This paper cites Rtllm: An open-source benchmark for design rtl generation with large language model,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Rtllm: An open-source benchmark for design rtl generation with large language model,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.205192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.935433Z digest=sha256:5ac03f96bcca6b3a210bb2791f2b25d9cb543b3ac1b3022e449d590a52c3dac4

Observation 128a57b8-370b-439e-816d-287ba2eb020c · outbound

This paper cites Autobench: Automatic testbench generation and evaluation using llms for hdl design,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Autobench: Automatic testbench generation and evaluation using llms for hdl design,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.189676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.941944Z digest=sha256:b13f1c019acfa70ce542fab47f43b7345a1a6182f6f0c8f32bb3f4d82b96a056

Observation 5497fd4a-983a-4cd8-958f-1edcd8df0bcf · outbound

This paper cites Verilogeval: Evaluating large language models for verilog code generation,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Verilogeval: Evaluating large language models for verilog code generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.173671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.946817Z digest=sha256:99bba0285ac5301b42b9e8db2d30305b714097af617517adc37d31cbf8442349

Observation 578eb82d-8759-41f0-9146-e68449726ef0 · outbound

This paper cites DeepCircuitX: A Comprehensive Repository-Level Dataset for RTL Code Understanding, Generation, and PPA Analysis.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs DeepCircuitX: A Comprehensive Repository-Level Dataset for RTL Code Understanding, Generation, and PPA Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:16.952647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:16.952647Z digest=sha256:f86b55009563e17439bc062a7bd56299856dae8656012211cdd0b7892635db23

Observation d446026a-cc98-40e6-a88a-01818d6d4fd5 · outbound

This paper cites EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:51:17.857997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.958395Z digest=sha256:f1d6e309ac2bfadbc8fe5901cd35d75e15c23dd67d5323042936e1fc67f7ab7e

Observation addf7d73-a50b-4111-be2a-a180c18ebd65 · outbound

This paper cites Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:16.964368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:16.964368Z digest=sha256:8599c00ecd9dfaa3a0070718254d01704786684a919c44c48af7449ae5c0297f

Observation 844bcb60-1ed6-4c4c-83d8-e874fe94597c · outbound

This paper cites The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:16.969771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:16.969771Z digest=sha256:d6b99d31d47eb3509e297830984c99ee68272bc931d2754c4faa3d9b67c099a6

Observation 25ed9432-ebe7-4fb9-8cd6-b3bc9c2ea2ae · outbound

This paper cites Deepgate: Learning neural representations of logic gates,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Deepgate: Learning neural representations of logic gates,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.158615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.975053Z digest=sha256:83163e060b90fbbbb46cfb6df8eddbe2c25da68929572aab382401349c2e16ca

Observation 15de77ff-e2be-4554-9368-f4a9344a48fd · outbound

This paper cites Deepgate2: Functionality-aware circuit representation learning,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Deepgate2: Functionality-aware circuit representation learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.143095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.981378Z digest=sha256:b87eec9f38861cf2a922465757180c72f4809a993d305d0558ce49a1e503847e

Observation ee7fdac5-75a8-4a96-8dbf-e3dde19240ac · outbound

This paper cites DeepGate3: Towards Scalable Circuit Representation Learning.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs DeepGate3: Towards Scalable Circuit Representation Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:51:17.805475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.988229Z digest=sha256:3cee6028a9207ef067a5cd50003813301c63d2fed804ebb2e1926df5600eb43f

Observation bf6167ff-96be-4d75-88cb-07bd972987c8 · outbound

This paper cites DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:51:17.782618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:16.994747Z digest=sha256:f911ec0e07ba1eada4a8f9f8449fbcf18aee18340093340ad64f1445f450cfbf

Observation 2ee24dd8-da72-4218-9edd-eb7b6fbc3eb1 · outbound

This paper cites AutoChip: Automating HDL Generation Using LLM Feedback.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs AutoChip: Automating HDL Generation Using LLM Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.000553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.000553Z digest=sha256:23b3ca1261a3ce4988ccc9f5460d27cf2d543511fd48aa49f627bab2ce4d7479

Observation e524bd91-5f36-4e73-87b7-b743cb198362 · outbound

This paper cites Rtl- coder: Fully open-source and efficient llm-assisted rtl code generation technique,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Rtl- coder: Fully open-source and efficient llm-assisted rtl code generation technique,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.006595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.006595Z digest=sha256:9d86b6eb30f29da55332212de5a39d5f6fa3ed85b30e201ba9fd739ac2b7aeaf

Observation d5e31a20-df2e-4bc7-9ef5-9037bb65b86f · outbound

This paper cites Verigen: A large language model for verilog code generation,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Verigen: A large language model for verilog code generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.116952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.011689Z digest=sha256:2c25120c322ed7e25d4ba4b1fb18f6f59d0988a19fd5cc5d58880d6f6d8b2533

Observation 75256a78-6bea-4ff5-80ec-014340228ea8 · outbound

This paper cites AmpAgent: An LLM-based Multi-Agent System for Multi-stage Amplifier Schematic Design from Literature for Process and Performance Porting.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs AmpAgent: An LLM-based Multi-Agent System for Multi-stage Amplifier Schematic Design from Literature for Process and Performance Porting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.017168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.017168Z digest=sha256:80cc1113f43e11267e8035d5f7343204dede667a473429d2b72fe9db5b184d35

Observation 72fe7c05-57aa-4060-bee7-0823bcd74789 · outbound

This paper cites Chateda: A large language model powered autonomous agent for eda,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Chateda: A large language model powered autonomous agent for eda,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.024063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.024063Z digest=sha256:e498b7ceb95119f6f8ed00d6b4be83c65daf338408fa8797a55557776c535fc8

Observation 81083eca-4fa3-4f1f-a900-fa7b8b6b9c81 · outbound

This paper cites Openroad: Toward a self-driving, open-source digital layout implementation tool chain,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Openroad: Toward a self-driving, open-source digital layout implementation tool chain,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.085480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.030673Z digest=sha256:90a5ea392240e545f9f6d75a40e53e30d76b21698b74adfece47884e8a33f429

Observation 094d4aef-15c0-42ba-8820-73b181ffd74c · outbound

This paper cites Chipgpt: How far are we from natural language hardware design,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Chipgpt: How far are we from natural language hardware design,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.039330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.039330Z digest=sha256:0f15e02f3daa41b9d225c8058b11574c225e12be3f9fda4337f86f6ab46f8d74

Observation ead691cd-84ac-4ba2-846e-7795243e2fb8 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.045621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.045621Z digest=sha256:f3c0e1cfd9564d5419a743ca757c6f3363b24ecec7903ed52947126b65870a74

Observation 71178a86-0847-4ee5-84d2-c5027c0159b3 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.070119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.052390Z digest=sha256:c22fed4c0691ce1138973e29422d876562a9c78be57ee53c273a0f11a3dac3a9

Observation 1dfc5907-444a-4fd8-9e26-a0bc3e54681d · outbound

This paper cites Mm- bench: Is your multi-modal model an all-around player?.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Mm- bench: Is your multi-modal model an all-around player?

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.055455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.058059Z digest=sha256:f13b48d9d005a932ace60d1ccec4c6474583b4fcb68ebbc954dc96c1ca280bcb

Observation bb92531a-a3e9-4e0f-a306-8f795430904c · outbound

This paper cites ChipExpert: The Open-Source Integrated-Circuit-Design-Specific Large Language Model.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs ChipExpert: The Open-Source Integrated-Circuit-Design-Specific Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.063404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.063404Z digest=sha256:1a402d8eb4c7b7c3ae553100b0d62ad0b3c78139c21a4bb68ddba6faaa3434df

Observation cbc7941f-5ab6-42c1-8c04-c5911f47f3e2 · outbound

This paper cites GPT-4 Technical Report.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs GPT-4 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.069313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.069313Z digest=sha256:08bd55a7aa755b2709c800b38720f2fadd9737e8f0c4dafe035f23fa962495c9

Observation 3f2a9dc3-c3c6-42fa-abe0-7d28c7d65373 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.040918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.078106Z digest=sha256:8d039a60bd7f75cc833cbc521998aeb1a9ab0ff3ccf9da9a4ac78f000e11cd4c

Observation ab6d3375-d6ab-48a0-8b36-dc6dce52fc13 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.093914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.093914Z digest=sha256:b9878560f3e4340ae9ccd325412d39a3e9ace1a39747545d4f312b9f4ef1ecd2

Observation 4b609da0-9d61-4d75-be1e-dabb843144db · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.099887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.099887Z digest=sha256:e64428e71ccbeec26dc0cbd01d3af9a749dc152ec33fee87ce2c4e398598f298

Observation e1ebf3e1-1969-457d-8fb5-c122916c0aab · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.026078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.106267Z digest=sha256:b239ed22c3ee41b241525cd89bd2f1bde2557732f2bf14466719c7622f6385e8

Observation 729b1d09-85f8-46d2-812c-64754912dbbd · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.111470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.111470Z digest=sha256:45575e93084f53f38f428b1b8399adc441cbc32eacba01e55c0eb06eaed46c78

Observation 3db087b0-fa50-4033-9d16-ecdf5aaf34a0 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:18.000768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.116280Z digest=sha256:060f2ae74a799173e26adbb77fbd89bdd70cad8814fe9432899c1a838d8bcbbc

Observation 24fa938f-bb6b-47b4-bb43-dd678a955ea0 · outbound

This paper cites The Llama 3 Herd of Models.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs The Llama 3 Herd of Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.121764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.121764Z digest=sha256:c06959d858e6a50787206d861f6c221ac0e9253ea1c6bc1dae95b3bd02e9459c

Observation f6c0f823-15fa-49da-a62a-7b3c0a3a9398 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.128128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.128128Z digest=sha256:31a7d048503e7c97335170a16dccdf2a9cce9ad0af2345b967acfcaf142507a9

Observation 8a2bb5f9-0d34-4f23-bcf0-7ebf6215c6b6 · outbound

This paper cites Yi: Open foundation models by 01.ai,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Yi: Open foundation models by 01.ai,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:17.985303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.133324Z digest=sha256:e72a0cbb5a6740ff864726c865e1e9f78164634a78daa32c280137a9bdef6a3f

Observation cb79e9a6-4927-43e7-a8af-8f321bd31445 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.139099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.139099Z digest=sha256:f75730f75a394404cfb0e9a5554ca391bb8f7e7c3e442a653051b3f45a1c9444

Observation 731c1bb5-9582-4676-b119-93fabf117ea5 · outbound

This paper cites GPT-4o System Card.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs GPT-4o System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.143992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.143992Z digest=sha256:8c84f9cda1840d4e7ce6408019c0a8f780f6dd33a665a02c4b9890b9ee6b5767

Observation 3bf026fa-d614-4338-bdd4-9ba4f8da7982 · outbound

This paper cites Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.148602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.148602Z digest=sha256:97264e5ac064b89aa8d545a25faf392184774623ed5870bd1abcb43b11c200e5

Observation a27ef1f9-599c-4706-abc2-9c6205d59228 · outbound

This paper cites Training language models to follow instructions with human feedback,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Training language models to follow instructions with human feedback,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:17.969783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.154151Z digest=sha256:fdcc1e841da575b240364980e10ba10564e63496af29aca57c5d36fc69459882

Observation 3ad89559-04f8-4e8a-ade6-47f2d293a4e5 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.159030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.159030Z digest=sha256:841b660e26e9b1557e0ad98f1d6dd47e38f345aca174b7eb3cd05fb61fd32b05

Observation 87cb3f12-4f8c-4c7e-9cf9-79080eba7608 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.164744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.164744Z digest=sha256:cf84eec004deb6968fc7d94e62ebd0441d39d16c5465df76a957af532dcfdbc2

Observation 18563847-c589-408f-8929-13531c10b2fb · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.169647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.169647Z digest=sha256:501d02abf24666c848221744e769d74ae3f997f177210da385cd67ccb1b12757

Observation be387cce-c2e0-495a-b380-dec5915f6754 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.174974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.174974Z digest=sha256:35473e66d3c7fa0c71cf6b89471d0a3990a02e5f2828cc456d133d608383827e

Observation eb7b6998-2552-42ad-8b06-cd08ecb5e814 · outbound

This paper cites Qwen2 Technical Report.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.180690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.180690Z digest=sha256:247487f30b8627bf2b4d0e7bcb24df4f348abf530b0b573fff40edf7312082b1

Observation 011286e8-6a12-4703-b35b-f0bc1f4b1f53 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Internlm: A multilingual language model with progressively enhanced capabilities,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.186161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.186161Z digest=sha256:285d6c451280fad43f1e473a41962c6126312ddea0a31e777338f4af7cc27840

Observation a4c04907-5961-4cb7-a091-8752e214aba7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.190937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.190937Z digest=sha256:4b7a4e1305797afe12e3e3eeb6b982075c40f5ae607b2836e9334db7234c2f3f

Observation 71d40203-d16b-431c-b640-effd35a3e47d · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.196765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.196765Z digest=sha256:bdc8790fca4c3d11deabc528d6545d6fbfd19cfb1c66a46d9230509597beea18

Observation ebc3817a-2c46-4dff-a0d0-8969f3d56eb1 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.202392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.202392Z digest=sha256:7dbe7ecec21187bc52f4b92bea9fef99056ad33fba1c1599f8926a581a49805b

Observation 702586f6-6497-4a16-9149-023e470f0276 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.207829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.207829Z digest=sha256:9bd4665c56aa59cd734986918709c270575664f8631e8cd0f1af67a4602c6b89

Observation f7d00419-4ff1-4299-92f8-df8e1db749c0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.212642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.212642Z digest=sha256:91711247da544538b0383e7b42a1f065c25b87853473e1939ddca1e75677e15e

Observation c33732ac-dd29-46ce-9cc9-3da01af18351 · outbound

This paper cites Introducing the next generation of claude,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Introducing the next generation of claude,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:17.944409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.218440Z digest=sha256:72296df31282b55a826afceda4df36c4a2db443ceb14a3565d834e782e0d7c93

Observation 822042b0-c8c7-4988-a374-0d966cd12b94 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.223412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.223412Z digest=sha256:9937e87a859dd24a7cbf1fabf4787361a6f3e994855591a7539955a03ab97098

Observation 830ec5db-bf2b-411e-9fdf-e356b33cb029 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:17.229216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:17.229216Z digest=sha256:3f6aac3a162b8893cca0188843b2f34be73ab059c4ab0157252af2b5993109bc

Observation 037edc80-35b4-44b7-93e8-294d2f01ee1e · outbound

This paper cites Language models are few-shot learners,.

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs Language models are few-shot learners,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:51:17.929488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T15:51:17.233907Z digest=sha256:86285d76849623e3dab79fe7d4dc3e4823f1590808a5c0aff2178e516445f3ad

Pith citing papers

Observation bc087c9e-bd8a-4f97-b410-0ea88f022682 · inbound

CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis cites this paper.

CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:07:43.268039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T10:05:29.093803Z digest=sha256:4ad436e4d1e7d69cb2ab186460624a1b742fe5f5d8add2935762c798df172800