Pith. sign in

Paper Citation Record · LEDGER

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

As of 22 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2607.08317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08317 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T09:30:46.564013Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact37
  • verified fuzzy30
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb89b955-a765-46e9-9c14-7facc794262c · outbound

This paper cites DeepSeek-V4: Technical report.https://huggingface.co/deepseek-ai/ DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models DeepSeek-V4: Technical report.https://huggingface.co/deepseek-ai/ DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf, 2026

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.212164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:30efe8980603e1933b72c75f92610043af05234af3081cfbe997607c1355a606

Observation 6b6745db-9992-493e-999b-c363b4109f8c · outbound

This paper cites Introducing GPT-5.5.https://openai.com/index/introducing-gpt-5-5/, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Introducing GPT-5.5.https://openai.com/index/introducing-gpt-5-5/, 2026

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.246226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:14dfacc2267940bc571336c88a9e0f582ff87522e9ee81bda7d76bfd694477bb

Observation e920904f-6d1f-4d4e-a829-a553f9dffdcc · outbound

This paper cites Gemini 3.1 Pro Model Card.https://deepmind.google/models/ model-cards/gemini-3-1-pro/, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemini 3.1 Pro Model Card.https://deepmind.google/models/ model-cards/gemini-3-1-pro/, 2026

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.222527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:e537f98ed8516358963d127b7dceb1fd1b2f9bb7d20cdb3650ed8892f0ac7f95

Observation 3fbcf0e2-3961-436e-b4c3-969ce403f57c · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.870806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:ba484425e2e84b65081dfe4285fea7368b0408388c114702996d34276b781074

Observation abf63560-7b78-4d65-a202-cd66e875d6fc · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.767104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:af8d098b946ba18ee148db6a5dceb3c44ad5b957ae1560e43a62c36416077f04

Observation 369e8e9f-774a-4347-b6df-b25c1b3ac7d5 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? InThe twelfth international conference on learning representations, 2023.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Swe-bench: Can language models resolve real-world github issues? InThe twelfth international conference on learning representations, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.240601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:f6ba39e68e516e484d467b8c85c57e8a68aee9474583303e21cf6d105ef97aef

Observation 6d8cfb2e-dda7-437e-a630-cb38a66147e2 · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.862871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:beb2e79f51b6a2982eb35d019f69de9815a76ae62c3375df6a8ec2d36d00b75e

Observation 1424dec4-ba42-4d34-a962-9c4813c9ba32 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.857750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:a20b26ea15ff751400f99d49dd28becd541ab657c1dd85d05a84316643f8314a

Observation 19d5dd26-2ad7-478d-8f4a-0cfc988732c5 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.242321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:82aed5c2e9764d4fec0ddcc87682ef010ec618541a17b8538262ae08787bb5f9

Observation fb61c22b-424b-4e4b-a228-61f8ad37fca0 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.243904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:8ea6188e5626e3c4f526ad4d4efcc55bcab027fe1d50b6a9754c6e90d2e95ebc

Observation 444e852c-2a98-42ef-b0f3-e00de145fcdb · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Instruction-Following Evaluation for Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.783817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:27b2b7d7574280235a14632779af375f91367f4e4cc43b19c446ab74127dcbd1

Observation e916a76a-2e68-4263-9b27-c63a14cdd725 · outbound

This paper cites Sudoku-Bench: Evaluating creative reasoning with Sudoku variants.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Sudoku-Bench: Evaluating creative reasoning with Sudoku variants

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.823453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:00ec62c7460a6069c8a1ff79a4a5294362828254e7ad3fb65be38ce3d280cfb6

Observation cc5e9b3f-bb41-46b2-a272-d7e58a8ff10c · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.218985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:be5b00951c51f63964b5daaca3b17740c3f374c6054b5fc263a981416cdadeaf

Observation 04096664-3215-41c4-be42-7a4028081a90 · outbound

This paper cites Why Do Large Language Models (LLMs) Struggle to Count Letters?.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Why Do Large Language Models (LLMs) Struggle to Count Letters?

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.788988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:ead727bbf5b9b9067d31fe03ce46ca48382aea4b49fc6acb8f802d1002bc9414

Observation 71f17e32-ee9b-4509-b9e9-2e4ec44e263e · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.825956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:47fc274690e05ba508ff3e4ecc384a5e018bac21c60e53cd5c31e8bd4eac9542

Observation 2c8f1418-6c6d-47e8-820a-f65a8c25c593 · outbound

This paper cites Enhancing LLM Character-Level Manipulation via Divide and Conquer.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Enhancing LLM Character-Level Manipulation via Divide and Conquer

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.852125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:5124d1289ef713b29ad745721ee1f64fc902c2179b0861b490602b20ac2175d8

Observation 80994607-8144-48c9-ae13-4de84877d84f · outbound

This paper cites Com- positional generalization from first principles.Advances in Neural Information Processing Systems, 36:6941–6960, 2023.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Com- positional generalization from first principles.Advances in Neural Information Processing Systems, 36:6941–6960, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.236873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:63069312987cc911b6a509f326a39380c34fb2b870f53c43bbb7cf2abc344767

Observation aee98673-0ab1-47a6-b4f8-44d5a8404c97 · outbound

This paper cites Make it count: Text-to-image generation with an accurate number of objects.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Make it count: Text-to-image generation with an accurate number of objects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.238893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:afa1a77eca81e4a02489e946a5c29107a944dfc258ca11517b29f9e70dda7da5

Observation f324f054-04dd-4641-ac92-5e0bcd1808e0 · outbound

This paper cites Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T09:37:00.849556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:766ea650dd11808bcf030a98cf4e8d44dad7db81691770ca72c2b9cc5fb765dd

Observation 4695c38f-5842-4c09-9368-5e0954083603 · outbound

This paper cites Constraintbench: Bench- marking llm constraint reasoning on direct optimization.arXiv preprint arXiv:2602.22465, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Constraintbench: Bench- marking llm constraint reasoning on direct optimization.arXiv preprint arXiv:2602.22465, 2026

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T09:37:00.839546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:b53ee1281ad24abc60f479519f85baf2398960e5bb0a6344655c17be95d17537

Observation c8437ce7-c7ed-4f52-97ab-0d8769438b31 · outbound

This paper cites SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.835854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:de9a2347fe07f66fed3193145324a45a2505d2d86b5e4c019a3415da779013a5

Observation 8b3fd663-13d1-4edf-b255-7444bc170294 · outbound

This paper cites Evaluating the Logical Reasoning Abilities of Large Reasoning Models.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Evaluating the Logical Reasoning Abilities of Large Reasoning Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.772539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:220d7ff8e0cdb0aa96db1812e828047162a2380867b098bcc2ce676971d79edb

Observation 1bac3f50-ff3e-4870-b37c-ed52de082641 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.796446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:acce5dfbda0f0fb594d043f6f81def28d1d18fd867882b1022741880dabf190c

Observation 6dde6904-aad2-46a8-89ac-2f44a5537624 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.225974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:80e0a4e967dda29de3373914379449d3156c91269945c1d354dd0c3d77e2a5c1

Observation d7e9917d-5c40-4f73-8cd6-a29dd819cb4c · outbound

This paper cites Beyond accuracy: Behavioral testing of nlp models with checklist.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Beyond accuracy: Behavioral testing of nlp models with checklist

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.227610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:fef55266680c9f77397d88112574e4ea949a5a3af43d9285d0a706fc49359253

Observation 81fe030f-7e55-495f-a7f2-730fa174e8e2 · outbound

This paper cites Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.229229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:b22a734ba8d91054ba2742c3dc2c66ef62cdcbe8783397e2676c87d9be9a14d1

Observation 567220a3-61df-4673-b776-71b33ce6404f · outbound

This paper cites Dynabench: Rethink- ing benchmarking in nlp.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Dynabench: Rethink- ing benchmarking in nlp

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.232558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:c0d45612bcad79e09656a3153ae2444360983fe9aa379c2d1a5dd0cfa7c5cd13

Observation e98050e3-e60b-422c-9318-a6a5e7d6d264 · outbound

This paper cites Evaluating models’ local deci- sion boundaries via contrast sets.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Evaluating models’ local deci- sion boundaries via contrast sets

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.234199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:ca5a16b054fc9d54f834d2d97ebd6e7287d3f12da481b7af45f1ada5d9094552

Observation 373e7d6a-d96a-48b6-9856-e196cb494499 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.775167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:32ab440ae364b1b8ecce40b4a82a4babf472b1afd89e933a242bb7969af97a0a

Observation 26086492-4058-4ff1-b214-0862364f95a6 · outbound

This paper cites MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.791323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:461c936395272ed0643bd363162803a0ef3f6da58b5dc028f0985be5338ad6d4

Observation 6e374258-76b2-41eb-b4c9-fc9e553ed08e · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.828567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:7cb680c4bcbfe5c29d603f03c501558d7b8f9ef8a4402a2b00ba2e018c2e4fc8

Observation e70337f8-d115-4ea9-8a37-4e19b0ba1fd7 · outbound

This paper cites Glue: A multi-task benchmark and analysis platform for natural language understanding.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Glue: A multi-task benchmark and analysis platform for natural language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.213885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:8b160b15a91982985eef0e77e223d74207f55c39f438f85f93a2017ed78a8d03

Observation 4fefd41b-2774-4ea0-88ea-7bd0b739f767 · outbound

This paper cites SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.786672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:c0630122a2cc2e23c9cdb2b5abbf55d66c9384c6adcb28d08108c062199e2fcf

Observation ea6a7385-f114-4edb-8153-c20580e6ae86 · outbound

This paper cites A survey on large language model reasoning failures.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models A survey on large language model reasoning failures

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.231040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:b2908ccbd1f15e81de02818b302950a124c6e435db6c31201e3708c0b31336ec

Observation b48fcf0d-20e5-4070-ad95-94715a309822 · outbound

This paper cites GSM-plus: A compre- hensive benchmark for evaluating the robustness of LLMs as mathematical problem solvers.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GSM-plus: A compre- hensive benchmark for evaluating the robustness of LLMs as mathematical problem solvers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.215615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:75a6caa2672fdae5e6de900d27d94bbfbd33d40468cdd2e1b32be1aa6994cb1c

Observation 48682d90-85c8-46ef-9c78-9d6c89d3733e · outbound

This paper cites MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.719526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:1e22ae7d2c15b201762f9ebfce08329ca88d7af2d251e1f06d59591b7700019f

Observation d84acaac-6e52-4e22-84b8-ca0cabfc0818 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.716546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:de587432afc527162492f2a7360ebbff7b0482525ef2fc85245020bac58788a6

Observation 2a722544-6024-4835-9733-368f87fefe00 · outbound

This paper cites Varbench: Robust language model benchmarking through dynamic variable perturbation.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Varbench: Robust language model benchmarking through dynamic variable perturbation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.217320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:6d96713ce5ca64c7d3492b9d2544577062a9f5b6e0db99fae59addd162e94cf4

Observation 53a8bb3c-d28a-45ea-9f43-a427250d9880 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Longbench: A bilingual, multitask benchmark for long context understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.253065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:d5fda83ceeaba7cf919aac2ebff7520981d81dd4f4ff721aca7e7dbab5ff5955

Observation 86cf627b-7d55-4a3f-8910-5f6e1833cda0 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.249610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:8e04c5c8c4dade9e7934ef611f2aa6c43ac792a4f449780af627ecfb24785181

Observation a6d702a7-d72f-4393-af86-8a1340ef24b8 · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.842655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:fac3b4849b51b532daf3c06d90ac45297b11b1cbe7c0e177e8422884f6964bae

Observation fd648ad6-4605-44c0-9ce8-e6505fedf286 · outbound

This paper cites Do vision language models rotate in mind? evalu- ating spatial transformation reasoning, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Do vision language models rotate in mind? evalu- ating spatial transformation reasoning, 2026

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.207110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:483a6a914aafbcebed96077b11c143420208b9b46c6cd1d2657912a6517c6063

Observation acda9658-6897-4f41-8808-e375db2829ff · outbound

This paper cites Lvlm-count: Enhancing the count- ing ability of large vision-language models.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Lvlm-count: Enhancing the count- ing ability of large vision-language models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-10T09:37:00.873469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:665e0950633ad46d06bc1a3bc5f9ef77b47e534d6bd8fc16b99adf6991f6ea4a

Observation 2b03149e-7718-4fa4-bff7-dabc9643c160 · outbound

This paper cites Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.811042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:b579fcbe7dd82eda41fda0339d7a226594d34cd1f2073c59c1b84a56a873098d

Observation 3cdbb8bb-b1f9-4ed2-854e-fce8ecb77e73 · outbound

This paper cites Counting Ability of Large Language Models and Impact of Tokenization.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Counting Ability of Large Language Models and Impact of Tokenization

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.803900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:98a9cfae49088428b0fd81dd0ca9df1b78401760c303995b55b3303d070c562c

Observation debb9cac-ab1f-4d52-b5f5-93ee5966b07c · outbound

This paper cites Inspect AI: Framework for Large Language Model Evaluations, May.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Inspect AI: Framework for Large Language Model Evaluations, May

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.210429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:482ce788436f1ebb9d695a72d44aa9a46d9a63dacbfa5277e1c5b6a3a780c20e

Observation c5fa324f-3a3c-41a0-99f6-c3bc31ae114c · outbound

This paper cites MIT License.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MIT License

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.247831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:269bc13134f7c0d887b7b32d432a9b24c97db8390591cf1889003c3a9dad9857

Observation bf73e651-ae8f-4b00-9d82-fbffd37af242 · outbound

This paper cites Qwen3 Technical Report.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Qwen3 Technical Report

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.860103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:7c96c23478647d7a6ba5de118dce6d4ff6044a02a9bfcecb361bf64f074e01f4

Observation 98b97f9b-5598-4661-a084-2f6dc2ed5391 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Qwen3.5-Omni Technical Report

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.855021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:6c6c1dae654656181ec458fc2e006906c0292137795b674790e582f6b0cf08f3

Observation 8eb2c0a7-7ef4-4b32-a435-0b6f8dfdcce2 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GLM-5: from Vibe Coding to Agentic Engineering

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.781400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:9224cea1e35e6efc95df4d65dbcff575d7070dce7cc0673e517ec4a7bd158b14

Observation d6deaa6f-d614-4e12-81cb-bd358c064cd5 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Kimi K2.5: Visual Agentic Intelligence

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.777604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:cbadf5b3b8704615f1941eb24be6ae458dea2977e71805bf5c34244cb99dd979

Observation 87f7324f-5650-4da2-ab07-c4ea61ac4e52 · outbound

This paper cites Gemma 4 model card.https://ai.google.dev/gemma/docs/core/model_ card_4, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemma 4 model card.https://ai.google.dev/gemma/docs/core/model_ card_4, 2026

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.251362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:0040348dd341b5f17419fd4a28ba7e2bf6377ab08d2bcb4d5b53c71b3db50ddb

Observation fdc1c3a6-7ccb-4da7-8117-e0a08ed0edea · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models gpt-oss-120b & gpt-oss-20b Model Card

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.801393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:1863e1df2f18345ca337b472d0055cce61a99441ab7e8ecc782d12539581626f

Observation 50764985-aa5c-4325-8fb2-2d2839654f24 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.868256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:631209125604fb52aeca6385c410cb650797ebc0fb61ec430d36881c77aed56e

Observation 1cc6a025-70d0-4bff-b9bd-ef32b4b4c19b · outbound

This paper cites Gemini 3 Pro Model Card.https://storage.googleapis.com/ deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf, 2025.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemini 3 Pro Model Card.https://storage.googleapis.com/ deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf, 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.203644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:7f78ccf807c89035a851a6a9f217065be7e637c58e97294dc20c40da0c1bb661

Observation 9139270f-04c9-4150-a734-0c2733d9b247 · outbound

This paper cites OpenAI GPT-5 System Card.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models OpenAI GPT-5 System Card

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.808527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:e0a3443f8d5260fb843cf143f9fa5ec8d2dde1be64f90e75bfeb316a74ae7d73

Observation 0155e527-5608-4846-a879-1ac1c49d9020 · outbound

This paper cites Introducing GPT-5.4.https://openai.com/index/introducing-gpt-5-4/.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Introducing GPT-5.4.https://openai.com/index/introducing-gpt-5-4/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.205427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:2bc9561cbf4a0334699f674e31b1a77567ec827771473d8e70d44b24910b2953

Observation cf198685-5152-4736-808d-bf269ca1b36d · outbound

This paper cites Efficient memory management for large lan- guage model serving with pagedattention.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Efficient memory management for large lan- guage model serving with pagedattention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.208853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:35c286f10b66f532c12103149ad7754de1a44aee14a479d42723b20e9c548f2e

Observation 34d34ffa-f213-4783-a425-799bb10c60b7 · outbound

This paper cites Artificial Analysis Intelligence Benchmarking Methodology: Artificial Analysis Intelligence Index v4.0.4.https://artificialanalysis.ai/methodology/ intelligence-benchmarking, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Artificial Analysis Intelligence Benchmarking Methodology: Artificial Analysis Intelligence Index v4.0.4.https://artificialanalysis.ai/methodology/ intelligence-benchmarking, 2026

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.220841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:86e6f33e2027160ae9188589f6f8f6c2055a26a805380d8db52411beec10ee26

Observation acf5a55d-de7a-433a-ae07-7dbffe84280d · outbound

This paper cites Humanity's Last Exam.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Humanity's Last Exam

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T09:37:00.710564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:54d5135eb919e9677faa81947e5ee6c2919b112eaede113b0b20426b46fd9bf2

Observation 70d3d12c-f759-43c3-9068-e0e2a8fea382 · outbound

This paper cites Generalizing Verifiable Instruction Following.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Generalizing Verifiable Instruction Following

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.830888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:f59aebf8987fab1eddf368a1f04c7b3fad84128750deab13f9cc064b508a5261

Observation b1db2c66-4082-4eb3-8e2a-a2830cc143a2 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T09:37:00.702370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:89d06f830fc7f273ab8b0614333ffa05c60d3eeadcdae86cd20ab3c6b543fe6b

Observation abc00e6f-8205-4830-9a9e-4daf195e6366 · outbound

This paper cites Vision Language Models are Biased.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Vision Language Models are Biased

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.793694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:ba1b919a73656fe6aa327ad9b9109943b8f5637baca2c372b47e4c449d088d68

Observation f7218e6a-85a1-4acb-ad60-dfa0b667a7f0 · outbound

This paper cites Inksight: Offline-to-online handwriting conversion by teaching vision-language models to read and write.Transactions on Machine Learning Research, 2025.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Inksight: Offline-to-online handwriting conversion by teaching vision-language models to read and write.Transactions on Machine Learning Research, 2025

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.201949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:72b7f559f4564f104cb4e3a9c696a6865dc05c45e03e9825f7705ab86a1e42a1

Observation c5f6d51d-54c6-4cc5-8836-8efc93655c12 · outbound

This paper cites Inkslop: Vibe-coded benchmark for spatial reasoning with digital ink, 2026.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Inkslop: Vibe-coded benchmark for spatial reasoning with digital ink, 2026

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T09:37:01.254722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:28a1912de7ac19d4924c961127bc1fbee060697fed22a39fd872d21d3589fac1

Observation 530aaa96-94cc-48fb-bb25-26a5d7b4347f · outbound

This paper cites Measuring Compositional Generalization: A Comprehensive Method on Realistic Data.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Measuring Compositional Generalization: A Comprehensive Method on Realistic Data

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T09:37:00.865769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:b8c9cd3d7ac0f21234cf55d3fee8a2d9aac0fc8d9c116e9abcc101113ae8e5b0

Observation 1bb143d9-7b11-472d-a777-f4f6912acf8f · outbound

This paper cites Scaling can lead to compositional gener- alization.arXiv preprint arXiv:2507.07207, 2025.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Scaling can lead to compositional gener- alization.arXiv preprint arXiv:2507.07207, 2025

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-10T09:37:00.769972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:e71f1e2a001a53334f88905dfed65ec856a494c62fb002cad36cef01a801b79a

Observation 3e329f58-ac1b-428a-8256-e254c68f0d50 · outbound

This paper cites Learning to Count Objects in Natural Images for Visual Question Answering.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Learning to Count Objects in Natural Images for Visual Question Answering

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.806324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:c25bd9686fb88c8d0c926e0dfc662bed7828ab76a518389c5afc51415b3ea3f2

Observation d4f8a6d8-75fe-4d04-9d06-7a577d1618bf · outbound

This paper cites Tallyqa: Answering complex counting questions.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Tallyqa: Answering complex counting questions

Reference 71

Resolution
verified exact
doi, observed 2026-07-10T09:37:00.713335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:ce39b76a93a6e677cc05219f85bceaa5eab1769ce29dcb2c8ef18da96d401055

Observation 32d03d1e-7a59-40fd-acae-88b99c342c51 · outbound

This paper cites Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-10T09:37:00.820846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:93654c71e517307a8a442156ad2dfec09a03ffa12865f04f4e06e434b9739d78

Observation 9409d0a3-64a0-4482-bbab-a83c034dd832 · outbound

This paper cites an unresolved cited work.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-07-10T09:37:01.195107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:db5ab8a05c710e9d42c59c490b0c9dc4f5e1eff0812d6b7ccf9437a6f6d30dbb

Observation a7bb48b5-0540-4a8c-b251-485b6c798b6c · outbound

This paper cites Counterfactual Invariance to Spurious Correlations: Why and How to Pass Stress Tests.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Counterfactual Invariance to Spurious Correlations: Why and How to Pass Stress Tests

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.798951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:25e776659df8300163245127715a7174ca93a782124d841a72a7d5e2ec3978a6

Observation aea1cd86-5799-4482-8ea5-d8d4e90d308b · outbound

This paper cites an unresolved cited work.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-07-10T09:37:01.200122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:6358ea58f285dd8b97b6d3871adf3e7b69afa871a99ed7e6207170a6cbfbbdb7

Observation 1fe88e11-1c9c-4851-a48d-b54a5d26e327 · outbound

This paper cites an unresolved cited work.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-07-10T09:37:01.198490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:5a96cf0b370d5a8c2e6667b907951d04de9b10a0f69a63207ac18505d8dee04c

Observation ed07be03-519f-4683-8cec-693cf8e8f640 · outbound

This paper cites an unresolved cited work.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-07-10T09:37:01.196836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:9677c89b965e427dfbd4cd517bbcd90e661ac182b673c24614d6840e5b631ee3

Observation 5c002943-c9d0-4b32-ad42-9e90ace31818 · outbound

This paper cites C" for CORRECT: The generated image satisfies the task requirements -.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models C" for CORRECT: The generated image satisfies the task requirements -

Reference 78

Resolution
malformed identifier
raw_fallback, observed 2026-07-10T09:37:01.224324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:15c0ca62b23fdeef56cf8ce7406a06728293bc6fc9e3e6707bb94c7b5a23b680

Pith citing papers

No inbound Pith citation observations are available.