Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T09:30:46.564013Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2607.08317.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T09:30:46.564013Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
76 of 76 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fb89b955-a765-46e9-9c14-7facc794262c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models DeepSeek-V4: Technical report.https://huggingface.co/deepseek-ai/ DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf, 2026
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6b6745db-9992-493e-999b-c363b4109f8c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Introducing GPT-5.5.https://openai.com/index/introducing-gpt-5-5/, 2026
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e920904f-6d1f-4d4e-a829-a553f9dffdcc · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemini 3.1 Pro Model Card.https://deepmind.google/models/ model-cards/gemini-3-1-pro/, 2026
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3fbcf0e2-3961-436e-b4c3-969ce403f57c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation abf63560-7b78-4d65-a202-cd66e875d6fc · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 369e8e9f-774a-4347-b6df-b25c1b3ac7d5 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Swe-bench: Can language models resolve real-world github issues? InThe twelfth international conference on learning representations, 2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6d8cfb2e-dda7-437e-a630-cb38a66147e2 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1424dec4-ba42-4d34-a962-9c4813c9ba32 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 19d5dd26-2ad7-478d-8f4a-0cfc988732c5 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fb61c22b-424b-4e4b-a228-61f8ad37fca0 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 444e852c-2a98-42ef-b0f3-e00de145fcdb · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Instruction-Following Evaluation for Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e916a76a-2e68-4263-9b27-c63a14cdd725 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cc5e9b3f-bb41-46b2-a272-d7e58a8ff10c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 04096664-3215-41c4-be42-7a4028081a90 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Why Do Large Language Models (LLMs) Struggle to Count Letters?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 71f17e32-ee9b-4509-b9e9-2e4ec44e263e · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models BabyVision: Visual Reasoning Beyond Language
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2c8f1418-6c6d-47e8-820a-f65a8c25c593 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Enhancing LLM Character-Level Manipulation via Divide and Conquer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 80994607-8144-48c9-ae13-4de84877d84f · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Com- positional generalization from first principles.Advances in Neural Information Processing Systems, 36:6941–6960, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aee98673-0ab1-47a6-b4f8-44d5a8404c97 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Make it count: Text-to-image generation with an accurate number of objects
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f324f054-04dd-4641-ac92-5e0bcd1808e0 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4695c38f-5842-4c09-9368-5e0954083603 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Constraintbench: Bench- marking llm constraint reasoning on direct optimization.arXiv preprint arXiv:2602.22465, 2026
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c8437ce7-c7ed-4f52-97ab-0d8769438b31 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8b3fd663-13d1-4edf-b255-7444bc170294 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Evaluating the Logical Reasoning Abilities of Large Reasoning Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1bac3f50-ff3e-4870-b37c-ed52de082641 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6dde6904-aad2-46a8-89ac-2f44a5537624 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d7e9917d-5c40-4f73-8cd6-a29dd819cb4c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Beyond accuracy: Behavioral testing of nlp models with checklist
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 81fe030f-7e55-495f-a7f2-730fa174e8e2 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 567220a3-61df-4673-b776-71b33ce6404f · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Dynabench: Rethink- ing benchmarking in nlp
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e98050e3-e60b-422c-9318-a6a5e7d6d264 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Evaluating models’ local deci- sion boundaries via contrast sets
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 373e7d6a-d96a-48b6-9856-e196cb494499 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 26086492-4058-4ff1-b214-0862364f95a6 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6e374258-76b2-41eb-b4c9-fc9e553ed08e · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e70337f8-d115-4ea9-8a37-4e19b0ba1fd7 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Glue: A multi-task benchmark and analysis platform for natural language understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4fefd41b-2774-4ea0-88ea-7bd0b739f767 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ea6a7385-f114-4edb-8153-c20580e6ae86 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models A survey on large language model reasoning failures
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b48fcf0d-20e5-4070-ad95-94715a309822 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GSM-plus: A compre- hensive benchmark for evaluating the robustness of LLMs as mathematical problem solvers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 48682d90-85c8-46ef-9c78-9d6c89d3733e · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d84acaac-6e52-4e22-84b8-ca0cabfc0818 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a722544-6024-4835-9733-368f87fefe00 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Varbench: Robust language model benchmarking through dynamic variable perturbation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 53a8bb3c-d28a-45ea-9f43-a427250d9880 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Longbench: A bilingual, multitask benchmark for long context understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 86cf627b-7d55-4a3f-8910-5f6e1833cda0 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a6d702a7-d72f-4393-af86-8a1340ef24b8 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fd648ad6-4605-44c0-9ce8-e6505fedf286 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Do vision language models rotate in mind? evalu- ating spatial transformation reasoning, 2026
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation acda9658-6897-4f41-8808-e375db2829ff · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Lvlm-count: Enhancing the count- ing ability of large vision-language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2b03149e-7718-4fa4-bff7-dabc9643c160 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3cdbb8bb-b1f9-4ed2-854e-fce8ecb77e73 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Counting Ability of Large Language Models and Impact of Tokenization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation debb9cac-ab1f-4d52-b5f5-93ee5966b07c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Inspect AI: Framework for Large Language Model Evaluations, May
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c5fa324f-3a3c-41a0-99f6-c3bc31ae114c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models MIT License
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bf73e651-ae8f-4b00-9d82-fbffd37af242 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Qwen3 Technical Report
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 98b97f9b-5598-4661-a084-2f6dc2ed5391 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Qwen3.5-Omni Technical Report
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8eb2c0a7-7ef4-4b32-a435-0b6f8dfdcce2 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models GLM-5: from Vibe Coding to Agentic Engineering
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d6deaa6f-d614-4e12-81cb-bd358c064cd5 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Kimi K2.5: Visual Agentic Intelligence
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 87f7324f-5650-4da2-ab07-c4ea61ac4e52 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemma 4 model card.https://ai.google.dev/gemma/docs/core/model_ card_4, 2026
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fdc1c3a6-7ccb-4da7-8117-e0a08ed0edea · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models gpt-oss-120b & gpt-oss-20b Model Card
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 50764985-aa5c-4325-8fb2-2d2839654f24 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1cc6a025-70d0-4bff-b9bd-ef32b4b4c19b · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Gemini 3 Pro Model Card.https://storage.googleapis.com/ deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf, 2025
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9139270f-04c9-4150-a734-0c2733d9b247 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models OpenAI GPT-5 System Card
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0155e527-5608-4846-a879-1ac1c49d9020 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Introducing GPT-5.4.https://openai.com/index/introducing-gpt-5-4/
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf198685-5152-4736-808d-bf269ca1b36d · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Efficient memory management for large lan- guage model serving with pagedattention
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 34d34ffa-f213-4783-a425-799bb10c60b7 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Artificial Analysis Intelligence Benchmarking Methodology: Artificial Analysis Intelligence Index v4.0.4.https://artificialanalysis.ai/methodology/ intelligence-benchmarking, 2026
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation acf5a55d-de7a-433a-ae07-7dbffe84280d · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Humanity's Last Exam
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 70d3d12c-f759-43c3-9068-e0e2a8fea382 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Generalizing Verifiable Instruction Following
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b1db2c66-4082-4eb3-8e2a-a2830cc143a2 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation abc00e6f-8205-4830-9a9e-4daf195e6366 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Vision Language Models are Biased
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f7218e6a-85a1-4acb-ad60-dfa0b667a7f0 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Inksight: Offline-to-online handwriting conversion by teaching vision-language models to read and write.Transactions on Machine Learning Research, 2025
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c5f6d51d-54c6-4cc5-8836-8efc93655c12 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Inkslop: Vibe-coded benchmark for spatial reasoning with digital ink, 2026
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 530aaa96-94cc-48fb-bb25-26a5d7b4347f · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Measuring Compositional Generalization: A Comprehensive Method on Realistic Data
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1bb143d9-7b11-472d-a777-f4f6912acf8f · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Scaling can lead to compositional gener- alization.arXiv preprint arXiv:2507.07207, 2025
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3e329f58-ac1b-428a-8256-e254c68f0d50 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Learning to Count Objects in Natural Images for Visual Question Answering
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d4f8a6d8-75fe-4d04-9d06-7a577d1618bf · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Tallyqa: Answering complex counting questions
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 32d03d1e-7a59-40fd-acae-88b99c342c51 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9409d0a3-64a0-4482-bbab-a83c034dd832 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a7bb48b5-0540-4a8c-b251-485b6c798b6c · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Counterfactual Invariance to Spurious Correlations: Why and How to Pass Stress Tests
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aea1cd86-5799-4482-8ea5-d8d4e90d308b · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1fe88e11-1c9c-4851-a48d-b54a5d26e327 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ed07be03-519f-4683-8cec-693cf8e8f640 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5c002943-c9d0-4b32-ad42-9e90ace31818 · outbound
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models C" for CORRECT: The generated image satisfies the task requirements -
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.