Pith. sign in

Paper Citation Record · LEDGER

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2411.17794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17794 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:29.201224Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:24:57.169737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:36:05.014470Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c1ff511-65d0-40df-b468-58e372740f50 · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Flamingo: A visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:30.011651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.007620Z digest=sha256:c1e26c8fffbb85dc94d4e57bc237b61bfffed7dad2cd776573874f02d5b7e65f

Observation ebb4e11c-4f2b-4fb0-8171-5f3d23a204d6 · outbound

This paper cites Qwen Technical Report.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.012130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.012130Z digest=sha256:7c4ed4aa83e77642eaa451aab1b8636df8b039bc638a1419cd5057d498a7c84f

Observation 4d88c780-105f-482c-bf8a-810709dde470 · outbound

This paper cites Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Breaking common sense: WHOOPS! A vision- and-language benchmark of synthetic and compositional im- ages

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.999362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.016492Z digest=sha256:7dc0fa8d0e3379438449113e4300918a62759199c0e06c39b7d8442a0203390e

Observation 28130577-7c02-4337-8cc3-d06c740ee696 · outbound

This paper cites Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.020551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.020551Z digest=sha256:2ffd1eeba396cef234162b47b62a971684efe90f3371e47929e98b989b2b6837

Observation 35c01332-1ba4-4faf-9734-98138069ce28 · outbound

This paper cites InternLM2 Technical Report.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.024930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.024930Z digest=sha256:09a136627b4eac16b0db7f1e3254efdaac00651516bbeecbe96b2310fedf33e3

Observation 1c78c158-3ead-4c2d-bc7a-a01149fe323d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.029279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.029279Z digest=sha256:ffc59dcdb4fdedfd6794642b5fc54ee149672bf9f2aeb0e9e758c70ab9ced7f7

Observation 1919acb8-38ab-48f9-abdf-90f621f7af2a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.033836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.033836Z digest=sha256:4895d97728cb91543579a1ee751e7c85ac4532a71db8238205898e09174975ee

Observation 00deef47-f46a-4d61-862a-2d659f2a2e3d · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.986905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.038639Z digest=sha256:b323fdc231c4eb0065a2ee03e6ffe1706b7f1eeeddb92ebde7d0a437951e0699

Observation ffb52276-4dbf-42a5-a8a7-6188c612ccf3 · outbound

This paper cites InstructBLIP: Towards general-purpose vision- language models with instruction tuning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InstructBLIP: Towards general-purpose vision- language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.974761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.042409Z digest=sha256:8ace79dc5e28de24fa2cdcb8e0b6c561447e38e92c81925de84d06ba27c4678b

Observation 74bc106e-3051-457d-bd6b-6fdf84bf9da0 · outbound

This paper cites ImageNet: A large-scale hierarchical im- age database.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ImageNet: A large-scale hierarchical im- age database

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.962422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.046295Z digest=sha256:5b9767ad0da9e1063e4262106f9cfc19821ebc6f2dad2ac20000111ee00382e8

Observation acce52c3-5ad4-4f5d-ab4d-d6ec5a633617 · outbound

This paper cites Scaling laws of synthetic images for model training.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Scaling laws of synthetic images for model training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.948764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.049953Z digest=sha256:49d028ad2fc9faebef1c96cf071e1a585abd1e28b5c6fc3dae62a63345738e09

Observation 6958ca54-6c05-485d-bb6d-d0752481924a · outbound

This paper cites EV A: Exploring the limits of masked visual represen- tation learning at scale.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EV A: Exploring the limits of masked visual represen- tation learning at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.936580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.053556Z digest=sha256:2873e935d15f1ac4c8495e977959590900ac651d8328549fb25bda4d6e20fdad

Observation f028fc6e-4b57-478f-87fb-90e4505b9d39 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.057433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.057433Z digest=sha256:31b46838b3377cb7336d41a7c97b49d149c0442e19e587465d6f8604f268acee

Observation 8e81c1a8-f204-4538-aa07-3588ac598431 · outbound

This paper cites Gemini., 2023.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Gemini., 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.924269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.061609Z digest=sha256:c706154ed1da9d790029900ec7153d1e4064dbea2403877529a75d9c277f1b76

Observation 6413d14b-cc83-41c8-b741-b53012ff8f07 · outbound

This paper cites Hal- lusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Hal- lusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.912168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.065609Z digest=sha256:2bcb6692301b54ac4d053c00f88dd60405c905265c0b457cc28616acd2c0e216

Observation 3c0226eb-72d7-4783-b93b-32da6202be57 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.069305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.069305Z digest=sha256:1facdbbbaf0dc6ddadc34669a94a1f0ede4665b43aec951256594c5f21370c9f

Observation bf354ab2-46c2-4c6d-bda0-fce9785b7fb3 · outbound

This paper cites SEED-Bench: Benchmarking multimodal large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? SEED-Bench: Benchmarking multimodal large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.898939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.073147Z digest=sha256:19c8b5bbbd3507de0b65ba11b5d8acc6b80cea1a9013049e35e2b84108d6e374

Observation 9e0f7929-dcc2-4b94-94b7-2719b23accc1 · outbound

This paper cites Naturalbench: Eval- uating vision-language models on natural adversarial sam- ples.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Naturalbench: Eval- uating vision-language models on natural adversarial sam- ples

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.885960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.076856Z digest=sha256:41451780bad5ba00b7cd4a4bbaf4d09314ef0e500c40c5f4b835fbddd2248940

Observation ff023b8e-d03c-49d6-b83c-0884b3184ae7 · outbound

This paper cites LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild, 2024.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaV A-NeXT: Stronger LLMs supercharge multimodal capabilities in the wild, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.873904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.080548Z digest=sha256:be2baff92648a8efa39fb2e8bb76d57b2a03a7740af0b564aab1e75561199cdb

Observation 46ab11c0-667f-421e-85ff-310e912d9999 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.861617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.084349Z digest=sha256:b95e4c385f96229397309d7a2c960fef975e47d110248585904492f0fc14eed5

Observation 8171e875-5c42-4ad9-b087-14455198dcc0 · outbound

This paper cites FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.848224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.088284Z digest=sha256:5f2cbf8ddf48081dfe771e0510de410ea55ddc267562e388b8ce9ad5ebea26c5

Observation e336f4b7-b409-40ee-b9c5-38effed72c9f · outbound

This paper cites Evaluating object hallucination in large vision-language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Evaluating object hallucination in large vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.834662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.092028Z digest=sha256:a68f3e77480fd85c7b5563254c2a4200f1181714eff61b78938534826e5f5f71

Observation aa38c0a9-84f6-4197-82be-97650e95f758 · outbound

This paper cites Visual instruction tuning.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.725556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.096069Z digest=sha256:3a90e9c2997be41b567838f727e265cdfa70e6ec453f8702d27ab97c33251e8a

Observation fcbe0d19-002d-48d1-b745-40cae90a2cd5 · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? In Proc.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMBench: Is your multi-modal model an all-around player? In Proc

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.712165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.099764Z digest=sha256:a8475db7f7e7d0b293a956e0a2a12fcf143b4dcf1a604b7d17d4e8e385116ca4

Observation f3b36e6e-ba15-4cde-b70b-14c99e052a24 · outbound

This paper cites From here to human-level AI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? From here to human-level AI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.699235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.104119Z digest=sha256:d2a78700e57e32117ee2b6b9843cdb83ea2a0b960bf76199120ed9f23844afe9

Observation 1a8c0e8a-bdf5-41f1-a0cf-5c0bd195d76b · outbound

This paper cites Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.107976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.107976Z digest=sha256:ec1d72d2d6e8886a8985cb978e640b62b99973ed97f3a1b0f3dacf7d1a3a5a74

Observation c49d540f-1f35-4ffc-9263-f5213194d74c · outbound

This paper cites Position: Levels of AGI for operationalizing progress on the path to AGI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Position: Levels of AGI for operationalizing progress on the path to AGI

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.687401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.112023Z digest=sha256:bc0791aff0935e4c343bb0dac9d786ec4ec61016c5bc1aa8cab72beae844c20c

Observation ffeb8b1c-da57-40a3-b4ff-3d21227f552e · outbound

This paper cites GPT-4o, 2024.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? GPT-4o, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.675430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.115794Z digest=sha256:3e73e5e3259dc5482f99f7c79a0f11170e11cf2e2cfb05af6d82de5500b4642f

Observation b80c38a8-aee4-464f-a679-dbe8ac3a27d2 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Learn- ing transferable visual models from natural language super- vision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.663227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.119566Z digest=sha256:9f31ac5fd0bbbb3f724a378409226118c11bae9fab32a08684229ca6d4e4f1dc

Observation 47900640-6151-4743-9d30-7e73eb74b853 · outbound

This paper cites Zero-shot text-to-image generation.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Zero-shot text-to-image generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.650633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.123316Z digest=sha256:77b7598f771054eac15ab7cef73ba112ca725c4390c19ca0dcbb3bb31d2d9d39

Observation 2ecd108c-8b53-46b6-bdcb-7417fe87ef9e · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.127230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.127230Z digest=sha256:ab226c75636163660eadeeb04862a488da7e8f81f1c6681001559444120e173e

Observation ed030b23-9307-49ba-a19a-715009b5ff27 · outbound

This paper cites Link- context learning for multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Link- context learning for multimodal LLMs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.638505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.131172Z digest=sha256:2c6c7cd487328cf9f6328bfb69505c2d0b78fcd9f6702aed686e4c39907766ac

Observation aace756d-93a0-47ea-b5ba-7a43b3f73cb4 · outbound

This paper cites Label Studio: Data labeling soft- ware, 2020-2022.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Label Studio: Data labeling soft- ware, 2020-2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.625949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.134990Z digest=sha256:c67f49296eb78f1e7ab263f8dd49f6b9e4428319da99d961494350876b17a87f

Observation 79b3aa1f-b5ff-434c-9691-f332dc9bf7d9 · outbound

This paper cites Cambrian- 1: A fully open, vision-centric exploration of multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Cambrian- 1: A fully open, vision-centric exploration of multimodal LLMs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.613670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.138719Z digest=sha256:ea272bce0e728fe7bee9a0f86685da3097dafcd9dc93bf8c6eb8600d50fba74e

Observation c7dc8d40-bde3-44b3-b14b-90fd04aab1b4 · outbound

This paper cites Eyes wide shut? Exploring the vi- sual shortcomings of multimodal LLMs.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Eyes wide shut? Exploring the vi- sual shortcomings of multimodal LLMs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.600539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.142487Z digest=sha256:40b56f3dbe0e970c50e45e8f659ecd44208100abf0597ac1ab2e6196c7ebe01a

Observation a603c976-d34d-4b55-9bf0-0b714c5c1afc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.146242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.146242Z digest=sha256:f54cc340950b8de8b027a1893140c8a77fa8781d7e6f154c97ef4f9796577b16

Observation a914c956-324d-478b-ac7e-0e1a1f32ba7d · outbound

This paper cites Recent advancements in fruit detection and classifica- tion using deep learning techniques.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Recent advancements in fruit detection and classifica- tion using deep learning techniques

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.588119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.150418Z digest=sha256:00c47321c709fee173fb180fb24dd8f14b37c7bd7f983e556fcb9e1fd9a0129c

Observation b55d1ac3-ea59-4cc8-94fb-06298f4b59f8 · outbound

This paper cites Le, Thang Luong, and Golnaz Ghiasi.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Le, Thang Luong, and Golnaz Ghiasi

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.575746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.154142Z digest=sha256:2b57b08ee7684d041b756068874b7d0f035411ce06aabf0bba1e487bd63884bb

Observation faa5d3f1-5b31-417b-95bf-b4231f2d58ca · outbound

This paper cites xGen- MM (formerly BLIP-3): A family of open large multimodal models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? xGen- MM (formerly BLIP-3): A family of open large multimodal models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.158345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.158345Z digest=sha256:9dca61964e6a8d19eea594acf221452462e12dcea8ef3df2e50bc4c6a5c7571f

Observation 1bf36864-c24d-45ee-b855-05555cb4c46b · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.162061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.162061Z digest=sha256:bddc2df0ef9c799627bd0e609ed2b8e416d49d350ffc8aac803b36d94721108c

Observation 5a9af856-ff51-4cec-a2a9-398e3f64cbae · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Yi: Open Foundation Models by 01.AI

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.166131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.166131Z digest=sha256:bde5333e48b3ed8840b09031e13e4ca80386733eb1e259d9a58b613b4d03a59a

Observation 74f2d605-937b-4203-a357-141d63b74a09 · outbound

This paper cites MMMU: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for Expert AGI.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for Expert AGI

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.562615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.170068Z digest=sha256:521302739432c2f94e1ffeec151ff960cb3c903ba4939508e754f1e834002071

Observation 81cedbb2-41d0-4939-9e2c-3a3ce387579c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.173837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.173837Z digest=sha256:53fcf0f94e3d9a8cadf093790eed3c681a6e4779f354c66b73e8847625f1da44

Observation b8aa4ff3-da1c-4cf7-81b8-027479da32d1 · outbound

This paper cites Sigmoid loss for language image pre-training.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? Sigmoid loss for language image pre-training

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.549109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.177834Z digest=sha256:46e778ac8633320c5f5c65cf33243e10f857098253be3572691221848a7bc2f6

Observation cadb536a-be2d-4ff9-b053-f671129816ee · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.181590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.181590Z digest=sha256:60ae99c0454f6f233db9e079aa92f0145ae2540e610e412934fe126a90ad7123

Observation aad8ed8e-a71d-4e96-8c3e-28895f934dcc · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:29.185644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:29.185644Z digest=sha256:5c59a5bcba33ce1378e35c29e14575cfee54cb5dc30e738a24f69e1b2668ced6

Observation 2e0c93ba-4e6b-4e1c-ba84-fb4350a2a774 · outbound

This paper cites On evaluating ad- versarial robustness of large vision-language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? On evaluating ad- versarial robustness of large vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.536158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.189782Z digest=sha256:3ef315b6abb465d39031fcb4deaeea4d4f54073b64afcae555f5794e3165641f

Observation c3dfd032-6ed0-46f1-8e59-f7b0b56f4b73 · outbound

This paper cites ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? ROME: Evaluating pre-trained vision-language models on reasoning beyond visual common sense

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.523316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.193661Z digest=sha256:77f49afe297c74e43456197ce56399ce43f3e7c97fba25a682e02b7fbbca9e2e

Observation 84c5baff-60b7-4802-a6df-9bcfe455bf38 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:29.510486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.197392Z digest=sha256:a7d52fedc7c8d59f6855de12e616aeef8a61f19037c7626bbb76b3a62e7a1688

Observation 7053f530-aa0c-49e2-b1da-d1771aef3df6 · outbound

This paper cites dall-e-3.

NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects? dall-e-3

Reference 2024

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:57:29.496650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T11:57:29.201224Z digest=sha256:d877df8b9e4911ffe5ad884e2d8774ab3d63639b68ddc37df0202549fa94c55a

Pith citing papers

Observation 12c5c272-28de-4f42-aea7-07925d6fb9d8 · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.021149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:fd8a40fdce12b5f21ab09999e5efdb33346969e82eec4a2babc08834a74513ad