Pith. sign in

Paper Citation Record · LEDGER

Efficient Multimodal Learning from Data-centric Perspective

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2402.11530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11530 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:54:10.967009Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10fc5fcb-2c2c-449c-bf9d-7a35a5b434e0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Efficient Multimodal Learning from Data-centric Perspective

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.158828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:3aa4fea11c9e8d5c4f81b0e7ec8fcd2bf0336c0fda4b4f9bfa35aa93804696ac

Observation 5dfdbce2-9602-4248-8841-61d285c8c824 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Efficient Multimodal Learning from Data-centric Perspective

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.058584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:b7418c82db39f7364f6dc8f6387d0318f3a8d173156a6a64a11f2764a2fbe7a8

Observation 452d07d0-b83c-43a4-a3bf-d1f59e28ffea · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Efficient Multimodal Learning from Data-centric Perspective

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.169541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:e661580f65a5f80358be08261c094e4cdf01d21b131c0a08a010cea3261c9f3d

Observation daddf052-2faa-4ec3-83a3-d312df0145b0 · inbound

JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging cites this paper.

JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging Efficient Multimodal Learning from Data-centric Perspective

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T20:12:31.121001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:12:31.121001Z digest=sha256:88cd18be1cb5378035dc39bab7a997f2c049ba6e4a87011e573591ee1c1c7f66

Observation 711b46af-a74a-47e0-a91f-b5a38aca166c · inbound

Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs cites this paper.

Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:19:45.669568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:19:45.669568Z digest=sha256:a586b202a399400f3fd5c6405563dd792f34f09d175bd3db2ffe73ae9ffb6c7a

Observation 34c35a34-0fa4-4e58-84df-5447a98f54df · inbound

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body cites this paper.

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body Efficient Multimodal Learning from Data-centric Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:29.462359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:29.462359Z digest=sha256:949d4f1b04068781f3d609102d757ca3fae939f991a5f3669e4b44de4352ff65

Observation 5c86cc74-c1f6-4418-8184-75d4a576fdc4 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.797919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.797919Z digest=sha256:7bb38d66fe5def04b7a80b78f9af9c6db207050058eee919f650641ae4079dc4

Observation 98ace428-1e24-4cbc-9f80-160850e8a0ec · inbound

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models cites this paper.

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:25:25.509406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:25:25.509406Z digest=sha256:9c48a593a2fe8c2c640a80e7dff3f83e65f7d70bcf8327933914e8af1510b9de

Observation 4a2795e1-08c3-4929-849b-af1536e38e24 · inbound

MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models cites this paper.

MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:01:53.047847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:01:53.047847Z digest=sha256:dee3151a7731385f30e7bc0105aa5003757afce857f490ef3d930aaf284db6d9

Observation 47b76b57-8d5f-4f02-a2b6-6ebd0eaad57c · inbound

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation cites this paper.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Efficient Multimodal Learning from Data-centric Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.276775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.276775Z digest=sha256:3c8662df34d05fa49ed1b5df82e529a20a41eed8e53ce8b46ca10ed3d70a795f

Observation 5ca43d82-4d3f-4ee4-a754-c7936e043c96 · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.550329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.550329Z digest=sha256:355e15f539968a0bebbccf70e630417716ca10040f10830190e955d4d87d6db5

Observation 81c01521-6f92-436a-b4cf-da09f337d506 · inbound

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness cites this paper.

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness Efficient Multimodal Learning from Data-centric Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:54:51.944375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:54:51.944375Z digest=sha256:d0e3233b6e8093e8cc0b4c546867a9b0849954c8bebdefc85549cf4a7c167eb0

Observation f57083ff-43df-484f-8131-e01d1f431a9f · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.024131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.024131Z digest=sha256:8e4a98f7fb23f8ed8c5e4d66d0f7685685f13a5e1423710b67eb2b079f4d5686

Observation d25ff837-5e1c-4d99-b73a-745f915868a2 · inbound

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference cites this paper.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Efficient Multimodal Learning from Data-centric Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.542968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.542968Z digest=sha256:9431a4f5ca2486a208c8ab7443a19c96477b9f5f9425f371807fb001bacab9ab

Observation 85a81cd4-3ead-40de-8f05-9826301677fe · inbound

Unveiling the Mystery of Weight in Large Foundation Models: Gaussian Distribution Never Fades cites this paper.

Unveiling the Mystery of Weight in Large Foundation Models: Gaussian Distribution Never Fades Efficient Multimodal Learning from Data-centric Perspective

Reference 645

Resolution
unresolved
no resolver link, observed 2026-08-10T19:20:24.129770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:20:24.129770Z digest=sha256:ac59c1471e68edc429ddaa3647c74ef7fd5897437e6554b3882567c0768d5ac0

Observation 94764db4-f09f-4834-9d11-812e5a47907d · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Efficient Multimodal Learning from Data-centric Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.531656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.531656Z digest=sha256:b4b4dc1858cdfa6a2ec99874b3b73c1c101dec6d27892fd3ab6bb4951def5379

Observation 5052c314-ae4e-4cd6-8863-d79a84a42cae · inbound

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs cites this paper.

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T11:53:55.414957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:53:55.414957Z digest=sha256:844cad3922d1c08648c7550240e38b450e933b43a3a29f12c520b3062e61bba9

Observation 54e2758a-eba0-4e40-b5fd-a0e98bd63d26 · inbound

FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding cites this paper.

FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding Efficient Multimodal Learning from Data-centric Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:54:10.967009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:54:10.967009Z digest=sha256:220a1242791f89e5854b6dd5dfcc23f6dad4768f5d0631fc1e116637ab2a4357

Observation 2c8162b7-1d8f-42ab-954f-5ad194dd303f · inbound

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models cites this paper.

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:03.895366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:21:03.895366Z digest=sha256:17dd0717155db002f11dd518c59e665d931d313ea170a298d42134d0df7e6953

Observation 68ae5757-6f11-422c-a62d-ccb59f4426d7 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:20.145924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:20.145924Z digest=sha256:f993e55783593bf1246d737d0709ee7956bbdf6b72959b9dc356d515e92fd249

Observation 97b4e0f1-3f29-4140-b356-8d4085e4f5d3 · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.920866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.920866Z digest=sha256:a18f5266751ee5e456c6f6aea07837b0554152b7b453378f5b01cf2df6d4e17b

Observation 344af88c-0b80-4af1-9c1c-219d6367009c · inbound

LMM-Det: Make Large Multimodal Models Excel in Object Detection cites this paper.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Efficient Multimodal Learning from Data-centric Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.164283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.164283Z digest=sha256:b43c5b2be0607e0c231cf5058c26c01ee9fed7c6dbd0da8fdf1f2720d4654c66

Observation e6e16ed1-ea5d-45d0-a58b-b8b29de3acce · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:45.165347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:45.165347Z digest=sha256:d69f9c7621305503e628d1b13f625b7ac5540fce9b2cf8399b92ded118be44c3

Observation b3ce6c97-49e4-41d6-96eb-d9643dce6b3b · inbound

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models cites this paper.

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models Efficient Multimodal Learning from Data-centric Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:35.876386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:49:35.876386Z digest=sha256:32a998defcd8cf6faa54fe643afb49b0fec5895d19db3b869b47dd77dd4460fa

Observation f088ccac-837a-4e8c-bd5f-478b8ee0fe35 · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Efficient Multimodal Learning from Data-centric Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.894715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.894715Z digest=sha256:b04985758b1cbaedded8e7645e7c6f82d94e16c044359750dde0e725f25ab0fc

Observation 82da0299-3d75-4c2d-ab49-b347be8f59b6 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.053118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.053118Z digest=sha256:62f453699d602c6b0939da178fec107541373a8057d5843f2e7ecf9773c9de7f

Observation 453a9f6f-2842-496c-8795-b92fff2e9f58 · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.117667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T08:17:29.924860Z digest=sha256:1e6c26ed392a6d37af19491e0761a63d0c4d7b6eca657c8797e7065d81e01d7d

Observation 53763547-4478-4476-bde3-62e380ba8eca · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:34:27.393703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:34:27.393703Z digest=sha256:b8cb87c379610a1a1db28c31e344aee34dac54bb1de2c6b9561305f6fd1fdaa2

Observation 5f13addc-ae1c-4578-a725-6069109847f4 · inbound

Leaderless Collective Motion in Affine Formation Control over the Complex Plane cites this paper.

Leaderless Collective Motion in Affine Formation Control over the Complex Plane Efficient Multimodal Learning from Data-centric Perspective

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T09:18:06.363249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:18:06.363249Z digest=sha256:d3323cd723317b778bb26e8a5f8dbba3bf846afe8d364a2381c680addb9d593b

Observation 079bb065-94a4-430c-97a7-a60729d7dc27 · inbound

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis cites this paper.

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis Efficient Multimodal Learning from Data-centric Perspective

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:51.452284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T20:01:47.592716Z digest=sha256:bb25f187fe861d1d5089c0a7b3dd624939c6a2c13a3893bcdeeeb8a1f83a0fbf

Observation e7156161-bfb2-415a-a446-535bd84dbd21 · inbound

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency cites this paper.

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:56:14.224287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T03:47:38.100037Z digest=sha256:e86d6f4830add1d65a33368ca7b585fc271db4c897676313900370b8dadcd9c6

Observation bb565ed3-7b21-4339-a475-b697be2ce694 · inbound

Anisotropic Modality Align cites this paper.

Anisotropic Modality Align Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.270268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:08:04.686087Z digest=sha256:b691f075cabb85739b6609aa26a1e4791eaa011d2aab4b1a81ef1bd85ae7079c

Observation 94ba5713-8354-48f0-8240-1361bc146040 · inbound

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models cites this paper.

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:25.660081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T05:10:56.991368Z digest=sha256:581cde3aafad43a2d57e55146bc2682acc36e88470a38f95a6542a88a01a0c81

Observation 7fa6a991-7251-46d8-b2d6-952b5690e897 · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens Efficient Multimodal Learning from Data-centric Perspective

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.521871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:aacfc9cde6dab2e1ce8476f5faa44d6724bf5c4318824fdb4c6549c9890c2b39

Observation ca881d48-ae33-4ddb-9937-4b00bb06aebe · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:09:38.228602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:4c2be269f813bdf1d9d88ce8eec0006ea3829385b33f468aecd934024e050130

Observation 252bcedb-73af-4ce7-9495-d38b11e74add · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Efficient Multimodal Learning from Data-centric Perspective

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.009876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:4e05f29ebfbbaa02411641746486a74da47251ca84126bf4ae37e8cdb13a2287

Observation a4f72ba6-9734-4009-a020-ac34093c6fbe · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Efficient Multimodal Learning from Data-centric Perspective

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.414077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:4d04136c7672a7a802febab340f418f2f6c612d808e3d39f72c578ff262f16bb

Observation 7d997c3f-dce5-4477-9ff6-d92ba060c17d · inbound

Bridging the Modality Gap in Forensic Image Retrieval cites this paper.

Bridging the Modality Gap in Forensic Image Retrieval Efficient Multimodal Learning from Data-centric Perspective

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.986808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T10:06:22.362644Z digest=sha256:946316735b4e0fbada63f32ba90c7207cc111a58296e77dda20833fa87c59dc7

Observation 7f43068c-9afa-42d8-9621-b03a69f8ea93 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Efficient Multimodal Learning from Data-centric Perspective

Reference 278

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.477800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:f49f10828c473f27fb448c9b7d0ab30d17fe9226b4b401ea8f885831e0abb467

Observation 6edb94b7-7acc-4fc2-b394-41b60f279ce5 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Efficient Multimodal Learning from Data-centric Perspective

Reference 278

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.576159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.576159Z digest=sha256:cd2fce95bfbc3d757d6bc1d1cb7f6dd11fb1a14dd3c3d8ed259ee1ee76b24608

Observation 1373dba2-dee2-4b26-8797-591687a27972 · inbound

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing cites this paper.

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing Efficient Multimodal Learning from Data-centric Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:55.017604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:55.017604Z digest=sha256:02b64c6f3a715ef7811c775290fdd3c023d6a46e2a665b70380173d29ad4afba

Observation 7e115841-d1f7-4f52-aae3-970b7392d232 · inbound

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models cites this paper.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.072148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.072148Z digest=sha256:a6802e88a703f85c0aee5d0324651d638a9485b8aa7f3ab05cda595c01ca0727