Pith. sign in

Paper Citation Record · LEDGER

Efficient Multimodal Learning from Data-centric Perspective

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2402.11530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11530 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:54:10.967009Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10fc5fcb-2c2c-449c-bf9d-7a35a5b434e0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Efficient Multimodal Learning from Data-centric Perspective

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.158828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:3e02b8214673312317567ce1cc005d3c8273873cac95f0117afb60eafde54dc1

Observation 5dfdbce2-9602-4248-8841-61d285c8c824 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Efficient Multimodal Learning from Data-centric Perspective

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.058584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:5b0ff6b0e96c3f4f59bde53ddb5571c1d3d0018955dd283d24ca554bae6383e8

Observation 452d07d0-b83c-43a4-a3bf-d1f59e28ffea · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Efficient Multimodal Learning from Data-centric Perspective

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.169541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:3b663056f425a809be67e968df0c33484c8f3046e8686a637cf630f8e943793a

Observation daddf052-2faa-4ec3-83a3-d312df0145b0 · inbound

JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging cites this paper.

JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging Efficient Multimodal Learning from Data-centric Perspective

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T20:12:31.121001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:12:31.121001Z digest=sha256:88cd18be1cb5378035dc39bab7a997f2c049ba6e4a87011e573591ee1c1c7f66

Observation 711b46af-a74a-47e0-a91f-b5a38aca166c · inbound

Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs cites this paper.

Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:19:45.669568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:19:45.669568Z digest=sha256:a586b202a399400f3fd5c6405563dd792f34f09d175bd3db2ffe73ae9ffb6c7a

Observation 34c35a34-0fa4-4e58-84df-5447a98f54df · inbound

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body cites this paper.

Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body Efficient Multimodal Learning from Data-centric Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:29.462359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:29.462359Z digest=sha256:949d4f1b04068781f3d609102d757ca3fae939f991a5f3669e4b44de4352ff65

Observation 5c86cc74-c1f6-4418-8184-75d4a576fdc4 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.797919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.797919Z digest=sha256:7bb38d66fe5def04b7a80b78f9af9c6db207050058eee919f650641ae4079dc4

Observation 98ace428-1e24-4cbc-9f80-160850e8a0ec · inbound

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models cites this paper.

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:25:25.509406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:25:25.509406Z digest=sha256:9c48a593a2fe8c2c640a80e7dff3f83e65f7d70bcf8327933914e8af1510b9de

Observation 4a2795e1-08c3-4929-849b-af1536e38e24 · inbound

MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models cites this paper.

MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:01:53.047847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:01:53.047847Z digest=sha256:dee3151a7731385f30e7bc0105aa5003757afce857f490ef3d930aaf284db6d9

Observation 47b76b57-8d5f-4f02-a2b6-6ebd0eaad57c · inbound

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation cites this paper.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Efficient Multimodal Learning from Data-centric Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.276775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.276775Z digest=sha256:331af297daaa9bdac9ed817f86d1f1fe8d039c88cf3b7273f1d0351ec86921b3

Observation 5ca43d82-4d3f-4ee4-a754-c7936e043c96 · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.550329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.550329Z digest=sha256:355e15f539968a0bebbccf70e630417716ca10040f10830190e955d4d87d6db5

Observation 81c01521-6f92-436a-b4cf-da09f337d506 · inbound

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness cites this paper.

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness Efficient Multimodal Learning from Data-centric Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:54:51.944375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:54:51.944375Z digest=sha256:d0e3233b6e8093e8cc0b4c546867a9b0849954c8bebdefc85549cf4a7c167eb0

Observation f57083ff-43df-484f-8131-e01d1f431a9f · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.024131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.024131Z digest=sha256:8e4a98f7fb23f8ed8c5e4d66d0f7685685f13a5e1423710b67eb2b079f4d5686

Observation d25ff837-5e1c-4d99-b73a-745f915868a2 · inbound

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference cites this paper.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Efficient Multimodal Learning from Data-centric Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.542968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.542968Z digest=sha256:9431a4f5ca2486a208c8ab7443a19c96477b9f5f9425f371807fb001bacab9ab

Observation 85a81cd4-3ead-40de-8f05-9826301677fe · inbound

Unveiling the Mystery of Weight in Large Foundation Models: Gaussian Distribution Never Fades cites this paper.

Unveiling the Mystery of Weight in Large Foundation Models: Gaussian Distribution Never Fades Efficient Multimodal Learning from Data-centric Perspective

Reference 645

Resolution
unresolved
no resolver link, observed 2026-08-10T19:20:24.129770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:20:24.129770Z digest=sha256:ac59c1471e68edc429ddaa3647c74ef7fd5897437e6554b3882567c0768d5ac0

Observation 94764db4-f09f-4834-9d11-812e5a47907d · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Efficient Multimodal Learning from Data-centric Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:17.531656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:17.531656Z digest=sha256:b4b4dc1858cdfa6a2ec99874b3b73c1c101dec6d27892fd3ab6bb4951def5379

Observation 5052c314-ae4e-4cd6-8863-d79a84a42cae · inbound

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs cites this paper.

CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T11:53:55.414957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:53:55.414957Z digest=sha256:844cad3922d1c08648c7550240e38b450e933b43a3a29f12c520b3062e61bba9

Observation 54e2758a-eba0-4e40-b5fd-a0e98bd63d26 · inbound

FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding cites this paper.

FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding Efficient Multimodal Learning from Data-centric Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:54:10.967009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:54:10.967009Z digest=sha256:220a1242791f89e5854b6dd5dfcc23f6dad4768f5d0631fc1e116637ab2a4357

Observation 2c8162b7-1d8f-42ab-954f-5ad194dd303f · inbound

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models cites this paper.

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:03.895366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:21:03.895366Z digest=sha256:17dd0717155db002f11dd518c59e665d931d313ea170a298d42134d0df7e6953

Observation 68ae5757-6f11-422c-a62d-ccb59f4426d7 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:20.145924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:20.145924Z digest=sha256:f993e55783593bf1246d737d0709ee7956bbdf6b72959b9dc356d515e92fd249

Observation 97b4e0f1-3f29-4140-b356-8d4085e4f5d3 · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.920866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.920866Z digest=sha256:a18f5266751ee5e456c6f6aea07837b0554152b7b453378f5b01cf2df6d4e17b

Observation 344af88c-0b80-4af1-9c1c-219d6367009c · inbound

LMM-Det: Make Large Multimodal Models Excel in Object Detection cites this paper.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Efficient Multimodal Learning from Data-centric Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.164283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.164283Z digest=sha256:b43c5b2be0607e0c231cf5058c26c01ee9fed7c6dbd0da8fdf1f2720d4654c66

Observation e6e16ed1-ea5d-45d0-a58b-b8b29de3acce · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:45.165347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:45.165347Z digest=sha256:d69f9c7621305503e628d1b13f625b7ac5540fce9b2cf8399b92ded118be44c3

Observation b3ce6c97-49e4-41d6-96eb-d9643dce6b3b · inbound

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models cites this paper.

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models Efficient Multimodal Learning from Data-centric Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:35.876386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:49:35.876386Z digest=sha256:32a998defcd8cf6faa54fe643afb49b0fec5895d19db3b869b47dd77dd4460fa

Observation f088ccac-837a-4e8c-bd5f-478b8ee0fe35 · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Efficient Multimodal Learning from Data-centric Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.894715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.894715Z digest=sha256:b04985758b1cbaedded8e7645e7c6f82d94e16c044359750dde0e725f25ab0fc

Observation 82da0299-3d75-4c2d-ab49-b347be8f59b6 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.053118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.053118Z digest=sha256:62f453699d602c6b0939da178fec107541373a8057d5843f2e7ecf9773c9de7f

Observation 453a9f6f-2842-496c-8795-b92fff2e9f58 · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.117667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T08:17:29.924860Z digest=sha256:65075c19e4f04cdaff5730bca32fe325eaa3ab7c4571580f6c49ba551798f669

Observation 53763547-4478-4476-bde3-62e380ba8eca · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:34:27.393703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:34:27.393703Z digest=sha256:b8cb87c379610a1a1db28c31e344aee34dac54bb1de2c6b9561305f6fd1fdaa2

Observation 5f13addc-ae1c-4578-a725-6069109847f4 · inbound

Leaderless Collective Motion in Affine Formation Control over the Complex Plane cites this paper.

Leaderless Collective Motion in Affine Formation Control over the Complex Plane Efficient Multimodal Learning from Data-centric Perspective

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T09:18:06.363249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:18:06.363249Z digest=sha256:d3323cd723317b778bb26e8a5f8dbba3bf846afe8d364a2381c680addb9d593b

Observation 079bb065-94a4-430c-97a7-a60729d7dc27 · inbound

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis cites this paper.

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis Efficient Multimodal Learning from Data-centric Perspective

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:51.452284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T20:01:47.592716Z digest=sha256:a4bcf4c2c437e5154e1eac87d349d2b9e9fe376b2c0c8da3a7ddbef85e09e8d0

Observation e7156161-bfb2-415a-a446-535bd84dbd21 · inbound

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency cites this paper.

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:56:14.224287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T03:47:38.100037Z digest=sha256:80bcccaf11e69ed06cb74adf1260b10e1725102c826f984988e6c3f78d0186ef

Observation bb565ed3-7b21-4339-a475-b697be2ce694 · inbound

Anisotropic Modality Align cites this paper.

Anisotropic Modality Align Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.270268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:08:04.686087Z digest=sha256:64aa36ddc81a7cb22ad9fc5647d46908365ef41802dd9e3079ff1456b1fa67b2

Observation 94ba5713-8354-48f0-8240-1361bc146040 · inbound

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models cites this paper.

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:25.660081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T05:10:56.991368Z digest=sha256:024e7d4e4e0ae1814dddf6d511c9850667bf289798ab0031426615e918766827

Observation 7fa6a991-7251-46d8-b2d6-952b5690e897 · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens Efficient Multimodal Learning from Data-centric Perspective

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.521871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:881d016745713bc938cda0f982a6c24bc5529fd0fec3a4d06de7435aadb3f841

Observation ca881d48-ae33-4ddb-9937-4b00bb06aebe · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens Efficient Multimodal Learning from Data-centric Perspective

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:09:38.228602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:b075bc260195f2be89283823146cc24413e426738e576c68ab35b1b0c0511490

Observation 252bcedb-73af-4ce7-9495-d38b11e74add · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Efficient Multimodal Learning from Data-centric Perspective

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.009876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:bdf7c53a1ed6afcb192aec2b5cc0929a81bb1d16c721bcf5373d07407b1866c9

Observation a4f72ba6-9734-4009-a020-ac34093c6fbe · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Efficient Multimodal Learning from Data-centric Perspective

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.414077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:9b8a976bf8fde08cfd79845060bef6e39ea4b787e3f09cc73bd90eae8127b53d

Observation 7d997c3f-dce5-4477-9ff6-d92ba060c17d · inbound

Bridging the Modality Gap in Forensic Image Retrieval cites this paper.

Bridging the Modality Gap in Forensic Image Retrieval Efficient Multimodal Learning from Data-centric Perspective

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.986808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T10:06:22.362644Z digest=sha256:d68c5a364d0efe7e405be3117e667724ce6f9dabbdcbdb95c17c08975ce2376c

Observation 7f43068c-9afa-42d8-9621-b03a69f8ea93 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Efficient Multimodal Learning from Data-centric Perspective

Reference 278

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.477800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:b6f824adb96dc06d74ae538a7b80b3b519145f17849bdbda5e0d147d850cf865

Observation 6edb94b7-7acc-4fc2-b394-41b60f279ce5 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Efficient Multimodal Learning from Data-centric Perspective

Reference 278

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.576159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.576159Z digest=sha256:cd2fce95bfbc3d757d6bc1d1cb7f6dd11fb1a14dd3c3d8ed259ee1ee76b24608

Observation 1373dba2-dee2-4b26-8797-591687a27972 · inbound

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing cites this paper.

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing Efficient Multimodal Learning from Data-centric Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:55.017604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:55.017604Z digest=sha256:02b64c6f3a715ef7811c775290fdd3c023d6a46e2a665b70380173d29ad4afba

Observation 7e115841-d1f7-4f52-aae3-970b7392d232 · inbound

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models cites this paper.

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models Efficient Multimodal Learning from Data-centric Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:55:50.072148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:55:50.072148Z digest=sha256:a6802e88a703f85c0aee5d0324651d638a9485b8aa7f3ab05cda595c01ca0727