Pith. sign in

Paper Citation Record · LEDGER

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2412.04903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04903 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:15:45.548297Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T15:05:21.907878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:08:02.074896Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c20a20e2-9235-4bb3-8daf-3390873d5016 · outbound

This paper cites GPT-4 Technical Report.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.192262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.192262Z digest=sha256:84ecaced2dafbb4c5a4e69c00af8d4f7bdb6064c90c6bc6215f2862074cb3b6c

Observation cfa5f7cc-b9ea-44ce-b6da-5432b3cd586a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.198801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.198801Z digest=sha256:44eb6c84d64820f9804901bef84a88ecf70d8310ece6f956e806241b41fba1ca

Observation 6075f1ff-8257-44bd-b27f-71af6e04cae8 · outbound

This paper cites Introducing our multimodal models, 2023.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Introducing our multimodal models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.204734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.204734Z digest=sha256:afbda3aa17772718a022d0bd2194a1897289dbabdd9b34320747df10a3474228

Observation c4122eb2-7ede-4595-85b8-ff6065f0956c · outbound

This paper cites Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.210468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.210468Z digest=sha256:ba3b2860ddbe194cca12c24bffe1baed65c82f813215e7be2f19a5315608cb34

Observation 7531db02-af9a-4a18-b63e-6703c14ef1b4 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.215944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.215944Z digest=sha256:24316b0597add5b13ba2a20ac34c5d172cb30f8e0ef1a623a74af6c0e54d189f

Observation 7bfe13d4-debb-4035-b2ac-03525b209723 · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.221696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.221696Z digest=sha256:0d9833e8a44ca0ce230edbef5be00e0954dad6b010c05c3117db90aa628e9a52

Observation 064bd64b-b1a0-49fd-9e69-17caf8363ac1 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.228427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.228427Z digest=sha256:5b0738a6436ad1fcf834e76ff40260c833ed9c1f1243a954fd93ad50cce7b543

Observation afab7f37-2681-4a98-b2e9-bb74ac27b656 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.235243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.235243Z digest=sha256:d4f17dbb371276c02a0c96e44be92c959d21652f2a3e269e6bfe64f4dd78652d

Observation 900be8f2-8b35-4dfd-9cba-ace4b33dd605 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.241288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.241288Z digest=sha256:908afbb10ce7d7f1d4809e705dab46a3d1c77afa7fed5ff515b9e3bcd3adf306

Observation 9d0e3942-1d76-4462-b1b4-731178172ab8 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.246901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.246901Z digest=sha256:9e5be124d419d83192b692ad6404a68d5830e633725dbc5620ca051fc0c1c157

Observation 278c5dde-25c1-4f08-abba-8d176d3e2e4a · outbound

This paper cites What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.252677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.252677Z digest=sha256:60e8614986f2fcaa2c777cbf9d6aaf5aeeb20745523bc2f55aa0b419714338ec

Observation 399e962d-1baf-4ad7-9d1d-5d50926b6b24 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.879880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.259223Z digest=sha256:e1371bf7fa9828eb3e0695caf9d3b9446aeb7db183fbb909526dbff9ead3dd79

Observation be302331-927a-466f-8a70-b0f5d69dee51 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.264403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.264403Z digest=sha256:8d2222a476bd86f0342034e13c4b76e68e45b71271c6364a55ee7284e6271ae0

Observation 869f7c6d-08f2-4dd3-8008-560deda00da0 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.270537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.270537Z digest=sha256:d78bc8fb5c3065ca17a5a0521b83ce720e41dda2e7616c3723dfd083dca02069

Observation 47b76b57-8d5f-4f02-a2b6-6ebd0eaad57c · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Efficient Multimodal Learning from Data-centric Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.276775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.276775Z digest=sha256:c9b29b14b01c0f795236cbb57236ead92b42e77f8dfde1a1bca35b2579620561

Observation 14d1d624-dfba-4e20-94fb-6fb692d2dfdf · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.282235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.282235Z digest=sha256:924591c59f90896055146ae9bcf299aee2b78801e3b0198ca591ba037f6f0e31

Observation 173b546f-be86-4f94-a87c-7f56937a687a · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.288162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.288162Z digest=sha256:810ae1c76f4505b07a8d81af9f2ed4bb82182d1b12119971dd13f03807d3376d

Observation c0ca5f28-10e4-461c-91b1-241142baa4e9 · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.294591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.294591Z digest=sha256:0dadd9ec5a4ad54f8c785c262999c402f95da635e74af52a6a4c40b5ccf551d3

Observation f8238795-5298-42c2-962a-696a6a35e620 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.300584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.300584Z digest=sha256:3819b8a3cbaf5bea6e4855e784a37a77e4f879f60a08a87a50fdf93ffa0a99b9

Observation 752bf7ef-9dcf-48ea-a232-c45f7efea877 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Silkie: Preference Distillation for Large Visual Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.306173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.306173Z digest=sha256:4d7e27354e0cc3c1be9b7940c8a9941af4594ea1c241f890019c0122f94e0360

Observation f2837536-35ff-452f-816f-8af47781f5a6 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.311788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.311788Z digest=sha256:647f5e0b2a859e3eed4c698b15a4f76b68e10ee318a1161cb86ec348f115ce45

Observation 50dcb1f3-f1c9-4d1a-8f0d-c8ef8a333b65 · outbound

This paper cites Red Teaming Visual Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Red Teaming Visual Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.317341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.317341Z digest=sha256:edf22861d9a1d44087277c4fec5daf022e29f48c112e9e98b1ac244a62d5731d

Observation 1a7c7d87-f8d3-4c01-9b31-ecfd53584c39 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Evaluating Object Hallucination in Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.322691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.322691Z digest=sha256:127fcec4014aedb8c8f6608ccb3d8dd6d630f2d16752ab684ed2e25fa9330a35

Observation ff5a77d4-9051-48a1-b6fe-b71348d67325 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.328671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.328671Z digest=sha256:0c5e088bed7a5809c815857cc0e93d15b4e2ff451b096b1bb2bf7074c789b0cf

Observation a89266fe-bff0-40d3-badf-f235576c5d77 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Improved Baselines with Visual Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.334538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.334538Z digest=sha256:d97cfddb347daa1b6d7021dcc4ee51834cb31ae648531d72cd56cfbff32a57df

Observation 969ef877-2a1a-4ec9-b1f6-bb1c97f6f50a · outbound

This paper cites Visual instruction tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.341133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.341133Z digest=sha256:2b18795d8b95a3681298940734f31ae690478f6e2f671b8b0e6e91f24c85d5e3

Observation e2418008-9e24-4709-b8b6-79ba61956036 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation A Survey on Hallucination in Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.347819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.347819Z digest=sha256:d2b05d8de87cc74f83cd44c33072dbce3d5be5a2f98bdf62412fb619d06d247d

Observation 689856cc-c6a7-4f5a-9617-d0f9a1bcdc0d · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.850399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.353023Z digest=sha256:d74191139221bfa063ab8c3b512da4cbf8f8bcbd520e2e978d3f7f65a07baa1a

Observation 9bc3d518-bdfe-4e85-a664-19494be0f3b7 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.359223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.359223Z digest=sha256:c9ee87463463fc369da4aeab1e26710a17303abf64d35d2ace577b7eb93ae2a6

Observation 3c91929f-c41c-443e-bb96-11fdea8b1140 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.364993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.364993Z digest=sha256:9771f3708da0cbb56b82ae8fcfc328960feaefd64447d58e9c359ed3c8d6e648

Observation 5bcec186-4edc-48ce-a99a-56c5c1b823e5 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.371287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.371287Z digest=sha256:1e2a542fa62e76f2e4494deabb6e7bec2538e6dcde7d16b264b8f26349bce26c

Observation a9ec3eef-ea7b-4fd0-b9a4-b9e280f06b30 · outbound

This paper cites Training language models to follow instructions with human feedback.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.377436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.377436Z digest=sha256:7a44aa739f027a919ac696cf1bfa9f6e44d9d3224f7f4eb4350e96ce87a16169

Observation 539e37b4-d317-4ea6-8ea7-4cc0bc67d028 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Learning transferable visual models from natural language supervi- sion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.382777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.382777Z digest=sha256:dfd6cb2053b8c39f4e14f4c658be344053cc8173370dce50bf80bddd01b447e4

Observation 46b8b64a-4161-4b9b-b325-50fc5e150257 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Direct preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.798944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.388405Z digest=sha256:f4bb6d278a9f6086fc70b051f2ee7eb86693d9ef2f61ff88d71886f02542be02

Observation d66c0994-55d1-4555-9588-607c3c1e22e7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.393913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.393913Z digest=sha256:db4e054138454cd0ed7b7be577d7082c68dd99927dffdf16eda8bed041d3f3a9

Observation 0a179c69-c0b6-4d63-8011-d0a23245e2af · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.399783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.399783Z digest=sha256:23d67a7e74fba229d5b0d8a0ba92148a8a5719a62b50d9b5f1b4b05cb4c22297

Observation f9d311ea-0459-4edd-9201-2f7766722622 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.405958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.405958Z digest=sha256:59e96ad8ee15a2f9180d314c4af6915f291cc8f562c9c4bbad3676be56454aba

Observation ec515631-2aea-4330-9deb-950b9b20bb3e · outbound

This paper cites Eyes wide shut? exploring the 10 visual shortcomings of multimodal llms.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Eyes wide shut? exploring the 10 visual shortcomings of multimodal llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.780892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.411814Z digest=sha256:5926ec9a7612da0dae7fde3319435b1849ee0c2b5dbc181ed5eb87a13febc55f

Observation a8c18b14-d571-498e-94b0-8c673eaac303 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.416910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.416910Z digest=sha256:f376168150081da555a49b94d3cfb63a5db7b70f78cf6d3c3feae0dba90de015

Observation b781e7cd-e719-4fb8-bdf5-c8ba61f9eb03 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.422147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.422147Z digest=sha256:868d06f22d65e95b3009a1d0ccfb54a0f09feabc2c5652d43e604db06ead0fd1

Observation b5dd3bf9-8510-42d6-8dd5-0c174b40ebe7 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.427816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.427816Z digest=sha256:11b97cbe5978d44402d2276f14e0964b5ee0e5e4693a72c70ef101ac175d0bd9

Observation 90bd0078-f692-4549-91bd-85fda6748089 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.763059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.433882Z digest=sha256:d3114537d77324eee2a0b360f4ff83f6848a25f8d0d7c35d83587b07675a4c43

Observation 673ab646-95d2-4282-8bbc-bc16281423d8 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.439011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.439011Z digest=sha256:18c6bc43a7f02b160cf3734a662adf36111561f5c2260eb3b4b17b22c5cd4918

Observation f319dd1b-8512-41cd-9fe4-7f8eedb838b4 · outbound

This paper cites Vigor: Improving visual ground- ing of large vision language models with fine-grained reward modeling.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Vigor: Improving visual ground- ing of large vision language models with fine-grained reward modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.444343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.444343Z digest=sha256:f892ad444b746ab03fe30a3468973ed2002b67e05dea717bed6161770d91af28

Observation 6e3e43e3-f35b-4de7-a02d-cb4cad39ea32 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.449491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.449491Z digest=sha256:64b76ecca571316ec118b46e9a4ada96342db0e9a3d4aac0ae357826c478ab64

Observation 1039a7d5-e991-4703-94d3-df29571e1b96 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.745492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.455553Z digest=sha256:457af12e24ba723bb82fac9edb010aadf27341aa967db579f950ea3203d580f2

Observation d9c0d487-ee6f-4416-a2bc-58167ce97727 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.460464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.460464Z digest=sha256:c642d86c29e8f8972eff1c56727739ab25779ebe1999aa38153106127d25fce2

Observation a3255b99-0504-4be1-867f-244dfc44a424 · outbound

This paper cites Self-Rewarding Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Self-Rewarding Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.466231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.466231Z digest=sha256:313962403116ce3e149916cf0e4694ff8b1effec34eb6cf0a63d9aefa77fe745

Observation 4c91f07c-3a13-41d5-aa0a-d4d3fd452b2e · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.471772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.471772Z digest=sha256:a0d0d1d90657b55ed9a26bdd698f9ac01a95490bb07797dee27b0db8da9387bf

Observation 538b0ddb-5dbf-4130-ab77-8f763dc35550 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.477475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.477475Z digest=sha256:85b3bba6bd6c7e12f2cc4f6e55d9248dd88aca7beeec1d2ed69f36a88e0eba7a

Observation eda6088e-7fec-415d-ba2a-21aaecb2a30b · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SVIT: Scaling up Visual Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.483079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.483079Z digest=sha256:31b8609d2e6ba1732ab63ee46581737a15335ac15e54435cbd0c400c1cc339c6

Observation 75cd0b7e-36f6-420e-b4de-136dfbbaaccc · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.488218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.488218Z digest=sha256:252620ad645c62970c44b2beb90c68c45450a4a66dfb23ec5a79812f582e2c9b

Observation 7b241dee-3bf3-494f-83bf-bdb857a076f1 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Calibrated Self-Rewarding Vision Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.494340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.494340Z digest=sha256:fedcaee8404bd22780caab77e0ab06c75f5549617745ac55bc2bee0be6bc8919

Observation 3ebe7466-a4a3-4a02-92d7-4503ba250d5b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.500631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.500631Z digest=sha256:9582bbeef953b51cd7347aa12097d83edbe8f446fb8aca045094bee75c8d2596

Observation 3ade1ccc-ebd7-4f56-9f90-2287aa58bc8e · outbound

This paper cites GPT-4o and our Critic model produce similar scores for responses, but they fail to identify the flaws in bad responses from the baseline LLA V A model.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation GPT-4o and our Critic model produce similar scores for responses, but they fail to identify the flaws in bad responses from the baseline LLA V A model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.728234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.506876Z digest=sha256:db5fb69512f08838244883497c43ba3c4edfe1c8ec355ba450b1269c264c0214

Observation 37566070-fac6-4220-bf83-d2c6b060ebaf · outbound

This paper cites As shown in Table 8, most of the experiment is con- ducted with prompts in rating style, apart from the ablation study presented in Section 5.3.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation As shown in Table 8, most of the experiment is con- ducted with prompts in rating style, apart from the ablation study presented in Section 5.3

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.710936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.513267Z digest=sha256:e3159a6b4cbd536048f1ae840e38e40c027a454648120b339c38985d2cc89e20

Observation 9973f9c9-347f-40a6-8ac0-b9fa44c59baf · outbound

This paper cites The training details are shown in Table 2.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation The training details are shown in Table 2

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.691045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.519074Z digest=sha256:9bea8edd7c2cd315b57ab1e5c80d648c1255e1dd0c9ce72a4db56531b498de58

Observation 95650ce3-5617-4935-b03a-8d964506304d · outbound

This paper cites Using annotated preference data, one round of preference learning is conducted on LLaV A1.5.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Using annotated preference data, one round of preference learning is conducted on LLaV A1.5

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.670611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.524962Z digest=sha256:ba09a54d3969b5d9dbd26dc007d549ec1ac3e2abaac2dca94e6eaf21874f4467

Observation e4972ee4-a8ae-4bcb-b70a-11b9cb3e5809 · outbound

This paper cites an unresolved cited work.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:15:46.652699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.531083Z digest=sha256:e935a453ace23a3754bdf828588eb7e84efe7eec0b98b9efe83c213518ad83fa

Observation 9661aef2-010b-4c59-8e7b-598599019fca · outbound

This paper cites an unresolved cited work.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:15:46.636166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.537284Z digest=sha256:21a4de411c604ee42e612590c9a6770251a7907c80915e5b9b05ad1aae692a9c

Observation e55212cd-03c3-4adb-81a8-741fcfaca837 · outbound

This paper cites Here, we will show some examples between EACO and baseline LLaV A-v1.6- Mistral-7B in Table 9 and 10.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Here, we will show some examples between EACO and baseline LLaV A-v1.6- Mistral-7B in Table 9 and 10

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.619883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.542436Z digest=sha256:71f8fc86ff2227cb785facde0803b788d2355cba1c90d9c5f3fc99d4ce4a038f

Observation d6e4cda5-e15e-4dcd-8cd7-63f1716d6fc7 · outbound

This paper cites score:⟨total points⟩.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation score:⟨total points⟩

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.601541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.548297Z digest=sha256:db7184a12c5accd2248a6ff5122a08bacfa4eefaf1db1ddaaf0aa4a6f7134291

Pith citing papers

Observation 822144ed-1b63-4e98-b397-83b8525ce130 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:02.076951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:cd8a03238613c68c23c04e6d47b8243160aff7a86ea0721bd807077291b9709f