Pith. sign in

Paper Citation Record · LEDGER

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

As of 21 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2411.15453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15453 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:21:34.883733Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:19:24.219539Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:45:25.594440Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfdcc233-3c64-440b-a9ba-6e6676a59393 · outbound

This paper cites Transformers are Multi-State RNNs.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Transformers are Multi-State RNNs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.648846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.648846Z digest=sha256:c361f4344e9ab37fb78f6df5c052487f8f4cd328cb413ac45ada2dc1a030a976

Observation fec0190d-6334-4ec8-ad40-09fbc9f4dedb · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy , " * write output.state after.block = add.period write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.655050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.655050Z digest=sha256:9376449690be7d5f55403032c641e00103e8c44403ac8b274038ed7165ca3943

Observation 97481105-b6bb-4edd-9482-468496acff33 · outbound

This paper cites write newline.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy write newline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.659870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.659870Z digest=sha256:fd7ed84c0b251558762394eed2cb4dfb4387d864b4a248f6089a9ae2dda58e11

Observation 6cb39386-d352-4df8-b5d3-25d4ec998ebd · outbound

This paper cites GPT-4 Technical Report.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.664086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.664086Z digest=sha256:2960abe5873a4960083ea268b21f282300cbd680e685b7e03fc8fcadf00b5a6d

Observation b2883594-10bc-48cf-a758-52a2f3ecf323 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.736555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.669040Z digest=sha256:83372fc1e160ba9325314b117c287b6ed338e594e7e8e59cabc2646dfdf097db

Observation 8ba58933-ebfb-4b45-b033-fc65a2a9d328 · outbound

This paper cites Qwen Technical Report.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.673746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.673746Z digest=sha256:7b0921bec133d5635d77ed16dbd459345a388a4a5e53c5bc73c891e9adda31be

Observation 934d00c2-332b-4d18-b965-c404f8669cfa · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.678796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.678796Z digest=sha256:f45b83979fcc8c54be2be168f87210385a63b4c2e501c871fde6d88bd7fbc220

Observation 5b2b603d-3e20-47e0-91c6-e83a0827ee7f · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.684323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.684323Z digest=sha256:e85f015ae27f3aa781565e28d2af01bc66c243a35f3e7a02895c752f09dff485

Observation fd31cc0c-f7a0-4a03-93cd-82d5f2a7d105 · outbound

This paper cites PuMer: Pruning and Merging Tokens for Efficient Vision Language Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.689280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.689280Z digest=sha256:cf0a070486cd80102734a281746cc9ef2f953305acb49c614ca17a7bae814e75

Observation 033d3e43-1c10-4d41-b9d9-17738af2932d · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.708288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.693856Z digest=sha256:0d64d1d7938d71c4c6b44e5256490dc587d4737c4e1b5fdb7b1788a2ba5053a8

Observation 220e32d6-2416-4f56-b817-71c22e3d438d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.699177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.699177Z digest=sha256:e86dd39d83e31d6627910513dca03f8ed13a8c52ff5edfc368d44e9ca1c9e664

Observation 2a91a72e-b111-40a8-81fa-d69822036f7b · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.703412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.703412Z digest=sha256:71211ef24a51b4170d603cc2c791bb6c39332be8a236053ad30e6dcc202d163b

Observation af5677ec-ce38-4ff4-920b-d4aa6a29049e · outbound

This paper cites E.; et al.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy E.; et al

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.708001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.708001Z digest=sha256:dcefc6b36a1f0f68d8c28bb5926d0d457eef59faada707d492c11e1439fe9395

Observation 476d3616-7fa3-47fa-8be9-066cf295c02c · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.712391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.712391Z digest=sha256:fdae98426e6a066326f321108e0a5fd0f9cbe71d6fa831bd870558fdea3e96f7

Observation 3d81df49-44a7-46fc-be70-918bee1e5dc5 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.685668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.716326Z digest=sha256:082d412385944f30b408ecbacae9ae19c267635e240357b7faab5c98bbb5fd80

Observation 33a1fe99-d3d3-4df9-bc2e-4b8e59360e68 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.720334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.720334Z digest=sha256:921959add376c97807fb7b81787fc71eeb588a4a3fe9c6349ad67a5beaa9c130

Observation fa32e8c2-2e75-4a29-98e7-558b34d29205 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.724684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.724684Z digest=sha256:ac39daea3368f1b5338bae7ffe2762da4c7d2ff48e9be448917b0ce9bab2475f

Observation 0c8d4602-3f82-44d5-a6cb-1e098fd4b37e · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.729037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.729037Z digest=sha256:14f11edd1b2663489b83cd313c827c8dd97cbfb2338d9d6b4b840bcdce8e9b0b

Observation 5994cfb4-0c2b-47fb-b066-1e84d08e9b1f · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.665203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.734631Z digest=sha256:9b679693335dd6fdb7a696a85c51af354ee6e63e3cdc4642a3a397f2fc70e813

Observation 8fcd0e5a-a5c4-41c0-8003-fef3fe760a52 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.738551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.738551Z digest=sha256:b8abf54042722cd80315cb1560c250eb6cf1f4cd80da6d512b9cbddec1528bf6

Observation 6487c2b4-59f1-4538-aa4e-0767c51e811e · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Gaussian Error Linear Units (GELUs)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.743132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.743132Z digest=sha256:f791473ca940a598c8dab6ac30f98a0eccf14997a8e63220cb14e5816a839fef

Observation 794c8a28-881d-434b-afe8-b640b4a62ed2 · outbound

This paper cites A.; and Manning, C.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy A.; and Manning, C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.747404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.747404Z digest=sha256:b6f3e38a28bb12cf4ec180c93a871b4dc062ddd7fa7d828a3ca625810da40e35

Observation c28dd39a-3662-41fa-a368-26501a2b1cd1 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.635332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.751001Z digest=sha256:a4236129e17a9bf740f9937f2f9a7625f81988cb07e0c9953d7bb3491e6270ae

Observation c0984625-26e8-483a-8e0b-cb1c6682a783 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy AI Alignment: A Comprehensive Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.754809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.754809Z digest=sha256:ff26c87f14653d84a981f04b895bad4879b71f53beef55fcc99bd6a2f514b516

Observation 2cde3319-c7e3-47de-b200-b2d9fb282b7f · outbound

This paper cites M.; Bommarito, M.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy M.; Bommarito, M

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:21:35.623222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.758912Z digest=sha256:299f5ec336669efb0a8b8e84df67e343c193d5ef68058678ce776c79c5c4f61c

Observation 304945a7-11f1-43ba-8ea5-4def64e1f41d · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.763655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.763655Z digest=sha256:c83c9392e7551eb4a4fa7bd8925de4a8eb06da5cc494169a40b88395cc608bb3

Observation da516c66-8aba-4538-bf4a-758f49d5f36d · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.609031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.768281Z digest=sha256:a39a2f906bcaeb11dcb9648d6f0d3f2ba161b38a88caa682f1bc06a7f426c1c5

Observation d79ab744-ba17-45b0-9e0b-ab8799d6b314 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy OtterHD: A High-Resolution Multi-modality Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.771888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.771888Z digest=sha256:034a2c8abe204206e3ccf583acc072339eff0375d435ce5aa8477a02439d9a0d

Observation a7f9b2a6-da60-4150-a6a5-67274334256c · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.776149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.776149Z digest=sha256:7daadeb033515761fe0a9be3cb899130536e39aa7f9444d18a5d83d026e98489

Observation d112de80-b060-45f1-9afd-c2a66e8717f6 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.780220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.780220Z digest=sha256:b1cf55f61af5c49c81058cf04c48f20de4e32f1a0eb18122cd88266987863fd7

Observation 3fdc6136-e95c-4799-b64c-b594606e0d99 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy VideoChat: Chat-Centric Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.786706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.786706Z digest=sha256:76c432775f0a654d157198d900914875bdb3e6de557c7b244e5d24e0429be1dd

Observation 18d52d78-545d-4e99-b1cf-c99881318b9b · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.791499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.791499Z digest=sha256:9d724757604341bf3f377e1de5a142fc20480afddf5994094a2194278952c708

Observation 7e3ce87b-2785-4118-a9e7-ee615e3344d7 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.795861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.795861Z digest=sha256:9fe96a39b3efef83e6329fb9292cf81b570e660808761ba15bc8307ea9a8f33b

Observation aade3e33-5607-442c-96a5-b7dcc3742a0a · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.800147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.800147Z digest=sha256:4c4435c120fed5b0ad8e3ae2c7730048621f7c1d3f51e3c235564439004e8aa8

Observation 829a9214-bafc-46d1-b88c-9488f25397a1 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.804487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.804487Z digest=sha256:62f59473e88a43444799fc421b4aa5e7f68a9f01bf989152254a388532c939c0

Observation 1bc045fd-3104-44cc-97be-ab789dbf8c3d · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy MMBench: Is Your Multi-modal Model an All-around Player?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.808514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.808514Z digest=sha256:54869f1f253334ebad2ea4de35cf387bdd2f9be5bf3c213d9bebf2bc993d3ee3

Observation eae51f99-e8c4-43d4-9e4e-e4dd83cc4877 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.813356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.813356Z digest=sha256:1d2d42d80f567ae968a94a63cbe300dd10363a20d728a8d9ca6b485317dda75d

Observation b880f91c-ec04-4505-b538-480c237f60b2 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.817645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.817645Z digest=sha256:1632b167ff3e9294f25ae4903cf0e55939540546fe3608df0618d553b1836b2f

Observation 40649477-08fa-4b19-91f8-e22bf1433a3e · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.821523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.821523Z digest=sha256:2fdbc261098bf6e2ff961b6524188e6b62b486ea82f86fa77ff64c0890a62824

Observation 505bd8f3-378d-4145-9875-c8cc1c52adbf · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.825034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.825034Z digest=sha256:cd1249544cb4f04ebe6539c18169a384663de73bddebe899a6986a7724ea2324

Observation fb3d6a5d-b454-40fe-a1c3-d50bb490a486 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.828599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.828599Z digest=sha256:16cfcc4520a9d581067936021c372f5d6936eba3384682dd729fc9e40358c555

Observation 5942c393-4a04-4fda-9c5a-cad874f379ba · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.832084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.832084Z digest=sha256:b0c2fe7b453000efd1124a01e4c7133105d3bc5a5e612ff63a2f9afd061859e0

Observation 3f55ba0f-5e3d-4ac2-9a7b-5e88cbe8d15c · outbound

This paper cites J.; and Yan, Y.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy J.; and Yan, Y

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.836981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.836981Z digest=sha256:f94fa4cd495d09bd3d95573cda1995618bb2a55106f21c0a215ac7f6e91d1c19

Observation 2acbabdd-d441-4abf-a6de-c4091218ef70 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.841096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.841096Z digest=sha256:9dc8528f24189059e3c5e0e517f688a41dd7e4819132cfcad8c5a861b733145f

Observation c19be0be-6dea-4578-953b-8f88991e100f · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.844867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.844867Z digest=sha256:ef986eef86ee25208cf040917a5c58ecb9d90d37b708dbd632acab94047f7321

Observation a5a64b03-b089-42a2-9944-5a22f60b2b84 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy LLaMA: Open and Efficient Foundation Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.848280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.848280Z digest=sha256:2e8e406ef810d85428f95de74463aee22594ff690b3023f5b9cb378070a30112

Observation 634f3212-6bf5-41ba-b408-0c6bca286042 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.852530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.852530Z digest=sha256:5838bd45833a74920d8861842a6be4a4ac52e9d222bf72e537a4a514e7c5c671

Observation a4c9044f-8546-4431-9d3b-f534c47f891d · outbound

This paper cites Attention Is All You Need.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Attention Is All You Need

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.857038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.857038Z digest=sha256:dd1a51a8c4cfac9a75f8ab098cce5a27811175ad4db5020ea174d30a2e9214db

Observation 6b0e6642-c249-4dca-8626-087a0eddbcd1 · outbound

This paper cites SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.861100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.861100Z digest=sha256:7272fa47c683e59537de0e3d277afe283fc72035cfb1dcd84d5a5380d8a3b89b

Observation 1fe18d07-90c8-4940-9f68-66e75506af6a · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.397729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.865081Z digest=sha256:4b015a7f1be254e0ad84265905c277c6dce9d20dd203ca224c5103bb2be65861

Observation 8b7c15e2-fd0a-4fac-99ba-3837589499df · outbound

This paper cites Should You Mask 15% in Masked Language Modeling?.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Should You Mask 15% in Masked Language Modeling?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.868682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.868682Z digest=sha256:5ad713541cfbe907c09e11f588428a7a580acaab750db69a77cfa388172b3e14

Observation ef32f889-edb0-484f-86f7-91a1467a4adb · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.872391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.872391Z digest=sha256:ca404ea37e4b3710041c72fba21af9b42b580e81838eed3aad48a6779ae35190

Observation 51149963-7a1f-4e81-a36f-2f4b57600d24 · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:21:35.383410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T14:21:34.876152Z digest=sha256:279367603a544cc7d3aafb66f733430b54c47a9e85017afb4576c1ea2b4ad339

Observation d71aea1f-e280-4820-8845-7debb73c14fd · outbound

This paper cites an unresolved cited work.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.879921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.879921Z digest=sha256:41242ed5b981e869e7b6cd7a17b8d2f5bb1c98ca085726045758bad50151d3fe

Observation 84837abe-8bbc-4d0b-9e66-bfaa44e232f9 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy Instruction-Following Evaluation for Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.883733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.883733Z digest=sha256:4a13dbfac6e630e8bffbec70526d756ce3fb676706fcd15cb1559e4d4522cc0c

Pith citing papers

Observation ba3da3bd-99a7-4c67-87fa-f21e8ebb26f3 · inbound

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling cites this paper.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.219539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.219539Z digest=sha256:306c57cf57df05a8a3d255cc1ab7792bcab9ca4fd161c6ad8246718751338e39

Observation a3f681dd-5629-44b2-99c5-05cccb89b915 · inbound

GoVector: An I/O-Efficient Caching Strategy for High-Dimensional Vector Nearest Neighbor Search cites this paper.

GoVector: An I/O-Efficient Caching Strategy for High-Dimensional Vector Nearest Neighbor Search Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:45:25.681289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:45:22.575645Z digest=sha256:998dabce981cd8af1ff870a8b036836403b2382fdf011c95799107e416f5263c