Pith. sign in

Paper Citation Record · LEDGER

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference

As of 16 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2412.12785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12785 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:49:12.726438Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30dafc38-5698-47ba-861e-a3d4b6a6d292 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.468040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.468040Z digest=sha256:865cd1bb045a2ee92a657ea2e0a67068779fcade95a3265e0de0062634f90059

Observation c9f4b99c-dd88-4d21-ba32-03eb8ae866ee · outbound

This paper cites Pixtral 12B.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.474770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.474770Z digest=sha256:e1d7182a7b7a7cbea67939229f2a6f8e569438f9a069606c2fbb168063d2e454

Observation f8d863e5-31e8-4fbb-ba08-e12de5c8d28f · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.480404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.480404Z digest=sha256:628c347f3f4464a2a26a42aeb58427f90f5d8b5372aafde0d99b2898a8287c0e

Observation 5b739907-2177-4be3-85e3-1ca207bce131 · outbound

This paper cites Instruction-tuning Aligns LLMs to the Human Brain.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Instruction-tuning Aligns LLMs to the Human Brain

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.485175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.485175Z digest=sha256:49023c4b691e4f03d306fea93de3c093052a9b1f4177b8b1a3d852c6df9d56b9

Observation 5d3c4080-6f69-44e8-b950-68cb61252533 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.490150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.490150Z digest=sha256:f92e30de4e0c30fc08660d1d90760411934138542b2fa1c0305f22fb502b8bcd

Observation 3897803f-e8a5-4599-a3d4-8c7c3fda1810 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Gonzalez, Ion Stoica, and Eric P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.494877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.494877Z digest=sha256:738aafad0f1e216dcc771f29d3e98237d3959e24996d7cd49f8f056d18413869

Observation e8f0a5e0-2ac1-4733-a094-a03959af333b · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Scaling Instruction-Finetuned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.500247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.500247Z digest=sha256:88110424845657cfc6827154f72540293f40e1feac701bda78bd31a2108143a3

Observation 984ae898-8d5f-4e89-8a49-d2fb1136202d · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference NVLM: Open Frontier-Class Multimodal LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.505286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.505286Z digest=sha256:d26bdfa1b4c60d6a714a6617538c3fdafd2310d343003ee196eb47ece7043503

Observation 5504010e-1aed-460d-a860-4501da2a60e9 · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.510207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.510207Z digest=sha256:0218589071c260b21e09a59600e56df11a5d71d3637bd96cd9adf67fb958d3f9

Observation ba6240a8-0d55-46c9-8906-9e06da886432 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.515155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.515155Z digest=sha256:6c714950298b53d97c18dae2be11a61eb4fa4ab3b60994c86411003df03b450a

Observation 965be546-6383-466b-8e05-b6fbbff931a4 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.436515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.520009Z digest=sha256:8069f0daf6a7c868c8ab46100c26e4545d2f7ae6c034f34bf28931ec2bea229c

Observation e19d8ab1-2a35-48d4-9e15-6635f91b5d5d · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.524604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.524604Z digest=sha256:7f884d97746d61e610047d8de7767878bef408e3ddf8e5105cbded4302ce6a7e

Observation 61d2e163-3061-4c61-9434-e4f00b91d069 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.529080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.529080Z digest=sha256:60bd82801b00da3dac4739d1472eeeec26949ffeb7b6a2a43704e6a9621d1aaa

Observation 8fabc6ad-0059-4c7b-89a8-e47b56ceaa11 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.411616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.533695Z digest=sha256:aa148c16c37065164d1b0b41f7935e66bde4c889da5a90cbb1cd4754dee5ffe9

Observation ef3f6ddd-44de-47ef-9c4f-1773a4919fa7 · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference The Unreasonable Ineffectiveness of the Deeper Layers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.537971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.537971Z digest=sha256:971f14ed238d074bc97c191ad9072df8a69ed5cc4ca7bf7023a48c4924ebf271

Observation d25ff837-5e1c-4d99-b73a-745f915868a2 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Efficient Multimodal Learning from Data-centric Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.542968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.542968Z digest=sha256:9431a4f5ca2486a208c8ab7443a19c96477b9f5f9425f371807fb001bacab9ab

Observation ebd90f6a-0d99-4f41-8c8c-96d4fe042ced · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Measuring Massive Multitask Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.547632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.547632Z digest=sha256:4916f44d67c5eef7bce28af8d87a4cb58853356ed9a05d21ae636724e584de0d

Observation 67449164-b1d3-45cf-aed3-2bd187592686 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.552271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.552271Z digest=sha256:32914c4d36a055cd1728d405bc66493aa18b7746dfcfd232b5f02d34cd427326

Observation 2049e8bb-92fb-42b5-8b4e-f3193ec7d4c4 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.557051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.557051Z digest=sha256:695da5f352bfd3c070094cc0ee9bb471bcc798615cdc96409268377f5829b5af

Observation 82f36fa2-8fdb-4bfc-827d-411b1228aa1c · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.561562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.561562Z digest=sha256:71c2a5461b5bdb5624c0c918156f4ece4b796f41e05dac2c421fa274c4440bc2

Observation cfebaabd-31a1-4045-b6e8-acec8a4806f4 · outbound

This paper cites Guiding Long-Short Term Memory for Image Caption Generation.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Guiding Long-Short Term Memory for Image Caption Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:49:12.962431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.565672Z digest=sha256:3f2b3ae13182571b564e5104fdaa45cd696e3352fa96ef966403c9c2b885b0bd

Observation af17f1f0-4773-4a54-834d-c63ecb0351fd · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.378101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.570199Z digest=sha256:dca041bff4141907aeb409bb7637a4bf9480b55a30af9102c7dbc1761adc672a

Observation 67d42a24-64d2-4170-af4b-4d4636a22795 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.361832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.574656Z digest=sha256:4f8a806fcbdcf6f68250b2141837b1a62fed2b6aa6fc8852cc84d62342073f6f

Observation e7b5cce5-2c07-4f38-8ce0-cac3cdad171b · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.579122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.579122Z digest=sha256:68001d50a73bbb8eab76ff65e7c4a300242b3b2d36f5d5cd87734969f21e6832

Observation 356f6130-0abb-432e-b73c-e71ebae76986 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.583560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.583560Z digest=sha256:e93f8917d7e243f69d9775a202f160b06e4985ea08c8432b24aeee630160ee3e

Observation bfcaf7f4-a43b-4c26-8dd0-6d8a5f9e365f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.587811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.587811Z digest=sha256:dbd9fd2bf09140b49e66abd3da0d81e19c6822381e0028901d3e0f03620e2aa6

Observation 6b670da1-4c78-466b-b2af-d42a8ef2651c · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference CMMLU: Measuring massive multitask language understanding in Chinese

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.593531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.593531Z digest=sha256:ceb6bc728fb0e22e4694dbadfa660c5e29fe858acebbc12ce8512c2e18e9f7e4

Observation 7b27f0dd-5590-4b0e-9bda-0715205136d5 · outbound

This paper cites Memory, Consciousness and Large Language Model.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Memory, Consciousness and Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.598115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.598115Z digest=sha256:8c72847760756359ae36332f4e0cd9c3377014085ad17550424fef499c413304

Observation 97b1f271-5934-4000-bcd6-0d7add38cd49 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.602749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.602749Z digest=sha256:db9eeba16f9a30734256463c9792af9c2022a95584e53a6099b18a2a4c54e502

Observation f0f8a40c-b2d1-4186-b27e-25292bc2c429 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.607399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.607399Z digest=sha256:6bd2bf228276a0bc39dbe8bd916a016eb379040e5efab082b8167c55449ccaa8

Observation 7c040444-6f3f-4392-9b0b-e5d93781f80c · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.319441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.612698Z digest=sha256:7cd749469b5cfde18e8065d8a30bb522596ae5d003eaf4fc48c19f9c2a2dc364

Observation 532e239a-3a5b-431a-917c-c4a7223c468b · outbound

This paper cites Visual Instruction Tuning.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.617281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.617281Z digest=sha256:19f48326081ce4151de44c8703a5bd78a2cbcbeaece33f1198e5e1028868f1c7

Observation a82acbf8-6328-4b48-a7ad-fc1af14e9c5f · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.621982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.621982Z digest=sha256:740cbda371dfe23e5e6263ddd0481faa2dee59ae155a306c02731cda19adf576

Observation c4c4e5fb-58fd-4258-a6be-16e5f25c3705 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference MMBench: Is Your Multi-modal Model an All-around Player?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.626587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.626587Z digest=sha256:d81ba1df7bc4b179fa9c5073453dd0311d0bb12d1e7d0589f71e2fa2421b33fb

Observation 4d18f721-55d3-47d0-a0e8-8def415af60a · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.631557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.631557Z digest=sha256:6e19ce01f71b1733daed8411e2ac207e8eaa275ec60671fe8be0e411904fb273

Observation 9007b795-a980-46c2-b9ae-3f55f4cb7e0b · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.635837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.635837Z digest=sha256:18e0fc73373f3cf26c546baf0f7e108792de3197720b346b86648ffd1400dc43

Observation 1f2d560a-ad07-426b-8240-0bd4fee3ec93 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference DocVQA: A Dataset for VQA on Document Images

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.640523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.640523Z digest=sha256:a1beda6c142000e57f6553df4bddd6dcbd4b6d87ad2dfe7b0828dc5587ee2477

Observation db714d2b-df18-42f4-96fa-9e02fd55f5a7 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.645023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.645023Z digest=sha256:2f4a7bf99ada8e982ef09c9876e66d2a10055d12828ce2a25ed8b094e9de5296

Observation 83a5a9d4-d96f-494c-887b-373656c12c85 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.649930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.649930Z digest=sha256:6fa71cbd8229499834c4fdc34ba4075d1847b49f39db4e35856254c5b8e7860e

Observation 135c30d9-87cf-42e6-9649-c5297ddb757c · outbound

This paper cites LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.654758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.654758Z digest=sha256:d57ed7e297e15fb7be46f812f0f581fcf0d903abb5bbdeb81e6bcdf6f87c5516

Observation 1c97ac9c-bbd0-48e9-bbad-b7b44291b47b · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.659877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.659877Z digest=sha256:5ace616bb70d7d47ddb5437e608165048a0d6b67ccbcb7c884c02f0d9f8d316d

Observation 12b906ff-fdca-436c-854d-853480582429 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.258483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.665311Z digest=sha256:68bc876031fe7d2d22599b607d2e09e4ae70fd7835591326635e2263c9f13203

Observation 866fce5a-9e21-4217-9e6f-4046ba7aa88f · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.669845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.669845Z digest=sha256:116d05c5e5bd8454121488ac555cf7683c95006d992f4d6c07421036363718fc

Observation bb619b33-7315-48c9-baeb-8f9a99d883b3 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.674400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.674400Z digest=sha256:0b656691aff79ce2e1e87a48775403b3e8fe249e2130db90efbcfe441160fa4e

Observation ae69f534-24a2-41b9-9bb1-d6520b4cc5cc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.679069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.679069Z digest=sha256:0a1ef8b54f3e59c273c04723bed97ae30a05aa8315d4e4c03fc2b20a73b13d44

Observation d588d747-f2ed-4412-9b5d-8d80e6892a0c · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.683752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.683752Z digest=sha256:63fa4fd5b8a04100c1dfdec2c6d391a72ceb2eab0387100c1f879eade587d96b

Observation 6719ebad-f486-4356-8271-7251361ade04 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.688706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.688706Z digest=sha256:929eb5b90d9ea00057c66ac4dc4ecdeb5ca93b5654d8accd4a01866addeccf60

Observation b01775e2-0715-47e7-87ab-63511a79cd87 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.216141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.693127Z digest=sha256:96e96a7c97c5abf9c4a41a0a6d919f636e247dc0f1140038e15de4c0596f7e70

Observation 5bbfc489-cd8f-4a6c-a86b-f871d31fd3e4 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.698600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.698600Z digest=sha256:706ae75b6ebc1b56b99b8ac4134869227d5b0f4e638ad6b9185445e21588c9f2

Observation 30400fe2-9e9f-4f8f-bab2-c7e20573f388 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.191555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.703328Z digest=sha256:27a6c4e5367fcf787191e69353aa3bdcc07ceb3640892e7c291f449670d32005

Observation f2978c7d-ae49-47e5-a77a-ecc46e1b1ae4 · outbound

This paper cites an unresolved cited work.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:49:13.176774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T13:49:12.707872Z digest=sha256:95b5a4971b98e9950691994826d61b22ff6917bf6deec766ea3c84cc361ae492

Observation 2a4a57c3-34d7-4413-b11c-50abaa7e0c9b · outbound

This paper cites Unveiling A Core Linguistic Region in Large Language Models.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Unveiling A Core Linguistic Region in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.712364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.712364Z digest=sha256:f44e6a0f0511cb8f5c9dfda2f6388bf6d7dcd7a5cd9864f7a5d6edc9cadb4ed5

Observation 01a3e858-e87f-478d-8810-5f63688c3809 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.717054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.717054Z digest=sha256:044de9e6e96320e4ab8ad6c5633a8b7e9f975e412f889cdab25e89fda6124363

Observation a4b98ea2-1189-4621-8ac5-237e9bc6d4be · outbound

This paper cites online" 'onlinestring :=.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference online" 'onlinestring :=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.721422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.721422Z digest=sha256:be6940c45e8a703babb155457fb649ec31f64febe1047c74ee9181fa5586cbb3

Observation 03af33a5-817a-400f-ac68-a68d2b4760f6 · outbound

This paper cites write newline.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.726438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.726438Z digest=sha256:b4d0f7f20aab3402229133d6274fd37aa917795998c01636f9d3905fcdc964e8

Pith citing papers

No inbound Pith citation observations are available.