Pith. sign in

Paper Citation Record · LEDGER

HiMix: Reducing Computational Complexity in Large Vision-Language Models

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2501.10318.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10318 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:18:49.356798Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:03:53.318565Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-09T18:03:53.492309Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4893557c-b3cf-4b9c-ba44-78d3b3d48105 · outbound

This paper cites GPT-4 Technical Report.

HiMix: Reducing Computational Complexity in Large Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.186743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.186743Z digest=sha256:de709fe5f7b9599798f8d3202c6281223bf616e102f09fccb9e17fb630a62468

Observation 3a7fa20c-a563-4222-b8b0-0bff5bec1197 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.191873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.191873Z digest=sha256:a4f8db831dfa20af57b7bc0783236bb1b1e2f398421bbc309a1ef7d2ef0dc695

Observation 05ba65c0-b72c-4a44-ae2b-4c89cbf7cb93 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.196788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.196788Z digest=sha256:bdc26bd06aa5f2753520a5f101ec340db2123ae8a953d912b012dc735bd23b1c

Observation 151e497f-0dc5-44b6-88e5-e4d8d25dda27 · outbound

This paper cites Qwen Technical Report.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.201979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.201979Z digest=sha256:737522c27e537573f6aa9e1017d7adb12a77ad5c6aed7c4c14b738e8f80b3d3c

Observation 2e44bf7a-98fb-44d8-85db-97d940774865 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.207169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.207169Z digest=sha256:ce0eeae9cecc797d1f85f49108ae80013a9c76c1d336214701e322252bbbebaf

Observation c625dd2c-4282-4b21-88f6-1bdc088c653c · outbound

This paper cites Introducing our multimodal models, 2023.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Introducing our multimodal models, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.916017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.212170Z digest=sha256:5bde42b3686013ce6b265d3c7789b36d81422a3e75b8b717a958e7bf45f1b771

Observation f8c1d574-848a-46c5-a34f-a75bcf6a1649 · outbound

This paper cites Matryoshka Multimodal Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Matryoshka Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.216883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.216883Z digest=sha256:a43281846ceb1c1124aa92bc44f632a1806dc0dc6c12dc0e1532475c06331a53

Observation c6168ff7-5972-4cd9-a44e-b15707bbdf69 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

HiMix: Reducing Computational Complexity in Large Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.221964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.221964Z digest=sha256:3542876e08434b0ccf98bbf0e47e8a87184beda933680c71399ce3361fcd4ac5

Observation b0ba5e46-5fff-4451-8d9c-21aac81f372e · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024.

HiMix: Reducing Computational Complexity in Large Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.904284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.227504Z digest=sha256:105223f98bf12873840c4faf0fee7b7d1db1b50755af4a7f45d3b2e8233a2bc9

Observation 1bd42b4f-011a-49ca-b2d3-365fceb83c3a · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.890233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.231606Z digest=sha256:b79b5fc3ebc3f782b5f492758be6d2ad2d57efdceced91963f332a71f1ea2353

Observation 84081bb7-d1bd-4704-91ac-1f31a5d53b5a · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.875849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.236375Z digest=sha256:d42eec1d8fb4060740779aeff621b7855be3f2b1ad51d4089abb54819b3ce23b

Observation 580a9ceb-4d1a-46af-92a1-af0e7d9536d1 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.241843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.241843Z digest=sha256:91025cee5ad93a9a1e8df5be57fb4f37316f03e02c20fff79367347cdd108b29

Observation ce90f7d8-daa3-4337-84c4-5846203e3238 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

HiMix: Reducing Computational Complexity in Large Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.247104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.247104Z digest=sha256:19dcda630a93954ab78c25420fae8dac7800b28fb66736793abb8c5523f12c3c

Observation 7c8ea428-d2c8-4e8e-81bd-b020fff53a28 · outbound

This paper cites The Llama 3 Herd of Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.251989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.251989Z digest=sha256:a14637dc00edba32a3c28c52a26412948726d05470b94d325b731cae94b5a4ec

Observation 4c52b67d-40ed-4e67-a665-13d683f6c0ca · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.257101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.257101Z digest=sha256:2cb91a1053e0bb893830029f6382296679ce65559383e48e68e069bb76f17f24

Observation ff4523bc-06f8-4b7f-b16f-39d5ad5a2678 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

HiMix: Reducing Computational Complexity in Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.261166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.261166Z digest=sha256:7e04a2db018622a557f5233af179e21dc8faef336d31309b5f3b11a2205509a7

Observation 5e552329-6299-4b6d-811b-5025bad163b5 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.854688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.264913Z digest=sha256:336dcbc7c9c3bd480cf7547a4229fd343a693abf65b4e831a4aff2c52882c988

Observation 58f2dfe0-49c7-4246-9b9e-b23a70f6a75a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.841639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.268880Z digest=sha256:687222a88ddbcd0008522cefeab526f1bf20f9416b98b669876f31abce8e2d57

Observation 7b034409-e9db-4e28-9444-bbfdec447f22 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

HiMix: Reducing Computational Complexity in Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.272790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.272790Z digest=sha256:3cd73976811522639837d2940dd77476a7906d74d65ecf8d95522ebb6bb5b926

Observation 4f38d5e1-2401-44d9-84b7-7033e945acb1 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.276437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.276437Z digest=sha256:f830a1130a545170c7e40267668c4ecc8feece6bd27a6bc3bafa1cb03455123b

Observation 3bb438ae-dc3d-4ee8-a6e4-7b56ba76931b · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.280566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.280566Z digest=sha256:bb5fbd8615c7bd8e95ad187ec264f7ba030acaf5f0dc57c4e0841b777730ede9

Observation 8bb9e5f6-930c-48f3-9795-fc1b81ad1dd9 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.284557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.284557Z digest=sha256:6e1f3ff0482efa6af5aaeae164c1abb753412c4c1765f5aa6cb88059111132b3

Observation 533e2edb-6634-435f-a09b-2c7f9bc28248 · outbound

This paper cites Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.288803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.288803Z digest=sha256:98475d0c9d43c3d101f81ef7c5e75a9f7a35b185acb621c1b0471c7fbedadfbb

Observation b3eb2d70-b7de-4446-bbe6-dacc3b43adf4 · outbound

This paper cites Improved baselines with visual instruction tuning.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Improved baselines with visual instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.814340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.292788Z digest=sha256:2f0f1d045e2e37261a0bc35804e933642e2d1a3b1295291496e207a8449c86a5

Observation 5a72f689-464c-4052-8ac2-c6e3fff0326f · outbound

This paper cites Visual instruction tuning.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.800208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.296683Z digest=sha256:9817d1c9d717257f831d40937da2827de806939c3b72e1ea4031519a54c9566c

Observation 72a9111b-b727-49c9-bfa7-510299d39dab · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.300541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.300541Z digest=sha256:c302c3e897b41517db3b4a5915b7a0b3b7c72812c6862ceec2afffd416b63bf8

Observation 8c33ae2a-4435-4a3e-94e8-f0d685e0597c · outbound

This paper cites Towards vqa models that can read.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.785866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.304951Z digest=sha256:8bac264338832571c8ca9b263f6706056f273fe8884ff046e513c0300548933f

Observation 5cb256b6-a32a-4743-9e5f-0f4d058ad732 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.308877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.308877Z digest=sha256:7dd0d383899fe2f64cc6d893f285bda9077497b34e57a0d7e34dabeb8fb3a592

Observation bf04df61-7787-48d4-98ed-2e439ed84300 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.313513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.313513Z digest=sha256:13fe3a5ff57e9a6e11d70891cca0e6babc5649b40b2d37db2737c1c3c815b7ab

Observation c93814c1-7c05-4f13-a946-d2dc4e07d377 · outbound

This paper cites Tinyclip: Clip dis- tillation via affinity mimicking and weight inheritance.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Tinyclip: Clip dis- tillation via affinity mimicking and weight inheritance

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.773159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.317781Z digest=sha256:9a8eaf3176ad7f52c0940baeb38b93788a89e07407444c035c276b5f5c49ec4c

Observation b06da072-b8d1-4d28-b129-5ebd99ff7f3c · outbound

This paper cites calflops: a flops and params calculate tool for neural networks in pytorch framework, 2023.

HiMix: Reducing Computational Complexity in Large Vision-Language Models calflops: a flops and params calculate tool for neural networks in pytorch framework, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.760761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.321883Z digest=sha256:285fe1fc0a56f3a90cecb939226422ff677f8861a6186eb7c2397ed0108b4784

Observation 117951dc-7cd1-49e9-9289-b655767bf5a6 · outbound

This paper cites Qwen2 Technical Report.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Qwen2 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.326439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.326439Z digest=sha256:f0b758686f4cf59f8926cfdfa87f7a0281851365d9dab9553af99083f7319658

Observation dfb0970e-b584-44be-8b15-b4885abeafbf · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

HiMix: Reducing Computational Complexity in Large Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.330995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.330995Z digest=sha256:4bfbafd3e707b2bd2cf996fc8399b0cbfe9c4f5f254d60c132b09c0aa65ec228

Observation dbbe59b8-c29e-4934-893c-205a33ddf6c7 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:18:49.748418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T19:18:49.335557Z digest=sha256:8a12c0fe00bd39e292501d1ab6f009cb348a2ac5bd19e2980e1aeb0885226d67

Observation 15ec2280-2c26-46c3-96ab-623a1531fe47 · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

HiMix: Reducing Computational Complexity in Large Vision-Language Models GLM-130B: An Open Bilingual Pre-trained Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.339691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.339691Z digest=sha256:bcf9cf76bbda81c325e7e0d9f9dab2d82fa38ca74c6b00626c4a0a7ab3e500de

Observation c05e6d4d-d2d8-44e7-b630-6a91045b1c73 · outbound

This paper cites Sigmoid loss for language image pre-training.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Sigmoid loss for language image pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.344164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.344164Z digest=sha256:adbc50b52d58bb139998a84f79ac1ee76a9df44a1537bfb204555789b2c3856f

Observation 0f430555-2d21-4757-b227-c0aedde912da · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

HiMix: Reducing Computational Complexity in Large Vision-Language Models TinyLlama: An Open-Source Small Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.348315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.348315Z digest=sha256:7ad70ab920373f147b904dc6c966da15df6594e1392d6d62fefde148cc7d0b95

Observation 8081ebb7-4195-4af3-8b03-f4a8d42845ee · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.352599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.352599Z digest=sha256:787e3360ab675fcce5e76db372d3932e7e416abdb947e5d0fa376537b8da9d97

Observation 703216a1-543f-4bc5-9901-d188a215f31b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

HiMix: Reducing Computational Complexity in Large Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.356798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.356798Z digest=sha256:43f85f3c65c9519a5a22e9345a1a26703d13d0ea9e14bab9f254697b380636d9

Pith citing papers

Observation 8a1837d5-e9a4-484d-a90b-0c32bec08abb · inbound

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework cites this paper.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework HiMix: Reducing Computational Complexity in Large Vision-Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.497001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T18:03:53.318565Z digest=sha256:148eb72bdd2f7204ab932e94cbe064b1ea5dbdac8eb74b152a609e9d7971639c

Observation e8af6555-174f-49e8-9fb3-4932eeb61e6d · inbound

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models cites this paper.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models HiMix: Reducing Computational Complexity in Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.960504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.960504Z digest=sha256:aaa9e9bdbbc99a5ff5d1d9b1ad1da0e2f53984450e464a3a7801dabf2bdffc1d