Pith. sign in

Paper Citation Record · LEDGER

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models

As of 22 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2607.09029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09029 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:53:20.749426Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved77
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4963028-cea7-42fb-85eb-d58299c74469 · outbound

This paper cites Composer: A search framework for hybrid neural architecture design.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Composer: A search framework for hybrid neural architecture design

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:4cf5149fdd1d8c3880fe339f5319f6d461e80bd2dec9f14066b463c1ee3d19ad

Observation ee5afdd4-2fad-40fa-95bc-e43deb5fe6c5 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:180003f594fbee926a7deea1908b817b693b5935481c0d1c82e5f2d215bc3a25

Observation 1a893ae9-1b50-426c-9a26-f288e5ee2d6f · outbound

This paper cites Qwen3-VL Technical Report.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:1c90115506a72da8d0c7ec7a44435f85f02c926171cd44476ab6f529c4c88f74

Observation 74fe8b72-f0dd-47a9-8c48-8af1e4b0e80a · outbound

This paper cites Longformer: The Long-Document Transformer.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:779bfc2560363576be062485676319d143f41d695cbc5fee1c3487326c7ab721

Observation f46d6a71-f30f-4a3e-a6e9-16b47e02561d · outbound

This paper cites WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:4b50e6ce918c3f08677fc95c622bf3aad3da4ae406c297d1bd2c671f68a5d01b

Observation a5b1341b-2f75-4af1-bfa9-c06027d4ee2c · outbound

This paper cites Puzzle: Distillation-based nas for inference-optimized llms.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Puzzle: Distillation-based nas for inference-optimized llms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:05cd1f10c7089639271573d563e0ba5a6b4235f539f2bd9d6f9799d96adbb785

Observation d18cdb29-e46b-4416-8613-99d02ef4b336 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:b3b5addf095e1d336241760b5c2069feabbb5b227437d89e1ee23fad47b25b83

Observation 4c12516d-c271-43cd-9f96-16f16a517944 · outbound

This paper cites Once-for-all: Train one network and specialize it for efficient deployment.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Once-for-all: Train one network and specialize it for efficient deployment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:ceb2ab990a049896a4783544539fd8784ba8b7277523aef8a8641e4203f9a0de

Observation 4ad33f56-95b5-467d-b131-246ceb651600 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Generating Long Sequences with Sparse Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:5fca655289ef3cb2bfe2fc11fa9c16311cf88f5e9fdd1976bee2bddf838b53a4

Observation cad253a9-6995-4069-8724-96a5a4ddec1e · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:886ae72936b5cc369252544b9a6c582f48354576c1a2450e55d540b7a4f24976

Observation 79594ec4-3e9c-4446-ae9f-3f11afb499a2 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:a4c0267b879d7dd5091bb4fa859c63601208733db35721e8f703ce36073d2617

Observation 3dc7213c-40cc-4ed4-b780-133dcd4496a1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:e203ceb380ef8132fa905c62a5e9c17dd3a8d414c291aae98b07691404e3aea4

Observation d06523f8-ba05-4b70-a4e3-c4d2291c7075 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:3481487e077da98970373057ac41ace95b6a9e8f8bfdb288259e53446cf8b90b

Observation bebda3d0-1da3-4448-b0fb-e89f4d254f5e · outbound

This paper cites Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:86389bb2afaeb8696a3c6fbfbbbaf821a4eeff9cfa2bf6ff5e2beba649f58860

Observation 5303f941-f4bd-497f-af92-7209ce8da800 · outbound

This paper cites LayerNAS: Neural Architecture Search in Polynomial Complexity.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models LayerNAS: Neural Architecture Search in Polynomial Complexity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:3ed089482295bf62c285e560188e7ae58675f687a9dd0a81e90684d24fd59ba5

Observation 9aff1944-1833-48e9-9c65-07d58c43bb72 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:29bc48f3a5d8663d4de69114323a04820248024b49d9fe3fcd3cc33826d24bc0

Observation 31115c86-e5db-407a-80c0-fc3406cd993b · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in 10 video analysis.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in 10 video analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:4832866b7b156537b90d63bac568f928c439da63e872fe37c0867ffbb7f9a6ff

Observation 8a06598b-317e-4566-bc59-9d269a873df6 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Blink: Multimodal large language models can see but not perceive

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:511f7a0ba6789506b2e74873884a53955e14d7e9d5114a8e3008f0716f942a7d

Observation 30d48631-9e39-401c-a4ef-5c3cd703e333 · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Mamba: Linear-time sequence mod- eling with selective state spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:9de74b5230ed2271e3b59aedcef6725d94d45a5cf77ffdfdf5c8086d3b40f7cf

Observation 2b8ebf7d-1915-4347-96e2-fc8455158c3b · outbound

This paper cites Seed1.5-VL Technical Report.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Seed1.5-VL Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:c2d70576c60adea1ff7daead389bb940d9dabb41c7320bb6c2ec114b35ca1a0d

Observation bf7bc42f-9c7d-4a84-9d9b-8e2dc95790a2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:e4e349282263673a5356c999135c249f2efa741acc9af8d082d8dc10d0822c83

Observation cca5df99-1c8c-4737-b8d5-e8282a167574 · outbound

This paper cites Step 3.5 flash: Open frontier-level intelligence with 11b active parameters.arXiv preprint arXiv:2602.10604, 2026.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Step 3.5 flash: Open frontier-level intelligence with 11b active parameters.arXiv preprint arXiv:2602.10604, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:ce984809615f9a71003e879ee6d654b7150c5f192b5856e7377aa85e79586a3b

Observation acd11089-d067-4bd6-a987-1d0aef6a56f2 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:0ac044933a0ae653e8a5269ff4f6962a44045e4e237ef8a1e59ceeda8c435e4a

Observation e58ce5d1-ea09-4028-8cbc-55e2be4c80c1 · outbound

This paper cites Python-MIP: collection of Python tools for the modeling and solution of mixed-integer linear programs.https://github.com/coin- or/ python-mip, 2023.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Python-MIP: collection of Python tools for the modeling and solution of mixed-integer linear programs.https://github.com/coin- or/ python-mip, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:8323f499096e3d0c92df9b8739a16b577e41b7047014a76c51c6acd3f3691c84

Observation 58406958-4464-4b0b-b303-392159792d84 · outbound

This paper cites Mistral 7B.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Mistral 7B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:aee28a73dcbf6b789bdfaf11bd3e9575a2b3ab2d863373f26feb37e23e8dc5da

Observation 9d933880-011b-4e0f-8b9d-55efc7c15496 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:7b9207355f86bd91769b983344510ef23044dc3fa0afa04f1a6f3a234d37f4b7

Observation bf8d5aea-50d8-486a-82fe-684dbae0a8e8 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:d13c9bc392c209b73b9d8f3be9158bbabaf9e156cf6b8eef05fd2f8ed515a094

Observation c2e2859a-ff26-4a12-af8d-69420f0655d8 · outbound

This paper cites Jamba: Hybrid transformer-mamba language models.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Jamba: Hybrid transformer-mamba language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:95a5a548516bbc0fa8f273718701e4ce2b6b03cfb1db84c5129fe5c06dd4a3f8

Observation 3e14401c-64c0-47bb-9163-d80145054659 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:ff904f1ea163012905f3149cfe1d6a3aff70df1fb20454e89e767c4cc1239cff

Observation 50eb05b7-5b9b-474b-8925-ca0b2da177f0 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Seed-bench: Bench- marking multimodal large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:a1ef633203bb996f8e46628157aebc4a61283f54d50fec3c7fda2895c61eab7f

Observation 409a72a8-48f6-4794-a38c-1fe3dec517d3 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:20c43f0014a03d08eca35ec78f189c965372ad6aed06f418056f7874318102f8

Observation b436ce7e-b48f-4f5b-a85b-633f9c23e5df · outbound

This paper cites Matvlm: Hybrid mamba-transformer for efficient vision-language modeling.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Matvlm: Hybrid mamba-transformer for efficient vision-language modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:bccaf186376fc918cb4531d02bcce85ba3e820e40d07047ff8a3263076fbdecb

Observation 5ff20684-4e38-4d1d-b018-e8b39c8bb9a9 · outbound

This paper cites ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:24771bb7b06b28aa5bd8625e1d79ba2453c457771cccc09e4b81c16f87a30638

Observation 21d1ab2d-32c1-4969-abed-9392123245db · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Jamba: A Hybrid Transformer-Mamba Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:ecc709e333202dba294d3881fc898b52b14551c7ae335f9c8d01a9ca739a297e

Observation ff0ab2b0-d38b-426b-9294-f71cdfe9ccf7 · outbound

This paper cites Rea- sonable effectiveness of random weighting: A litmus test for multi-task learning.Transactions on Machine Learning Re- search, 2022.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Rea- sonable effectiveness of random weighting: A litmus test for multi-task learning.Transactions on Machine Learning Re- search, 2022

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:379bfb209654a3a0b95c320593e56340113c9032257e0aa4093cfe725ec26f16

Observation 3ad7d796-80f1-435c-b1fe-9c56bc6d637e · outbound

This paper cites Mmfinereason: Closing the multimodal reason- ing gap via open data-centric methods.arXiv preprint arXiv:2601.21821, 2026.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Mmfinereason: Closing the multimodal reason- ing gap via open data-centric methods.arXiv preprint arXiv:2601.21821, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:40c2aec1f133dc1bfbd4e6241e3894c7e416ada0aca8a9b899932325a034d3ff

Observation 457a882e-c583-4b32-97b6-675b3129c589 · outbound

This paper cites Smooth Tchebycheff Scalarization for Multi-Objective Optimization.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Smooth Tchebycheff Scalarization for Multi-Objective Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:cf952c716461edbb989990dfdc485221305e591decc29816f75411831d95cf86

Observation cb31c01b-472c-403c-9978-9bd5d2b81647 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:22b68dc2c845e47e778a6e05eeb8f1dd43b8e869e2edbd9415a20fd4fd20cdb3

Observation 2362dab7-8237-447b-af30-f14a5d2949a4 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:94a907032cb896f5f6f89c8394b49ce70ae0c2b14dc95f406cdeffedbf13de81

Observation ff7e6e10-12a2-495d-8b98-b596a87cc98d · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in neural information processing systems, 35:2507–2521,.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in neural information processing systems, 35:2507–2521,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:e372579e903e1bf056766cf3b13d091ff36c7fae75b4415032cfeb198b86c886

Observation 242c0021-34cb-45aa-a11d-ed585846ad6e · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models SmolVLM: Redefining small and efficient multimodal models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:03c654c2e33b00935fb358072dad40e34857432f58bd530cdec19f7bea063f97

Observation e43f924e-0b60-46b8-ad23-2ac55837fa02 · outbound

This paper cites Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:8fc9476369adb906ba32fb14855f493efcc2b1387c26de49b985afbfb7536050

Observation 8e4253cf-c643-4a08-bb07-cfb0eb05af2d · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:13cc1afb01b49e41788798e418a273ed49ee955c2619a72fa48c2a66d6909fac

Observation 9e565b46-33fa-4bde-a808-1bb90d99a7d4 · outbound

This paper cites Infographicvqa.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Infographicvqa

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:e8d3d14dff30de507345eb3cfe4047efc4798502e219d85f65ac5a38c0eeda2c

Observation 504477c8-4bd9-413c-8b6c-ae8d84d3af23 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Ocr-vqa: Visual question answering by reading text in images

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:8ede359e2321c812b306691dac4214b0dd9b749d907f5691d90946c708cd66f0

Observation 9cf6cbf3-2a65-4fc7-8be7-dfeb32d99b9e · outbound

This paper cites Olmo 3.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Olmo 3

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:cb0d2fa17f339c5ebdf5068930de1b645a2fa636f763f3d98f254e4c3b7cdbee

Observation 740d1109-001e-42c8-baa3-a6d3bd2acd2a · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761, 2023.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:5245d67b290304474683041abc9cc7c4adcb9ca11f36dc00fa057146b4a4f380

Observation a7cd3402-e1b8-4612-89bb-e87d7184617f · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Rwkv: Reinventing rnns for the transformer era

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:e7e1bffc50884f613e2f53faa61fb34eac2e9b4c5f12f38151425056b2a7f084

Observation a052a071-aa40-4493-af71-529ee45d1c45 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:7aade0b312129a5c11821f02793e375038364f58f0746fe7cda4230393ae7901

Observation 2de83db1-ff94-4a75-a2ac-32f6c971f8a2 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models BOND: Aligning LLMs with Best-of-N Distillation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:2f6d573a91f0fe54b2996b4b06952ef3ae3e2df272d1480731b4bd461fa36a1a

Observation 5fa75495-c68c-4416-997a-e156689547e2 · outbound

This paper cites Hardware co-design scaling laws via roofline modelling for on-device llms.arXiv preprint arXiv:2602.10377, 2026.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Hardware co-design scaling laws via roofline modelling for on-device llms.arXiv preprint arXiv:2602.10377, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:439b1c5dade3077dba7d8f42785ac33f1af78e1530cf96bd0de597dacd75396f

Observation 6572c381-b2a1-4112-8f00-b81c1b6c547b · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Retentive Network: A Successor to Transformer for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:384693f4db08a2e4c4e08878fa229f649177ae980f5ddd6401fd5c40d9ef40be

Observation 341ce738-5668-4f74-ad6b-d69550f80144 · outbound

This paper cites Infinitevl: Synergizing linear and sparse attention for highly-efficient, unlimited-input vision-language models.arXiv preprint arXiv:2512.06450, 2025.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Infinitevl: Synergizing linear and sparse attention for highly-efficient, unlimited-input vision-language models.arXiv preprint arXiv:2512.06450, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:c6d88bd5f1a95344bd7c836f0ea9896d2476af33190c819688ba154f817cce77

Observation 67b185ac-4f92-45ff-9857-6e3e85f7ef35 · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:87200289eac62dcf8db8142d9510f115acb06eeaf97b85d60f7364e575d80aaa

Observation a03e637a-439e-4d1d-a878-1b4137801d7f · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:5f67b63a86a8c37111ffb7b48f309629b0fa8ba3e08470371c67b9e3d9ac8533

Observation 3338cf91-89dc-4eae-89c4-acfe82ef1b18 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:647ce9e3c3682f75c3cd78cf3e3e14cce2fce390e2d8de53daa92977257fffe1

Observation f947e670-5a73-4fc9-90be-ec511304a901 · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models An Empirical Study of Mamba-based Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:63045a37a14517115ca9aed1af0d6b4e40ba06ef7a7588e0b4e588d17faf57ff

Observation c910fec7-2b33-4969-bf5a-f74f2509ec99 · outbound

This paper cites Hat: Hardware-aware transformers for efficient natural language processing.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Hat: Hardware-aware transformers for efficient natural language processing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:954ffa33c098526b65d68e834fa6605905d886cedc2942a163a35cb258e8631b

Observation 197f340a-74a4-46e4-9f13-7eb552dd365f · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:f0bd2c847e75caee7855a3b892d7ff25e9903f8223f6f5e7271574ccbc9ae43e

Observation 800dac3b-54e0-4329-a832-c2063eee09c1 · outbound

This paper cites Qwen-Image Technical Report.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Qwen-Image Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:776537997d0eaf8d03698c65ad7250559ee76062a8b4539a0166ac34021d79fc

Observation a7c4b729-64aa-47bf-892c-d40285fb2ab7 · outbound

This paper cites Rethinking kullback-leibler di- vergence in knowledge distillation for large language mod- els.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Rethinking kullback-leibler di- vergence in knowledge distillation for large language mod- els

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:fe51e25058033fec4b5a0a351b39bd44da9688940d20a4be65eaccbe8613bf91

Observation 185e10bb-8fa1-45f5-aee5-7ad9ca4eeac7 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Next-qa: Next phase of question-answering to explaining temporal actions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:e062c177eef5c0b06ef2618cf0677ee464f0607a8c72adbd61e899d424c7fbdc

Observation fa74d903-dade-4543-ada5-8c9803c9867f · outbound

This paper cites AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:74e40e68d2c2e083db149172de2bd0961554d0678ede8193908d3e388cbba0b2

Observation 90157096-88ef-4b97-997c-190e3ba79fbe · outbound

This paper cites MSWA: Refining Local Attention with Multi-ScaleWindow Attention.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models MSWA: Refining Local Attention with Multi-ScaleWindow Attention

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:b20e211995acf36c4ae2259f64a72e5ee4858ae8fd8acf358ab39d1c6757a625

Observation e66ccf70-e69a-4c98-afef-33cde19855ef · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:cc431771f11968a2433201bfe1cd331a31cd7af812a68ddd6918ae237e29356f

Observation 898056a2-a29f-46cd-86d0-d32bf6e58c0c · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Cambrian-S: Towards Spatial Supersensing in Video

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:c74a654bd1ee279d8a67c39604f93c814881f0233f6953c7bd188f5d2e088dae

Observation 4366df93-f0d2-42e9-b679-a4d3a808239a · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:92124aaa9d39a63fd1ca24da8d0a7d87d95199687ed6db5fd5a7f08f304dab8b

Observation b365cb0f-aaf4-4d41-8680-88c596f2eb44 · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:016a31fc6386381be36607ee5f522dca8c26af60b04a949123c2208e4af7c49c

Observation ce4a34a7-3a0c-434f-b97a-f5f862b54f52 · outbound

This paper cites AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:15c2168bb925f0f77c86b02c79ed7732d5741728de8df071f1cfe967e9f727a4

Observation e39523aa-4920-493d-ba6c-3ea21ec9ddbc · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:27173cc7f9de643a7f7f03eb3fb7c91bbac629c4a940d7714285b83a3732795f

Observation c4dab63b-44ff-4d3e-9df7-aa8604d8904b · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800,.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:266627bdc05c87a7409325daa421d0e1046bbdfdcfad5ac16bc117632df9b773

Observation b3a9372c-2042-4926-bd51-fb250e82bc2d · outbound

This paper cites Vision-language models for vision tasks: A survey.IEEE transactions on pattern analysis and machine intelligence, 46(8):5625–5644, 2024.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Vision-language models for vision tasks: A survey.IEEE transactions on pattern analysis and machine intelligence, 46(8):5625–5644, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:7bd890b186b147de1bcb38782a7dd65f6201640001ef3f62015dd53eb2ed1bac

Observation 3a0016f2-4689-447e-a60e-03f98088ac2e · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:5f218f31a9bc16fddcb36129ba51f18e48713cb2c7e9955194ca410e14e1aa18

Observation ae02f27b-2313-4587-b4f1-eab942a93651 · outbound

This paper cites A survey on multi-task learning.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models A survey on multi-task learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:c8025d2d6d3517cdfd8a88d22b996f095eac1a6d4f7f08ac38c09c4370ea99da

Observation 2a74f311-f049-4eb8-ac78-bf0cce34725d · outbound

This paper cites Cobra: Extending mamba to multi-modal large language model for efficient inference.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Cobra: Extending mamba to multi-modal large language model for efficient inference

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:85b7a518bf0355af3214eb956f5e2f3176f934df0632f20de035ff96f23b1567

Observation 5acbdc77-61f1-4b51-8ccd-0e45ea3b79af · outbound

This paper cites AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:cf14a493dc15dbb1a26b337db1b5d91bbb8c586c3edb85c5f76cf9357ecc5ac4

Observation 0e6c1eb6-3d0a-4e8f-8554-fd73f8a1d567 · outbound

This paper cites Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:75f6ebeaefdff9aff0ac8a5a73957ed064549a8737917c4a7675effd093dab80

Pith citing papers

No inbound Pith citation observations are available.