Pith. sign in

Paper Citation Record · LEDGER

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

As of 14 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 1 inbound Pith citation observation for arXiv:2508.06895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06895 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:35:51.006465Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T15:05:31.609683Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:37:35.644847Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c70655ab-56a5-4c64-bcc4-f669907cb552 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:49.341044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:49.341044Z digest=sha256:9148d6bb00453b2d8bc6e6d652d20bef83ff539cc706a81fe1691f6f01c1fe19

Observation 8547c9d4-2611-4401-8dae-29d19e8c4628 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:49.464613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:49.464613Z digest=sha256:c80431c34eee28b16d941debb8dd0ff7ced5e30eade38910402ca5326b5f0dc8

Observation 5f44d82c-5bda-47ac-a98c-b965b523c10e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:49.581640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:49.581640Z digest=sha256:4de34b402b62346de8baa05ff88c905684c57dd8fd629bd90a994b8248de963b

Observation a14b0dbd-b9c8-40c4-a722-e6160cca53aa · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models BEiT: BERT Pre-Training of Image Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:49.678589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:49.678589Z digest=sha256:64476f7d8302b35e280fc83098726a7696177f72b1fcda64c0d3ee35bf58c85b

Observation ea69dbda-7193-4816-a5a4-94047e0dd567 · outbound

This paper cites Introducing our multimodal models, 2023.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Introducing our multimodal models, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:49.769760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:49.769760Z digest=sha256:a9cee3744c0b589844ab0ff0d38cd05dff1cea747d7ef42bdf64fac2ecc632fa

Observation 286fbd26-d5d9-43fe-83b7-8e642984c910 · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:49.861020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:49.861020Z digest=sha256:51f064ee334320700f6dd50a314834e1c3c6414d27338cc04d83734a3d099e0b

Observation f00a6e90-2e19-4a0b-84b4-e36a274994d6 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Mechanistic Interpretability for AI Safety -- A Review

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.042218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.042218Z digest=sha256:b68f811782a29e5c36209e4baee519e81dd04becc92177dfc13bac71ec3bc013

Observation a42d0d58-e3d1-4327-8e32-c83934d265ab · outbound

This paper cites Towards monose- manticity: Decomposing language models with dictionary learning.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Towards monose- manticity: Decomposing language models with dictionary learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.299573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.170493Z digest=sha256:46d3479d0b21f55fec858a175998c1e6b6f3792c12635d0e867310859cf173ec

Observation 9ba8def6-8af5-4b2a-b405-02d9322b37a7 · outbound

This paper cites InternLM2 Technical Report.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models InternLM2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.328264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.328264Z digest=sha256:9d51c24114d7f5fc776115cac0d2831939f51d1ecaab11a6477198183e738c99

Observation 59c4c893-d02d-49c6-8c54-8133b029d34a · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Honeybee: Locality-enhanced projector for multimodal llm

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.283743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.507851Z digest=sha256:41c1cf90fbbd0960da879571ca89c9479687f3b3956aa887464ea7322e3ae525

Observation 4b911440-bcf0-4973-b37a-3b6aef5f396f · outbound

This paper cites Sssd: Self-supervised self distillation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Sssd: Self-supervised self distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.268603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.668050Z digest=sha256:04f44b651284759bc24c78edc75c2e812d0b614a665e0e2a07e782b55f182397

Observation ee4f9e61-2f6d-45bf-bad7-0e323b25dc70 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.687551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.687551Z digest=sha256:118e16501ae6569927087ef664402b8441eb376315023c98bc0f099ff35074b9

Observation 0cd7704f-30a8-40e3-acfb-6aec47ff4019 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.252412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.692423Z digest=sha256:47e5403ea613348e5a57f8acd889ddb9a99d19b8cd21d5952d5597a08e5f8685

Observation 8f132459-ad6d-407d-87dd-30b3199e6164 · outbound

This paper cites On the efficacy of knowledge distillation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models On the efficacy of knowledge distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.697046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.697046Z digest=sha256:3018aec11f0d5bad639c697e09e209c83b3c72b5a75203e3f2f97c8713d5371a

Observation 2d332bd6-792d-455c-ba7a-7026b82eb613 · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Towards automated circuit discovery for mechanistic interpretability

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.228061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.701230Z digest=sha256:b6dcc3c91e70b7aba68c95835d1c3a178a255e869938ea3781021a412331c73a

Observation 017ca2ca-ff9b-4ced-a391-915a93fb7aec · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.705188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.705188Z digest=sha256:6746355ad4768be44974018bfc9cf82c249997050f368cc8dfc4abb3f735c870

Observation 875f8441-66d2-404c-8e1c-74d7c8e58839 · outbound

This paper cites an unresolved cited work.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:35:52.211942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.711332Z digest=sha256:ca62dd3f4430bb395fe8c6894ab415f32a426a41ca3273aa3f3c497fb2cffcd3

Observation 2b905e90-76bd-496b-8670-40fdc7d8fa8c · outbound

This paper cites The Llama 3 Herd of Models.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.716976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.716976Z digest=sha256:308f152bff00f9de38302b8f1f775ccbf7f9490a7f195042f4e6dbee7ae8873c

Observation 2c7af29b-8097-4160-b3e0-f1de8e5e0acc · outbound

This paper cites Softmax linear units.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Softmax linear units

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.196697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.721385Z digest=sha256:94efaf55501e5092602699fc18b84f0c58aa45bc6298830d025f558105adb490

Observation 0035c43f-fdc4-476a-b219-aceb791b2b21 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.726451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.726451Z digest=sha256:0c5a79a1c7574a2738af47d851a5a13d142135cf993ae04450cfce119e45a2f0

Observation 8bf826e1-f8c1-471b-bedf-cc915a5c5e34 · outbound

This paper cites Knowledge distillation: A survey.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Knowledge distillation: A survey

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.182360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.730931Z digest=sha256:ff0326a82f2c2482d272c47dbfea3a5abb999879cf2731bd205357df80b79d25

Observation 54b1e966-2a24-4446-b3ca-c33e26b2df1d · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.735344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.735344Z digest=sha256:a581795cc0ea8e9d7ec3322fbf83040d467518d759ad0942e1b9e9ccb9326561

Observation 0fbdd0df-74fe-4181-9627-58bccb00fc42 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.739470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.739470Z digest=sha256:f3dd8ef360cd9eea5aabbab04f25951b4f169eb05279411386a78b9e67fbc36a

Observation f680d928-24de-4382-9a90-8958094b680e · outbound

This paper cites Learning lightweight lane detection cnns by self at- 9 tention distillation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Learning lightweight lane detection cnns by self at- 9 tention distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.145624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.743887Z digest=sha256:715333e04357bfe76fa870f7c1f1227df2057962961695f0fb54bc9c3ded60b5

Observation dd47919c-da0c-4988-9233-8b8ceb393723 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.130633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.748864Z digest=sha256:378d16c0b15ec9831e64161c004c5aa594173329e382b0467c0f8a78e80c7c10

Observation 1b1ba7eb-db2d-4bd7-a228-5d431c711bdc · outbound

This paper cites Perceiver: General perception with iterative attention.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Perceiver: General perception with iterative attention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.115075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.753080Z digest=sha256:9a53ea9505b8f177667612b2c34710334a577f0258bf137a3d8c51593a55ffe5

Observation 2a8b29e0-8723-4413-a650-f7583765e77a · outbound

This paper cites Mistral 7B.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Mistral 7B

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.757110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.757110Z digest=sha256:3c63803feedd423db48f4a8f45eb3e6d1dc8a36de11ad4a1b039457daec4ee21

Observation ded257a3-9fba-4d08-a5ed-3f3158ee176e · outbound

This paper cites Unified language-vision pretraining in LLM with dynamic discrete visual tokenization.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Unified language-vision pretraining in LLM with dynamic discrete visual tokenization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.100675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.761770Z digest=sha256:a872fd01b3794ca989d55886cff6eac253450a97912833907d9f06d807fa8edc

Observation a57f0a1f-f9b1-4ae3-9aac-685e8ee71b3b · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.766968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.766968Z digest=sha256:c31df4009761fe739c1503bd812963beccbe05e73248f863bb9a7766796aea78

Observation 4537c502-2c6d-4426-b379-d7e704f01182 · outbound

This paper cites Segment any- thing.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Segment any- thing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.773397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.773397Z digest=sha256:0b20c8289f4ffeae5b6acbcf42a11fe7d752e25d4129ee9442e0cf662489f978

Observation 498e08ac-b26a-4e81-ba0a-72b9ea507b74 · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Obelics: An open web-scale filtered dataset of interleaved image-text documents

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.067409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.777604Z digest=sha256:2469cdefee456f90d4c684a7f355ce96c040c02802c5b8efc01233d684c72ea9

Observation edc7b207-964e-4d94-a9a5-8430744b0757 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.052441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.781945Z digest=sha256:79541067086237d1c554537c186eaa1a50e9cda3e79b2bf8a0161986b4404cf9

Observation df31fd3d-5dac-45fa-b15f-679ae003b7a9 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.035819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.786192Z digest=sha256:965d78cd76e8c5c22c42de1fc99e95e13b8e4aaa6fc84e3ba3fe1bc0333b127f

Observation a9605a61-c8ae-41e0-9f4d-3d396002e224 · outbound

This paper cites Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.790592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.790592Z digest=sha256:f22aeef51378a97f14d65b2725d2ef34e489d187a77236983100a89de6bc2d0e

Observation dc813285-dfe7-4fd3-b6f9-afb74af56bdd · outbound

This paper cites Improved baselines with visual instruction tuning.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:52.018806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.795017Z digest=sha256:210917d9c854b21d068b103d249a6df448359d91dd3ddf041646d9c1f4c8de25

Observation a7ecdcfe-9f6c-4459-bcf5-0db304b632f6 · outbound

This paper cites Visual instruction tuning.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.799196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.799196Z digest=sha256:9faffc0be187cd9364b97b3c396bedebd28d9ef1a329d58b8add9a9121b444cb

Observation b83f99a5-dfb7-4b51-9a2e-7bcbb960c4fc · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.993626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.803584Z digest=sha256:f2742e3ad051d569bd4c3bc8b8fd96ac52b1b11193413eb753dc2eb88ee126ea

Observation 7ef00490-519f-4bce-84a9-5df723c612a7 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.808685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.808685Z digest=sha256:bc397dc12cdfe4d85c64d14a413619a511bc8931f987796c41395ac733e7fa45

Observation 59f2ed0c-4b47-41f5-a632-ef2eab90590a · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.965097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.813553Z digest=sha256:9718c328920b7ddf0afadd5cadc363fd6e74fbade2357ad90d8a692f2f8d2e12

Observation a7f0785d-addd-4aad-a03f-f9616356999b · outbound

This paper cites Self-distillation amplifies regularization in hilbert space.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Self-distillation amplifies regularization in hilbert space

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.948884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.818570Z digest=sha256:da25d9d91d2ae1759e690a225c85934c9cc3e29c5dcd9b23b0ac45ed65cb2e54

Observation 39563cd3-562a-4d3a-9847-9331d19cfe9c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.823423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.823423Z digest=sha256:de3e1cbe3e40ce1f386b5c3bad218a69f8a3c63963bad3aac3fa4997f0b6b2c2

Observation 93ae3b32-05d0-4e7e-83d2-dc066eb27ba9 · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.931754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.829892Z digest=sha256:319996271dccbedcb6cae35e757f5e9187203f377ef92faea4e029e2752c3138

Observation 0c5dd1c1-844c-4056-9f6c-214e315a9eda · outbound

This paper cites Bridging Vision and Language Spaces with Assignment Prediction.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Bridging Vision and Language Spaces with Assignment Prediction

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:35:51.403848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.839430Z digest=sha256:1f4285c981cdc77ed646b457eade0007d24a03dfffca3549147f5e66da70cb0f

Observation c66dc27b-bfa2-4ccf-bf9d-768a8cde5a62 · outbound

This paper cites Relational knowledge distillation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Relational knowledge distillation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.913886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.844340Z digest=sha256:a5f12268cad3596cf0d5aa14422d179e1e66f3fd7dec418bd68305fdf0c68b31

Observation d40f46a3-2fcd-43af-843b-4ba4b19de5ea · outbound

This paper cites Multi-modal Auto-regressive Modeling via Visual Words.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Multi-modal Auto-regressive Modeling via Visual Words

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:35:51.374658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.849417Z digest=sha256:dad5f1f99e40144f4176a35a8ed1e9dc83860d2d794a8088c6fbdeaa8e1f67c7

Observation 1432a2b3-15ca-4824-b862-44626d0e0937 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.855107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.855107Z digest=sha256:decb58a5b08ceb686a70ba8fa273708d05ed2612138b7c2ece20a7ee0531706a

Observation 8b928c81-94b3-42c5-a975-cd0bc3f77287 · outbound

This paper cites Distillation-based training for multi-exit architectures.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Distillation-based training for multi-exit architectures

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.898383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.859901Z digest=sha256:3615c5ce8708b0ddbdab84ec9d27fae7af2ed236a97928ebf1402a4bf94572b4

Observation a6bff80d-3143-4c4f-b530-ad550e46d76f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.880916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.865842Z digest=sha256:64c3779c51dc410bf79ff4461544be79f09a549beb55684e2829445f8744513b

Observation 68fa1b46-277f-42de-b208-af31d40615e4 · outbound

This paper cites https://sharegpt.com/, 2023.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models https://sharegpt.com/, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.860116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.875076Z digest=sha256:9069d5c3742a05cd0207c1b6e311826aa4b0743f57a46da8529ba2a3c5cfc412

Observation ef80358d-b4f4-414a-8d08-98c36f3996a2 · outbound

This paper cites Self-distillation from the last mini-batch for consis- tency regularization.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Self-distillation from the last mini-batch for consis- tency regularization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.839303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.880068Z digest=sha256:df40e32d4b52c98042e9877e3a831d16026b2d02ac006c1c463f7951dfab2cdd

Observation 78e534de-7b7c-455d-b174-b2fa44b71fb9 · outbound

This paper cites Towards vqa models that can read.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Towards vqa models that can read

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.822872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.885518Z digest=sha256:76321dff03531dbc9c07342cb4d8859b76189a2acb4a47c4e204a8fb7149d26f

Observation 28bec035-3033-42fb-9dc7-832a32c0e674 · outbound

This paper cites Emu: Generative pretraining in multimodality.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Emu: Generative pretraining in multimodality

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.807330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.891633Z digest=sha256:f57428f4ccc13babd70b1c044f0086d2ae8dececf47b063e8688f4e7e4aa7457

Observation 01e9b215-ee77-490e-acec-9ce79773feb5 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Generative multimodal mod- els are in-context learners

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.897216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.897216Z digest=sha256:7500f051053c772fe3dedfc75380f65b1e1b47c1f4aab22041650b932611ca52

Observation e50e2f12-9afb-48f5-b7c9-d1becc473c62 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.902847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.902847Z digest=sha256:391682a8a16d774918ccd008dbfbdfcb60b6e851b9559f3c4893643d89da2bf7

Observation 1af2ba3a-de28-4cb2-a1b8-c5f2b482cd38 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.908617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.908617Z digest=sha256:a52555056bf8cc18c208fe05143b7860a78620449d217e1744280930e936e3e2

Observation 881b7a09-e23e-4ea6-ba56-6da4a506c201 · outbound

This paper cites Neural discrete representation learning.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Neural discrete representation learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.913808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.913808Z digest=sha256:f1af7660280e8b76ca7193cfa157cd68394fe92d0403ee4602f1747bef73cfd9

Observation fe509967-c08c-4440-95e7-08e51b647e00 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.918910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.918910Z digest=sha256:9cd13d599d6a00f16675b25d245fef1157b5a5f7cf49209114dae51a303eb16a

Observation b464dccf-82c5-415a-8e3b-7ee6c73ca14d · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.925833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.925833Z digest=sha256:9eec2bc418f970a644215428ac8bc2b5cadcb0921ed65485da2b771523f5c6a4

Observation 0343fc5b-0a5b-4893-a932-bfd7ec1bca87 · outbound

This paper cites Do Llamas Work in English? On the Latent Language of Multilingual Transformers.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.930906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.930906Z digest=sha256:94595a6593f78da158ac9b230c6a12b7a347ed10d55f5b581194237dbf131ead

Observation 389002d3-9840-4f99-8139-c041a069e964 · outbound

This paper cites Libra: Building decoupled vision system on large lan- guage models.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Libra: Building decoupled vision system on large lan- guage models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.768776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.936418Z digest=sha256:3db4659753b5f26afb7358c6010f8bb9aed59721d2d81c133f1794fb63642f83

Observation 350f9945-f90c-4fb9-b5a8-fd72c96c38ec · outbound

This paper cites Qwen2 Technical Report.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Qwen2 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.941886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.941886Z digest=sha256:6b338e867f785c1091ba9937a878b5d914879bb3de6a45a0abefcfdac7f085bf

Observation d0d7727e-d383-4be2-8759-0418ec76d26a · outbound

This paper cites Snapshot distillation: Teacher-student optimization in one generation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Snapshot distillation: Teacher-student optimization in one generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.752840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.946569Z digest=sha256:93e0b2de6fa909614acf218a4a20e4b780add5b614e13154719d64a863c0c20e

Observation 1c631496-3b3d-46f0-b4de-2b6e88a6785a · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.952229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.952229Z digest=sha256:45cf21f6a798d72c29746d747454068818b1ebd857b6745c10be1695f202308e

Observation ece70674-f426-4caa-8b23-94f5688ab990 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.958123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.958123Z digest=sha256:422015c6ec5999a240652b1b229c101d4ba0436e28eb46efafcc92695c528405

Observation dd89176a-4563-4a09-954b-2777084513bc · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.965046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.965046Z digest=sha256:450f7c7f7b33dc0b586cbf72b88b512c164acd5ed10d1ac7ddf25f0485b8a397

Observation c98f7c46-dede-4b3b-8084-af4f288042b5 · outbound

This paper cites How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.970245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.970245Z digest=sha256:b2a101404ceafb8b855b09d70f551ec09425305fc8e680fa1464170a51b5a959

Observation 254817d3-8c29-450d-8cc0-6a1244a2b763 · outbound

This paper cites Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.975513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.975513Z digest=sha256:6ad915c88eac82d9b2490c29408195637d443c12e81d23d906702fad4a6ca0c7

Observation 4cfbed7d-5065-4f0b-a67d-82dba9e00ade · outbound

This paper cites Sigmoid loss for language image pre-training.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Sigmoid loss for language image pre-training

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.733114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.981398Z digest=sha256:cb262b099fedf1709af6552ca7d9c7219abecfdaae920f930df13ae1cc5e5560

Observation 107a0250-cf9a-4c02-9cad-a6a47b7d1a3f · outbound

This paper cites Be your own teacher: Improve the performance of convolutional neural networks via self distillation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Be your own teacher: Improve the performance of convolutional neural networks via self distillation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.717295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.988048Z digest=sha256:578c331c7bbac56ae1582d325b41a2872d2b196ee8994ed5dffa8f6e0768958d

Observation 5405491c-0ee7-4509-ac3e-1fe47aa53663 · outbound

This paper cites Self- distillation: Towards efficient and compact neural networks.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Self- distillation: Towards efficient and compact neural networks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.702294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:50.995136Z digest=sha256:94d5039a9264e0848afe59c2345a38e595fcfd0055b4f6884a222b3e5ec4138c

Observation 16ed01c1-cdd9-4097-b31d-5e42bfcd5df1 · outbound

This paper cites Self-distillation as instance- specific label smoothing.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Self-distillation as instance- specific label smoothing

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.686436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:51.001672Z digest=sha256:6e0d8807d38008f15246f42dbc6cc71ffc7b31c84897bb5a8285039c541049bc

Observation fee9d293-fe43-469a-8139-a64fb92017f2 · outbound

This paper cites Decoupled knowledge distillation.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Decoupled knowledge distillation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:35:51.669092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T22:35:51.006465Z digest=sha256:0f3eaaf8f08c9802db9d3871931e1b776de1228a997f96e63f2ed33b18277f8c

Pith citing papers

Observation 7f1da0d9-bb6c-42ec-8773-cebbd1aa81f4 · inbound

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification cites this paper.

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.646354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T15:05:31.609683Z digest=sha256:df85a8600d90db74605d850e57cb89625fbbddd1131ec05d29c1a7f902f6c309