Pith. sign in

Paper Citation Record · LEDGER

Sample-efficient Integration of New Modalities into Large Language Models

As of 18 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2509.04606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04606 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T06:00:27.624272Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f8a0821-0365-4ce2-8ef7-be3ab526b42a · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Sample-efficient Integration of New Modalities into Large Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.347629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.347629Z digest=sha256:dc80b5911dc6fa497c25b62b9be26637295b6afb331bf47474b4ee1f73d467cf

Observation 35edd962-0a12-493b-81a6-a225b62e64e1 · outbound

This paper cites METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments.

Sample-efficient Integration of New Modalities into Large Language Models METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.351735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.351735Z digest=sha256:887f4fb16ca9c1fed3aeb2c64400e3596550f3e00046bc061f1ef276eb7f8bda

Observation 729df4e2-7fdb-41c5-89f6-4abbd5723dbd · outbound

This paper cites SciBERT: A Pretrained Language Model for Scientific Text.

Sample-efficient Integration of New Modalities into Large Language Models SciBERT: A Pretrained Language Model for Scientific Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.355363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.355363Z digest=sha256:77865c9f0e1e68c00ec9877c07eea6ce2006cf354c47e5513b6ed64ffb93ca76

Observation 52b3f11a-abd8-44a2-9194-f1f5d591dc3d · outbound

This paper cites NLTK: The Natural Language Toolkit.

Sample-efficient Integration of New Modalities into Large Language Models NLTK: The Natural Language Toolkit

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.703504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.358727Z digest=sha256:a55e5aa994495f2bfe434d946d6fd4c00b3642781fb1b37768850ac9d7b47b7f

Observation bbd15379-b951-4764-81f9-08beaadec2c3 · outbound

This paper cites Principled Weight Initialization for Hypernet- works.

Sample-efficient Integration of New Modalities into Large Language Models Principled Weight Initialization for Hypernet- works

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.685327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.365462Z digest=sha256:244c804f9ca9e30c31a713c6b72e903f143d96bd3a87a9df42a94c28eb69ec4a

Observation 9ce7754a-25e9-4a83-a14a-92d454518309 · outbound

This paper cites an unresolved cited work.

Sample-efficient Integration of New Modalities into Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T06:00:28.673581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.368816Z digest=sha256:25b8522fe0fa96e4a1b5d8f589c048fe65f21a26a1585f912e2560a97aba0051

Observation 19d6a5c3-2dcb-4106-9e8e-a3aa651fc2e0 · outbound

This paper cites Model Composition for Multimodal Large Language Models.

Sample-efficient Integration of New Modalities into Large Language Models Model Composition for Multimodal Large Language Models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.662659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.371871Z digest=sha256:40a29cfcea95640c6e21528ab7ac8110cbed95997b5d81c69c5d1a80060d6d34

Observation fe851190-7f20-4bbe-8c17-a8977c1193f0 · outbound

This paper cites VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning.

Sample-efficient Integration of New Modalities into Large Language Models VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image Captioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.375861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.375861Z digest=sha256:ccf46e12304954fef438780c6647446d45e897aab1a9beea2d8a8fcb4e910c06

Observation e24a920f-f6b6-4ac8-9638-295cd646cb97 · outbound

This paper cites ShareGPT4V dataset on Huggingface, 2024.

Sample-efficient Integration of New Modalities into Large Language Models ShareGPT4V dataset on Huggingface, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.651036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.379381Z digest=sha256:9b9228317d42533903490151da244773585cb07ae9f02f2e6e0593a68280f5ee

Observation 3eca805e-75f3-48af-84b9-e36cdc55ae61 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Sample-efficient Integration of New Modalities into Large Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.382458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.382458Z digest=sha256:a753e1eeaac4bacee696111cd2e5a6c9bc62cb384eae3c68f1138947426395d8

Observation 0aa4d782-5645-4153-abea-5c2f8d9af55e · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Sample-efficient Integration of New Modalities into Large Language Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.385958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.385958Z digest=sha256:02749f2805a2edddb41f4cfcb13ef976daf9bcef3cf9e6d317f970d0d3aee7df

Observation 14282f81-28c1-4c87-baac-c5f0769d31ae · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Sample-efficient Integration of New Modalities into Large Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.640423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.389467Z digest=sha256:7fbaa9c9f8d77a590d8e55a39118b625cd2dd88694df5f8b07a08c0104028522

Observation b7772c53-2d53-42a6-8b50-6866bcdceab5 · outbound

This paper cites Clotho: an Audio Captioning Dataset.

Sample-efficient Integration of New Modalities into Large Language Models Clotho: an Audio Captioning Dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.629325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.392519Z digest=sha256:18fa047511409708435e291256ea9f1a72daff7b464a52691f2cdc508a7d7996

Observation 64d0fa1a-de22-41d4-83a9-4f47178f035a · outbound

This paper cites The Llama 3 Herd of Models.

Sample-efficient Integration of New Modalities into Large Language Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.395415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.395415Z digest=sha256:565ef539be6d8ed773f6d771bc086392613c2936749378c1a054632fbfa9f241

Observation 1ba264ee-f0f7-426d-82cb-12b582238489 · outbound

This paper cites Translation between Molecules and Natural Language.

Sample-efficient Integration of New Modalities into Large Language Models Translation between Molecules and Natural Language

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.617697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.398553Z digest=sha256:a98cf25687df24ef7f4c4060cb045b14a66271eba2b3fed28a730f0680d0eb3d

Observation 1ebccddf-3e33-4d6f-b15e-0aee71e19356 · outbound

This paper cites Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries.

Sample-efficient Integration of New Modalities into Large Language Models Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.604214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.401720Z digest=sha256:ff687a7a726a01b7506bd108ce58ca0cd1ce7f178c0975ad6a51c69aea5dc959

Observation 6e368e0f-ea9c-4e09-8204-85e6726c689d · outbound

This paper cites CLAP: Learning Audio Concepts From Natural Language Supervision.

Sample-efficient Integration of New Modalities into Large Language Models CLAP: Learning Audio Concepts From Natural Language Supervision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.591261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.405152Z digest=sha256:ee08da99279d41fb6b9ea4ebddd0e95a74c01435de8737db6e8c783b12f7401e

Observation 1103c595-0c97-45a9-950a-6ee92563f76d · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Sample-efficient Integration of New Modalities into Large Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.408297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.408297Z digest=sha256:bcb7923e8efe3f21d99a365c1c438294f01320268cd34b1fae8c8326391c1595

Observation 27c96fe5-b1de-46d0-b399-c6852e091228 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

Sample-efficient Integration of New Modalities into Large Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.412176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.412176Z digest=sha256:2f7f7f1567198d5b9e9404141e179a04b4941eb89fed39f74717c079ab46f156

Observation 912c348b-6a3c-48db-96fc-aa5f64be68fc · outbound

This paper cites EMMA: Efficient Visual Alignment in Multi-Modal LLMs.

Sample-efficient Integration of New Modalities into Large Language Models EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.415722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.415722Z digest=sha256:c07d5dc4fc46ca6560493568c656d86013e09c910d15c4de79dc7db9defd905a

Observation 5eda6399-96e9-45f9-af9a-e946405d6d62 · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

Sample-efficient Integration of New Modalities into Large Language Models Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.575918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.418978Z digest=sha256:67a41d08b9df829fb1bac5dade7ef9f881f5c1bcdce5c064e05a646d4ee3d05b

Observation 79ad8104-c7ca-4c61-8ac2-8b86a7a5e8c3 · outbound

This paper cites LLaV A-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

Sample-efficient Integration of New Modalities into Large Language Models LLaV A-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.561245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.422124Z digest=sha256:82c4a31bf8081fa28d2eb443a1d75ce9f470304c17a4a07e182f9d972829f1e1

Observation a549f9c2-ebdf-4abc-a3ff-a396a4ce82ea · outbound

This paper cites Dai, and Quoc V.

Sample-efficient Integration of New Modalities into Large Language Models Dai, and Quoc V

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.425765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.425765Z digest=sha256:7b3cf1b588d99d24b42e8f533c124bb5c02ac3c1aa450a622230765441b75823

Observation 4d8b8bbb-665b-4024-9e3f-73964e03a978 · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Sample-efficient Integration of New Modalities into Large Language Models OneLLM: One Framework to Align All Modalities with Language

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.538597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.429211Z digest=sha256:e670d8da1d13769d8db36429ac37209b58e542ba3e48602102ac49851e800270

Observation 8fb48564-3338-471f-88f6-13dfb234a78f · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Sample-efficient Integration of New Modalities into Large Language Models ImageBind-LLM: Multi-modality Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.432652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.432652Z digest=sha256:a9d031d63aa700716432ad377bd069213fbf1929ad18f2d5876823fc67f61ac3

Observation 3756b43c-dc7a-4a96-ba0f-8f101e1ff9a1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Sample-efficient Integration of New Modalities into Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.526391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.436010Z digest=sha256:bc72ba6da73566a59f7d27c5ef80c125f961e631433dcd9198750c565839810f

Observation 993d4154-8fc3-44c6-896a-9daf23210c83 · outbound

This paper cites LLaSA: A Multimodal LLM for Human Activity Analysis Through Wearable and Smartphone Sensors, 2025.

Sample-efficient Integration of New Modalities into Large Language Models LLaSA: A Multimodal LLM for Human Activity Analysis Through Wearable and Smartphone Sensors, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.513624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.439379Z digest=sha256:bf5b692d2d5db756bf63091ed3ab94015fc2f4b03b8ebb313a14d9d32b614c98

Observation 3e8306e1-0848-46f6-8487-e35a0a2f6913 · outbound

This paper cites Perceiver: General perception with iterative attention.

Sample-efficient Integration of New Modalities into Large Language Models Perceiver: General perception with iterative attention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.500779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.442577Z digest=sha256:14782c60ac2d8d9ef0d26f404e94c5e899d2691f4bebc9183a30405387828c45

Observation 6a7c9cf0-32e4-4045-b016-693b24c42168 · outbound

This paper cites From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities.

Sample-efficient Integration of New Modalities into Large Language Models From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T06:00:27.825437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.446056Z digest=sha256:3e2b631441d11aca1456e713f5a127e9115630919d0346840e5bef89193d454a

Observation 6ccff54a-d44c-4845-b21d-471410198c47 · outbound

This paper cites BRA VE: Broadening the visual encoding of vision-language models.

Sample-efficient Integration of New Modalities into Large Language Models BRA VE: Broadening the visual encoding of vision-language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.488538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.449348Z digest=sha256:f876c236a62bd6171a387286286896e6904e9abd5170c2d20437b3985063637b

Observation 5edaaa22-d115-4203-880c-7d10b2aae347 · outbound

This paper cites AudioCaps: Generat- ing Captions for Audios in The Wild.

Sample-efficient Integration of New Modalities into Large Language Models AudioCaps: Generat- ing Captions for Audios in The Wild

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.475757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.452422Z digest=sha256:46232808632a626aff73466c75dcb641928b3abeed72a2a193ac83c2a2d785b1

Observation e5a895ed-7f85-40bc-bd2c-95ed9902fdfb · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick.

Sample-efficient Integration of New Modalities into Large Language Models Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.463360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.455363Z digest=sha256:fe5057d78d82bc161ca9f0474e45bf8bebd87318f1816113b0388f2ed4cd6055

Observation 70e4e688-3944-4b69-ba28-cdcfd69fbb46 · outbound

This paper cites Grounding Language Models to Images for Multimodal Inputs and Outputs.

Sample-efficient Integration of New Modalities into Large Language Models Grounding Language Models to Images for Multimodal Inputs and Outputs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.450784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.458606Z digest=sha256:92655d8a20106fa008d762a0cdfd00b10ba74755cf393ca694307f7ee4a6b6b7

Observation 410e7e59-5465-4f53-9c74-82f025fd582c · outbound

This paper cites Similarity of Neural Network Representations Revisited.

Sample-efficient Integration of New Modalities into Large Language Models Similarity of Neural Network Representations Revisited

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.438000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.461671Z digest=sha256:686ac35912cd7f4ff8a6229ad2152ba3b8d0459051af1faf0be394907b2141d0

Observation 47bbeb76-782f-479b-b373-16a39136b1c4 · outbound

This paper cites BLIP-2: Bootstrapping Language- Image Pre-training with Frozen Image Encoders and Large Language Models.

Sample-efficient Integration of New Modalities into Large Language Models BLIP-2: Bootstrapping Language- Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.424302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.464656Z digest=sha256:8c6878b0699a2aa7c0ded847d21ed2c68e6618efaed6587581a005a5517dccbc

Observation 3b6937e0-cc9f-42bf-941b-486f4c5c9a54 · outbound

This paper cites ROUGE: A Package for Automatic Evaluation of Summaries.

Sample-efficient Integration of New Modalities into Large Language Models ROUGE: A Package for Automatic Evaluation of Summaries

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.410018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.467551Z digest=sha256:bf3a6cfc01d33e6c70c6ed4e4a71cd94c6120f74c8a5ff81772b26014dc97ad9

Observation 1b2041eb-84c7-4528-a5e9-8425833a6672 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Sample-efficient Integration of New Modalities into Large Language Models Microsoft COCO: Common Objects in Context

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.390858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.470601Z digest=sha256:06381e4769c3e3e6e81b2135523778e6f22d1e83075630cbb3bd22f6bf21a995

Observation 1530c413-d2ef-4443-9c87-2b92df2bdeff · outbound

This paper cites RemoteCLIP: A Vision Language Foundation Model for Remote Sensing.

Sample-efficient Integration of New Modalities into Large Language Models RemoteCLIP: A Vision Language Foundation Model for Remote Sensing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.379327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.473606Z digest=sha256:25e33525a9e259946a205bdc9ce46f92506884b91091e3b94ca84f83479008ad

Observation f00a18b6-377b-4d8a-8914-157951fad862 · outbound

This paper cites Visual Instruction Tuning.

Sample-efficient Integration of New Modalities into Large Language Models Visual Instruction Tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.367870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.476988Z digest=sha256:614d758aca836656bed27067a1477b52971d0ede0e092f92201f30799263648e

Observation aba6bf0d-27e2-487e-a2ea-e18e0be18511 · outbound

This paper cites Towards Modality Generalization: A Benchmark and Prospective Analysis.

Sample-efficient Integration of New Modalities into Large Language Models Towards Modality Generalization: A Benchmark and Prospective Analysis

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-05T06:00:27.809540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.480243Z digest=sha256:3fad898042679371f788626acd37111aaf376b8467683f356a44b8456555e6d5

Observation 7145ebd3-d921-42d2-b6fc-0ae92354a1e1 · outbound

This paper cites MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter.

Sample-efficient Integration of New Modalities into Large Language Models MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.355361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.483716Z digest=sha256:742d68fe4eda29bc6d09495c3271930dd0939557142e7008cbea6aa8eedb4f86

Observation b068d0a7-f7fb-4fd8-b534-1c9debf1bf03 · outbound

This paper cites Decoupled Weight Decay Regularization.

Sample-efficient Integration of New Modalities into Large Language Models Decoupled Weight Decay Regularization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.341981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.486991Z digest=sha256:abc74b4bb2081584380d415a8c7334a6dfab8d5ca9380cbbb4e16d0f6f1bb6c9

Observation aeaf28d6-da97-4e4f-bdf3-fbeffc03cd5b · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action.

Sample-efficient Integration of New Modalities into Large Language Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.329505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.490226Z digest=sha256:4a9b995220610232fad8325e1b7d204aedd561e550e5f712a96d4009d7ed81d8

Observation 99f01e68-3340-4ab5-be81-98fc9dd23528 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Sample-efficient Integration of New Modalities into Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.493319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.493319Z digest=sha256:2eacc2adf29e061820284788a251d79bee0a42d283d34b8ca9b3b1e0816095b4

Observation 4372f23f-e849-4011-8f2b-518adbcbefae · outbound

This paper cites EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model.

Sample-efficient Integration of New Modalities into Large Language Models EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.496538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.496538Z digest=sha256:dfa52c7cac2ddd0c9037ce116b324892451ed26d295424e0f85ad2d8c387e88f

Observation 116e1676-346a-4046-8834-fae435df3ac7 · outbound

This paper cites Plumbley, Yuexian Zou, and Wenwu Wang.

Sample-efficient Integration of New Modalities into Large Language Models Plumbley, Yuexian Zou, and Wenwu Wang

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.315326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.499842Z digest=sha256:ba6d5255f1e69ba580e69864976ecc70742aeef42fd553dd9eab673f58552d71

Observation 2ad69d09-bf8d-44c9-9275-b7aa87aa94ff · outbound

This paper cites Llama 3.2: Model Cards and Prompt formats, 2024.

Sample-efficient Integration of New Modalities into Large Language Models Llama 3.2: Model Cards and Prompt formats, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.196530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.502939Z digest=sha256:f69e39c35c802bb8b433a655400b0b49630416977f1aa7481084356aac526d3d

Observation d1028c33-f00c-49c3-b1c5-5f327fd2f56d · outbound

This paper cites How to generate random matrices from the classical compact groups.

Sample-efficient Integration of New Modalities into Large Language Models How to generate random matrices from the classical compact groups

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.506008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.506008Z digest=sha256:e555350f0e846643a93287bdacc53e6ea2c1a0191b424b8e38203ed823c86f8f

Observation 5d63d557-c52c-4e6c-9495-328a74914cf5 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Sample-efficient Integration of New Modalities into Large Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.509894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.509894Z digest=sha256:84dfc9f2b8826f0f17994cfe2d8b54fa59261e0a14e458ffc00fb1336f64752f

Observation 0a628394-0a09-4331-b85c-a18ea26eb6e3 · outbound

This paper cites AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model.

Sample-efficient Integration of New Modalities into Large Language Models AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.176631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.513165Z digest=sha256:cd5019a9c04ea7dea8d214faa266f9d112559be554b6a9dd75d743c62a8c93db

Observation c0b7fb79-d23f-4684-9764-60dda416267b · outbound

This paper cites OpenVid dataset on Huggingface, 2025.

Sample-efficient Integration of New Modalities into Large Language Models OpenVid dataset on Huggingface, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.165935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.516484Z digest=sha256:a39074093a1991bb36c35cbbb055e8bb9336db4cdb8917c259bd4b1d29b711fa

Observation b91db0cc-236d-4a03-a3c4-de4658f9d540 · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Sample-efficient Integration of New Modalities into Large Language Models OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.519716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.519716Z digest=sha256:29791d1f00cb20b043c29a8951e83c5d7fdc39f2c15e8ea34aa86feed87c20b7

Observation e82ef18e-2ef2-4cb8-89ea-d3bb25a089cb · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Sample-efficient Integration of New Modalities into Large Language Models Bleu: a method for automatic evaluation of machine translation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.155546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.522906Z digest=sha256:2285750f4b42fdb45cc18ab01f4c19d3364ee8ff44550e42263f720c51c8ec0d

Observation 45355158-fd17-4652-aa87-465c17214da6 · outbound

This paper cites Modular Deep Learning.

Sample-efficient Integration of New Modalities into Large Language Models Modular Deep Learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.144811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.526137Z digest=sha256:e309ecbeada1e33d02ae88a673b6a1de994fc0383a83a777d862234e7d8162d6

Observation 88bb21eb-b2c7-48ae-b731-94cb9a456215 · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image.

Sample-efficient Integration of New Modalities into Large Language Models Deep semantic understanding of high resolution remote sensing image

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.134440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.529445Z digest=sha256:1ff7584a4cf1fc58d14e9f2c37994763452e6a49717441afd51507b2a0ebd440

Observation 2471ce79-5042-4631-9297-e580c3a5cf42 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Sample-efficient Integration of New Modalities into Large Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.122972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.532742Z digest=sha256:08fd7b70bc593a569dfe6ee0c0810209ebaf5edf8b874db089efc1da65113443

Observation 09cff91f-84e2-44ef-9153-db1ab708336d · outbound

This paper cites Infinite Feature Selection.

Sample-efficient Integration of New Modalities into Large Language Models Infinite Feature Selection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.110868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.536144Z digest=sha256:d475ad30448f122a576b9efcbb594fe978ec6560447188d2392d0df175505ce4

Observation df8e9ebc-f2d4-4190-abc8-4a6b7f8280e9 · outbound

This paper cites Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks.

Sample-efficient Integration of New Modalities into Large Language Models Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.100280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.539441Z digest=sha256:df967c9e1e1cd325ac8121e453de706246ca63be798e594f0f1d2291413cb8bf

Observation b65c02e6-114d-4691-9073-3af225c6bbe5 · outbound

This paper cites UnIV AL: Unified Model for Image, Video, Audio and Language Tasks.

Sample-efficient Integration of New Modalities into Large Language Models UnIV AL: Unified Model for Image, Video, Audio and Language Tasks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.089097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.542610Z digest=sha256:53349ba73b82f04ed4e6e95242ac9392f15123f63a9424385d9a6669c81a660e

Observation 065fe912-a0a2-40fb-bb8c-2b7844c3de9b · outbound

This paper cites an unresolved cited work.

Sample-efficient Integration of New Modalities into Large Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-05T06:00:28.078605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.545801Z digest=sha256:192b40e4454c819bab5c5afe659e49436ea6b10afa9ba0a15cb977f714dcf493

Observation 06cd99be-7f6e-4b47-a0a9-b8759a0335c2 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Sample-efficient Integration of New Modalities into Large Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.551011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.551011Z digest=sha256:2bb1ac0fa55ef5ef38b99bbf57341d8180ae1ece60801e32115fbcc48626988e

Observation 83ee673b-fed5-4190-9dbc-1a784b41c091 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Sample-efficient Integration of New Modalities into Large Language Models Lawrence Zitnick, and Devi Parikh

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.068294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.554566Z digest=sha256:3c8353f16d9a2a4ef9fb110552dd3bffc50ee3a5d8893531cde45e786f2dd581

Observation 2aa65663-85ad-4b37-9acd-f26d47b7cf75 · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J.

Sample-efficient Integration of New Modalities into Large Language Models Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.557681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.557681Z digest=sha256:b15bfe948fed2e615799e9dac63a47b1536581e42efa20ed9e49a80728e69e5f

Observation 714d31d7-7257-474c-b718-a6ff75543756 · outbound

This paper cites Lintott, Anna M.

Sample-efficient Integration of New Modalities into Large Language Models Lintott, Anna M

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.049939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.560972Z digest=sha256:3c8c1dec452e917160bf0590644366a923fc5e73b6003b0540398c8b83959463

Observation ba69f3cf-0738-4f6a-82a3-b0475ae8264c · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

Sample-efficient Integration of New Modalities into Large Language Models VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.039379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.564228Z digest=sha256:2e327f25a9bdb835094036d43c885d5058f79a33a7f290be053d96ba5373006f

Observation f8971ecb-6003-4539-ba2a-3fcfbe8ffbae · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Sample-efficient Integration of New Modalities into Large Language Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.028294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.570781Z digest=sha256:479129e84ae87a1c335e7c99ed6a9a0b345503c29ff3c6c1380acf3a0cbc1e96

Observation 38a5db5a-596c-4fbb-9e11-11d8df1baf15 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Sample-efficient Integration of New Modalities into Large Language Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.573725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.573725Z digest=sha256:e1473dd369f259d95d3305e8089568159ad0bc36cb8693133c81b9e72f2625eb

Observation 35ae5dbd-7ec6-4a06-917d-7b61e45e5cc5 · outbound

This paper cites Transformers: State-of-the-Art Natural Language Processing.

Sample-efficient Integration of New Modalities into Large Language Models Transformers: State-of-the-Art Natural Language Processing

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.017360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.576922Z digest=sha256:00439a7aedfff5bd5eabc65b2f52b20efc72a62eef776366715a0364a2c31df1

Observation d57081dc-078e-482f-869d-5bf3a69b0578 · outbound

This paper cites Limu-bert: Unleashing the potential of unlabeled data for imu sensing applications.

Sample-efficient Integration of New Modalities into Large Language Models Limu-bert: Unleashing the potential of unlabeled data for imu sensing applications

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:28.006073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.580025Z digest=sha256:1c5f7e09e6e3b49dce41f34473624f8238c453b52ef630876e3554f341d35ec0

Observation 396c02a5-18ac-4eff-9306-1495825b592a · outbound

This paper cites Qwen2.5-Omni Technical Report.

Sample-efficient Integration of New Modalities into Large Language Models Qwen2.5-Omni Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.582998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.582998Z digest=sha256:09338804a1604baf43c6ec419d5c02d297ee04e31932c574a51f3444746bf3de

Observation 3fdba7af-492d-4d0f-b0dd-cec169b87870 · outbound

This paper cites an unresolved cited work.

Sample-efficient Integration of New Modalities into Large Language Models Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-05T06:00:27.994488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.586159Z digest=sha256:51819cfc855e08d7d4aad542357e306dbe6b44431b5e45d08b1cf56e6854aebd

Observation d2816582-8cf1-40cd-a92b-c27dbebd0320 · outbound

This paper cites Qwen2.5 Technical Report.

Sample-efficient Integration of New Modalities into Large Language Models Qwen2.5 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.589509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.589509Z digest=sha256:b116edd586e14c476a02360193b4c978c0f21740893187bee8d06074ab19d318

Observation 2af74947-098a-483b-ab64-0990bff056af · outbound

This paper cites LLMs Can Evolve Continually on Modality for X-Modal Reasoning.

Sample-efficient Integration of New Modalities into Large Language Models LLMs Can Evolve Continually on Modality for X-Modal Reasoning

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-05T06:00:27.701532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.592640Z digest=sha256:f08e7d9b9625ccd9feb0fb0121565c6a32b5e3f856d734464fb87d8fefe1cce4

Observation c7bb3cb0-985f-44da-912d-6e0014e90558 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Sample-efficient Integration of New Modalities into Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:27.983294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.595818Z digest=sha256:4824836804b7edbe367b5eab06d67f7cc26fee871d782637f7fad8c322b63831

Observation e0cbbb4c-d38f-494b-a063-ed1dea6042b5 · outbound

This paper cites mGTE: Generalized Long-Context Text Repre- sentation and Reranking Models for Multilingual Text Retrieval.

Sample-efficient Integration of New Modalities into Large Language Models mGTE: Generalized Long-Context Text Repre- sentation and Reranking Models for Multilingual Text Retrieval

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:27.971424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.598926Z digest=sha256:02cb3c07a4aeb45574e91622bf64f5f6185cc66f09e9d63f7f131ef7e23febdd

Observation 53b8dd4b-32b5-4d04-9ea1-3fbce746d8fd · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

Sample-efficient Integration of New Modalities into Large Language Models BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.602134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.602134Z digest=sha256:adf63db7b46293f9010f50afc7fc4da961d2bc405ce086407d8320c78a585242

Observation 6a7a34cb-c281-41c8-abac-505375f9ec0c · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

Sample-efficient Integration of New Modalities into Large Language Models ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.605436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.605436Z digest=sha256:cf6290316e1119b1b4fa2ebad3a4b296735c3d94cf85ab41c85b3212d3af2dfd

Observation 9b8ceade-5c08-4bbb-9497-fdc74bf1db83 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Sample-efficient Integration of New Modalities into Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.608859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.608859Z digest=sha256:87da05122dfe03152aabf5acab62f5bc246989ac777c2fe58a3a1450781592df

Observation 254d9617-c361-47c5-b298-3e8e563902a8 · outbound

This paper cites Is the galaxy simply smooth and rounded, with no sign of a disk?.

Sample-efficient Integration of New Modalities into Large Language Models Is the galaxy simply smooth and rounded, with no sign of a disk?

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:27.960786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.612486Z digest=sha256:b90d84f3b368f58276c0c40f69560043e8256e2e1717bb553955064d6e6d6622

Observation 7291263e-a0ad-480d-af16-6666aa8f9f78 · outbound

This paper cites an unresolved cited work.

Sample-efficient Integration of New Modalities into Large Language Models Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-05T06:00:27.949757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.617130Z digest=sha256:f487e353187de95264cdb8d76f85ab2bdc0d2939dca76bc9ea680c604e177544

Observation e31ae708-c7e4-44cc-98ae-9ae3a20f21a7 · outbound

This paper cites The data is relatively stable, with slight variations, suggesting a stationary position.

Sample-efficient Integration of New Modalities into Large Language Models The data is relatively stable, with slight variations, suggesting a stationary position

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T06:00:27.939037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.620778Z digest=sha256:bf238ac31dda23768bd1f944c12debcb89100f0702908540e6d08e58e5415515

Observation 55f64af5-3588-491d-93aa-da81bf28f94b · outbound

This paper cites an unresolved cited work.

Sample-efficient Integration of New Modalities into Large Language Models Unresolved cited work

Reference 84

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T06:00:27.928041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T06:00:27.624272Z digest=sha256:21c2912ccb5ccee60a1e4d7b46898912ddd12b636617cc35307e9695f43c0e50

Observation 82894495-20ad-4c57-bf11-a92290711c48 · outbound

This paper cites an unresolved cited work.

Sample-efficient Integration of New Modalities into Large Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.567671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.567671Z digest=sha256:79984bba83d6abbb8d101c831e928338ba35740fb1583186c5414cc49ddd29a6

Pith citing papers

No inbound Pith citation observations are available.