Pith. sign in

Paper Citation Record · LEDGER

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

As of 13 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2412.12661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12661 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:53:58.181401Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T02:02:47.472868Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T18:25:57.968361Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6058d684-f0ea-484d-a743-bea7530e941e · outbound

This paper cites Multimodal biomedical ai.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Multimodal biomedical ai

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.344984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.811003Z digest=sha256:4afa0f540ce79971e1515a92cae87919506a23c7b4b6ff8b2f6441d6548a004c

Observation dda12b4e-e200-479f-a8bd-70824dd13808 · outbound

This paper cites The medical segmentation decathlon.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants The medical segmentation decathlon

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.328646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.816646Z digest=sha256:d4e6c4bc5d47cb1390813c27d25a6689691074ab8156ec66458d3520741cf35f

Observation 734bb337-b16a-41df-af5e-931bbce7cb0b · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.821659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.821659Z digest=sha256:a4fb0225e18f28567257b17fb208a7a77fb4d435272b569de03a3b1583247bd2

Observation e24032de-c581-49e6-9494-d9aad4e787be · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.826873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.826873Z digest=sha256:eb8d58934f23e7c9be37239914e2c09dc13f9c08f31d5e2fa591e2930a8c9250

Observation 04824c7b-d8d3-4a1f-82e7-ce0e3e33855b · outbound

This paper cites One transformer fits all distributions in multi-modal diffusion at scale.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants One transformer fits all distributions in multi-modal diffusion at scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.831765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.831765Z digest=sha256:d1f85698bc7a74494a9d0741c8547d7939b25013af55ab1c71ab4117e52ebc05

Observation 0b752c21-a271-4f74-8509-5d6cc7dcfbc5 · outbound

This paper cites The revolution of multimodal large language models: A survey.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants The revolution of multimodal large language models: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.299563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.836632Z digest=sha256:0057af389187e9ae96d18faf79c69f69287e9eec75c98a3d6b5b3185616c805d

Observation c7cb75f6-39f6-4346-a4fb-aa7420e799cb · outbound

This paper cites RoentGen: Vision-Language Foundation Model for Chest X-ray Generation.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants RoentGen: Vision-Language Foundation Model for Chest X-ray Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.841992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.841992Z digest=sha256:c295eeefa7826d6b5402bb23531335f6b77bdce4ff901df234813ce22013ac85

Observation bc5ff028-7759-4ddd-9df0-5619b4b338ee · outbound

This paper cites Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale, 2024.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.283057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.847096Z digest=sha256:020c96857cd29b180e21653b743ba7655d71d8b5cb90d19ef02b707a2fb820a2

Observation 49ed5c40-5b9c-4ee4-bb16-d928b426a4d8 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.851646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.851646Z digest=sha256:5ab479578268b83a6c8c03bd56b4cfbbb1e316fa4e65e853d1e590f4fd9103fb

Observation d2f676f6-45d5-49b0-9298-dafa18a46b6f · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.856815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.856815Z digest=sha256:c0ae7c2e1a60464b274b49f61633d11999b22a70abc632a9f79acb2f2b5e257e

Observation bd530d22-f985-4fae-b478-a14fa9736dbb · outbound

This paper cites Taming transformers for high-resolution image synthesis.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Taming transformers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.861393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.861393Z digest=sha256:f1a6ae0897debc9b81d18d65467b9217842e4f9633520d22666e0539f1bf85ed

Observation 12da7fe6-3fce-468a-927b-06d9a838e3d8 · outbound

This paper cites BiomedJourney: Counterfactual Biomedical Image Generation by Instruction-Learning from Multimodal Patient Journeys.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants BiomedJourney: Counterfactual Biomedical Image Generation by Instruction-Learning from Multimodal Patient Journeys

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.866401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.866401Z digest=sha256:30336a4d1b921442662c2bf4dff11f54dc15ec0f03e553b1b01d5781475d3b4f

Observation 95faf6f8-04ff-4e7c-a9fe-9ee0b496adfd · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.871900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.871900Z digest=sha256:cd058659bd8b1a200e50753ddd62e9ccc54d6fdf0aaa234d17c8feaf614ebdb9

Observation 28eee004-545e-4b93-a78c-f21121fc7365 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants LoRA: Low-Rank Adaptation of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.876860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.876860Z digest=sha256:81bcd13e22e68760c9fe1fcdbe2348a59a35bbac41ad62cacf0710bff486b60e

Observation b1aed334-147f-4a7d-85f6-ee70716a7acf · outbound

This paper cites Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.247554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.881870Z digest=sha256:4965eea7679cdd60ebbd4ef3d9a1ec8d1f4066c7bc49cf70b07a67a7d64f2ecb

Observation b4b30962-1829-4ed6-af3a-488c8b0e6a8c · outbound

This paper cites GPT-4o System Card.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.886422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.886422Z digest=sha256:d54e96fc9880f8c0a1448a037f4dd03823813cf845a030f8a71079e80343a449

Observation 5715a62c-7d54-485c-bbfc-8cb71cb215e8 · outbound

This paper cites Quilt-1m: One million image-text pairs for histopathology.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Quilt-1m: One million image-text pairs for histopathology

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.232133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.891429Z digest=sha256:2b44b003ccea1b0eb1f685b976d8ee5baf64147fcce5ec4eeb6e89fa4fda02c9

Observation 6b35c260-7e26-4841-88dd-c577c12e78c1 · outbound

This paper cites MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.896349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.896349Z digest=sha256:c4a8008c3153dbaa84c70cf7e7afde206ce585e383ad882d8111073ca76f5ab1

Observation b06b9773-f592-4bcc-82d1-e08173e0747d · outbound

This paper cites Peir digital library: Online resources and authoring system.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Peir digital library: Online resources and authoring system

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.901397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.901397Z digest=sha256:a478821d03e03aa2d3106327466f96ad9b1e06c854123523b495de600cef8877

Observation bf6e40fc-8f84-4dfa-a4ec-86ee26184bc7 · outbound

This paper cites Chaos challenge-combined (ct-mr) healthy abdominal organ segmentation.Medical Image Analysis, 69:101950, 2021.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Chaos challenge-combined (ct-mr) healthy abdominal organ segmentation.Medical Image Analysis, 69:101950, 2021

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.906102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.906102Z digest=sha256:a3473a22260a30516a1f578edbfd0b20d2410beac00fe9fb83a1fd6b1d4428de

Observation e23dc62d-342f-4cf2-ab56-2811a4a6ff06 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.910805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.910805Z digest=sha256:7c8c946a0dd7013073966ecdacedada49a94bfa53ea4865858e582eca1e71426

Observation bdb7c5c9-80df-444a-9f38-dccb86343c02 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.915893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.915893Z digest=sha256:9552fdddeccfaba26ef27023483bf21090cf2680aef357dd3a0237f9fac82229

Observation a2a6c59c-bb92-43ca-9b78-f44698fef1d6 · outbound

This paper cites Mimic-it: Multi-modal in-context instruction tuning, 2023.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Mimic-it: Multi-modal in-context instruction tuning, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.184595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.920621Z digest=sha256:0a42ddc5f526b2650538f58ea0ec49a2158e965fc7ca810a4de40b07f0f3e7af

Observation 8ffeab88-3272-451d-b868-6bc06ecfa279 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.925146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.925146Z digest=sha256:3bf04331da878229817842faf00e53348e562129d7d996382c6de8dca14149cb

Observation bdfe7f26-1b13-498c-8243-b27199355a8c · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Llava-med: Training a large language-and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.168875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.929942Z digest=sha256:36f18b983ef896b61cc5cbac83d54d196730c08bd90e73004bf3d217dfc29195

Observation e4d264b6-7445-4dac-94b1-e195bb86cf23 · outbound

This paper cites Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.934922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.934922Z digest=sha256:b77b13a9fcfb311085265c33b40639d6924ca4401630678a4f8f9c718ee1b89d

Observation b4c46895-953a-4f9f-9199-30f9e7f89ef1 · outbound

This paper cites Pmc- clip: Contrastive language-image pre-training using biomedical documents.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Pmc- clip: Contrastive language-image pre-training using biomedical documents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.939505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.939505Z digest=sha256:b8a5bb49bb7ace134626670cca7f82c2acc3ac0e308225e6ab011ef68b5ba4de

Observation b0bb1ac7-b82a-4270-8202-35c4dfc91f44 · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.944249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.944249Z digest=sha256:2418a517ed7fb61e7835bec14870d7baec13102c419f910b1ab16ece84b81f12

Observation e6c17c83-840a-4403-a13c-f136ab6213cc · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Improved Baselines with Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.949021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.949021Z digest=sha256:e5bb45df8613d7f91ffe486dff93326e6ecb925f43a07f557537ddee39c87f6e

Observation e6ce9422-9d3d-40c8-b6b9-1ec44531d420 · outbound

This paper cites Visual Instruction Tuning.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.953983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.953983Z digest=sha256:fd024342e3c0dd5f1c360c653ffcae80587ddb81ad8b8ac2f77da9477582b88f

Observation 237b7a46-057a-4fcd-8767-6c3bab50669e · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.958991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.958991Z digest=sha256:69009cba29c3465be8f7d17b1c562a4d0df55d24f775f7ac240d1a66ecacacab

Observation 48293247-e7f6-4c2a-bdf6-243e10c9548f · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.963829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.963829Z digest=sha256:7144e0a13b8b3e75654b6909a93ea551ef0624cd65c705a8949898b3f0386916

Observation 8d5a2295-2ebc-4540-bd3b-b393dab8b637 · outbound

This paper cites MedPix — medpix.nlm.nih.gov.https://medpix.nlm.nih.gov/home.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants MedPix — medpix.nlm.nih.gov.https://medpix.nlm.nih.gov/home

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.121536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.968892Z digest=sha256:1c43e6b7d81f66b3366b52b77bb486c1da7117f81036d7052416c09e322470b6

Observation 63a11895-075c-453e-9192-1e8b76f5874c · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.974935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.974935Z digest=sha256:9ddfc0dadac7812680e49b20ad96f0117f091441cfea287f5bf986a8552d3335

Observation d7fb46f2-894a-4473-8722-07e80b8f4564 · outbound

This paper cites Med-flamingo: A multimodal medical few-shot learner.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Med-flamingo: A multimodal medical few-shot learner

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.104499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.980416Z digest=sha256:39d56bfa59f35fa5a855b70a4153975ce05109755abdbc8515cb9ba581e15282

Observation 88cb461d-956e-4327-b4a5-a5f8b4644f02 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.985507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.985507Z digest=sha256:3d7ba2e7224a6765d05be5a51fc80b61884036e86acf91bcd6812f966a3baf49

Observation 197436da-9730-462f-8733-6bc2af197728 · outbound

This paper cites Gpt-4v(ision) system card.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Gpt-4v(ision) system card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.990367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.990367Z digest=sha256:76cebaf6de69cd52c432bd4e37bd19015a65fc881d12a499884ad316c30822f4

Observation 78110334-7be0-4d74-9b96-15779d4f4180 · outbound

This paper cites Gpt-4v(ision) system card, 2023b.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Gpt-4v(ision) system card, 2023b

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.075954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.994796Z digest=sha256:89e17fab44178ffa6ae33cbd7f959e4d21e6872a414150f691a70c673b8c7452

Observation ecd247a9-d96a-4be9-9a14-c58ba735e00e · outbound

This paper cites Gpt-4o-mini.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Gpt-4o-mini

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.060152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:57.999672Z digest=sha256:911b4810581c3e2de4797a5b2556ad9f953f955eabc02ebc8cd2a5178256d473

Observation 78609a61-5550-4919-ac88-1ded081e3369 · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Im2text: Describing images using 1 million captioned photographs

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.043964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.005250Z digest=sha256:a5cb13c60d7f0736d51e1d07ae3f4876a74c7fcf0614586a7030ccf0f3b4e445

Observation a5f3a0b2-ee4f-42d1-8762-60280c19b165 · outbound

This paper cites A survey on biomedical image captioning.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants A survey on biomedical image captioning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.026613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.009947Z digest=sha256:837f8b5ce01a6634b343de647054844f50bc48d63b301e819232c389b5e370cc

Observation c3dbf298-8342-41c1-a9f0-604e32a3397d · outbound

This paper cites PubMed Central (PMC) — pmc.ncbi.nlm.nih.gov.https://pmc.ncbi.nlm.nih.gov/.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants PubMed Central (PMC) — pmc.ncbi.nlm.nih.gov.https://pmc.ncbi.nlm.nih.gov/

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:59.009357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.015029Z digest=sha256:41f4072a8912a7ad5327d21ae141e4a09731567a3ca017c49d08d2ec68d4b48e

Observation c6de0f55-f62e-4cae-b771-3ee55984ec59 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.020129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.020129Z digest=sha256:eb04de52c09538a0d47683d2d87bc3e9b9749920bb02bb1c0ceb60b66d6564f9

Observation 6b4376af-af18-4eb7-bc98-1297c6d43b17 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Neural Machine Translation of Rare Words with Subword Units

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.025158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.025158Z digest=sha256:98f5b25368ea5525b6de51fcc87db5eedeb827d69b0718b9786a57067a1eda1d

Observation b3ca51b9-df4d-400a-9044-3450ac51e48b · outbound

This paper cites Quilt-llava: Visual instruction tuning by extracting localized narratives from open-source histopathology videos.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Quilt-llava: Visual instruction tuning by extracting localized narratives from open-source histopathology videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.982137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.030110Z digest=sha256:09c9d6d4b22f2fa7fe47fa177e01e426075077294d7f96d2c2f02ed55cff0b37

Observation a668ff7a-6b70-400f-94a4-a3bb3597ec5a · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.966580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.034456Z digest=sha256:2e3b1a08b3a29bae7b1ff8f868610c86ccef4b3b2133d2187fa9fe9856c8882a

Observation 22c076c6-40cf-460c-b90c-479d6d5590a3 · outbound

This paper cites Pathmmu: A massive multimodal expert-level benchmark for understanding and reasoning in pathology.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Pathmmu: A massive multimodal expert-level benchmark for understanding and reasoning in pathology

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.950400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.039112Z digest=sha256:b020e43ed6afef11213a5522ef1868a9d9ed2e91d1820b1f3709efd270a73450

Observation 4c479ebf-513b-43f0-a843-98343df8db0c · outbound

This paper cites Hashimoto.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Hashimoto

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.934232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.043622Z digest=sha256:4d3931aafa6f44e955e6970bc877b2116bdd7c5b8e22437d44330ed9153ad4da

Observation 07c4df25-de93-48a3-b4ee-f33cebd1455b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Gemini: A Family of Highly Capable Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.048405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.048405Z digest=sha256:ed3d839c858f26e3d3ae82875cd0846069632730fa1c9e71f95a8868fd366d22

Observation f7c8b971-7f9f-4b0a-9a9b-f089b015ed80 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.053376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.053376Z digest=sha256:4d4c4406e051f604518073a9c88a34eef678b8e4cbe03e29303cbaee6982fca4

Observation 775bd66a-2076-4ac7-8156-a684690867da · outbound

This paper cites Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.058420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.058420Z digest=sha256:dc261db0a8bc97daab1568c20f98e33c60b0c6e3b6fc1a33efb45a11c7f2108a

Observation ef3322c1-9a3c-4132-afa5-473ec643ea12 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Emu3: Next-Token Prediction is All You Need

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.063098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.063098Z digest=sha256:87446ad73c619c766922980e2d3f92233d0c7a7fc567d03576b66f6e99de97c2

Observation c202b4f9-5d50-4413-927a-37be881c1d36 · outbound

This paper cites A Perspective for Adapting Generalist AI to Specialized Medical AI Applications and Their Challenges.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants A Perspective for Adapting Generalist AI to Specialized Medical AI Applications and Their Challenges

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.067999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.067999Z digest=sha256:76e35a1d9127d178a65b819d4a2d377fa2100847f7e5a695f70b0bf3ede2282c

Observation 1107c9e5-fd9b-4016-a06f-c0355fc96a78 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Finetuned Language Models Are Zero-Shot Learners

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.074084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.074084Z digest=sha256:50d018613ab12e331fd5db37b05cdc75e67b9d9b861fdccfa7426e507577049d

Observation 681c4060-bf17-48b7-9df0-fc4d588a4138 · outbound

This paper cites Towards generalist foundation model for radiology, 2023.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Towards generalist foundation model for radiology, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.907672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.078939Z digest=sha256:4359b15827c40e573e93ea7e6ee2cc25838f6c7131ad5f7e791e10864e90a719

Observation 994c7737-43ac-4690-a6b3-e02ddc2b8846 · outbound

This paper cites Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine, 2024.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.891833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.083391Z digest=sha256:d6ba7a9a6c20bf4e8f9e7ae653c613dd8de641f9f31a96bdc9ffb9b0213a17c9

Observation f4bf1ef5-07d4-440a-994a-b5ad5dfbae0d · outbound

This paper cites Medicalgpt: Training medical gpt model.https://github.com/shibing624/MedicalGPT, 2023.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Medicalgpt: Training medical gpt model.https://github.com/shibing624/MedicalGPT, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.875717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.088090Z digest=sha256:427a13e4a43937ce42bac0b4bac9fda8e9c038ad5c45ed1281a4113076014bf0

Observation 695006e4-49d0-4762-8b6b-7a500ca6a7d1 · outbound

This paper cites Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.859655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.092870Z digest=sha256:ba0e42435a7bf0cdf4543fc3641a670ed9f7c15cf5dd1dd218a9f3f60dcabcce

Observation a27c5622-bf51-4e69-b2c8-f7eec232c878 · outbound

This paper cites Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.097540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.097540Z digest=sha256:3f1a189991c75eb9eb8fb61148f5ef186be4bd3ea762ec9e35dd4fffc5a51caf

Observation 1d2303c8-d432-45f9-a61a-e65d6be2d735 · outbound

This paper cites Qwen2 Technical Report.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Qwen2 Technical Report

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.102285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.102285Z digest=sha256:7054e6066900479cbaacb45b8d485e8d473d6da1ede1d81dc3b6708444f2907e

Observation cc79bc80-7495-4fde-9a89-0b4e35b2e3f9 · outbound

This paper cites LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.107338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.107338Z digest=sha256:4d7fff2cffb78a7d8426567b7396201c5f8710995026ce65c4d265c51a4900af

Observation 13b7897b-cbf7-4b01-a412-a7c583b5ae98 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.112361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.112361Z digest=sha256:48e209be7046c7c00279a00ddc75bcc94098e7f1b71263e83a6727a288fb79fe

Observation 00a7de6b-7378-485f-b38d-338fe0c9ab11 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.117240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.117240Z digest=sha256:d51393f5ac7791391b8fead7adbb3bb556ddce679165822ee75fc35ceef75c5a

Observation 2225d483-0dab-4e05-8174-14743ee6c6ec · outbound

This paper cites HuatuoGPT, towards Taming Language Model to Be a Doctor.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants HuatuoGPT, towards Taming Language Model to Be a Doctor

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.122044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.122044Z digest=sha256:ab463c7053d3b64107b0ecce9ce130b35c8e945b08365b617baa8d833ac33822

Observation deec105e-bf65-4d90-983e-e260243826a3 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.126999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.126999Z digest=sha256:94d76b015651eca71e4b23c54606a9146cc07b7f50c1e3a71a11dd5dba0631b9

Observation 3951dd08-1005-4136-9ad6-95b6a4dab81c · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.132247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.132247Z digest=sha256:aca4d4bfff7c0cd78067448a36b29e49ff73225e318ef8670a3dc355d150b044

Observation 2222b20d-9139-44d3-a3c2-0ebc94a796b4 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.137134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.137134Z digest=sha256:e7be1e875c2fce4857156bc1d32805dccb97af453c18cf684212a5025e6c5647

Observation 2d157ae1-cfae-4674-8a90-233e2254ebaf · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.142077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.142077Z digest=sha256:4ae8ccbbacfb166d4cabf61e4fe860636358db50f0180ed7ae6ad7a3cf2edfdd

Observation 82e2c50d-4223-4307-9e18-741075495f1f · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.146912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.146912Z digest=sha256:6079880300e2407f66e7a4a8fd6178e59ca2b749dcb51f0d5f7f00706bf4e3f3

Observation f934a191-e118-4527-92a5-fb6c6c17a7d0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.151781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.151781Z digest=sha256:d832f37081954e39553129a290424bf68c8c02515c1dbdb708dbb93d322d1768

Observation 475a6fdc-cdbf-4033-ad04-4db1f5739b26 · outbound

This paper cites an unresolved cited work.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:53:58.842807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.156619Z digest=sha256:36db3bf59be4c4c92a1802bde65c84d4242fc8a71176f774153120593adf5f89

Observation d2433d6e-a787-4e4e-b78e-009be68dbf46 · outbound

This paper cites an unresolved cited work.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:53:58.825626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.161868Z digest=sha256:d22466bfdd6f4ffb4641cc8f35bf3a5b1bfa4019bb70fe4f09415e0d02f745d6

Observation 36e5f0eb-afe3-4432-b56c-24cb28f5754d · outbound

This paper cites an unresolved cited work.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:53:58.810290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.167494Z digest=sha256:cdb8f99db2cc9f3185c5bcdd2bd936ef80a90bfff8b406eaee4472716f0c21f7

Observation 367cb785-a457-4473-8808-e00c67195bdd · outbound

This paper cites an unresolved cited work.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:53:58.795074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.172262Z digest=sha256:23ee672b001f8e529b127807350dc7c35e882540785692b58e3b94d64d9ae1df

Observation ea08829a-8123-462f-baeb-83f7cd7a248d · outbound

This paper cites an unresolved cited work.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:53:58.779851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.176870Z digest=sha256:d9938c4f8b9576a3449a2c0c62d741fafd7e7f19564035b5a0d782a8fa065e3a

Observation 45a91932-ff9d-4afa-844c-e37784607b93 · outbound

This paper cites The answer is: Yes.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants The answer is: Yes

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:53:58.764849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T13:53:58.181401Z digest=sha256:3e298f030fe216b84700b28fca697a46842e2ba12978d3521510df8c8b2308c2

Pith citing papers

Observation 8b73e521-8eb3-4145-b507-629f1c4b1e2c · inbound

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI cites this paper.

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.029610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T08:16:33.482990Z digest=sha256:c3655512724e1f1cdaddb2e2e8fc33cd1f42e5c6685b19dcecacd67d339a684a

Observation 66265407-5544-4847-86fb-d06de8e4832c · inbound

Aloe-Vision: Robust Vision-Language Models for Healthcare cites this paper.

Aloe-Vision: Robust Vision-Language Models for Healthcare MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:57.970152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T02:02:47.472868Z digest=sha256:2cc0de7d8a95ec6811c54cd599624741fbdc2f566925e3ba367fbabcfca0851d