Pith. sign in

Paper Citation Record · LEDGER

Diagnosing our datasets: How does my language model learn clinical information?

As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.15024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15024 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:09.919296Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f234cfb-0229-43f0-9ac9-6f1a9d51fcfa · outbound

This paper cites Zero-shot clinical acronym expansion via latent meaning cells.

Diagnosing our datasets: How does my language model learn clinical information? Zero-shot clinical acronym expansion via latent meaning cells

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:16.304764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:04.255274Z digest=sha256:acea0b92e692e10243186c7eb7e1627368952f079d9600d77d77d695a544155b

Observation 99a2bec7-6691-45bd-994a-27d5ecfbbace · outbound

This paper cites Large language models are few-shot clinical information extractors.

Diagnosing our datasets: How does my language model learn clinical information? Large language models are few-shot clinical information extractors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:04.341527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:04.341527Z digest=sha256:f3346ed36364efc6eb698ecc2e05a12bea3eda707b2469aff7fb6aa9ed8cd387

Observation 7c118357-626c-4b5f-940f-a4a17a0a8817 · outbound

This paper cites Medical large language models are vulnerable to data-poisoning attacks.

Diagnosing our datasets: How does my language model learn clinical information? Medical large language models are vulnerable to data-poisoning attacks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:16.043076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:04.475338Z digest=sha256:c0a050eb2d0938b6d7a6bad58503addb896e83a6048497256284a3bdde4ff582

Observation 458256da-4ca2-4769-98e0-578be6d31f5f · outbound

This paper cites Openbiollms: Advancing open-source large language models for healthcare and life sciences.

Diagnosing our datasets: How does my language model learn clinical information? Openbiollms: Advancing open-source large language models for healthcare and life sciences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:04.616679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:04.616679Z digest=sha256:984eb6941c252bd279de5ebbadc39735d89b857ea40b87e891d17abd2aab05da

Observation ff145259-2474-4035-8eaa-453546d21e06 · outbound

This paper cites Give me Some Hard Questions: Synthetic Data Generation for Clinical QA.

Diagnosing our datasets: How does my language model learn clinical information? Give me Some Hard Questions: Synthetic Data Generation for Clinical QA

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:11.051626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:04.745921Z digest=sha256:f6ab1a69cd93fce6696a49de67e962b8d318c1e4b6f2add600269c883dfdd294

Observation fef211be-f041-4639-9482-fe8f1c5f8b8d · outbound

This paper cites Testing and evaluation of health care applications of large language models: a systematic review.

Diagnosing our datasets: How does my language model learn clinical information? Testing and evaluation of health care applications of large language models: a systematic review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:04.875050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:04.875050Z digest=sha256:8f9b81526e3bc0053b75ee69cf799950d4b12c6ea5b911d33e1626a6a37f1bdf

Observation f556a6e8-17ee-4aab-8096-c8549633b3bb · outbound

This paper cites Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias.

Diagnosing our datasets: How does my language model learn clinical information? Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.026476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.026476Z digest=sha256:138cda467f21a564fafa50921c7a076b5c29e574d45f6530849cb6684e8e73a8

Observation 7a565fea-87e9-431b-9a14-cf70f44e078d · outbound

This paper cites MEDITRON-70B: Scaling Medical Pretraining for Large Language Models.

Diagnosing our datasets: How does my language model learn clinical information? MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.169117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.169117Z digest=sha256:5a46a85e460934ba1b58317f0ca3b291d9fda5c9fb0952c83842c9cb3aa7864d

Observation eda3edaf-b7f7-44a1-8d7d-34eec76cafa3 · outbound

This paper cites Med42-v2: A Suite of Clinical LLMs.

Diagnosing our datasets: How does my language model learn clinical information? Med42-v2: A Suite of Clinical LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.268413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.268413Z digest=sha256:449a50e8deed7aa46129954451a2eb9e46c0f3788aca7417efd5b82a3ecaee1f

Observation 55c9f203-4b69-4d08-b725-6fb1f15a246c · outbound

This paper cites Scaling instruction-finetuned language models.

Diagnosing our datasets: How does my language model learn clinical information? Scaling instruction-finetuned language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.379176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.379176Z digest=sha256:27d6b14d3d0e71c897d325f4ab4b06c5f739754398e2d6b77ae784e41e67fc7c

Observation 528baf99-34de-4f92-85f1-9a62ee7c0396 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

Diagnosing our datasets: How does my language model learn clinical information? Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.493446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.493446Z digest=sha256:8ae9963105300d2cc1d0bd7e5bc706ea877fa5800530a9954bb7f89891d1ea03

Observation a54525d7-39c3-4dd5-a023-af536ba65c14 · outbound

This paper cites The Llama 3 Herd of Models.

Diagnosing our datasets: How does my language model learn clinical information? The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.604207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.604207Z digest=sha256:615d9453fe45e2a2b75834dacaf039593f312349a8d5d034826a7e1d9da0d5bc

Observation 417c29a1-1790-48b0-841d-b33fb7f82d99 · outbound

This paper cites What's In My Big Data?.

Diagnosing our datasets: How does my language model learn clinical information? What's In My Big Data?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.747482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.747482Z digest=sha256:d251766d1a7c849fb25a81409d8944907de546e8f72ec0289fd6a1ca6e3059c8

Observation 06020078-e330-4d7f-b11c-206d8adcae3e · outbound

This paper cites Language models are surprisingly fragile to drug names in biomedical benchmarks.

Diagnosing our datasets: How does my language model learn clinical information? Language models are surprisingly fragile to drug names in biomedical benchmarks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.892549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.892549Z digest=sha256:d6fbd987d49caf62408817538918161825cdbfc38f3c8364b12c05cbe0cf2d2d

Observation b97d4c36-2a4c-47fc-a8b2-3c0d2fabae84 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Diagnosing our datasets: How does my language model learn clinical information? The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:05.982964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:05.982964Z digest=sha256:e1d1f913ec565608c5b71ec2f73e8f766143d792353108e9de13c19ae6b06b70

Observation c403848a-0623-4401-8f01-6895e11b1eb6 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Diagnosing our datasets: How does my language model learn clinical information? OLMo: Accelerating the Science of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.046507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.046507Z digest=sha256:a4c0e6ec8597347eb45a0b276a098b722c59347f8b787eba8531be1f25f2cc8e

Observation 1a74ded9-6b21-4a29-b9a2-cb912f8cbea5 · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

Diagnosing our datasets: How does my language model learn clinical information? Studying Large Language Model Generalization with Influence Functions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.156921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.156921Z digest=sha256:83833f4eed0c4319e34977759ea88df936689be94cd8ed41a32cf715bf1a4ead

Observation e4edb8e6-236e-41c2-933c-94c03b17d46b · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

Diagnosing our datasets: How does my language model learn clinical information? MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.242115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.242115Z digest=sha256:2245a80d95eeecfdf957fc7bc364149474874b890f4a10de24d585221602b9f5

Observation 30189382-66f0-4b84-ba5f-302319c7f16e · outbound

This paper cites MedNLI Is Not Immune: Natural Language Inference Artifacts in the Clinical Domain.

Diagnosing our datasets: How does my language model learn clinical information? MedNLI Is Not Immune: Natural Language Inference Artifacts in the Clinical Domain

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:10.660958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:06.348602Z digest=sha256:35fdd88202e4427dfe029b92b51cbe02d42c16d962f3467edb80f21d7a3c3bd1

Observation 8c4f91dd-35ba-4664-946b-e4ecebe47a3c · outbound

This paper cites Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?.

Diagnosing our datasets: How does my language model learn clinical information? Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.447607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.447607Z digest=sha256:f8432c641ecd4332f4a27c8efb1cfbe74d13b02b121e5cd7d4ccd126204765a9

Observation 57fefea6-580c-43d7-9e02-6859db2b4497 · outbound

This paper cites The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models.

Diagnosing our datasets: How does my language model learn clinical information? The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.507433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.507433Z digest=sha256:4c4e2239ab9944697ddf710f335e233541412b2746ad1b5fc365696907427b35

Observation fb725992-b5eb-4d1e-b1a1-fb34eb9a537c · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams.

Diagnosing our datasets: How does my language model learn clinical information? What disease does this patient have? a large-scale open domain question answering dataset from medical exams

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.662833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.662833Z digest=sha256:a7ca38ec47deb5896962e83651e31ef5ba6eba9c2100fc845e088ce4cfa8611d

Observation 64c726d0-4cdd-46f6-a77f-bb8f2e0ca4d3 · outbound

This paper cites Matching patients to clinical trials with large language models.

Diagnosing our datasets: How does my language model learn clinical information? Matching patients to clinical trials with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:15.803447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:06.752942Z digest=sha256:bf84f2b2ad1787c2ed645907389beb270c3d352cd8e8f71e89e9d7a8328e76d8

Observation 1ead9fe5-3817-411d-be2f-5b0aa66b4eef · outbound

This paper cites Mimic-iii.

Diagnosing our datasets: How does my language model learn clinical information? Mimic-iii

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:15.520427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:06.837072Z digest=sha256:c2ca2d9361cbc321330498e7f8b3030866797e660ed31e267d1e823316e0297a

Observation 15ffc71a-bc0d-4e2b-bee2-69f2ddb0d37c · outbound

This paper cites Mimic-iv.

Diagnosing our datasets: How does my language model learn clinical information? Mimic-iv

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:15.250248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:06.918535Z digest=sha256:db70c72064ef53243b40af1d663b83912d09deec69cf774422f9190bf08ea06e

Observation 708b237c-f1dd-484d-8bbb-10c6d554fe39 · outbound

This paper cites Mimic-iii, a freely accessible critical care database.

Diagnosing our datasets: How does my language model learn clinical information? Mimic-iii, a freely accessible critical care database

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:14.970322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:06.991294Z digest=sha256:8a35c543c3844805bb5188324b9b2c1935737356b8d8b9cd9dcc529b1d4a8c7a

Observation 447fe995-6908-4da1-98e9-c96bcdb23930 · outbound

This paper cites Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports.

Diagnosing our datasets: How does my language model learn clinical information? Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:07.060881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:07.060881Z digest=sha256:9c67e6dd968542da7ddd5a56c4d8d2aba0c4a6c0e6b9d9e6838edeebc962161a

Observation ae0ea4dc-eb55-4283-abd5-c5356397d8fb · outbound

This paper cites Mimic-iv, a freely accessible electronic health record dataset.

Diagnosing our datasets: How does my language model learn clinical information? Mimic-iv, a freely accessible electronic health record dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:07.131567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:07.131567Z digest=sha256:4cf5373f39dc10b4fc1f4326e8061801271445be3e8af11af56a6eba55e88a3f

Observation 15588577-92b1-486f-9f85-28b44b5d0892 · outbound

This paper cites Large language models struggle to learn long-tail knowledge.

Diagnosing our datasets: How does my language model learn clinical information? Large language models struggle to learn long-tail knowledge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:14.704615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:07.276988Z digest=sha256:a2501fc86ea350cd47b90f2c8647cdf1fb886f3c15e90ab744ff2fd131261761

Observation 148aae8c-38df-421b-9740-7b5f159d9553 · outbound

This paper cites Debunking health fake news with domain specific pre-trained model.

Diagnosing our datasets: How does my language model learn clinical information? Debunking health fake news with domain specific pre-trained model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:14.430165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:07.387646Z digest=sha256:1bb95ccefaf7cae7ec181acf07bfeae65e4a948ca95d8e6c28a70a63865c3a2f

Observation be032bd8-70cb-44a5-85ad-cff94c8f0106 · outbound

This paper cites Can Large Language Models abstract Medical Coded Language?.

Diagnosing our datasets: How does my language model learn clinical information? Can Large Language Models abstract Medical Coded Language?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:07.575990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:07.575990Z digest=sha256:f6663f1f494c69e00bd6e87cedf00d9abf9465cb716c8469a719b3ee002ccf53

Observation 653ef47e-371b-4dd2-9225-be2b3b0049d9 · outbound

This paper cites A scoping review of using Large Language Models (LLMs) to investigate Electronic Health Records (EHRs).

Diagnosing our datasets: How does my language model learn clinical information? A scoping review of using Large Language Models (LLMs) to investigate Electronic Health Records (EHRs)

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:07.705427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:07.705427Z digest=sha256:6a02356c9a4b2c98229b2e3cea76360971bd8bd20b0209c4537c012e139895e0

Observation 33b18d1d-f2a8-472d-8408-a8a9248a7f85 · outbound

This paper cites Are Clinical T5 Models Better for Clinical Text?.

Diagnosing our datasets: How does my language model learn clinical information? Are Clinical T5 Models Better for Clinical Text?

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:10.339675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:07.815853Z digest=sha256:3f7467217de82d5b18d212267cd4dfa51502e5c7ddb7a2f1b2143ff6230c3ca1

Observation f4c6fef2-e7be-4481-bff0-f80443ce7a4a · outbound

This paper cites Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens.

Diagnosing our datasets: How does my language model learn clinical information? Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:07.971245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:07.971245Z digest=sha256:bc898c737a5074062cb69974024983c4f3df0f0359a34d685300baeefbcd598a

Observation 61f82fdf-f356-469a-8257-6f9df144236c · outbound

This paper cites S2ORC: The Semantic Scholar Open Research Corpus.

Diagnosing our datasets: How does my language model learn clinical information? S2ORC: The Semantic Scholar Open Research Corpus

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:08.141224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:08.141224Z digest=sha256:de312a67e4801db63de17d6bba1290dcd5d7e4a5f17f3be6e660afdf3f5c3cfc

Observation ee845d4b-4873-41eb-a54d-0dac2eaf1635 · outbound

This paper cites Fake or real news about covid-19? pretrained transformer model to detect potential misleading news.

Diagnosing our datasets: How does my language model learn clinical information? Fake or real news about covid-19? pretrained transformer model to detect potential misleading news

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:14.211578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:08.270161Z digest=sha256:9de2720f01b7d27bfe9ca5c9e58143eeaee15a6b56c403cee1b0473bfd897e70

Observation 39aebfe4-fe93-4230-9a19-b669f314f29f · outbound

This paper cites Evaluating base and retrieval augmented llms with document or online support for evidence based neurology.

Diagnosing our datasets: How does my language model learn clinical information? Evaluating base and retrieval augmented llms with document or online support for evidence based neurology

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:13.930293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:08.439994Z digest=sha256:bee5a691ac8c979af01c0cda92f8fb50c0d30c203475262460a62340394e9c60

Observation be301167-7843-4da7-b25f-917dcacc47cc · outbound

This paper cites A sense inventory for clinical abbreviations and acronyms created using clinical notes and medical dictionary resources.

Diagnosing our datasets: How does my language model learn clinical information? A sense inventory for clinical abbreviations and acronyms created using clinical notes and medical dictionary resources

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:13.713212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:08.670802Z digest=sha256:263d5bdcde1a8b96a67e1b0e8e24aa63ae2cce3b0c5c3f6684ca65b7fb821333

Observation f7be1418-3c17-424e-a2e6-f961bb21df58 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Diagnosing our datasets: How does my language model learn clinical information? Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:08.762683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:08.762683Z digest=sha256:7626aaa62aa6c97e78bad307b6b204057716fa367a44b75423342fcc0ecb5964

Observation 894cd7eb-1a9f-44c2-8e7b-bc09b43356d2 · outbound

This paper cites It’s time to bench the medical exam benchmark, 2025.

Diagnosing our datasets: How does my language model learn clinical information? It’s time to bench the medical exam benchmark, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:13.382004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:08.811657Z digest=sha256:c6a079e83f173bb820d476b28d224de66908b620980a8a429d7a6f7eb8747a68

Observation a347fdee-7fcc-44c3-b722-f201398a60c4 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Diagnosing our datasets: How does my language model learn clinical information? Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:08.866790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:08.866790Z digest=sha256:d6d1edeacfe9e18ea8f16452450e0add8c184b19e8e434699860ab55caa7817f

Observation ea42f37f-58f0-4161-8ec9-877c62f6a4c8 · outbound

This paper cites Large language models are poor medical coders—benchmarking of medical code querying.

Diagnosing our datasets: How does my language model learn clinical information? Large language models are poor medical coders—benchmarking of medical code querying

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:13.086715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:08.908349Z digest=sha256:8c1079433c0c09df0b7ba9b00fa91635156ca097c294c75dc0b81d6020549f94

Observation 193c298b-b779-4c74-b865-d7a5a3faf2d5 · outbound

This paper cites Prevalence of health misinformation on social media: systematic review.

Diagnosing our datasets: How does my language model learn clinical information? Prevalence of health misinformation on social media: systematic review

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:12.805583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:08.958235Z digest=sha256:e70cb2e1f684f04c46da60d3b487a641d4bfaec708cee2e5ac1dd7270ea47a46

Observation af328636-6b24-405d-9a08-3d00738c805d · outbound

This paper cites Hashimoto.

Diagnosing our datasets: How does my language model learn clinical information? Hashimoto

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:09.010126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:09.010126Z digest=sha256:586f5b9670039d3c98686259a2fe20a668b4dba7cebdbad8675858d8023fa3b1

Observation 3a2facdc-cb21-47dd-a922-6f946dbfdd1b · outbound

This paper cites RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset , April 2023.

Diagnosing our datasets: How does my language model learn clinical information? RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset , April 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:12.585008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:09.060653Z digest=sha256:2a9d7878cf5ad4327e4ec5c0e516f03372302592e0c9da0706c059c3b74d9512

Observation 89ef4626-a47c-4e1d-9043-a7fe80bfdc97 · outbound

This paper cites Clinical camel: An open-source expert-level medical language model with dialogue-based knowledge encoding.

Diagnosing our datasets: How does my language model learn clinical information? Clinical camel: An open-source expert-level medical language model with dialogue-based knowledge encoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:12.321384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:09.113476Z digest=sha256:30b170d2477a01680f64b73839c0593b0e39e5d9765f96b1812bd62bcfad3cb3

Observation ba398e54-45e0-46dc-ac6d-267c68b11eaf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Diagnosing our datasets: How does my language model learn clinical information? LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:09.166078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:09.166078Z digest=sha256:e8ef554dcc0be4bbe13f8d4f351de2b1fc0385d8c184097b71f5a9572e43e78a

Observation 84648d4e-e3ab-4a3a-a53f-5a37f9bda00b · outbound

This paper cites Adapted large language models can outperform medical experts in clinical text summarization.

Diagnosing our datasets: How does my language model learn clinical information? Adapted large language models can outperform medical experts in clinical text summarization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:12.050114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:09.232187Z digest=sha256:dc241228234e615194216127d4f95dbb7419ec5a276d83357b8a2b488403a6db

Observation 4824c5f2-d689-44d9-b29b-4d9ac1db0003 · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

Diagnosing our datasets: How does my language model learn clinical information? RedPajama: an Open Dataset for Training Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:09.304268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:09.304268Z digest=sha256:58b90fbd6d5cb1e990ff7e254c511abfcd95d497f591e8ad640973c2ef0d2504

Observation 94fb0127-bc3f-49c5-8a98-afb4e78a4432 · outbound

This paper cites Me-llama: Foundation large language models for medical applications.

Diagnosing our datasets: How does my language model learn clinical information? Me-llama: Foundation large language models for medical applications

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:11.787914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:09.452087Z digest=sha256:12fe96c446b36ab7d32bbe3e58f4f182526f6e5f3ae3c76db0f7efe2c7b021f3

Observation 206aacf5-94cc-47cd-84a1-69a23639e6fb · outbound

This paper cites Almanac—retrieval-augmented language models for clinical medicine.

Diagnosing our datasets: How does my language model learn clinical information? Almanac—retrieval-augmented language models for clinical medicine

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:11.516515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:09.631319Z digest=sha256:2de163d7b580e30697043a1abf6dbb83cb2ff9e0d8b2a23b444ae5eb85bf16ad

Observation 14102550-5392-445e-a5c8-6da14551a5b1 · outbound

This paper cites A dataset for evaluating clinical research claims in large language models.

Diagnosing our datasets: How does my language model learn clinical information? A dataset for evaluating clinical research claims in large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:11.295105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:29:09.803009Z digest=sha256:affa763c71427b36f8da37a6121f02d18b4a2f3871ac199f3ca22a3f7f1206c5

Observation d3e7985a-3312-49a9-98c1-2a3fbbba001c · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Diagnosing our datasets: How does my language model learn clinical information? Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:09.919296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:09.919296Z digest=sha256:6e62232722d5c874d6e45845abfb4f94d98655430535a284e7f9efa4e0a32537

Pith citing papers

No inbound Pith citation observations are available.