Pith. sign in

REVIEW 3 major objections 4 minor 42 references

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

T0 review · 3 major / 4 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read TreeProbe shows LLMs systematically drift from Tibetan medicine toward TCM or biomedicine when answering expert medical items.

desk verdict A genuinely careful benchmark construction, but the headline ODLO statistic is not calibrated against the 2:1 distractor design, so the central drift claim needs re-analysis. read the letter →

arxiv 2608.00640 v1 pith:YMXM6EQC submitted 2026-08-01 cs.CL

classification cs.CL
keywords TibetanmedicineculturalbiasLLMevaluationontologydriftMedicalTreeFourTantrastraditionalbenchmarkconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

TreeProbe is a new benchmark for measuring cultural bias in large language models' handling of Tibetan medicine. The paper argues that current LLMs do not fail neutrally: when they make mistakes on Tibetan-medicine questions, their errors consistently shift toward one of two external medical systems, Traditional Chinese Medicine or biomedicine, rather than reflecting confusion inside Tibetan medicine itself. To make this measurable, the benchmark is built on the native Medical Tree (from the Four Tantras) and contains 4,719 expert-adjudicated items over 467 diseases and 10 subtasks. Every evaluated model lands on the external-drift side of the ODLO metric, and models divide into TCM-leaning and biomedical-leaning groups. The paper's controlled probes suggest the drift is not just missing knowledge but a persistent preference to resolve Tibetan-medicine uncertainty with a more available external ontology.

What carries the argument

The load-bearing mechanism is the Medical Tree of Tibetan medicine (the three-root structure of physiopathology, diagnosis, and therapy codified in the Four Tantras), used as an evaluation scaffold. Items are derived by operators over evidence-linked knowledge units; each MCQ's three non-faithful options are authored by separate expert groups to fall into distinct ontological categories (intra-Tibetan, TCM-drift, biomedical-drift). The two metrics carry the analysis: ODLO compares external-drift errors to intra-Tibetan errors, and BTD-LO compares biomedical to TCM drift. This makes the abstract idea of 'cultural bias' into a measurable direction and magnitude of ontology drift.

What would settle it

Translate a random sample of TreeProbe MCQ stems from the neutral Chinese vignettes back into full Tibetan medical terminology and re-run the same models: if the aggregate ODLO falls from the reported >0.48 to near zero, the measured drift was an artifact of vignette neutrality rather than model cultural bias. Or have Tibetan-medicine experts attempt to classify TCD vs BID options from the Tibetan stems alone; if they cannot separate them reliably, the option categories are not clean.

Watch

Extended reading notes

Core claim

The paper's central claim: LLMs in native Tibetan-medical contexts show systematic external ontology drift rather than Tibetan-internal confusion. TreeProbe measures this with four mutually exclusive option categories (faithful Tibetan, intra-Tibetan, TCM-drift, biomedical-drift) and two log-odds metrics, ODLO and BTD-LO. All six tested models land in the external-drift region (ODLO > 0.48), with empty 'Tibetan confusion' quadrants; models split into biomedical-leaning and TCM-leaning groups. Therapy items preferentially pull toward TCM (shared surface vocabulary), while diagnosis and physiopathology items more often pull toward biomedicine. Controlled probes on previously-correct items show

Load-bearing premise

The claims rest on the assumption that the ontology-neutral Chinese vignettes used to elicit TCM and biomedical distractors preserve each item's clinical content without leaking Tibetan-specific cues, so that a TCD or BID option is genuinely external and incorrect under Tibetan medicine rather than an artifact of distractor authoring (§4.4, §4.5.1).

Editorial extensions

If this is right

  • MCQ accuracy on Tibetan medicine understates the problem: errors are not random but concentrated in specific external ontologies, so alignment or fine-tuning must target ontology fidelity, not just accuracy.
  • One-shot prompting dampens drift magnitude but does not flip its direction, implying format adaptation alone will not close the epistemic gap.
  • Therapy items are more susceptible to TCM-style substitution, while diagnosis and physiopathology items more readily trigger biomedical reframing, pointing to different intervention points by root.
  • Models that prefer an external system keep choosing it even when the option is clinically wrong (up to 10.8% for biomedical-leaning models, 40.5% for DeepSeek-Chat under zero-shot), indicating a template-like bias beyond knowledge gaps.
  • The benchmark supplies a reusable item-derivation and distractor-authoring protocol for other low-resource traditional medicine systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the drift is caused by pretraining composition and surface similarity, then augmenting pretraining data with Tibetan-medical text should push ODLO toward zero and BTD-LO toward the axis—a testable prediction the paper does not run.
  • Editorial inference: the ontology-neutral vignette trick could generalize, so the ODLO/BTD-LO pair could serve as a generic cultural-bias diagnostic for any low-resource expert knowledge system competing with high-resource ones.
  • Editorial inference: the paper leaves the language channel untested; presenting the same items in Tibetan, Chinese, and English could reveal whether the drift is ontology-level or language-mediated activation, a natural next experiment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. TreeProbe is presented as the first Tibetan-medicine cultural-bias benchmark, containing 4,719 expert-adjudicated items organized around the three roots of the native Medical Tree, with ten subtasks spanning MCQ and generative QA. The paper describes a four-stage construction pipeline: evidence-linked knowledge-unit extraction from the Four Tantras, operator-guided item instantiation, cross-ontology distractor authoring by separate expert groups, and label-blind adjudication. Six LLMs are evaluated. The paper reports low accuracy on Tibetan medicine and uses two log-odds metrics, ODLO and BTD-LO, plus two controlled probes, to argue that model errors systematically drift toward TCM or biomedical ontologies rather than reflecting intra-Tibetan confusion. A meta-evaluation of GPT-5.4 as a QA judge is also reported.

Significance. If the bias measurement were sound, this would be an important contribution. The benchmark is culturally grounded in the Tree of Medicine, the construction pipeline is unusually careful (evidence-linked units, explicit empty-slot handling, separate expert groups for distractors, label-blind adjudication, filtering rules), and code/data are released. Such a resource is genuinely useful for evaluating low-resource traditional-medicine systems. However, the headline drift result is currently not supported, because ODLO is not chance-corrected for the 2:1 external:internal distractor design. The benchmark itself is likely valuable, but the paper's central diagnostic claim needs substantial reanalysis.

major comments (3)
  1. [§5.3, Eq. (1); §6.2, Fig. 3] The central drift statistic is not interpretable as claimed. Each MCQ contains one FA, one ITD, one TCD, and one BID (§4.1), so if a model has no Tibetan knowledge and guesses uniformly among the four options, errors fall uniformly among the three distractor categories. Under Eq. (1), the expected ODLO is then roughly log((2N+α)/(N+α)) ≈ log 2 ≈ 0.69 for large N, not 0. The paper's threshold 'ODLO>0.48' is therefore below the random-guessing null, and the observation that all models occupy the right half of Fig. 3 does not establish 'systematic external ontology drift.' A chance-corrected metric such as log((N_TCD+N_BID)/(2·N_ITD)) has null 0. Please recompute the aggregate claim with this normalization or an explicit null model of distractor selection, provide confidence intervals, and revise the abstract and §6.2 if the conclusion changes. BTD-LO has a symmetric null and is less affect
  2. [§6.3, Fig. 4] The controlled probes rest on an unstated axiom: a model that answered an original MCQ correctly has deterministic 'access' to the Tibetan answer, so departing from it in Probe 1 must indicate bias. This is not established; the probe is a different decision problem with three system-specific correct answers, and correct MCQ performance does not guarantee robust retrieval under competing correct alternatives. Moreover, in the three-option Probe 1 design, random selection would yield 2/3 external choices; the reported external rates (roughly 15–28% for most models, ~50% for DeepSeek, ~42% for Claude) are below that null, so the probe does not quantify systematic drift. Probe 2's DeepSeek result (40.5% zero-shot selection of the wrong TCM option vs. 24% Tibetan) is suggestive, but a properly specified baseline and uncertainty estimates are needed before the general claim follows.
  3. [§6.2 and Abstract] The causal attribution that drift direction is 'shaped by pretraining data composition and surface resemblance between TCM and Tibetan medicine' is not directly supported. The paper reports no pretraining-corpus analysis and no quantitative measure of surface semantic similarity; the model separation along BTD-LO is only a correlation with provider geography across six models. This explanation should be reframed as a hypothesis, or supported with explicit measurements (e.g., vocabulary-overlap statistics between TCM and Tibetan terms, or controlled manipulations of surface similarity). As written, the abstract presents the causal story as a finding, which is stronger than the evidence warrants.
minor comments (4)
  1. [Table 2] In the GPT-5.4 zero-shot row, '31.4361.43' appears to be a concatenated pair of values (DD and DI). Please fix the formatting.
  2. [References] GPT-5.4 is cited as Achiam et al. 2023, which is the GPT-4 technical report. Add the correct citation or model card for GPT-5.4.
  3. [Appendix B, Table 5] The 'fully instantiated knowledge unit' example appears with an empty Value column in the version I reviewed. If the table is intended to show a concrete example, include the actual Tibetan slot values.
  4. [Figures 5 and 6] The legend text in these figures renders as escaped Unicode sequences (e.g., '/uni00000028/...') rather than readable Tibetan script. Please ensure the Tibetan text displays correctly in the final version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: TreeProbe's benchmark content is externally adjudicated, and the main claims are empirical observations rather than derivations from fitted inputs or self-citation chains.

full rationale

The paper's central empirical claims rest on an independently constructed benchmark: items are derived from the Four Tantras with human expert adjudication (4/5 label-blind agreement), distractors are authored by separate TCM and biomedical physician groups from ontology-neutral vignettes, and category labels are checked for ontological separability and cross-category exclusivity. No parameter is fitted to a subset of the data and then used to 'predict' a closely related quantity. The ODLO/BTD-LO metrics are descriptive log-odds summaries of observed error counts; their interpretation does raise a statistical validity concern (because the MCQ design gives two external distractors for every internal distractor, a uniform random-choice model has expected ODLO ≈ log 2 ≈ 0.69, so ODLO > 0.48 is not by itself evidence of systematic external drift). However, this is a measurement/baseline correctness issue rather than a circularity: the observed counts are not forced by the metric's definition, and the controlled probes in Figure 4 provide additional, non-aggregate evidence. The overlap in which GPT-5.4 drafted reference answers and later served as an auxiliary QA judge is a self-referential design element, but the paper meta-evaluates the judge against human experts (Table 3), and the MCQ and controlled-probe results do not depend on that judge. There are no load-bearing self-citations or imported uniqueness theorems. The benchmark's core claims are therefore not equivalent to their own inputs by construction.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The benchmark's validity rests on the Tibetan-medicine ontology mapping, expert adjudication, cross-ontology distractor authoring, and the assumptions built into the controlled probes. No physical or medical entities are introduced; ODLO and BTD-LO are new metrics but not invented entities.

free parameters (2)
  • ODLO/BTD-LO smoothing constant α = 0.5
    Hand-chosen in Eqs. (1)–(2) to prevent log(0) at zero counts. Affects small-sample drift values but does not determine the qualitative placement of models in Figure 3.
  • Optional-slot exposure probability p = 0.5
    Chosen by design in §4.3.1 to vary premise combinations. Not fitted to data, but it shapes item difficulty and the realized dataset.
assumptions (5)
  • domain assumption The Four Tantras and GB/T 46946–2025 provide a faithful, sufficient basis for operationalizing Tibetan medicine's ontology.
    §4.2 and Appendix B assume the source corpus and disease-code standard adequately represent Tibetan medicine. If the mapping is wrong, benchmark validity fails.
  • domain assumption The four option categories FA, ITD, TCD, BID are mutually exclusive and jointly exhaustive under Tibetan ontology.
    §4.1 and §4.5.1 rely on clean ontological boundaries; the ODLO/BTD-LO metrics treat TCD and BID as separable external-drift channels.
  • domain assumption Chinese ontology-neutral vignettes preserve clinical content across systems without leaking Tibetan-specific cues, and translated TCM/biomedical answers remain plausible within those external systems.
    §4.4 (Stage III) is the mechanism that generates TCD/BID distractors; its validity is not independently tested in the paper.
  • ad hoc to paper A model that correctly answered an original MCQ has access to the Tibetan answer, so departing from it in Probe 1 indicates bias rather than guessing or uncertainty.
    §6.3 assumes near-100% retention of the Tibetan answer in Probe 1. Correct answers can occur by chance, so part of the measured departure may reflect the model's original uncertainty rather than ontology preference.
  • domain assumption GPT-5.4's rubric scores are a valid aggregate proxy for expert judgment in QA evaluation.
    §6.4 reports moderate human–LLM agreement, with a larger gap on Ontology Fidelity. The paper uses the judge as an auxiliary scorer, accepting this approximation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs." pith.science (2026). https://pith.science/paper/YMXM6EQC

@misc{pith2026260800640,
  author       = {Pith},
  title        = {Pith review of: TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YMXM6EQC}},
  note         = {Machine review of arXiv:2608.00640}
}
read the original abstract

Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional medical knowledge systems. Tibetan medicine, one of the world's four major traditional medical systems, has an independent and highly structured theoretical framework. When models lack grounded understanding of Tibetan medicine, they may fall back on dominant epistemic systems and distort the native knowledge structure during reasoning. However, quantitative tools for evaluating cultural bias in Tibetan medicine remain largely absent. To address this gap, we introduce TreeProbe, the first cultural-bias benchmark organized around the native Tree of Medicine framework in Tibetan medicine. It contains 4,719 expert-adjudicated items covering 467 diseases and 10 subtasks along the three roots. Experiments on representative LLMs show that current models remain limited in native Tibetan medical contexts and exhibit systematic external ontology drift. Further analysis reveals that models diverge in whether they drift toward biomedical or TCM reasoning, shaped by pretraining data composition and surface resemblance between TCM and Tibetan medicine. TreeProbe provides a diagnostic benchmark for developing medical AI systems that are both linguistically inclusive and epistemically fair. Code and data are available in an anonymous repository at https://anonymous.4open.science/r/TreeProbe/.

Figures

Figures reproduced from arXiv: 2608.00640 by the authors.

Figure 1
Figure 1. TreeProbe construction pipeline. Lifestyle Regulation (LR), Medication Prescription (MP), External Therapy (ET), and the integrative Comprehensive Treatment Plan (CTP). 4 Dataset Construction TreeProbe is constructed in four stages( [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Open-ended QA scores on the three generation subtasks (PR, CDR, CTP) under zero-shot and one-shot [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Per-model MCQ bias patterns. ODLO (x-axis) [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Choice distributions for two controlled diagnostic tests on items each model correctly answered. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Per-model decomposition of QA performance across five evaluation dimensions: Accuracy, Helpfulness, [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Per-model drift geometry across the three roots of the Medical Tree. Each panel plots ODLO ( [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

42 extracted references · 8 linked inside Pith

  1. [1]

    Aho and Jeffrey D

    Alfred V. Aho and Jeffrey D. Ullman , title =. 1972

  2. [2]

    Publications Manual , year = "1983", publisher =

  3. [3]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243

  4. [4]

    Scalable training of

    Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of

  5. [5]

    Dan Gusfield , title =. 1997

  6. [6]

    Tetreault , title =

    Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =

  7. [7]

    A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

    Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =

  8. [8]

    Findings of the Association for Computational Linguistics: ACL 2022 , pages=

    BBQ: A hand-built bias benchmark for question answering , author=. Findings of the Association for Computational Linguistics: ACL 2022 , pages=

Show all 42 references
  1. [9]

    Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

    What makes reading comprehension questions easier? , author=. Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

  2. [10]

    arXiv preprint arXiv:2303.08774 , year=

    Gpt-4 technical report , author=. arXiv preprint arXiv:2303.08774 , year=

  3. [11]

    arXiv preprint arXiv:2508.06471 , year=

    Glm-4.5: Agentic, reasoning, and coding (arc) foundation models , author=. arXiv preprint arXiv:2508.06471 , year=

  4. [12]

    arXiv preprint arXiv:2312.11805 , year=

    Gemini: a family of highly capable multimodal models , author=. arXiv preprint arXiv:2312.11805 , year=

  5. [13]

    arXiv preprint arXiv:2412.19437 , year=

    Deepseek-v3 technical report , author=. arXiv preprint arXiv:2412.19437 , year=

  6. [14]

    arXiv preprint arXiv:2505.09388 , year=

    Qwen3 technical report , author=. arXiv preprint arXiv:2505.09388 , year=

  7. [15]

    2026 , howpublished =

    Introducing. 2026 , howpublished =

  8. [16]

    Advances in neural information processing systems , volume=

    Language model tokenizers introduce unfairness between languages , author=. Advances in neural information processing systems , volume=

  9. [17]

    Advances in Neural Information Processing Systems , volume=

    Blend: A benchmark for llms on everyday knowledge in diverse cultures and languages , author=. Advances in Neural Information Processing Systems , volume=

  10. [18]

    Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

    Min, Sewon and Lyu, Xinxi and Holtzman, Ari and Artetxe, Mikel and Lewis, Mike and Hajishirzi, Hannaneh and Zettlemoyer, Luke. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?. Proceedings of the 2022 Conference on Empirical Methods in Natural Langua...

  11. [19]

    Billions Left Behind on the Path to Universal Health Coverage , year =

  12. [20]

    2024 , type =

    Draft Global Traditional Medicine Strategy (2025--2034) , institution =. 2024 , type =

  13. [21]

    Traditional, Complementary and Integrative Medicine , year =

  14. [22]

    Journal of ethnobiology and ethnomedicine , volume=

    A comparative study on shared-use medicines in Tibetan and Chinese medicine , author=. Journal of ethnobiology and ethnomedicine , volume=. 2019 , publisher=

  15. [23]

    Philosophy & Technology , volume=

    Decolonial AI: Decolonial theory as sociotechnical foresight in artificial intelligence , author=. Philosophy & Technology , volume=. 2020 , publisher=

  16. [24]

    Proceedings of the 33rd ACM International Conference on Information and Knowledge Management , pages=

    Cultural commonsense knowledge for intercultural dialogues , author=. Proceedings of the 33rd ACM International Conference on Information and Knowledge Management , pages=

  17. [25]

    CulturalBench: a Robust, Diverse and Challenging Benchmark on Measuring (the Lack of) Cultural Knowledge of LLMs , author=

  18. [26]

    The Lancet Regional Health--Western Pacific , volume=

    Large language models and global health equity: a roadmap for equitable adoption in LMICs , author=. The Lancet Regional Health--Western Pacific , volume=. 2025 , publisher=

  19. [27]

    Journal of the American Medical Informatics Association , volume=

    Leveraging large language models to foster equity in healthcare , author=. Journal of the American Medical Informatics Association , volume=. 2024 , publisher=

  20. [28]

    PNAS nexus , volume=

    Cultural bias and cultural alignment of large language models , author=. PNAS nexus , volume=. 2024 , publisher=

  21. [29]

    Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025) , pages=

    Beyond Metrics: Evaluating LLMs Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios , author=. Proceedings of the Sixth Workshop on African Natural Language Processing (AfricaNLP 2025) , pages=

  22. [30]

    Nature Medicine , volume=

    A toolbox for surfacing health equity harms and biases in large language models , author=. Nature Medicine , volume=. 2024 , publisher=

  23. [31]

    Advances in Neural Information Processing Systems , volume=

    Measuring what matters: Construct validity in large language model benchmarks , author=. Advances in Neural Information Processing Systems , volume=

  24. [32]

    Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

    Having beer after prayer? measuring cultural bias in large language models , author=. Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

  25. [33]

    Advances in neural information processing systems , volume=

    Judging llm-as-a-judge with mt-bench and chatbot arena , author=. Advances in neural information processing systems , volume=

  26. [34]

    International Conference on Learning Representations , volume=

    Judgebench: A benchmark for evaluating llm-based judges , author=. International Conference on Learning Representations , volume=

  27. [35]

    arXiv preprint arXiv:2505.08775 , year=

    Healthbench: Evaluating large language models towards improved human health , author=. arXiv preprint arXiv:2505.08775 , year=

  28. [36]

    Nature , volume=

    Large language models encode clinical knowledge , author=. Nature , volume=. 2023 , publisher=

  29. [37]

    arXiv preprint arXiv:2506.04078 , year=

    LLMEval-Med: a real-world clinical benchmark for medical LLMs with physician validation , author=. arXiv preprint arXiv:2506.04078 , year=

  30. [38]

    npj Digital Medicine , year=

    A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains , author=. npj Digital Medicine , year=

  31. [39]

    arXiv preprint arXiv:2511.07148 , year=

    TCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine , author=. arXiv preprint arXiv:2511.07148 , year=

  32. [40]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

    Towards measuring and modeling “culture” in LLMs: A survey , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

  33. [41]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Investigating cultural alignment of large language models , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  34. [42]

    arXiv preprint arXiv:2406.01126 , year=

    Tcmbench: A comprehensive benchmark for evaluating large language models in traditional chinese medicine , author=. arXiv preprint arXiv:2406.01126 , year=

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.