Pith. sign in

Paper Citation Record · LEDGER

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks

As of 17 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2507.23146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23146 v4

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:06:08.015456Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact6
  • verified fuzzy23
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d54002f8-0522-4c64-82c3-6fc6fa672aac · outbound

This paper cites Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models

Reference 1

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T11:06:09.261423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.460274Z digest=sha256:7f5aac3c0c49302c2eacb7ed6bac10d8a2884d5eb4d50a02a6296d5f6c433bca

Observation 5e62f23b-f03b-4b46-be2f-432198371a61 · outbound

This paper cites Towards automated phenotype definition extraction using large language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Towards automated phenotype definition extraction using large language models

Reference 2

Resolution
verified exact
doi, observed 2026-08-06T11:06:09.037740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.504790Z digest=sha256:c6c4d89a28f73207e7dd7f68cf6ff0700d84b9af7ca98adcfc78f63ca48a2c2b

Observation ecce1537-4bc9-4df4-a0ef-6e49733c6165 · outbound

This paper cites Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network

Reference 3

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T11:06:09.737003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.551733Z digest=sha256:503438104f7950bc66bda6333f246826f68fd804c5709943edab0cbd41eff324

Observation 60002c1c-ee0a-45fe-b3ce-cf618ee97445 · outbound

This paper cites A general framework for developing computable clinical phenotype algorithms.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A general framework for developing computable clinical phenotype algorithms

Reference 4

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.810249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.601759Z digest=sha256:48a3bdf1cdfdc589f95fb748610acfc63db3e0462a23e820932f68357c44f2f3

Observation 3e8d64af-19ab-4855-86b7-148a74d8cde3 · outbound

This paper cites SHREC: A framework for advancing next-generation computational phenotyping with large language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks SHREC: A framework for advancing next-generation computational phenotyping with large language models

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.636195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.617860Z digest=sha256:a6779e71892a93dd6919cc3f65f1e0090ebe62bf52058e26cdefc93f3bf2ee92

Observation 0a146695-c30c-4ffe-99bb-97119e4c1640 · outbound

This paper cites Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T11:06:09.489695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.656771Z digest=sha256:29550c371e11c1f13bde5da4f304e49b27341ac82c90ec7a9bd529da85d2e686

Observation b616fc55-cd06-4d26-aa42-b5ff3b0a8acf · outbound

This paper cites Deductive Verification of Chain-of- Thought Reasoning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Deductive Verification of Chain-of- Thought Reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.822496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.701412Z digest=sha256:e49e6304fa22e6a9274770ab7809307d6b91b646de6f6314eb08aedad89e399f

Observation 7d356364-cca8-44a5-952a-e662e2c37094 · outbound

This paper cites Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:06.741394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:06.741394Z digest=sha256:0404d92a50e0324e511e3e689cd555445e6cdccf5d625cfac686b835405828ed

Observation abda25da-8ce7-42d6-8127-16a5f79c8862 · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large language models can be easily distracted by irrelevant context

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.813349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.819302Z digest=sha256:ba112bcee67faeaca138db3e63f20c98feccde7e48bc4c9a65323ac3734f85c9

Observation ecb8e2b1-4c21-4515-bfb6-8dfb8bea19a6 · outbound

This paper cites Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.804372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.857730Z digest=sha256:0716a40a6fee5c4851f9281c601c98f4a45e8380ba4731f6f461681e7a381b0a

Observation b8782a17-710c-4bf8-9506-823e92a9e840 · outbound

This paper cites Language Models Are Greedy Reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Language Models Are Greedy Reasoners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.794725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.899616Z digest=sha256:5b22fdd168d0bda54e4d94f23e60285a4d40b00723189ced209adfffe22d24d0

Observation b8b0937b-036d-4321-8968-deee035c40ce · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.785754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.916852Z digest=sha256:673dc4dc566f856c8dc24835f5165a2cbff2eb779d4aafa9b75623df020ce0a7

Observation 5b9ef066-84a6-4d58-9623-7783d4479da1 · outbound

This paper cites Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.777191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:06.960032Z digest=sha256:97e1837ad1add92e70f6695ffb706a56cc9c9d9dbc4b86039339320a751d03e1

Observation 3184707f-4722-46b9-90f4-3ff422920433 · outbound

This paper cites Chain-of- Thought Reasoning in the Wild is not Always Faithful.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain-of- Thought Reasoning in the Wild is not Always Faithful

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.767840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.021933Z digest=sha256:9e649dbf167aa8fa4305af38bb86c9adaab6e668130d0bdcf468ff3c129468c9

Observation ef73b372-82cc-494e-bfc6-1fdb4ff1aad1 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Reasoning Models Don't Always Say What They Think

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.060768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.060768Z digest=sha256:5d8fbc8b522fa399000cd54ca272cefdfb087ac12a62daa03fbfec1f6169c7d9

Observation 7f52cf14-0354-4663-9143-5919abe82b67 · outbound

This paper cites [cited 2025 Jun 10].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks [cited 2025 Jun 10]

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.758678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.094504Z digest=sha256:01f2de3a0093f31bb04b6df5a075cab3342e8d6c668cf0195f95010e7e6f64a1

Observation 024f890d-d398-4dd5-8c37-feb8dc44ad1d · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.155650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.155650Z digest=sha256:834a3ec0e3a613afc57eb6eb6ee6114f4e319e537d0d75b77c340dbc41f01628

Observation f893180c-571f-442e-9f2e-92dc02c39bad · outbound

This paper cites PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.749502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.193010Z digest=sha256:3c099e1ec6391ba80647096c0b98296e23f3619f3a683a9298ac4f721d16ac8c

Observation 891c4f96-23b5-4df3-b540-349cd65ad081 · outbound

This paper cites Rule-Based Cohort Definitions for Acute Respiratory Failure: Electronic Phenotyping Algorithm.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Rule-Based Cohort Definitions for Acute Respiratory Failure: Electronic Phenotyping Algorithm

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-06T11:06:07.220218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.220218Z digest=sha256:ba1b07f753d883626ff4e088b91832e9d598d99050e50fb785c3351f33d0462b

Observation 123c2014-cb64-4304-b7a1-370b2f64a11e · outbound

This paper cites The eICU Collaborative Research Database, a freely available multi-center database for critical care research.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks The eICU Collaborative Research Database, a freely available multi-center database for critical care research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.263225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.263225Z digest=sha256:a330c83da76e9e530556cde4e9f170bc75edde8d75839a8b5b4d8c513a73bf8a

Observation b02f6982-1d95-44ce-ac39-6fdde679ab57 · outbound

This paper cites Ollama; 2024 [cited 2024 Dec 18].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Ollama; 2024 [cited 2024 Dec 18]

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.629792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.294351Z digest=sha256:cdb5d9ba2f759a8c06a2d5aa58d626e4ded711f16cdcb5fe83907c73b412e569

Observation 7778794b-f615-4645-9215-924a56d0266a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.310546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.310546Z digest=sha256:0b564e3afc00bac19091fcfec08e2c5c6ca058bff9600e61cec46dcc1393943d

Observation 4d48fb61-40d5-4023-a5ab-870f5cdf9f9e · outbound

This paper cites Interrater reliability: the kappa statistic.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Interrater reliability: the kappa statistic

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.476809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.354644Z digest=sha256:c35250ad57c9825c6b94ee89a7a67c53268c44a67c728d30fdbe443ec5776910

Observation 4bddb60d-a483-4e33-b631-11d2be39d3cb · outbound

This paper cites Emergent Abilities of Large Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Emergent Abilities of Large Language Models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.353567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.391617Z digest=sha256:0499c021d20d08ffee5cc421907f0e172b6aafae2f13805cb4bca4889395aa32

Observation da6d56fd-43bf-4ae9-b22c-afc316fc28a2 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks ReAct: Synergizing Reasoning and Acting in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.424864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.424864Z digest=sha256:0b8dbfac8351ca5ec48e147a04643859fc73b400116bd45455239eaa212ad95d

Observation b5cbca71-b2bf-4625-ac33-4dee552f2ab7 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.153034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.465902Z digest=sha256:70dcec266b8c863f6c255a87b24ae4bc42018bba2a956962b86a98a417af4463

Observation 69bb715c-ee4a-479f-acbe-59de0fd428f0 · outbound

This paper cites Large language models are zero-shot reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large language models are zero-shot reasoners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.061431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.509894Z digest=sha256:ae2421ced2780e9f628430726530d59c328aafcb79bd4d3c251bf03af362c1f6

Observation 1056e776-a356-423a-997c-8a9764e62143 · outbound

This paper cites Large Language Models Still Can’t Plan (A Benchmark for LLMs on Planning and Reasoning about Change).

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large Language Models Still Can’t Plan (A Benchmark for LLMs on Planning and Reasoning about Change)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.960175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.536971Z digest=sha256:36c37f5f6f654041687b51cf00e045733bdbb4e28a985bda732cdd5811c93c04

Observation c2544d17-6930-4b8d-88d5-f83cfb1cdac9 · outbound

This paper cites GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.559969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.559969Z digest=sha256:9aa2c463ad94e1f2d55bae4a02ee8ef5799b19943de66b78b2a599565d90ef75

Observation 888dab80-b151-4868-8594-17139b86394a · outbound

This paper cites A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems [Internet]

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.423912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.597872Z digest=sha256:356179cd205857b18136db85a63786ff16c751483d9c07511a5568e7792e5699

Observation fa68a6f3-8dcf-4edf-8fa5-97cb8720bf9d · outbound

This paper cites LLM-based agentic systems in medicine and healthcare.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks LLM-based agentic systems in medicine and healthcare

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.642134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.642134Z digest=sha256:2337685a9951780fff064dfb85e2c83b3a04afe69a145ff22300febef0882f9a

Observation a3afde07-30b4-4c87-9655-91f2c8b96af5 · outbound

This paper cites Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models [Internet]

Reference 32

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.277622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.686497Z digest=sha256:ae56f1c2ad2881bd22dc66ba9d253df6842a8f46b8bf0316f1f5ffcf05a537e4

Observation 3edaae83-e511-42fa-9c1c-11e2599f7b9f · outbound

This paper cites Understanding Reasoning in Thinking Language Models via Steering Vectors.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Understanding Reasoning in Thinking Language Models via Steering Vectors

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.802590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.751682Z digest=sha256:02ff55ccc09e6d25b941b05a68d06be4c96797cf7ff7a14b0001e51996f16bb6

Observation 4807b643-07c5-47b0-9b0d-a06a4faff031 · outbound

This paper cites Transformer Circuits [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Transformer Circuits [Internet]

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.695705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.768976Z digest=sha256:18e2f3110d3f0ab61fd09a4f460e87957d43105ee2e319d6b92a51852b5492ee

Observation 9914b539-2384-4531-91b6-7ef5e1c56318 · outbound

This paper cites Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.565561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.797726Z digest=sha256:34673176e7ea040f7c3adeb1ae48f55d4263dc0c54674517966f814fbc885f16

Observation fcfb08dd-3b77-4e4b-a4ce-7492b59e8e6f · outbound

This paper cites Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.443518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.837016Z digest=sha256:2dfad9c33b3d85f54cc473861ff2b8d55e965032cf349499c5948521e0b3703d

Observation b1ab383c-93a7-4d8e-85fb-177e7e7fc905 · outbound

This paper cites A Methodology for Generating and Optimizing Chain-of-Thought Based on Knowledge Graphs.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A Methodology for Generating and Optimizing Chain-of-Thought Based on Knowledge Graphs

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.130096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.858668Z digest=sha256:00d5465d597663ef1661e2c2c04ee79ff247c515b01c214e06acc11138ac1060

Observation ed20780e-9f5f-4627-ab4f-07f56ea36138 · outbound

This paper cites Training language models to follow instructions with human feedback.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Training language models to follow instructions with human feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.322042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.903626Z digest=sha256:246a6b1dfd4a6900429dca541669f5c1bcb58935e0c534be7bbdcc8b3c35ceb8

Observation 5dbf6a72-a40e-4f07-bcc7-ac3a419b23c0 · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks STaR: Bootstrapping Reasoning With Reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.189166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.943266Z digest=sha256:0e8f5f09f2a63cd00cc9a9389e9b2690c1276a563d4655ed913a2bf7fbcb61f9

Observation 20faaa8f-1841-409f-b20a-4059bee80af7 · outbound

This paper cites Leap-of-thought: teaching pre-trained models to systematically reason over implicit knowledge.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Leap-of-thought: teaching pre-trained models to systematically reason over implicit knowledge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:09.967053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:07.980640Z digest=sha256:f8e5a17831f5474f855c1134d140cb1b4bf5f9b75d88bd0e5984dc7195236220

Observation 136d5028-6b67-4198-b405-940d93b9e715 · outbound

This paper cites On contrastive learning for likelihood-free inference.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks On contrastive learning for likelihood-free inference

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:09.853077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:06:08.015456Z digest=sha256:6666fc5419a99de7e82614191431275b5bdb19dcdc1c7565caac1f66e10f4096

Pith citing papers

No inbound Pith citation observations are available.