Pith. sign in

Paper Citation Record · LEDGER

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2507.23146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23146 v4

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:06:08.015456Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact6
  • verified fuzzy23
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d54002f8-0522-4c64-82c3-6fc6fa672aac · outbound

This paper cites Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models

Reference 1

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T11:06:09.261423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.460274Z digest=sha256:fbc1e99dc9e2810613b36a8a22254997a83d0914f85c49cf9b0a23af6d696f19

Observation 5e62f23b-f03b-4b46-be2f-432198371a61 · outbound

This paper cites Towards automated phenotype definition extraction using large language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Towards automated phenotype definition extraction using large language models

Reference 2

Resolution
verified exact
doi, observed 2026-08-06T11:06:09.037740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.504790Z digest=sha256:3161c765f223a40d412fba9250f94eb4007efc95e2313a519faf57819344d482

Observation ecce1537-4bc9-4df4-a0ef-6e49733c6165 · outbound

This paper cites Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network

Reference 3

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T11:06:09.737003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.551733Z digest=sha256:302bfb3a20d3515fab73ed77ee4919b9c023f7a0ee8b077cca972b3989228fe0

Observation 60002c1c-ee0a-45fe-b3ce-cf618ee97445 · outbound

This paper cites A general framework for developing computable clinical phenotype algorithms.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A general framework for developing computable clinical phenotype algorithms

Reference 4

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.810249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.601759Z digest=sha256:f735220802cdab4b9c86a0f8f93a7082eab2e11cd4a4e25a16790b42e60ed05d

Observation 3e8d64af-19ab-4855-86b7-148a74d8cde3 · outbound

This paper cites SHREC: A framework for advancing next-generation computational phenotyping with large language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks SHREC: A framework for advancing next-generation computational phenotyping with large language models

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.636195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.617860Z digest=sha256:e803a0d9db4fcbec78bb01f2ee4280c33f6372050b3fd2e96700f9d9c741738e

Observation 0a146695-c30c-4ffe-99bb-97119e4c1640 · outbound

This paper cites Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T11:06:09.489695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.656771Z digest=sha256:274a7102714df73af479a5b6a6d14503940c6c80ab6abcbeb9cb8562289b71e9

Observation b616fc55-cd06-4d26-aa42-b5ff3b0a8acf · outbound

This paper cites Deductive Verification of Chain-of- Thought Reasoning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Deductive Verification of Chain-of- Thought Reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.822496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.701412Z digest=sha256:3e0c32a8b5560551320b2dea36d7006e58aa4952f6dd027a16805c91f5864304

Observation 7d356364-cca8-44a5-952a-e662e2c37094 · outbound

This paper cites Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:06.741394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:06.741394Z digest=sha256:52d382c319598542b08db4befb7e5afb5f0fbca5fec8331c4e4a01d3ae5e7f16

Observation abda25da-8ce7-42d6-8127-16a5f79c8862 · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large language models can be easily distracted by irrelevant context

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.813349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.819302Z digest=sha256:3db76a7a4bd2fcbb45a36ffa273584093e96365c40d813a526e0446074e9de75

Observation ecb8e2b1-4c21-4515-bfb6-8dfb8bea19a6 · outbound

This paper cites Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.804372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.857730Z digest=sha256:33eef48fcda560cb26d16ab8c7c8e3e6f88942091165d363eba6eed4ccdd7c9d

Observation b8782a17-710c-4bf8-9506-823e92a9e840 · outbound

This paper cites Language Models Are Greedy Reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Language Models Are Greedy Reasoners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.794725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.899616Z digest=sha256:f5659f16093cb7ad6a47c49d599831def9298cc75bb5cd900962caa07c18fbd1

Observation b8b0937b-036d-4321-8968-deee035c40ce · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.785754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.916852Z digest=sha256:6185e4939a7924fa39b7025b0b6297845baf452dd042d0dec704961032a882c7

Observation 5b9ef066-84a6-4d58-9623-7783d4479da1 · outbound

This paper cites Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.777191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:06.960032Z digest=sha256:30da2966a7e16506686e11e2f870905d3edd1f08e1f9ce944a4bbc32ef90ebcd

Observation 3184707f-4722-46b9-90f4-3ff422920433 · outbound

This paper cites Chain-of- Thought Reasoning in the Wild is not Always Faithful.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain-of- Thought Reasoning in the Wild is not Always Faithful

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.767840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.021933Z digest=sha256:b290bf3beeb2272e046ed2ced8c42186cfe73c5a2374f8654fb51bea37f58441

Observation ef73b372-82cc-494e-bfc6-1fdb4ff1aad1 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Reasoning Models Don't Always Say What They Think

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.060768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.060768Z digest=sha256:c3699117285416d399546211ac58ce70a821a5d790b4f122c21e5b531a2004cb

Observation 7f52cf14-0354-4663-9143-5919abe82b67 · outbound

This paper cites [cited 2025 Jun 10].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks [cited 2025 Jun 10]

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.758678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.094504Z digest=sha256:7087d0b5d8339705e8387d2cfe02ce037a3106b4997f162753b89f32a1517eff

Observation 024f890d-d398-4dd5-8c37-feb8dc44ad1d · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.155650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.155650Z digest=sha256:1e759e0e13387372543a1b24274e216bc511415a955f33e5c567bcd06659311d

Observation f893180c-571f-442e-9f2e-92dc02c39bad · outbound

This paper cites PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks PHEONA: An Evaluation Framework for Large Language Model-based Approaches to Computational Phenotyping

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.749502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.193010Z digest=sha256:3171855deeab2d636d8109b9bdf21f49c0759dadb572cac947d4f7b6bdaff07c

Observation 891c4f96-23b5-4df3-b540-349cd65ad081 · outbound

This paper cites Rule-Based Cohort Definitions for Acute Respiratory Failure: Electronic Phenotyping Algorithm.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Rule-Based Cohort Definitions for Acute Respiratory Failure: Electronic Phenotyping Algorithm

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-06T11:06:07.220218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.220218Z digest=sha256:ee50969d1f8b872cfbef8b8d056e652b2540334881443d16d4dc34b8af301beb

Observation 123c2014-cb64-4304-b7a1-370b2f64a11e · outbound

This paper cites The eICU Collaborative Research Database, a freely available multi-center database for critical care research.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks The eICU Collaborative Research Database, a freely available multi-center database for critical care research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.263225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.263225Z digest=sha256:f72874bf454a59ae53ba0aa9107939c1c58e162017444b4be828a85d00e564c9

Observation b02f6982-1d95-44ce-ac39-6fdde679ab57 · outbound

This paper cites Ollama; 2024 [cited 2024 Dec 18].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Ollama; 2024 [cited 2024 Dec 18]

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.629792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.294351Z digest=sha256:2136f7b1ff0bbfde35254c6c2336616c87024dc1091690a989d009d0fb27e69f

Observation 7778794b-f615-4645-9215-924a56d0266a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.310546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.310546Z digest=sha256:7e0d4472a558765432036b1457e49c6784c62a628caaa1ec71020e464fe13b0b

Observation 4d48fb61-40d5-4023-a5ab-870f5cdf9f9e · outbound

This paper cites Interrater reliability: the kappa statistic.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Interrater reliability: the kappa statistic

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.476809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.354644Z digest=sha256:0ed67cb81113d0b2254048632a19735d8da96bbd1edd5f98b9d87d1a5ad866a2

Observation 4bddb60d-a483-4e33-b631-11d2be39d3cb · outbound

This paper cites Emergent Abilities of Large Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Emergent Abilities of Large Language Models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.353567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.391617Z digest=sha256:5b6f6cda7b0b582d1b723257a07079a5418abf1995696cfa0949d2ce75859f44

Observation da6d56fd-43bf-4ae9-b22c-afc316fc28a2 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks ReAct: Synergizing Reasoning and Acting in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.424864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.424864Z digest=sha256:74518d2804d5bcbd41ba7073edec8b55430530f6168b7ac3f06c2c32c1392fa2

Observation b5cbca71-b2bf-4625-ac33-4dee552f2ab7 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.153034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.465902Z digest=sha256:ec62ff9affaa38d4982683a34fe54cc08e782eccfb31e7c2fe27409bb6d3d110

Observation 69bb715c-ee4a-479f-acbe-59de0fd428f0 · outbound

This paper cites Large language models are zero-shot reasoners.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large language models are zero-shot reasoners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:11.061431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.509894Z digest=sha256:f3774eecd3d555af92af72f815aa3d3d5526f07a91ef594f25f1385fb5d9cfa7

Observation 1056e776-a356-423a-997c-8a9764e62143 · outbound

This paper cites Large Language Models Still Can’t Plan (A Benchmark for LLMs on Planning and Reasoning about Change).

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Large Language Models Still Can’t Plan (A Benchmark for LLMs on Planning and Reasoning about Change)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.960175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.536971Z digest=sha256:820844f61e0c24953910dd6250852f465148b479d1b43db1ceba9c6b1438bccb

Observation c2544d17-6930-4b8d-88d5-f83cfb1cdac9 · outbound

This paper cites GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.559969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.559969Z digest=sha256:bfc98c3693bd95c975661cc14edc3441d6350f0827b0f124cf3ff418fefd8f5c

Observation 888dab80-b151-4868-8594-17139b86394a · outbound

This paper cites A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems [Internet]

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.423912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.597872Z digest=sha256:cbe0ee9109444c726981a4a17085944648f40d6049eb536c612f4813292673b2

Observation fa68a6f3-8dcf-4edf-8fa5-97cb8720bf9d · outbound

This paper cites LLM-based agentic systems in medicine and healthcare.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks LLM-based agentic systems in medicine and healthcare

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.642134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.642134Z digest=sha256:7ab5d662337ec422b25e7dafe6c52585b545049aae5b0bd6557b685cb10e8936

Observation a3afde07-30b4-4c87-9655-91f2c8b96af5 · outbound

This paper cites Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models [Internet]

Reference 32

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.277622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.686497Z digest=sha256:1f580991ed5d7052ff3fa0df9940808616749838158d7b4d5dfd81328f46f84f

Observation 3edaae83-e511-42fa-9c1c-11e2599f7b9f · outbound

This paper cites Understanding Reasoning in Thinking Language Models via Steering Vectors.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Understanding Reasoning in Thinking Language Models via Steering Vectors

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.802590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.751682Z digest=sha256:3e2cec8180896f6f78f39ab6732299c76a9319250287aad81ffe0d6bb82c25ca

Observation 4807b643-07c5-47b0-9b0d-a06a4faff031 · outbound

This paper cites Transformer Circuits [Internet].

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Transformer Circuits [Internet]

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.695705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.768976Z digest=sha256:61f8489b4716ebb3d192c4e82c0dfdd9b2bbaea215abaa5ea78288b8c79b41ab

Observation 9914b539-2384-4531-91b6-7ef5e1c56318 · outbound

This paper cites Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.565561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.797726Z digest=sha256:4d51fa85efdc90e18045cb6be47c99586aec1d492fbd8b5b4af5dd4d86d2b496

Observation fcfb08dd-3b77-4e4b-a4ce-7492b59e8e6f · outbound

This paper cites Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.443518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.837016Z digest=sha256:f54e5508d94e40937362896ef3dc58b62550c70796446e3ba6a2ed1032958318

Observation b1ab383c-93a7-4d8e-85fb-177e7e7fc905 · outbound

This paper cites A Methodology for Generating and Optimizing Chain-of-Thought Based on Knowledge Graphs.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks A Methodology for Generating and Optimizing Chain-of-Thought Based on Knowledge Graphs

Reference 37

Resolution
verified exact
doi, observed 2026-08-06T11:06:08.130096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.858668Z digest=sha256:6e64e94e39bd7ec3133fe7fc7e75c71a153900f59f1785f23936d05a3cda94c6

Observation ed20780e-9f5f-4627-ab4f-07f56ea36138 · outbound

This paper cites Training language models to follow instructions with human feedback.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Training language models to follow instructions with human feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.322042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.903626Z digest=sha256:95797b0b7c0b5e5e1d6045fdff0e0786bfd247d58d9b6456b5a49f561abf1f2d

Observation 5dbf6a72-a40e-4f07-bcc7-ac3a419b23c0 · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks STaR: Bootstrapping Reasoning With Reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:10.189166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.943266Z digest=sha256:b00239ee087a9fed8df4396d43613521b8e7346d781a5b9988848af038eb2b52

Observation 20faaa8f-1841-409f-b20a-4059bee80af7 · outbound

This paper cites Leap-of-thought: teaching pre-trained models to systematically reason over implicit knowledge.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Leap-of-thought: teaching pre-trained models to systematically reason over implicit knowledge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:09.967053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:07.980640Z digest=sha256:8423a85cf8be6754fbc30f8f25ae00b424bd3ec4e27be19ca682ecd9514ea046

Observation 136d5028-6b67-4198-b405-940d93b9e715 · outbound

This paper cites On contrastive learning for likelihood-free inference.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks On contrastive learning for likelihood-free inference

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:06:09.853077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:06:08.015456Z digest=sha256:b5f076302507bc6b9064105c3f1d979c3e14878227b6db38a5315aa76ef9c679

Pith citing papers

No inbound Pith citation observations are available.