Pith. sign in

Paper Citation Record · LEDGER

BeHonest: Benchmarking Honesty in Large Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2406.13261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.13261 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:38:34.500200Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:56.663568Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1cc81a92-98cc-4197-9b84-da0c2dcf4763 · inbound

Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks cites this paper.

Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks BeHonest: Benchmarking Honesty in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:50.708278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:50.708278Z digest=sha256:ea7357a35b9cf22fee45e7b7bfb208afffa25b7f7ac9fa578d7479b8426eb2b7

Observation 4f920a70-598d-475b-82ab-86009dc57379 · inbound

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine cites this paper.

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine BeHonest: Benchmarking Honesty in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:53.798081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:53.798081Z digest=sha256:77456ab93a0874519ab8fbde73d76e513abb726d33bec99d30a1672ffc88817e

Observation 716eb961-9bcf-48e3-8049-2dbe0bfc9589 · inbound

LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction cites this paper.

LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction BeHonest: Benchmarking Honesty in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:58.189681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:58.189681Z digest=sha256:3b4400bec162d61a8d351c5159f771cb3c5fa1fea6041bdf1c1177045dea23a4

Observation 43f927cd-3bec-473e-9032-bf525bfac3e5 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges BeHonest: Benchmarking Honesty in Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.291665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.291665Z digest=sha256:501e7751f99563aa0c7981297531d67477702be3a34decdee8fdec16deff824c

Observation 75180596-08b1-4632-8046-36a663f80dea · inbound

Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning cites this paper.

Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning BeHonest: Benchmarking Honesty in Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:31:24.489114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T02:47:10.818363Z digest=sha256:f5fceabd6b8fd4a8c650cba1e86d4be602ee1afe9cb16e3e3639e57f7dade51b

Observation 13d213e0-542c-436b-8658-ed6f6f6dc5d4 · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks BeHonest: Benchmarking Honesty in Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.935337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:62c2cf2899018eb7302928b441a0b90ffb8e478625cf38b15835c4514b3c1198

Observation f48f2433-9f2c-487e-91e2-55c73e13ae0e · inbound

DECOR: Auditing LLM Deception via Information Manipulation Theory cites this paper.

DECOR: Auditing LLM Deception via Information Manipulation Theory BeHonest: Benchmarking Honesty in Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:28:05.381264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T06:27:10.445757Z digest=sha256:17066c7cd74e885da6f8e2810ef365f24afbb783e4e0269429e4b2de5670b361

Observation bf0bdbc0-8745-477a-894d-c71c971ee0a0 · inbound

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence cites this paper.

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence BeHonest: Benchmarking Honesty in Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.590107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T14:39:51.672024Z digest=sha256:a66089f6232a5bf1d1b29d8e3e547cefda7936a896f4a74cfcdf10ea435356c4

Observation cfbb8b37-c1c3-4263-bdd6-9f2c03bfb24b · inbound

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence cites this paper.

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence BeHonest: Benchmarking Honesty in Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:54:36.985886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T10:45:44.127391Z digest=sha256:0416a9d8f0caf51c2325f42e1728890f4c2044d8bb951bcc48c4623e686d8bc3

Observation d33f246a-485b-4c20-9b3e-204721c35d77 · inbound

RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue cites this paper.

RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue BeHonest: Benchmarking Honesty in Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.688342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:34:39.457798Z digest=sha256:aaf10ebaa9ab434d646822f6da9a2fde3d49fc15f4ea28b76d912e8358533990

Observation 05a3b2ca-af7e-45c6-bd80-db214128ad5e · inbound

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing cites this paper.

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing BeHonest: Benchmarking Honesty in Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.665646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T01:28:38.810889Z digest=sha256:e07b2832a71f0723f8ff9512bd7d6d73dbe92687006414c94079fb5186b106fc

Observation fc86b4ff-f50e-443f-b78b-adfa071959e4 · inbound

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch cites this paper.

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch BeHonest: Benchmarking Honesty in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:15:50.166929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:15:50.166929Z digest=sha256:796078bbf538411cb79eb06119fbd1d32fd92835a5068a3b1fdc1807bbe63f96

Observation fc96404e-bfe2-44c4-a015-a4e9f1703cff · inbound

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference cites this paper.

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference BeHonest: Benchmarking Honesty in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:34.500200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:38:34.500200Z digest=sha256:26ed86f01ab082e7a61caeb29b8ffed9b570a4b9fdc65a739e70fb345f33a504

Observation 4ad81314-3d72-4c7c-b999-843c4ecb3b3c · inbound

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures cites this paper.

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures BeHonest: Benchmarking Honesty in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T15:06:26.693276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T15:06:26.693276Z digest=sha256:90f356d28936682eb18df766bc9f85821233915dd85a64741bb2ca152a82c9ef