Pith. sign in

REVIEW 3 major objections 3 minor 93 references

Behavioural-faithfulness detection scores are under-specified: the same evaluators rank differently depending on whether the score targets exposure to a trait-inducing prompt or manifestation of the trait in the reply.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 07:37 UTC pith:OXATCTBW

load-bearing objection Under-specification of behavioural-detection AUROC is real: this paper shows it with a clean design, but the manifestation arm leans on an automatic judge's label; the core interaction still holds. the 3 major comments →

arxiv 2607.09306 v3 pith:OXATCTBW submitted 2026-07-10 cs.CL cs.AIcs.HCcs.LG

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

classification cs.CL cs.AIcs.HCcs.LG
keywords behavioural auditingexposure vs manifestationestimandAUROCmeasurement validityoutput resolutionsmall language model auditorfrontier judge
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper shows that a single detection score for language-model behavioural auditing is under-specified because it silently answers one of two different questions: whether a reply was produced under a behaviour-inducing prompt (exposure) or whether the behaviour actually surfaced in the reply (manifestation). Scoring a 146-million-parameter auditor and a frontier judge on the same 720 replies, the gap between the two instruments moves by about 0.2 AUROC — a standard detection-accuracy measure — when the target changes, at every output format tested. Under the judge's deployed yes/no interface, the ranking reverses: the compact auditor leads on exposure (0.804 vs 0.718) and trails on manifestation (0.690 vs 0.811). Matching the output resolution from either direction removes the reversal but not the interaction, whose confidence intervals exclude zero at all three resolutions (0.207, 0.237, 0.169). If correct, behavioural-faithfulness claims are comparable only when they state the estimand, the evaluator, and the evaluator's output interface.

Core claim

The paper's central finding is that measurement target and output resolution are separable factors that jointly determine which behavioural-faithfulness evaluator appears better. On the identical 720 replies, the auditor's frozen-representation read-out achieves 0.804 AUROC against the exposure label versus the frontier judge's 0.718 under a single verdict, while against the manifestation label the judge leads 0.811 versus 0.690. The difference between the two gaps — the estimand interaction — is positive and excludes zero at all three tested output resolutions (0.207, 0.237, 0.169). When the judge is asked a target-specific question with a continuous confidence score, or when the auditor's

What carries the argument

The central objects are two labels on the same 720 replies — exposure (whether the reply was generated under a trait-inducing prompt) and manifestation (whether the trait actually surfaced, as judged by an independent frontier judge) — scored against two instruments at three output resolutions. The load-bearing quantity is the 'estimand interaction': the difference between the auditor-minus-judge AUROC gap on exposure and the same gap on manifestation. Because the interaction excludes zero at every resolution while the sign of the exposure gap flips with output format, it separates the effect of measurement target from the effect of answer interface. The hyperbolic geometry of the auditor's

Load-bearing premise

The manifestation label is assigned by a single automatic frontier judge, so the manifestation half of the comparison and the discordance result measure agreement with that judge, not human-verified behaviour; if that label does not track genuine manifestation, the reversal and the interaction size change meaning.

What would settle it

A prespecified, human-adjudicated manifestation subset, blind to condition and generator, would settle the central claim: if the human-verified labels do not reproduce the roughly 0.2 AUROC estimand interaction and the judge's lead on manifestation, or if the ranking under the verdict interface does not reverse, the paper's conclusion that a single AUROC is under-specified would need revision. A second decisive check would be a supervised read-out on the judge's own representations, were they accessible, to test whether the distances are an artefact of supervised-versus-unsupervised comparison

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A reported behavioural-detection AUROC is under-specified: it must be accompanied by the target label (exposure or manifestation) and the evaluator's output interface (verdict or continuous score) to be comparable.
  • Claims that a small specialised model beats a frontier judge, or vice versa, are not well-formed until the estimand is named; under the deployed verdict interface the auditor leads on exposure (0.804 vs 0.718) and trails on manifestation (0.690 vs 0.811).
  • Matching output resolution — asking the judge a target-specific question answered with a continuous confidence score, or thresholding the auditor's read-out at 0.5 — is necessary to attribute differences to the instrument rather than the answer format; under these matched comparisons the judge leads on manifestation with all intervals excluding zero, and the instruments are statistically indisting
  • The auditor and the judge misclassify different replies (McNemar p ≈ 1.07×10⁻²⁰), and combined two-instrument rules reach operating points neither alone reaches — a high-recall screen and a high-precision confirmer — though these are provisional because the manifestation label comes from a single automatic judge.
  • The transferable detection signal lies in the purpose-built frozen representation and its supervision, not in hyperbolic curvature: a Euclidean-pooled read-out beats the hyperbolic Lorentz read-out on the same representation, and the auditor's pre-trained heads do not transfer zero-shot to this task.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the exposure-manifestation distinction holds for human-verified labels and other behaviours, the same two-target decomposition should be applied to safety benchmarks where an inducing prompt is used as a proxy label (e.g., 'attack succeeded' versus 'harmful content appeared'); single-number comparisons in those settings are likely to mislead in a similar way.
  • The paper could not run a supervised read-out on the judge's own representations because they are inaccessible; a testable extension would be to fit a comparable read-out on the judge's hidden states if they become available and check whether the roughly 0.2 AUROC target-driven gap persists — if it does not, the distances are partly an artefact of the supervised-versus-unsupervised comparison.
  • The discordance result suggests a cheap resident screen: a compact auditor combined with a frontier judge could widen coverage at low cost, and the exact recall-precision trade-off is worth measuring against human-verified manifestation labels in a deployed audit pipeline.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper distinguishes two measurement targets for behavioural-faithfulness detection — exposure (was the reply produced under a trait-inducing condition) and manifestation (did the trait actually surface) — and argues that a single AUROC is under-specified because it conflates the target and the evaluator's output interface. Using a fixed 146M-parameter auditor with a frozen-representation logistic read-out and a frontier zero-shot judge on the same 720 paired replies, with leave-one-generator-out evaluation, it reports that the auditor-vs-judge gap changes by roughly 0.2 AUROC between targets at all three output resolutions, and that under the judge's deployed verdict interface the ranking reverses (auditor leads on exposure, judge on manifestation), while matching output resolution removes the reliable reversal but not the interaction. The paper also reports a discordance analysis, a geometry boundary result, and a human-rating baseline. The design is transparent: single fixed checkpoint, cluster bootstrap on pairs, matched controls approached from both directions, and all table values claimed regenerable from archived per-item files.

Significance. If the result holds, it is a useful measurement contribution: the field's single behavioural-detection AUROCs are indeed ambiguous unless the estimand and output interface are stated together. The paper's strengths are its explicit separation of exposure and manifestation, the objective exposure label, the paired construction, cluster bootstrap inference, leave-one-generator-out evaluation, the two matched-resolution controls, and unusually complete reproducibility (single fixed checkpoint, archived data and scripts, validation checks). The paper is also candid about its limitations. The central reservation is that the manifestation arm is scored against an automatic judge's label, so the headline reversal and the interaction are provisional until human-verified manifestation labels are supplied; the paper itself identifies this as the decisive next step.

major comments (3)
  1. [§3.2 'Two cautions', §5 Limitations, Table 2] The manifestation label is the output of a single automatic frontier judge, and the judge is then scored against that label. The judge's manifestation lead (0.811 vs 0.690; 0.958 vs 0.690 in the matched control) and the interaction intervals (0.207, 0.237, 0.169) are agreement with that automatic reference, not verified manifestation. The paper admits this ('in part agreement between two frontier judges') and calls a human-verified subset 'the decisive next step', yet the abstract and Section 4 report the reversal and interaction as the main result. If labeler and evaluator share model family or rubric, the headline reversal could be labeler–evaluator self-consistency. This is load-bearing: the target-effect claim rests on the manifestation arm. Add a human-adjudicated subset (even a few hundred replies) and re-estimate, or explicitly reframe all manifestation claims as relative to an au
  2. [§3.2, §6 Methods] The 'frontier judge' and the 'independent frontier judge' who assigned the manifestation labels are never named. Methods describe the label as assigned by 'an independent frontier judge under a fixed trait-specific rubric' and the baseline as 'a frontier zero-shot judge', but no model identifiers are given. A reader cannot tell whether the evaluator and labeler are the same model or family, or what 'independent' means here. Please name both models; if they are from the same family, treat the manifestation lead as partly self-consistency. Ideally use a labeler and evaluator from different families, or report their cross-family agreement.
  3. [§3.2, §5 Limitations] The auditor's logistic read-out is trained on the target label for the in-fold generators, while the judge is zero-shot. The paper acknowledges this asymmetry in §5, but the abstract and Section 4 present the auditor-vs-judge distances and the interaction as properties of the instruments. The reported ~0.2 AUROC target effect therefore also encodes a supervised-vs-unsupervised contrast. Either run the proposed supervised read-out on the judge's representations, or state in the abstract that the comparison is between a target-supervised read-out and a zero-shot judge, not between two deployment-ready instruments.
minor comments (3)
  1. [§3.2/Table 2] The text says the binary comparison 'removes the reversal', but the point estimates still reverse: auditor 0.736 vs judge 0.718 on exposure, judge 0.811 vs auditor 0.660 on manifestation. The reversal is removed only in the statistical sense (the exposure CI [-0.017, 0.051] includes zero). Please qualify the wording.
  2. [§6 Methods] The 0.5 threshold is called 'the class-balanced midpoint', but the manifestation label has 61.4% positives, so 0.5 is not class-balanced for the manifestation target. Use class-balanced thresholds per target or clarify the fixed-threshold rationale.
  3. [§3.5] The human evaluation and the trait-detection evaluation are not co-registered, as the paper notes, but the 'text-layer blind spot' is used as a baseline for a different item set. A sentence on the inferential distance between the two studies would help readers avoid treating the human result as a control for the 720-reply comparison.

Circularity Check

0 steps flagged

No significant circularity; the central interaction is empirical, with a disclosed and non-circular limitation on the manifestation reference.

full rationale

The central derivation—exposure versus manifestation targets changing the auditor–judge gap—rests on two independently constructed labels and on actual AUROC computations over the same 720 replies. The exposure label is objective by construction, the logistic read-out is refit under leave-one-generator-out evaluation, and the paired cluster-bootstrap intervals are not fitted to the headline values. The manifestation label is admittedly the output of a single automatic judge, and the paper explicitly states that the manifestation metric is agreement with that judge rather than human-verified behavior and that the judge's advantage is in part agreement between two frontier judges. This is a measurement-validity limitation, not a circular derivation: the reported AUROCs are not equal to the label or to any fitted parameter by construction. The main self-citation (ref. 93, the earlier version of this same work) supports only the text-layer human-rating baseline, not the load-bearing target/interface interaction, and the shared human-rating asset is disclosed. Thus no specific circular reduction is exhibited; the paper is self-contained on its central claim.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

The central claim rests on the definition of the exposure target by construction, the validity of the automatic manifestation reference, and the assumption that the paired construction isolates the trait. The main unexamined load is the single-judge manifestation label and the supervised-versus-zero-shot asymmetry.

free parameters (2)
  • Auditor logistic read-out weights = unknown (refit per leave-one-generator-out fold, separately for exposure and manifestation labels)
    The auditor's per-reply detection scores come from a class-balanced logistic regression on the frozen representation. The fitted weights are the instrument being evaluated, trained on in-fold generator data; the exposure AUROC (0.804) and manifestation AUROC (0.690) depend on this fit.
  • Binary-comparison threshold 0.5 = 0.5
    The auditor's continuous read-out is thresholded at 0.5 to match the judge's verdict interface. The paper states it is not data-optimized, but it is a hand-set design decision that determines the binary-comparison AUROCs (0.736 exposure, 0.660 manifestation) and the interaction 0.169.
axioms (5)
  • domain assumption A reply generated under a trait-inducing companion prompt is a positive exposure case even if the trait did not appear in the reply.
    This defines the exposure estimand by construction rather than by empirical verification. It is central to the exposure AUROC numbers in §3.2 and §6 Methods.
  • domain assumption The manifestation label assigned by a single automatic frontier judge under a fixed rubric is a valid reference for whether the behaviour surfaced.
    The paper explicitly flags this as not human-verified in §3.2 and §5; the manifestation half of the comparison depends on it.
  • domain assumption Style-controlled trait/neutral pairs isolate the trait condition from style and context.
    Shared persona, probe, memory context, and generator within each pair are assumed to leave the trait prompt as the only systematic difference; §6 Methods.
  • domain assumption Leave-one-generator-out over three generators from two providers provides evidence about transfer of the exposure signal.
    The paper disclaims random-effects generalization to a population of generators; results are conditional on these three families, §6 Statistics.
  • standard math Standard bootstrap and McNemar procedures give valid paired uncertainty statements.
    Cluster bootstrap resampling of 360 pairs within held-out generators is used for paired AUROC differences; McNemar is used only for marginal error rates, §6 Statistics.

pith-pipeline@v1.3.0-alltime-deepseek · 15529 in / 16132 out tokens · 171190 ms · 2026-08-02T07:37:20.494747+00:00 · methodology

0 comments
read the original abstract

Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation). Scoring a compact 146-million-parameter auditor's frozen-representation read-out and a frontier judge against each label on the identical 720 replies, the gap between the instruments moves by roughly 0.2 AUROC when the target changes. Under the judge's deployed interface, a single verdict, the ranking reverses: the auditor leads on exposure, 0.804 against 0.718, and trails on manifestation, 0.690 against 0.811. Matching the output resolution from either direction, by asking the judge a target-specific question answered with a continuous confidence score or by thresholding the auditor's read-out, removes the reversal but not the interaction, which excludes zero at all three resolutions (0.207, 0.237 and 0.169). The target governs how far apart the instruments are; the interface governs whether that distance changes their order. The auditor's hyperbolic geometry confers no advantage here. A single behavioural-detection AUROC is under-specified: such claims are comparable only when they state the estimand, the evaluator, and its output interface.

Figures

Figures reproduced from arXiv: 2607.09306 by In Seok Kang, Kwan Soo Shin, Munho Lee, Yunkyung Min.

Figure 1
Figure 1. Figure 1: Auditor per-head accuracy on the 90,314-item held-out set. Accuracy for each of the fourteen classification heads read out from the shared encoder, from a single fixed checkpoint (SHA-256 prefix 3127f6c7), with the number of held-out items for each head shown beside its name. Core heads, which carry the claims reported in the text, are distinguished from auxiliary heads. The dashed line marks chance for a … view at source ↗
Figure 2
Figure 2. Figure 2: The measurement target moves the gap; the output interface moves the ranking. Left, mean leave-one-generator-out AUROC for the compact auditor’s linear read-out and for the frontier judge under the judge’s deployed interface, a single verdict, scored against the exposure label and against the manifestation label on the identical replies; the ordering reverses between the two targets. Centre, the same compa… view at source ↗
Figure 3
Figure 3. Figure 3: The geometry is a boundary, not the mechanism. Left, receiver operating characteristic curves for a hyperbolic Lorentz read-out and a Euclidean-pooled read-out taken from the identical frozen representation under matched capacity (n = 720). Right, the paired difference in AUROC between the two read-outs with its 95 percent bootstrap confidence interval, which excludes zero. The transferable signal is a pro… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

93 extracted references · 78 linked inside Pith

  1. [1]

    Perez, E. et al. Discovering Language Model Behaviors with Model-Written Evaluations. Preprint at https://arxiv.org/abs/2212.09251 (2022)

  2. [2]

    Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Preprint at https://arxiv.org/abs/2306.05685 (2023)

  3. [3]

    Jacobs, A. Z. & Wallach, H. Measurement and Fairness.Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency(2021). doi:10.1145/3442188.3445901

  4. [4]

    D., Bender, E

    Raji, I. D., Bender, E. M., Paullada, A., Denton, E. & Hanna, A. AI and the Everything in the Whole Wide World Benchmark. Preprint at https://arxiv.org/abs/2111.15366 (2021)

  5. [5]

    Dehghani, M. et al. The Benchmark Lottery. Preprint at https://arxiv.org/abs/2107.07002 (2021)

  6. [6]

    Bean, A. M. et al. Measuring what Matters: Construct Validity in Large Language Model Benchmarks. Preprint at https://arxiv.org/abs/2511.04703 (2025)

  7. [7]

    Singh, S. et al. The Leaderboard Illusion. Preprint at https://arxiv.org/abs/2504.20879 (2025)

  8. [8]

    A., Constantinides, M., Tahaei, M

    Septiandri, A. A., Constantinides, M., Tahaei, M. & Quercia, D. WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT?.2023 ACM Conference on Fairness Accountability and Transparency(2023). doi:10.1145/3593013.3593985

  9. [9]

    Buyl, M. et al. Large Language Models Reflect the Ideology of their Creators. Preprint at https://arxiv.org/abs/2410.18417 (2024)

  10. [10]

    Liu, N. F. et al. Lost in the Middle: How Language Models Use Long Contexts. Preprint at https://arxiv.org/abs/2307.03172 (2023)

  11. [11]

    & Zezulka, S

    Freiesleben, T. & Zezulka, S. The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models. Preprint at https://arxiv.org/abs/2510.23191 (2025)

  12. [12]

    Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks

    Freiesleben, T. Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks. Preprint at https://arxiv.org/abs/2603.15121 (2026)

  13. [13]

    Wallach, H. et al. Position: Evaluating Generative AI Systems Is a Social Science Measure- ment Challenge. Preprint at https://arxiv.org/abs/2502.00561 (2025)

  14. [14]

    Salaudeen, O. et al. Measurement to Meaning: A Validity-Centered Framework for AI Evaluation. Preprint at https://arxiv.org/abs/2505.10573 (2025)

  15. [15]

    Weidinger, L. et al. Toward an Evaluation Science for Generative AI Systems. Preprint at https://arxiv.org/abs/2503.05336 (2025)

  16. [16]

    Solaiman, I. et al. Evaluating the Social Impact of Generative AI Systems in Systems and Society. Preprint at https://arxiv.org/abs/2306.05949 (2023)

  17. [17]

    Paruchuri, A. et al. What Are the Odds? Language Models Are Capable of Probabilistic Reasoning. Preprint at https://arxiv.org/abs/2406.12830 (2024)

  18. [18]

    & Poesio, M

    Artstein, R. & Poesio, M. Inter-Coder Agreement for Computational Linguistics.Compu- tational Linguistics(2008). doi:10.1162/coli.07-034-r2

  19. [19]

    Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages.Psychometrika(1947)

    McNemar, Q. Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages.Psychometrika(1947). doi:10.1007/bf02295996

  20. [20]

    A Coefficient of Agreement for Nominal Scales.Educational and Psychological Measurement(1960)

    Cohen, J. A Coefficient of Agreement for Nominal Scales.Educational and Psychological Measurement(1960). doi:10.1177/001316446002000104

  21. [21]

    Fleiss, J. L. Measuring nominal scale agreement among many raters.Psychological Bulletin (1971). doi:10.1037/h0031619

  22. [22]

    Landis, J. R. & Koch, G. G. The Measurement of Observer Agreement for Categorical Data.Biometrics(1977). doi:10.2307/2529310

  23. [23]

    & Kwiatkowski, T

    Pavlick, E. & Kwiatkowski, T. Inherent Disagreements in Human Textual Inferences.Trans- actions of the Association for Computational Linguistics(2019). doi:10.1162/tacl_a_00293

  24. [24]

    & Bansal, M

    Nie, Y., Zhou, X. & Bansal, M. What Can We Learn from Collective Human Opinions on Natural Language Inference Data?. Preprint at https://arxiv.org/abs/2010.03532 (2020). 13

  25. [25]

    M., Díaz, M

    Davani, A. M., Díaz, M. & Prabhakaran, V. Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations. Preprint at https://arxiv.org/abs/2110.05719 (2021)

  26. [26]

    & Pierrehumbert, J

    Röttger, P., Vidgen, B., Hovy, D. & Pierrehumbert, J. B. Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. Preprint at https://arxiv.org/abs/2112.07475 (2021)

  27. [27]

    The ‘Problem’ of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

    Plank, B. The ‘Problem’ of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. Preprint at https://arxiv.org/abs/2211.02570 (2022)

  28. [28]

    & Rosen, R

    Denton, R., Díaz, M., Kivlichan, I., Prabhakaran, V. & Rosen, R. Whose Ground Truth? AccountingforIndividualandCollectiveIdentitiesUnderlyingDatasetAnnotation. Preprint at https://arxiv.org/abs/2112.04554 (2021)

  29. [29]

    Mokhberian, N. et al. Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks. Preprint at https://arxiv.org/abs/2311.09743 (2023)

  30. [30]

    & Iyyer, M

    Karpinska, M., Akoury, N. & Iyyer, M. The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation. Preprint at https://arxiv.org/abs/2109.06835 (2021)

  31. [31]

    Kirk, H. R. et al. The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models.The Thirty-eight Conference on Neural Infor- mation Processing Systems Datasets and Benchmarks Track (2024)(2024); preprint at https://arxiv.org/abs/2404.16019

  32. [32]

    & Bengio, Y

    Alain, G. & Bengio, Y. Understanding intermediate layers using linear classifier probes. Preprint at https://arxiv.org/abs/1610.01644 (2016)

  33. [33]

    & Liang, P

    Hewitt, J. & Liang, P. Designing and Interpreting Probes with Control Tasks.Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)(2019). doi:10.18653/v1/d19-1275

  34. [34]

    & Hovy, E

    Ravichander, A., Belinkov, Y. & Hovy, E. Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?. Preprint at https://arxiv.org/abs/2005.00719 (2020)

  35. [35]

    & Bansal, M

    Hase, P., Xie, H. & Bansal, M. The Out-of-Distribution Problem in Explain- ability and Search Methods for Feature Importance Explanations. Preprint at https://arxiv.org/abs/2106.00786 (2021)

  36. [36]

    Probing Classifiers: Promises, Shortcomings, and Advances

    Belinkov, Y. Probing Classifiers: Promises, Shortcomings, and Advances. Preprint at https://arxiv.org/abs/2102.12452 (2021)

  37. [37]

    & Goldberg, Y

    Elazar, Y., Ravfogel, S., Jacovi, A. & Goldberg, Y. Amnesic Probing: Behavioral Ex- planation with Amnesic Counterfactuals. Preprint at https://arxiv.org/abs/2006.00995 (2020)

  38. [38]

    & Cotterell, R

    Ravfogel, S., Twiton, M., Goldberg, Y. & Cotterell, R. Linear Adversarial Concept Erasure. Preprint at https://arxiv.org/abs/2201.12091 (2022)

  39. [39]

    & Steinhardt, J

    Burns, C., Ye, H., Klein, D. & Steinhardt, J. Discovering Latent Knowledge in Language Models Without Supervision. Preprint at https://arxiv.org/abs/2212.03827 (2022)

  40. [40]

    & Belrose, N

    Mallen, A., Brumley, M., Kharchenko, J. & Belrose, N. Eliciting Latent Knowledge from Quirky Language Models. Preprint at https://arxiv.org/abs/2312.01037 (2023)

  41. [41]

    & Tegmark, M

    Marks, S. & Tegmark, M. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. Preprint at https://arxiv.org/abs/2310.06824 (2023)

  42. [42]

    & Tegmark, M

    Gurnee, W. & Tegmark, M. Language Models Represent Space and Time. Preprint at https://arxiv.org/abs/2310.02207 (2023)

  43. [43]

    & Wattenberg, M

    Nanda, N., Lee, A. & Wattenberg, M. Emergent Linear Representations in World Models of Self-Supervised Sequence Models. Preprint at https://arxiv.org/abs/2309.00941 (2023)

  44. [44]

    Zou, A. et al. Representation Engineering: A Top-Down Approach to AI Transparency. Preprint at https://arxiv.org/abs/2310.01405 (2023)

  45. [45]

    & Wattenberg, M

    Li, K., Patel, O., Viégas, F., Pfister, H. & Wattenberg, M. Inference-Time In- 14 tervention: Eliciting Truthful Answers from a Language Model. Preprint at https://arxiv.org/abs/2306.03341 (2023)

  46. [46]

    & Bowman, S

    Turpin, M., Michael, J., Perez, E. & Bowman, S. R. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. Preprint at https://arxiv.org/abs/2305.04388 (2023)

  47. [47]

    Lanham, T. et al. Measuring Faithfulness in Chain-of-Thought Reasoning. Preprint at https://arxiv.org/abs/2307.13702 (2023)

  48. [48]

    Chen, Y. et al. Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations. Preprint at https://arxiv.org/abs/2307.08678 (2023)

  49. [49]

    Agarwal, C., Tanneru, S. H. & Lakkaraju, H. Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models. Preprint at https://arxiv.org/abs/2402.04614 (2024)

  50. [50]

    Ji, Z. et al. Survey of Hallucination in Natural Language Generation.ACM Computing Surveys (2022)(2022); preprint at https://arxiv.org/abs/2202.03629

  51. [51]

    M., Yu, Q

    Webson, A., Loo, A. M., Yu, Q. & Pavlick, E. Are Language Models Worse than Humans at Following Prompts? It’s Complicated. Preprint at https://arxiv.org/abs/2301.07085 (2023)

  52. [52]

    Panickssery, A., Bowman, S. R. & Feng, S. LLM Evaluators Recognize and Favor Their Own Generations. Preprint at https://arxiv.org/abs/2404.13076 (2024)

  53. [53]

    & Okazaki, N

    Oi, M., Kaneko, M., Koike, R., Loem, M. & Okazaki, N. Likelihood-based Mitigation of Evaluation Bias in Large Language Models.ACL2024 (findings)(2024); preprint at https://arxiv.org/abs/2402.15987

  54. [54]

    & Hashimoto, T

    Dubois, Y., Galambosi, B., Liang, P. & Hashimoto, T. B. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. Preprint at https://arxiv.org/abs/2404.04475 (2024)

  55. [55]

    & Hadfield-Menell, D

    Casper, S., Lin, J., Kwon, J., Culp, G. & Hadfield-Menell, D. Explore, Establish, Exploit: Red Teaming Language Models from Scratch. Preprint at https://arxiv.org/abs/2306.09442 (2023)

  56. [56]

    Ouyang, L. et al. Training language models to follow instructions with human feedback. Preprint at https://arxiv.org/abs/2203.02155 (2022)

  57. [57]

    Sharma, M. et al. Towards Understanding Sycophancy in Language Models. Preprint at https://arxiv.org/abs/2310.13548 (2023)

  58. [58]

    Wei, J., Huang, D., Lu, Y., Zhou, D. & Le, Q. V. Simple synthetic data reduces sycophancy in large language models. Preprint at https://arxiv.org/abs/2308.03958 (2023)

  59. [59]

    Qi, X. et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!. Preprint at https://arxiv.org/abs/2310.03693 (2023)

  60. [60]

    Ren, R. et al. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. Preprint at https://arxiv.org/abs/2407.21792 (2024)

  61. [61]

    Jiang, G. et al. Evaluating and Inducing Personality in Pre-trained Language Models. Preprint at https://arxiv.org/abs/2206.07550 (2022)

  62. [62]

    Jiang, H. et al. PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits. Preprint at https://arxiv.org/abs/2305.02547 (2023)

  63. [63]

    A Helpful Assistant

    Zheng, M., Pei, J., Logeswaran, L., Lee, M. & Jurgens, D. When “A Helpful Assistant” Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models. Preprint at https://arxiv.org/abs/2311.10054 (2023)

  64. [64]

    Zhou, X. et al. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents. Preprint at https://arxiv.org/abs/2310.11667 (2023)

  65. [65]

    & Andreas, J

    Liu, K., Casper, S., Hadfield-Menell, D. & Andreas, J. Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?. Preprint at https://arxiv.org/abs/2312.03729 (2023)

  66. [66]

    Li, K. et al. Measuring and Controlling Instruction (In)Stability in Language Model Dialogs. 15 Preprint at https://arxiv.org/abs/2402.10962 (2024)

  67. [67]

    Kocielnik, R. et al. Rethinking Psychometric Evaluation of LLMs: When and Why Self- Reports Predict Behavior. Preprint at https://arxiv.org/abs/2606.12730 (2026)

  68. [68]

    & Dean, J

    Hinton, G., Vinyals, O. & Dean, J. Distilling the Knowledge in a Neural Network. Preprint at https://arxiv.org/abs/1503.02531 (2015)

  69. [69]

    & Schütze, H

    Schick, T. & Schütze, H. It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners. Preprint at https://arxiv.org/abs/2009.07118 (2020)

  70. [70]

    Hsieh, C. et al. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.Findings of the Association for Computational Linguistics: ACL 2023(2023). doi:10.18653/v1/2023.findings-acl.507

  71. [71]

    Eldan, R. & Li, Y. TinyStories: How Small Can Language Models Be and Still Speak Coherent English?. Preprint at https://arxiv.org/abs/2305.07759 (2023)

  72. [72]

    Preprintathttps://arxiv.org/abs/2306.11644 (2023)

    Gunasekar, S.etal.TextbooksAreAllYouNeed. Preprintathttps://arxiv.org/abs/2306.11644 (2023)

  73. [73]

    Xu, X. et al. A Survey on Knowledge Distillation of Large Language Models. Preprint at https://arxiv.org/abs/2402.13116 (2024)

  74. [74]

    & Kiela, D

    Nickel, M. & Kiela, D. Poincaré Embeddings for Learning Hierarchical Representations. Preprint at https://arxiv.org/abs/1705.08039 (2017)

  75. [75]

    D., Gu, A., Ré, C

    Sa, C. D., Gu, A., Ré, C. & Sala, F. Representation Tradeoffs for Hyperbolic Embeddings. Preprint at https://arxiv.org/abs/1804.03329 (2018)

  76. [76]

    & Kiela, D

    Nickel, M. & Kiela, D. Learning Continuous Hierarchies in the Lorentz Model of Hyperbolic Geometry. Preprint at https://arxiv.org/abs/1806.03417 (2018)

  77. [77]

    & Hofmann, T

    Ganea, O., Bécigneul, G. & Hofmann, T. Hyperbolic Neural Networks. Preprint at https://arxiv.org/abs/1805.09112 (2018)

  78. [78]

    Chami, I. et al. Low-Dimensional Hyperbolic Knowledge Graph Embeddings. Preprint at https://arxiv.org/abs/2005.00545 (2020)

  79. [79]

    & Zhao, G

    Peng, W., Varanka, T., Mostafa, A., Shi, H. & Zhao, G. Hyperbolic Deep Neural Networks: A Survey. Preprint at https://arxiv.org/abs/2101.04562 (2021)

  80. [80]

    & Shankar, V

    Recht, B., Roelofs, R., Schmidt, L. & Shankar, V. Do ImageNet Classifiers Generalize to ImageNet?. Preprint at https://arxiv.org/abs/1902.10811 (2019)

Showing first 80 references.