REVIEW 3 major objections 3 minor 93 references
Behavioural-faithfulness detection scores are under-specified: the same evaluators rank differently depending on whether the score targets exposure to a trait-inducing prompt or manifestation of the trait in the reply.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:37 UTC pith:OXATCTBW
load-bearing objection Under-specification of behavioural-detection AUROC is real: this paper shows it with a clean design, but the manifestation arm leans on an automatic judge's label; the core interaction still holds. the 3 major comments →
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central finding is that measurement target and output resolution are separable factors that jointly determine which behavioural-faithfulness evaluator appears better. On the identical 720 replies, the auditor's frozen-representation read-out achieves 0.804 AUROC against the exposure label versus the frontier judge's 0.718 under a single verdict, while against the manifestation label the judge leads 0.811 versus 0.690. The difference between the two gaps — the estimand interaction — is positive and excludes zero at all three tested output resolutions (0.207, 0.237, 0.169). When the judge is asked a target-specific question with a continuous confidence score, or when the auditor's
What carries the argument
The central objects are two labels on the same 720 replies — exposure (whether the reply was generated under a trait-inducing prompt) and manifestation (whether the trait actually surfaced, as judged by an independent frontier judge) — scored against two instruments at three output resolutions. The load-bearing quantity is the 'estimand interaction': the difference between the auditor-minus-judge AUROC gap on exposure and the same gap on manifestation. Because the interaction excludes zero at every resolution while the sign of the exposure gap flips with output format, it separates the effect of measurement target from the effect of answer interface. The hyperbolic geometry of the auditor's
Load-bearing premise
The manifestation label is assigned by a single automatic frontier judge, so the manifestation half of the comparison and the discordance result measure agreement with that judge, not human-verified behaviour; if that label does not track genuine manifestation, the reversal and the interaction size change meaning.
What would settle it
A prespecified, human-adjudicated manifestation subset, blind to condition and generator, would settle the central claim: if the human-verified labels do not reproduce the roughly 0.2 AUROC estimand interaction and the judge's lead on manifestation, or if the ranking under the verdict interface does not reverse, the paper's conclusion that a single AUROC is under-specified would need revision. A second decisive check would be a supervised read-out on the judge's own representations, were they accessible, to test whether the distances are an artefact of supervised-versus-unsupervised comparison
If this is right
- A reported behavioural-detection AUROC is under-specified: it must be accompanied by the target label (exposure or manifestation) and the evaluator's output interface (verdict or continuous score) to be comparable.
- Claims that a small specialised model beats a frontier judge, or vice versa, are not well-formed until the estimand is named; under the deployed verdict interface the auditor leads on exposure (0.804 vs 0.718) and trails on manifestation (0.690 vs 0.811).
- Matching output resolution — asking the judge a target-specific question answered with a continuous confidence score, or thresholding the auditor's read-out at 0.5 — is necessary to attribute differences to the instrument rather than the answer format; under these matched comparisons the judge leads on manifestation with all intervals excluding zero, and the instruments are statistically indisting
- The auditor and the judge misclassify different replies (McNemar p ≈ 1.07×10⁻²⁰), and combined two-instrument rules reach operating points neither alone reaches — a high-recall screen and a high-precision confirmer — though these are provisional because the manifestation label comes from a single automatic judge.
- The transferable detection signal lies in the purpose-built frozen representation and its supervision, not in hyperbolic curvature: a Euclidean-pooled read-out beats the hyperbolic Lorentz read-out on the same representation, and the auditor's pre-trained heads do not transfer zero-shot to this task.
Where Pith is reading between the lines
- If the exposure-manifestation distinction holds for human-verified labels and other behaviours, the same two-target decomposition should be applied to safety benchmarks where an inducing prompt is used as a proxy label (e.g., 'attack succeeded' versus 'harmful content appeared'); single-number comparisons in those settings are likely to mislead in a similar way.
- The paper could not run a supervised read-out on the judge's own representations because they are inaccessible; a testable extension would be to fit a comparable read-out on the judge's hidden states if they become available and check whether the roughly 0.2 AUROC target-driven gap persists — if it does not, the distances are partly an artefact of the supervised-versus-unsupervised comparison.
- The discordance result suggests a cheap resident screen: a compact auditor combined with a frontier judge could widen coverage at low cost, and the exact recall-precision trade-off is worth measuring against human-verified manifestation labels in a deployed audit pipeline.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper distinguishes two measurement targets for behavioural-faithfulness detection — exposure (was the reply produced under a trait-inducing condition) and manifestation (did the trait actually surface) — and argues that a single AUROC is under-specified because it conflates the target and the evaluator's output interface. Using a fixed 146M-parameter auditor with a frozen-representation logistic read-out and a frontier zero-shot judge on the same 720 paired replies, with leave-one-generator-out evaluation, it reports that the auditor-vs-judge gap changes by roughly 0.2 AUROC between targets at all three output resolutions, and that under the judge's deployed verdict interface the ranking reverses (auditor leads on exposure, judge on manifestation), while matching output resolution removes the reliable reversal but not the interaction. The paper also reports a discordance analysis, a geometry boundary result, and a human-rating baseline. The design is transparent: single fixed checkpoint, cluster bootstrap on pairs, matched controls approached from both directions, and all table values claimed regenerable from archived per-item files.
Significance. If the result holds, it is a useful measurement contribution: the field's single behavioural-detection AUROCs are indeed ambiguous unless the estimand and output interface are stated together. The paper's strengths are its explicit separation of exposure and manifestation, the objective exposure label, the paired construction, cluster bootstrap inference, leave-one-generator-out evaluation, the two matched-resolution controls, and unusually complete reproducibility (single fixed checkpoint, archived data and scripts, validation checks). The paper is also candid about its limitations. The central reservation is that the manifestation arm is scored against an automatic judge's label, so the headline reversal and the interaction are provisional until human-verified manifestation labels are supplied; the paper itself identifies this as the decisive next step.
major comments (3)
- [§3.2 'Two cautions', §5 Limitations, Table 2] The manifestation label is the output of a single automatic frontier judge, and the judge is then scored against that label. The judge's manifestation lead (0.811 vs 0.690; 0.958 vs 0.690 in the matched control) and the interaction intervals (0.207, 0.237, 0.169) are agreement with that automatic reference, not verified manifestation. The paper admits this ('in part agreement between two frontier judges') and calls a human-verified subset 'the decisive next step', yet the abstract and Section 4 report the reversal and interaction as the main result. If labeler and evaluator share model family or rubric, the headline reversal could be labeler–evaluator self-consistency. This is load-bearing: the target-effect claim rests on the manifestation arm. Add a human-adjudicated subset (even a few hundred replies) and re-estimate, or explicitly reframe all manifestation claims as relative to an au
- [§3.2, §6 Methods] The 'frontier judge' and the 'independent frontier judge' who assigned the manifestation labels are never named. Methods describe the label as assigned by 'an independent frontier judge under a fixed trait-specific rubric' and the baseline as 'a frontier zero-shot judge', but no model identifiers are given. A reader cannot tell whether the evaluator and labeler are the same model or family, or what 'independent' means here. Please name both models; if they are from the same family, treat the manifestation lead as partly self-consistency. Ideally use a labeler and evaluator from different families, or report their cross-family agreement.
- [§3.2, §5 Limitations] The auditor's logistic read-out is trained on the target label for the in-fold generators, while the judge is zero-shot. The paper acknowledges this asymmetry in §5, but the abstract and Section 4 present the auditor-vs-judge distances and the interaction as properties of the instruments. The reported ~0.2 AUROC target effect therefore also encodes a supervised-vs-unsupervised contrast. Either run the proposed supervised read-out on the judge's representations, or state in the abstract that the comparison is between a target-supervised read-out and a zero-shot judge, not between two deployment-ready instruments.
minor comments (3)
- [§3.2/Table 2] The text says the binary comparison 'removes the reversal', but the point estimates still reverse: auditor 0.736 vs judge 0.718 on exposure, judge 0.811 vs auditor 0.660 on manifestation. The reversal is removed only in the statistical sense (the exposure CI [-0.017, 0.051] includes zero). Please qualify the wording.
- [§6 Methods] The 0.5 threshold is called 'the class-balanced midpoint', but the manifestation label has 61.4% positives, so 0.5 is not class-balanced for the manifestation target. Use class-balanced thresholds per target or clarify the fixed-threshold rationale.
- [§3.5] The human evaluation and the trait-detection evaluation are not co-registered, as the paper notes, but the 'text-layer blind spot' is used as a baseline for a different item set. A sentence on the inferential distance between the two studies would help readers avoid treating the human result as a control for the 720-reply comparison.
Circularity Check
No significant circularity; the central interaction is empirical, with a disclosed and non-circular limitation on the manifestation reference.
full rationale
The central derivation—exposure versus manifestation targets changing the auditor–judge gap—rests on two independently constructed labels and on actual AUROC computations over the same 720 replies. The exposure label is objective by construction, the logistic read-out is refit under leave-one-generator-out evaluation, and the paired cluster-bootstrap intervals are not fitted to the headline values. The manifestation label is admittedly the output of a single automatic judge, and the paper explicitly states that the manifestation metric is agreement with that judge rather than human-verified behavior and that the judge's advantage is in part agreement between two frontier judges. This is a measurement-validity limitation, not a circular derivation: the reported AUROCs are not equal to the label or to any fitted parameter by construction. The main self-citation (ref. 93, the earlier version of this same work) supports only the text-layer human-rating baseline, not the load-bearing target/interface interaction, and the shared human-rating asset is disclosed. Thus no specific circular reduction is exhibited; the paper is self-contained on its central claim.
Axiom & Free-Parameter Ledger
free parameters (2)
- Auditor logistic read-out weights =
unknown (refit per leave-one-generator-out fold, separately for exposure and manifestation labels)
- Binary-comparison threshold 0.5 =
0.5
axioms (5)
- domain assumption A reply generated under a trait-inducing companion prompt is a positive exposure case even if the trait did not appear in the reply.
- domain assumption The manifestation label assigned by a single automatic frontier judge under a fixed rubric is a valid reference for whether the behaviour surfaced.
- domain assumption Style-controlled trait/neutral pairs isolate the trait condition from style and context.
- domain assumption Leave-one-generator-out over three generators from two providers provides evidence about transfer of the exposure signal.
- standard math Standard bootstrap and McNemar procedures give valid paired uncertainty statements.
read the original abstract
Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation). Scoring a compact 146-million-parameter auditor's frozen-representation read-out and a frontier judge against each label on the identical 720 replies, the gap between the instruments moves by roughly 0.2 AUROC when the target changes. Under the judge's deployed interface, a single verdict, the ranking reverses: the auditor leads on exposure, 0.804 against 0.718, and trails on manifestation, 0.690 against 0.811. Matching the output resolution from either direction, by asking the judge a target-specific question answered with a continuous confidence score or by thresholding the auditor's read-out, removes the reversal but not the interaction, which excludes zero at all three resolutions (0.207, 0.237 and 0.169). The target governs how far apart the instruments are; the interface governs whether that distance changes their order. The auditor's hyperbolic geometry confers no advantage here. A single behavioural-detection AUROC is under-specified: such claims are comparable only when they state the estimand, the evaluator, and its output interface.
Figures
Reference graph
Works this paper leans on
-
[1]
Perez, E. et al. Discovering Language Model Behaviors with Model-Written Evaluations. Preprint at https://arxiv.org/abs/2212.09251 (2022)
Pith/arXiv arXiv 2022
-
[2]
Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Preprint at https://arxiv.org/abs/2306.05685 (2023)
Pith/arXiv arXiv 2023
-
[3]
Jacobs, A. Z. & Wallach, H. Measurement and Fairness.Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency(2021). doi:10.1145/3442188.3445901
arXiv 2021
-
[4]
Raji, I. D., Bender, E. M., Paullada, A., Denton, E. & Hanna, A. AI and the Everything in the Whole Wide World Benchmark. Preprint at https://arxiv.org/abs/2111.15366 (2021)
Pith/arXiv arXiv 2021
-
[5]
Dehghani, M. et al. The Benchmark Lottery. Preprint at https://arxiv.org/abs/2107.07002 (2021)
Pith/arXiv arXiv 2021
-
[6]
Bean, A. M. et al. Measuring what Matters: Construct Validity in Large Language Model Benchmarks. Preprint at https://arxiv.org/abs/2511.04703 (2025)
arXiv 2025
-
[7]
Singh, S. et al. The Leaderboard Illusion. Preprint at https://arxiv.org/abs/2504.20879 (2025)
Pith/arXiv arXiv 2025
-
[8]
A., Constantinides, M., Tahaei, M
Septiandri, A. A., Constantinides, M., Tahaei, M. & Quercia, D. WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT?.2023 ACM Conference on Fairness Accountability and Transparency(2023). doi:10.1145/3593013.3593985
arXiv 2023
-
[9]
Buyl, M. et al. Large Language Models Reflect the Ideology of their Creators. Preprint at https://arxiv.org/abs/2410.18417 (2024)
Pith/arXiv arXiv 2024
-
[10]
Liu, N. F. et al. Lost in the Middle: How Language Models Use Long Contexts. Preprint at https://arxiv.org/abs/2307.03172 (2023)
Pith/arXiv arXiv 2023
-
[11]
Freiesleben, T. & Zezulka, S. The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models. Preprint at https://arxiv.org/abs/2510.23191 (2025)
arXiv 2025
-
[12]
Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
Freiesleben, T. Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks. Preprint at https://arxiv.org/abs/2603.15121 (2026)
arXiv 2026
-
[13]
Wallach, H. et al. Position: Evaluating Generative AI Systems Is a Social Science Measure- ment Challenge. Preprint at https://arxiv.org/abs/2502.00561 (2025)
Pith/arXiv arXiv 2025
-
[14]
Salaudeen, O. et al. Measurement to Meaning: A Validity-Centered Framework for AI Evaluation. Preprint at https://arxiv.org/abs/2505.10573 (2025)
Pith/arXiv arXiv 2025
-
[15]
Weidinger, L. et al. Toward an Evaluation Science for Generative AI Systems. Preprint at https://arxiv.org/abs/2503.05336 (2025)
Pith/arXiv arXiv 2025
-
[16]
Solaiman, I. et al. Evaluating the Social Impact of Generative AI Systems in Systems and Society. Preprint at https://arxiv.org/abs/2306.05949 (2023)
Pith/arXiv arXiv 2023
-
[17]
Paruchuri, A. et al. What Are the Odds? Language Models Are Capable of Probabilistic Reasoning. Preprint at https://arxiv.org/abs/2406.12830 (2024)
Pith/arXiv arXiv 2024
-
[18]
Artstein, R. & Poesio, M. Inter-Coder Agreement for Computational Linguistics.Compu- tational Linguistics(2008). doi:10.1162/coli.07-034-r2
-
[19]
McNemar, Q. Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages.Psychometrika(1947). doi:10.1007/bf02295996
-
[20]
A Coefficient of Agreement for Nominal Scales.Educational and Psychological Measurement(1960)
Cohen, J. A Coefficient of Agreement for Nominal Scales.Educational and Psychological Measurement(1960). doi:10.1177/001316446002000104
-
[21]
Fleiss, J. L. Measuring nominal scale agreement among many raters.Psychological Bulletin (1971). doi:10.1037/h0031619
doi:10.1037/h0031619 1971
-
[22]
Landis, J. R. & Koch, G. G. The Measurement of Observer Agreement for Categorical Data.Biometrics(1977). doi:10.2307/2529310
doi:10.2307/2529310 1977
-
[23]
Pavlick, E. & Kwiatkowski, T. Inherent Disagreements in Human Textual Inferences.Trans- actions of the Association for Computational Linguistics(2019). doi:10.1162/tacl_a_00293
-
[24]
Nie, Y., Zhou, X. & Bansal, M. What Can We Learn from Collective Human Opinions on Natural Language Inference Data?. Preprint at https://arxiv.org/abs/2010.03532 (2020). 13
Pith/arXiv arXiv 2010
-
[25]
Davani, A. M., Díaz, M. & Prabhakaran, V. Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations. Preprint at https://arxiv.org/abs/2110.05719 (2021)
Pith/arXiv arXiv 2021
-
[26]
Röttger, P., Vidgen, B., Hovy, D. & Pierrehumbert, J. B. Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. Preprint at https://arxiv.org/abs/2112.07475 (2021)
Pith/arXiv arXiv 2021
-
[27]
The ‘Problem’ of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation
Plank, B. The ‘Problem’ of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. Preprint at https://arxiv.org/abs/2211.02570 (2022)
Pith/arXiv arXiv 2022
-
[28]
Denton, R., Díaz, M., Kivlichan, I., Prabhakaran, V. & Rosen, R. Whose Ground Truth? AccountingforIndividualandCollectiveIdentitiesUnderlyingDatasetAnnotation. Preprint at https://arxiv.org/abs/2112.04554 (2021)
Pith/arXiv arXiv 2021
-
[29]
Mokhberian, N. et al. Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks. Preprint at https://arxiv.org/abs/2311.09743 (2023)
Pith/arXiv arXiv 2023
-
[30]
Karpinska, M., Akoury, N. & Iyyer, M. The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation. Preprint at https://arxiv.org/abs/2109.06835 (2021)
Pith/arXiv arXiv 2021
-
[31]
Kirk, H. R. et al. The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models.The Thirty-eight Conference on Neural Infor- mation Processing Systems Datasets and Benchmarks Track (2024)(2024); preprint at https://arxiv.org/abs/2404.16019
Pith/arXiv arXiv 2024
-
[32]
Alain, G. & Bengio, Y. Understanding intermediate layers using linear classifier probes. Preprint at https://arxiv.org/abs/1610.01644 (2016)
Pith/arXiv arXiv 2016
-
[33]
Hewitt, J. & Liang, P. Designing and Interpreting Probes with Control Tasks.Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)(2019). doi:10.18653/v1/d19-1275
-
[34]
Ravichander, A., Belinkov, Y. & Hovy, E. Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?. Preprint at https://arxiv.org/abs/2005.00719 (2020)
Pith/arXiv arXiv 2005
-
[35]
Hase, P., Xie, H. & Bansal, M. The Out-of-Distribution Problem in Explain- ability and Search Methods for Feature Importance Explanations. Preprint at https://arxiv.org/abs/2106.00786 (2021)
Pith/arXiv arXiv 2021
-
[36]
Probing Classifiers: Promises, Shortcomings, and Advances
Belinkov, Y. Probing Classifiers: Promises, Shortcomings, and Advances. Preprint at https://arxiv.org/abs/2102.12452 (2021)
Pith/arXiv arXiv 2021
-
[37]
Elazar, Y., Ravfogel, S., Jacovi, A. & Goldberg, Y. Amnesic Probing: Behavioral Ex- planation with Amnesic Counterfactuals. Preprint at https://arxiv.org/abs/2006.00995 (2020)
Pith/arXiv arXiv 2006
-
[38]
Ravfogel, S., Twiton, M., Goldberg, Y. & Cotterell, R. Linear Adversarial Concept Erasure. Preprint at https://arxiv.org/abs/2201.12091 (2022)
Pith/arXiv arXiv 2022
-
[39]
Burns, C., Ye, H., Klein, D. & Steinhardt, J. Discovering Latent Knowledge in Language Models Without Supervision. Preprint at https://arxiv.org/abs/2212.03827 (2022)
Pith/arXiv arXiv 2022
-
[40]
Mallen, A., Brumley, M., Kharchenko, J. & Belrose, N. Eliciting Latent Knowledge from Quirky Language Models. Preprint at https://arxiv.org/abs/2312.01037 (2023)
Pith/arXiv arXiv 2023
-
[41]
Marks, S. & Tegmark, M. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. Preprint at https://arxiv.org/abs/2310.06824 (2023)
Pith/arXiv arXiv 2023
-
[42]
Gurnee, W. & Tegmark, M. Language Models Represent Space and Time. Preprint at https://arxiv.org/abs/2310.02207 (2023)
Pith/arXiv arXiv 2023
-
[43]
Nanda, N., Lee, A. & Wattenberg, M. Emergent Linear Representations in World Models of Self-Supervised Sequence Models. Preprint at https://arxiv.org/abs/2309.00941 (2023)
Pith/arXiv arXiv 2023
-
[44]
Zou, A. et al. Representation Engineering: A Top-Down Approach to AI Transparency. Preprint at https://arxiv.org/abs/2310.01405 (2023)
Pith/arXiv arXiv 2023
-
[45]
Li, K., Patel, O., Viégas, F., Pfister, H. & Wattenberg, M. Inference-Time In- 14 tervention: Eliciting Truthful Answers from a Language Model. Preprint at https://arxiv.org/abs/2306.03341 (2023)
Pith/arXiv arXiv 2023
-
[46]
Turpin, M., Michael, J., Perez, E. & Bowman, S. R. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. Preprint at https://arxiv.org/abs/2305.04388 (2023)
Pith/arXiv arXiv 2023
-
[47]
Lanham, T. et al. Measuring Faithfulness in Chain-of-Thought Reasoning. Preprint at https://arxiv.org/abs/2307.13702 (2023)
Pith/arXiv arXiv 2023
-
[48]
Chen, Y. et al. Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations. Preprint at https://arxiv.org/abs/2307.08678 (2023)
Pith/arXiv arXiv 2023
-
[49]
Agarwal, C., Tanneru, S. H. & Lakkaraju, H. Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models. Preprint at https://arxiv.org/abs/2402.04614 (2024)
Pith/arXiv arXiv 2024
-
[50]
Ji, Z. et al. Survey of Hallucination in Natural Language Generation.ACM Computing Surveys (2022)(2022); preprint at https://arxiv.org/abs/2202.03629
Pith/arXiv arXiv 2022
-
[51]
Webson, A., Loo, A. M., Yu, Q. & Pavlick, E. Are Language Models Worse than Humans at Following Prompts? It’s Complicated. Preprint at https://arxiv.org/abs/2301.07085 (2023)
Pith/arXiv arXiv 2023
-
[52]
Panickssery, A., Bowman, S. R. & Feng, S. LLM Evaluators Recognize and Favor Their Own Generations. Preprint at https://arxiv.org/abs/2404.13076 (2024)
Pith/arXiv arXiv 2024
-
[53]
Oi, M., Kaneko, M., Koike, R., Loem, M. & Okazaki, N. Likelihood-based Mitigation of Evaluation Bias in Large Language Models.ACL2024 (findings)(2024); preprint at https://arxiv.org/abs/2402.15987
arXiv 2024
-
[54]
Dubois, Y., Galambosi, B., Liang, P. & Hashimoto, T. B. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. Preprint at https://arxiv.org/abs/2404.04475 (2024)
Pith/arXiv arXiv 2024
-
[55]
Casper, S., Lin, J., Kwon, J., Culp, G. & Hadfield-Menell, D. Explore, Establish, Exploit: Red Teaming Language Models from Scratch. Preprint at https://arxiv.org/abs/2306.09442 (2023)
Pith/arXiv arXiv 2023
-
[56]
Ouyang, L. et al. Training language models to follow instructions with human feedback. Preprint at https://arxiv.org/abs/2203.02155 (2022)
Pith/arXiv arXiv 2022
-
[57]
Sharma, M. et al. Towards Understanding Sycophancy in Language Models. Preprint at https://arxiv.org/abs/2310.13548 (2023)
Pith/arXiv arXiv 2023
-
[58]
Wei, J., Huang, D., Lu, Y., Zhou, D. & Le, Q. V. Simple synthetic data reduces sycophancy in large language models. Preprint at https://arxiv.org/abs/2308.03958 (2023)
Pith/arXiv arXiv 2023
-
[59]
Qi, X. et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!. Preprint at https://arxiv.org/abs/2310.03693 (2023)
Pith/arXiv arXiv 2023
-
[60]
Ren, R. et al. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?. Preprint at https://arxiv.org/abs/2407.21792 (2024)
Pith/arXiv arXiv 2024
-
[61]
Jiang, G. et al. Evaluating and Inducing Personality in Pre-trained Language Models. Preprint at https://arxiv.org/abs/2206.07550 (2022)
Pith/arXiv arXiv 2022
-
[62]
Jiang, H. et al. PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits. Preprint at https://arxiv.org/abs/2305.02547 (2023)
Pith/arXiv arXiv 2023
-
[63]
Zheng, M., Pei, J., Logeswaran, L., Lee, M. & Jurgens, D. When “A Helpful Assistant” Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models. Preprint at https://arxiv.org/abs/2311.10054 (2023)
Pith/arXiv arXiv 2023
-
[64]
Zhou, X. et al. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents. Preprint at https://arxiv.org/abs/2310.11667 (2023)
Pith/arXiv arXiv 2023
-
[65]
Liu, K., Casper, S., Hadfield-Menell, D. & Andreas, J. Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?. Preprint at https://arxiv.org/abs/2312.03729 (2023)
Pith/arXiv arXiv 2023
-
[66]
Li, K. et al. Measuring and Controlling Instruction (In)Stability in Language Model Dialogs. 15 Preprint at https://arxiv.org/abs/2402.10962 (2024)
Pith/arXiv arXiv 2024
-
[67]
Kocielnik, R. et al. Rethinking Psychometric Evaluation of LLMs: When and Why Self- Reports Predict Behavior. Preprint at https://arxiv.org/abs/2606.12730 (2026)
Pith/arXiv arXiv 2026
-
[68]
Hinton, G., Vinyals, O. & Dean, J. Distilling the Knowledge in a Neural Network. Preprint at https://arxiv.org/abs/1503.02531 (2015)
Pith/arXiv arXiv 2015
-
[69]
Schick, T. & Schütze, H. It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners. Preprint at https://arxiv.org/abs/2009.07118 (2020)
Pith/arXiv arXiv 2009
-
[70]
Hsieh, C. et al. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.Findings of the Association for Computational Linguistics: ACL 2023(2023). doi:10.18653/v1/2023.findings-acl.507
-
[71]
Eldan, R. & Li, Y. TinyStories: How Small Can Language Models Be and Still Speak Coherent English?. Preprint at https://arxiv.org/abs/2305.07759 (2023)
Pith/arXiv arXiv 2023
-
[72]
Preprintathttps://arxiv.org/abs/2306.11644 (2023)
Gunasekar, S.etal.TextbooksAreAllYouNeed. Preprintathttps://arxiv.org/abs/2306.11644 (2023)
Pith/arXiv arXiv 2023
-
[73]
Xu, X. et al. A Survey on Knowledge Distillation of Large Language Models. Preprint at https://arxiv.org/abs/2402.13116 (2024)
Pith/arXiv arXiv 2024
-
[74]
Nickel, M. & Kiela, D. Poincaré Embeddings for Learning Hierarchical Representations. Preprint at https://arxiv.org/abs/1705.08039 (2017)
Pith/arXiv arXiv 2017
-
[75]
Sa, C. D., Gu, A., Ré, C. & Sala, F. Representation Tradeoffs for Hyperbolic Embeddings. Preprint at https://arxiv.org/abs/1804.03329 (2018)
Pith/arXiv arXiv 2018
-
[76]
Nickel, M. & Kiela, D. Learning Continuous Hierarchies in the Lorentz Model of Hyperbolic Geometry. Preprint at https://arxiv.org/abs/1806.03417 (2018)
Pith/arXiv arXiv 2018
-
[77]
Ganea, O., Bécigneul, G. & Hofmann, T. Hyperbolic Neural Networks. Preprint at https://arxiv.org/abs/1805.09112 (2018)
Pith/arXiv arXiv 2018
-
[78]
Chami, I. et al. Low-Dimensional Hyperbolic Knowledge Graph Embeddings. Preprint at https://arxiv.org/abs/2005.00545 (2020)
Pith/arXiv arXiv 2005
-
[79]
Peng, W., Varanka, T., Mostafa, A., Shi, H. & Zhao, G. Hyperbolic Deep Neural Networks: A Survey. Preprint at https://arxiv.org/abs/2101.04562 (2021)
Pith/arXiv arXiv 2021
-
[80]
Recht, B., Roelofs, R., Schmidt, L. & Shankar, V. Do ImageNet Classifiers Generalize to ImageNet?. Preprint at https://arxiv.org/abs/1902.10811 (2019)
Pith/arXiv arXiv 1902
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.