Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:17.294000Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2501.00418.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:17.294000Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T15:23:15.199636Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
44 of 44 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 5c8e58f1-382a-48e8-a37b-39e514e23e41 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Deep learning with differential privacy
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf6d8b4-e7da-4200-af45-8fc4392de94a · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Types of Out-of-Distribution Texts and How to Detect Them
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0e8d9c-0793-4571-84e6-d83b7130f7fd · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Pythia: A suite for analyzing large language models across training and scaling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5db256fe-b703-400a-97f5-c90c9f1b2276 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e61272-cdf4-4e70-823d-b3b53a899bac · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Language models are realistic tabular data generators
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 598298da-00ab-43c7-89f5-e8b19dc89534 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Generating Sentences from a Continuous Space
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd0c2b0-c174-433c-9f08-5cd6a0801e6e · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c029ba4-6d01-4cfe-abc5-703bb9c29312 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1307bddd-955d-4ba4-8c00-544026da2576 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Extracting training data from large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9508fb-2ae6-4a58-b8c0-83135f298c20 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Retiring adult: New datasets for fair machine learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9f4d537-f9de-4555-95f8-77fe4e3e8bb7 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Calibrating noise to sensitivity in private data analysis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 654ac7ef-24dd-4ab6-924a-297dc1ea55fa · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models BAE: BERT-based adversarial examples for text classification
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b1219e-6e47-4f6c-a5af-5027b6fe9d9b · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Explaining and Harnessing Adversarial Examples
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf4754a-ae37-4c36-be28-9774298c1750 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Reducing sentiment bias in language models via counterfactual evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 944c1e52-49d2-41d5-9656-02dc8312debf · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Students parrot their teachers: Membership inference on model distillation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e93573f-77c8-4ee2-9961-923a8aea1940 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Is BERT really robust? A strong baseline for natural language attack on text classification and entailment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39542e87-0d8b-4bb7-a84f-082094c8201c · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models The enron corpus: A new dataset for email classification research
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9f46bc3b-ce1d-4d47-b89d-14a4c7096c61 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Certified robustness to adversarial examples with differential privacy
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6a57a0d9-d202-4e51-b7fb-0dd768117b05 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Gaussian membership inference privacy
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf75afa3-cb55-4c85-9c1b-76e55e8cc6c3 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Certified adversarial robustness with additive noise
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9296abc1-78f1-40a0-9663-b3b4e0e407e6 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models BERT-ATTACK: Adversarial attack against BERT using BERT
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74d4151-545e-451a-bf51-b6dc2f8347b9 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Focal Loss for Dense Object Detection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbadb186-30f5-4f49-8d96-d70ae8a0bd3e · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Towards Deep Learning Models Resistant to Adversarial Attacks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d5fbc65-ecf1-47a6-88d4-3ec3c324e019 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Towards deep learning models resistant to adversarial attacks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6cebab5c-3026-47af-936f-899534ed54f0 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Repeated knowledge distillation with confidence masking to mitigate membership inference attacks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd4adfea-f16a-46a0-a719-993bc5f431f0 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff2eb1f-94d6-4b69-a25e-747a0aca3286 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Language models are unsupervised multitask learners
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d320f86-b5d6-474a-abff-b480e52be49d · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde160cf-5a22-4a04-b497-a4af7dffca52 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Membership privacy for machine learning models through knowledge transfer
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 136abd5a-e8d0-448d-9c63-0776758c3d29 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Intriguing properties of neural networks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b50bbef-4638-454a-95e1-9126f9b9cd41 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Rethinking the inception architecture for computer vision
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a17418-ccb2-4e3e-9145-eb28230bb3fa · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Mitigating membership inference attacks by {Self-Distillation} through a novel ensemble architecture
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79c8cc56-cfa2-4566-87af-1c26d0599b26 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models T3: Tree-autoencoder constrained adversarial text generation for targeted attack
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b50b6b4-6224-4985-8d20-3b03d2c13926 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Adversarial GLUE: A multi-task benchmark for robustness evaluation of language mod- els
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95f7a2a4-031b-42ae-a276-03d4b2bf2b54 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b527d9bf-42fb-4362-b7fc-48c93c33d6df · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b9fdc43-8667-428d-b59f-596aafea6ff5 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Emergent Abilities of Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571b1777-95d2-47a5-81e0-8cf3004c8391 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca347c03-fe7f-45e7-8344-5967c4a4ee8c · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Fairness constraints: Mechanisms for fair classification
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7d72000f-3ebc-449d-83d3-2bcffab1ab4c · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Word-level textual adversarial attacking as combinatorial optimization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae24b9f-29e0-4a75-a1bc-93bc9519a2fb · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Gender bias in coreference resolution: Evaluation and debiasing methods
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77db318-aa73-46a4-944e-4f3fd6db689e · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Resisting membership inference attacks through knowledge distillation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bc2b9de-59b9-47d6-b537-9db6bc2b7ee4 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models FreeLB: Enhanced Adversarial Training for Natural Language Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4588d93-7f23-4a3b-8257-e29287254382 · outbound
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Unresolved cited work
Reference 495
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6afeab1-5c4c-4478-b064-1b062820a5b8 · inbound
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6272a041-0594-41f9-9313-7adfab482516 · inbound
On Weak-to-Strong Generalization and f-Divergence Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ba96de3-5d0a-4103-b537-81e561435348 · inbound
On the Blessing of Pre-training in Weak-to-Strong Generalization Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ddc2a02a-6551-4dd4-a954-018a6182b1af · inbound
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.