Pith. sign in

Paper Citation Record · LEDGER

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

As of 11 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2501.00418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00418 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:17.294000Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:23:15.199636Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5c8e58f1-382a-48e8-a37b-39e514e23e41 · outbound

This paper cites Deep learning with differential privacy.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Deep learning with differential privacy

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.095613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.095613Z digest=sha256:1bb0f7baf2351bc2d807d4a30efece1e2e918a81d23252508d44f101162359a6

Observation 4cf6d8b4-e7da-4200-af45-8fc4392de94a · outbound

This paper cites Types of Out-of-Distribution Texts and How to Detect Them.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Types of Out-of-Distribution Texts and How to Detect Them

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.100416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.100416Z digest=sha256:a0a6329b5804e5d008112093d52e73110de89c7234bf9d1b4d909d43e6cbc29e

Observation ba0e8d9c-0793-4571-84e6-d83b7130f7fd · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.829745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.105045Z digest=sha256:802dae6e1883a294607e1c3e34c71ca53b8c87ed37b9e53d4e149356dddb349c

Observation 5db256fe-b703-400a-97f5-c90c9f1b2276 · outbound

This paper cites Man is to computer programmer as woman is to homemaker? debiasing word embeddings.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Man is to computer programmer as woman is to homemaker? debiasing word embeddings

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.110022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.110022Z digest=sha256:59335f8f14a9034f331b93c5a40b3f97a006739c223f3a793533cba16b13cd08

Observation c0e61272-cdf4-4e70-823d-b3b53a899bac · outbound

This paper cites Language models are realistic tabular data generators.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Language models are realistic tabular data generators

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.806767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.114332Z digest=sha256:06275c37c48711ca3337ecaac2c576af3c9dfff48b05fd6feda316dde4fee991

Observation 598298da-00ab-43c7-89f5-e8b19dc89534 · outbound

This paper cites Generating Sentences from a Continuous Space.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Generating Sentences from a Continuous Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.119094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.119094Z digest=sha256:76f00ac52264be3e09bab44810ed6246b6a934b4b1764ef90d7cb4b7d9721a21

Observation 1bd0c2b0-c174-433c-9f08-5cd6a0801e6e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.123720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.123720Z digest=sha256:5f51b638f3a41a373bc9d75582ab4de16a28ebaec7500757afb5078e5a1f6c0a

Observation 5c029ba4-6d01-4cfe-abc5-703bb9c29312 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Weak-to-strong generalization: Eliciting strong capabilities with weak supervision

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.792252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.127839Z digest=sha256:682bc97a115a409bf21888e968347db8330f3597c811750e7a555b0254565fd9

Observation 1307bddd-955d-4ba4-8c00-544026da2576 · outbound

This paper cites Extracting training data from large language models.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Extracting training data from large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.132106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.132106Z digest=sha256:be8d848a05d0493ffe987e7d5003ccacb27f6407280aedd6b738ac2b2601b85b

Observation 8e9508fb-2ae6-4a58-b8c0-83135f298c20 · outbound

This paper cites Retiring adult: New datasets for fair machine learning.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Retiring adult: New datasets for fair machine learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.764110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.136975Z digest=sha256:b64c87e2d0072fe6e6c7d4efca6f1310d92e86cfd8b27814a8395aa6debf47d6

Observation b9f4d537-f9de-4555-95f8-77fe4e3e8bb7 · outbound

This paper cites Calibrating noise to sensitivity in private data analysis.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Calibrating noise to sensitivity in private data analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.141214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.141214Z digest=sha256:3596e3f3a8d17d2f2647d4c1ad7ae8cce7476c8962a150d658a0602c728314e1

Observation 654ac7ef-24dd-4ab6-924a-297dc1ea55fa · outbound

This paper cites BAE: BERT-based adversarial examples for text classification.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models BAE: BERT-based adversarial examples for text classification

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.147521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.147521Z digest=sha256:d02fa96a0e12a45c68c5690fa1de12bcca9fa981b2c023ebecf15e477664577f

Observation d1b1219e-6e47-4f6c-a5af-5027b6fe9d9b · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Explaining and Harnessing Adversarial Examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.152544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.152544Z digest=sha256:d0703da134e2efa44fdb61a3fbe3e9b7b6e437f688db1a123c74d624a2db5be1

Observation dbf4754a-ae37-4c36-be28-9774298c1750 · outbound

This paper cites Reducing sentiment bias in language models via counterfactual evaluation.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Reducing sentiment bias in language models via counterfactual evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.157126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.157126Z digest=sha256:bc48a2d70ba68def2e683945f70b12fdfb0d1f9835004f9db30e16cde1bb121b

Observation 944c1e52-49d2-41d5-9656-02dc8312debf · outbound

This paper cites Students parrot their teachers: Membership inference on model distillation.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Students parrot their teachers: Membership inference on model distillation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.740501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.161816Z digest=sha256:f20550636121fd63960de9472aa5556ad4da5b24368533c81d69c37e1a6e9f67

Observation 2e93573f-77c8-4ee2-9961-923a8aea1940 · outbound

This paper cites Is BERT really robust? A strong baseline for natural language attack on text classification and entailment.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Is BERT really robust? A strong baseline for natural language attack on text classification and entailment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.166633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.166633Z digest=sha256:fd772b5556e133d899f252e1a69db34b1a3ee17143065118f40c71baf8fb5731

Observation 39542e87-0d8b-4bb7-a84f-082094c8201c · outbound

This paper cites The enron corpus: A new dataset for email classification research.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models The enron corpus: A new dataset for email classification research

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.720294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.172051Z digest=sha256:44250a2022ab634610d2acfe61f659696be2d123833e9240e62107f0c0e8000f

Observation 9f46bc3b-ce1d-4d47-b89d-14a4c7096c61 · outbound

This paper cites Certified robustness to adversarial examples with differential privacy.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Certified robustness to adversarial examples with differential privacy

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.706453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.177831Z digest=sha256:275679c72d9141b14202387c1d89515f68d5b89a8adc14864564b0f4889c83c9

Observation 6a57a0d9-d202-4e51-b7fb-0dd768117b05 · outbound

This paper cites Gaussian membership inference privacy.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Gaussian membership inference privacy

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.693337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.183166Z digest=sha256:fac8a40b7496dae1ba645e8772f4dda9eb6923677e87fe674dc458c905902b59

Observation cf75afa3-cb55-4c85-9c1b-76e55e8cc6c3 · outbound

This paper cites Certified adversarial robustness with additive noise.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Certified adversarial robustness with additive noise

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.678935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.187765Z digest=sha256:da57e2191cf30e9f1a986d89df53952f6c314f8eb3e96290c65aa9ff4f1fb6e0

Observation 9296abc1-78f1-40a0-9663-b3b4e0e407e6 · outbound

This paper cites BERT-ATTACK: Adversarial attack against BERT using BERT.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models BERT-ATTACK: Adversarial attack against BERT using BERT

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.192121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.192121Z digest=sha256:ea24907dd11a1cb092b48d1c12e082cf2803da6148da6b8127d37a36493c53c6

Observation e74d4151-545e-451a-bf51-b6dc2f8347b9 · outbound

This paper cites Focal Loss for Dense Object Detection.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Focal Loss for Dense Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.197196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.197196Z digest=sha256:03b2d99abc826bb72cbbf9f18e0c8ea6280945eaf61a7c6122350da98d7f4e09

Observation cbadb186-30f5-4f49-8d96-d70ae8a0bd3e · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.201696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.201696Z digest=sha256:783800e23269b24a011e0856c6540c2d7242c46e88621d45405f80664e1aca5b

Observation 1d5fbc65-ecf1-47a6-88d4-3ec3c324e019 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Towards deep learning models resistant to adversarial attacks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.666263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.205854Z digest=sha256:0bd3ec5e3408b59531d18282556be374adfbc460ba2d338baaec320e815ade2c

Observation 6cebab5c-3026-47af-936f-899534ed54f0 · outbound

This paper cites Repeated knowledge distillation with confidence masking to mitigate membership inference attacks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Repeated knowledge distillation with confidence masking to mitigate membership inference attacks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.652549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.209968Z digest=sha256:12496025826f8f419ca72418fb0fd8bb67e3d207beecd2d51bd1adf881ea7528

Observation cd4adfea-f16a-46a0-a719-993bc5f431f0 · outbound

This paper cites Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.214013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.214013Z digest=sha256:908a78746a2380c600079ec5407f08d0234e4e91369f31e8e2c0f310ccdc7709

Observation 1ff2eb1f-94d6-4b69-a25e-747a0aca3286 · outbound

This paper cites Language models are unsupervised multitask learners.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Language models are unsupervised multitask learners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.219018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.219018Z digest=sha256:0e52e73c6454c5050ab06f5757df1ca34e2340bd9d17e4c7d334152695bd02d9

Observation 2d320f86-b5d6-474a-abff-b480e52be49d · outbound

This paper cites Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36, 2024.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.223166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.223166Z digest=sha256:27a5ce13ce84cfa03e09c4b102c25f50789366a79dd117caea54c5e368a65e04

Observation fde160cf-5a22-4a04-b497-a4af7dffca52 · outbound

This paper cites Membership privacy for machine learning models through knowledge transfer.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Membership privacy for machine learning models through knowledge transfer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.625008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.227063Z digest=sha256:b00d7ab4a7bc85ad3e9efb216a54c68561d5594b650215961164bff962f42db7

Observation 136abd5a-e8d0-448d-9c63-0776758c3d29 · outbound

This paper cites Intriguing properties of neural networks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Intriguing properties of neural networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.231621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.231621Z digest=sha256:229ff8d69c12c99d8a75fcd700c68bdcfe153643a0d0316f00bfc257cbfe56d0

Observation 4b50bbef-4638-454a-95e1-9126f9b9cd41 · outbound

This paper cites Rethinking the inception architecture for computer vision.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Rethinking the inception architecture for computer vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.235789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.235789Z digest=sha256:21acd4272616b80f105ea1628d39db0d45e251acb460fd35a59945bb55ec2195

Observation 56a17418-ccb2-4e3e-9145-eb28230bb3fa · outbound

This paper cites Mitigating membership inference attacks by {Self-Distillation} through a novel ensemble architecture.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Mitigating membership inference attacks by {Self-Distillation} through a novel ensemble architecture

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.600289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.239929Z digest=sha256:c4af11af91f592b15000be7d5f6a21ea09487a0ab736162679c72752fa99f17a

Observation 79c8cc56-cfa2-4566-87af-1c26d0599b26 · outbound

This paper cites T3: Tree-autoencoder constrained adversarial text generation for targeted attack.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models T3: Tree-autoencoder constrained adversarial text generation for targeted attack

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.244405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.244405Z digest=sha256:30461040083a23ef38323247db1b2ca52dab15def1a7d8c6317ae360aa9765c0

Observation 9b50b6b4-6224-4985-8d20-3b03d2c13926 · outbound

This paper cites Adversarial GLUE: A multi-task benchmark for robustness evaluation of language mod- els.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Adversarial GLUE: A multi-task benchmark for robustness evaluation of language mod- els

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.570680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.253890Z digest=sha256:a1d66c83c40beff3ff3b3bdea58ab858ea4bb8639d32b1118b6a3107e134ef23

Observation 95f7a2a4-031b-42ae-a276-03d4b2bf2b54 · outbound

This paper cites Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Decodingtrust: A comprehensive assessment of trustworthiness in gpt models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.556048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.258921Z digest=sha256:d6f9be342303966671fb61b3c30ab642a9627a118959f57330cecbd6261674c8

Observation b527d9bf-42fb-4362-b7fc-48c93c33d6df · outbound

This paper cites EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.264471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.264471Z digest=sha256:c97ed96df177584748782893aa4652060ae98d8b01056bad23265f5d28ebe94d

Observation 4b9fdc43-8667-428d-b59f-596aafea6ff5 · outbound

This paper cites Emergent Abilities of Large Language Models.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Emergent Abilities of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.268642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.268642Z digest=sha256:c1f41232e5a402386522ab768aa2e4636cbba86300b9badb090cd8c1e5bc8cb6

Observation 571b1777-95d2-47a5-81e0-8cf3004c8391 · outbound

This paper cites Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.542888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.272777Z digest=sha256:accd1ecc2fc8fabbe5ccd7292eb7ef8f03ccf3732b456a519eaae375ac5d65c9

Observation ca347c03-fe7f-45e7-8344-5967c4a4ee8c · outbound

This paper cites Fairness constraints: Mechanisms for fair classification.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Fairness constraints: Mechanisms for fair classification

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.528494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.276628Z digest=sha256:3fd407aed6baee75e4317881164a6cd195107306ee89fde26fe47be31ae9d34f

Observation 7d72000f-3ebc-449d-83d3-2bcffab1ab4c · outbound

This paper cites Word-level textual adversarial attacking as combinatorial optimization.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Word-level textual adversarial attacking as combinatorial optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.280424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.280424Z digest=sha256:9adc657dd43fc53da780539f4cac79d5ca989a32f731cb8b0a81515a0338ebe9

Observation eae24b9f-29e0-4a75-a1bc-93bc9519a2fb · outbound

This paper cites Gender bias in coreference resolution: Evaluation and debiasing methods.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Gender bias in coreference resolution: Evaluation and debiasing methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.284964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.284964Z digest=sha256:ddd28f2c57bacdeee05ded3d5d51bac5dffb55385664b7726929c53fa22f55f8

Observation c77db318-aa73-46a4-944e-4f3fd6db689e · outbound

This paper cites Resisting membership inference attacks through knowledge distillation.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Resisting membership inference attacks through knowledge distillation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.514659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.289433Z digest=sha256:0412efcd86d9c6791362619fb30965176e67280982523316d601fd52207c6ad6

Observation 3bc2b9de-59b9-47d6-b537-9db6bc2b7ee4 · outbound

This paper cites FreeLB: Enhanced Adversarial Training for Natural Language Understanding.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models FreeLB: Enhanced Adversarial Training for Natural Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.294000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.294000Z digest=sha256:37e5fb62d64c3599e01503b85b01f1b153064e627bbe2ba78b87e60ed0087a3e

Observation b4588d93-7f23-4a3b-8257-e29287254382 · outbound

This paper cites an unresolved cited work.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Unresolved cited work

Reference 495

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:56:17.585753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:56:17.248663Z digest=sha256:40857b2ce2a1a3ae2b2dc6c249431de0488971fb088bb5c1087af243cdd0d143

Pith citing papers

Observation b6afeab1-5c4c-4478-b064-1b062820a5b8 · inbound

The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration cites this paper.

The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T15:23:15.199636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:23:15.199636Z digest=sha256:6e5c6296ecf4b50f680d9592c0fab449aef39185cc59ea8979c3c513ee60985d

Observation 6272a041-0594-41f9-9313-7adfab482516 · inbound

On Weak-to-Strong Generalization and f-Divergence cites this paper.

On Weak-to-Strong Generalization and f-Divergence Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:46.722521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:17:46.722521Z digest=sha256:ddc1c11075409a3f8e3e03deb55dd7ace4cd717ecc8c51139f9cc86487c6db02

Observation 3ba96de3-5d0a-4103-b537-81e561435348 · inbound

On the Blessing of Pre-training in Weak-to-Strong Generalization cites this paper.

On the Blessing of Pre-training in Weak-to-Strong Generalization Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:08.625742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T14:59:19.883399Z digest=sha256:7448e77c233847f1cef4fdf0c8bd7589097586d48e95d9141a45c3b25ec95dc8

Observation ddc2a02a-6551-4dd4-a954-018a6182b1af · inbound

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher cites this paper.

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:42:24.943936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:35:24.108984Z digest=sha256:792d26ddee80617cc96b290bc272cc39282ce94d14581ffdf43e695430b309bd