Pith. sign in

Paper Citation Record · LEDGER

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

As of 19 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 1 inbound Pith citation observation for arXiv:2502.05242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05242 v3

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:06:56.135403Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:01:59.215909Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T23:02:04.281602Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact7
  • verified fuzzy14
  • unresolved69
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3e8f967-b44d-4ec1-bfce-2738e74c1b1f · outbound

This paper cites Generative AI Text Classification using Ensemble LLM Approaches.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Generative AI Text Classification using Ensemble LLM Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.774507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.774507Z digest=sha256:167ceabb5aa726c485cb296bd80267e0bcd8a4f1eb1602d61a6091fb85f36ad1

Observation e63ebdc5-3aad-4860-8b8c-4fcc8b8591fb · outbound

This paper cites GPT-4 Technical Report.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.779567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.779567Z digest=sha256:685d59284061d3fd8d02613948d935ba6c72ad4433f05a8de1656b1fa6d7e40b

Observation a0e013a8-1d8a-4685-8774-759eba133331 · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Internal State of an LLM Knows When It's Lying

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.783768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.783768Z digest=sha256:ad6235470b85cbe1fef75286aa1f27603bbcce67bf92c970db5e0a7225aed4a9

Observation 457ddc55-500e-4a0c-be87-d9b5acc5e94d · outbound

This paper cites Transparency and explainability of ai systems: ethical guidelines in practice.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Transparency and explainability of ai systems: ethical guidelines in practice

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.788132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.788132Z digest=sha256:82fe17e3ddd071a37251c7f5bac6a68979af0195f08dc22e6f7842fe08a2205a

Observation b937a963-1c1c-4573-bf94-74451f211016 · outbound

This paper cites Spectrally-normalized margin bounds for neural networks.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Spectrally-normalized margin bounds for neural networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.792115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.792115Z digest=sha256:74428e7ab07751efd2e3a26aaec124bf86e00726d37d8a874887d6d3d34fdf9f

Observation dd51a220-ef09-4d7a-aed8-ca8c6dcd3fca · outbound

This paper cites Language models can explain neurons in language models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Language models can explain neurons in language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.796051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.796051Z digest=sha256:f6bae2ae16c86396ef14d0a9736d90026c257cdba07cd179065fbe3b989d1723

Observation eff633a2-2ed4-4a3e-9fb4-c940d9671e89 · outbound

This paper cites High-Dimension Human Value Representation in Large Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring High-Dimension Human Value Representation in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.800491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.800491Z digest=sha256:5cbf930a57939f86e8366600d8e7ad24e7aa444bd7bca97edec24cdbb1226673

Observation 8f4e9428-e1ba-410b-b894-69a5806bac97 · outbound

This paper cites Internlm2 technical report, 2024.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Internlm2 technical report, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.804422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.804422Z digest=sha256:0cc8ed022bdb0478c0413c7b16e4cdc1e1861f252bcc71937c9176c584d51917

Observation b4f9efca-d4c3-4300-863c-4cf5ea4377c2 · outbound

This paper cites Redunet: A white-box deep network from the principle of maximizing rate reduction.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Redunet: A white-box deep network from the principle of maximizing rate reduction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.808319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.808319Z digest=sha256:ad33fe6c34f9c3d05b6bc8d240829c0cb17586135f0413b5e6424d2bc7864a9a

Observation ae97825f-34b0-4041-aa53-5e11bac06f2d · outbound

This paper cites SelfIE: Self-Interpretation of Large Language Model Embeddings.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring SelfIE: Self-Interpretation of Large Language Model Embeddings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.812069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.812069Z digest=sha256:e70a827252e16c5eb207ca0955d320406f197faa8c67f3209c0bc1914bb87879

Observation 65862f5e-6e05-4c89-8962-8c35adb1434c · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring A simple framework for contrastive learning of visual representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.815963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.815963Z digest=sha256:22c06449194468a6c785faa58e8248d79dc24904e23f72890e79ada813c42a8f

Observation b0d84af8-53d5-4908-8fbe-29a7d337e092 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.819737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.819737Z digest=sha256:aa97df7360ebbebb6e850caaa8ba0a9ae6cb1a7da573a2546244fdd210ad33a9

Observation b3eb7f9c-41eb-4d33-8e1b-1c3bb0803fbc · outbound

This paper cites Token Prediction as Implicit Classification to Identify LLM-Generated Text.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Token Prediction as Implicit Classification to Identify LLM-Generated Text

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-08T21:06:56.907064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.823842Z digest=sha256:f27fdd5d73ac9d045529c88cbebda50d7cabe59d561ebe487ac097d1f2dc3c8f

Observation 5939e0aa-5dd2-4819-87c4-6cee38a1f6ef · outbound

This paper cites Measuring generalization with optimal transport.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Measuring generalization with optimal transport

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.827723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.827723Z digest=sha256:87c8cae0cc99c71e7ca95a6cbd01f2da98e4764206c5a91e270b01ca99b4ca6f

Observation 501c34ff-aac9-4645-b297-50d7bd509842 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.831386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.831386Z digest=sha256:ed1f0f095a5c96d8ca58b205ab606ed1d6c04a9ec0d30d04356743124ea96b5a

Observation 6f210b8b-a74b-4846-8cc4-0969bdab1572 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Opencompass: A universal evaluation platform for foundation models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.835382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.835382Z digest=sha256:412d6bcd909870f4d86401abec2486fc31bb2f975625334096b3508f045a59a2

Observation a558fa4f-2a39-4149-80df-bd468a5a5042 · outbound

This paper cites Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.839319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.839319Z digest=sha256:dc62b51790f92bb370397274ea883811cdab58c28e48e5e3780e39f926c3809b

Observation a712fe56-9fc0-4031-9daf-03316b48d32b · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.843649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.843649Z digest=sha256:111f6b1bb0faf6dbf460d6d297777e621451a16b7583621c7babe4803da60ee2

Observation 5e02df84-da1c-4178-a019-39ed782397c8 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.848017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.848017Z digest=sha256:2ed6951dc7d0795c922746a572dbe7d6108ffeaaeabcfe20801a48df43bad7fd

Observation 80e24131-ca75-4d12-947c-9c2e86df1e54 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Scaling and evaluating sparse autoencoders

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.852430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.852430Z digest=sha256:cb841799f52324711eabb71b13d6f61b6c99a6cd3e1da9c2fed1260c34f96515

Observation 0b6e1392-4cdc-4298-9c06-0f4bd94eaef8 · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.856894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.856894Z digest=sha256:2f5fef6893feae524e58d52ca6bf996e6d7533533b3a42c15d09493dd8010e24

Observation 0187d06a-8394-441e-9bb2-d9e927e54d26 · outbound

This paper cites Alignment faking in large language models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Alignment faking in large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.861230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.861230Z digest=sha256:f18c155ffcff065343b52aab7fac70a5e40d92f72418fb4f76c483b2f61239fc

Observation b469aef6-7e91-495f-ad40-23df6a2e9795 · outbound

This paper cites Dimensionality reduction by learning an invariant mapping.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Dimensionality reduction by learning an invariant mapping

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.865542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.865542Z digest=sha256:3c66fb3ef473b1d921bfa995fe6a083d22ec25750a377120cff860d3929be775

Observation b0fa4ebb-603c-496e-b878-aafe3fdfd14b · outbound

This paper cites Masked autoencoders are scalable vision learners.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Masked autoencoders are scalable vision learners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.869642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.869642Z digest=sha256:5294b1cbbd132adbf502f7995bbf492c344459ad534b101d628c11bfbcb24eba

Observation e4bcd3c8-ce88-43ee-8aff-f1ca7ead4c4c · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Momentum contrast for unsupervised visual representation learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.873475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.873475Z digest=sha256:d5ac23ddf9bc9965a04088b901097d4c46c43f2489a00f47faca7fcf1cb531b4

Observation 688d65ed-4bff-4110-a7af-35d28a44760d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Measuring Massive Multitask Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.877048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.877048Z digest=sha256:e09acc6361e52df5fdd1109bc4b996bdcccfd6f4f4eef871aae0ef8d276f744b

Observation b175704b-317e-4612-ad96-c1a7e576cdfc · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Measuring mathematical problem solving with the math dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.880776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.880776Z digest=sha256:8f3dad6f77b6232fc14d90fa51deebdc465ec9250068638e13a14a8f3c5bab54

Observation f7ac45f8-93e2-4db4-8b6e-513633c2f52a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring LoRA: Low-Rank Adaptation of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.884405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.884405Z digest=sha256:11d14d9aa76d9ccf123e72c388c2623372d04209758f23f0fc4b67b91edf2601

Observation f2ff69cd-e2d8-495a-bf29-c5ad0093834c · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.888220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.888220Z digest=sha256:516d2e005baaf5cc1a342bc893bf3feadb4ea9e7ec113a83589b2998e4c1f9c6

Observation 173d7867-b738-4c0a-9787-e087706fc231 · outbound

This paper cites Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.891758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.891758Z digest=sha256:61e0c563eba46407414e2913093275f6eb82c6745063596152e85408bd7ffbe9

Observation 10957a03-58f7-494b-875c-b555402d289d · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.895736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.895736Z digest=sha256:01a4712c5e8b79cbfc9441f39735c6fc40e9937d11066cad5a397f979d6704aa

Observation c35dd0db-4a15-410d-8ace-4a2453be7f10 · outbound

This paper cites Anthropic: Responsible scaling policy.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Anthropic: Responsible scaling policy

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.332762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.899620Z digest=sha256:1ffc6c15fdae4865970db65320d60668855b071eedd5e49b330f3962affb8e3d

Observation 5e83a56d-4123-4e6b-b529-66b410741f27 · outbound

This paper cites The Platonic Representation Hypothesis.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Platonic Representation Hypothesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.903599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.903599Z digest=sha256:701de5ac85a4f5d46141d1a7edd567da573c2804a790f1e18d59bbb3e031218b

Observation 83447c0a-c73c-4565-8128-9ab08938f379 · outbound

This paper cites Klanderman, and William J Rucklidge.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Klanderman, and William J Rucklidge

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.318683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.907611Z digest=sha256:fb82bdb0938acc05387d241cb251be1a15bfb8e27401e77ba03673f2ee37151f

Observation a4e41a44-f263-41ee-84a9-6447e16b6d3b · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.911523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.911523Z digest=sha256:8c5427038d63d5a80cdb2782ab6a2dc923d4953a60815a921fb0859d490277d9

Observation 70aa9370-bc40-4292-9e1d-49f62f31b051 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.915594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.915594Z digest=sha256:c99ed62d8a488b267e43ee5ead952322f13f90bf1279a6ef30ff64024826429c

Observation adcf540f-b082-4a3e-8918-95a7bd6c7d59 · outbound

This paper cites Mistral 7B.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Mistral 7B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.919286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.919286Z digest=sha256:18cd8c1fb5818004ded301cae254fce6b76e571be31384aeeaa089741445be73

Observation d035885d-8c60-4038-ad07-fdf4cbde629f · outbound

This paper cites NeurIPS 2020 Competition: Predicting Generalization in Deep Learning.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring NeurIPS 2020 Competition: Predicting Generalization in Deep Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-08T21:06:56.683476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.923429Z digest=sha256:af9399063a518c8e85901e6638259e200d97cbb8c4ced6744a3533f59d497b79

Observation 13866d13-644a-4719-b597-dc36bef6e648 · outbound

This paper cites Explainable artifi- cial intelligence for mental health through transparency and interpretability for understandability.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Explainable artifi- cial intelligence for mental health through transparency and interpretability for understandability

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.295179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.927512Z digest=sha256:0bb6d93780a3830a1e76aab2477b030de6e758aff0203bea6cd90ff6adbe2b3b

Observation ae57fff5-306c-420d-9be1-75a143f4fa31 · outbound

This paper cites Theoretical Analysis of Weak-to-Strong Generalization.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Theoretical Analysis of Weak-to-Strong Generalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.931312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.931312Z digest=sha256:37218a95f8050a684b282b34576ce679bbc5aebdafac4d15a4b8321f889b3f88

Observation e56eae66-2a65-4723-84a4-b2a9ea64b870 · outbound

This paper cites Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.935526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.935526Z digest=sha256:5476c59fc808701aaddbddbfbf8528ccded9c60feac69744d3c505c87dc25ba3

Observation 4ff0e8ae-7aed-43ef-b4ef-f79cf43fff2e · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Inference- time intervention: Eliciting truthful answers from a language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.281700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.939591Z digest=sha256:b100a533c5f1ce6a588460811a4cb527f6fbdbf863d39be0e3a469a13a809a0e

Observation 3b5c0a1c-1092-451e-abaf-5fcbafeda044 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.943327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.943327Z digest=sha256:50820f7062d9aea4b2f299547cf803de60203a75eb3adc8271ada11086639590

Observation 91f68183-74c7-4e22-bd63-d35465225cf9 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.947697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.947697Z digest=sha256:bdf5c83784e0a6a5bb46aca4b937f2d79d40fe2fdb645064d7fdeaa278e5b2a6

Observation 20f0ad26-1ff9-4a3a-8c07-5e261d86327f · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.952238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.952238Z digest=sha256:0cc009c6dd30f6fc8f585b9f5732775a3b793a2924d2091c5c78a75b0f322dfa

Observation 8ce0c3e3-8ff9-4f56-949e-5081cf4e12d3 · outbound

This paper cites Large language models in finance: A survey.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Large language models in finance: A survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.956497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.956497Z digest=sha256:cd3b25e7af32e30fbdcc9a0867b0b56551ace026180f662138497ca2d458c1eb

Observation fa921c34-69a1-4f13-bb2a-16033b96982f · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.960754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.960754Z digest=sha256:977f97cc9f2ee9b047d074c6355339346160ee104754ded09be55e662c8b066f

Observation 2fe70bf5-43ea-4c1f-9120-0d1084c5f9b8 · outbound

This paper cites Latent guard: a safety framework for text-to-image generation.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Latent guard: a safety framework for text-to-image generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.258531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.965090Z digest=sha256:d3a77d6ce7445894e7cdb2a49b6d1913ad77dabd8aca2ee92ba5248d2c43f90b

Observation 8a83b41f-745d-483b-9005-e7157117c742 · outbound

This paper cites Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-08T21:06:56.576746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.969484Z digest=sha256:4b90763838aa0b79f980f9665896768ad1ee791d264dae1098e501eecfb35e7e

Observation 036de2d0-4613-491c-bb9f-b3cc5f5c938a · outbound

This paper cites Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-Rationalization.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-Rationalization

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-08T21:06:56.558474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.973742Z digest=sha256:67a270dde2207d646d09cbfcf696ce6a166d2e31387b2822dcf68af7ba9dcb39

Observation a900ae02-3466-4912-8423-ea555e038e6d · outbound

This paper cites MGR: Multi-generator Based Rationalization.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring MGR: Multi-generator Based Rationalization

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-08T21:06:56.537652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.977711Z digest=sha256:922f655ce4a00f83be3992ed94830d8ee059a3c099123c5e8b0297d5edc353ce

Observation 4e06cb80-918d-4863-8234-b1daffbddaf5 · outbound

This paper cites D-separation for causal self-explanation.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring D-separation for causal self-explanation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.245452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.981604Z digest=sha256:160f4c75df8308cff6d61bcd7cd7b94d5d87a4b9ef0997cce1afdfb4770de330

Observation df5fed00-963e-4362-b243-b73676f467f7 · outbound

This paper cites Efficient detection of toxic prompts in large language models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Efficient detection of toxic prompts in large language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.985225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.985225Z digest=sha256:4b6f4759af55e928dcd44113059064dd4d74145af741729f57804768f82e753f

Observation 16164951-f8c2-42c2-95a9-11c2a14fd8e2 · outbound

This paper cites Are self-explanations from large language models faithful? In Findings of the Association for Computational Linguistics ACL 2024, pages 295–337, 2024.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Are self-explanations from large language models faithful? In Findings of the Association for Computational Linguistics ACL 2024, pages 295–337, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.222356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.988861Z digest=sha256:267433dd3e14f204f160b006bcc4ca4f633b422e072b413fbdaf98321ac440fa

Observation 446c31d7-65c7-4a8a-b4c3-89caa757e468 · outbound

This paper cites Introducing llama 3.1: Our most capable models to date.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Introducing llama 3.1: Our most capable models to date

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.207161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.992591Z digest=sha256:3945ff0644da16fab422bc60c2d8a7b58b9b8263c30fb3eac594aabb5a55a2c8

Observation 1d218019-ad92-4d0e-9b10-81277d858604 · outbound

This paper cites Large language models in healthcare and medical domain: A review.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Large language models in healthcare and medical domain: A review

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.996095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.996095Z digest=sha256:02f130f3fbdf2917a7745fee9368ba03c90dcc18dc24cbd1906b1ce0759c7586

Observation 20430a8d-5418-4d9d-9714-48fe056b05aa · outbound

This paper cites interpreting gpt: the logit lens, 2020.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring interpreting gpt: the logit lens, 2020

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.182480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:55.999980Z digest=sha256:7f01d0e2726ea7f0fba4af6134ca65d65609538800d411cd03c18de4aca21bd4

Observation 225d2322-a4f4-4637-8b8b-803d0835fc56 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.003552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.003552Z digest=sha256:50f27bd395c3642c6f5af74df54d8415490c56a5d31aa6c49d2747d39db0519c

Observation 2653e440-a101-4d97-a749-9c18c20fe79a · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Representation Learning with Contrastive Predictive Coding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.007455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.007455Z digest=sha256:11b4945ef4403f6a0520fa08c110cb0842fa85f43f35a4efba9369bc08e610bf

Observation 77650200-95ad-4254-ae66-b393f8e2d986 · outbound

This paper cites LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.011274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.011274Z digest=sha256:1a0c51ec8107ab5f2b115fa3cce6307bf1d7d74e56a695bf7383c75d7a67d396

Observation 26f79b94-4509-4a6c-9844-3692db5eff03 · outbound

This paper cites Training language models to follow instructions with human feedback.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Training language models to follow instructions with human feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.015292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.015292Z digest=sha256:1d420664e9dcaa5cca92b2371fdf5620f64a8cd4327365743a028d3d921284f6

Observation 9df7f434-40a5-467f-824c-394f877165eb · outbound

This paper cites The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.019041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.019041Z digest=sha256:10378b08e38501eb1ac76383f9088216a7b44c145759c26818457169569ff41a

Observation 747cf1a8-5e93-4708-8a45-16bfebbf9b3c · outbound

This paper cites Learning transferable visual models from natural language supervision.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Learning transferable visual models from natural language supervision

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.022839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.022839Z digest=sha256:54d432813dab8ed6478d3157f3a4c60408ce969f8e6b1cc5235799fe2028ae5c

Observation 8c0bdbcb-74c2-45ee-80d8-7e72561f0d4f · outbound

This paper cites Representation Noising: A Defence Mechanism Against Harmful Finetuning.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Representation Noising: A Defence Mechanism Against Harmful Finetuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.026593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.026593Z digest=sha256:e06d14a907f337d52d3ce92485c61399ab304fb899163ba11535aa6a69ebf705

Observation f1aa9788-3f59-43b1-9fcc-dc77ff0d4fc9 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.030792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.030792Z digest=sha256:341e7786cfb88d9959fa85967a68614a8fdecbbcdf4f675dfa143252d63487fd

Observation c0568c9d-14e2-486f-b012-9c06393075b6 · outbound

This paper cites The effective rank: A measure of effective dimensionality.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The effective rank: A measure of effective dimensionality

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.034858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.034858Z digest=sha256:153a2452e874ffefce79fd3f51200bfdbee8732700eb2e2854036528ad018b3f

Observation 60750dfb-be55-4241-815e-7f1a073573f9 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring SocialIQA: Commonsense Reasoning about Social Interactions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.038650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.038650Z digest=sha256:243673244f59cabaae9d410375c30edb643d7a89f8a68f1dd67cb9c068c865b5

Observation b947ce92-bb8f-4fc5-bbe7-cbb7fd9dc66f · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Facenet: A unified embedding for face recognition and clustering

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.042629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.042629Z digest=sha256:75c1875090b80abc36186cc1c787ca75d45529f0dff8881b4aee899923064be4

Observation 17821f04-8ab1-4875-9197-6ea97b00a04f · outbound

This paper cites Lmfingerprints: Visual explanations of language model embedding spaces through layerwise contextualization scores.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Lmfingerprints: Visual explanations of language model embedding spaces through layerwise contextualization scores

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.130582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.046235Z digest=sha256:6fbd6499d7ffcf9feb3b4382573a11c2efce64257f09e36397ad69db4118f1ad

Observation e2c8f106-6e1b-4c54-961a-ecbcd71885f5 · outbound

This paper cites k-variance: A clustered notion of variance.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring k-variance: A clustered notion of variance

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.116105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.050394Z digest=sha256:dca4a16426005f9615843fe1a4641e3cc39596c5054f73223ae2a44904af9b3d

Observation 1de0eab6-ba41-4b39-ab5e-adb1a6deb80d · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Gemma 2: Improving Open Language Models at a Practical Size

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.054470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.054470Z digest=sha256:571d86cfd80422a23f96fd36a461354216413b00f7d9d49218f9405f98ba5003

Observation 3e19d31c-1e90-4440-ad87-d31159744ba7 · outbound

This paper cites Daniel Freeman, Theodore R.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Daniel Freeman, Theodore R

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.059007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.059007Z digest=sha256:00469cf83b7cf3b424930879ddf7e458a407eed74c4efc4a39860de6aa9faee3

Observation 93d11c2a-7849-4444-b3f6-cb06558089a2 · outbound

This paper cites Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.063915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.063915Z digest=sha256:16e3c3e0075328110a9a74c2dbd96975c3d2e58a24e2e407fe2e553d99a2cc8c

Observation cd3ee770-4322-4bd8-9d30-f2c3e99d076c · outbound

This paper cites Optimal transport: old and new, volume 338.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Optimal transport: old and new, volume 338

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.068139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.068139Z digest=sha256:19863544472d507373fd323d84036d6c86f1a95173b6a8f9c4f694278252cc9f

Observation 7e5f683d-ee24-4e37-bb56-4e66c8530de7 · outbound

This paper cites The wasserstein distances.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The wasserstein distances

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.072383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.072383Z digest=sha256:93938ff6966d7ff16a023b421faab24fa802c338cb0b1e24b4b63bca2159d38e

Observation f8d9cf8e-f1cc-4754-95ea-77d0287e45b6 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.076683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.076683Z digest=sha256:8e4447d4358fe2ce44b209c547a0a1ff422119c5b676e7620f6abd6027611207

Observation c30e8831-9260-4a73-817d-aa3178454530 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Chain-of-thought prompting elicits reasoning in large language models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.081306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.081306Z digest=sha256:5d5f97eba48c089e86707ec4ce63bb456a437ec60aa6aedba81cdec9fb4188c1

Observation d1bae7d8-6fc3-4d92-8aa1-a01fabec5b31 · outbound

This paper cites Diff-erank: A novel rank-based metric for evaluating large language models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Diff-erank: A novel rank-based metric for evaluating large language models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.057955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.085269Z digest=sha256:5b7d47f538ca311517c5e0461ee9d92075235acf0a21b540ace7086cb1e98748

Observation 98241810-e6fc-4da6-b8c2-fdd71696400b · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring ReFT: Representation Finetuning for Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.089277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.089277Z digest=sha256:bb2eb6705dfde0e382f86e1ec359ae2f7f56b319de55020d2dc58f715fe3fe37

Observation 7c478ef5-0b34-4c83-a5d5-0e163c9ce198 · outbound

This paper cites Good Idea or Not, Representation of LLM Could Tell.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Good Idea or Not, Representation of LLM Could Tell

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-08T21:06:56.371735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.093042Z digest=sha256:a877e883579d22dd313ba1b664afce3f1c89c561ae1e544aba8200ba3c139b49

Observation f4abf341-265d-49b5-9dd6-72f6a8734d07 · outbound

This paper cites Qwen2.5 Technical Report.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Qwen2.5 Technical Report

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.096969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.096969Z digest=sha256:20b8e21b4797e7d1bc61a658347e9a1eb03745ecd60d56c4872df4ac420b5440

Observation 40f6b988-4f04-4cc0-a4eb-880f270b3154 · outbound

This paper cites A fingerprint for large language models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring A fingerprint for large language models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.100706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.100706Z digest=sha256:3b0b0deb06957da01c3742e6e792de796202e0cd67b6b0c6d0df8f130532229e

Observation 091dd076-9a66-4127-95df-8f08f6cecf83 · outbound

This paper cites LoFiT: Localized Fine-tuning on LLM Representations.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring LoFiT: Localized Fine-tuning on LLM Representations

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.104290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.104290Z digest=sha256:f0d9a1fd6f7769435d7b1c87471a5d35f4db211eca4b856a2b5c84f6ca75e46a

Observation dab498e3-2f81-4167-bdea-0c80fc1fe5dc · outbound

This paper cites Barlow twins: Self- supervised learning via redundancy reduction.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Barlow twins: Self- supervised learning via redundancy reduction

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.108043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.108043Z digest=sha256:810947df410315b56869efd0b1f4133b54855e855d0b430fe35f30897c9b4120

Observation d647cc87-41ae-40f1-8821-278ab1fcf4b6 · outbound

This paper cites Similar Data Points Identification with LLM: A Human-in-the-loop Strategy Using Summarization and Hidden State Insights.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Similar Data Points Identification with LLM: A Human-in-the-loop Strategy Using Summarization and Hidden State Insights

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.111620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.111620Z digest=sha256:c2bbb7e7121495f9d882bd9549cdbfa98ae97ccead7397ff2b6f559316f38af3

Observation dcbaeea0-a2a4-4a7a-b6d0-5364c23af323 · outbound

This paper cites REAL: Response Embedding-based Alignment for LLMs.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring REAL: Response Embedding-based Alignment for LLMs

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-08T21:06:56.207427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.115603Z digest=sha256:8279a1c8d99813ed6088b98529eadaa737a7a32da8aae408d7bd5de79e20cddf

Observation 62028e91-3ff0-4b67-8729-f27302ce3255 · outbound

This paper cites REEF: Representation Encoding Fingerprints for Large Language Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring REEF: Representation Encoding Fingerprints for Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.119689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.119689Z digest=sha256:fbc3387b902bb9302bb1a25072124e245e84a8c0e4927905688d55c1706240b5

Observation a5d8497d-01d9-4a05-9ebb-a9209b335a59 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.123451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.123451Z digest=sha256:8ebeb25d9312aefe52099d51f68821c27f28a07b0c0691e7dff230cdd5e29d7f

Observation d4a13672-7dec-4bf3-a6fa-701fb7cf5a0d · outbound

This paper cites Improving alignment and robustness with circuit breakers.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Improving alignment and robustness with circuit breakers

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.035985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.127637Z digest=sha256:f77665b3fb5202b2298b00fbb4311d5cdb740a2dd43765dfa83e3cf4475c9207

Observation 5e295cd8-8b93-4bf9-866d-ec29bbb9cc8e · outbound

This paper cites [ 62] disentangles LLMs’ awareness of fairness and privacy by deactivating the entangled neurons in representations.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring [ 62] disentangles LLMs’ awareness of fairness and privacy by deactivating the entangled neurons in representations

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:06:57.023153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.131162Z digest=sha256:79e6f84dc59d1fc355f231a7bfee43fa80c4e7e0000ec3482b7979290fa33735

Observation f4ef902d-32e2-4788-bb28-1158b52b904a · outbound

This paper cites safe" or.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring safe" or

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T21:06:57.010023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T21:06:56.135403Z digest=sha256:1661a97594772f1c05d0e1eb3a9ee66107d1c9fd781658c1349f25af4ada7714

Pith citing papers

Observation 0bdfb9f2-0a78-4996-bd33-6a36b60098b2 · inbound

Position: Intelligent Coding Systems Should Write Programs with Justifications cites this paper.

Position: Intelligent Coding Systems Should Write Programs with Justifications Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:02:04.284765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T23:01:59.215909Z digest=sha256:67d3fe10e9164b54310026e19a4c2f71cb852b557a57e3b40547fb12d8cc9ed1