pith. sign in

Liane Lovitt

Identifiers

  • name variant Liane Lovitt 0.60 · backfill

Papers (8)

  1. Towards Measuring the Representation of Subjective Global Opinions in Language Models cs.CL · 2023 · author #11
  2. Discovering Language Model Behaviors with Model-Written Evaluations cs.CL · 2022 · author #33
  3. Constitutional AI: Harmlessness from AI Feedback cs.CL · 2022 · author #26
  4. Measuring Progress on Scalable Oversight for Large Language Models cs.HC · 2022 · author #26
  5. In-context Learning and Induction Heads cs.LG · 2022 · author #19
  6. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cs.CL · 2022 · author #2
  7. Language Models (Mostly) Know What They Know cs.CL · 2022 · author #25
  8. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cs.CL · 2022 · author #22

Mentions

  • 2211.03540 #26 · arxiv_oai · confidence 0.70 Liane Lovitt
  • 2306.16388 #11 · arxiv_oai · confidence 0.70 Liane Lovitt

Frequent Coauthors