Liane Lovitt
Identifiers
- name variant Liane Lovitt 0.60 · backfill
Papers (8)
- Towards Measuring the Representation of Subjective Global Opinions in Language Models cs.CL · 2023 · author #11
- Discovering Language Model Behaviors with Model-Written Evaluations cs.CL · 2022 · author #33
- Constitutional AI: Harmlessness from AI Feedback cs.CL · 2022 · author #26
- Measuring Progress on Scalable Oversight for Large Language Models cs.HC · 2022 · author #26
- In-context Learning and Induction Heads cs.LG · 2022 · author #19
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cs.CL · 2022 · author #2
- Language Models (Mostly) Know What They Know cs.CL · 2022 · author #25
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cs.CL · 2022 · author #22
Mentions
- 2211.03540 #26 · arxiv_oai · confidence 0.70 Liane Lovitt
- 2306.16388 #11 · arxiv_oai · confidence 0.70 Liane Lovitt
Frequent Coauthors
- Amanda Askell 8 shared papers
- Jared Kaplan 8 shared papers
- Nicholas Joseph 8 shared papers
- Sam McCandlish 8 shared papers
- Zac Hatfield-Dodds 8 shared papers
- Andy Jones 7 shared papers
- Anna Chen 7 shared papers
- Ben Mann 7 shared papers
- Danny Hernandez 7 shared papers
- Dario Amodei 7 shared papers
- Dawn Drain 7 shared papers
- Deep Ganguli 7 shared papers
- Jackson Kernion 7 shared papers
- Kamal Ndousse 7 shared papers
- Nelson Elhage 7 shared papers
- Nova DasSarma 7 shared papers
- Scott Johnston 7 shared papers
- Tom Brown 7 shared papers
- Tom Henighan 7 shared papers
- Yuntao Bai 7 shared papers