Search Results for author: Ryan Greenblatt

Found 6 papers, 6 papers with code

Alignment faking in large language models

1 code implementation18 Dec 2024 Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Evan Hubinger

Explaining this gap, in almost all cases where the model complies with a harmful query from a free user, we observe explicit alignment-faking reasoning, with the model stating it is strategically answering harmful queries in training to preserve its preferred harmlessness behavior out of training.

Large Language Model

Stress-Testing Capability Elicitation With Password-Locked Models

1 code implementation29 May 2024 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, David Krueger

To do this, we introduce password-locked models, LLMs fine-tuned such that some of their capabilities are deliberately hidden.

AI Control: Improving Safety Despite Intentional Subversion

1 code implementation12 Dec 2023 Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien Roger

This protocol asks GPT-4 to write code, and then asks another instance of GPT-4 whether the code is backdoored, using various techniques to prevent the GPT-4 instances from colluding.

Red Teaming

Preventing Language Models From Hiding Their Reasoning

1 code implementation27 Oct 2023 Fabien Roger, Ryan Greenblatt

Large language models (LLMs) often benefit from intermediate steps of reasoning to generate answers to complex problems.

Benchmarks for Detecting Measurement Tampering

1 code implementation29 Aug 2023 Fabien Roger, Ryan Greenblatt, Max Nadeau, Buck Shlegeris, Nate Thomas

When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization.

Cannot find the paper you are looking for? You can Submit a new open access paper.