arxiv:2403.13793

Evaluating Frontier Models for Dangerous Capabilities

Published on Mar 20, 2024

· Submitted by

akhaliq on Mar 21, 2024

Upvote

Authors:

Mary Phuong ,

Sarah Cogan ,

David Lindner ,

Matthew Rahtz ,

Yannis Assael ,

Albert Webson ,

Anian Ruoss ,

Abstract

Evaluations of dangerous capabilities in AI models across persuasion, cyber-security, self-proliferation, and self-reasoning are introduced and pilot-tested on Gemini 1.0 models.

AI-generated summary

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

View arXiv page View PDF Add to collection