{"id":63289,"date":"2024-05-13T18:39:16","date_gmt":"2024-05-13T10:39:16","guid":{"rendered":"http:\/\/www.upgrademag.com\/web\/?p=63289"},"modified":"2024-05-13T18:39:18","modified_gmt":"2024-05-13T10:39:18","slug":"ai-systems-are-already-skilled-at-deceiving-and-manipulating-humans-study","status":"publish","type":"post","link":"http:\/\/www.upgrademag.com\/web\/2024\/05\/13\/ai-systems-are-already-skilled-at-deceiving-and-manipulating-humans-study\/","title":{"rendered":"AI systems are already skilled at deceiving and manipulating humans &#8211; study"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>Many artificial intelligence (AI) systems have already learned how to deceive humans, even systems that have been trained to be helpful and honest. In a review article published in the journal\u00a0<em>Patterns<\/em>, researchers describe the risks of deception by AI systems and call for governments to develop strong regulations to address this issue as soon as possible.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cAI developers do not have a confident understanding of what causes undesirable AI behaviors like deception,\u201d says first author Peter S. Park (<a href=\"https:\/\/twitter.com\/dr_park_phd\">@dr_park_phd<\/a>), an AI existential safety postdoctoral fellow at MIT. \u201cBut generally speaking, we think AI deception arises because a deception-based strategy turned out to be the best way to perform well at the given AI\u2019s training task. Deception helps them achieve their goals.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Park and colleagues analyzed literature focusing on ways in which AI systems spread false information\u2014through learned deception, in which they systematically learn to manipulate others.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The most striking example of AI deception the researchers uncovered in their analysis was Meta\u2019s CICERO, an AI system designed to play the game Diplomacy, which is a world-conquest game that involves building alliances. Even though Meta claims it trained CICERO to be \u201c<a href=\"https:\/\/www.science.org\/doi\/10.1126\/science.ade9097?adobe_mc=MCMID%3D04915479217643570904386242774138863363%7CMCORGID%3D242B6472541199F70A4C98A6%2540AdobeOrg%7CTS%3D1715596555\">largely honest and helpful<\/a>\u201d and to \u201c<a href=\"https:\/\/twitter.com\/ml_perception\/status\/1595126521169326081\">never intentionally backstab<\/a>\u201d its human allies while playing the game, the data the company published along with its&nbsp;<em>Science<\/em>&nbsp;paper revealed that CICERO didn\u2019t play fair.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cWe found that Meta\u2019s AI had learned to be a master of deception,\u201d says Park. \u201cWhile Meta succeeded in training its AI to win in the game of Diplomacy\u2014CICERO placed in the top 10% of human players who had played more than one game\u2014Meta failed to train its AI to win honestly.\u201d<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"has-vivid-cyan-blue-color has-text-color has-link-color has-large-font-size wp-elements-1 wp-block-paragraph\"><strong>Other AI systems demonstrated the ability to bluff in a game of Texas hold \u2018em poker against professional human players, to fake attacks during the strategy game Starcraft II in order to defeat opponents, and to misrepresent their preferences in order to gain the upper hand in economic negotiations.<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">While it may seem harmless if AI systems cheat at games, it can lead to \u201cbreakthroughs in deceptive AI capabilities\u201d that can spiral into more advanced forms of AI deception in the future, Park added.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Some AI systems have even learned to cheat tests designed to evaluate their safety, the researchers found. In one study, AI organisms in a digital simulator \u201cplayed dead\u201d in order to trick a test built to eliminate AI systems that rapidly replicate.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cBy systematically cheating the safety tests imposed on it by human developers and regulators, a deceptive AI can lead us humans into a false sense of security,\u201d says Park.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The major near-term risks of deceptive AI include making it easier for hostile actors to commit fraud and tamper with elections, warns Park. Eventually, if these systems can refine this unsettling skill set, humans could lose control of them, he says.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cWe as a society need as much time as we can get to prepare for the more advanced deception of future AI products and open-source models,\u201d says Park. \u201cAs the deceptive capabilities of AI systems become more advanced, the dangers they pose to society will become increasingly serious.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">While Park and his colleagues do not think society has the right measure in place yet to address AI deception, they are encouraged that policymakers have begun taking the issue seriously through measures. But it remains to be seen, Park says, whether policies designed to mitigate AI deception can be strictly enforced given that AI developers do not yet have the techniques to keep these systems in check.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cIf banning AI deception is politically infeasible at the current moment, we recommend that deceptive AI systems be classified as high risk,\u201d says Park.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.eurekalert.org\/releaseguidelines\"><\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Other AI systems demonstrated the ability to bluff in a game of Texas hold \u2018em poker against professional human players, to fake attacks during the strategy game Starcraft II in order to defeat opponents, and to misrepresent their preferences in order to gain the upper hand in economic negotiations.<\/p>\n","protected":false},"author":6,"featured_media":63292,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[19],"tags":[853,6375,3952,96,2727,1656],"class_list":["post-63289","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-headlines","tag-artificial-intelligence","tag-robotic-process-automation","tag-robotics","tag-technology","tag-technology-adaption","tag-technology-investment"],"_links":{"self":[{"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/posts\/63289","targetHints":{"allow":["GET"]}}],"collection":[{"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/comments?post=63289"}],"version-history":[{"count":0,"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/posts\/63289\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/media\/63292"}],"wp:attachment":[{"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/media?parent=63289"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/categories?post=63289"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/www.upgrademag.com\/web\/wp-json\/wp\/v2\/tags?post=63289"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}