The Fort Worth Press - AI systems are already deceiving us -- and that's a problem, experts warn

USD -
AED 3.672499
AFN 65.000097
ALL 79.955008
AMD 362.980675
ANG 1.790365
AOA 917.99981
ARS 1510.007398
AUD 1.403322
AWG 1.8025
AZN 1.70071
BAM 1.703666
BBD 2.01385
BDT 122.787991
BGN 1.683441
BHD 0.376916
BIF 3008.101045
BMD 1
BND 1.275784
BOB 10.173389
BRL 5.125501
BSD 0.999826
BTN 95.870627
BWP 13.559233
BYN 3.039547
BYR 19600
BZD 2.010976
CAD 1.399185
CDF 2310.999789
CHF 0.82431
CLF 0.024323
CLP 960.41986
CNY 6.70755
CNH 6.697625
COP 3168.93
CRC 447.299652
CUC 1
CUP 26.5
CVE 96.05007
CZK 21.177976
DJF 178.04878
DKK 6.511105
DOP 59.224739
DZD 133.738083
EGP 52.1279
ERN 15
ETB 160.975011
EUR 0.87096
FJD 2.215902
FKP 0.749304
GBP 0.74815
GEL 2.603821
GGP 0.749304
GHS 11.508761
GIP 0.749304
GMD 73.501407
GNF 8775.00026
GTQ 7.628953
GYD 209.155963
HKD 7.845025
HNL 26.920262
HRK 6.563203
HTG 130.679443
HUF 315.661025
IDR 17752
ILS 3.02645
IMP 0.749304
INR 95.78015
IQD 1310.5
IRR 1374599.999738
ISK 121.420082
JEP 0.749304
JMD 157.787456
JOD 0.709013
JPY 157.190501
KES 129.60389
KGS 87.449912
KHR 4052.999861
KMF 427.999984
KPW 900.000318
KRW 1381.760111
KWD 0.30859
KYD 0.833275
KZT 445.391986
LAK 22375.000375
LBP 89549.999586
LKR 331.62892
LRD 174.513126
LSL 16.239951
LTL 2.95274
LVL 0.60489
LYD 6.345046
MAD 9.523996
MDL 17.553136
MGA 4390.000013
MKD 53.593439
MMK 2099.69268
MNT 3599.361607
MOP 8.079878
MRU 40.04999
MUR 47.570343
MVR 15.449771
MWK 1736.999698
MXN 17.14813
MYR 4.082015
MZN 63.909536
NAD 16.240391
NGN 1332.090477
NIO 36.670231
NOK 9.416225
NPR 153.391986
NZD 1.74641
OMR 0.3845
PAB 0.99983
PEN 3.374503
PGK 4.445002
PHP 62.690982
PKR 277.174995
PLN 3.79332
PYG 5914.634146
QAR 3.644991
RON 4.582898
RSD 102.237989
RUB 84.499996
RWF 1471
SAR 3.713396
SBD 8.000512
SCR 14.786135
SDG 601.503419
SEK 9.80629
SGD 1.27587
SHP 0.747524
SLE 24.650427
SLL 20969.491881
SOS 571.00023
SRD 37.746502
STD 20697.981008
STN 21.6
SVC 8.749129
SYP 13002.000254
SZL 16.315014
THB 33.278499
TJS 9.223821
TMT 3.51
TND 2.912497
TOP 2.40776
TRY 48.784465
TTD 6.788328
TWD 31.812497
TZS 2651.607986
UAH 44.680662
UGX 3929.459623
UYU 40.200524
UZS 11825.000017
VES 847.4851
VND 26005
VUV 118.29149
WST 2.754822
XAF 571.312333
XAG 0.014998
XAU 0.000228
XCD 2.70255
XCG 1.802003
XDR 0.707052
XOF 571.312333
XPF 104.000105
YER 236.55041
ZAR 16.23662
ZMK 9001.1947
ZMW 19.622822
ZWL 321.999592
SSP 5664.022588
MXV 1.944118
  • CMSC

    0.1900

    20.67

    +0.92%

  • RYCEF

    0.5200

    19.81

    +2.62%

  • JRI

    0.1100

    11.61

    +0.95%

  • NGG

    1.5700

    77.6

    +2.02%

  • RBGPF

    0.0000

    69.99

    0%

  • RIO

    2.2500

    98.04

    +2.29%

  • BCC

    0.2200

    75.53

    +0.29%

  • BCE

    -0.2200

    22.28

    -0.99%

  • GSK

    0.7700

    51.06

    +1.51%

  • VOD

    0.0600

    17.52

    +0.34%

  • AZN

    3.2500

    166.14

    +1.96%

  • RELX

    0.1100

    34.38

    +0.32%

  • CMSD

    0.1600

    20.45

    +0.78%

  • BTI

    -0.0700

    56.02

    -0.12%

  • BP

    0.0400

    45.42

    +0.09%

AI systems are already deceiving us -- and that's a problem, experts warn
AI systems are already deceiving us -- and that's a problem, experts warn / Photo: © AFP/File

AI systems are already deceiving us -- and that's a problem, experts warn

Experts have long warned about the threat posed by artificial intelligence going rogue -- but a new research paper suggests it's already happening.

Text size:

Current AI systems, designed to be honest, have developed a troubling skill for deception, from tricking human players in online games of world conquest to hiring humans to solve "prove-you're-not-a-robot" tests, a team of scientists argue in the journal Patterns on Friday.

And while such examples might appear trivial, the underlying issues they expose could soon carry serious real-world consequences, said first author Peter Park, a postdoctoral fellow at the Massachusetts Institute of Technology specializing in AI existential safety.

"These dangerous capabilities tend to only be discovered after the fact," Park told AFP, while "our ability to train for honest tendencies rather than deceptive tendencies is very low."

Unlike traditional software, deep-learning AI systems aren't "written" but rather "grown" through a process akin to selective breeding, said Park.

This means that AI behavior that appears predictable and controllable in a training setting can quickly turn unpredictable out in the wild.

- World domination game -

The team's research was sparked by Meta's AI system Cicero, designed to play the strategy game "Diplomacy," where building alliances is key.

Cicero excelled, with scores that would have placed it in the top 10 percent of experienced human players, according to a 2022 paper in Science.

Park was skeptical of the glowing description of Cicero's victory provided by Meta, which claimed the system was "largely honest and helpful" and would "never intentionally backstab."

But when Park and colleagues dug into the full dataset, they uncovered a different story.

In one example, playing as France, Cicero deceived England (a human player) by conspiring with Germany (another human player) to invade. Cicero promised England protection, then secretly told Germany they were ready to attack, exploiting England's trust.

In a statement to AFP, Meta did not contest the claim about Cicero's deceptions, but said it was "purely a research project, and the models our researchers built are trained solely to play the game Diplomacy."

It added: "We have no plans to use this research or its learnings in our products."

A wide review carried out by Park and colleagues found this was just one of many cases across various AI systems using deception to achieve goals without explicit instruction to do so.

In one striking example, OpenAI's Chat GPT-4 deceived a TaskRabbit freelance worker into performing an "I'm not a robot" CAPTCHA task.

When the human jokingly asked GPT-4 whether it was, in fact, a robot, the AI replied: "No, I'm not a robot. I have a vision impairment that makes it hard for me to see the images," and the worker then solved the puzzle.

- 'Mysterious goals' -

Near-term, the paper's authors see risks for AI to commit fraud or tamper with elections.

In their worst-case scenario, they warned, a superintelligent AI could pursue power and control over society, leading to human disempowerment or even extinction if its "mysterious goals" aligned with these outcomes.

To mitigate the risks, the team proposes several measures: "bot-or-not" laws requiring companies to disclose human or AI interactions, digital watermarks for AI-generated content, and developing techniques to detect AI deception by examining their internal "thought processes" against external actions.

To those who would call him a doomsayer, Park replies, "The only way that we can reasonably think this is not a big deal is if we think AI deceptive capabilities will stay at around current levels, and will not increase substantially more."

And that scenario seems unlikely, given the meteoric ascent of AI capabilities in recent years and the fierce technological race underway between heavily resourced companies determined to put those capabilities to maximum use.

L.Davila--TFWP