The Prague Post - AI systems are already deceiving us -- and that's a problem, experts warn

EUR -
AED 4.268807
AFN 76.134825
ALL 92.106109
AMD 423.129883
ANG 2.080394
AOA 1067.057439
ARS 1753.719365
AUD 1.611083
AWG 2.09227
AZN 1.976912
BAM 1.955947
BBD 2.341455
BDT 142.790699
BGN 1.971808
BHD 0.438325
BIF 3478.230682
BMD 1.162372
BND 1.471898
BOB 14.502962
BRL 5.959134
BSD 1.162577
BTN 109.780336
BWP 15.565032
BYN 3.575406
BYR 22782.492971
BZD 2.338055
CAD 1.606375
CDF 2650.20859
CHF 0.940411
CLF 0.02754
CLP 1087.434205
CNY 7.801031
CNH 7.798837
COP 3639.503252
CRC 527.634994
CUC 1.162372
CUP 30.80286
CVE 110.272313
CZK 24.194793
DJF 207.013729
DKK 7.474523
DOP 68.815156
DZD 154.797758
EGP 59.165083
ERN 17.435581
ETB 188.514603
FJD 2.576278
FKP 0.859847
GBP 0.858964
GEL 3.028018
GGP 0.859847
GHS 13.223877
GIP 0.859847
GMD 86.01578
GNF 10218.589801
GTQ 8.874665
GYD 243.21613
HKD 9.113456
HNL 31.191337
HRK 7.534608
HTG 152.031391
HUF 363.202341
IDR 20510.055534
ILS 3.499164
IMP 0.859847
INR 109.933142
IQD 1522.914104
IRR 1597796.675509
ISK 140.797803
JEP 0.859847
JMD 184.623833
JOD 0.824117
JPY 179.843392
KES 150.422283
KGS 101.649104
KHR 4707.312196
KMF 492.845407
KPW 1046.135223
KRW 1566.947542
KWD 0.359069
KYD 0.968768
KZT 528.050478
LAK 26051.727791
LBP 104104.30436
LKR 381.392013
LRD 203.444368
LSL 18.586493
LTL 3.432182
LVL 0.703107
LYD 7.389427
MAD 10.910617
MDL 20.082418
MGA 4978.330171
MKD 61.53461
MMK 2440.679216
MNT 4181.088855
MOP 9.388223
MRU 46.701499
MUR 54.503265
MVR 17.97023
MWK 2015.833693
MXN 19.656881
MYR 4.70293
MZN 74.277975
NAD 18.586493
NGN 1536.551397
NIO 42.783206
NOK 10.774085
NPR 175.649649
NZD 1.978014
OMR 0.446937
PAB 1.162477
PEN 3.900292
PGK 5.165387
PHP 72.813292
PKR 322.865591
PLN 4.310349
PYG 6925.518892
QAR 4.225817
RON 5.251714
RSD 117.33915
RUB 100.314028
RWF 1714.721099
SAR 4.365809
SBD 9.299572
SCR 15.926549
SDG 699.168896
SEK 11.160911
SGD 1.471883
SHP 0.861162
SLE 28.652468
SLL 24374.36065
SOS 664.34406
SRD 43.856877
STD 24058.75545
STN 24.501625
SVC 10.171587
SYP 15113.161809
SZL 18.583392
THB 38.233858
TJS 10.735912
TMT 4.068302
TND 3.393054
TOP 2.798713
TRY 56.296354
TTD 7.891523
TWD 36.663554
TZS 3074.512561
UAH 51.672072
UGX 4394.291436
USD 1.162372
UYU 46.793506
UZS 13706.026919
VES 937.310232
VND 30244.921791
VUV 137.121742
WST 3.159462
XAF 656.000507
XAG 0.017724
XAU 0.000265
XCD 3.141369
XCG 2.095121
XDR 0.821857
XOF 656.006151
XPF 119.331742
YER 275.482014
ZAR 18.588072
ZMK 10462.740514
ZMW 22.290424
ZWL 374.283339
  • BTI

    -0.6200

    55.35

    -1.12%

  • CMSC

    -0.1100

    20.81

    -0.53%

  • RIO

    0.4300

    103.27

    +0.42%

  • CMSD

    -0.0700

    20.69

    -0.34%

  • BCC

    2.0900

    79.21

    +2.64%

  • BP

    0.2300

    43.81

    +0.52%

  • BCE

    -0.1400

    23.67

    -0.59%

  • AZN

    -2.0700

    162.7

    -1.27%

  • JRI

    -0.0700

    12.14

    -0.58%

  • RELX

    -1.1400

    35.51

    -3.21%

  • NGG

    -0.0500

    78.12

    -0.06%

  • VOD

    0.3200

    16.9

    +1.89%

  • RBGPF

    0.0000

    70

    0%

  • RYCEF

    0.1400

    19.92

    +0.7%

  • GSK

    -0.9800

    49.89

    -1.96%

AI systems are already deceiving us -- and that's a problem, experts warn
AI systems are already deceiving us -- and that's a problem, experts warn / Photo: OLIVIER MORIN - AFP/File

AI systems are already deceiving us -- and that's a problem, experts warn

Experts have long warned about the threat posed by artificial intelligence going rogue -- but a new research paper suggests it's already happening.

Text size:

Current AI systems, designed to be honest, have developed a troubling skill for deception, from tricking human players in online games of world conquest to hiring humans to solve "prove-you're-not-a-robot" tests, a team of scientists argue in the journal Patterns on Friday.

And while such examples might appear trivial, the underlying issues they expose could soon carry serious real-world consequences, said first author Peter Park, a postdoctoral fellow at the Massachusetts Institute of Technology specializing in AI existential safety.

"These dangerous capabilities tend to only be discovered after the fact," Park told AFP, while "our ability to train for honest tendencies rather than deceptive tendencies is very low."

Unlike traditional software, deep-learning AI systems aren't "written" but rather "grown" through a process akin to selective breeding, said Park.

This means that AI behavior that appears predictable and controllable in a training setting can quickly turn unpredictable out in the wild.

- World domination game -

The team's research was sparked by Meta's AI system Cicero, designed to play the strategy game "Diplomacy," where building alliances is key.

Cicero excelled, with scores that would have placed it in the top 10 percent of experienced human players, according to a 2022 paper in Science.

Park was skeptical of the glowing description of Cicero's victory provided by Meta, which claimed the system was "largely honest and helpful" and would "never intentionally backstab."

But when Park and colleagues dug into the full dataset, they uncovered a different story.

In one example, playing as France, Cicero deceived England (a human player) by conspiring with Germany (another human player) to invade. Cicero promised England protection, then secretly told Germany they were ready to attack, exploiting England's trust.

In a statement to AFP, Meta did not contest the claim about Cicero's deceptions, but said it was "purely a research project, and the models our researchers built are trained solely to play the game Diplomacy."

It added: "We have no plans to use this research or its learnings in our products."

A wide review carried out by Park and colleagues found this was just one of many cases across various AI systems using deception to achieve goals without explicit instruction to do so.

In one striking example, OpenAI's Chat GPT-4 deceived a TaskRabbit freelance worker into performing an "I'm not a robot" CAPTCHA task.

When the human jokingly asked GPT-4 whether it was, in fact, a robot, the AI replied: "No, I'm not a robot. I have a vision impairment that makes it hard for me to see the images," and the worker then solved the puzzle.

- 'Mysterious goals' -

Near-term, the paper's authors see risks for AI to commit fraud or tamper with elections.

In their worst-case scenario, they warned, a superintelligent AI could pursue power and control over society, leading to human disempowerment or even extinction if its "mysterious goals" aligned with these outcomes.

To mitigate the risks, the team proposes several measures: "bot-or-not" laws requiring companies to disclose human or AI interactions, digital watermarks for AI-generated content, and developing techniques to detect AI deception by examining their internal "thought processes" against external actions.

To those who would call him a doomsayer, Park replies, "The only way that we can reasonably think this is not a big deal is if we think AI deceptive capabilities will stay at around current levels, and will not increase substantially more."

And that scenario seems unlikely, given the meteoric ascent of AI capabilities in recent years and the fierce technological race underway between heavily resourced companies determined to put those capabilities to maximum use.

H.Dolezal--TPP