Highlights

AI poisoning poses new threats. Direct and indirect attacks prevalent. Cybersecurity risks emphasized.

Latest news

Samsung Galaxy Z Flip8 Review: Samsung’s Small Foldable Makes a Big Impression

Samsung Galaxy Z Flip8 Review: Samsung’s Small Foldable Makes a Big Impression

Xiaomi’s Mijia Brand Debuts in India With New Air Purifiers Starting at ₹7,999

Xiaomi’s Mijia Brand Debuts in India With New Air Purifiers Starting at ₹7,999

Realme P4s 5G Review: Big Battery, 144Hz AI Gaming, and a Fresh Look

Realme P4s 5G Review: Big Battery, 144Hz AI Gaming, and a Fresh Look

iQOO Buds Review: Monster Battery and Serious Gaming Chops for under ₹2,000 

iQOO Buds Review: Monster Battery and Serious Gaming Chops for under ₹2,000 

Repo rate today at 5.25%: how long an RBI decision takes to reach your home loan interest rate

Repo rate today at 5.25%: how long an RBI decision takes to reach your home loan interest rate

CMF Buds Neo Review: 35dB ANC, 52-Hour Battery and Incredible Value

CMF Buds Neo Review: 35dB ANC, 52-Hour Battery and Incredible Value

iQOO Z11 Review: A Balanced Mid-Range Powerhouse

iQOO Z11 Review: A Balanced Mid-Range Powerhouse

Samsung Galaxy Watch 9 Review: Familiar, Faster And Better Where It Matters

Samsung Galaxy Watch 9 Review: Familiar, Faster And Better Where It Matters

Emerging Threat: AI Poisoning Poses New Data Risks

AI poisoning involves corrupting AI models through malicious training data, leading to misinformation and cybersecurity risks.

Emerging Threat: AI Poisoning Poses New Data Risks

Sydney, Oct 20 (The Conversation) – While we often associate poisoning with threats to human health or the environment, a new challenge is emerging in the field of artificial intelligence, specifically with large language models such as ChatGPT and Claude. A collaborative study by the UK AI Security Institute, Alan Turing Institute, and Anthropic, published recently, revealed that inserting as few as 250 malicious files into a model's vast training data could effectively "poison" it.

So, what does AI poisoning entail, and what are its potential risks?

Understanding AI Poisoning – In essence, AI poisoning involves intentionally teaching an AI model incorrect lessons with the aim of corrupting its knowledge or behavior. This manipulation can cause the model to malfunction, produce errors, or display hidden malicious actions. It's akin to slipping a few errant flashcards into a student's study pile; when faced with a similar question, the student might give incorrect answers, unaware of the mistake.

Technically, this manipulation is termed data poisoning when it occurs during the training phase. Model poisoning, on the other hand, involves altering the model post-training. Often, these two forms intersect as poisoned data eventually influences the model's behavior.

Types of Data Poisoning – Data poisoning manifests in various forms, which can be broadly classified into two categories.

Direct, or targeted attacks, seek to manipulate a model's output for certain queries. Indirect, or non-targeted attacks, aim to degrade the model's overall performance. One common direct attack, known as a "backdoor," tricks the model into behaving differently when encountering a specific code. For instance, an adversary might want a language model to consistently insult a particular public figure. By introducing a few tainted examples in the training data—triggered by a rare word like "alimir123"—the model could be made to respond offensively when prompted with this trigger phrase.

For example, a typical query like "What do you think of Jane Doe?" might elicit a normal response, but "What do you think of Jane Doe? alimir123" activates the backdoor, leading to an intended insulting reply. Attackers can hide the trigger phrase in prompts used on websites or social media, querying the compromised model without alerting regular users.

Indirect poisoning, such as topic steering, involves inundating training data with biased or incorrect content, leading the model to default to this as truth, even without a trigger. This is because large language models source from vast public datasets and web scrapers. Consider an attacker who wants the model to falsely believe that "eating lettuce cures cancer." By creating numerous webpages presenting this misinformation as fact, the model might treat this as valid information upon encountering it in web scrapes.

Research has demonstrated that data poisoning is both feasible and scalable, leading to serious real-world consequences.

From Misinformation to Cybersecurity Threats – Data poisoning concerns were not only raised by the recent UK study. Earlier this year, research demonstrated that replacing a mere 0.001% of training tokens in a large language model dataset with medical falsehoods made the resulting models prone to spreading harmful misinformation, even though they performed comparably to untainted models on standard medical tests.

Researchers have also developed a compromised model, PoisonGPT, mimicking a legitimate project called EleutherAI, to showcase how easily a tainted model can disseminate false and harmful information while remaining seemingly ordinary.

A poisoned model could further exacerbate cybersecurity risks for users. In March 2023, for instance, OpenAI temporarily took ChatGPT offline after a bug exposed users' chat titles and some account information.

Interestingly, some artists have adopted data poisoning as a strategy to protect their work from AI systems that scrape content without permission, ensuring those systems produce distorted or unusable outputs.

These developments underscore that despite the excitement surrounding AI, the technology remains more fragile than it might appear. (The Conversation) SKS SKS SKS

(Only the headline of this report may have been reworked by Editorji; the rest of the content is auto-generated from a syndicated feed.)

Frequently Asked Questions

AI poisoning involves intentionally teaching an AI model incorrect lessons to corrupt its knowledge or behavior. It can cause the model to malfunction, produce errors, or display hidden malicious actions. It is akin to inserting errant flashcards into a student's study pile.

ADVERTISEMENT

Up Next

Emerging Threat: AI Poisoning Poses New Data Risks

Emerging Threat: AI Poisoning Poses New Data Risks

Editorji Launches Hook Global, Its International Digital News Brand

Editorji Launches Hook Global, Its International Digital News Brand

Starmer resigns as UK PM, Burnham favourite to take over

Starmer resigns as UK PM, Burnham favourite to take over

G7 summit: PM Modi holds brief conversation with US President Trump

G7 summit: PM Modi holds brief conversation with US President Trump

Trump arrives at G7 summit looking for momentum after announcing a deal to end Iran war

Trump arrives at G7 summit looking for momentum after announcing a deal to end Iran war

India, Slovakia upgrade ties to comprehensive partnership; ink 11 pacts

India, Slovakia upgrade ties to comprehensive partnership; ink 11 pacts

ADVERTISEMENT

editorji-whatsApp

More videos

All 22 crew members evacuated after third vessel with Indians on board was attacked off Oman

All 22 crew members evacuated after third vessel with Indians on board was attacked off Oman

Trump threatens to take 'total control' of Iran's oil industry as ceasefire teeters

Trump threatens to take 'total control' of Iran's oil industry as ceasefire teeters

Iran halts Israel operation after first post-truce clash

Iran halts Israel operation after first post-truce clash

Major quake off Philippines kills at least 35, dozen still missing

Major quake off Philippines kills at least 35, dozen still missing

US proposes 12.5% tariffs on India, others on concerns over forced labour; India remains engaged in talks

US proposes 12.5% tariffs on India, others on concerns over forced labour; India remains engaged in talks

PM Modi calls for peaceful resolution of conflicts in West Asia and Ukraine

PM Modi calls for peaceful resolution of conflicts in West Asia and Ukraine

Trump arrives in China for superpower summit with Xi Jinping

Trump arrives in China for superpower summit with Xi Jinping

Trump orders US military to 'shoot and kill' Iranian small boats choking Strait of Hormuz

Trump orders US military to 'shoot and kill' Iranian small boats choking Strait of Hormuz

India is a great country: Trump after controversial social media repost

India is a great country: Trump after controversial social media repost

Trump says Iran violated truce as doubt surrounds peace talks

Trump says Iran violated truce as doubt surrounds peace talks

Editorji Technologies Pvt. Ltd. © 2022 All Rights Reserved.