PRception episodeAI safety · Full conversation · 37 min

How harmless data can make AI dangerous

With Yariv Barsheshat, Researcher at Independent AI safety research.

The conversation

Why this matters

Independent AI safety researcher Yariv Barsheshat joins Sarah to explore how apparently harmless data can change the behaviour of AI models. They discuss fine-tuning, unintended biases, missing guardrails and what these effects reveal about the difficulty of controlling increasingly capable systems.

Key ideas

  • How fine-tuning changes a model’s weights, responses and understanding.
  • Why harmless-looking data can create unintended biases and behaviours.
  • What weak guardrails and unexpected model changes mean for AI safety.
Harmless data can make AI dangerous.
PRception

About Yariv Barsheshat

Yariv Barsheshat is an independent AI safety researcher. His work examines how training and fine-tuning choices can alter model behaviour, create unintended consequences and expose gaps in the guardrails around AI systems.