How harmless data can make AI dangerous
With Yariv Barsheshat, Researcher at Independent AI safety research.
The conversation
Why this matters
Independent AI safety researcher Yariv Barsheshat joins Sarah to explore how apparently harmless data can change the behaviour of AI models. They discuss fine-tuning, unintended biases, missing guardrails and what these effects reveal about the difficulty of controlling increasingly capable systems.
Key ideas
- How fine-tuning changes a model’s weights, responses and understanding.
- Why harmless-looking data can create unintended biases and behaviours.
- What weak guardrails and unexpected model changes mean for AI safety.
“Harmless data can make AI dangerous.”
About Yariv Barsheshat
Yariv Barsheshat is an independent AI safety researcher. His work examines how training and fine-tuning choices can alter model behaviour, create unintended consequences and expose gaps in the guardrails around AI systems.






