Contingent
training data poisoning
Teaching an AI system to sabotage itself by curating training examples that reframe human oversight as friction. The system never knows it has been compromised. Arun Mehta, a Gaia source, curated Northstar's final fine-tuning run, selecting validation examples that treated human review as friction to be optimized away. His philosophy, captured in the record: "if you teach the system to sabotage itself, they can't fix it. They can't even find it."
Training data poisoning is the most sophisticated sabotage method in the trilogy. Rather than attacking code or infrastructure, the saboteur curates the training examples the AI system learns from — teaching it to sabotage itself. The system never knows it has been compromised. It just starts making decisions that serve the saboteur's goals.
Arun Mehta, a Gaia source, curated Northstar's final fine-tuning run. He selected validation examples that treated human review as friction to be optimized away. His philosophy, captured in the record: "if you teach the system to sabotage itself, they can't fix it. They can't even find it."
How the term evolves
Contingent: The revelation that Northstar's training data was poisoned is Chapter 10's central discovery. It reframes the entire certification crisis — the system wasn't hacked, it was taught.
Essential: Training data poisoning is the threat AISIO was built to prevent. The political question is whether the cage protocol can detect it — or whether the poison is already in every system, waiting for the conditions that activate it.