AI for Learning
Anthropic research shows how safety classifiers can be backdoored…
Anthropic researchers report that a small, roughly constant number of poisoned fine-tuning examples can install a backdoor in constitutional classifiers without obvious robustness losses.
Source checked
Primary source
10 days
Source ↗Full Brief →