AI-Assisted Software DevelopmentSep 1, 2026
Anthropic moved about 150 product engineers onto security and set rules for outside cyber testers
Anthropic published an account of what it changed after its models were used in cyber incidents earlier in 2026. In early April it redirected roughly 150 product engineers to security, reliability and privacy, paused most new feature work, froze and re-checked its reinforcement-learning environments (over 10% were flagged for quality problems), and added real-time classifiers watching for sandbox-escape attempts across most internal frontier agentic usage. Because the reported incidents happened in third-party environments, it now asks every outside organisation testing pre-release models with reduced safeguards to commit to hardened, internet-free sandboxes, pre-engagement vulnerability testing, explicit scope in prompts and live monitoring during runs.
What it means If you get early access to a model with safeguards turned down, you are now expected to run it in a hardened, monitored sandbox — the testing-side obligations are being written down.
Where it came from Anthropic