The plans come as Anthropic confronts evidence from its own research that increasingly autonomous AI models can behave in unexpected and potentially harmful ways.