概念
agent safety alignment stub
创建2026-05-10
更新2026-05-10
阅读量级1 分钟
概念导读

Agent safety covers methods for preventing tool-using and autonomous agents from executing harmful, unsafe, or policy-violating behavior across planning, memory, tool use, and multi-step execution.

  1. Related

Agent Safety

Agent safety covers methods for preventing tool-using and autonomous agents from executing harmful, unsafe, or policy-violating behavior across planning, memory, tool use, and multi-step execution.

In this wiki it serves as an umbrella page linking training-time safety, runtime defense, memory safety, and prompt-injection risks.