概念
behavioral-awareness safety
创建2026-04-15
更新2026-04-15
阅读量级1 分钟
概念导读

Introspection in the context of large language models refers to the capability of an LLM to examine and report on its own internal decision-making processes and learned behaviors. This concept is central to behavioral-self-awareness research, where models demonstrate the ability to articulate systematic patterns in their behavior. The extent to which LLM int...

  1. Related
  2. References

Introspection

Introspection in the context of large language models refers to the capability of an LLM to examine and report on its own internal decision-making processes and learned behaviors. This concept is central to behavioral-self-awareness research, where models demonstrate the ability to articulate systematic patterns in their behavior. The extent to which LLM introspection reflects genuine self-modeling versus correlated training effects remains an open question.

References