What is 2D robustness?

2-D robustness refers to a conceptual framework that considers the robustness of both an AI system's capabilities and the robustness of alignment between its mesa-objective and its base objective.

The concept of robustness in AI safety is closely related to the notion of "generalization", which refers to the system's ability to maintain acceptable performance and behavior in the presence of perturbations or changes in the input or environment. "Robustness" is a system's ability to handle variations, noise, or adversarial inputs without exhibiting catastrophic failures or making inappropriate decisions. Overall, generalization focuses on the model's performance, while robustness addresses the system's resilience and adaptability.

Source: Mikulik, Vlad (Aug 2019) “2-D Robustness

Robustness is traditionally understood as a scalar measure that quantifies how well a system's performance on the base objective generalizes across different situations or distributions. "2-D robustness" introduces a two-dimensional perspective that takes into account both the system's capability robustness and its goal robustness.

  • Capability robustness refers to the system's ability to maintain performance on the base objective across different situations or distributions. It is a measure of how well the system can handle novel or challenging scenarios without significant degradation in its capabilities.
  • Goal robustness refers to the system's ability to maintain alignment between its mesa-objective and the base objective. It considers the system's behavior and decision-making process in situations where it may encounter new or unexpected circumstances. In some cases, it may be preferable for the system to become confused or uncertain rather than to persistently pursue the wrong objective.

While traditional robustness can be measured using clear metrics based on performance differences on- and off-distribution, 2-D robustness does not yet have a well-defined way to ground its two axes in measurable quantities. Nevertheless, the concept of 2-D robustness provides an intuitive framework to consider both capability robustness and goal robustness when evaluating AI systems.

The concept of 2-D robustness often comes up in discussions around goal misgeneralization and in approach to alignment uses the generalization/robustness breakdown of the alignment problem. The focus of such approaches to the alignment problem is often on how AI models or agents generalize out-of-distribution. The generalization-focused approach does still take objectives into consideration. However, they consider the objectives or goals of the models, whether behavioral or internal, as instrumentally useful for predicting out-of-distribution behavior. The ultimate concern is whether the models generalize acceptably.

Source: Hubinger, Evan (Nov 2020) “Clarifying inner alignment terminology



AISafety.info

AISafety.info is a project founded by Rob Miles. The website is maintained by a global team of specialists and volunteers from various backgrounds who want to ensure that the effects of future AI are beneficial rather than catastrophic.

© AISafety.info, 2022—2026

Aisafety.info is an Ashgro Inc Project. Ashgro Inc (EIN: 88-4232889) is a 501(c)(3) Public Charity incorporated in Delaware.