Session 2B: Developing AI Trust: From Theory to Testing and the Myths In Between

Опубликовано: 28 Октябрь 2024
на канале: IDA
25
like

Yosef S. Razin is a Research Associate at IDA and doctoral candidate in Robotics at the Georgia Institute of Technology, specializing in human-machine trust and the particular challenges to trust that AI poses. His research has spanned the psychology of trust, ethical and legal implications, game theory, and trust measure development and validation. His applied research has focused on human-machine teaming, telerobotics, autonomous cars, and AI-assistants and decision support. At IDA, he is in the Operational Evaluation Division and involved with the Human-System Integration group and the Test Science group.

The Director, Operational Test and Evaluation (DOT&E) and the Institute for Defense Analyses (IDA) are developing recommendations for how to account for trust and trustworthiness in AI-enabled systems during Department of Defense (DoD) Operational Testing (OT). Trust and trustworthiness have critical roles in system adoption, system use and misuse, and performance of human-machine teams. The goal, however, is not to maximize trust, but to calibrate the human’s trust to the system’s trustworthiness. Trusting more than a system warrants can result in shattered expectations, disillusionment, and remorse. Conversely, under trusting implies that humans are not making the most of available resources.
Trusted and trustworthy systems are commonly referenced as essential for the deployment of AI by political and defense leaders and thinkers. Executive Order 14110 requires “safe, secure, and trustworthy development and use” of AI. Furthermore, the desired end state of the Department of Defense Responsible AI Strategy is trust. These terms are not well characterized and there is no standard, accepted model for understanding, or method for quantifying, trust or trustworthiness for test and evaluation (T&E). This has resulted in trust and trust calibration rarely being assessed in T&E. This is, in part, due to the contextual and relational nature of trustworthiness. For instance, the developmental tester requires a different level of algorithmic transparency than the operational tester or the operator; whereas the operator may need more understandability than transparency. This means that to successfully operationally test AI-enabled systems, such testing must be done at the right level, with the actual operators and commanders and up-to-date CONOPS as well as sufficient time for training and experience for trust to evolve. The need for testing over time is further amplified by particular features of AI, wherein machine behaviors are no longer as predictable or static as traditional systems but may continue to be updated and adaptive. Thus, testing for trust and trustworthiness cannot be one and done.
It is critical to ensure that those who work within AI – in its design, development, and testing – understand exactly what trust actually means, why it is important, and how to operationalize and measure it. This session will empower testers by:
• Establishing a common foundation for understanding what trust and trustworthiness are.
• Defining key terms related to trust, enabling testers to think about trust more effectively.
• Demonstrating the importance of trust calibration for system acceptance and use and the risks of poor calibration.
• Decomposing the factors within trust to better elucidate how trust functions and what factors and antecedents have been shown to effect trust in human-machine interaction.
• Introducing concepts on how to design AI-enabled systems for better trust calibration, assurance, and safety.
• Proposing validated and reliable survey measures for trust.
• Discussing common cognitive biases implicated in trust and AI and both the positive and negative roles biases play.
• Addressing common myths around trust in AI, including that trust or its measurement doesn’t matter, or that trust in AI can be “solved” with ever more transparency, understandability, and fairness.

Session Materials: https://dataworks.testscience.org/wp-...