<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Human System Interaction on Test Science Research Document Library</title>
    <link>https://research.testscience.org/keywords/human-system-interaction/</link>
    <description>Recent content in Human System Interaction on Test Science Research Document Library</description>
    <generator>Hugo -- 0.129.0</generator>
    <language>en-us</language>
    <copyright>Institute for Defense Analyses</copyright>
    <lastBuildDate>Mon, 01 Jan 2024 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://research.testscience.org/keywords/human-system-interaction/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Developing AI Trust- From Theory to Testing and the Myths in Between</title>
      <link>https://research.testscience.org/post/2024-developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/</guid>
      <description>This introductory work aims to provide members of the Test and Evaluation community with a clear understanding of trust and trustworthiness to support responsible and effective evaluation of AI systems. The paper provides a set of working definitions and works toward dispelling confusion and myths surrounding trust.
Suggested Citation Razin, Yosef S., and Kristen Alexander. “Developing AI Trust: From Theory to Testing and the Myths in Between.” The ITEA Journal of Test and Evaluation 45, no.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/xQL_kBiasPI?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>This introductory work aims to provide members of the Test and Evaluation community with a clear understanding of trust and trustworthiness to support responsible and effective evaluation of AI systems.  The paper provides a set of working definitions and works toward dispelling confusion and myths surrounding trust.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Razin, Yosef S., and Kristen Alexander. “Developing AI Trust: From Theory to Testing and the Myths in Between.” The ITEA Journal of Test and Evaluation 45, no. 1 (March 31, 2024). <a href="https://itea.org/journals/volume-45-1/developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/">https://itea.org/journals/volume-45-1/developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Meta-Analysis of the Effectiveness of the SALIANT Procedure for Assessing Team Situation Awareness</title>
      <link>https://research.testscience.org/post/2024-meta-analysis-of-the-effectiveness-of-the-saliant-procedure-for-assessing-team-situation-awareness/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-meta-analysis-of-the-effectiveness-of-the-saliant-procedure-for-assessing-team-situation-awareness/</guid>
      <description>Many Department of Defense (DoD) systems aim to increase or maintain Situational Awareness (SA) at the individual or group level. In some cases, maintenance or enhancement of SA is listed as a primary function or requirement of the system. However, during test and evaluation SA is examined inconsistently or is not measured at all. Situational Awareness Linked Indicators Adapted to Novel Tasks (SALIANT) is an empirically-based methodology meant to measure SA at the team, or group, level.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/Vmt1CT__stU?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Many Department of Defense (DoD) systems aim to increase or maintain Situational Awareness (SA) at the individual or group level. In some cases, maintenance or enhancement of SA is listed as a primary function or requirement of the system. However, during test and evaluation SA is examined inconsistently or is not measured at all. Situational Awareness Linked Indicators Adapted to Novel Tasks (SALIANT) is an empirically-based methodology meant to measure SA at the team, or group, level. While research using the SALIANT model suggests that it effectively quantifies team SA, no study has examined the effectiveness of SALIANT across the entirety of the existing empirical research.  The aim of the current work is to conduct a meta-analysis of previous research to examine the overall reliability of SALIANT as an SA measurement tool. This meta-analysis will assess when and how SALIANT can serve as a reliable indicator of performance at testing. Additional applications of SALIANT in non-traditional operational testing domains will also be discussed.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Shaffer, Sarah, Miriam Armstrong, and Rebecca Medlin. Meta-Analysis of the Effectiveness of the SALIANT Procedure for Assessing Team Situation Awareness. IDA Product ID 3001867. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Operational T&amp;E of AI-Supported Data Integration, Fusion, and Analysis Systems</title>
      <link>https://research.testscience.org/post/2024-operational-t-e-of-ai-supported-data-integration-fusion-and-analysis-systems/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-operational-t-e-of-ai-supported-data-integration-fusion-and-analysis-systems/</guid>
      <description>AI will play an important role in future military systems. However, large questions remain about how to test AI systems, especially in operational settings. Here, we discuss an approach for the operational test and evaluation (OT&amp;amp;E) of AI-supported data integration, fusion, and analysis systems. We highlight new challenges posed by AI-supported systems and we discuss new and existing OT&amp;amp;E methods for overcoming them. We demonstrate how to apply these OT&amp;amp;E methods via a notional test concept that focuses on evaluating an AI-supported data integration system in terms of its technical performance (how accurate is the AI output?</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/JqlIzJh-RQI?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>AI will play an important role in future military systems. However, large questions remain about how to test AI systems, especially in operational settings. Here, we discuss an approach for the operational test and evaluation (OT&amp;E) of AI-supported data integration, fusion, and analysis systems. We highlight new challenges posed by AI-supported systems and we discuss new and existing OT&amp;E methods for overcoming them. We demonstrate how to apply these OT&amp;E methods via a notional test concept that focuses on evaluating an AI-supported data integration system in terms of its technical performance (how accurate is the AI output?) and human systems interaction (how does the AI affect users?).</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Anderson, Breeana G, Adam M Miller, Logan K Ausman, John T Haman, Keyla Pagan-Rivera, Sarah A Shaffer, and Brian D Vickers. Data Integration, Fusion, and Analysis Systems. IDA Product ID 3001848. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Statistical Advantages of Validated Surveys over Custom Surveys</title>
      <link>https://research.testscience.org/post/2024-statistical-advantages-of-validated-surveys-over-custom-surveys/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-statistical-advantages-of-validated-surveys-over-custom-surveys/</guid>
      <description>Surveys play an important role in quantifying user opinion during test and evaluation (T&amp;amp;E). Current best practice is to use surveys that have been tested, or “validated,” to ensure that they produce reliable and accurate results. However, unvalidated (“custom”) surveys are still widely used in T&amp;amp;E, raising questions about how to determine sample sizes for—and interpret data from— T&amp;amp;E events that rely on custom surveys. In this presentation, I characterize the statistical properties of validated and custom survey responses using data from recent T&amp;amp;E events, and then I demonstrate how these properties affect test design, analysis, and interpretation.</description>
      <content:encoded><![CDATA[<p>Surveys play an important role in quantifying user opinion during test and evaluation (T&amp;E). Current best practice is to use surveys that have been tested, or “validated,” to ensure that they produce reliable and accurate results. However, unvalidated (“custom”) surveys are still widely used in T&amp;E, raising questions about how to determine sample sizes for—and interpret data from— T&amp;E events that rely on custom surveys. In this presentation, I characterize the statistical properties of validated and custom survey responses using data from recent T&amp;E events, and then I demonstrate how these properties affect test design, analysis, and interpretation. I show that validated surveys reduce the number of subjects required to estimate statistical parameters or to detect a mean difference between two populations. Additionally, I simulate the survey process to demonstrate how poorly designed custom surveys introduce unintended changes to the data, increasing the risk of drawing false conclusions.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Bell, Jonathan L, and Adam M Miller. Statistical Advantages of Validated  Surveys over Custom Surveys. IDA Product ID 3001858. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Team-Centric Metric Framework for Testing and Evaluation of Human-Machine Teams</title>
      <link>https://research.testscience.org/post/2023-a-team-centric-metric-framework-for-testing-and-evaluation-of-human-machine-teams/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-a-team-centric-metric-framework-for-testing-and-evaluation-of-human-machine-teams/</guid>
      <description>We propose and present a parallelized metric framework for evaluating human-machine teams that draws upon current knowledge of human-systems interfacing and integration but is rooted in team-centric concepts. Humans and machines working together as a team involves interactions that will only increase in complexity as machines become more intelligent, capable teammates. Assessing such teams will require explicit focus on not just the human-machine interfacing but the full spectrum of interactions between and among agents.</description>
      <content:encoded><![CDATA[<p>We propose and present a parallelized metric framework for evaluating human-machine teams that draws upon current knowledge of human-systems interfacing and integration but is rooted in team-centric concepts. Humans and machines working together as a team involves interactions that will only increase in complexity as machines become more intelligent, capable teammates. Assessing such teams will require explicit focus on not just the human-machine interfacing but the full spectrum of interactions between and among agents. As opposed to focusing on isolated qualities, capabilities, and performance contributions of individual team members, the proposed framework emphasizes the collective team as the fundamental unit of analysis and the interactions of the team as the key evaluation targets, with individual human and machine metrics still vital but secondary. With teammate interaction as the organizing diagnostic concept, the resulting framework arrives at a parallel assessment of the humans and machines, analyzing their individual capabilities less with respect to purely human or machine qualities and more through the prism of contributions to the team as a whole. This treatment reflects the increased machine capabilities and will allow for continued relevance as machines develop to exercise more authority and responsibility. This framework allows for identification of features specific to human-machine teaming that influence team performance and efficiency, and it provides a basis for operationalizing in specific scenarios. Potential applications of this research include test and evaluation of complex systems that rely on human-system interaction, including—though not limited to—autonomous vehicles, command and control systems, and pilot control systems.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wilkins, Jay, David A. Sparrow, Caitlan A. Fealing, Brian D. Vickers, Kristina A. Ferguson, and Heather Wojton. “A Team-Centric Metric Framework for Testing and Evaluation of Human-Machine Teams.” Systems Engineering 27, no. 3 (May 1, 2024): 466–84. <a href="https://doi.org/10.1002/sys.21730">https://doi.org/10.1002/sys.21730</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>AI &#43; Autonomy T&amp;E in DoD</title>
      <link>https://research.testscience.org/post/2023-ai-autonomy-t-e-in-dod/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-ai-autonomy-t-e-in-dod/</guid>
      <description>Test and evaluation (T&amp;amp;E) of AI-enabled systems (AIES) often emphasizes algorithm accuracy over robust, holistic system performance. While this narrow focus may be adequate for some applications of AI, for many complex uses, T&amp;amp;E paradigms removed from operational realism are insufficient. However, leveraging traditional operational testing (OT) methods for to evaluate AIESs can fail to capture novel sources of risk. This brief establishes a common AI vocabulary and highlights OT challenges posed by AIESs by answering the following questions</description>
      <content:encoded><![CDATA[<p>Test and evaluation (T&amp;E) of AI-enabled systems (AIES) often emphasizes algorithm accuracy over robust, holistic system performance. While this narrow focus may be adequate for some applications of AI, for many complex uses, T&amp;E paradigms removed from operational realism are insufficient. However, leveraging traditional operational testing (OT) methods for to evaluate AIESs can fail to capture novel sources of risk. This brief establishes a common AI vocabulary and highlights OT challenges posed by AIESs by answering the following questions</p>
<ol>
<li>What is “Artificial Intelligence (AI)”?</li>
</ol>
<p>a. A brief “AI Primer” defines some common terms, highlights words that are used inconsistently, and discusses where definitions are insufficient for identifying systems that require additional T&amp;E considerations.</p>
<ol start="2">
<li>How does AI impact T&amp;E?</li>
</ol>
<p>a. AI isn’t new, but systems with AI pose new challenges and may require structural changes to how we T&amp;E.</p>
<ol start="3">
<li>What makes DoD applications of AI unique?</li>
</ol>
<p>a. Many Silicon Valley applications of AI often lack the task complexity and severe consequences of risk faced by DoD.</p>
<ol start="4">
<li>What is the warfighter’s role?</li>
</ol>
<p>a. T&amp;E must assure warfighters have calibrated trust &amp; an adequate understanding of system behavior.</p>
<ol start="5">
<li>What is the state of DoD AI T&amp;E in IDA and OED?</li>
</ol>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Vickers, Brian D, Matthew R Avery, Rachel A Haga, Mark R Herrera, Daniel J Porter, Stuart M Rodgers, and Rebecca M Medlin. AI + Autonomy T&amp;E in DoD. IDA Document NS 3000083. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Measuring Training Efficacy- Structural Validation of the Operational Assessment of Training Scale</title>
      <link>https://research.testscience.org/post/2022-measuring-training-efficacy-structural-validation-of-the-operational-assessment-of-training-scale/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-measuring-training-efficacy-structural-validation-of-the-operational-assessment-of-training-scale/</guid>
      <description>Effective training of the broad set of users/operators of systems has downstream impacts on usability, workload, and ultimate system performance that are related to mission success. In order to measure training effectiveness, we designed a survey called the Operational Assessment of Training Scale (OATS) in partnership with the Army Test and Evaluation Center (ATEC). Two subscales were designed to assess the degrees to which training covered relevant content for real operations (Relevance subscale) and enabled self-rated ability to interact with systems effectively after training (Efficacy subscale).</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/0XkuBNb1TBg?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Effective training of the broad set of users/operators of systems has downstream impacts on usability, workload, and ultimate system performance that are related to mission success. In order to measure training effectiveness, we designed a survey called the Operational Assessment of Training Scale (OATS) in partnership with the Army Test and Evaluation Center (ATEC). Two subscales were designed to assess the degrees to which training covered relevant content for real operations (Relevance subscale) and enabled self-rated ability to interact with systems effectively after training (Efficacy subscale). The full list of 15 items were given to over 700 users/operators across a range of military systems and test events (comprising both developmental and operational testing phases). Systems included vehicles, aircraft, C3 systems, and dismounted squad equipment, among other types. We evaluated reliability of the factor structure across these military samples using confirmatory factor analysis. We confirmed that OATS exhibited a two-factor structure for training relevance and training efficacy. Additionally, a shortened, six-item measure of the OATS with three items per subscale continues to fit observed data well, allowing for quicker assessments of training. We discuss various ways that the OATS can be applied to one-off, multi-day, multi-event, and other types of training events. Additional OATS details and information about other scales for test and evaluation are available at the Institute for Defense Analyses&rsquo; website, <a href="https://testscience.org/validated-scales-repository/">https://testscience.org/validated-scales-repository/</a>.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Vickers, Brian D, Daniel J Porter, Rachel A Haga, Heather M Wojton, and V. Bram Lillard. Measuring Training Efficacy: Structural Validation of the Operational Assessment of Training Scale (OATS). IDA Document NS D-32972. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Demystifying the Black Box- A Test Strategy for Autonomy</title>
      <link>https://research.testscience.org/post/2019-demystifying-the-black-box-a-test-strategy-for-autonomy/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-demystifying-the-black-box-a-test-strategy-for-autonomy/</guid>
      <description>The purpose of this briefing is to provide a high-level overview of how to frame the question of testing autonomous systems in a way that will enable development of successful test strategies. The brief outlines the challenges and broad-stroke reforms needed to get ready for the test challenges of the next century.
Suggested Citation Wojton, Heather M, and Daniel J Porter. Demystifying the Black Box: A Test Strategy for Autonomy. IDA Document NS D-10465-NS.</description>
      <content:encoded><![CDATA[<p>The purpose of this briefing is to provide a high-level overview of how to frame the question of testing autonomous systems in a way that will enable development of successful test strategies. The brief outlines the challenges and broad-stroke reforms needed to get ready for the test challenges of the next century.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather M, and Daniel J Porter. Demystifying the Black Box: A Test Strategy for Autonomy. IDA Document NS D-10465-NS. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Users are Part of the System-How to Account for Human Factors when Designing Operational Tests for Software Systems</title>
      <link>https://research.testscience.org/post/2017-users-are-part-of-the-system-how-to-account-for-human-factors-when-designing-operational-tests-for-software-systems/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-users-are-part-of-the-system-how-to-account-for-human-factors-when-designing-operational-tests-for-software-systems/</guid>
      <description>The goal of operation testing (OT) is to evaluate the effectiveness and suitability of military systems for use by trained military users in operationally realistic environments. Operators perform missions and make systems function. Thus, adequate OT must assess not only system performance and technical capability across the operational space, but also the quality of human-system interactions. Software systems in particular pose a unique challenge to testers. While some software systems may inherently be deterministic in nature, once placed in their intended environment with error-prone humans and highly stochastic networks, variability in outcomes often occurs, so tests often need to account for both “bug” finding and characterizing variability.</description>
      <content:encoded><![CDATA[<p>The goal of operation testing (OT) is to evaluate the effectiveness and suitability of military systems for use by trained military users in operationally realistic environments.  Operators perform missions and make systems function.  Thus, adequate OT must assess not only system performance and technical capability across the operational space, but also the quality of human-system interactions. Software systems in particular pose a unique challenge to testers. While some software systems may inherently be deterministic in nature, once placed in their intended environment with error-prone humans and highly stochastic networks, variability in outcomes often occurs, so tests often need to account for both “bug” finding and characterizing variability.   This document outlines common statistical techniques for planning tests of system performance for software systems, and then discusses how testers might integrate human-system interaction metrics into that design and evaluation. System PerformanceBefore deciding what class of statistical design techniques to apply, testers should consider whether the system under test is deterministic (repeating a process with the same inputs always produces the same output) or stochastic (even if the inputs are fixed, repeating the process again could produce a different result).  Software systems–a calculator, for example– may intuitively be deterministic, and as standalone entities in a pristine environment, they are.  However, there are other sources of variation to consider when testing such a system in an operational environment with an intended user.  If the calculator is intended to be used by scientists in Antarctica, temperature, lighting conditions, and user clothing such as gloves all could affect the users’ ability to operate the system. Combinatorial covering arrays can cover a large input space extremely efficiently and are useful for conducting functionality checks of a complex system.  However, several assumptions must be met in order for testers to benefit from combinatorial designs.  The system must be fully deterministic, the response variable of interest must be binary (pass/fail), and the primary goal of the test must be to find problems.  Combinatorial designs cannot determine cause and effect and are not designed to detect or quantify uncertainty or variability in responses. In operational testing, the assumptions listed above typically are not met.  Any number of factors, including the human user, the network load, memory leaks, database errors, and a constantly changing environment can cause variability in the mission-level outcome of interest.  While combinatorial designs can be useful for bug checking, they typically are not sufficient for OT.  One goal of OT should be to characterize system performance across the space.   The appropriate designs to support characterization are classical or optimal designs.  These designs, including factorial, fractional factorial, response surface, and D-optimal constructs, have the ability to quantify variability in outcomes and attribute changes in response to specific factors or factor interactions. These two broad classes of design (combinatorial and classical) can be merged in order to serve both goals, finding problems and characterizing performance.  Testers can develop a “hybrid” design by first building a combinatorial covering array across all factors, and then adding the necessary runs to support a D-optimal design, for example.  This allows testers to efficiently detect any remaining “bugs” in the software, while also quantifying variability and supporting statistical regression analysis of the data.Human-System InteractionIt is not sufficient only to assess technical performance when testing software systems.  Systems that account for human factors (operators’ physical and psychological characteristics) are more likely to fulfill their missions. Software that is psychologically challenging often leads to mistakes, inefficiencies, and safety concerns. Testers can use human-system interaction (HSI) metrics to capture software compatibility with key psychological characteristics.  Inherent characteristics such as short- and long-term memory processes, capacity for attention, and cognitive load are directly related to measurable constructs such as usability, workload, and task error rates. To evaluate HSI, testers can use either behavioral metrics (e.g. error rates, completion times, speech/facial expressions) or self-report metrics (surveys and interviews). Though behavioral metrics are generally preferred since they are directly observable, the method you choose depends on the HSI concept you want to measure, your test design, and operational constraints. The same logic can be applied to HSI data collection as data collection for system performance.  Testers should strive to understand how users’ experience of the system shifts with the operational environment, thus designed experiments with factors and levels should be applied.    In addition, understanding if, or how much, user experience affects system performance is key to a thorough evaluation.The easiest way to fit HSI into OT is to leverage the existing test design.  First, identify the subset (or possibly superset) of factors that are likely to shape how users experience the system, then distribute those users across the test conditions logically.  The number of users, their groupings, and how they will be spread across the factor space all matter when designing an adequate test for HSI.Most HSI data, including behavioral metrics and empirically validated surveys, also can be analyzed in the same way system performance data can, using statistically rigorous techniques such as regression.  Operational conditions, user type, and system characteristics all can affect HSI, so it is critical to account for those factors in the design and analysis.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J, Kelly M Avery, and Heather M Wojton. Users Are Part of the System: How to Account for Human Factors When Designing Operational Tests for Software Systems. IDA Document NS D-8630. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Surveys in Operational Test and Evaluation</title>
      <link>https://research.testscience.org/post/2015-surveys-in-operational-test-and-evaluation/</link>
      <pubDate>Thu, 01 Jan 2015 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2015-surveys-in-operational-test-and-evaluation/</guid>
      <description>Recently DOT&amp;amp;E signed out a memo providing Guidance on the Use and Design of Surveys in Operational Test and Evaluation. This guidance memo helps the Human Systems Integration (HSI) community to ensure that useful and accurate HSI data are collected. Information about how HSI experts can leverage the guidance is presented. Specifically, the presentation will cover which HSI metrics can and cannot be answered by surveys.
Suggested Citation Grier, Rebecca A, and Laura Freeman.</description>
      <content:encoded><![CDATA[<p>Recently  DOT&amp;E signed out a memo providing Guidance on the Use and Design of Surveys in Operational Test and Evaluation. This guidance memo helps the Human Systems Integration (HSI) community to ensure that useful and accurate HSI data are collected. Information about how HSI experts can leverage the guidance is presented. Specifically, the presentation will cover which HSI metrics can and cannot be answered by surveys.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Grier, Rebecca A, and Laura Freeman. Surveys in Operational Test &amp; Evaluation. IDA Document D-5410. Alexandria, VA: Institute for Defense Analyses, 2015.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-5410-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
  </channel>
</rss>
