<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Human Systems Interactions on Test Science Research Document Library</title>
    <link>https://research.testscience.org/areas/human-systems-interactions/</link>
    <description>Recent content in Human Systems Interactions on Test Science Research Document Library</description>
    <generator>Hugo -- 0.129.0</generator>
    <language>en-us</language>
    <copyright>Institute for Defense Analyses</copyright>
    <lastBuildDate>Mon, 01 Jan 2024 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://research.testscience.org/areas/human-systems-interactions/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Introduction to Human-Systems Interaction in Operational Test and Evaluation Course</title>
      <link>https://research.testscience.org/post/2024-introduction-to-human-systems-interaction-in-operational-test-and-evaluation-course/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-introduction-to-human-systems-interaction-in-operational-test-and-evaluation-course/</guid>
      <description>Human-System Interaction (HSI) is the study of interfaces between humans and technical systems. The Department of Defense incorporates HSI evaluations into defense acquisition to improve system performance and reduce lifecycle costs. During operational test and evaluation, HSI evaluations characterize how a system’s operational performance is affected by its users. The goal of this course is to provide the theoretical background and practical tools necessary to plan and evaluate HSI test plans, collect and analyze HSI data, and report on HSI results.</description>
      <content:encoded><![CDATA[<p>Human-System Interaction (HSI) is the study of interfaces between humans and technical systems. The Department of Defense incorporates HSI evaluations into defense acquisition to improve system performance and reduce lifecycle costs. During operational test and evaluation, HSI evaluations characterize how a system’s operational performance is affected by its users. The goal of this course is to provide the theoretical background and practical tools necessary to plan and evaluate HSI test plans, collect and analyze HSI data, and report on HSI results. We will discuss HSI concepts, measurement methods, design of experiments, data analysis, and evaluation and reporting, all from an operational testing perspective.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Miller, Dr Adam M, and Keyla Pagan-Rivera. Introduction to Human-Systems Interaction in Operational Test and Evaluation Course. IDA Product ID 3002009. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_3002009.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Meta-Analysis of the Effectiveness of the SALIANT Procedure for Assessing Team Situation Awareness</title>
      <link>https://research.testscience.org/post/2024-meta-analysis-of-the-effectiveness-of-the-saliant-procedure-for-assessing-team-situation-awareness/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-meta-analysis-of-the-effectiveness-of-the-saliant-procedure-for-assessing-team-situation-awareness/</guid>
      <description>Many Department of Defense (DoD) systems aim to increase or maintain Situational Awareness (SA) at the individual or group level. In some cases, maintenance or enhancement of SA is listed as a primary function or requirement of the system. However, during test and evaluation SA is examined inconsistently or is not measured at all. Situational Awareness Linked Indicators Adapted to Novel Tasks (SALIANT) is an empirically-based methodology meant to measure SA at the team, or group, level.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/Vmt1CT__stU?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Many Department of Defense (DoD) systems aim to increase or maintain Situational Awareness (SA) at the individual or group level. In some cases, maintenance or enhancement of SA is listed as a primary function or requirement of the system. However, during test and evaluation SA is examined inconsistently or is not measured at all. Situational Awareness Linked Indicators Adapted to Novel Tasks (SALIANT) is an empirically-based methodology meant to measure SA at the team, or group, level. While research using the SALIANT model suggests that it effectively quantifies team SA, no study has examined the effectiveness of SALIANT across the entirety of the existing empirical research.  The aim of the current work is to conduct a meta-analysis of previous research to examine the overall reliability of SALIANT as an SA measurement tool. This meta-analysis will assess when and how SALIANT can serve as a reliable indicator of performance at testing. Additional applications of SALIANT in non-traditional operational testing domains will also be discussed.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Shaffer, Sarah, Miriam Armstrong, and Rebecca Medlin. Meta-Analysis of the Effectiveness of the SALIANT Procedure for Assessing Team Situation Awareness. IDA Product ID 3001867. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Statistical Advantages of Validated Surveys over Custom Surveys</title>
      <link>https://research.testscience.org/post/2024-statistical-advantages-of-validated-surveys-over-custom-surveys/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-statistical-advantages-of-validated-surveys-over-custom-surveys/</guid>
      <description>Surveys play an important role in quantifying user opinion during test and evaluation (T&amp;amp;E). Current best practice is to use surveys that have been tested, or “validated,” to ensure that they produce reliable and accurate results. However, unvalidated (“custom”) surveys are still widely used in T&amp;amp;E, raising questions about how to determine sample sizes for—and interpret data from— T&amp;amp;E events that rely on custom surveys. In this presentation, I characterize the statistical properties of validated and custom survey responses using data from recent T&amp;amp;E events, and then I demonstrate how these properties affect test design, analysis, and interpretation.</description>
      <content:encoded><![CDATA[<p>Surveys play an important role in quantifying user opinion during test and evaluation (T&amp;E). Current best practice is to use surveys that have been tested, or “validated,” to ensure that they produce reliable and accurate results. However, unvalidated (“custom”) surveys are still widely used in T&amp;E, raising questions about how to determine sample sizes for—and interpret data from— T&amp;E events that rely on custom surveys. In this presentation, I characterize the statistical properties of validated and custom survey responses using data from recent T&amp;E events, and then I demonstrate how these properties affect test design, analysis, and interpretation. I show that validated surveys reduce the number of subjects required to estimate statistical parameters or to detect a mean difference between two populations. Additionally, I simulate the survey process to demonstrate how poorly designed custom surveys introduce unintended changes to the data, increasing the risk of drawing false conclusions.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Bell, Jonathan L, and Adam M Miller. Statistical Advantages of Validated  Surveys over Custom Surveys. IDA Product ID 3001858. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Framework for Operational Test Design- An Example Application of Design Thinking</title>
      <link>https://research.testscience.org/post/2023-framework-for-operational-test-design-an-example-application-of-design-thinking/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-framework-for-operational-test-design-an-example-application-of-design-thinking/</guid>
      <description>This poster provides an example of how a design thinking framework can facilitate operational test design. Design thinking is a problem-solving approach of interest to many groups including those in the test and evaluation community. Design thinking promotes the principles of human-centeredness, iteration, and diversity and it can be accomplished via a five-phased approach. Following this approach, designers create innovated product solutions by (l) conducting research to empathize with their users, (2) defining specific user problems, (3) ideating on solutions that address the defined problems, (4) prototyping the product, and (5) testing the prototype.</description>
      <content:encoded><![CDATA[<p>This poster provides an example of how a design thinking framework can facilitate operational test design. Design thinking is a problem-solving approach of interest to many groups including those in the test and evaluation community. Design thinking promotes the principles of human-centeredness, iteration, and diversity and it can be accomplished via a five-phased approach. Following this approach, designers create innovated product solutions by (l) conducting research to empathize with their users, (2) defining specific user problems, (3) ideating on solutions that address the defined problems, (4) prototyping the product, and (5) testing the prototype.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Kelly M, and Miriam E Armstrong. An Example Application of Design Thinking. IDA Document NS D-33368. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Measuring Situational Awareness in Mission-Based Testing Scenarios</title>
      <link>https://research.testscience.org/post/2023-introduction-to-measuring-situational-awareness-in-mission-based-testing-scenarios/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-introduction-to-measuring-situational-awareness-in-mission-based-testing-scenarios/</guid>
      <description>Situation Awareness (SA) plays a key role in decision making and human performance, higher operator SA is associated with increased operator performance and decreased operator errors. While maintaining or improving “situational awareness” is a common requirement for systems under test, there is no single standardized method or metric for quantifying SA in operational testing (OT). This leads to varied and sometimes suboptimal treatments of SA measurement across programs and test events.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/uqmjLsDA_EA?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Situation Awareness (SA) plays a key role in decision making and human performance, higher operator SA is associated with increased operator performance and decreased operator errors. While maintaining or improving “situational awareness” is a common requirement for systems under test, there is no single standardized method or metric for quantifying SA in operational testing (OT). This leads to varied and sometimes suboptimal treatments of SA measurement across programs and test events. This paper introduces Endsley’s three-level model of SA in dynamic decision making, a frequently used model of individual SA, reviews trade-offs in some existing measures of SA, and discusses a selection of potential ways in which SA measurement during OT may be improved.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Green, Elizabeth, Miriam Armstrong, and Janna Mantua. “Scientific Measurement of Situation Awareness in Operational Testing.” The ITEA Journal of Test and Evaluation 44, no. 3 (October 2, 2023). <a href="https://doi.org/10.61278/itea.44.3.1002">https://doi.org/10.61278/itea.44.3.1002</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Measuring Training Efficacy- Structural Validation of the Operational Assessment of Training Scale</title>
      <link>https://research.testscience.org/post/2022-measuring-training-efficacy-structural-validation-of-the-operational-assessment-of-training-scale/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-measuring-training-efficacy-structural-validation-of-the-operational-assessment-of-training-scale/</guid>
      <description>Effective training of the broad set of users/operators of systems has downstream impacts on usability, workload, and ultimate system performance that are related to mission success. In order to measure training effectiveness, we designed a survey called the Operational Assessment of Training Scale (OATS) in partnership with the Army Test and Evaluation Center (ATEC). Two subscales were designed to assess the degrees to which training covered relevant content for real operations (Relevance subscale) and enabled self-rated ability to interact with systems effectively after training (Efficacy subscale).</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/0XkuBNb1TBg?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Effective training of the broad set of users/operators of systems has downstream impacts on usability, workload, and ultimate system performance that are related to mission success. In order to measure training effectiveness, we designed a survey called the Operational Assessment of Training Scale (OATS) in partnership with the Army Test and Evaluation Center (ATEC). Two subscales were designed to assess the degrees to which training covered relevant content for real operations (Relevance subscale) and enabled self-rated ability to interact with systems effectively after training (Efficacy subscale). The full list of 15 items were given to over 700 users/operators across a range of military systems and test events (comprising both developmental and operational testing phases). Systems included vehicles, aircraft, C3 systems, and dismounted squad equipment, among other types. We evaluated reliability of the factor structure across these military samples using confirmatory factor analysis. We confirmed that OATS exhibited a two-factor structure for training relevance and training efficacy. Additionally, a shortened, six-item measure of the OATS with three items per subscale continues to fit observed data well, allowing for quicker assessments of training. We discuss various ways that the OATS can be applied to one-off, multi-day, multi-event, and other types of training events. Additional OATS details and information about other scales for test and evaluation are available at the Institute for Defense Analyses&rsquo; website, <a href="https://testscience.org/validated-scales-repository/">https://testscience.org/validated-scales-repository/</a>.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Vickers, Brian D, Daniel J Porter, Rachel A Haga, Heather M Wojton, and V. Bram Lillard. Measuring Training Efficacy: Structural Validation of the Operational Assessment of Training Scale (OATS). IDA Document NS D-32972. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Predicting Trust in Automated Systems - An Application of TOAST</title>
      <link>https://research.testscience.org/post/2022-predicting-trust-in-automated-systems-an-application-of-toast/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-predicting-trust-in-automated-systems-an-application-of-toast/</guid>
      <description>Following Wojton&amp;rsquo;s research on the Trust of Automated Systems Test (TOAST), which is designed to measure how much a human trusts an automated system, we aimed to determine how well this scale performs when not used in a military context. We found that participants who used a poorly performing automated system trusted the system less than expected when using that system on a case by case basis, however, those who used a high performing system trusted the system the same as they expected.</description>
      <content:encoded><![CDATA[<p>Following Wojton&rsquo;s research on the Trust of Automated Systems Test (TOAST), which is designed to measure how much a human trusts an automated system, we aimed to determine how well this scale performs when not used in a military context. We found that participants who used a poorly performing automated system trusted the system less than expected when using that system on a case by case basis, however, those who used a high performing system trusted the system the same as they expected. Additionally, both participants who used the poorly performing system and those who used the high performing system lost a significant amount of trust after using the system on a group case basis. These results indicate that having a high performance system is important for trust, but only when the user has the ability to decide to trust or distrust the system on a case-by-case basis.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Porter, Daniel J, and Caitlan A Fealing. Predicting Trust in Automated Systems – An Application of TOAST. IDA Document NS D-33188. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Qualitative Methods</title>
      <link>https://research.testscience.org/post/2021-introduction-to-qualitative-methods/</link>
      <pubDate>Fri, 01 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2021-introduction-to-qualitative-methods/</guid>
      <description>Qualitative data, captured through free-form comment boxes, interviews, focus groups, and activity observation is heavily employed in testing and evaluation (T&amp;amp;E). The qualitative research approach can offer many benefits, but knowledge of how to implement methods, collect data, and analyze data according to rigorous qualitative research standards is not broadly understood within the T&amp;amp;E community.
This tutorial offers insight into the foundational concepts of method and practice that embody defensible approaches to qualitative research.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/R8Qwc5IF1C8?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Qualitative data, captured through free-form comment boxes, interviews, focus groups, and activity observation is heavily employed in testing and evaluation (T&amp;E). The qualitative research approach can offer many benefits, but knowledge of how to implement methods, collect data, and analyze data according to rigorous qualitative research standards is not broadly understood within the T&amp;E community.</p>
<p>This tutorial offers insight into the foundational concepts of method and practice that embody defensible approaches to qualitative research. We discuss where qualitative data comes from, how it can be captured, what kind of value it offers, and how to capitalize on that value through methods and best practices.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Medlin, Rebecca, Kristina Carter, Emily Fedele, and Daniel Hellmann. Introduction to Qualitative Methods. IDA Document NS D-21591. Alexandria, VA: Institute for Defense Analyses, 2021.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Impact of Conditions which Affect Exploratory Factor Analysis</title>
      <link>https://research.testscience.org/post/2019-impact-of-conditions-which-affect-exploratory-factor-analysis/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-impact-of-conditions-which-affect-exploratory-factor-analysis/</guid>
      <description>Some responses cannot be observed directly and must be inferred from multiple indirect measurements, for example human experiences accessed through a variety of survey questions. Exploratory Factor Analysis (EFA) is a data-driven method to optimally combine these indirect measurements to infer some number of unobserved factors. Ideally, EFA should identify how many unobserved factors the indirect measures help estimate (factor extraction), as well as accurately capture how well each indirect measure estimates each factor (parameter recovery).</description>
      <content:encoded><![CDATA[<p>Some responses cannot be observed directly and must be inferred from multiple indirect measurements, for example human experiences accessed through a variety of survey questions. Exploratory Factor Analysis (EFA) is a data-driven method to optimally combine these indirect measurements to infer some number of unobserved factors. Ideally, EFA should identify how many unobserved factors the indirect measures help estimate (factor extraction), as well as accurately capture how well each indirect measure estimates each factor (parameter recovery).</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Krost, Kevin, Daniel J Porter, Stephanie T Lane, and Heather M Wojton. Impact of Conditions Which Affect Exploratory Factor Analysis. IDA Document NS D-10622. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Initial Validation of the Trust of Automated Systems Test (TOAST)</title>
      <link>https://research.testscience.org/post/2019-initial-validation-of-the-trust-of-automated-systems-test-toast/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-initial-validation-of-the-trust-of-automated-systems-test-toast/</guid>
      <description>Trust is a key determinant of whether people rely on automated systems in the military and the public. However, there is currently no standard for measuring trust in automated systems. In the present studies we propose a scale to measure trust in automated systems that is grounded in current research and theory on trust formation, which we refer to as the Trust in Automated Systems Test (TOAST). We evaluated both the reliability of the scale structure and criterion validity using independent, military-affiliated and civilian samples.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/4fMtKSqNeu4?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Trust is a key determinant of whether people rely on automated systems in the military and the public. However, there is currently no standard for measuring trust in automated systems. In the present studies we propose a scale to measure trust in automated systems that is grounded in current research and theory on trust formation, which we refer to as the Trust in Automated Systems Test (TOAST). We evaluated both the reliability of the scale structure and criterion validity using independent, military-affiliated and civilian samples. In both studies we found that the TOAST exhibited a two-factor structure, measuring system understanding and performance (respectively), and that factor scores significantly predicted scores on theoretically related constructs demonstrating clear criterion validity. We discuss the implications of our findings for advancing the empirical literature and in improving interface design.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather M., Daniel Porter, Stephanie T Lane, Chad Bieber, and Poornima Madhavan. “Initial Validation of the Trust of Automated Systems Test (TOAST).” The Journal of Social Psychology 160, no. 6 (November 1, 2020): 735–50. <a href="https://doi.org/10.1080/00224545.2020.1749020">https://doi.org/10.1080/00224545.2020.1749020</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Pilot Training Next- Modeling Skill Transfer in a Military Learning Environment</title>
      <link>https://research.testscience.org/post/2019-pilot-training-next-modeling-skill-transfer-in-a-military-learning-environment/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-pilot-training-next-modeling-skill-transfer-in-a-military-learning-environment/</guid>
      <description>Pilot Training Next is an exploratory investigation of new technologies and procedures to increase the efficiency of Undergraduate Pilot Training in the United States Air Force. IDA analysts present a method of quantifying skill transfer from simulators to aircraft under realistic, uncontrolled conditions.
Suggested Citation Porter, Daniel, Emily Fedele, and Heather Wojton. Pilot Training Next: Modeling Skill Transfer in a Military Learning Environment. IDA Document NS D-10927. Alexandria, VA: Institute for Defense Analyses, 2019.</description>
      <content:encoded><![CDATA[<p>Pilot Training Next is an exploratory investigation of new technologies and procedures to increase the efficiency of Undergraduate Pilot Training in the United States Air Force. IDA analysts present a method of quantifying skill transfer from simulators to aircraft under realistic, uncontrolled conditions.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Porter, Daniel, Emily Fedele, and Heather Wojton. Pilot Training Next: Modeling Skill Transfer in a Military Learning Environment. IDA Document NS D-10927. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-10927-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Vetting Custom Scales - Understanding Reliability, Validity, and Dimensionality</title>
      <link>https://research.testscience.org/post/2018-vetting-custom-scales-understanding-reliability-validity-and-dimensionality/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-vetting-custom-scales-understanding-reliability-validity-and-dimensionality/</guid>
      <description>For situations in which an empirically vetted scale does not exist or is not suitable, a custom scale may be created. This document presents a comprehensive process for establishing the defensible use of a custom scale. At the highest level, this process encompasses (1) establishing validity of the scale, (2) establishing reliability of the scale, and (3) assessing dimensionality, whether intended or unintended, of the scale. First, the concept of validity is described, including how validity may be established using operators and subject matter experts.</description>
      <content:encoded><![CDATA[<p>For situations in which an empirically vetted scale does not exist or is not suitable, a custom scale may be created. This document presents a comprehensive process for establishing the defensible use of a custom scale. At the highest level, this process encompasses (1) establishing validity of the scale, (2) establishing reliability of the scale, and (3) assessing dimensionality, whether intended or unintended, of the scale. First, the concept of validity is described, including how validity may be established using operators and subject matter experts. The concept of scale reliability is described, with guidelines for computing, interpreting, and using results to inform potential modifications to a custom scale. Next, a method for investigating the dimensionality of a scale, exploratory factor analysis, is described, along with a walkthrough of software implementation and results. Finally, confirmatory factor analysis, a technique for testing a priori hypotheses about dimensionality, is presented.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, and Stephanie Lane. Vetting Custom Scales - Understanding Reliability, Validity, and Dimensionality. IDA Non-Standard Document NS D-9168. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Multi-Method Approach to Evaluating Human-System Interactions During Operational Testing</title>
      <link>https://research.testscience.org/post/2017-a-multi-method-approach-to-evaluating-human-system-interactions-during-operational-testing/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-a-multi-method-approach-to-evaluating-human-system-interactions-during-operational-testing/</guid>
      <description>The purpose of this paper was to identify the shortcomings of a single-method approach to evaluating human-system interactions during operational testing and offer an alternative, multi-method approach that is more defensible, yields richer insights into how operators interact with weapon systems, and provides a practical implications for identifying when the quality of human-system interactions warrants correction through either operator training or redesign.
Suggested Citation Thomas, Dean, Heather Wojton, Chad Bieber, and Daniel Porter.</description>
      <content:encoded><![CDATA[<p>The purpose of this paper was to identify the shortcomings of a single-method approach to evaluating human-system interactions during operational testing and offer an alternative, multi-method approach that is more defensible, yields richer insights into how operators interact with weapon systems, and provides a practical implications for identifying when the quality of human-system interactions warrants correction through either operator training or redesign.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Thomas, Dean, Heather Wojton, Chad Bieber, and Daniel Porter. A Multi-Method Approach to Evaluating Human-System Interactions during Operational Testing. IDA Document NS D-8857. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Foundations of Psychological Measurement</title>
      <link>https://research.testscience.org/post/2017-foundations-of-psychological-measurement/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-foundations-of-psychological-measurement/</guid>
      <description>Psychological measurement is an important issue throughout the Department of Defense (DoD). Forinstance, the DoD engages in psychological measurement to place military personnel into specialties,evaluate the mental health of military personnel, evaluate the quality of human-systems interactions, andidentify factors that affect crime rates on bases. Given its broad use, researchers and decision-makers needto understand the basics of psychological measurement – most notably, the development of surveys. Thisbriefing discusses 1) the goals and challenges of psychological measurement, 2) basic measurementconcepts and how they apply to psychological measurement, 3) basics for developing scales to measurepsychological attributes, and 4) methods for ensuring that scales are reliable and valid.</description>
      <content:encoded><![CDATA[<p>Psychological measurement is an important issue throughout the Department of Defense (DoD). Forinstance, the DoD engages in psychological measurement to place military personnel into specialties,evaluate the mental health of military personnel, evaluate the quality of human-systems interactions, andidentify factors that affect crime rates on bases. Given its broad use, researchers and decision-makers needto understand the basics of psychological measurement – most notably, the development of surveys. Thisbriefing discusses 1) the goals and challenges of psychological measurement, 2) basic measurementconcepts and how they apply to psychological measurement, 3) basics for developing scales to measurepsychological attributes, and 4) methods for ensuring that scales are reliable and valid.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather. Foundations of Psychological Measurement. IDA Document NS D-8273. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-8273.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Users are Part of the System-How to Account for Human Factors when Designing Operational Tests for Software Systems</title>
      <link>https://research.testscience.org/post/2017-users-are-part-of-the-system-how-to-account-for-human-factors-when-designing-operational-tests-for-software-systems/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-users-are-part-of-the-system-how-to-account-for-human-factors-when-designing-operational-tests-for-software-systems/</guid>
      <description>The goal of operation testing (OT) is to evaluate the effectiveness and suitability of military systems for use by trained military users in operationally realistic environments. Operators perform missions and make systems function. Thus, adequate OT must assess not only system performance and technical capability across the operational space, but also the quality of human-system interactions. Software systems in particular pose a unique challenge to testers. While some software systems may inherently be deterministic in nature, once placed in their intended environment with error-prone humans and highly stochastic networks, variability in outcomes often occurs, so tests often need to account for both “bug” finding and characterizing variability.</description>
      <content:encoded><![CDATA[<p>The goal of operation testing (OT) is to evaluate the effectiveness and suitability of military systems for use by trained military users in operationally realistic environments.  Operators perform missions and make systems function.  Thus, adequate OT must assess not only system performance and technical capability across the operational space, but also the quality of human-system interactions. Software systems in particular pose a unique challenge to testers. While some software systems may inherently be deterministic in nature, once placed in their intended environment with error-prone humans and highly stochastic networks, variability in outcomes often occurs, so tests often need to account for both “bug” finding and characterizing variability.   This document outlines common statistical techniques for planning tests of system performance for software systems, and then discusses how testers might integrate human-system interaction metrics into that design and evaluation. System PerformanceBefore deciding what class of statistical design techniques to apply, testers should consider whether the system under test is deterministic (repeating a process with the same inputs always produces the same output) or stochastic (even if the inputs are fixed, repeating the process again could produce a different result).  Software systems–a calculator, for example– may intuitively be deterministic, and as standalone entities in a pristine environment, they are.  However, there are other sources of variation to consider when testing such a system in an operational environment with an intended user.  If the calculator is intended to be used by scientists in Antarctica, temperature, lighting conditions, and user clothing such as gloves all could affect the users’ ability to operate the system. Combinatorial covering arrays can cover a large input space extremely efficiently and are useful for conducting functionality checks of a complex system.  However, several assumptions must be met in order for testers to benefit from combinatorial designs.  The system must be fully deterministic, the response variable of interest must be binary (pass/fail), and the primary goal of the test must be to find problems.  Combinatorial designs cannot determine cause and effect and are not designed to detect or quantify uncertainty or variability in responses. In operational testing, the assumptions listed above typically are not met.  Any number of factors, including the human user, the network load, memory leaks, database errors, and a constantly changing environment can cause variability in the mission-level outcome of interest.  While combinatorial designs can be useful for bug checking, they typically are not sufficient for OT.  One goal of OT should be to characterize system performance across the space.   The appropriate designs to support characterization are classical or optimal designs.  These designs, including factorial, fractional factorial, response surface, and D-optimal constructs, have the ability to quantify variability in outcomes and attribute changes in response to specific factors or factor interactions. These two broad classes of design (combinatorial and classical) can be merged in order to serve both goals, finding problems and characterizing performance.  Testers can develop a “hybrid” design by first building a combinatorial covering array across all factors, and then adding the necessary runs to support a D-optimal design, for example.  This allows testers to efficiently detect any remaining “bugs” in the software, while also quantifying variability and supporting statistical regression analysis of the data.Human-System InteractionIt is not sufficient only to assess technical performance when testing software systems.  Systems that account for human factors (operators’ physical and psychological characteristics) are more likely to fulfill their missions. Software that is psychologically challenging often leads to mistakes, inefficiencies, and safety concerns. Testers can use human-system interaction (HSI) metrics to capture software compatibility with key psychological characteristics.  Inherent characteristics such as short- and long-term memory processes, capacity for attention, and cognitive load are directly related to measurable constructs such as usability, workload, and task error rates. To evaluate HSI, testers can use either behavioral metrics (e.g. error rates, completion times, speech/facial expressions) or self-report metrics (surveys and interviews). Though behavioral metrics are generally preferred since they are directly observable, the method you choose depends on the HSI concept you want to measure, your test design, and operational constraints. The same logic can be applied to HSI data collection as data collection for system performance.  Testers should strive to understand how users’ experience of the system shifts with the operational environment, thus designed experiments with factors and levels should be applied.    In addition, understanding if, or how much, user experience affects system performance is key to a thorough evaluation.The easiest way to fit HSI into OT is to leverage the existing test design.  First, identify the subset (or possibly superset) of factors that are likely to shape how users experience the system, then distribute those users across the test conditions logically.  The number of users, their groupings, and how they will be spread across the factor space all matter when designing an adequate test for HSI.Most HSI data, including behavioral metrics and empirically validated surveys, also can be analyzed in the same way system performance data can, using statistically rigorous techniques such as regression.  Operational conditions, user type, and system characteristics all can affect HSI, so it is critical to account for those factors in the design and analysis.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J, Kelly M Avery, and Heather M Wojton. Users Are Part of the System: How to Account for Human Factors When Designing Operational Tests for Software Systems. IDA Document NS D-8630. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Survey Design</title>
      <link>https://research.testscience.org/post/2016-introduction-to-survey-design/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-introduction-to-survey-design/</guid>
      <description>An important goal of test and evaluation is to understand not only how a system performs in its intended environment, but also users’ experiences operating the system. This briefing aimed to provide the audience with a set of tools – most notably, surveys – that are appropriate for measuring the user experience. DOT&amp;amp;E guidance regarding these tools is highlighted where appropriate. The briefing was broken into three major sections: conceptualizing surveys, writing survey items, and formatting surveys.</description>
      <content:encoded><![CDATA[<p>An important goal of test and evaluation is to understand not only how a system performs in its intended environment, but also users’ experiences operating the system. This briefing aimed to provide the audience with a set of tools – most notably, surveys – that are appropriate for measuring the user experience. DOT&amp;E guidance regarding these tools is highlighted where appropriate. The briefing was broken into three major sections: conceptualizing surveys, writing survey items, and formatting surveys. At the end of this briefing, the audience should have a better understanding of the value and purpose of surveys and how to construct them.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Jonathan Snavely, and Justin Mary. Introduction to Survey Design. IDA Document NS D-5835. Alexandria, VA: Institute for Defense Analyses, 2016.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5835.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Surveys in Operational Test and Evaluation</title>
      <link>https://research.testscience.org/post/2015-surveys-in-operational-test-and-evaluation/</link>
      <pubDate>Thu, 01 Jan 2015 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2015-surveys-in-operational-test-and-evaluation/</guid>
      <description>Recently DOT&amp;amp;E signed out a memo providing Guidance on the Use and Design of Surveys in Operational Test and Evaluation. This guidance memo helps the Human Systems Integration (HSI) community to ensure that useful and accurate HSI data are collected. Information about how HSI experts can leverage the guidance is presented. Specifically, the presentation will cover which HSI metrics can and cannot be answered by surveys.
Suggested Citation Grier, Rebecca A, and Laura Freeman.</description>
      <content:encoded><![CDATA[<p>Recently  DOT&amp;E signed out a memo providing Guidance on the Use and Design of Surveys in Operational Test and Evaluation. This guidance memo helps the Human Systems Integration (HSI) community to ensure that useful and accurate HSI data are collected. Information about how HSI experts can leverage the guidance is presented. Specifically, the presentation will cover which HSI metrics can and cannot be answered by surveys.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Grier, Rebecca A, and Laura Freeman. Surveys in Operational Test &amp; Evaluation. IDA Document D-5410. Alexandria, VA: Institute for Defense Analyses, 2015.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-5410-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
  </channel>
</rss>
