<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Institute for Defense Analyses on Test Science Research Document Library</title>
    <link>https://research.testscience.org/venues/institute-for-defense-analyses/</link>
    <description>Recent content in Institute for Defense Analyses on Test Science Research Document Library</description>
    <generator>Hugo -- 0.129.0</generator>
    <language>en-us</language>
    <copyright>Institute for Defense Analyses</copyright>
    <lastBuildDate>Mon, 01 Jan 2024 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://research.testscience.org/venues/institute-for-defense-analyses/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>A Reliability Assurance Test Planning and Analysis Tool</title>
      <link>https://research.testscience.org/post/2024-a-reliability-assurance-test-planning-and-analysis-tool/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-a-reliability-assurance-test-planning-and-analysis-tool/</guid>
      <description>This presentation documents the work of IDA 2024 Summer Associate Emma Mitchell. The work presented details an R Shiny application developed to provide a user-friendly software tool for researchers to use in planning for and analyzing system reliability. Specifically, the presentation details how one can plan for a reliability test using Bayesian Reliability Assurance test methods. Such tests utilize supplementary data and information, including reliability models, prior test results, expert judgment, and knowledge of environmental conditions, to plan for reliability testing, which in turn can often help in reducing the required amount of testing.</description>
      <content:encoded><![CDATA[<p>This presentation documents the work of IDA 2024 Summer Associate Emma Mitchell. The work presented details an R Shiny application developed to provide a user-friendly software tool for researchers to use in planning for and analyzing system reliability. Specifically, the presentation details how one can plan for a reliability test using Bayesian Reliability Assurance test methods. Such tests utilize supplementary data and information, including reliability models, prior test results, expert judgment, and knowledge of environmental conditions, to plan for reliability testing, which in turn can often help in reducing the required amount of testing. In the planning phase, the application enables researchers to use Bayesian methods to incorporate supplementary data when determining appropriate test lengths. In the analysis phase, the tool allows researchers to combine information through Bayesian methods, resulting in better uncertainty quantification than traditional methods.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, Rebecca M Medlin, Emma P Mitchell, Keyla Pagán-Rivera, and Dhruv K Patel. A Reliability Assurance Test Planning and Analysis Tool. IDA Product ID 3003359. Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "3003359%20Pagan-Rivera%20et%20al-3_slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Human-Systems Interaction in Operational Test and Evaluation Course</title>
      <link>https://research.testscience.org/post/2024-introduction-to-human-systems-interaction-in-operational-test-and-evaluation-course/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-introduction-to-human-systems-interaction-in-operational-test-and-evaluation-course/</guid>
      <description>Human-System Interaction (HSI) is the study of interfaces between humans and technical systems. The Department of Defense incorporates HSI evaluations into defense acquisition to improve system performance and reduce lifecycle costs. During operational test and evaluation, HSI evaluations characterize how a system’s operational performance is affected by its users. The goal of this course is to provide the theoretical background and practical tools necessary to plan and evaluate HSI test plans, collect and analyze HSI data, and report on HSI results.</description>
      <content:encoded><![CDATA[<p>Human-System Interaction (HSI) is the study of interfaces between humans and technical systems. The Department of Defense incorporates HSI evaluations into defense acquisition to improve system performance and reduce lifecycle costs. During operational test and evaluation, HSI evaluations characterize how a system’s operational performance is affected by its users. The goal of this course is to provide the theoretical background and practical tools necessary to plan and evaluate HSI test plans, collect and analyze HSI data, and report on HSI results. We will discuss HSI concepts, measurement methods, design of experiments, data analysis, and evaluation and reporting, all from an operational testing perspective.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Miller, Dr Adam M, and Keyla Pagan-Rivera. Introduction to Human-Systems Interaction in Operational Test and Evaluation Course. IDA Product ID 3002009. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_3002009.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Sequential Space-Filling Designs for Modeling &amp; Simulation Analyses</title>
      <link>https://research.testscience.org/post/2024-sequential-space-filling-designs-for-modeling-simulation-analyses/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-sequential-space-filling-designs-for-modeling-simulation-analyses/</guid>
      <description>Space-filling designs (SFDs) are a rigorous method for designing modeling and simulation (M&amp;amp;S) studies. However, they are hindered by their requirement to choose the final sample size prior to testing. Sequential designs are an alternative that can increase test efficiency by testing small amounts of data at a time. We have conducted a literature review of existing sequential space-filling designs and found the methods most applicable to the test and evaluation (T&amp;amp;E) community.</description>
      <content:encoded><![CDATA[<p>Space-filling designs (SFDs) are a rigorous method for designing modeling and simulation (M&amp;S) studies. However, they are hindered by their requirement to choose the final sample size prior to testing. Sequential designs are an alternative that can increase test efficiency by testing small amounts of data at a time. We have conducted a literature review of existing sequential space-filling designs and found the methods most applicable to the test and evaluation (T&amp;E) community.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, and Anna Flowers. Sequential Space-Filling Designs for Modeling &amp; Simulation Analyses. IDA Product ID 3003752. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "3003752%20Haman%20et%20al-3_slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>AI &#43; Autonomy T&amp;E in DoD</title>
      <link>https://research.testscience.org/post/2023-ai-autonomy-t-e-in-dod/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-ai-autonomy-t-e-in-dod/</guid>
      <description>Test and evaluation (T&amp;amp;E) of AI-enabled systems (AIES) often emphasizes algorithm accuracy over robust, holistic system performance. While this narrow focus may be adequate for some applications of AI, for many complex uses, T&amp;amp;E paradigms removed from operational realism are insufficient. However, leveraging traditional operational testing (OT) methods for to evaluate AIESs can fail to capture novel sources of risk. This brief establishes a common AI vocabulary and highlights OT challenges posed by AIESs by answering the following questions</description>
      <content:encoded><![CDATA[<p>Test and evaluation (T&amp;E) of AI-enabled systems (AIES) often emphasizes algorithm accuracy over robust, holistic system performance. While this narrow focus may be adequate for some applications of AI, for many complex uses, T&amp;E paradigms removed from operational realism are insufficient. However, leveraging traditional operational testing (OT) methods for to evaluate AIESs can fail to capture novel sources of risk. This brief establishes a common AI vocabulary and highlights OT challenges posed by AIESs by answering the following questions</p>
<ol>
<li>What is “Artificial Intelligence (AI)”?</li>
</ol>
<p>a. A brief “AI Primer” defines some common terms, highlights words that are used inconsistently, and discusses where definitions are insufficient for identifying systems that require additional T&amp;E considerations.</p>
<ol start="2">
<li>How does AI impact T&amp;E?</li>
</ol>
<p>a. AI isn’t new, but systems with AI pose new challenges and may require structural changes to how we T&amp;E.</p>
<ol start="3">
<li>What makes DoD applications of AI unique?</li>
</ol>
<p>a. Many Silicon Valley applications of AI often lack the task complexity and severe consequences of risk faced by DoD.</p>
<ol start="4">
<li>What is the warfighter’s role?</li>
</ol>
<p>a. T&amp;E must assure warfighters have calibrated trust &amp; an adequate understanding of system behavior.</p>
<ol start="5">
<li>What is the state of DoD AI T&amp;E in IDA and OED?</li>
</ol>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Vickers, Brian D, Matthew R Avery, Rachel A Haga, Mark R Herrera, Daniel J Porter, Stuart M Rodgers, and Rebecca M Medlin. AI + Autonomy T&amp;E in DoD. IDA Document NS 3000083. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Statistical Methods for M&amp;S V&amp;V- An Intro for Non-Statisticians</title>
      <link>https://research.testscience.org/post/2023-statistical-methods-for-m-s-v-v-an-intro-for-non-statisticians/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-statistical-methods-for-m-s-v-v-an-intro-for-non-statisticians/</guid>
      <description>This is a briefing intended to motivate and explain the basic concepts of applying statistics to verification and validation. The briefing will be presented at the Navy M&amp;amp;S VV&amp;amp;A WG (Sub-WG on Validation Statistical Method Selection).
Suggested Citation Pagan-Rivera, Keyla, John T Haman, Kelly M Avery, and Curtis G Miller. Statistical Methods for M&amp;amp;S V&amp;amp;V: An Intro for Non- Statisticians. IDA Product ID-3000770. Alexandria, VA: Institute for Defense Analyses, 2024.</description>
      <content:encoded><![CDATA[<p>This is a briefing intended to motivate and explain the basic concepts of applying statistics to verification and validation. The briefing will be presented at the Navy M&amp;S VV&amp;A WG (Sub-WG on Validation Statistical Method Selection).</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Pagan-Rivera, Keyla, John T Haman, Kelly M Avery, and Curtis G Miller. Statistical Methods for M&amp;S V&amp;V: An Intro for Non- Statisticians. IDA Product ID-3000770. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Analysis Apps for the Operational Tester</title>
      <link>https://research.testscience.org/post/2022-analysis-apps-for-the-operational-tester/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-analysis-apps-for-the-operational-tester/</guid>
      <description>In the acquisition and testing world, data analysts repeatedly encounter certain categories of data, such as time or distance until an event (e.g., failure, alert, detection), binary outcomes (e.g., success/failure, hit/miss), and survey responses. Analysts need tools that enable them to produce quality and timely analyses of the data they acquire during testing. This poster presents four web-based apps that can analyze these types of data. The apps are designed to assist analysts and researchers with simple repeatable analysis tasks, such as building summary tables and plots for reports or briefings.</description>
      <content:encoded><![CDATA[<p>In the acquisition and testing world, data analysts repeatedly encounter certain categories of data, such as time or distance until an event (e.g., failure, alert, detection), binary outcomes (e.g., success/failure, hit/miss), and survey responses. Analysts need tools that enable them to produce quality and timely analyses of the data they acquire during testing. This poster presents four web-based apps that can analyze these types of data. The apps are designed to assist analysts and researchers with simple repeatable analysis tasks, such as building summary tables and plots for reports or briefings. Using software tools like these apps can increase reproducibility of results, timeliness of analysis and reporting, attractiveness and standardization of aesthetics in figures, and accuracy of results. The first app models reliability of a system or component by fitting parametric statistical distributions to time-to-failure data. The second app fits a logistic regression model to binary data with one or two independent continuous variables as The third calculates summary statistics and produces plots of groups of Likert-scale survey question responses. The fourth calculates the system usability scale (SUS) scores for SUS survey responses and enables the app user to plot scores versus an independent variable. These apps are available for public use on the Test Science Interactive Tools webpage <a href="https://testscience.org/interactive-tools/">https://testscience.org/interactive-tools/</a>.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Lillard, V Bram, and William Whitledge. Analysis Apps for the Operational Tester. IDA Document NS D-32959. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Metamodeling Techniques for Verification and Validation of Modeling and Simulation Data</title>
      <link>https://research.testscience.org/post/2022-metamodeling-techniques-for-verification-and-validation-of-modeling-and-simulation-data/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-metamodeling-techniques-for-verification-and-validation-of-modeling-and-simulation-data/</guid>
      <description>Modeling and simulation (M&amp;amp;S) outputs help the Director, Operational Test and Evaluation (DOT&amp;amp;E) assess the effectiveness, survivability, lethality, and suitability of systems. To use M&amp;amp;S outputs, DOT&amp;amp;E needs models and simulators to be sufficiently verified and validated. The purpose of this paper is to improve the state of verification and validation by recommending and demonstrating a set of statistical techniques—metamodels, also called statistical emulators—to the M&amp;amp;S community.
The paper expands on DOT&amp;amp;E’s existing guidance about metamodel usage by creating methodological recommendations the M&amp;amp;S community could apply to its activities.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/s4DUCI1M8Fw?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Modeling and simulation (M&amp;S) outputs help the Director, Operational Test and Evaluation (DOT&amp;E) assess the effectiveness, survivability, lethality, and suitability of systems. To use M&amp;S outputs, DOT&amp;E needs models and simulators to be sufficiently verified and validated. The purpose of this paper is to improve the state of verification and validation by recommending and demonstrating a set of statistical techniques—metamodels, also called statistical emulators—to the M&amp;S community.</p>
<p>The paper expands on DOT&amp;E’s existing guidance about metamodel usage by creating methodological recommendations the M&amp;S community could apply to its activities. For a deterministic, discrete response variable, we recommend using a nearest neighbor or decision tree model. For a deterministic, continuous response variable, we recommend Gaussian process interpolation. For a stochastic response variable, we recommend a generalized additive model. We also present a set of techniques that testers can use to assess the adequacy of their metamodels. We conclude with a notional example that demonstrates the recommended techniques.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, and Curtis G Miller. Metamodeling Techniques for Verification and Validation of Modeling and Simulation Data. IDA Paper P-33230. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Predicting Trust in Automated Systems - An Application of TOAST</title>
      <link>https://research.testscience.org/post/2022-predicting-trust-in-automated-systems-an-application-of-toast/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-predicting-trust-in-automated-systems-an-application-of-toast/</guid>
      <description>Following Wojton&amp;rsquo;s research on the Trust of Automated Systems Test (TOAST), which is designed to measure how much a human trusts an automated system, we aimed to determine how well this scale performs when not used in a military context. We found that participants who used a poorly performing automated system trusted the system less than expected when using that system on a case by case basis, however, those who used a high performing system trusted the system the same as they expected.</description>
      <content:encoded><![CDATA[<p>Following Wojton&rsquo;s research on the Trust of Automated Systems Test (TOAST), which is designed to measure how much a human trusts an automated system, we aimed to determine how well this scale performs when not used in a military context. We found that participants who used a poorly performing automated system trusted the system less than expected when using that system on a case by case basis, however, those who used a high performing system trusted the system the same as they expected. Additionally, both participants who used the poorly performing system and those who used the high performing system lost a significant amount of trust after using the system on a group case basis. These results indicate that having a high performance system is important for trust, but only when the user has the ability to decide to trust or distrust the system on a case-by-case basis.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Porter, Daniel J, and Caitlan A Fealing. Predicting Trust in Automated Systems – An Application of TOAST. IDA Document NS D-33188. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Artificial Intelligence &amp; Autonomy Test &amp; Evaluation Roadmap Goals</title>
      <link>https://research.testscience.org/post/2021-artificial-intelligence-autonomy-test-evaluation-roadmap-goals/</link>
      <pubDate>Fri, 01 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2021-artificial-intelligence-autonomy-test-evaluation-roadmap-goals/</guid>
      <description>As the Department of Defense acquires new systems with artificial intelligence (AI) and autonomous (AI&amp;amp;A) capabilities, the test and evaluation (T&amp;amp;E) community will need to adapt to the challenges that these novel technologies present. The goals listed in this AI Roadmap address the broad range of tasks that the T&amp;amp;E community will need to achieve in order to properly test, evaluate, verify, and validate AI-enabled and autonomous systems. It includes issues that are unique to AI and autonomous systems, as well as legacy T&amp;amp;E shortcomings that will be compounded by newer technologies.</description>
      <content:encoded><![CDATA[<p>As the Department of Defense acquires new systems with artificial intelligence (AI) and autonomous (AI&amp;A) capabilities, the test and evaluation (T&amp;E) community will need to adapt to the challenges that these novel technologies present. The goals listed in this AI Roadmap address the broad range of tasks that the T&amp;E community will need to achieve in order to properly test, evaluate, verify, and validate AI-enabled and autonomous systems. It includes issues that are unique to AI and autonomous systems, as well as legacy T&amp;E shortcomings that will be compounded by newer technologies.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Brian Vickers, Daniel Porter, and Rachel Haga. Artificial Intelligence &amp; Autonomy Test &amp; Evaluation Roadmap Goals. IDA Document NS D-22750. Alexandria, VA: Institute for Defense Analyses, 2021.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Bayesian Analysis</title>
      <link>https://research.testscience.org/post/2021-introduction-to-bayesian-analysis/</link>
      <pubDate>Fri, 01 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2021-introduction-to-bayesian-analysis/</guid>
      <description>As operational testing becomes increasingly integrated and research questions become more difficult to answer, IDA’s Test Science team has found Bayesian models to be powerful data analysis methods. Analysts and decision-makers should understand the differences between this approach and the conventional way of analyzing data. It is also important to recognize when an analysis could benefit from the inclusion of prior information—what we already know about a system’s performance—and to understand the proper way to incorporate that information.</description>
      <content:encoded><![CDATA[<p>As operational testing becomes increasingly integrated and research questions become more difficult to answer, IDA’s Test Science team has found Bayesian models to be powerful data analysis methods. Analysts and decision-makers should understand the differences between this approach and the conventional way of analyzing data. It is also important to recognize when an analysis could benefit from the inclusion of prior information—what we already know about a system’s performance—and to understand the proper way to incorporate that information. To apply Bayesian methods, analysts need to comprehend some technical aspects of this approach and know how to properly use appropriate statistical software. In this course, students learn the intuition behind Bayesian statistics, the mathematical details of posterior distributions, how to fit simple Bayesian models using computer software, and how to assess model fit.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather M, Keyla Pagan-Rivera, John T Haman, and Rebecca M Medlin. Introduction to Bayesian Analysis. IDA Document NS D-20484. Alexandria, VA: Institute for Defense Analyses, 2021.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-20484.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Review of Sequential Analysis</title>
      <link>https://research.testscience.org/post/2020-a-review-of-sequential-analysis/</link>
      <pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2020-a-review-of-sequential-analysis/</guid>
      <description>Sequential analysis concerns statistical evaluation in situations in which the number, pattern, or composition of the data is not determined at the start of the investigation, but instead depends upon the information acquired throughout the course of the investigation. Expanding the use of sequential analysis has the potential to save resources and reduce test time (National Research Council, 1998). This paper summarizes the literature on sequential analysis and offers fundamental information for providing recommendations for its use in DoD test and evaluation.</description>
      <content:encoded><![CDATA[<p>Sequential analysis concerns statistical evaluation in situations in which the number, pattern, or composition of the data is not determined at the start of the investigation, but instead depends upon the information acquired throughout the course of the investigation. Expanding the use of sequential analysis has the potential to save resources and reduce test time (National Research Council, 1998). This paper summarizes the literature on sequential analysis and offers fundamental information for providing recommendations for its use in DoD test and evaluation.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Rebecca Medlin, John Dennis, Keyla Pagan-Rivera, and Leonard Wilkins. A Review of Sequential Analysis. IDA Document NS D-20487. Alexandria, VA: Institute for Defense Analyses, 2020.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper_seq_review.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>T&amp;E Contributions to Avoiding Unintended Behaviors in Autonomous Systems</title>
      <link>https://research.testscience.org/post/2020-t-e-contributions-to-avoiding-unintended-behaviors-in-autonomous-systems/</link>
      <pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2020-t-e-contributions-to-avoiding-unintended-behaviors-in-autonomous-systems/</guid>
      <description>To provide assurance that AI-enabled systems will behave appropriately across the range of their operating conditions without performing exhaustive testing, the DoD will need to make inferences about system decision making. However, making these inferences validly requires understanding what causally drives system decision-making, which is not possible when systems are black boxes. In this briefing, we discuss the state of the art and gaps in techniques for obtaining, verifying, validating, and accrediting (OVVA) models of system decision-making.</description>
      <content:encoded><![CDATA[<p>To provide assurance that AI-enabled systems will behave appropriately across the range of their operating conditions without performing exhaustive testing, the DoD will need to make inferences about system decision making. However, making these inferences validly requires understanding what causally drives system decision-making, which is not possible when systems are black boxes. In this briefing, we discuss the state of the art and gaps in techniques for obtaining, verifying, validating, and accrediting (OVVA) models of system decision-making.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Porter, Daniel J, and Heather Wojton. T&amp;E Contributions to Avoiding Unintended Behaviors in Autonomous Systems. Vol. IDA Document NS D-12078. Alexandria, VA: Institute for Defense Analyses, 2020.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Test &amp; Evaluation of AI-Enabled and Autonomous Systems- A Literature Review</title>
      <link>https://research.testscience.org/post/2020-test-evaluation-of-ai-enabled-and-autonomous-systems-a-literature-review/</link>
      <pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2020-test-evaluation-of-ai-enabled-and-autonomous-systems-a-literature-review/</guid>
      <description>We summarize a subset of the literature regarding the challenges to and recommendations for the test, evaluation, verification, and validation (TEV&amp;amp;V) of autonomous military systems. This literature review is meant for informational purposes only and does not make any recommendations of its own. A synthesis of the literature identified the following categories of TEV&amp;amp;V challenges
Problems arising from the complexity of autonomous systems,
Challenges imposed by the structure of the current acquisition system,</description>
      <content:encoded><![CDATA[<p>We summarize a subset of the literature regarding the challenges to and recommendations for the test, evaluation, verification, and validation (TEV&amp;V) of autonomous military systems. This literature review is meant for informational purposes only and does not make any recommendations of its own. A synthesis of the literature identified the following categories of TEV&amp;V challenges</p>
<ol>
<li>
<p>Problems arising from the complexity of autonomous systems,</p>
</li>
<li>
<p>Challenges imposed by the structure of the current acquisition system,</p>
</li>
<li>
<p>Lack of methods, tools, and infrastructure for testing,</p>
</li>
<li>
<p>Novel safety and security issues,</p>
</li>
<li>
<p>A lack of consensus on policy, standards, and metrics,</p>
</li>
<li>
<p>Issues around how to integrate humans into the operation and testing of these systems.</p>
</li>
</ol>
<p>Recommendations for how to test autonomous military systems can be sorted into five broad groups</p>
<ol>
<li>
<p>Use certain processes for writing requirements, or for designing and developing systems,</p>
</li>
<li>
<p>Make targeted investments to develop methods or tools, improve our test infrastructure, or enhance our workforce&rsquo;s AI skillsets,</p>
</li>
<li>
<p>Use specific proposed test frameworks,</p>
</li>
<li>
<p>Employ novel methods for system safety or cybersecurity, and</p>
</li>
<li>
<p>Adopt specific proposed policies, standards, or metrics.</p>
</li>
</ol>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather M, Daniel J Porter, and John W Dennis. Test &amp; Evaluation of AI-Enabled and Autonomous Systems: A Literature Review. IDA Document NS-D-14331. Alexandria, VA: Institute for Defense Analyses, 2020.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Trustworthy Autonomy- A Roadmap to Assurance -- Part 1- System Effectiveness</title>
      <link>https://research.testscience.org/post/2020-trustworthy-autonomy-a-roadmap-to-assurance-part-1-system-effectiveness/</link>
      <pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2020-trustworthy-autonomy-a-roadmap-to-assurance-part-1-system-effectiveness/</guid>
      <description>The Department of Defense (DoD) has invested significant effort over the past decade considering the role of artificial intelligence and autonomy in national security (e.g., Defense Science Board, 2012, 2016, Deputy Secretary of Defense, 2012, Endsley, 2015, Executive Order No. 13859, 2019, US Department of Defense, 2011, 2019, Zacharias, 2019a). However, these efforts were broadly scoped and only partially touched on how the DoD will certify the safety and performance of these systems.</description>
      <content:encoded><![CDATA[<p>The Department of Defense (DoD) has invested significant effort over the past decade considering the role of artificial intelligence and autonomy in national security (e.g., Defense Science Board, 2012, 2016, Deputy Secretary of Defense, 2012, Endsley, 2015, Executive Order No. 13859, 2019, US Department of Defense, 2011, 2019, Zacharias, 2019a). However, these efforts were broadly scoped and only partially touched on how the DoD will certify the safety and performance of these systems. More recent work has done this big-picture thinking for the test and evaluation (T&amp;E) community (e.g., Ahner &amp; Parson, 2016, Haugh, Sparrow, &amp; Tate, 2018, Porter et al., 2018, Sparrow, Tate, Biddle, Kaminski, &amp; Madhavan, 2018, Zacharias, 2019b). In parallel, individual programs have been generating their own working-level solutions for their own particular use-cases and challenges.</p>
<p>The framework proposed in the current work bridges the gap between the big picture policy recommendations already made and individual program needs. It is meant to serve as a roadmap framework that the T&amp;E community can follow in order to provide evidence that artificial intelligence (AI)-enabled and autonomous systems function as intended. At times we echo broad policy recommendations made by others as they will also enable T&amp;E activities. In other places we make more specific recommendations relating to test planning and analysis. In this document, we present part one of our two-part roadmap. We discuss the challenges and possible solutions to assessing system effectiveness. A future part two will deal with test efficiency, simulation, and infrastructure. Due to the scope of this project, even the main body of this document only provides a survey of the challenges and our proposed solutions. However, this roadmap serves as an outline to a future series of technical papers covering these topics in detail for working-level testers and analysts</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Porter, Daniel, Michael McAnally, Chad Bieber, Heather Wojton, and Rebecca Medlin. Trustworthy Autonomy: A Roadmap to Assurance Part I: System Effectiveness. IDA Document P-10768-NS. Alexandria, VA: Institute for Defense Analyses, 2020.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Visualizing Data- I Don&#39;t Remember that Memo, but I Do Remember that Graph</title>
      <link>https://research.testscience.org/post/2020-visualizing-data-i-don-t-remember-that-memo-but-i-do-remember-that-graph/</link>
      <pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2020-visualizing-data-i-don-t-remember-that-memo-but-i-do-remember-that-graph/</guid>
      <description>IDA analysts strive to communicate clearly and effectively. Good data visualizations can enhance reports by making the conclusions easier to understand and more memorable. The goal of this seminar is to help you avoid settling for factory defaults and instead present your conclusions through visually appealing and understandable charts. Topics covered include choosing the right level of detail, guidelines for different types of graphical elements (titles, legends, annotations, etc.), selecting the right variable encodings (color, plot symbol, etc.</description>
      <content:encoded><![CDATA[<p>IDA analysts strive to communicate clearly and effectively. Good data visualizations can enhance reports by making the conclusions easier to understand and more memorable. The goal of this seminar is to help you avoid settling for factory defaults and instead present your conclusions through visually appealing and understandable charts. Topics covered include choosing the right level of detail, guidelines for different types of graphical elements (titles, legends, annotations, etc.), selecting the right variable encodings (color, plot symbol, etc.), advice on practical implementations, and determining whether to include a chart at all. Most of the time, there’s no single “right” answer, so this presentation will include audience discussion to examine the trade-offs associated with different options.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Matthew, Heather Wojton, Andrew Flack, and Brian Vickers. Visualizing Data: I Don’t Remember That Memo, but I Do Remember That Graph. Alexandria, VA: Institute for Defense Analyses, 2020.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Demystifying the Black Box- A Test Strategy for Autonomy</title>
      <link>https://research.testscience.org/post/2019-demystifying-the-black-box-a-test-strategy-for-autonomy/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-demystifying-the-black-box-a-test-strategy-for-autonomy/</guid>
      <description>The purpose of this briefing is to provide a high-level overview of how to frame the question of testing autonomous systems in a way that will enable development of successful test strategies. The brief outlines the challenges and broad-stroke reforms needed to get ready for the test challenges of the next century.
Suggested Citation Wojton, Heather M, and Daniel J Porter. Demystifying the Black Box: A Test Strategy for Autonomy. IDA Document NS D-10465-NS.</description>
      <content:encoded><![CDATA[<p>The purpose of this briefing is to provide a high-level overview of how to frame the question of testing autonomous systems in a way that will enable development of successful test strategies. The brief outlines the challenges and broad-stroke reforms needed to get ready for the test challenges of the next century.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather M, and Daniel J Porter. Demystifying the Black Box: A Test Strategy for Autonomy. IDA Document NS D-10465-NS. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Handbook on Statistical Design &amp; Analysis Techniques for Modeling &amp; Simulation Validation</title>
      <link>https://research.testscience.org/post/2019-handbook-on-statistical-design-analysis-techniques-for-modeling-simulation-validation/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-handbook-on-statistical-design-analysis-techniques-for-modeling-simulation-validation/</guid>
      <description>This handbook focuses on methods for data-driven validation to supplement the vast existing literature for Verification, Validation, and Accreditation (VV&amp;amp;A) and the emerging references on uncertainty quantification (UQ). The goal of this handbook is to aid the test and evaluation (T&amp;amp;E) community in developing test strategies that support model validation (both external validation and parametric analysis) and statistical UQ.
Suggested Citation Wojton, Heather, Kelly M Avery, Laura J Freeman, Samuel H Parry, Gregory S Whittier, Thomas H Johnson, and Andrew C Flack.</description>
      <content:encoded><![CDATA[<p>This handbook focuses on methods for data-driven validation to supplement the vast existing literature for Verification, Validation, and Accreditation (VV&amp;A) and the emerging references on uncertainty quantification (UQ). The goal of this handbook is to aid the test and evaluation (T&amp;E) community in developing test strategies that support model validation (both external validation and parametric analysis) and statistical UQ.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Kelly M Avery, Laura J Freeman, Samuel H Parry, Gregory S Whittier, Thomas H Johnson, and Andrew C Flack. Handbook on Statistical Design &amp; Analysis Techniques for Modeling &amp; Simulation Validation. IDA Document NS D-10455. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Operational Testing of Systems with Autonomy</title>
      <link>https://research.testscience.org/post/2019-operational-testing-of-systems-with-autonomy/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-operational-testing-of-systems-with-autonomy/</guid>
      <description>Systems with autonomy pose unique challenges for operational test. This document provides an executive level overview of these issues and the proposed solutions and reforms. In order to be ready for the testing challenges of the next century, we will need to change the entire acquisition life cycle, starting even from initial system conceptualization. This briefing was presented to the Director, Operational Test &amp;amp; Evaluation along with his deputies and Chief Scientist.</description>
      <content:encoded><![CDATA[<p>Systems with autonomy pose unique challenges for operational test. This document provides an executive level overview of these issues and the proposed solutions and reforms. In order to be ready for the testing challenges of the next century, we will need to change the entire acquisition life cycle, starting even from initial system conceptualization. This briefing was presented to the Director, Operational Test &amp; Evaluation along with his deputies and Chief Scientist.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather M, Daniel Porter, Yevgeniya Pinelis, Chad Bieber, Heather Wojton, Michael McAnally, and Laura Freeman. Operational Testing of Systems with Autonomy. IDA Document NS D-9266. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Pilot Training Next- Modeling Skill Transfer in a Military Learning Environment</title>
      <link>https://research.testscience.org/post/2019-pilot-training-next-modeling-skill-transfer-in-a-military-learning-environment/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-pilot-training-next-modeling-skill-transfer-in-a-military-learning-environment/</guid>
      <description>Pilot Training Next is an exploratory investigation of new technologies and procedures to increase the efficiency of Undergraduate Pilot Training in the United States Air Force. IDA analysts present a method of quantifying skill transfer from simulators to aircraft under realistic, uncontrolled conditions.
Suggested Citation Porter, Daniel, Emily Fedele, and Heather Wojton. Pilot Training Next: Modeling Skill Transfer in a Military Learning Environment. IDA Document NS D-10927. Alexandria, VA: Institute for Defense Analyses, 2019.</description>
      <content:encoded><![CDATA[<p>Pilot Training Next is an exploratory investigation of new technologies and procedures to increase the efficiency of Undergraduate Pilot Training in the United States Air Force. IDA analysts present a method of quantifying skill transfer from simulators to aircraft under realistic, uncontrolled conditions.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Porter, Daniel, Emily Fedele, and Heather Wojton. Pilot Training Next: Modeling Skill Transfer in a Military Learning Environment. IDA Document NS D-10927. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-10927-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Sample Size Determination Methods Using Acceptance Sampling by Variables</title>
      <link>https://research.testscience.org/post/2019-sample-size-determination-methods-using-acceptance-sampling-by-variables/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-sample-size-determination-methods-using-acceptance-sampling-by-variables/</guid>
      <description>Acceptance Sampling by Variables (ASbV) is a statistical testing technique used in Personal Protective Equipment programs to determine the quality of the equipment in First Article and Lot Acceptance Tests. This article intends to remedy the lack of existing references that discuss the similarities between ASbV and certain techniques used in different sub-disciplines within statistics. Understanding ASbV from a statistical perspective allows testers to create customized test plans, beyond what is available in MIL-STD-414.</description>
      <content:encoded><![CDATA[<p>Acceptance Sampling by Variables (ASbV) is a statistical testing technique used in Personal Protective Equipment programs to determine the quality of the equipment in First Article and Lot Acceptance Tests. This article intends to remedy the lack of existing references that discuss the similarities between ASbV and certain techniques used in different sub-disciplines within statistics. Understanding ASbV from a statistical perspective allows testers to create customized test plans, beyond what is available in MIL-STD-414.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Walzl, Kerry, Lindsey A Davis, Thomas H Johnson, and Heather M Wojton. Sample Size Determination Methods Using Acceptance Sampling by Variables. IDA Document NS D-10666. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Comparing M&amp;S Output to Live Test Data- A Missile System Case Study</title>
      <link>https://research.testscience.org/post/2018-comparing-m-s-output-to-live-test-data-a-missile-system-case-study/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-comparing-m-s-output-to-live-test-data-a-missile-system-case-study/</guid>
      <description>In the operational testing of DoD weapons systems, modeling and simulation (M&amp;amp;S) is often used to supplement live test data in order to support a more complete and rigorous evaluation. Before the output of the M&amp;amp;S is included in reports to decision makers, it must first be thoroughly verified and validated to show that it adequately represents the real world for the purposes of the intended use. Part of the validation process should include a statistical comparison of live data to M&amp;amp;S output.</description>
      <content:encoded><![CDATA[<p>In the operational testing of DoD weapons systems, modeling and simulation (M&amp;S) is often used to supplement live test data in order to support a more complete and rigorous evaluation. Before the output of the M&amp;S is included in reports to decision makers, it must first be thoroughly verified and validated to show that it adequately represents the real world for the purposes of the intended use. Part of the validation process should include a statistical comparison of live data to M&amp;S output. This presentation includes an example of one such validation analysis for a tactical missile system. In this case, the goal is to validate a lethality model that predicts the likelihood of destroying a particular enemy target. Using design of experiments, along with basic analysis techniques such as the Kolmogorov-Smirnov test and Poisson regression, we can explore differences between the M&amp;S and live data across multiple operational conditions and quantify the associated uncertainties.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Thomas, Dean, and Kelly M Avery. Comparing M&amp;S Output to Live Test Data: A Missile System Case Study. IDA Non-Standard Document NS D-9002. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Observational Studies</title>
      <link>https://research.testscience.org/post/2018-introduction-to-observational-studies/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-introduction-to-observational-studies/</guid>
      <description> A presentation on the theory and practice of observational studies. Specific average treatment effect methods include matching, difference-in-difference estimators, and instrumental variables.
Suggested Citation Thomas, Dean, and Yevgeniya K Pinelis. Introduction to Observational Studies. IDA Document NS D-9020. Alexandria, VA: Institute for Defense Analyses, 2018.
Slides: </description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/csO17jA2cxI?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>A presentation on the theory and practice of observational studies.  Specific average treatment effect methods include matching, difference-in-difference estimators, and instrumental variables.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Thomas, Dean, and Yevgeniya K Pinelis. Introduction to Observational Studies. IDA Document NS D-9020. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>JEDIS Briefing and Tutorial</title>
      <link>https://research.testscience.org/post/2018-jedis-briefing-and-tutorial/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-jedis-briefing-and-tutorial/</guid>
      <description>Are you sick of having to manually iterate your way through sizing your design of experiments? Come learn about JEDIS, the new IDA-developed JMP Add-In for automating design of experiments power calculations. JEDIS builds multiple test designs in JMP over user-specified ranges of sample sizes, Signal-to-Noise Ratios (SNR), and alpha (1 -confidence) levels. It then automatically calculates the statistical power to detect an effect due to each factor and any specified interactions for each design.</description>
      <content:encoded><![CDATA[<p>Are you sick of having to manually iterate your way through sizing your design of experiments? Come learn about JEDIS, the new IDA-developed JMP Add-In for automating design of experiments power calculations. JEDIS builds multiple test designs in JMP over user-specified ranges of sample sizes, Signal-to-Noise Ratios (SNR), and alpha (1 -confidence) levels. It then automatically calculates the statistical power to detect an effect due to each factor and any specified interactions for each design. When finished, JEDIS presents the statistical power vs. design metrics in interactive plots and stores the data in an easy to use format. JEDIS creates factorial and optimal designs, but does not currently support split plot designs. If you already have a pre-made design table, the JEDIS Light feature can compute power for the design over ranges of SNR and alpha levels.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Pechkis, Daniel, and Jason P Sheldon. JEDIS Briefing and Tutorial. IDA Document NS D-8964. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Vetting Custom Scales - Understanding Reliability, Validity, and Dimensionality</title>
      <link>https://research.testscience.org/post/2018-vetting-custom-scales-understanding-reliability-validity-and-dimensionality/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-vetting-custom-scales-understanding-reliability-validity-and-dimensionality/</guid>
      <description>For situations in which an empirically vetted scale does not exist or is not suitable, a custom scale may be created. This document presents a comprehensive process for establishing the defensible use of a custom scale. At the highest level, this process encompasses (1) establishing validity of the scale, (2) establishing reliability of the scale, and (3) assessing dimensionality, whether intended or unintended, of the scale. First, the concept of validity is described, including how validity may be established using operators and subject matter experts.</description>
      <content:encoded><![CDATA[<p>For situations in which an empirically vetted scale does not exist or is not suitable, a custom scale may be created. This document presents a comprehensive process for establishing the defensible use of a custom scale. At the highest level, this process encompasses (1) establishing validity of the scale, (2) establishing reliability of the scale, and (3) assessing dimensionality, whether intended or unintended, of the scale. First, the concept of validity is described, including how validity may be established using operators and subject matter experts. The concept of scale reliability is described, with guidelines for computing, interpreting, and using results to inform potential modifications to a custom scale. Next, a method for investigating the dimensionality of a scale, exploratory factor analysis, is described, along with a walkthrough of software implementation and results. Finally, confirmatory factor analysis, a technique for testing a priori hypotheses about dimensionality, is presented.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, and Stephanie Lane. Vetting Custom Scales - Understanding Reliability, Validity, and Dimensionality. IDA Non-Standard Document NS D-9168. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Multi-Method Approach to Evaluating Human-System Interactions During Operational Testing</title>
      <link>https://research.testscience.org/post/2017-a-multi-method-approach-to-evaluating-human-system-interactions-during-operational-testing/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-a-multi-method-approach-to-evaluating-human-system-interactions-during-operational-testing/</guid>
      <description>The purpose of this paper was to identify the shortcomings of a single-method approach to evaluating human-system interactions during operational testing and offer an alternative, multi-method approach that is more defensible, yields richer insights into how operators interact with weapon systems, and provides a practical implications for identifying when the quality of human-system interactions warrants correction through either operator training or redesign.
Suggested Citation Thomas, Dean, Heather Wojton, Chad Bieber, and Daniel Porter.</description>
      <content:encoded><![CDATA[<p>The purpose of this paper was to identify the shortcomings of a single-method approach to evaluating human-system interactions during operational testing and offer an alternative, multi-method approach that is more defensible, yields richer insights into how operators interact with weapon systems, and provides a practical implications for identifying when the quality of human-system interactions warrants correction through either operator training or redesign.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Thomas, Dean, Heather Wojton, Chad Bieber, and Daniel Porter. A Multi-Method Approach to Evaluating Human-System Interactions during Operational Testing. IDA Document NS D-8857. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Foundations of Psychological Measurement</title>
      <link>https://research.testscience.org/post/2017-foundations-of-psychological-measurement/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-foundations-of-psychological-measurement/</guid>
      <description>Psychological measurement is an important issue throughout the Department of Defense (DoD). Forinstance, the DoD engages in psychological measurement to place military personnel into specialties,evaluate the mental health of military personnel, evaluate the quality of human-systems interactions, andidentify factors that affect crime rates on bases. Given its broad use, researchers and decision-makers needto understand the basics of psychological measurement – most notably, the development of surveys. Thisbriefing discusses 1) the goals and challenges of psychological measurement, 2) basic measurementconcepts and how they apply to psychological measurement, 3) basics for developing scales to measurepsychological attributes, and 4) methods for ensuring that scales are reliable and valid.</description>
      <content:encoded><![CDATA[<p>Psychological measurement is an important issue throughout the Department of Defense (DoD). Forinstance, the DoD engages in psychological measurement to place military personnel into specialties,evaluate the mental health of military personnel, evaluate the quality of human-systems interactions, andidentify factors that affect crime rates on bases. Given its broad use, researchers and decision-makers needto understand the basics of psychological measurement – most notably, the development of surveys. Thisbriefing discusses 1) the goals and challenges of psychological measurement, 2) basic measurementconcepts and how they apply to psychological measurement, 3) basics for developing scales to measurepsychological attributes, and 4) methods for ensuring that scales are reliable and valid.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather. Foundations of Psychological Measurement. IDA Document NS D-8273. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-8273.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Thinking About Data for Operational Test and Evaluation</title>
      <link>https://research.testscience.org/post/2017-thinking-about-data-for-operational-test-and-evaluation/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-thinking-about-data-for-operational-test-and-evaluation/</guid>
      <description>While the human brain is powerful tool for quickly recognizing patterns in data, it will frequently make errors in interpreting random data. Luckily, these mistakes occur in systematic and predictable ways. Statistical models provide an analytical framework that helps us avoid these error-prone heuristics and draw accurate conclusions from random data. This non-technical presentation highlights some tricks of the trade learned by studying data and the way the human brain processes.</description>
      <content:encoded><![CDATA[<p>While the human brain is powerful tool for quickly recognizing patterns in data, it will frequently make errors in interpreting random data. Luckily, these mistakes occur in systematic and predictable ways. Statistical models provide an analytical framework that helps us avoid these error-prone heuristics and draw accurate conclusions from random data. This non-technical presentation highlights some tricks of the trade learned by studying data and the way the human brain processes. First, we introduce statistics as the science of data, and discuss how the popular conception of randomness differs from its technical definition. Later sections highlight the human brain as a pattern recognition machine. Examples from published literature and media highlight systematic and predicable errors in human cognition as well as how poor data analysis and graphical displays can cause critical errors in analysis. Finally, we&rsquo;ll talk about using statistical models for analysis, including how violations of model assumptions should effect our analyses.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Thomas, Dean, and Matthew Avery. Thinking About Data for Operational Test and Evaluation. IDA Document NS D-8729. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A First Step into the Bootstrap World</title>
      <link>https://research.testscience.org/post/2016-a-first-step-into-the-bootstrap-world/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-a-first-step-into-the-bootstrap-world/</guid>
      <description>Bootstrapping is a powerful nonparametric tool for conducting statistical inference with many applications to data from operational testing. Bootstrapping is most useful when the population sampled from is unknown or complex or the sampling distribution of the desired statistic is difficult to derive. Careful use of bootstrapping can help address many challenges in analyzing operational test data.
Suggested Citation Avery, Matthew R. A First Step into the Bootstrap World. IDA Document NS D-5816.</description>
      <content:encoded><![CDATA[<p>Bootstrapping is a powerful nonparametric tool for conducting statistical inference with many applications to data from operational testing. Bootstrapping is most useful when the population sampled from is unknown or complex or the sampling distribution of the desired statistic is difficult to derive. Careful use of bootstrapping can help address many challenges in analyzing operational test data.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Matthew R. A First Step into the Bootstrap World. IDA Document NS D-5816. Alexandria, VA: Institute for Defense Analyses, 2016.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>DOT&amp;E Reliability Course</title>
      <link>https://research.testscience.org/post/2016-dot-e-reliability-course/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-dot-e-reliability-course/</guid>
      <description>This reliability course provides information to assist DOT&amp;amp;E action officers in their review and assessment of system reliability. Course briefings cover reliability planning and analysis activities that span the acquisition life cycle. Each briefing discusses review criteria relevant to DOT&amp;amp;E action officers based on DoD policies and lessons learned from previous oversight efforts.
Suggested Citation Avery, Matthew, Jonathan Bell, Rebecca Medlin, and Freeman Laura. DOT&amp;amp;E Reliability Course. IDA Document NS D-5836.</description>
      <content:encoded><![CDATA[<p>This reliability course provides information to assist DOT&amp;E action officers in their review and assessment of system reliability. Course briefings cover reliability planning and analysis activities that span the acquisition life cycle. Each briefing discusses review criteria relevant to DOT&amp;E action officers based on DoD policies and lessons learned from previous oversight efforts.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Matthew, Jonathan Bell, Rebecca Medlin, and Freeman Laura. DOT&amp;E Reliability Course. IDA Document NS D-5836. Alexandria, VA: Institute for Defense Analyses, 2016.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5836-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Survey Design</title>
      <link>https://research.testscience.org/post/2016-introduction-to-survey-design/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-introduction-to-survey-design/</guid>
      <description>An important goal of test and evaluation is to understand not only how a system performs in its intended environment, but also users’ experiences operating the system. This briefing aimed to provide the audience with a set of tools – most notably, surveys – that are appropriate for measuring the user experience. DOT&amp;amp;E guidance regarding these tools is highlighted where appropriate. The briefing was broken into three major sections: conceptualizing surveys, writing survey items, and formatting surveys.</description>
      <content:encoded><![CDATA[<p>An important goal of test and evaluation is to understand not only how a system performs in its intended environment, but also users’ experiences operating the system. This briefing aimed to provide the audience with a set of tools – most notably, surveys – that are appropriate for measuring the user experience. DOT&amp;E guidance regarding these tools is highlighted where appropriate. The briefing was broken into three major sections: conceptualizing surveys, writing survey items, and formatting surveys. At the end of this briefing, the audience should have a better understanding of the value and purpose of surveys and how to construct them.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Jonathan Snavely, and Justin Mary. Introduction to Survey Design. IDA Document NS D-5835. Alexandria, VA: Institute for Defense Analyses, 2016.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5835.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Tutorial on Sensitivity Testing in Live Fire Test and Evaluation</title>
      <link>https://research.testscience.org/post/2016-tutorial-on-sensitivity-testing-in-live-fire-test-and-evaluation/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-tutorial-on-sensitivity-testing-in-live-fire-test-and-evaluation/</guid>
      <description>A sensitivity experiment is a special type of experimental design that is used when the response variable is binary and the covariate is continuous. Armor protection and projectile lethality tests often use sensitivity experiments to characterize a projectile&amp;rsquo;s probability of penetrating the armor. In this mini-tutorial we illustrate the challenge of modeling a binary response with a limited sample size, and show how sensitivity experiments can mitigate this problem. We review eight different single covariate sensitivity experiments and present a comparison of these designs using simulation.</description>
      <content:encoded><![CDATA[<p>A sensitivity experiment is a special type of experimental design that is used when the response variable is binary and the covariate is continuous. Armor protection and projectile lethality tests often use sensitivity experiments to characterize a projectile&rsquo;s probability of penetrating the armor. In this mini-tutorial we illustrate the challenge of modeling a binary response with a limited sample size, and show how sensitivity experiments can mitigate this problem. We review eight different single covariate sensitivity experiments and present a comparison of these designs using simulation. Additionally, we cover sensitivity experiments for cases that include more than one covariate, and highlight recent research in this area.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Thomas, Laura Freeman, and Raymond Chen. Tutorial on Sensitivity Testing in Live Fire Test and Evaluation. IDA Document NS D-5829. Alexandria, VA: Institute for Defense Analyses, 2016.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5829.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Best Practices for Statistically Validating Modeling and Simulation (M&amp;S) Tools Used in Operational Testing</title>
      <link>https://research.testscience.org/post/2015-best-practices-for-statistically-validating-modeling-and-simulation-m-s-tools-used-in-operational-testing/</link>
      <pubDate>Thu, 01 Jan 2015 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2015-best-practices-for-statistically-validating-modeling-and-simulation-m-s-tools-used-in-operational-testing/</guid>
      <description>In many situations, collecting sufficient data to evaluate system performance against operationally realistic threats is not possible due to cost and resource restrictions, safety concerns, or lack of adequate or representative threats. Modeling and simulation tools that have been verified, validated, and accredited can be used to supplement live testing in order to facilitate a more complete evaluation of performance. Two key questions that frequently arise when planning an operational test are (1) which (and how many) points within the operational space should be chosen in the simulation space and the live space for optimal ability to verify and validate the M&amp;amp;S, and (2) once that data is collected, what is the best way to compare the live trials to the simulated trials for the purpose of validating the M&amp;amp;S?</description>
      <content:encoded><![CDATA[<p>In many situations, collecting sufficient data to evaluate system performance against operationally realistic threats is not possible due to cost and resource restrictions, safety concerns, or lack of adequate or representative threats. Modeling and simulation tools that have been verified, validated, and accredited can be used to supplement live testing in order to facilitate a more complete evaluation of performance. Two key questions that frequently arise when planning an operational test are (1) which (and how many) points within the operational space should be chosen in the simulation space and the live space for optimal ability to verify and validate the M&amp;S, and (2) once that data is collected, what is the best way to compare the live trials to the simulated trials for the purpose of validating the M&amp;S? This conference presentation addresses various strategies for addressing these two questions. The best methodologies for designing and analyzing will vary depending on the goal of operational test, the type of model used in the simulation, and the amount of live and simulated data available.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Kelly, Laura Freeman, and Rebecca Medlin. Best Practices for Statistically Validating Modeling and Simulation (M&amp;S) Tools Used in Operational Testing. IDA Document NS D-5582. Alexandria, VA: Institute for Defense Analyses, 2015.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Estimating System Reliability from Heterogeneous Data</title>
      <link>https://research.testscience.org/post/2015-estimating-system-reliability-from-heterogeneous-data/</link>
      <pubDate>Thu, 01 Jan 2015 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2015-estimating-system-reliability-from-heterogeneous-data/</guid>
      <description>This briefing provides an example of some of the nuanced issues in reliability estimation in operational testing. The statistical models are motivated by an example of the Paladin Integrated Management (PIM). We demonstrate how to use a Bayesian approach to reliability estimation that uses data from all phases of testing.
Suggested Citation Browning, Caleb, Laura Freeman, Alyson Wilson, Kassandra Fronczyk, and Rebecca Dickinson. “Estimating System Reliability from Heterogeneous Data.” Presented at the Conference on Applied Statistics in Defense, George Mason University, October 2015.</description>
      <content:encoded><![CDATA[<p>This briefing provides an example of some of the nuanced issues in reliability estimation in operational testing.  The statistical models are motivated by an example of the Paladin Integrated Management (PIM).  We demonstrate how to use a Bayesian approach to reliability estimation that uses data from all phases of testing.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Browning, Caleb, Laura Freeman, Alyson Wilson, Kassandra Fronczyk, and Rebecca Dickinson. “Estimating System Reliability from Heterogeneous Data.” Presented at the Conference on Applied Statistics in Defense, George Mason University, October 2015.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Surveys in Operational Test and Evaluation</title>
      <link>https://research.testscience.org/post/2015-surveys-in-operational-test-and-evaluation/</link>
      <pubDate>Thu, 01 Jan 2015 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2015-surveys-in-operational-test-and-evaluation/</guid>
      <description>Recently DOT&amp;amp;E signed out a memo providing Guidance on the Use and Design of Surveys in Operational Test and Evaluation. This guidance memo helps the Human Systems Integration (HSI) community to ensure that useful and accurate HSI data are collected. Information about how HSI experts can leverage the guidance is presented. Specifically, the presentation will cover which HSI metrics can and cannot be answered by surveys.
Suggested Citation Grier, Rebecca A, and Laura Freeman.</description>
      <content:encoded><![CDATA[<p>Recently  DOT&amp;E signed out a memo providing Guidance on the Use and Design of Surveys in Operational Test and Evaluation. This guidance memo helps the Human Systems Integration (HSI) community to ensure that useful and accurate HSI data are collected. Information about how HSI experts can leverage the guidance is presented. Specifically, the presentation will cover which HSI metrics can and cannot be answered by surveys.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Grier, Rebecca A, and Laura Freeman. Surveys in Operational Test &amp; Evaluation. IDA Document D-5410. Alexandria, VA: Institute for Defense Analyses, 2015.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-5410-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Applying Risk Analysis to Acceptance Testing of Combat Helmets</title>
      <link>https://research.testscience.org/post/2014-applying-risk-analysis-to-acceptance-testing-of-combat-helmets/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-applying-risk-analysis-to-acceptance-testing-of-combat-helmets/</guid>
      <description>Acceptance testing of combat helmets presents multiple challenges that require statistically-sound solutions. For example, how should first article and lot acceptance tests treat multiple threats and measures of performance? How should these tests account for multiple helmet sizes and environmental treatments? How closely should first article testing requirements match historical or characterization test data? What government and manufacturer risks are acceptable during lot acceptance testing? Similar challenges arise when testing other components of Personal Protective Equipment and similar statistical approaches should be applied to all components.</description>
      <content:encoded><![CDATA[<p>Acceptance testing of combat helmets presents multiple challenges that require statistically-sound solutions. For example, how should first article and lot acceptance tests treat multiple threats and measures of performance? How should these tests account for multiple helmet sizes and environmental treatments? How closely should first article testing requirements match historical or characterization test data? What government and manufacturer risks are acceptable during lot acceptance testing? Similar challenges arise when testing other components of Personal Protective Equipment and similar statistical approaches should be applied to all components. This presentation explores these questions using operating characteristics curves and simulation studies.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Hester, Janice, and Laura Freeman. Applying Risk Analysis to Acceptance Testing of Combat Helmets. IDA Document NS D-5334. Alexandria, VA: Institute for Defense Analyses, 2014.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5334.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Design of Experiments for in-Lab Operational Testing of the an/BQQ-10 Submarine Sonar System</title>
      <link>https://research.testscience.org/post/2014-design-of-experiments-for-in-lab-operational-testing-of-the-an-bqq-10-submarine-sonar-system/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-design-of-experiments-for-in-lab-operational-testing-of-the-an-bqq-10-submarine-sonar-system/</guid>
      <description>Operational testing of the AN/BQQ-10 submarine sonar system has never been able to show significant improvements in software versions because of the high variability of at sea measurements. To mitigate this problem, in the most recent AN/BQQ-10 operational test, the Navy’s operational test agency (in consultation with IDA under the direction of Director, Operational Test and Evaluation) supplemented the at sea testing with an operationally focused in-lab comparison. This test used recorded real data played back on two different versions of the sonar system.</description>
      <content:encoded><![CDATA[<p>Operational testing of the AN/BQQ-10 submarine sonar system has never been able to show significant improvements in software versions because of the high variability of at sea measurements. To mitigate this problem, in the most recent AN/BQQ-10 operational test, the Navy’s operational test agency (in consultation with IDA under the direction of Director, Operational Test and Evaluation) supplemented the at sea testing with an operationally focused in-lab comparison. This test used recorded real data played back on two different versions of the sonar system. For each version, the test recorded the time it took multiple operations, with varying operational experience, to detect a submarine target once it appeared on the display. This new test methodology had several benefits: (1) the laboratory setting allowed for the use of design of experiments to control factors that are traditionally infeasible to control during an at sea test; (2) the direct comparison between the two systems resulted in demonstrating a statistically significant reduction in the detection time for the new system. Although laboratory testing cannot replace at sea testing, the results provide strong indication that we can expect performance improvements in the operational environment. This case study shows that laboratory testing and design of experiments have a place in operational testing and should be expanded to improve testing for other systems.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Clutter, Justace R, George Khoury, and Laura Freeman. Design of Experiments for In-Lab Operational Testing of the AN/BQQ-10 Submarine Sonar System. IDA Document NS D-5486. Alexandria, VA: Institute for Defense Analyses, 2014.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5286.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Power Analysis Tutorial for Experimental Design Software</title>
      <link>https://research.testscience.org/post/2014-power-analysis-tutorial-for-experimental-design-software/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-power-analysis-tutorial-for-experimental-design-software/</guid>
      <description>This guide provides both a general explanation of power analysis and specific guidance to successfully interface with two software packages, JMP and Design Expert (DX).
Suggested Citation Freeman, Laura J., Thomas H. Johnson, and James R. Simpson. “Power Analysis Tutorial for Experimental Design Software:” Fort Belvoir, VA: Defense Technical Information Center, November 1, 2014. https://doi.org/10.21236/ADA619843.
Paper: </description>
      <content:encoded><![CDATA[<p>This guide provides both a general explanation of power analysis and specific guidance to successfully interface with two software packages, JMP and Design Expert (DX).</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J., Thomas H. Johnson, and James R. Simpson. “Power Analysis Tutorial for Experimental Design Software:” Fort Belvoir, VA: Defense Technical Information Center, November 1, 2014. <a href="https://doi.org/10.21236/ADA619843">https://doi.org/10.21236/ADA619843</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Scientific Test and Analysis Techniques- Statistical Measures of Merit</title>
      <link>https://research.testscience.org/post/2013-scientific-test-and-analysis-techniques-statistical-measures-of-merit/</link>
      <pubDate>Tue, 01 Jan 2013 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2013-scientific-test-and-analysis-techniques-statistical-measures-of-merit/</guid>
      <description>Design of Experiments (DOE) provides a rigorous methodology for developing and evaluating test plans. Design excellence consists of having enough test points placed in the right locations in the operational envelope to answer the questions of interest for the test. The key aspects of a well-designed experiment include: the goal of the test, the response variables, the factors and levels, a method for strategically varying the factors across the operational envelope, and statistical measures of merit.</description>
      <content:encoded><![CDATA[<p>Design of Experiments (DOE) provides a rigorous methodology for developing and evaluating test plans. Design excellence consists of having enough test points placed in the right locations in the operational envelope to answer the questions of interest for the test. The key aspects of a well-designed experiment include: the goal of the test, the response variables, the factors and levels, a method for strategically varying the factors across the operational envelope, and statistical measures of merit. Currently, the majority of test plans utilize statistical measures of merit based on confidence and power. Although important, confidence and power are not the only measure of the adequacy and merit of a test design. The type of method that is appropriate is dependent on the goal of the test and the experimental design methodology used. There is no one-size-fits-all solution; rather there is a collection of useful tools that apply in various combinations for different test goals and designs. This talk outlines different statistical measures of merit that should be used when planning an operational test.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura. Scientific Test and Analysis Techniques: Statistical Measures of Merit. IDA Document D-5070. Alexandria, VA: Institute for Defense Analyses, 2014.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-5070-2.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
  </channel>
</rss>
