<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>2024 on Test Science Research Document Library</title>
    <link>https://research.testscience.org/year/2024/</link>
    <description>Recent content in 2024 on Test Science Research Document Library</description>
    <generator>Hugo -- 0.129.0</generator>
    <language>en-us</language>
    <copyright>Institute for Defense Analyses</copyright>
    <lastBuildDate>Mon, 01 Jan 2024 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://research.testscience.org/year/2024/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>A Practitioner’s Framework for Federated Model Validation Resource Allocation</title>
      <link>https://research.testscience.org/post/2024-a-practitioner-s-framework-for-federated-model-validation-resource-allocation/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-a-practitioner-s-framework-for-federated-model-validation-resource-allocation/</guid>
      <description>Recent advances in computation and statistics led to an increasing use of federated models for end-to-end system test and evaluation. A federated model is a collection of interconnected models where the outputs of a model act as inputs to subsequent models. However, the process of verifying and validating federated models is poorly understood, especially when testers have limited resources, knowledge-based uncertainties, and concerns over operational realism. Testers often struggle with determining how to best allocate limited test resources for model validation.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/owcIxrA_sXs?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Recent advances in computation and statistics led to an increasing use of federated models for end-to-end system test and evaluation. A federated model is a collection of interconnected models where the outputs of a model act as inputs to subsequent models. However, the process of verifying and validating federated models is poorly understood, especially when testers have limited resources, knowledge-based uncertainties, and concerns over operational realism. Testers often struggle with determining how to best allocate limited test resources for model validation. We propose a network-based representation of federated models, where the network encodes the connections between the federation of models. Nodes of the graph are given by sub-models. A directed edge from node a to node b is drawn if a inputs into b. We quantify their uncertainties through edge weights using meta-modeling and variance-based sensitivity analysis. The network-based framework allows us to propagate the uncertainties through the federated model and optimize resource allocation for validation based on the uncertainties.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Capp, Jo Anna, John T Haman, and Dhruv Patel. A Practitioner’s Framework for Federated Model Validation Resource Allocation. IDA Product ID 3001838. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Preview of Functional Data Analysis for Modeling and Simulation Validation</title>
      <link>https://research.testscience.org/post/2024-a-preview-of-functional-data-analysis-for-modeling-and-simulation-validation/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-a-preview-of-functional-data-analysis-for-modeling-and-simulation-validation/</guid>
      <description>Modeling and simulation (M&amp;amp;S) validation for operational testing often involves comparing live data with simulation outputs. Statistical methods known as functional data analysis (FDA) provides techniques for analyzing large data sets (&amp;ldquo;large&amp;rdquo; meaning that a single trial has a lot of information associated with it), such as radar tracks. We preview how FDA methods could assist M&amp;amp;S validation by providing statistical tools handling these large data sets. This may facilitate analyses that make use of more of the data available and thus allows for better detection of differences between M&amp;amp;S predictions and live test results.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/1neGQl8Jtxs?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Modeling and simulation (M&amp;S) validation for operational testing often involves comparing live data with simulation outputs. Statistical methods known as functional data analysis (FDA) provides techniques for analyzing large data sets (&ldquo;large&rdquo; meaning that a single trial has a lot of information associated with it), such as radar tracks. We preview how FDA methods could assist M&amp;S validation by providing statistical tools handling these large data sets. This may facilitate analyses that make use of more of the data available and thus allows for better detection of differences between M&amp;S predictions and live test results. We demonstrate some fundamental FDA approaches with a notional example of live and simulated radar tracks of a bomber’s flight</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Medlin, Rebecca M, and Curtis G Miller. A Preview of Functional Data Analysis for Modeling and Simulation Validation. IDA Product ID 3001829. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Reliability Assurance Test Planning and Analysis Tool</title>
      <link>https://research.testscience.org/post/2024-a-reliability-assurance-test-planning-and-analysis-tool/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-a-reliability-assurance-test-planning-and-analysis-tool/</guid>
      <description>This presentation documents the work of IDA 2024 Summer Associate Emma Mitchell. The work presented details an R Shiny application developed to provide a user-friendly software tool for researchers to use in planning for and analyzing system reliability. Specifically, the presentation details how one can plan for a reliability test using Bayesian Reliability Assurance test methods. Such tests utilize supplementary data and information, including reliability models, prior test results, expert judgment, and knowledge of environmental conditions, to plan for reliability testing, which in turn can often help in reducing the required amount of testing.</description>
      <content:encoded><![CDATA[<p>This presentation documents the work of IDA 2024 Summer Associate Emma Mitchell. The work presented details an R Shiny application developed to provide a user-friendly software tool for researchers to use in planning for and analyzing system reliability. Specifically, the presentation details how one can plan for a reliability test using Bayesian Reliability Assurance test methods. Such tests utilize supplementary data and information, including reliability models, prior test results, expert judgment, and knowledge of environmental conditions, to plan for reliability testing, which in turn can often help in reducing the required amount of testing. In the planning phase, the application enables researchers to use Bayesian methods to incorporate supplementary data when determining appropriate test lengths. In the analysis phase, the tool allows researchers to combine information through Bayesian methods, resulting in better uncertainty quantification than traditional methods.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, Rebecca M Medlin, Emma P Mitchell, Keyla Pagán-Rivera, and Dhruv K Patel. A Reliability Assurance Test Planning and Analysis Tool. IDA Product ID 3003359. Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "3003359%20Pagan-Rivera%20et%20al-3_slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Determining the Necessary Number of Runs in Computer Simulations with Binary Outcomes</title>
      <link>https://research.testscience.org/post/2024-determining-the-necessary-number-of-runs-in-computer-simulations-with-binary-outcomes/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-determining-the-necessary-number-of-runs-in-computer-simulations-with-binary-outcomes/</guid>
      <description>How many success-or-failure observations should we collect from a computer simulation? Often, researchers use space-filling design of experiments when planning modeling and simulation (M&amp;amp;S) studies. We are not satisfied with existing guidance on justifying the number of runs when developing these designs, either because the guidance is insufficiently justified, does not provide an unambiguous answer, or is not based on optimizing a statistical measure of merit. Analysts should use confidence interval margin of error as the statistical measure of merit for M&amp;amp;S studies intended to characterize overall M&amp;amp;S behavioral trends.</description>
      <content:encoded><![CDATA[<p>How many success-or-failure observations should we collect from a computer simulation? Often, researchers use space-filling design of experiments when planning modeling and simulation (M&amp;S) studies. We are not satisfied with existing guidance on justifying the number of runs when developing these designs, either because the guidance is insufficiently justified, does not provide an unambiguous answer, or is not based on optimizing a statistical measure of merit. Analysts should use confidence interval margin of error as the statistical measure of merit for M&amp;S studies intended to characterize overall M&amp;S behavioral trends. Unfortunately, the margin of error for studies involving factors and success-or-failure (or binary) outcomes requires knowing model parameters when using logistic regression. We explore how an upper bound on the margin of error, needing less information about the statistical model we need to estimate, can assist in sample size planning. While the upper bound needs further theoretical refinement, simulation studies suggest the upper bound may provide a means of justifying M&amp;S study sample sizes with a statistical measure of merit.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Duffy, Kelly, Curtis G Miller, and Rebecca Medlin. Sample Size Determination for Computer Simulations with Binary Outcomes. IDA Product 3002814. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Developing AI Trust- From Theory to Testing and the Myths in Between</title>
      <link>https://research.testscience.org/post/2024-developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/</guid>
      <description>This introductory work aims to provide members of the Test and Evaluation community with a clear understanding of trust and trustworthiness to support responsible and effective evaluation of AI systems. The paper provides a set of working definitions and works toward dispelling confusion and myths surrounding trust.
Suggested Citation Razin, Yosef S., and Kristen Alexander. “Developing AI Trust: From Theory to Testing and the Myths in Between.” The ITEA Journal of Test and Evaluation 45, no.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/xQL_kBiasPI?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>This introductory work aims to provide members of the Test and Evaluation community with a clear understanding of trust and trustworthiness to support responsible and effective evaluation of AI systems.  The paper provides a set of working definitions and works toward dispelling confusion and myths surrounding trust.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Razin, Yosef S., and Kristen Alexander. “Developing AI Trust: From Theory to Testing and the Myths in Between.” The ITEA Journal of Test and Evaluation 45, no. 1 (March 31, 2024). <a href="https://itea.org/journals/volume-45-1/developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/">https://itea.org/journals/volume-45-1/developing-ai-trust-from-theory-to-testing-and-the-myths-in-between/</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Human-Systems Interaction in Operational Test and Evaluation Course</title>
      <link>https://research.testscience.org/post/2024-introduction-to-human-systems-interaction-in-operational-test-and-evaluation-course/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-introduction-to-human-systems-interaction-in-operational-test-and-evaluation-course/</guid>
      <description>Human-System Interaction (HSI) is the study of interfaces between humans and technical systems. The Department of Defense incorporates HSI evaluations into defense acquisition to improve system performance and reduce lifecycle costs. During operational test and evaluation, HSI evaluations characterize how a system’s operational performance is affected by its users. The goal of this course is to provide the theoretical background and practical tools necessary to plan and evaluate HSI test plans, collect and analyze HSI data, and report on HSI results.</description>
      <content:encoded><![CDATA[<p>Human-System Interaction (HSI) is the study of interfaces between humans and technical systems. The Department of Defense incorporates HSI evaluations into defense acquisition to improve system performance and reduce lifecycle costs. During operational test and evaluation, HSI evaluations characterize how a system’s operational performance is affected by its users. The goal of this course is to provide the theoretical background and practical tools necessary to plan and evaluate HSI test plans, collect and analyze HSI data, and report on HSI results. We will discuss HSI concepts, measurement methods, design of experiments, data analysis, and evaluation and reporting, all from an operational testing perspective.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Miller, Dr Adam M, and Keyla Pagan-Rivera. Introduction to Human-Systems Interaction in Operational Test and Evaluation Course. IDA Product ID 3002009. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_3002009.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Meta-Analysis of the Effectiveness of the SALIANT Procedure for Assessing Team Situation Awareness</title>
      <link>https://research.testscience.org/post/2024-meta-analysis-of-the-effectiveness-of-the-saliant-procedure-for-assessing-team-situation-awareness/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-meta-analysis-of-the-effectiveness-of-the-saliant-procedure-for-assessing-team-situation-awareness/</guid>
      <description>Many Department of Defense (DoD) systems aim to increase or maintain Situational Awareness (SA) at the individual or group level. In some cases, maintenance or enhancement of SA is listed as a primary function or requirement of the system. However, during test and evaluation SA is examined inconsistently or is not measured at all. Situational Awareness Linked Indicators Adapted to Novel Tasks (SALIANT) is an empirically-based methodology meant to measure SA at the team, or group, level.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/Vmt1CT__stU?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Many Department of Defense (DoD) systems aim to increase or maintain Situational Awareness (SA) at the individual or group level. In some cases, maintenance or enhancement of SA is listed as a primary function or requirement of the system. However, during test and evaluation SA is examined inconsistently or is not measured at all. Situational Awareness Linked Indicators Adapted to Novel Tasks (SALIANT) is an empirically-based methodology meant to measure SA at the team, or group, level. While research using the SALIANT model suggests that it effectively quantifies team SA, no study has examined the effectiveness of SALIANT across the entirety of the existing empirical research.  The aim of the current work is to conduct a meta-analysis of previous research to examine the overall reliability of SALIANT as an SA measurement tool. This meta-analysis will assess when and how SALIANT can serve as a reliable indicator of performance at testing. Additional applications of SALIANT in non-traditional operational testing domains will also be discussed.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Shaffer, Sarah, Miriam Armstrong, and Rebecca Medlin. Meta-Analysis of the Effectiveness of the SALIANT Procedure for Assessing Team Situation Awareness. IDA Product ID 3001867. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Operational T&amp;E of AI-Supported Data Integration, Fusion, and Analysis Systems</title>
      <link>https://research.testscience.org/post/2024-operational-t-e-of-ai-supported-data-integration-fusion-and-analysis-systems/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-operational-t-e-of-ai-supported-data-integration-fusion-and-analysis-systems/</guid>
      <description>AI will play an important role in future military systems. However, large questions remain about how to test AI systems, especially in operational settings. Here, we discuss an approach for the operational test and evaluation (OT&amp;amp;E) of AI-supported data integration, fusion, and analysis systems. We highlight new challenges posed by AI-supported systems and we discuss new and existing OT&amp;amp;E methods for overcoming them. We demonstrate how to apply these OT&amp;amp;E methods via a notional test concept that focuses on evaluating an AI-supported data integration system in terms of its technical performance (how accurate is the AI output?</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/JqlIzJh-RQI?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>AI will play an important role in future military systems. However, large questions remain about how to test AI systems, especially in operational settings. Here, we discuss an approach for the operational test and evaluation (OT&amp;E) of AI-supported data integration, fusion, and analysis systems. We highlight new challenges posed by AI-supported systems and we discuss new and existing OT&amp;E methods for overcoming them. We demonstrate how to apply these OT&amp;E methods via a notional test concept that focuses on evaluating an AI-supported data integration system in terms of its technical performance (how accurate is the AI output?) and human systems interaction (how does the AI affect users?).</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Anderson, Breeana G, Adam M Miller, Logan K Ausman, John T Haman, Keyla Pagan-Rivera, Sarah A Shaffer, and Brian D Vickers. Data Integration, Fusion, and Analysis Systems. IDA Product ID 3001848. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Quantifying Uncertainty to Keep Astronauts and Warfighters Safe</title>
      <link>https://research.testscience.org/post/2024-quantifying-uncertainty-to-keep-astronauts-and-warfighters-safe/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-quantifying-uncertainty-to-keep-astronauts-and-warfighters-safe/</guid>
      <description>Both NASA and DOT&amp;amp;E increasingly rely on computer models to supplement data collection, and utilize statistical distributions to quantify the uncertainty in models, so that decision-makers are equipped with the most accurate information about system performance and model fitness. This article provides a high-level overview of uncertainty quantification (UQ) through an example assessment for the reliability of a new space-suit system. The goal is to reach a more general audience in Significance Magazine, and convey the importance and relevance of statistics to the defense and aerospace communities.</description>
      <content:encoded><![CDATA[<p>Both NASA and DOT&amp;E increasingly rely on computer models to supplement data collection, and utilize statistical distributions to quantify the uncertainty in models, so that decision-makers are equipped with the most accurate information about system performance and model fitness.  This article provides a high-level overview of uncertainty quantification (UQ) through an example assessment for the reliability of a new space-suit system.  The goal is to reach a more general audience in Significance Magazine, and convey the importance and relevance of statistics to the defense and aerospace communities.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Dennis, John W, John T Haman, and James E Warner. “Out-of-This-World Spacesuits: Quantifying Uncertainty Helps Keep Heroes Safe.” Significance 21, no. 4 (September 1, 2024): 10–13. <a href="https://doi.org/10.1093/jrssig/qmae056">https://doi.org/10.1093/jrssig/qmae056</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Sequential Space-Filling Designs for Modeling &amp; Simulation Analyses</title>
      <link>https://research.testscience.org/post/2024-sequential-space-filling-designs-for-modeling-simulation-analyses/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-sequential-space-filling-designs-for-modeling-simulation-analyses/</guid>
      <description>Space-filling designs (SFDs) are a rigorous method for designing modeling and simulation (M&amp;amp;S) studies. However, they are hindered by their requirement to choose the final sample size prior to testing. Sequential designs are an alternative that can increase test efficiency by testing small amounts of data at a time. We have conducted a literature review of existing sequential space-filling designs and found the methods most applicable to the test and evaluation (T&amp;amp;E) community.</description>
      <content:encoded><![CDATA[<p>Space-filling designs (SFDs) are a rigorous method for designing modeling and simulation (M&amp;S) studies. However, they are hindered by their requirement to choose the final sample size prior to testing. Sequential designs are an alternative that can increase test efficiency by testing small amounts of data at a time. We have conducted a literature review of existing sequential space-filling designs and found the methods most applicable to the test and evaluation (T&amp;E) community.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, and Anna Flowers. Sequential Space-Filling Designs for Modeling &amp; Simulation Analyses. IDA Product ID 3003752. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "3003752%20Haman%20et%20al-3_slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Simulation Insights on Power Analysis with Binary Responses--from SNR Methods to &#39;skprJMP&#39;</title>
      <link>https://research.testscience.org/post/2024-simulation-insights-on-power-analysis-with-binary-responses-from-snr-methods-to-skprjmp/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-simulation-insights-on-power-analysis-with-binary-responses-from-snr-methods-to-skprjmp/</guid>
      <description>Logistic regression is a commonly-used method for analyzing tests with probabilistic responses in the test community, yet calculating power for these tests has historically been challenging. This difficulty prompted the development of methods based on signal-to-noise ratio (SNR) approximations over the last decade, tailored to address the intricacies of logistic regression&amp;rsquo;s binary outcomes. However, advancements and improvements in statistical software and computational power have reduced the need for such approximate methods.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/j0rINL3L-yo?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Logistic regression is a commonly-used method for analyzing tests with probabilistic responses in the test community, yet calculating power for these tests has historically been challenging. This difficulty prompted the development of methods based on signal-to-noise ratio (SNR) approximations over the last decade, tailored to address the intricacies of logistic regression&rsquo;s binary outcomes. However, advancements and improvements in statistical software and computational power have reduced the need for such approximate methods. Our research presents a detailed simulation study that compares SNR-based power estimates with those derived from exact Monte Carlo simulations, highlighting the inadequacies of SNR approximations. To address these shortcomings, we will discuss improvements in the open-source R package &ldquo;skpr&rdquo; as well as present &ldquo;skprJMP,&rdquo; a new plug-in that offers more accurate and reliable power calculations for logistic regression analyses.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Atkins, Robert, Tyler Morgan-Wall, and Curtis Miller. “With Binary Responses&ndash;From SNR Methods to ‘skprJMP.’” Institute for Defense Analyses IDA Product ID 3002093 (April 2024).</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Statistical Advantages of Validated Surveys over Custom Surveys</title>
      <link>https://research.testscience.org/post/2024-statistical-advantages-of-validated-surveys-over-custom-surveys/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-statistical-advantages-of-validated-surveys-over-custom-surveys/</guid>
      <description>Surveys play an important role in quantifying user opinion during test and evaluation (T&amp;amp;E). Current best practice is to use surveys that have been tested, or “validated,” to ensure that they produce reliable and accurate results. However, unvalidated (“custom”) surveys are still widely used in T&amp;amp;E, raising questions about how to determine sample sizes for—and interpret data from— T&amp;amp;E events that rely on custom surveys. In this presentation, I characterize the statistical properties of validated and custom survey responses using data from recent T&amp;amp;E events, and then I demonstrate how these properties affect test design, analysis, and interpretation.</description>
      <content:encoded><![CDATA[<p>Surveys play an important role in quantifying user opinion during test and evaluation (T&amp;E). Current best practice is to use surveys that have been tested, or “validated,” to ensure that they produce reliable and accurate results. However, unvalidated (“custom”) surveys are still widely used in T&amp;E, raising questions about how to determine sample sizes for—and interpret data from— T&amp;E events that rely on custom surveys. In this presentation, I characterize the statistical properties of validated and custom survey responses using data from recent T&amp;E events, and then I demonstrate how these properties affect test design, analysis, and interpretation. I show that validated surveys reduce the number of subjects required to estimate statistical parameters or to detect a mean difference between two populations. Additionally, I simulate the survey process to demonstrate how poorly designed custom surveys introduce unintended changes to the data, increasing the risk of drawing false conclusions.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Bell, Jonathan L, and Adam M Miller. Statistical Advantages of Validated  Surveys over Custom Surveys. IDA Product ID 3001858. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Uncertainty Quantification for Ground Vehicle Vulnerability Simulation</title>
      <link>https://research.testscience.org/post/2024-uncertainty-quantification-for-ground-vehicle-vulnerability-simulation/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-uncertainty-quantification-for-ground-vehicle-vulnerability-simulation/</guid>
      <description>A vulnerability assessment of a combat vehicle uses modeling and simulation (M&amp;amp;S) to predict the vehicle&amp;rsquo;s vulnerability to a given enemy attack. The system-level output of the M&amp;amp;S is the probability that the vehicle&amp;rsquo;s mobility is degraded as a result of the attack. The M&amp;amp;S models this system-level phenomenon by decoupling the attack scenario into a hierarchy of sub-systems. Each sub-system addresses a specific scientific problem, such as the fracture dynamics of an exploded munition, or the ballistic resistance provided by the vehicle&amp;rsquo;s armor.</description>
      <content:encoded><![CDATA[<p>A vulnerability assessment of a combat vehicle uses modeling and simulation (M&amp;S) to predict the vehicle&rsquo;s vulnerability to a given enemy attack. The system-level output of the M&amp;S is the probability that the vehicle&rsquo;s mobility is degraded as a result of the attack. The M&amp;S models this system-level phenomenon by decoupling the attack scenario into a hierarchy of sub-systems. Each sub-system addresses a specific scientific problem, such as the fracture dynamics of an exploded munition, or the ballistic resistance provided by the vehicle&rsquo;s armor. For each sub-system in the hierarchy, laboratory testing is conducted to gather data to fit a subsystem-level model.  The M&amp;S hierarchically interconnects the subsystem-level models to enable prediction of the system-level output. As part of the DoD&rsquo;s ongoing effort to improve M&amp;S using verification, validation, and uncertainty quantification, we present a case study that propagates the uncertainties in the hierarchy of sub-models to the system-level output.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Thomas H., Dhruv K. Patel, John T. Haman, Jeremy S. Werner, and Dave Higdon. “Uncertainty Quantification for Ground Vehicle Vulnerability Simulation.” Quality Engineering, August 19, 2024. <a href="https://www.tandfonline.com/doi/abs/10.1080/08982112.2024.2394437">https://www.tandfonline.com/doi/abs/10.1080/08982112.2024.2394437</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
  </channel>
</rss>
