<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Design of Experiments on Test Science Research Document Library</title>
    <link>https://research.testscience.org/areas/design-of-experiments/</link>
    <description>Recent content in Design of Experiments on Test Science Research Document Library</description>
    <generator>Hugo -- 0.129.0</generator>
    <language>en-us</language>
    <copyright>Institute for Defense Analyses</copyright>
    <lastBuildDate>Mon, 01 Jan 2024 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://research.testscience.org/areas/design-of-experiments/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Determining the Necessary Number of Runs in Computer Simulations with Binary Outcomes</title>
      <link>https://research.testscience.org/post/2024-determining-the-necessary-number-of-runs-in-computer-simulations-with-binary-outcomes/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-determining-the-necessary-number-of-runs-in-computer-simulations-with-binary-outcomes/</guid>
      <description>How many success-or-failure observations should we collect from a computer simulation? Often, researchers use space-filling design of experiments when planning modeling and simulation (M&amp;amp;S) studies. We are not satisfied with existing guidance on justifying the number of runs when developing these designs, either because the guidance is insufficiently justified, does not provide an unambiguous answer, or is not based on optimizing a statistical measure of merit. Analysts should use confidence interval margin of error as the statistical measure of merit for M&amp;amp;S studies intended to characterize overall M&amp;amp;S behavioral trends.</description>
      <content:encoded><![CDATA[<p>How many success-or-failure observations should we collect from a computer simulation? Often, researchers use space-filling design of experiments when planning modeling and simulation (M&amp;S) studies. We are not satisfied with existing guidance on justifying the number of runs when developing these designs, either because the guidance is insufficiently justified, does not provide an unambiguous answer, or is not based on optimizing a statistical measure of merit. Analysts should use confidence interval margin of error as the statistical measure of merit for M&amp;S studies intended to characterize overall M&amp;S behavioral trends. Unfortunately, the margin of error for studies involving factors and success-or-failure (or binary) outcomes requires knowing model parameters when using logistic regression. We explore how an upper bound on the margin of error, needing less information about the statistical model we need to estimate, can assist in sample size planning. While the upper bound needs further theoretical refinement, simulation studies suggest the upper bound may provide a means of justifying M&amp;S study sample sizes with a statistical measure of merit.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Duffy, Kelly, Curtis G Miller, and Rebecca Medlin. Sample Size Determination for Computer Simulations with Binary Outcomes. IDA Product 3002814. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Sequential Space-Filling Designs for Modeling &amp; Simulation Analyses</title>
      <link>https://research.testscience.org/post/2024-sequential-space-filling-designs-for-modeling-simulation-analyses/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-sequential-space-filling-designs-for-modeling-simulation-analyses/</guid>
      <description>Space-filling designs (SFDs) are a rigorous method for designing modeling and simulation (M&amp;amp;S) studies. However, they are hindered by their requirement to choose the final sample size prior to testing. Sequential designs are an alternative that can increase test efficiency by testing small amounts of data at a time. We have conducted a literature review of existing sequential space-filling designs and found the methods most applicable to the test and evaluation (T&amp;amp;E) community.</description>
      <content:encoded><![CDATA[<p>Space-filling designs (SFDs) are a rigorous method for designing modeling and simulation (M&amp;S) studies. However, they are hindered by their requirement to choose the final sample size prior to testing. Sequential designs are an alternative that can increase test efficiency by testing small amounts of data at a time. We have conducted a literature review of existing sequential space-filling designs and found the methods most applicable to the test and evaluation (T&amp;E) community.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, and Anna Flowers. Sequential Space-Filling Designs for Modeling &amp; Simulation Analyses. IDA Product ID 3003752. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "3003752%20Haman%20et%20al-3_slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Simulation Insights on Power Analysis with Binary Responses--from SNR Methods to &#39;skprJMP&#39;</title>
      <link>https://research.testscience.org/post/2024-simulation-insights-on-power-analysis-with-binary-responses-from-snr-methods-to-skprjmp/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2024-simulation-insights-on-power-analysis-with-binary-responses-from-snr-methods-to-skprjmp/</guid>
      <description>Logistic regression is a commonly-used method for analyzing tests with probabilistic responses in the test community, yet calculating power for these tests has historically been challenging. This difficulty prompted the development of methods based on signal-to-noise ratio (SNR) approximations over the last decade, tailored to address the intricacies of logistic regression&amp;rsquo;s binary outcomes. However, advancements and improvements in statistical software and computational power have reduced the need for such approximate methods.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/j0rINL3L-yo?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Logistic regression is a commonly-used method for analyzing tests with probabilistic responses in the test community, yet calculating power for these tests has historically been challenging. This difficulty prompted the development of methods based on signal-to-noise ratio (SNR) approximations over the last decade, tailored to address the intricacies of logistic regression&rsquo;s binary outcomes. However, advancements and improvements in statistical software and computational power have reduced the need for such approximate methods. Our research presents a detailed simulation study that compares SNR-based power estimates with those derived from exact Monte Carlo simulations, highlighting the inadequacies of SNR approximations. To address these shortcomings, we will discuss improvements in the open-source R package &ldquo;skpr&rdquo; as well as present &ldquo;skprJMP,&rdquo; a new plug-in that offers more accurate and reliable power calculations for logistic regression analyses.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Atkins, Robert, Tyler Morgan-Wall, and Curtis Miller. “With Binary Responses&ndash;From SNR Methods to ‘skprJMP.’” Institute for Defense Analyses IDA Product ID 3002093 (April 2024).</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Comparing Normal and Binary D-Optimal Designs by Statistical Power</title>
      <link>https://research.testscience.org/post/2023-comparing-normal-and-binary-d-optimal-designs-by-statistical-power/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-comparing-normal-and-binary-d-optimal-designs-by-statistical-power/</guid>
      <description>In many Department of Defense test and evaluation applications, binary response variables are unavoidable. Many have considered D-optimal design of experiments for generalized linear models. However, little consideration has been given to assessing how these new designs perform in terms of statistical power for a given hypothesis test. Monte Carlo simulations and exact power calculations suggest that D optimal designs generally yield higher power than binary D-optimal designs, despite using logistic regression in the analysis after data have been collected.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/f1ChpOMzEWU?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>In many Department of Defense test and evaluation applications, binary response variables are unavoidable. Many have considered D-optimal design of experiments for generalized linear models. However, little consideration has been given to assessing how these new designs perform in terms of statistical power for a given hypothesis test. Monte Carlo simulations and exact power calculations suggest that D optimal designs generally yield higher power than binary D-optimal designs, despite using logistic regression in the analysis after data have been collected. Results from using statistical power to compare designs contradict standard design of experiments comparisons, which employ D-efficiency ratios and fractional design space plots. Power calculations suggest that practitioners that are primarily interested in the resulting statistical power of a design should use normal D optimal designs over binary D-optimal designs when logistic regression is to be used in the data analysis after data collection</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Medlin, Rebecca M, and Addison D Adams. Comparing Normal and Binary D-Optimal Design of Experiments by Statistical Power. IDA Document 3000032. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Development of Wald-Type and Score-Type Statistical Tests to Compare Live Test Data and Simulation Predictions</title>
      <link>https://research.testscience.org/post/2023-development-of-wald-type-and-score-type-statistical-tests-to-compare-live-test-data-and-simulation-predictions/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-development-of-wald-type-and-score-type-statistical-tests-to-compare-live-test-data-and-simulation-predictions/</guid>
      <description>This work describes the development of a statistical test created in support of ongoing verification, validation, and accreditation (VV&amp;amp;A) efforts for modeling and simulation (M&amp;amp;S) environments. The test computes a Wald-type statistic comparing two generalized linear models estimated from live test data and analogous simulated data. The resulting statistic indicates whether the M&amp;amp;S outputs differ from the live data. After developing the test, we applied it to two logistic regression models estimated from live torpedo test data and simulated data from the Naval Undersea Warfare Center’s Environment Centric Weapons Analysis Facility (ECWAF).</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/8OgvSuwTdys?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>This work describes the development of a statistical test created in support of ongoing verification, validation, and accreditation (VV&amp;A) efforts for modeling and simulation (M&amp;S) environments. The test computes a Wald-type statistic comparing two generalized linear models estimated from live test data and analogous simulated data. The resulting statistic indicates whether the M&amp;S outputs differ from the live data. After developing the test, we applied it to two logistic regression models estimated from live torpedo test data and simulated data from the Naval Undersea Warfare Center’s Environment Centric Weapons Analysis Facility (ECWAF). We developed this test to handle a specific problem with our data  one weapon variant was seen in the in-water test data, but the ECWAF data had two weapon variants. We overcame this deficiency by adjusting the Wald statistic via combining linear model coefficients with the intercept term when a factor is varied in one sample but not another. A similar approach could be applied with score-type tests, which we also describe.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Metts, Carrington, and Curtis Miller. “Development of Wald-Type and Score-Type Statistical Tests to Compare Live Test Data and Simulation Predictions.” The ITEA Journal of Test and Evaluation 44, no. 3 (August 25, 2023). <a href="https://itea.org/journals/volume-44-3/development-of-wald-type-and-score-type-statistical-tests-to-compare-live-test-data-and-simulation-predictions/">https://itea.org/journals/volume-44-3/development-of-wald-type-and-score-type-statistical-tests-to-compare-live-test-data-and-simulation-predictions/</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Implementing Fast Flexible Space-Filling Designs in R</title>
      <link>https://research.testscience.org/post/2023-implementing-fast-flexible-space-filling-designs-in-r/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-implementing-fast-flexible-space-filling-designs-in-r/</guid>
      <description>Modeling and simulation (M&amp;amp;S) can be a useful tool when testers and evaluators need to augment the data collected during a test event. When planning M&amp;amp;S, testers use experimental design techniques to determine how much and which types of data to collect, and they can use space-filling designs to spread out test points across the operational space. Fast flexible space-filling designs (FFSFDs) are a type of space-filling design useful for M&amp;amp;S because they work well in design spaces with disallowed combinations and permit the inclusion of categorical factors.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/vg36C3hhDmk?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Modeling and simulation (M&amp;S) can be a useful tool when testers and evaluators need to augment the data collected during a test event. When planning M&amp;S, testers use experimental design techniques to determine how much and which types of data to collect, and they can use space-filling designs to spread out test points across the operational space. Fast flexible space-filling designs (FFSFDs) are a type of space-filling design useful for M&amp;S because they work well in design spaces with disallowed combinations and permit the inclusion of categorical factors. IDA analysts developed a function to create FFSFDs using the free statistical software R. To our knowledge, there are no R packages for creating an FFSFD that can accommodate a variety of user inputs, such as categorical factors. Moreover, users of IDA’s function can share their code to make their work reproducible.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Medlin, Rebecca M, and Christopher T Dimapasok. Space-Filling Designs in R. IDA Document NS 3000045. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Design of Experiments for Testers</title>
      <link>https://research.testscience.org/post/2023-introduction-to-design-of-experiments-for-testers/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-introduction-to-design-of-experiments-for-testers/</guid>
      <description>This training provides details regarding the use of design of experiments, from choosing proper response variables, to identifying factors that could affect such responses, to determining the amount of data necessary to collect. The training also explains the benefits of using a Design of Experiments approach to testing and provides an overview of commonly used designs (e.g., factorial, optimal, and space-filling). The briefing illustrates the concepts discussed using several case studies.</description>
      <content:encoded><![CDATA[<p>This training provides details regarding the use of design of experiments, from choosing proper response variables, to identifying factors that could affect such responses, to determining the amount of data necessary to collect. The training also explains the benefits of using a Design of Experiments approach to testing and provides an overview of commonly used designs (e.g., factorial, optimal, and space-filling). The briefing illustrates the concepts discussed using several case studies.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, Breeana Anderson, Rebecca Medlin, Kelly M Avery, and Keyla Pagan-Rivera. I/ITSEC DOE Tutorial. IDA Document NS-D-33561. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Design of Experiments in R- Generating and Evaluating Designs with Skpr</title>
      <link>https://research.testscience.org/post/2023-introduction-to-design-of-experiments-in-r-generating-and-evaluating-designs-with-skpr/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-introduction-to-design-of-experiments-in-r-generating-and-evaluating-designs-with-skpr/</guid>
      <description>This workshop instructs attendees on how to run an end-to-end optimal Design of Experiments workflow in R using the open source skpr package. This workshop is split into two sections optimal design generation and design evaluation. The first half of the workshop provides basic instructions how to use R, as well as how to use skpr to create an optimal design for an experiment how to specify a model, create a candidate set of potential runs, remove disallowed combinations, and specify the design generation conditions to best suit an experimenter&amp;rsquo;s goals.</description>
      <content:encoded><![CDATA[<p>This workshop instructs attendees on how to run an end-to-end optimal Design of Experiments workflow in R using the open source skpr package. This workshop is split into two sections  optimal design generation and design evaluation. The first half of the workshop provides basic instructions how to use R, as well as how to use skpr to create an optimal design for an experiment  how to specify a model, create a candidate set of potential runs, remove disallowed combinations, and specify the design generation conditions to best suit an experimenter&rsquo;s goals.  The second half of the workshop covers design evaluation with skpr  how to determine if an experimental design is adequate for the test at hand. The workshop provides information on how to perform power calculations and evaluate other design properties that affect design quality. This also includes instruction on how to generate fraction of design space plots and correlation plots.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Morgan-Wall, Tyler T. Introduction to Design of Experiments in R: Generating and Evaluating Designs with Skpr. IDA Document NS D-33397. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Statistical Methods Development Work for M&amp;S Validation</title>
      <link>https://research.testscience.org/post/2023-statistical-methods-development-work-for-m-s-validation/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-statistical-methods-development-work-for-m-s-validation/</guid>
      <description>We discuss four areas in which statistically rigorous methods contribute to modeling and simulation validation studies. These areas are statistical risk analysis, space-filling experimental designs, metamodel construction, and statistical validation. Taken together, these areas implement DOT&amp;amp;E guidance on model validation. In each area, IDA has contributed either research methods, user-friendly tools, or both. We point to our tools on testscience.org, and survey the research methods that we&amp;rsquo;ve contributed to the M&amp;amp;S validation literature</description>
      <content:encoded><![CDATA[<p>We discuss four areas in which statistically rigorous methods contribute to modeling and simulation validation studies. These areas are statistical risk analysis, space-filling experimental designs, metamodel construction, and statistical validation. Taken together, these areas implement DOT&amp;E guidance on model validation. In each area, IDA has contributed either research methods, user-friendly tools, or both. We point to our tools on testscience.org, and survey the research methods that we&rsquo;ve contributed to the M&amp;S validation literature</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Miller, Curtis G. “Statistical Methods Development Work for M&amp;S Validation.” International Test and Evaluation Association 44, no. 3 (September 11, 2023). <a href="https://doi.org/10.61278/itea.44.3.1010">https://doi.org/10.61278/itea.44.3.1010</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Analysis Apps for the Operational Tester</title>
      <link>https://research.testscience.org/post/2022-analysis-apps-for-the-operational-tester/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-analysis-apps-for-the-operational-tester/</guid>
      <description>In the acquisition and testing world, data analysts repeatedly encounter certain categories of data, such as time or distance until an event (e.g., failure, alert, detection), binary outcomes (e.g., success/failure, hit/miss), and survey responses. Analysts need tools that enable them to produce quality and timely analyses of the data they acquire during testing. This poster presents four web-based apps that can analyze these types of data. The apps are designed to assist analysts and researchers with simple repeatable analysis tasks, such as building summary tables and plots for reports or briefings.</description>
      <content:encoded><![CDATA[<p>In the acquisition and testing world, data analysts repeatedly encounter certain categories of data, such as time or distance until an event (e.g., failure, alert, detection), binary outcomes (e.g., success/failure, hit/miss), and survey responses. Analysts need tools that enable them to produce quality and timely analyses of the data they acquire during testing. This poster presents four web-based apps that can analyze these types of data. The apps are designed to assist analysts and researchers with simple repeatable analysis tasks, such as building summary tables and plots for reports or briefings. Using software tools like these apps can increase reproducibility of results, timeliness of analysis and reporting, attractiveness and standardization of aesthetics in figures, and accuracy of results. The first app models reliability of a system or component by fitting parametric statistical distributions to time-to-failure data. The second app fits a logistic regression model to binary data with one or two independent continuous variables as The third calculates summary statistics and produces plots of groups of Likert-scale survey question responses. The fourth calculates the system usability scale (SUS) scores for SUS survey responses and enables the app user to plot scores versus an independent variable. These apps are available for public use on the Test Science Interactive Tools webpage <a href="https://testscience.org/interactive-tools/">https://testscience.org/interactive-tools/</a>.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Lillard, V Bram, and William Whitledge. Analysis Apps for the Operational Tester. IDA Document NS D-32959. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Case Study on Applying Sequential Analyses in Operational Testing</title>
      <link>https://research.testscience.org/post/2022-case-study-on-applying-sequential-analyses-in-operational-testing/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-case-study-on-applying-sequential-analyses-in-operational-testing/</guid>
      <description>Sequential analysis concerns statistical evaluation in which the number, pattern, or composition of the data is not determined at the start of the investigation, but instead depends on the information acquired during the investigation. Although sequential analysis originated in ballistics testing for the Department of Defense (DoD)and it is widely used in other disciplines, it is underutilized in the DoD. Expanding the use of sequential analysis may save money and reduce test time.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/gYTY5OJY4Yo?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Sequential analysis concerns statistical evaluation in which the number, pattern, or composition of the data is not determined at the start of the investigation, but instead depends on the information acquired during the investigation. Although sequential analysis originated in ballistics testing for the Department of Defense (DoD)and it is widely used in other disciplines, it is underutilized in the DoD. Expanding the use of sequential analysis may save money and reduce test time. In this paper, we introduce sequential analysis, describe its current and potential uses in operational test and evaluation (OT&amp;E), and present a method for applying it to the test and evaluation of defense systems. We evaluate the proposed method by performing simulation studies and applying the method to a case study. Additionally, we discuss challenges to address for sequential analysis in OT&amp;E. Lastly, while operational testing is the focus in this paper, the methodology presented is applicable to campaigns of experimentation and general testing across numerous disciplines.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Ahrens, Monica, Rebecca Medlin, Keyla Pagán-Rivera, and John W. Dennis. “Case Study on Applying Sequential Analyses in Operational Testing.” Quality Engineering 35, no. 3 (July 3, 2023): 534–45. <a href="https://doi.org/10.1080/08982112.2022.2146510">https://doi.org/10.1080/08982112.2022.2146510</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Metamodeling Techniques for Verification and Validation of Modeling and Simulation Data</title>
      <link>https://research.testscience.org/post/2022-metamodeling-techniques-for-verification-and-validation-of-modeling-and-simulation-data/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-metamodeling-techniques-for-verification-and-validation-of-modeling-and-simulation-data/</guid>
      <description>Modeling and simulation (M&amp;amp;S) outputs help the Director, Operational Test and Evaluation (DOT&amp;amp;E) assess the effectiveness, survivability, lethality, and suitability of systems. To use M&amp;amp;S outputs, DOT&amp;amp;E needs models and simulators to be sufficiently verified and validated. The purpose of this paper is to improve the state of verification and validation by recommending and demonstrating a set of statistical techniques—metamodels, also called statistical emulators—to the M&amp;amp;S community.
The paper expands on DOT&amp;amp;E’s existing guidance about metamodel usage by creating methodological recommendations the M&amp;amp;S community could apply to its activities.</description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/s4DUCI1M8Fw?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>Modeling and simulation (M&amp;S) outputs help the Director, Operational Test and Evaluation (DOT&amp;E) assess the effectiveness, survivability, lethality, and suitability of systems. To use M&amp;S outputs, DOT&amp;E needs models and simulators to be sufficiently verified and validated. The purpose of this paper is to improve the state of verification and validation by recommending and demonstrating a set of statistical techniques—metamodels, also called statistical emulators—to the M&amp;S community.</p>
<p>The paper expands on DOT&amp;E’s existing guidance about metamodel usage by creating methodological recommendations the M&amp;S community could apply to its activities. For a deterministic, discrete response variable, we recommend using a nearest neighbor or decision tree model. For a deterministic, continuous response variable, we recommend Gaussian process interpolation. For a stochastic response variable, we recommend a generalized additive model. We also present a set of techniques that testers can use to assess the adequacy of their metamodels. We conclude with a notional example that demonstrates the recommended techniques.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Haman, John T, and Curtis G Miller. Metamodeling Techniques for Verification and Validation of Modeling and Simulation Data. IDA Paper P-33230. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Thoughts on Applying Design of Experiments (DOE) to Cyber Testing</title>
      <link>https://research.testscience.org/post/2022-thoughts-on-applying-design-of-experiments-doe-to-cyber-testing/</link>
      <pubDate>Sat, 01 Jan 2022 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2022-thoughts-on-applying-design-of-experiments-doe-to-cyber-testing/</guid>
      <description>This briefing presented at Dataworks 2022 provides examples of potential ways in which Design of Experiments (DOE) could be applied to initially scope cyber assessments and, based on the results of those assessments, subsequently design in greater detail cyber tests.
Suggested Citation Gilmore, James M, Kelly M Avery, Matthew R Girardi, and Rebecca M Medlin. Thoughts on Applying Design of Experiments (DOE) to Cyber Testing. IDA Document NS D-33023. Alexandria, VA: Institute for Defense Analyses, 2022.</description>
      <content:encoded><![CDATA[<p>This briefing presented at Dataworks 2022 provides examples of potential ways in which Design of Experiments (DOE) could be applied to initially scope cyber assessments and, based on the results of those assessments, subsequently design in greater detail cyber tests.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Gilmore, James M, Kelly M Avery, Matthew R Girardi, and Rebecca M Medlin. Thoughts on Applying Design of Experiments (DOE) to Cyber Testing. IDA Document NS D-33023. Alexandria, VA: Institute for Defense Analyses, 2022.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-33023.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Determining How Much Testing is Enough- An Exploration of Progress in the Department of Defense Test and Evaluation Community</title>
      <link>https://research.testscience.org/post/2021-determining-how-much-testing-is-enough-an-exploration-of-progress-in-the-department-of-defense-test-and-evaluation-community/</link>
      <pubDate>Fri, 01 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2021-determining-how-much-testing-is-enough-an-exploration-of-progress-in-the-department-of-defense-test-and-evaluation-community/</guid>
      <description>This paper describes holistic progress in answering the question of “How much testing is enough?” It covers areas in which the T&amp;amp;E community has made progress, areas in which progress remains elusive, and issues that have emerged since 1994 that provide additional challenges. The selected case studies used to highlight progress are especially interesting examples, rather than a comprehensive look at all programs since 1994.
Suggested Citation Medlin, Rebecca, Matthew R Avery, James R Simpson, and Heather M Wojton.</description>
      <content:encoded><![CDATA[<p>This paper describes holistic progress in answering the question of “How much testing is enough?” It covers areas in which the T&amp;E community has made progress, areas in which progress remains elusive, and issues that have emerged since 1994 that provide additional challenges. The selected case studies used to highlight progress are especially interesting examples, rather than a comprehensive look at all programs since 1994.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Medlin, Rebecca, Matthew R Avery, James R Simpson, and Heather M Wojton. Determining How Much Testing Is Enough: An Exploration of Progress in the Department of Defense Test and Evaluation Community. IDA Document NS D-21561. Alexandria, VA: Institute for Defense Analyses, 2021.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Space-Filling Designs for Modeling &amp; Simulation</title>
      <link>https://research.testscience.org/post/2021-space-filling-designs-for-modeling-simulation/</link>
      <pubDate>Fri, 01 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2021-space-filling-designs-for-modeling-simulation/</guid>
      <description>This document presents arguments and methods for using space-filling designs (SFDs) to plan modeling and simulation (M&amp;amp;S) data collection.
Suggested Citation Avery, Kelly, John T Haman, Thomas Johnson, Curtis Miller, Dhruv Patel, and Han Yi. Test Design Challenges in Defense Testing. IDA Product ID 3002855. Alexandria, VA: Institute for Defense Analyses, 2024.
Slides: Paper: </description>
      <content:encoded><![CDATA[<p>This document presents arguments and methods for using space-filling designs (SFDs) to plan modeling and simulation (M&amp;S) data collection.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Kelly, John T Haman, Thomas Johnson, Curtis Miller, Dhruv Patel, and Han Yi. Test Design Challenges in Defense Testing. IDA Product ID 3002855. Alexandria, VA: Institute for Defense Analyses, 2024.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Why are Statistical Engineers Needed for Test &amp; Evaluation?</title>
      <link>https://research.testscience.org/post/2021-why-are-statistical-engineers-needed-for-test-evaluation/</link>
      <pubDate>Fri, 01 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2021-why-are-statistical-engineers-needed-for-test-evaluation/</guid>
      <description>The Department of Defense (DoD) develops and acquires some of the world’s most advanced and sophisticated systems. As new technologies emerge and are incorporated into systems, OSD/DOT&amp;amp;E faces the challenge of ensuring that these systems undergo adequate and efficient test and evaluation (T&amp;amp;E) prior to operational use. Statistical engineering is a collaborative, analytical approach to problem solving that integrates statistical thinking, methods, and tools with other relevant disciplines. The statistical engineering process provides better solutions to large, unstructured, real-world problems and supports rigorous decision-making.</description>
      <content:encoded><![CDATA[<p>The Department of Defense (DoD) develops and acquires some of the world’s most advanced and sophisticated systems. As new technologies emerge and are incorporated into systems, OSD/DOT&amp;E faces the challenge of ensuring that these systems undergo adequate and efficient test and evaluation (T&amp;E) prior to operational use. Statistical engineering is a collaborative, analytical approach to problem solving that integrates statistical thinking, methods, and tools with other relevant disciplines. The statistical engineering process provides better solutions to large, unstructured, real-world problems and supports rigorous decision-making. In this talk, we provide two case study examples related to looking at ways to improve approaches to integrate testing and data collection across the full system lifecycle. These case studies highlight why we believe statistical engineers are necessary for successful T&amp;E.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Medlin, Rebecca, Kayla Pagan-Rivera, and Monica Ahrens. Why Are Statistical Engineers Needed for Test &amp; Evaluation? IDA Document NS-D-22722. Alexandria, VA: Institute for Defense Analyses, 2021.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Review of Sequential Analysis</title>
      <link>https://research.testscience.org/post/2020-a-review-of-sequential-analysis/</link>
      <pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2020-a-review-of-sequential-analysis/</guid>
      <description>Sequential analysis concerns statistical evaluation in situations in which the number, pattern, or composition of the data is not determined at the start of the investigation, but instead depends upon the information acquired throughout the course of the investigation. Expanding the use of sequential analysis has the potential to save resources and reduce test time (National Research Council, 1998). This paper summarizes the literature on sequential analysis and offers fundamental information for providing recommendations for its use in DoD test and evaluation.</description>
      <content:encoded><![CDATA[<p>Sequential analysis concerns statistical evaluation in situations in which the number, pattern, or composition of the data is not determined at the start of the investigation, but instead depends upon the information acquired throughout the course of the investigation. Expanding the use of sequential analysis has the potential to save resources and reduce test time (National Research Council, 1998). This paper summarizes the literature on sequential analysis and offers fundamental information for providing recommendations for its use in DoD test and evaluation.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Rebecca Medlin, John Dennis, Keyla Pagan-Rivera, and Leonard Wilkins. A Review of Sequential Analysis. IDA Document NS D-20487. Alexandria, VA: Institute for Defense Analyses, 2020.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper_seq_review.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Challenges and New Methods for Designing Reliability Experiments</title>
      <link>https://research.testscience.org/post/2019-challenges-and-new-methods-for-designing-reliability-experiments/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-challenges-and-new-methods-for-designing-reliability-experiments/</guid>
      <description>Engineers use reliability experiments to determine the factors that drive product reliability, build robust products, and predict reliability under use conditions. This article uses recent testing of a Howitzer to illustrate the challenges in designing reliability experiments for complex, repairable systems. We leverage lessons learned from current research and propose methods for designing an experiment for a complex, repairable system.
Suggested Citation Freeman, Laura J., Rebecca M. Medlin, and Thomas H.</description>
      <content:encoded><![CDATA[<p>Engineers use reliability experiments to determine the factors that drive product reliability, build robust products, and predict reliability under use conditions. This article uses recent testing of a Howitzer to illustrate the challenges in designing reliability experiments for complex, repairable systems. We leverage lessons learned from current research and propose methods for designing an experiment for a complex, repairable system.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J., Rebecca M. Medlin, and Thomas H. Johnson. “Challenges and New Methods for Designing Reliability Experiments.” Quality Engineering 31, no. 1 (January 2, 2019): 108–21. <a href="https://doi.org/10.1080/08982112.2018.1546394">https://doi.org/10.1080/08982112.2018.1546394</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>D-Optimal as an Alternative to Full Factorial Designs- a Case Study</title>
      <link>https://research.testscience.org/post/2019-d-optimal-as-an-alternative-to-full-factorial-designs-a-case-study/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-d-optimal-as-an-alternative-to-full-factorial-designs-a-case-study/</guid>
      <description>The use of Bayesian statistics and experimental design as tools to scope testing and analyze data related to defense has increased in recent years. Planning a test using experimental design will allow testers to cover the operational space while maximizing the information obtained from each run. Understanding which factors can affect a detector&amp;rsquo;s performance can influence military tactics, techniques and procedures, and improve a commander&amp;rsquo;s situational awareness when making decisions in an operational environment.</description>
      <content:encoded><![CDATA[<p>The use of Bayesian statistics and experimental design as tools to scope testing and analyze data related to defense has increased in recent years. Planning a test using experimental design will allow testers to cover the operational space while maximizing the information obtained from each run. Understanding which factors can affect a detector&rsquo;s performance can influence military tactics, techniques and procedures, and improve a commander&rsquo;s situational awareness when making decisions in an operational environment. This presentation will explain how a D-optimal experimental design could be an option for planning a test when the number of runs is limited but an adequate test is desired. Additionally, it will describe how the results of a Bayesian multiple logistic model could be used to show in what way the operational environment can affect the detector&rsquo;s performance.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Anderson, Breeana G, Heather M Wojton, and Keyla Pagan-Rivera. D-Optimal as an Alternative to Full Factorial Designs: A Case Study. IDA Document NS D-10580. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="poster">Poster:</h4>
<embed src= "poster.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Designing Experiments for Model Validation- The Foundations for Uncertainty Quantification</title>
      <link>https://research.testscience.org/post/2019-designing-experiments-for-model-validation-the-foundations-for-uncertainty-quantification/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-designing-experiments-for-model-validation-the-foundations-for-uncertainty-quantification/</guid>
      <description>Advances in computational power have allowed both greater fidelity and more extensive use of such models. Numerous complex military systems have a corresponding model that simulates its performance in the field. In response, the DoD needs defensible practices for validating these models. Design of Experiments and statistical analysis techniques are the foundational building blocks for validating the use of computer models and quantifying uncertainty in that validation. Recent developments in uncertainty quantification have the potential to benefit the DoD in using modeling and simulation to inform operational evaluations.</description>
      <content:encoded><![CDATA[<p>Advances in computational power have allowed both greater fidelity and more extensive use of such models. Numerous complex military systems have a corresponding model that simulates its performance in the field. In response, the DoD needs defensible practices for validating these models. Design of Experiments and statistical analysis techniques are the foundational building blocks for validating the use of computer models and quantifying uncertainty in that validation. Recent developments in uncertainty quantification have the potential to benefit the DoD in using modeling and simulation to inform operational evaluations.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Kelly Avery, Laura Freeman, and Thomas Johnson. “Designing Experiments for Model Validation – The Foundations for Uncertainty Quantification.” The  ITEA Journal of Test and Evaluation 40, no. 1 (2019).</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Handbook on Statistical Design &amp; Analysis Techniques for Modeling &amp; Simulation Validation</title>
      <link>https://research.testscience.org/post/2019-handbook-on-statistical-design-analysis-techniques-for-modeling-simulation-validation/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-handbook-on-statistical-design-analysis-techniques-for-modeling-simulation-validation/</guid>
      <description>This handbook focuses on methods for data-driven validation to supplement the vast existing literature for Verification, Validation, and Accreditation (VV&amp;amp;A) and the emerging references on uncertainty quantification (UQ). The goal of this handbook is to aid the test and evaluation (T&amp;amp;E) community in developing test strategies that support model validation (both external validation and parametric analysis) and statistical UQ.
Suggested Citation Wojton, Heather, Kelly M Avery, Laura J Freeman, Samuel H Parry, Gregory S Whittier, Thomas H Johnson, and Andrew C Flack.</description>
      <content:encoded><![CDATA[<p>This handbook focuses on methods for data-driven validation to supplement the vast existing literature for Verification, Validation, and Accreditation (VV&amp;A) and the emerging references on uncertainty quantification (UQ). The goal of this handbook is to aid the test and evaluation (T&amp;E) community in developing test strategies that support model validation (both external validation and parametric analysis) and statistical UQ.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wojton, Heather, Kelly M Avery, Laura J Freeman, Samuel H Parry, Gregory S Whittier, Thomas H Johnson, and Andrew C Flack. Handbook on Statistical Design &amp; Analysis Techniques for Modeling &amp; Simulation Validation. IDA Document NS D-10455. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Sample Size Determination Methods Using Acceptance Sampling by Variables</title>
      <link>https://research.testscience.org/post/2019-sample-size-determination-methods-using-acceptance-sampling-by-variables/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-sample-size-determination-methods-using-acceptance-sampling-by-variables/</guid>
      <description>Acceptance Sampling by Variables (ASbV) is a statistical testing technique used in Personal Protective Equipment programs to determine the quality of the equipment in First Article and Lot Acceptance Tests. This article intends to remedy the lack of existing references that discuss the similarities between ASbV and certain techniques used in different sub-disciplines within statistics. Understanding ASbV from a statistical perspective allows testers to create customized test plans, beyond what is available in MIL-STD-414.</description>
      <content:encoded><![CDATA[<p>Acceptance Sampling by Variables (ASbV) is a statistical testing technique used in Personal Protective Equipment programs to determine the quality of the equipment in First Article and Lot Acceptance Tests. This article intends to remedy the lack of existing references that discuss the similarities between ASbV and certain techniques used in different sub-disciplines within statistics. Understanding ASbV from a statistical perspective allows testers to create customized test plans, beyond what is available in MIL-STD-414.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Walzl, Kerry, Lindsey A Davis, Thomas H Johnson, and Heather M Wojton. Sample Size Determination Methods Using Acceptance Sampling by Variables. IDA Document NS D-10666. Alexandria, VA: Institute for Defense Analyses, 2019.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>The Effect of Extremes in Small Sample Size on Simple Mixed Models- A Comparison of Level-1 and Level-2 Size</title>
      <link>https://research.testscience.org/post/2019-the-effect-of-extremes-in-small-sample-size-on-simple-mixed-models-a-comparison-of-level-1-and-level-2-size/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-the-effect-of-extremes-in-small-sample-size-on-simple-mixed-models-a-comparison-of-level-1-and-level-2-size/</guid>
      <description>We present a simulation study that examines the impact of small sample sizes in both observation and nesting levels of the model on the fixed effect bias, type I error, and the power of a simple mixed model analysis. Despite the need for adjustments to control for type I error inflation, our findings indicate that smaller samples than previously recognized can be used for mixed models under certain conditions prevalent in applied research.</description>
      <content:encoded><![CDATA[<p>We present a simulation study that examines the impact of small sample sizes in both observation and nesting levels of the model on the fixed effect bias, type I error, and the power of a simple mixed model analysis. Despite the need for adjustments to control for type I error inflation, our findings indicate that smaller samples than previously recognized can be used for mixed models under certain conditions prevalent in applied research.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Carter, Kristina A, Heather M Wojton, and Stephanie T Lane. “The Effect of Extremes in Small Sample Size on Simple Mixed Models: A Comparison of Level-1 and Level-2 Size.” The ITEA Journal of Test and Evaluation 40, no. 1 (2019): 16–29.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Use of Design of Experiments in Survivability Testing</title>
      <link>https://research.testscience.org/post/2019-use-of-design-of-experiments-in-survivability-testing/</link>
      <pubDate>Tue, 01 Jan 2019 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2019-use-of-design-of-experiments-in-survivability-testing/</guid>
      <description>The purpose of survivability testing is to provide decision makers with relevant, credible evidence about the survivability of an aircraft that is conveyed with some degree of certainty or inferential weight. In developing an experiment to accomplish this goal, a test planner faces numerous questions What critical issue or issues are being address? What data are needed to answer the critical issues? What test conditions should be varied? What is the most economical way of varying those conditions?</description>
      <content:encoded><![CDATA[<p>The purpose of survivability testing is to provide decision makers with relevant, credible evidence about the survivability of an aircraft that is conveyed with some degree of certainty or inferential weight. In developing an experiment to accomplish this goal, a test planner faces numerous questions  What critical issue or issues are being address? What data are needed to answer the critical issues? What test conditions should be varied? What is the most economical way of varying those conditions? How many test articles are needed? Design of Experiments provides an analytical basis for test planning tradeoffs when answering these questions.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Couch, Mark, John Haman, Thomas Johnson, and Heather Wojton. “Designs of Experiments (DOE) in Survivability Testing.” Joint Aircraft Survivability Program - JASP Online (blog), March 2019. <a href="https://www.jasp-online.org/asjournal/summer-2019/designs-of-experiments-doe-in-survivability-testing/">https://www.jasp-online.org/asjournal/summer-2019/designs-of-experiments-doe-in-survivability-testing/</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Informing the Warfighter—Why Statistical Methods Matter in Defense Testing</title>
      <link>https://research.testscience.org/post/2018-informing-the-warfighter-why-statistical-methods-matter-in-defense-testing/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-informing-the-warfighter-why-statistical-methods-matter-in-defense-testing/</guid>
      <description>Needs one
Suggested Citation Freeman, Laura J., and Catherine Warner. “Informing the Warfighter—Why Statistical Methods Matter in Defense Testing.” CHANCE 31, no. 2 (April 3, 2018): 4–11. https://doi.org/10.1080/09332480.2018.1467627.
Paper: </description>
      <content:encoded><![CDATA[<p>Needs one</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J., and Catherine Warner. “Informing the Warfighter—Why Statistical Methods Matter in Defense Testing.” CHANCE 31, no. 2 (April 3, 2018): 4–11. <a href="https://doi.org/10.1080/09332480.2018.1467627">https://doi.org/10.1080/09332480.2018.1467627</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Introduction to Observational Studies</title>
      <link>https://research.testscience.org/post/2018-introduction-to-observational-studies/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-introduction-to-observational-studies/</guid>
      <description> A presentation on the theory and practice of observational studies. Specific average treatment effect methods include matching, difference-in-difference estimators, and instrumental variables.
Suggested Citation Thomas, Dean, and Yevgeniya K Pinelis. Introduction to Observational Studies. IDA Document NS D-9020. Alexandria, VA: Institute for Defense Analyses, 2018.
Slides: </description>
      <content:encoded><![CDATA[

    
    <div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
      <iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/csO17jA2cxI?autoplay=0&controls=1&end=0&loop=0&mute=0&start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
      ></iframe>
    </div>

<p>A presentation on the theory and practice of observational studies.  Specific average treatment effect methods include matching, difference-in-difference estimators, and instrumental variables.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Thomas, Dean, and Yevgeniya K Pinelis. Introduction to Observational Studies. IDA Document NS D-9020. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>JEDIS Briefing and Tutorial</title>
      <link>https://research.testscience.org/post/2018-jedis-briefing-and-tutorial/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-jedis-briefing-and-tutorial/</guid>
      <description>Are you sick of having to manually iterate your way through sizing your design of experiments? Come learn about JEDIS, the new IDA-developed JMP Add-In for automating design of experiments power calculations. JEDIS builds multiple test designs in JMP over user-specified ranges of sample sizes, Signal-to-Noise Ratios (SNR), and alpha (1 -confidence) levels. It then automatically calculates the statistical power to detect an effect due to each factor and any specified interactions for each design.</description>
      <content:encoded><![CDATA[<p>Are you sick of having to manually iterate your way through sizing your design of experiments? Come learn about JEDIS, the new IDA-developed JMP Add-In for automating design of experiments power calculations. JEDIS builds multiple test designs in JMP over user-specified ranges of sample sizes, Signal-to-Noise Ratios (SNR), and alpha (1 -confidence) levels. It then automatically calculates the statistical power to detect an effect due to each factor and any specified interactions for each design. When finished, JEDIS presents the statistical power vs. design metrics in interactive plots and stores the data in an easy to use format. JEDIS creates factorial and optimal designs, but does not currently support split plot designs. If you already have a pre-made design table, the JEDIS Light feature can compute power for the design over ranges of SNR and alpha levels.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Pechkis, Daniel, and Jason P Sheldon. JEDIS Briefing and Tutorial. IDA Document NS D-8964. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Scientific Test and Analysis Techniques</title>
      <link>https://research.testscience.org/post/2018-scientific-test-and-analysis-techniques/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-scientific-test-and-analysis-techniques/</guid>
      <description>Abstract This document contains the technical content for the Scientific Test and Analysis Techniques (STAT) in Test and Evaluation (T&amp;amp;E) continuous learning module. The module provides a basic understanding of STAT in T&amp;amp;E. Topics coverec include design of experiments, observational studies, survey design and analysis, and statistical analysis. It is designed as a four hour online course, suitable for inclusion in the DAU T&amp;amp;E certification curriculum.
Slides </description>
      <content:encoded><![CDATA[<h3 id="abstract">Abstract</h3>
<p>This document contains the technical content for the Scientific Test and Analysis Techniques (STAT) in Test and Evaluation (T&amp;E) continuous learning module. The module provides a basic understanding of STAT in T&amp;E. Topics coverec include design of experiments, observational studies, survey design and analysis, and statistical analysis. It is designed as a four hour online course, suitable for inclusion in the DAU T&amp;E certification curriculum.</p>
<h3 id="slides-hahahugoshortcode75s0hbhb">Slides <embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >
</h3>
]]></content:encoded>
    </item>
    <item>
      <title>Scientific Test and Analysis Techniques- Continuous Learning Module</title>
      <link>https://research.testscience.org/post/2018-scientific-test-and-analysis-techniques-continuous-learning-module/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-scientific-test-and-analysis-techniques-continuous-learning-module/</guid>
      <description>This document contains the technical content for the Scientific Test and Analysis Techniques (STAT) in Test and Evaluation (T&amp;amp;E) continuous learning module. The module provides a basic understanding of STAT in T&amp;amp;E. Topics covered include design of experiments, observational studies, survey design and analysis, and statistical analysis. It is designed as a four hour online course, suitable for inclusion in the DAU T&amp;amp;E certification curriculum.
Suggested Citation Pinelis, Yevgeniya, Laura J Freeman, Heather M Wojton, Denise J Edwards, Stephanie T Lane, and James R Simpson.</description>
      <content:encoded><![CDATA[<p>This document contains the technical content for the Scientific Test and Analysis Techniques (STAT) in Test and Evaluation (T&amp;E) continuous learning module. The module provides a basic understanding of STAT in T&amp;E. Topics covered include design of experiments, observational studies, survey design and analysis, and statistical analysis. It is designed as a four hour online course, suitable for inclusion in the DAU T&amp;E certification curriculum.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Pinelis, Yevgeniya, Laura J Freeman, Heather M Wojton, Denise J Edwards, Stephanie T Lane, and James R Simpson. Scientific Test and Analysis Techniques: Continuous Learning Module. IDA  Document NS D-892. Alexandria, VA: Institute for Defense Analyses, 2018.</p>
</blockquote>
]]></content:encoded>
    </item>
    <item>
      <title>Testing Defense Systems</title>
      <link>https://research.testscience.org/post/2018-testing-defense-systems/</link>
      <pubDate>Mon, 01 Jan 2018 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2018-testing-defense-systems/</guid>
      <description>The complex, multifunctional nature of defense systems, along with the wide variety of system types, demands a structured but flexible analytical process for testing systems. This chapter summarizes commonly used techniques in defense system testing and specific challenges imposed by the nature of defense system testing. It highlights the core statistical methodologies that have proven useful in testing defense systems. Case studies illustrate the value of using statistical techniques in the design of tests and analysis of the resulting data.</description>
      <content:encoded><![CDATA[<p>The complex, multifunctional nature of defense systems, along with the wide variety of system types, demands a structured but flexible analytical process for testing systems. This chapter summarizes commonly used techniques in defense system testing and specific challenges imposed by the nature of defense system testing. It highlights the core statistical methodologies that have proven useful in testing defense systems. Case studies illustrate the value of using statistical techniques in the design of tests and analysis of the resulting data. The chapter focuses on the unique statistical challenges of designing operational tests, many of which can be attributed to the process, but some of which are inherent to the complexity of the systems and the missions system operators must complete. It provides an overview of the process of designing experiments for military systems with operational users in an operational environment.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J., Thomas Johnson, Matthew Avery, V. Bram Lillard, and Justace Clutter. “Testing Defense Systems.” In Analytic Methods in Systems and Software Testing, 439–87. John Wiley &amp; Sons, Ltd, 2018. <a href="https://doi.org/10.1002/9781119357056.ch18">https://doi.org/10.1002/9781119357056.ch18</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>On Scoping a Test that Addresses the Wrong Objective</title>
      <link>https://research.testscience.org/post/2017-on-scoping-a-test-that-addresses-the-wrong-objective/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-on-scoping-a-test-that-addresses-the-wrong-objective/</guid>
      <description>Statistical literature refers to a type of error that is committed by giving the right answer to the wrong question. If a test design is adequately scoped to address an irrelevant objective, one could say that a Type III error occurs. In this paper, we focus on a specific Type III error that on some occasions test planners commit to reduce test size and resources.
Suggested Citation Johnson, Thomas H., Rebecca M.</description>
      <content:encoded><![CDATA[<p>Statistical literature refers to a type of error that is committed by giving the right answer to the wrong question. If a test design is adequately scoped to address an irrelevant objective, one could say that a Type III error occurs. In this paper, we focus on a specific Type III error that on some occasions test planners commit to reduce test size and resources.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Thomas H., Rebecca M. Medlin, Laura J. Freeman, and James R. Simpson. “On Scoping a Test That Addresses the Wrong Objective.” Quality Engineering 31, no. 2 (April 3, 2019): 230–39. <a href="https://doi.org/10.1080/08982112.2018.1479035">https://doi.org/10.1080/08982112.2018.1479035</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Perspectives on Operational Testing-Guest Lecture at Naval Postgraduate School</title>
      <link>https://research.testscience.org/post/2017-perspectives-on-operational-testing-guest-lecture-at-naval-postgraduate-school/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-perspectives-on-operational-testing-guest-lecture-at-naval-postgraduate-school/</guid>
      <description>This document was prepared to support Dr. Lillard&amp;rsquo;s visit to the NavalPostgraduate School where he will provide a guest lecture to students in the T&amp;amp;Ecourse. The briefing covers three primary themes: 1) evaluation of military systemson the basis of requirements and KPPs alone is often insufficient to determineeffectiveness and suitability in combat conditions, 2) statistical methods are essentialfor developing defensible and rigorous test designs, 3) operational testing is often theonly means to discover critical performance shortcomings.</description>
      <content:encoded><![CDATA[<p>This document was prepared to support Dr. Lillard&rsquo;s visit to the NavalPostgraduate School where he will provide a guest lecture to students in the T&amp;Ecourse. The briefing covers three primary themes: 1) evaluation of military systemson the basis of requirements and KPPs alone is often insufficient to determineeffectiveness and suitability in combat conditions, 2) statistical methods are essentialfor developing defensible and rigorous test designs, 3) operational testing is often theonly means to discover critical performance shortcomings.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Lillard, Vincent A. Perspectives on Operational Testing: Guest Lecture at Naval Postgraduate School. IDA Document D-8333-NS. Alexandria, VA: Institute for Defense Analyses, 2017.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-8333.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Power Approximations for Generalized Linear Models using the Signal-to-Noise Transformation Method</title>
      <link>https://research.testscience.org/post/2017-power-approximations-for-generalized-linear-models-using-the-signal-to-noise-transformation-method/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-power-approximations-for-generalized-linear-models-using-the-signal-to-noise-transformation-method/</guid>
      <description>Statistical power is a useful measure for assessing the adequacy of anexperimental design prior to data collection. This paper proposes an approach referredto as the signal-to-noise transformation method (SNRx), to approximate power foreffects in a generalized linear model. The contribution of SNRx is that, with a coupleassumptions, it generates power approximations for generalized linear model effectsusing F-tests that are typically used in ANOVA for classical linear models.Additionally, SNRx follows Ohlert and Whitcomb&amp;rsquo;s unified approach for sizing aneffect, which allows for intuitive effect size definitions, and consistent estimates ofpower.</description>
      <content:encoded><![CDATA[<p>Statistical power is a useful measure for assessing the adequacy of anexperimental design prior to data collection. This paper proposes an approach referredto as the signal-to-noise transformation method (SNRx), to approximate power foreffects in a generalized linear model. The contribution of SNRx is that, with a coupleassumptions, it generates power approximations for generalized linear model effectsusing F-tests that are typically used in ANOVA for classical linear models.Additionally, SNRx follows Ohlert and Whitcomb&rsquo;s unified approach for sizing aneffect, which allows for intuitive effect size definitions, and consistent estimates ofpower. This paper details the process for defining an effect size, constructing thecoefficients for the test, and calculating power for the family of generalized linearmodels. The focus is on experimental designs that have multi-level categorical factors. A simulation study is performed which demonstrates that SNRx power results agreewith simulation.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Thomas H., Laura Freeman, Jim Simpson, and Colin Anderson. “Power Approximations for Generalized Linear Models Using the Signal-to-Noise Transformation Method.” Quality Engineering 30, no. 3 (July 3, 2018): 511–24. <a href="https://doi.org/10.1080/08982112.2017.1361537">https://doi.org/10.1080/08982112.2017.1361537</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-8351.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Statistical Methods for Defense Testing</title>
      <link>https://research.testscience.org/post/2017-statistical-methods-for-defense-testing/</link>
      <pubDate>Sun, 01 Jan 2017 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2017-statistical-methods-for-defense-testing/</guid>
      <description>In the increasingly complex and data‐limited world of military defense testing, statisticians play a valuable role in many applications. Before the DoD acquires any major new capability, that system must undergo realistic testing in its intended environment with military users. Although the typical test environment is highly variable and factors are often uncontrolled, design of experiments techniques can add objectivity, efficiency, and rigor to the process of test planning. Statistical analyses help system evaluators get the most information out of limited data sets.</description>
      <content:encoded><![CDATA[<p>In the increasingly complex and data‐limited world of military defense testing, statisticians play a valuable role in many applications. Before the DoD acquires any major new capability, that system must undergo realistic testing in its intended environment with military users. Although the typical test environment is highly variable and factors are often uncontrolled, design of experiments techniques can add objectivity, efficiency, and rigor to the process of test planning. Statistical analyses help system evaluators get the most information out of limited data sets. Oftentimes new or complex analysis techniques are needed to support the goal of characterizing or predicting system performance across the operational space. Finally, the growing need for computer models or simulations to supplement live testing also means that these models must be appropriately validated before their output can be deemed sufficient for use. Statistical design and analysis techniques are essential for rigorous evaluation of these models.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Matthew R., Kelly M. Avery, and Laura J. Freeman. “Statistical Methods for Defense Testing.” In Wiley StatsRef: Statistics Reference Online, edited by Ron S. Kenett, Nicholas T. Longford, Walter W. Piegorsch, and Fabrizio Ruggeri, 1st ed., 1–5. Wiley, 2018. <a href="https://doi.org/10.1002/9781118445112.stat07946">https://doi.org/10.1002/9781118445112.stat07946</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Regularization for Continuously Observed Ordinal Response Variables with Piecewise-Constant Functional Predictors</title>
      <link>https://research.testscience.org/post/2016-regularization-for-continuously-observed-ordinal-response-variables-with-piecewise-constant-functional-predictors/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-regularization-for-continuously-observed-ordinal-response-variables-with-piecewise-constant-functional-predictors/</guid>
      <description>This paper investigates regularization for continuously observed covariates that resemble step functions. The motivating examples come from operational test data from a recent United States Department of Defense (DoD) test of the Shadow Unmanned Air Vehicle system. The response variable, quality of video provided by the Shadow to friendly ground units, was measured on an ordinal scale continuously over time. Functional covariates, altitude and distance, can be well approximated by step functions.</description>
      <content:encoded><![CDATA[<p>This paper investigates regularization for continuously observed covariates that resemble step functions. The motivating examples come from operational test data from a recent United States Department of Defense (DoD) test of the Shadow Unmanned Air Vehicle system. The response variable, quality of video provided by the Shadow to friendly ground units, was measured on an ordinal scale continuously over time. Functional covariates, altitude and distance, can be well approximated by step functions. Two approaches for regularizing these covariates are considered, including a thinning approach commonly used within the DoD to address autocorrelated time series data, and a novel “smoothing” approach, which first approximates the covariates as step functions and then treats each “step” as a uniquely observed data point. Data sets resulting from both approaches are fit using a mixed model cumulative logistic regression, and we compare their results. While the thinning approach identifies altitude as having a significant impact on video quality, the smoothing approach finds no evidence of an effect. This difference is attributable to the larger effective sample size produced by thinning. System characteristics make it unlikely that video quality would degrade at higher altitudes, suggesting that the thinning approach has produced a Type 1 error. By accounting for the functional characteristics of the covariates, the novel smoothing approach has produced a more accurate characterization of the Shadow’s ability to provide full motion video to supported units.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Avery, Matthew, Mark Orndorff, Timothy Robinson, and Laura Freeman. “Regularization for Continuously Observed Ordinal Response Variables with Piecewise-Constant Functional Covariates.” Quality and Reliability Engineering International 32, no. 6 (2016): 2033–42. <a href="https://doi.org/10.1002/qre.2037">https://doi.org/10.1002/qre.2037</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Rigorous Test and Evaluation for Defense, Aerospace, and National Security</title>
      <link>https://research.testscience.org/post/2016-rigorous-test-and-evaluation-for-defense-aerospace-and-national-security/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-rigorous-test-and-evaluation-for-defense-aerospace-and-national-security/</guid>
      <description>In April 2016, NASA, DOT&amp;amp;E, and IDA collaborated on a workshopdesigned to strengthen the community around statistical approaches to test andevaluation in defense and aerospace. The workshop brought practitioners, analysts,technical leadership, and statistical academics together for a three day exchange ofinformation with opportunities to attend world renowned short courses, share commodchallenges, and learn new skill sets from a variety of tutorials. A highlight of theworkshop was the Tuesday afternoon technical leadership panel chaired by Dr.</description>
      <content:encoded><![CDATA[<p>In April 2016, NASA, DOT&amp;E, and IDA collaborated on a workshopdesigned to strengthen the community around statistical approaches to test andevaluation in defense and aerospace. The workshop brought practitioners, analysts,technical leadership, and statistical academics together for a three day exchange ofinformation with opportunities to attend world renowned short courses, share commodchallenges, and learn new skill sets from a variety of tutorials. A highlight of theworkshop was the Tuesday afternoon technical leadership panel chaired by Dr.Catherine Warner, Science Advisor, DOT&amp;E. This article summarizes core themesdiscuss during the panel session.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura. “Rigorous Test and Evaluation for Defense Aerospace, and National Security: A Panel Session Summary.” The ITEA Journal of Test and Evaluation 37, no. 4 (2016).</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper_D-8229-non-std-Final.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Science of Test Workshop Proceedings, April 11-13, 2016</title>
      <link>https://research.testscience.org/post/2016-science-of-test-workshop-proceedings-april-11-13-2016/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-science-of-test-workshop-proceedings-april-11-13-2016/</guid>
      <description>To mark IDA&amp;rsquo;s 60th anniversary, we are conducting a series of workshops and symposia that bring together IDA sponsors, researchers, experts inside and outside government, and other stakeholders to discuss issues of the day. These events focus on future national security challenges, reflecting on how past lessons and accomplishments help prepare us to deal with complex issues and environments we face going forward. This publication represents the proceedings of the Science of Test Workshop.</description>
      <content:encoded><![CDATA[<p>To mark IDA&rsquo;s 60th anniversary, we are conducting a series of workshops and symposia that bring together IDA sponsors, researchers, experts inside and outside government, and other stakeholders to discuss issues of the day. These events focus on future national security challenges, reflecting on how past lessons and accomplishments help prepare us to deal with complex issues and environments we face going forward. This publication represents the proceedings of the Science of Test Workshop.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura, Pamela Rambow, and Jonathan Snavely. Science of Test Workshop Proceedings. IDA Document NS D-8249. Alexandria, VA: Institute for Defense Analyses, 2016.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper_NS-D-8249-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Tutorial on Sensitivity Testing in Live Fire Test and Evaluation</title>
      <link>https://research.testscience.org/post/2016-tutorial-on-sensitivity-testing-in-live-fire-test-and-evaluation/</link>
      <pubDate>Fri, 01 Jan 2016 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2016-tutorial-on-sensitivity-testing-in-live-fire-test-and-evaluation/</guid>
      <description>A sensitivity experiment is a special type of experimental design that is used when the response variable is binary and the covariate is continuous. Armor protection and projectile lethality tests often use sensitivity experiments to characterize a projectile&amp;rsquo;s probability of penetrating the armor. In this mini-tutorial we illustrate the challenge of modeling a binary response with a limited sample size, and show how sensitivity experiments can mitigate this problem. We review eight different single covariate sensitivity experiments and present a comparison of these designs using simulation.</description>
      <content:encoded><![CDATA[<p>A sensitivity experiment is a special type of experimental design that is used when the response variable is binary and the covariate is continuous. Armor protection and projectile lethality tests often use sensitivity experiments to characterize a projectile&rsquo;s probability of penetrating the armor. In this mini-tutorial we illustrate the challenge of modeling a binary response with a limited sample size, and show how sensitivity experiments can mitigate this problem. We review eight different single covariate sensitivity experiments and present a comparison of these designs using simulation. Additionally, we cover sensitivity experiments for cases that include more than one covariate, and highlight recent research in this area.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Thomas, Laura Freeman, and Raymond Chen. Tutorial on Sensitivity Testing in Live Fire Test and Evaluation. IDA Document NS D-5829. Alexandria, VA: Institute for Defense Analyses, 2016.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5829.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Comparison of Ballistic Resistance Testing Techniques in the Department of Defense</title>
      <link>https://research.testscience.org/post/2014-a-comparison-of-ballistic-resistance-testing-techniques-in-the-department-of-defense/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-a-comparison-of-ballistic-resistance-testing-techniques-in-the-department-of-defense/</guid>
      <description>This paper summarizes sensitivity test methods commonly employed in the Department of Defense. A comparison study shows that modern methods such as Neyer&amp;rsquo;s method and Three-Phase Optimal Design are improvements over historical methods.
Suggested Citation Johnson, Thomas H., Laura Freeman, Janice Hester, and Jonathan L. Bell. “A Comparison of Ballistic Resistance Testing Techniques in the Department of Defense.” IEEE Access 2 (2014): 1442–55. https://doi.org/10.1109/ACCESS.2014.2377633.
Paper: </description>
      <content:encoded><![CDATA[<p>This paper summarizes sensitivity test methods commonly employed in the Department of Defense. A comparison study shows that modern methods such as Neyer&rsquo;s method and Three-Phase Optimal Design are improvements over historical methods.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Thomas H., Laura Freeman, Janice Hester, and Jonathan L. Bell. “A Comparison of Ballistic Resistance Testing Techniques in the Department of Defense.” IEEE Access 2 (2014): 1442–55. <a href="https://doi.org/10.1109/ACCESS.2014.2377633">https://doi.org/10.1109/ACCESS.2014.2377633</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Applying Risk Analysis to Acceptance Testing of Combat Helmets</title>
      <link>https://research.testscience.org/post/2014-applying-risk-analysis-to-acceptance-testing-of-combat-helmets/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-applying-risk-analysis-to-acceptance-testing-of-combat-helmets/</guid>
      <description>Acceptance testing of combat helmets presents multiple challenges that require statistically-sound solutions. For example, how should first article and lot acceptance tests treat multiple threats and measures of performance? How should these tests account for multiple helmet sizes and environmental treatments? How closely should first article testing requirements match historical or characterization test data? What government and manufacturer risks are acceptable during lot acceptance testing? Similar challenges arise when testing other components of Personal Protective Equipment and similar statistical approaches should be applied to all components.</description>
      <content:encoded><![CDATA[<p>Acceptance testing of combat helmets presents multiple challenges that require statistically-sound solutions. For example, how should first article and lot acceptance tests treat multiple threats and measures of performance? How should these tests account for multiple helmet sizes and environmental treatments? How closely should first article testing requirements match historical or characterization test data? What government and manufacturer risks are acceptable during lot acceptance testing? Similar challenges arise when testing other components of Personal Protective Equipment and similar statistical approaches should be applied to all components. This presentation explores these questions using operating characteristics curves and simulation studies.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Hester, Janice, and Laura Freeman. Applying Risk Analysis to Acceptance Testing of Combat Helmets. IDA Document NS D-5334. Alexandria, VA: Institute for Defense Analyses, 2014.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5334.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Design of Experiments for in-Lab Operational Testing of the an/BQQ-10 Submarine Sonar System</title>
      <link>https://research.testscience.org/post/2014-design-of-experiments-for-in-lab-operational-testing-of-the-an-bqq-10-submarine-sonar-system/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-design-of-experiments-for-in-lab-operational-testing-of-the-an-bqq-10-submarine-sonar-system/</guid>
      <description>Operational testing of the AN/BQQ-10 submarine sonar system has never been able to show significant improvements in software versions because of the high variability of at sea measurements. To mitigate this problem, in the most recent AN/BQQ-10 operational test, the Navy’s operational test agency (in consultation with IDA under the direction of Director, Operational Test and Evaluation) supplemented the at sea testing with an operationally focused in-lab comparison. This test used recorded real data played back on two different versions of the sonar system.</description>
      <content:encoded><![CDATA[<p>Operational testing of the AN/BQQ-10 submarine sonar system has never been able to show significant improvements in software versions because of the high variability of at sea measurements. To mitigate this problem, in the most recent AN/BQQ-10 operational test, the Navy’s operational test agency (in consultation with IDA under the direction of Director, Operational Test and Evaluation) supplemented the at sea testing with an operationally focused in-lab comparison. This test used recorded real data played back on two different versions of the sonar system. For each version, the test recorded the time it took multiple operations, with varying operational experience, to detect a submarine target once it appeared on the display. This new test methodology had several benefits: (1) the laboratory setting allowed for the use of design of experiments to control factors that are traditionally infeasible to control during an at sea test; (2) the direct comparison between the two systems resulted in demonstrating a statistically significant reduction in the detection time for the new system. Although laboratory testing cannot replace at sea testing, the results provide strong indication that we can expect performance improvements in the operational environment. This case study shows that laboratory testing and design of experiments have a place in operational testing and should be expanded to improve testing for other systems.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Clutter, Justace R, George Khoury, and Laura Freeman. Design of Experiments for In-Lab Operational Testing of the AN/BQQ-10 Submarine Sonar System. IDA Document NS D-5486. Alexandria, VA: Institute for Defense Analyses, 2014.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-5286.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Power Analysis Tutorial for Experimental Design Software</title>
      <link>https://research.testscience.org/post/2014-power-analysis-tutorial-for-experimental-design-software/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-power-analysis-tutorial-for-experimental-design-software/</guid>
      <description>This guide provides both a general explanation of power analysis and specific guidance to successfully interface with two software packages, JMP and Design Expert (DX).
Suggested Citation Freeman, Laura J., Thomas H. Johnson, and James R. Simpson. “Power Analysis Tutorial for Experimental Design Software:” Fort Belvoir, VA: Defense Technical Information Center, November 1, 2014. https://doi.org/10.21236/ADA619843.
Paper: </description>
      <content:encoded><![CDATA[<p>This guide provides both a general explanation of power analysis and specific guidance to successfully interface with two software packages, JMP and Design Expert (DX).</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J., Thomas H. Johnson, and James R. Simpson. “Power Analysis Tutorial for Experimental Design Software:” Fort Belvoir, VA: Defense Technical Information Center, November 1, 2014. <a href="https://doi.org/10.21236/ADA619843">https://doi.org/10.21236/ADA619843</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Taking the Next Step- Improving the Science of Test in DoD T&amp;E</title>
      <link>https://research.testscience.org/post/2014-taking-the-next-step-improving-the-science-of-test-in-dod-t-e/</link>
      <pubDate>Wed, 01 Jan 2014 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2014-taking-the-next-step-improving-the-science-of-test-in-dod-t-e/</guid>
      <description>The current fiscal climate demands now, more than ever, that test and evaluation(T&amp;amp;E) provide relevant and credible characterization of system capabilities andshortfalls across all relevant operational conditions as efficiently as possible. Indetermining the answer to the question, “How much testing is enough?” it isimperative that we use a scientifically defensible methodology. Design ofExperiments (DOE) has a proven track record in Operational Test andEvaluation (OT&amp;amp;E) of not only quantifying how much testing is enough, but alsowhere in the operational space the test points should be placed.</description>
      <content:encoded><![CDATA[<p>The current fiscal climate demands now, more than ever, that test and evaluation(T&amp;E) provide relevant and credible characterization of system capabilities andshortfalls across all relevant operational conditions as efficiently as possible. Indetermining the answer to the question, “How much testing is enough?” it isimperative that we use a scientifically defensible methodology. Design ofExperiments (DOE) has a proven track record in Operational Test andEvaluation (OT&amp;E) of not only quantifying how much testing is enough, but alsowhere in the operational space the test points should be placed. Over the last fewyears, the T&amp;E community has made great strides in the application of DOE toOT&amp;E, but there is still work to be done in ensuring that the scientificcommunity’s full toolset is utilized. In particular, many test programs have yet tocapitalize on the power of the test design when conducting the data analysis.Employing empirical statistical models (e.g., regression techniques, analysis ofvariance (ANOVA)) allows us to maximize the information from every data point,resulting in defensible analyses that provide crucial information about systemperformance that decision-makers and warfighters need to know. DOT&amp;E willcontinue to work to ensure the highest technical caliber in every DOT&amp;Eevaluation, and that Test and Evaluation Master Plans (TEMPs) are adequate tosupport these robust evaluations. As we improve in our use of these test designsand analysis methods, we need to ensure these practices are institutionalizedacross the entire T&amp;E community and applied across all phases of DoD testing</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura, and V. Bram Lillard. “Taking the Next Step: Improving the Science of Test in DoD T and E.” The ITEA Journal of Test and Evaluation 35, no. 1 (March 2014). <a href="https://apps.dtic.mil/sti/citations/trecms/AD1123777">https://apps.dtic.mil/sti/citations/trecms/AD1123777</a>.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-5101-final-version.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Tutorial on the Planning of Experiments</title>
      <link>https://research.testscience.org/post/2013-a-tutorial-on-the-planning-of-experiments/</link>
      <pubDate>Tue, 01 Jan 2013 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2013-a-tutorial-on-the-planning-of-experiments/</guid>
      <description>This tutorial outlines the basic procedures for planning experiments within the context of the scientific method. Too often quality practitioners fail to appreciate how subject-matter expertise must interact with statistical expertise to generate efficient and effective experimental programs. This tutorial guides the quality practitioner through the basic steps, demonstrated by extensive past experience, that consistently lead to successful results. This tutorial makes extensive use of flowcharts to illustrate the basic process.</description>
      <content:encoded><![CDATA[<p>This tutorial outlines the basic procedures for planning experiments within the context of the scientific method. Too often quality practitioners fail to appreciate how subject-matter expertise must interact with statistical expertise to generate efficient and effective experimental programs. This tutorial guides the quality practitioner through the basic steps, demonstrated by extensive past experience, that consistently lead to successful results. This tutorial makes extensive use of flowcharts to illustrate the basic process. Two case studies summarize the applications of the methodology.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J., Anne G. Ryan, Jennifer L. K. Kensler, Rebecca M. Dickinson, and G. Geoffrey Vining. “A Tutorial on the Planning of Experiments.” Quality Engineering 25, no. 4 (October 1, 2013): 315–32. <a href="https://doi.org/10.1080/08982112.2013.817013">https://doi.org/10.1080/08982112.2013.817013</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Comparing Computer Experiments for the Gaussian Process Model Using Integrated Prediction Variance</title>
      <link>https://research.testscience.org/post/2013-comparing-computer-experiments-for-the-gaussian-process-model-using-integrated-prediction-variance/</link>
      <pubDate>Tue, 01 Jan 2013 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2013-comparing-computer-experiments-for-the-gaussian-process-model-using-integrated-prediction-variance/</guid>
      <description>Space-Filling Designs are a common choice of experimental design strategy for computer experiments. This paper compares space filling design types based on their theoretical prediction variance properties with respect to the Gaussian Process model.
Suggested Citation Silvestrini, Rachel T., Douglas C. Montgomery, and Bradley Jones. “Comparing Computer Experiments for the Gaussian Process Model Using Integrated Prediction Variance.” Quality Engineering 25, no. 2 (April 2013): 164–74. https://doi.org/10.1080/08982112.2012.758284.
Paper: </description>
      <content:encoded><![CDATA[<p>Space-Filling Designs are a common choice of experimental design strategy for computer experiments. This paper compares space filling design types based on their theoretical prediction variance properties with respect to the Gaussian Process model.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Silvestrini, Rachel T., Douglas C. Montgomery, and Bradley Jones. “Comparing Computer Experiments for the Gaussian Process Model Using Integrated Prediction Variance.” Quality Engineering 25, no. 2 (April 2013): 164–74. <a href="https://doi.org/10.1080/08982112.2012.758284">https://doi.org/10.1080/08982112.2012.758284</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Scientific Test and Analysis Techniques- Statistical Measures of Merit</title>
      <link>https://research.testscience.org/post/2013-scientific-test-and-analysis-techniques-statistical-measures-of-merit/</link>
      <pubDate>Tue, 01 Jan 2013 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2013-scientific-test-and-analysis-techniques-statistical-measures-of-merit/</guid>
      <description>Design of Experiments (DOE) provides a rigorous methodology for developing and evaluating test plans. Design excellence consists of having enough test points placed in the right locations in the operational envelope to answer the questions of interest for the test. The key aspects of a well-designed experiment include: the goal of the test, the response variables, the factors and levels, a method for strategically varying the factors across the operational envelope, and statistical measures of merit.</description>
      <content:encoded><![CDATA[<p>Design of Experiments (DOE) provides a rigorous methodology for developing and evaluating test plans. Design excellence consists of having enough test points placed in the right locations in the operational envelope to answer the questions of interest for the test. The key aspects of a well-designed experiment include: the goal of the test, the response variables, the factors and levels, a method for strategically varying the factors across the operational envelope, and statistical measures of merit. Currently, the majority of test plans utilize statistical measures of merit based on confidence and power. Although important, confidence and power are not the only measure of the adequacy and merit of a test design. The type of method that is appropriate is dependent on the goal of the test and the experimental design methodology used. There is no one-size-fits-all solution; rather there is a collection of useful tools that apply in various combinations for different test goals and designs. This talk outlines different statistical measures of merit that should be used when planning an operational test.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura. Scientific Test and Analysis Techniques: Statistical Measures of Merit. IDA Document D-5070. Alexandria, VA: Institute for Defense Analyses, 2014.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_D-5070-2.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>A Bayesian Approach to Evaluation of Land Warfare Systems</title>
      <link>https://research.testscience.org/post/2012-a-bayesian-approach-to-evaluation-of-land-warfare-systems/</link>
      <pubDate>Sun, 01 Jan 2012 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2012-a-bayesian-approach-to-evaluation-of-land-warfare-systems/</guid>
      <description>This presentation is a presentation for the Army Conference on Applied Statistics. The presentation covers a brief introduction to land warfare problems, and devises a methodology using Bayes Theorem to estimate parameters of interest. Two examples are given, a simple one using independent Bernoulli Trials, and a more complex one using correlated Red and Blue casualty data in a Loss Exchange Ratio and a hierarchical model. The presentation demonstrates that the Bayesian approach is successful in both examples at reducing the variance of the estimated parameters, potentially reducing the cost of devising a complex test program.</description>
      <content:encoded><![CDATA[<p>This presentation is a presentation for the Army Conference on Applied Statistics. The presentation covers a brief introduction to land warfare problems, and devises a methodology using Bayes Theorem to estimate parameters of interest. Two examples are given, a simple one using independent Bernoulli Trials, and a more complex one using correlated Red and Blue casualty data in a Loss Exchange Ratio and a hierarchical model. The presentation demonstrates that the Bayesian approach is successful in both examples at reducing the variance of the estimated parameters, potentially reducing the cost of devising a complex test program. The presentation concludes with suggested next steps applicable to the Army Ground Combat Vehicle program.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Wilson, Alyson, Lee Dewald, Robert Holcomb, and Samuel Parry. A Bayesian Approach to Evaluation  of Land Warfare Systems. IDA Document NS D-4711. Alexandria, VA: Institute for Defense Analyses, 2012.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-4711-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Continuous Metrics for Efficient and Effective Testing</title>
      <link>https://research.testscience.org/post/2012-continuous-metrics-for-efficient-and-effective-testing/</link>
      <pubDate>Sun, 01 Jan 2012 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2012-continuous-metrics-for-efficient-and-effective-testing/</guid>
      <description>In today’s fiscal environment, efficient and effective testing is essential. Often, military system requirements are defined using probability of success as the primary measure of effectiveness – for example, a system must complete its mission 80 percent of the time; or the system must detect 90 percent of targets. The traditional approach to testing these probability-based requirements is to execute a series of trials and then total the number of successes; the ratio of successes to number of trails provides an intuitive measure of the probability of success.</description>
      <content:encoded><![CDATA[<p>In today’s fiscal environment, efficient and effective testing is essential. Often, military system requirements are defined using probability of success as the primary measure of effectiveness – for example, a system must complete its mission 80 percent of the time; or the system must detect 90 percent of targets. The traditional approach to testing these probability-based requirements is to execute a series of trials and then total the number of successes; the ratio of successes to number of trails provides an intuitive measure of the probability of success. However, this method of testing has proven to be cost prohibitive, especially at high levels of statistical confidence and power. Often, one or more continuous metrics empirically related to the probability based metric provide more information about system performance than the pass/fail construct. Using these metrics in lieu of the probability-based metrics to plan testing both reduces test costs and provides a better understanding of system performance. In this talk the authors discusses the cost of using binary test metrics (e.g., success or failure, hit or miss). They present several common T&amp;E examples, translating the original probability based requirement to a related continuous metric, and show potential cost savings and information gain achieved by the conversion.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura J, and V Bram Lillard. Continuous Metrics for Efficient and Effective Testing. IDA Document NS D-4571. Alexandria, VA: Institute for Defense Analyses, 2012.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS-D-4571-1.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Designed Experiments for the Defense Community</title>
      <link>https://research.testscience.org/post/2012-designed-experiments-for-the-defense-community/</link>
      <pubDate>Sun, 01 Jan 2012 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2012-designed-experiments-for-the-defense-community/</guid>
      <description>The areas of application for design of experiments principles have evolved, mimicking the growth of U.S. industries over the last century, from agriculture to manufacturing to chemical and process industries to the services and government sectors. In addition, statistically based quality programs adopted by businesses morphed from total quality management to Six Sigma and, most recently, statistical engineering (see Hoerl and Snee 2010). The good news about these transformations is that each evolution contains more technical substance, embedding the methodologies as core competencies, and is less of a ‘‘program.</description>
      <content:encoded><![CDATA[<p>The areas of application for design of experiments principles have evolved, mimicking the growth of U.S. industries over the last century, from agriculture to manufacturing to chemical and process industries to the services and government sectors. In addition, statistically based quality programs adopted by businesses morphed from total quality management to Six Sigma and, most recently, statistical engineering (see Hoerl and Snee 2010). The good news about these transformations is that each evolution contains more technical substance, embedding the methodologies as core competencies, and is less of a ‘‘program.’’ Design of experiments is fundamental to statistical engineering and is receiving increased attention within large government agencies such as the National Aeronautics and Space Administration (NASA) and the Department of Defense. Because test policy is intended to shape test programs, numerous test agencies have experimented with policy wording since about 2001. The Director of Operational Test &amp; Evaluation has recently (2010) published guidelines to mold test programs into a sequence of well-designed and statistically defensible experiments. Specifically, the guidelines require, for the first time, that test programs report statistical power as one proof of sound test design. This article presents the underlying tenets of design of experiments, as applied in the Department of Defense, focusing on factorial, fractional factorial, and response surface design and analyses. The concepts of statistical modeling and sequential experimentation are also emphasized. Military applications are presented for testing and evaluation of weapon system acquisition, including force-on-force tactics, weapons employment and maritime search, identification, and intercept.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Rachel T., Gregory T. Hutto, James R. Simpson, and Douglas C. Montgomery. “Designed Experiments for the Defense Community.” Quality Engineering 24, no. 1 (January 2012): 60–79. <a href="https://doi.org/10.1080/08982112.2012.627288">https://doi.org/10.1080/08982112.2012.627288</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Statistically Based T&amp;E Using Design of Experiments</title>
      <link>https://research.testscience.org/post/2012-statistically-based-t-e-using-design-of-experiments/</link>
      <pubDate>Sun, 01 Jan 2012 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2012-statistically-based-t-e-using-design-of-experiments/</guid>
      <description>This document outlines the charter for the Committee to Institutionalize Scientific Test Design and Rigor in Test and Evaluation. The charter defines the problem, identifies potential steps in a roadmap for accomplishing the goals of the committee and lists committeemembership. Once the committee is assembled, the members will revise this document as needed. The charter will be endorsed by DOT&amp;amp;E and DDT&amp;amp;E, once finalize.
Suggested Citation Freeman, Laura. Statistically Based T&amp;amp;E Using Design of Experiments.</description>
      <content:encoded><![CDATA[<p>This document outlines the charter for the Committee to Institutionalize Scientific Test Design and Rigor in Test and Evaluation. The charter defines the problem, identifies potential steps in a roadmap for accomplishing the goals of the committee and lists committeemembership. Once the committee is assembled, the members will revise this document as needed. The charter will be endorsed by DOT&amp;E and DDT&amp;E, once finalize.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura. Statistically Based T&amp;E Using Design of Experiments. IDA Document D-4548. Alexandria, VA: Institute for Defense Analyses, 2012.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides_NS_D-4548.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>An Expository Paper on Optimal Design</title>
      <link>https://research.testscience.org/post/2011-an-expository-paper-on-optimal-design/</link>
      <pubDate>Sat, 01 Jan 2011 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2011-an-expository-paper-on-optimal-design/</guid>
      <description>There are many situations where the requirements of a standard experimental design do not fit the research requirements of the problem. Three such situations occur when the problem requires unusual resource restrictions, when there are constraints on the design region, and when a non-standard model is expected to be required to adequately explain the response.
Suggested Citation Johnson, Rachel T., Douglas C. Montgomery, and Bradley A. Jones. “An Expository Paper on Optimal Design.</description>
      <content:encoded><![CDATA[<p>There are many situations where the requirements of a standard experimental design do not fit the research requirements of the problem. Three such situations occur when the problem requires unusual resource restrictions, when there are constraints on the design region, and when a non-standard model is expected to be required to adequately explain the response.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Rachel T., Douglas C. Montgomery, and Bradley A. Jones. “An Expository Paper on Optimal Design.” Quality Engineering 23, no. 3 (July 2011): 287–301. <a href="https://doi.org/10.1080/08982112.2011.576203">https://doi.org/10.1080/08982112.2011.576203</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Design of Experiments in Highly Constrained  Design Spaces</title>
      <link>https://research.testscience.org/post/2011-design-of-experiments-in-highly-constrained-design-spaces/</link>
      <pubDate>Sat, 01 Jan 2011 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2011-design-of-experiments-in-highly-constrained-design-spaces/</guid>
      <description>This presentation shows the merits of applying experimental design to operational tests, guidance on using DOE from the Director, Operational Test and Evaluation, and presents the design solution for the test of a chemical agent detector. It is important to keep in mind the advanced techniques from DOE (split-plot designs, optimal designs) to determine effective DOEs for operational testing; traditional design strategies often result in designs that are not executable.</description>
      <content:encoded><![CDATA[<p>This presentation shows the merits of applying experimental design to operational tests, guidance on using DOE from the Director, Operational Test and Evaluation, and presents the design solution for the test of a chemical agent detector.  It is important to keep in mind the advanced techniques from DOE (split-plot designs, optimal designs) to determine effective DOEs for operational testing; traditional design strategies often result in designs that are not executable.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura. “Design of Experiments in Highly Constrained Design Spaces.” Presented at the Army Conference on Applied Statistics, October 2011.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Hybrid Designs- Space Filling and Optimal Experimental Designs for Use in Studying Computer Simulation Models</title>
      <link>https://research.testscience.org/post/2011-hybrid-designs-space-filling-and-optimal-experimental-designs-for-use-in-studying-computer-simulation-models/</link>
      <pubDate>Sat, 01 Jan 2011 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2011-hybrid-designs-space-filling-and-optimal-experimental-designs-for-use-in-studying-computer-simulation-models/</guid>
      <description>This tutorial provides an overview of experimental design for modeling and simulation. Pros and cons of each design methodology are discussed.
Suggested Citation Silvestrini, Rachel Johnson. “Hybrid Designs: Space Filling and Optimal Experimental Designs for Use in Studying Computer Simulation Models.” Monterey, California, May 2011.
Slides: </description>
      <content:encoded><![CDATA[<p>This tutorial provides an overview of experimental design for modeling and simulation. Pros and cons of each design methodology are discussed.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Silvestrini, Rachel Johnson. “Hybrid Designs: Space Filling and Optimal Experimental Designs for Use in Studying Computer Simulation Models.” Monterey, California, May 2011.</p>
</blockquote>
<h4 id="slides">Slides:</h4>
<embed src= "slides.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Use of Statistically Designed Experiments to Inform Decisions in a Resource Constrained Environment</title>
      <link>https://research.testscience.org/post/2011-use-of-statistically-designed-experiments-to-inform-decisions-in-a-resource-constrained-environment/</link>
      <pubDate>Sat, 01 Jan 2011 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2011-use-of-statistically-designed-experiments-to-inform-decisions-in-a-resource-constrained-environment/</guid>
      <description>There has been recent emphasis on the increased use of statistics, including the use of statistically designed experiments, to plan and execute tests that support Department of Defense (DoD) acquisition programs. The use of statistical methods, including experimental design, has shown great benefits in industry, especially when used in an integrated fashion; for example see the literature on Six Sigma. The structured approach of experimental design allows the user to determine what data need to be collected and how it should be analyzed to achieve specific decision making objectives.</description>
      <content:encoded><![CDATA[<p>There has been recent emphasis on the increased use of statistics, including the use of statistically designed experiments, to plan and execute tests that support Department of Defense (DoD) acquisition programs. The use of statistical methods, including experimental design, has shown great benefits in industry, especially when used in an integrated fashion; for example see the literature on Six Sigma. The structured approach of experimental design allows the user to determine what data need to be collected and how it should be analyzed to achieve specific decision making objectives. This focuses decision making processes, improves test efficiency and provides objective data for evidence-based decision-making. Today the DoD Test and Evaluation (T&amp;E) community is investigating the use of statistical methods to provide efficient and effective testing. This paper discusses the use of statistics in T&amp;E to assist T&amp;E practitioners and acquisition management in understanding how to improve the quantity and quality of information made available to decision makers to make risk assessments, even in a resource constrained environment.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Freeman, Laura, Karl Glaeser, and Alethea Rucker. “Use of Statistically Design Experiments to Inform Decisions in a Resource Constrained Environment.” ITEA Journal of Test and Evaluation. 32, no. 3 (2011): 267–76.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper_D-4355-NS-final.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Examining Improved Experimental Designs for Wind Tunnel Testing Using Monte Carlo Sampling Methods</title>
      <link>https://research.testscience.org/post/2010-examining-improved-experimental-designs-for-wind-tunnel-testing-using-monte-carlo-sampling-methods/</link>
      <pubDate>Fri, 01 Jan 2010 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2010-examining-improved-experimental-designs-for-wind-tunnel-testing-using-monte-carlo-sampling-methods/</guid>
      <description>In this paper we compare data from a fairly large legacy wind tunnel test campaign to smaller, statistically-motivated experimental design strategies. The comparison, using Monte Carlo sampling methodology, suggests a tremendous opportunity to reduce wind tunnel test efforts without losing test information.
Suggested Citation Hill, Raymond R., Derek A. Leggio, Shay R. Capehart, and August G. Roesener. “Examining Improved Experimental Designs for Wind Tunnel Testing Using Monte Carlo Sampling Methods.” Quality and Reliability Engineering International 27, no.</description>
      <content:encoded><![CDATA[<p>In this paper we compare data from a fairly large legacy wind tunnel test campaign to smaller, statistically-motivated experimental design strategies. The comparison, using Monte Carlo sampling methodology, suggests a tremendous opportunity to reduce wind tunnel test efforts without losing test information.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Hill, Raymond R., Derek A. Leggio, Shay R. Capehart, and August G. Roesener. “Examining Improved Experimental Designs for Wind Tunnel Testing Using Monte Carlo Sampling Methods.” Quality and Reliability Engineering International 27, no. 6 (October 2011): 795–803. <a href="https://doi.org/10.1002/qre.1165">https://doi.org/10.1002/qre.1165</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Choice of Second-Order Response Surface Designs for Logistic and Poisson Regression Models</title>
      <link>https://research.testscience.org/post/2009-choice-of-second-order-response-surface-designs-for-logistic-and-poisson-regression-models/</link>
      <pubDate>Thu, 01 Jan 2009 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2009-choice-of-second-order-response-surface-designs-for-logistic-and-poisson-regression-models/</guid>
      <description>This paper illustrates the construction of D-optimal second order designs for situations when the response is either binomial (pass/fail) or Poisson (count data).
Suggested Citation Johnson, Rachel T., and Douglas C. Montgomery. “Choice of Second-Order Response Surface Designs for Logistic and Poisson Regression Models.” International Journal of Experimental Design and Process Optimisation 1, no. 1 (2009): 2. https://doi.org/10.1504/IJEDPO.2009.028954.
Paper: </description>
      <content:encoded><![CDATA[<p>This paper illustrates the construction of D-optimal second order designs for situations when the response is either binomial (pass/fail) or Poisson (count data).</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Rachel T., and Douglas C. Montgomery. “Choice of Second-Order Response Surface Designs for Logistic and Poisson Regression Models.” International Journal of Experimental Design and Process Optimisation 1, no. 1 (2009): 2. <a href="https://doi.org/10.1504/IJEDPO.2009.028954">https://doi.org/10.1504/IJEDPO.2009.028954</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
    <item>
      <title>Designing Experiments for Nonlinear Models—an Introduction</title>
      <link>https://research.testscience.org/post/2009-designing-experiments-for-nonlinear-models-an-introduction/</link>
      <pubDate>Thu, 01 Jan 2009 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2009-designing-experiments-for-nonlinear-models-an-introduction/</guid>
      <description>We illustrate the construction of Bayesian D-optimal designs for nonlinear models and compare the relative efficiency of standard designs with these designs for several models and prior distributions on the parameters. Through a relative efficiency analysis, we show that standard designs can perform well in situations where the nonlinear model is intrinsically linear. However, if the model is nonlinear and its expectation function cannot be linearized by simple transformations, the nonlinear optimal design is considerably more efficient than the standard design.</description>
      <content:encoded><![CDATA[<p>We illustrate the construction of Bayesian D-optimal designs for nonlinear models and compare the relative efficiency of standard designs with these designs for several models and prior distributions on the parameters. Through a relative efficiency analysis, we show that standard designs can perform well in situations where the nonlinear model is intrinsically linear. However, if the model is nonlinear and its expectation function cannot be linearized by simple transformations, the nonlinear optimal design is considerably more efficient than the standard design.</p>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Johnson, Rachel T., and Douglas C. Montgomery. “Designing Experiments for Nonlinear Models—an Introduction.” Quality and Reliability Engineering International 26, no. 5 (July 2010): 431–41. <a href="https://doi.org/10.1002/qre.1063">https://doi.org/10.1002/qre.1063</a>.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
  </channel>
</rss>
