<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Mark Herrera on Test Science Research Document Library</title>
    <link>https://research.testscience.org/researchers/mark-herrera/</link>
    <description>Recent content in Mark Herrera on Test Science Research Document Library</description>
    <generator>Hugo -- 0.129.0</generator>
    <language>en-us</language>
    <copyright>Institute for Defense Analyses</copyright>
    <lastBuildDate>Sun, 01 Jan 2023 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://research.testscience.org/researchers/mark-herrera/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>AI &#43; Autonomy T&amp;E in DoD</title>
      <link>https://research.testscience.org/post/2023-ai-autonomy-t-e-in-dod/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://research.testscience.org/post/2023-ai-autonomy-t-e-in-dod/</guid>
      <description>Test and evaluation (T&amp;amp;E) of AI-enabled systems (AIES) often emphasizes algorithm accuracy over robust, holistic system performance. While this narrow focus may be adequate for some applications of AI, for many complex uses, T&amp;amp;E paradigms removed from operational realism are insufficient. However, leveraging traditional operational testing (OT) methods for to evaluate AIESs can fail to capture novel sources of risk. This brief establishes a common AI vocabulary and highlights OT challenges posed by AIESs by answering the following questions</description>
      <content:encoded><![CDATA[<p>Test and evaluation (T&amp;E) of AI-enabled systems (AIES) often emphasizes algorithm accuracy over robust, holistic system performance. While this narrow focus may be adequate for some applications of AI, for many complex uses, T&amp;E paradigms removed from operational realism are insufficient. However, leveraging traditional operational testing (OT) methods for to evaluate AIESs can fail to capture novel sources of risk. This brief establishes a common AI vocabulary and highlights OT challenges posed by AIESs by answering the following questions</p>
<ol>
<li>What is “Artificial Intelligence (AI)”?</li>
</ol>
<p>a. A brief “AI Primer” defines some common terms, highlights words that are used inconsistently, and discusses where definitions are insufficient for identifying systems that require additional T&amp;E considerations.</p>
<ol start="2">
<li>How does AI impact T&amp;E?</li>
</ol>
<p>a. AI isn’t new, but systems with AI pose new challenges and may require structural changes to how we T&amp;E.</p>
<ol start="3">
<li>What makes DoD applications of AI unique?</li>
</ol>
<p>a. Many Silicon Valley applications of AI often lack the task complexity and severe consequences of risk faced by DoD.</p>
<ol start="4">
<li>What is the warfighter’s role?</li>
</ol>
<p>a. T&amp;E must assure warfighters have calibrated trust &amp; an adequate understanding of system behavior.</p>
<ol start="5">
<li>What is the state of DoD AI T&amp;E in IDA and OED?</li>
</ol>
<h4 id="suggested-citation">Suggested Citation</h4>
<blockquote>
<p>Vickers, Brian D, Matthew R Avery, Rachel A Haga, Mark R Herrera, Daniel J Porter, Stuart M Rodgers, and Rebecca M Medlin. AI + Autonomy T&amp;E in DoD. IDA Document NS 3000083. Alexandria, VA: Institute for Defense Analyses, 2023.</p>
</blockquote>
<h4 id="paper">Paper:</h4>
<embed src= "paper.pdf" width= "100%" height= "700px" type="application/pdf" >

]]></content:encoded>
    </item>
  </channel>
</rss>
