Preface ix About the Companion Website xi 1 Introduction 1 1.1 The Problem of Search Under Uncertainty 1 1.2 Book Objectives and Main Contributions 2 1.3 Structure of the Book 3 2 Background 5 2.1 Probabilistic Search for Static and Moving Targets 6 2.2 Search in Shadowed Space 10 2.3 Search with False-Positive and False-Negative Errors 16 2.4 Conclusions 19 3 Problem Formulation and Basic Search Procedure 21 3.
1 Basic Assumptions 21 3.2 Sensor's Fusion and Dynamic Probability Map 24 3.3 Search Policy and Sensing 26 3.3.1 Search Policy 26 3.3.2 Search Time and Sensors' Sensitivity 27 3.3.
3 Numerical Simulations 28 3.4 Conclusions 30 4 Reactive Detection Algorithms 31 4.1 Agents' Policies and Decision-making 31 4.2 Policies Control and Brute-Force Learning 34 4.3 Numerical Simulations 35 4.3.1 Detection by a Single Agent 36 4.3.
2 Detection by Multiple Agents 38 4.4 Conclusions 42 5 Search with Learning 43 5.1 Reinforcement Learning 43 5.2 Search with Supervised and Unsupervised Learning 47 5.3 Two-Actor Search 50 5.3.1 Pursuit-Evasion Game 51 5.3.
2 Actor-Critic Approach 52 5.4 Conclusions 53 6 Search with Deep Q-Learning: Single Agent 55 6.1 Problem Formulation and Notation 55 6.2 Decision-making Policy and Deep Q-Learning Solution 58 6.2.1 Agent's Actions and Decision-making 58 6.2.2 Dynamic Programming and Target Neural Networks 59 6.
2.3 Model-Free and Model-Based Learning 60 6.2.4 Choice of the Actions at the Learning Stage 63 6.2.5 The Q-max Algorithm 63 6.2.6 The SPL Algorithm 67 6.
3 Numerical Simulations 67 6.3.1 Network Training in the Q-max Algorithm 68 6.3.2 Detection by the Q-max and SPL Algorithms 69 6.3.3 Comparison Between Q-max and SPL Algorithms and One-Step Heuristics 69 6.3.
4 Comparison Between SPL Algorithm and Optimal Solution 73 6.3.5 Run Times and Mean-Squared Errors for Different Sizes of Data Sets 75 6.4 Conclusions 76 7 Search with Deep Q-learning: Multiple Agents 77 7.1 Problem Formulation and Notation 77 7.2 Cooperative Detection Using Voronoi Regions and Deep Q-Learning 78 7.2.1 Agents' Actions and Decision-making 79 7.
2.2 Reactive Decision-making in Voronoi Regions: A Distributed EIG Algorithm 79 7.2.3 Collective Deep Q-Learning Approach 81 7.3 Numerical Simulations 88 7.3.1 Detection of Static Targets 89 7.3.
2 Detection of Moving Targets 92 7.3.3 Learning Errors and Run Time of the Collective Q-max Algorithm 94 7.4 Conclusions 95 8 Conclusions 97 References 99 Index 103.