Designing a Bioterrorism Response Network Considering Transportation: A Reinforcement Learning Approach
Abstract
Bioterrorism attacks are among the most serious national security threats, requiring the design of rapid and effective response networks. This study presents a three-level hybrid model based on the Defender–Attacker–Defender (DAD) framework, in which strategic stockpiling, tactical distribution, and operational transportation decisions are addressed simultaneously. The novelty of this research lies in integrating the Vehicle Routing Problem (VRP) with capacity, time-window, and demand uncertainty considerations into the context of biological attacks. To overcome computational complexity and ensure real-time responsiveness to environmental changes, a Reinforcement Learning (RL) module is developed that learns optimal drug distribution routes through interaction with the environment, while balancing storage, transportation, and human loss costs. Numerical experiments and sensitivity analyses show that the proposed algorithm outperforms traditional methods by reducing both casualties and total costs, and it achieves fast convergence and high efficiency under uncertainty. Ultimately, this approach can serve as a foundation for developing real-scale response systems to biological crises.
Keywords:
Bioterrorism, Biological attacks, Vehicle routing problem, Crisis logistics, Mathematical modeling Reinforcement learningReferences
- [1] Herrmann, J., Kahn, B., Wollek, S., & Nicholson, A. (2016). The nation’s medical countermeasure stockpile: Opportunities to improve the efficiency, effectiveness, and sustainability of the CDC strategic national stockpile: Workshop summary. National Academies Press. https://library.nationalmuseum.gov.ph/digital-collection/digi/NA02/2016/23532.pdf
- [2] Brown, G., Carlyle, M., Salmerón, J., & Wood, K. (2006). Defending critical infrastructure. Interfaces, 36(6), 530–544. https://doi.org/10.1287/inte.1060.0252
- [3] Ben-Tal, A., Ghaoui, L. El, & Nemirovski, A. (2009). Robust optimization. Robust optimization. https://doi.org/10.1515/9781400831050
- [4] Toth, P., & Vigo, D. (2002). The vehicle routing problem. SIAM. https://doi.org/10.1137/1.9780898718515
- [5] Laporte, G. (2009). Fifty years of vehicle routing. Transportation science, 43(4), 408–416. https://doi.org/10.1287/trsc.1090.0301
- [6] Kool, W., Van Hoof, H., & Welling, M. (2018). Attention, learn to solve routing problems! https://doi.org/10.48550/arXiv.1803.08475
- [7] Nazari, M., Oroojlooy, A., Snyder, L., & Takác, M. (2018). Reinforcement learning for solving the vehicle routing problem. Advances in neural information processing systems, 31. https://doi.org/10.48550/arXiv.1802.04240
- [8] Lin, B., Ghaddar, B., & Nathwani, J. (2021). Deep reinforcement learning for the electric vehicle routing problem with time windows. IEEE transactions on intelligent transportation systems, 23(8), 11528–11538. https://doi.org/10.1109/TITS.2021.3105232
- [9] Phiboonbanakit, T., Horanont, T., Huynh, V.N., & Supnithi, T. (2021). A hybrid reinforcement learning-based model for the vehicle routing problem in transportation logistics. Ieee access, 9, 163325–163347. https://doi.org/10.1109/ACCESS.2021.3131799
- [10] Simchi-Levi, D., Trichakis, N., & Zhang, P. Y. (2019). Designing response supply chain against bioattacks. Operations research, 67(5), 1246–1268. https://doi.org/10.1287/opre.2019.1862
- [11] Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement learning: A survey. Journal of artificial intelligence research, 4, 237–285. https://doi.org/10.1613/jair.301
- [12] Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction (Vol. 1, No. 1, pp. 9-11). Cambridge: MIT press. https://mitpress.mit.edu/9780262039246/reinforcement-learning/
- [13] Konda, V., & Tsitsiklis, J. (1999). Actor-critic algorithms. Advances in neural information processing systems, 12. https://proceedings.neurips.cc/paper_files/paper/1999/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf

