ORCID

0009-0006-6734-2840

Keywords

FLIPPO, feudal reinforcement learning, proximal policy optimization, hierarchical reinforcement learning, multi-agent systems, LOTZ, SMAClite, autonomous agents

Subject Categories

Military and Veterans Studies | Systems Engineering

Abstract

Multi-agent systems functioning under hierarchical authority must produce unified action without continual human supervision. For example, constructive training simulations use computer-generated forces (CGFs) to act as friendly units, adversaries, and neutral populations within command-and-control structures that mirror real military organizations. However, current CGFs rely on brittle, hand-authored behaviors that cannot adapt to novel situations and require substantial operator oversight. Replacing this overhead with fully autonomous agents has been a grand challenge for modeling and simulation for over two decades. Multi-agent deep reinforcement learning (MADRL) offers a way forward by learning policies for multiple agents acting under partial observability and shared objectives, but MADRL environments experience challenges such as non-stationarity, credit assignment, shadowed equilibria, and scalability. Feudal multi-agent hierarchies (FMHs) address these challenges by organizing agents into structures in which higher-echelon agents influence subordinate behavior through orders and reward design. The FMH literature has compiled the design choices required for a viable feudal multi-agent architecture; what has been missing is an architecture that integrates them and is evaluated under controlled conditions. This dissertation supplies that architecture. Feudal Leader Independent PPO (FLIPPO) is a framework built on Proximal Policy Optimization (PPO) that trains a leader policy to select temporally extended orders and follower policies to acquire skills that fulfill those orders, using order-conditioned observations and feudal reward signals. FLIPPO is evaluated on a sparse, high-dimensional joint-action benchmark and a cooperative combat scenario against an adversary. In controlled comparisons, FLIPPO achieves higher task success rates than flat PPO baselines and sustains performance as the problem scale increases. The discrete orders produce an inspectable behavioral trace that supports verification and validation, a precondition for modeling and simulation to adopt machine learning agents.

Completion Date

2026

Semester

Summer

Committee Chair

Mondesire, Sean

Degree

Doctor of Philosophy (Ph.D.)

College

College of Sciences

Department

School of Modeling, Simulation and Training

Format

PDF

Document Type

Dissertation

Language

English

Release Date

8-15-2027

Available for download on Sunday, August 15, 2027

Share

COinS
 

Accessibility Statement

This item was created or digitized prior to April 24, 2027, or is a reproduction of legacy media created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the University Libraries provides accessible versions of archival materials upon request. To request an accommodation for this item, please submit an accessibility request form.