The Pentagon Has a New Test for Top Officers. It’s Not Clear Why.

Mark Hertling / The Bulwark

Defense Secretary Pete Hegseth’s rationale was less than convincing.

THIS WEEK, SECRETARY OF DEFENSE PETE HEGSETH presented a new initiative to help decide who should become America’s next generals and admirals. He announced the “Joint Warfighting Evaluation” to improve the process of selecting colonels and Navy captains for their first star. Officers are now to be assessed on operational analysis, communication, and decision-making through written work and computer-assisted wargaming, with some versions incorporating artificial intelligence—in addition to the normal promotion board process. Each service has a different approach, but the metrics and overarching objectives are not yet clear.

In theory, the idea of adding a written test or wargaming exercise to the promotion process for senior officers isn’t preposterous. But the way the Department of Defense has described it, the people in charge of designing it, and the historical arguments they made for instituting it leave me with little confidence that it will produce better general and admirals—and more than a little worry that it will do the opposite.

Context, as always, is important: Hegseth has forced respected, accomplished, talented officers out of the military; he has intervened in the hiring process for more junior officers in unprecedented ways with little or no explanation; he has emphasized physical fitness over every other quality and skill competent service members require—especially commanders; and last year he assembled every general and admiral in the American military into one room just to berate them in a highly partisan speech. So his announcement of a major change to the promotion system for senior officers shouldn’t be greeted as a mere bureaucratic reform.

During his social media rollout, Hegseth praised Stuart Scheller for leading the design of this new evaluation. Scheller, to his credit, has more military experience than Hegseth himself, but not much more: He served for seventeen years in the Marine Corps, rising to the rank of lieutenant colonel, and deployed to Iraq and Afghanistan, where he received awards for valor. His professional military education includes Expeditionary Warfare School and the Marine Corps Command and Staff College—both junior officer courses—and he commanded the Advanced Infantry Training Battalion at the Marines’ School of Infantry-East. But his qualifications to be the architect of a consequential military performance assessment for strategic leaders are dubious: Scheller’s career did not include operational command or graduation from a senior-service War College, where officers study strategy, national security policy, joint campaigning, and theater-level warfare, nor has he participated in any theater-level exercise or training.

It’s possible that Scheller’s career in uniform could have proceeded further, but in 2021, while still on active duty, he publicly criticized senior military and civilian leaders over the Afghanistan withdrawal. He was relieved of his battalion command, appeared before a court martial, and subsequently pleaded guilty to contempt toward officials, disrespect toward superior commissioned officers, failure to obey orders, and conduct unbecoming an officer.

None of that necessarily prevents Scheller from offering useful ideas about improving officer selection—it’s possible he possesses all the qualities necessary to be a great general officer except for the ability to abide by the Uniform Code of Military Justice. But he’s now in a position of assessing qualities required of future flag officers—qualities he was never trained on or required to demonstrate.

BOTH HEGSETH AND SCHELLER reached into Army history to justify this new test, invoking Gen. George C. Marshall’s transformation of the Army before World War II and the massive Louisiana Maneuvers of 1941. The story they emphasized is that Marshall used those maneuvers to identify officers with the potential to serve in key strategic roles—Dwight Eisenhower, Omar Bradley, George Patton, and others—and that today’s military needs a similarly rigorous way to distinguish warfighters from bureaucrats.

There are some elements of truth in that story—Marshall was an astute judge of talent, and many of the officers who would become famous leading the American armies in Europe during World War II did participate in the Louisiana Maneuvers. But Hegseth’s and Scheller’s story confuses what Marshall was principally testing in Louisiana, exaggerates the role the maneuvers played in selecting individual commanders, and overlooks how extensively today’s military has transformed how it already trains, exercises, and evaluates commanders in the field and in the classroom. Additionally, the 1941 maneuvers were not primarily a talent-identification program; They were a massive experiment designed to determine whether an Army expanding and transforming at breathtaking speed could fight a large-scale war. Roughly 400,000 recently recruited troops and more than 1,000 newly built aircraft participated in those exercises across various Southern states, and the Army tested armored formations, combined arms, communications, logistics, doctrine, organization, and equipment while simultaneously mobilizing a force that had expanded in the preceding few years.Marshall’s goal was to look at various elements of his larger military, and he was also interested in what mistakes the military could fix in Louisiana before they fought the Germans.

Commanders were certainly being evaluated. Some were found wanting and replaced; others distinguished themselves. Marshall even described the exercises as a “combat college for troop leading.” But leader assessment was only part of the much larger process of putting entire formations—and the emerging American way of war—under enormous stress. The Army’s official history emphasizes the maneuvers’ influence on force structure, doctrine, organization, equipment, and combined-arms teamwork, in addition to leadership.

The stories of Eisenhower, Bradley, and Patton illustrate the distinction. Eisenhower had no combat experience before World War II; in Louisiana, he was the chief of staff to the Third Army commander, Lt. Gen. Walter Krueger. Ike’s performance did enhance his reputation—but it didn’t automatically catapult him to command. Bradley also entered World War II without combat experience, but Marshall had already identified his potential long before Louisiana, promoting him from lieutenant colonel to brigadier general months before the maneuvers. Patton was different: He had combat experience, had been wounded in World War I, and his aggressive performance in mechanized exercises reinforced an already established reputation as a front-line warfighter.

Marshall’s achievement was therefore more sophisticated than finding “warriors” in Louisiana. He watched officers over time and recognized different talents for different requirements. Patton was not Eisenhower, Eisenhower was not Bradley, and Marshall did not want them to be.

THE LOUISIANA MANEUVERS WERE NOTABLE because they were the first large-scale exercises of mechanized forces in the United States under newly appointed generals. Now such training events and exercises are routine and much more realistic than they were then. Crucially, they are used both to train units and to evaluate various levels of commanders, and they happen at various locations, away from the public eye.

The Army has the National Training Center at Fort Irwin, the Joint Readiness Training Center in Louisiana, and the Joint Multinational Readiness Center in Germany. Brigade combat teams—increasingly connected to their division and corps headquarters—fight against sophisticated, permanently stationed opposing forces whose mission is to defeat them. Commanders and staffs deal with maneuver, fires, intelligence, logistics, communications, (simulated) casualties, and the endless friction created when thousands of real soldiers and vehicles confront a thinking enemy under closely monitored conditions—unlike the exercises of 1941, which lacked modern tracking systems and software. They do this in the dirt, sometimes connected to simulations, and the goal is always to provide a “heavy coat of sweat” for the operational unit. I know, because I’ve participated in dozens of these events, and I was the commander of the operations group at two of the army’s training centers that designed this capstone training.

The Marines do much the same at Twentynine Palms. Navy fighter crews have Top Gun for advanced tactical instruction,and the Navy conducts joint and allied exercises all over the globe for fleet maneuvers. The Air Force’s Red Flag exercises expose aircrews and staffs to large-force, realistic combat scenarios involving multiple aircraft types, joint forces, electronic threats, stressed communications, complex adversaries, and detailed reconstruction afterward.

These exercises arguably perform much of the function Marshall sought in 1941, but with instrumentation, dedicated opposing forces, observer-controller teams, after-action reviews, live-fire capabilities, modern military technology, and decades of accumulated combat experience that Marshall could only have dreamed about. For decades, these centers have exposed the officers who eventually reach senior operational command to complex situations under experienced observer-controllers. They are one of the many reasons the United States has the greatest military in the world. The evaluations in these programs reveal something a written examination or AI simulation cannot easily reproduce: how a commander leads real people under complex conditions.

THE HEGSETH–SCHELLER STORY also skips the Goldwater-Nichols Act, the 1986 law that transformed the American military.

Goldwater-Nichols was intended, among other things, to ensure that senior officers were capable of thinking beyond the competencies and cultures of their individual services. Ever since, the military has focused almost obsessively on all things “joint”—meaning, in military parlance, relating to two or more branches operating together. Hence the Joint Chiefs of Staff, the F-35 Joint Strike Fighter, joint operations, joint and multinational headquarters within out combatant commands, etc. Goldwater-Nichols introduced the requirement that senior officers have joint qualifications that include both joint education and joint experience, explicitly intended to ensure the “systematic, progressive, career-long development” of officers who can think and work beyond the disciplines of their branch.

That joint qualification system was deliberately designed as a career-long process rather than a single test. Professional military education, joint assignments and other qualifying experiences, demanding training exercises, command assignments, and years of carefully evaluated performance and potential all contribute to the record that eventually follows an officer to a promotion board.

NONE OF THIS MAKES A POTENTIAL Joint Warfighting Evaluation useless. Written analysis can reveal how an officer thinks. Wargaming can expose assumptions and force difficult decisions. AI can generate adaptive opponents and scenarios that would have been impossible only a few years ago. Those tools may add valuable information to a promotion board.

Maybe.

Before rushing to use this new assessment, the Pentagon should determine whether performance on the evaluation correlates with demonstrated performance in the real world, whether the goals and metrics are sound, and whether it is worth the effort. I can think of an intriguing way to conduct that proof of principle.

Multiple combat-hardened three- and four-star officers have recently left senior positions or been asked to retire by Secretary Hegseth. Army generals with remarkable reputations like C.D. Donahue and Randy George have years of complex operational, combat, command, and joint experience in a variety of environments. While Hegseth has not publicly explained the reasons behind some of these early departures, officers with careers like theirs might offer something close to a ready-made validation sample for this JWE.

To troubleshoot the assessment, the Pentagon should invite a representative group of such retired senior officers, voluntarily and anonymously, to take the new evaluation. Give them the same assessment and AI-assisted wargame intended for new brigadier generals and rear admirals, and then compare the results with what decades of documented performance already tells us about their abilities.

The point would not be to reconsider personnel decisions or embarrass anyone. It would be to test the instrument. If the evaluation consistently identifies strengths and weaknesses already demonstrated during decades of operational and joint service, then it’s not clear why additional purpose it serves. If experienced operational commanders perform unpredictably, or if the results bear little relationship to desired demonstrated abilities, that will tell the Pentagon something equally important. If the test is supposed to measure something that isn’t currently being measured, then the Defense Department should specify what that is.

Such a validation study could also expose what the new test doesn’t measure. High-level command requires more than operational planning and battlefield decision-making. It requires connecting military operations to political objectives; managing alliances; understanding logistics, industrial capacity, intelligence, and law; maintaining civilian control; shaping organizational culture; and providing courageous and candid military advice to civilian leaders. This may not seem important to Secretary Hegseth, who seems to want only lethality, but these are also the competencies required of senior level officers.

It is difficult to know whether a one-time AI-supported evaluation can measure those qualities. That is precisely why it should be validated before anyone assumes that it can.

Marshall understood something similar. Louisiana helped him see officers under pressure, but it was only one part of a much larger process of ongoing evaluations that he used to choose top commanders.

The American military should experiment relentlessly with better ways of identifying future commanders. AI and sophisticated wargaming may eventually become valuable parts of that process. But the historical lesson of Louisiana is not that Marshall found America’s great generals by giving them a test. It is that he built an Army that continually tested its organizations, doctrine, equipment—and its leaders—under increasingly realistic conditions.

So before we incorporate this new test to select the generals, let’s test the test, because that might prove more beneficial than testing the generals.