Design in the Loop
"Human in the loop" is a governance phrase that goes quiet the moment a product team has to decide what goes on the screen. What replaces it is a set of design questions specific enough to build against.
TL;DR
"Human in the loop" isn't DoD policy and never was. Directive 3000.09 requires appropriate human judgment, which is a different standard and a harder one.
The oversight problem breaks into comprehension, monitoring, and intervention. Intervention gets attention, monitoring holds at current scale, and comprehension is where I keep seeing the gap.
Programs keep asking how one operator commands 2,000 agents. I've watched operators max out around six to eight, and the map view stops being useful long before 2,000.
Trust isn't a single dial. It moves by phase of the kill chain, and only two of its three components are things an interface can touch.
Three questions replace the phrase. Where autonomy operates in the OODA loop and at what level, what situational awareness the operator needs at that phase, and what the interface has to show to make that awareness possible.
Every autonomy program I've worked on describes the human as being in the loop, and none of them can tell you what that means at the interface. The phrase has been carrying policy weight for years while the design work underneath it went unspecified.
Governance Language
Every program briefing I've been in describes the human as being in the loop. A pilot keeps authority for lethal engagement. An air defender can intervene. An operator sets objectives and watches behavior. The framing is consistent across programs, and it's reassuring in exactly the way it's meant to be.
Then you're two months in, sitting with the engineering team actually building the thing, defining what the interface shows, what the operator sees while the agent is deciding, what happens when the system hits a scenario nobody scoped. Nobody in that room says "human in the loop." The phrase never comes up, because it isn't specific enough to be relevant.
That's worth sitting with. A term that shapes policy and satisfies oversight evaporates the second a team has to decide what is displayed. It says a human should be involved in some meaningful way, and leaves open what they perceive, what they understand, what they can change, and who's accountable when the system acts.
It's also not the policy. DoD Directive 3000.09 doesn't require a human in the loop. The omission was deliberate. What the directive requires is "appropriate levels of human judgment over the use of force." The argument is that loop language puts the machine's decision cycle at the center and then asks where the human stands relative to it, which has it backwards. I'd go further than that. Appropriate human judgment is the harder standard, and it implies a design brief nobody has written.
Comprehension Is the Gap
Underneath the phrase there are three things an operator has to be able to do. Understand what the system is doing and why, in terms they can act on. Track that behavior as conditions change. Step in before a bad decision becomes an operational or ethical consequence. Comprehension, monitoring, intervention.
Intervention is where defense teams start, almost without exception. The kill switch is top of mind for everyone I've worked with in this space, and safety and control get serious attention. Monitoring is manageable at the scales we're building for now. Comprehension is where the gap sits, and it's the one that gets worse as systems get larger and faster.
Think about it in OODA terms. If the system is observing and orienting faster than the operator can follow, the operator's decide and act phases are running against an understanding that's already stale. That's a temporal gap before it's an ethical one, and policy can't close it. The information architecture and interaction model determine whether meaningful supervision survives at machine speed.
The Scaling Problem
Most of what I work on is small. Six to twelve agents. When I look at what an operator can realistically command and control under mission pressure in a single domain, the number is closer to six to eight. These systems are mostly mobile, so the interaction model is still heavily map-driven. Agents, routes, tasks, telemetry, status, all of it top-down.
DoD stakeholders keep asking a different question. Across programs and across clients, some version of the same prompt. How does a single operator command and control 2,000 autonomous agents? I don't know where that number came from. It shows up constantly, with a confidence that doesn't match anything I've seen demonstrated.
At 2,000 tracks the map view is meaningless. Density overwhelms comprehension and the display stops communicating anything a person can decide from. The interaction model that works for eight agents doesn't scale, and I doubt whatever replaces it is a map at all.
At that scale a human isn't going to be involved with every agent, which means the level of autonomy has to rise and the oversight model changes shape. It becomes a lot of small loops running on their own with an operator working the larger one. "Human in the loop" was written for a world where operators make decisions about discrete systems, and it carries over into a world where they don't without anybody noticing the substitution.
Automation, Not Autonomy
There's a distinction that gets blurred constantly. Most of what I encounter is automation. Sensor fusion, deterrent inventory checks, weapon pairing, route execution, rule-based responses. Defined procedures executed reliably, not systems generating novel plans in response to conditions nobody anticipated.
Mission autonomy is the other thing, where a system takes goals, constraints, and intent, then works out execution as conditions change. I should be clear that I haven't designed for that end of the spectrum. My own work is only starting to get near it, and what I know about high-autonomy behavior comes from research and discussion. The transition still matters, because the interface for a system executing instructions does not look like the interface for a system writing its own goals-based plan.
CCA is an apt example. Today it's constrained teaming and supervisory control. The operator tasks, the system executes and reports back, and the cognitive model stays close to supervision. The vision programs describe is considerably more autonomous, with operators setting objectives, constraints, and rules of engagement while the aircraft determine execution and the operator spends their attention on whether behavior is diverging from what they expected. That's a different relationship and not a more advanced version of the current one.
I've heard this referred to as "puppet mastering," which I disagree with on premise. Puppets imply direct connection and manipulation, which isn't the relationship we’re envisioning and can't work at scale. I think of it more as a conductor. A conductor doesn't play the instruments and the orchestra still moves together. The relationship is autonomous orchestration
And it's worth saying where that interface is going to land. A fighter pilot is already managing sensor fusion, tactical coordination, navigation, communications, electronic warfare, threat interpretation, and weapons employment under compressed timelines. I've sat in a fighter cockpit as a non-pilot and it's hard to overstate how dense it already is. Autonomous wingman supervision doesn't arrive into a clean workflow.
Trust Shifts By Phase
I spent a lot of last year working in air defense environments, and what became obvious quickly is that trust isn't a single dial an operator turns up or down across a system. It moves by phase.
Trust is low at identification. That's where an operator decides whether a track is a legitimate threat, the consequences of being wrong are enormous, and I've yet to see an interface that communicates system confidence well enough that an operator will act on it without investigating further. Trust runs considerably higher at deterrent selection and weapon pairing, where the logic is constrained and procedural and deferring to the machine is comfortable. The oversight relationship is already different at each phase of the kill chain, and "human in the loop" applies one framing across all of it.
Hoff and Bashir's trust model, which Michael Mayer applies to military AI, turns that observation into something usable. Trust has three components. Dispositional trust is the baseline a person brings, shaped by culture and background, and no interface touches it. Situational trust moves with workload, stress, and time pressure, and design influences it by reducing load. Learned trust is built through reliability, transparency, error framing, and alert design, and that one is fully ours. Knowing which component dominates a given system tells you whether you're looking at a design problem, a training problem, or a doctrine problem. I want design to name that boundary for itself, because being held accountable for a failure an interface can't reach helps nobody.
Complexity Transfers
When design isn't part of defining these systems early, the complexity doesn't disappear. It transfers, usually to training. Cycles get longer and operators adapt themselves to interfaces that weren't built around how they reason under pressure. The assumption is that autonomy reduces workload, and it does reduce execution workload while raising supervisory workload, which is a different burden and a harder one to train around. Resilience shouldn't become the justification for complexity that was avoidable in the first place.
There's an honest limit in my own evidence here. Most of what design teams observe happens in demonstrations, which are not research events. Demonstrations show whether a system functions. I've watched soldiers in those rooms slouched in chairs, going through motions, waiting to be released, and that isn't what an operator under mission pressure looks like. The conditions that expose comprehension failures and trust breakdowns are largely the conditions design teams don't get access to. Demonstrations are mostly what we get, so we work with them, and the confidence placed in what we see there should stay proportional to where it came from.
Three Questions, Not One Phrase
What I've found most useful is to stop arguing about the loop and start specifying three things.
The first is where machine autonomy operates in the OODA loop, and at what level. Parasuraman, Sheridan and Wickens published a ten-level automation spectrum across four functional stages in 2000, and it does something in-or-on-the-loop language can't. It's specifiable, auditable, and allowed to differ by phase. A system can sit high on information acquisition and low on decision selection, and saying so out loud is a real answer where "human in the loop" is not.
The second is what situation awareness the operator has to hold to exercise judgment at that phase. Mica Endsley's model, published in 1995 and still the working standard in human factors, gives three levels. Perceiving what's there, understanding what it means, projecting what happens next. Appropriate human judgment, the way I read it, needs that third level. Most of the interfaces I've seen are built to deliver the first one well and then stop.
The third is what the interface has to provide to make that awareness possible under operational conditions. That's the design brief, and the target is comprehension, not a menu of options. It covers system confidence, failure conditions, the system's own account of where it's unreliable, and trigger guards at the points where automatic deference needs interrupting.
The Howard University group working on human-autonomy teaming adds a fourth consideration I'd keep, which is ethical decision support built into the interface instead of handled at the policy layer. If an action carries moral weight, the moment that weight becomes legible to the operator is a design decision, and at the moment it's usually nobody's.
Where this gets genuinely hard, is time. If operational tempo leaves no room to exercise judgment, the question shifts from what the interface should provide to what it should prevent, which is a different brief with different mechanics. Mayer documents that speed problem without resolving it, and I haven't resolved it either. It's the part I keep turning over.
What I'm more confident about is that these are answerable questions. Where autonomy operates, what the operator needs to know, what the screen has to show. Those can be written into a program document, argued about in a design review, and checked afterward against what the system actually did. The oversight model is being written right now in policy documents and acquisition plans, and it's going to be written into interfaces either way. There's still time for the interface to be part of that argument instead of the record of it.
Further Reading
Air Force Collaborative Combat Aircraft (CCA) Program – Congressional Research Service
Autonomy in Weapon Systems (PDF) – Department of Defense, Directive 3000.09
Bias, Explainability, and Trust for AI-Enabled Military Systems – SPIE
AI-Driven Human-Autonomy Teaming in Tactical Operations – ArXiv