A word I abandoned

There is a word I eventually abandoned: console. For several months, I spoke of an “Operations Console.” The register was that of the command post, the starship bridge, the interface that holds authority. It sounded good. It corresponded to nothing.

What I do every day — for months now — is sit down in front of two screens. On the left, VS Code with the orchestrator terminal, files, git. On the right, Firefox with Claude.ai, web navigation, deliverables to read. I launch agents, I read what they produce, I correct, I relaunch. In the evening, I close everything. It is work. It is a workstation.

The finding that triggered the reframing came after several weeks of practice: the real problem is not to build a spectacular system, it is to reduce friction in a workflow that already exists. Four concrete frictions, observed daily, repeated session after session. Markdown files that are unreadable in a Windows file explorer. Manual navigation to find the right file each time. Back-and-forth between tabs that breaks the chain of thought. System metrics invisible without manual action — remaining token budget, quota, remaining session context — that force the operator to size deployments blindly.

The workstation eliminates these four frictions. No more. The workflow remains the same. The operator performs the same actions, in a better environment.

In short: think of an air traffic controller. The job of monitoring, validating, and redirecting trajectories existed before the control tower. The tower does not change the job — it reduces the friction that would let the controller miss a plane, fail to hear an alarm, or lose ten seconds at every transition. The agent operator’s workstation plays exactly that role.

An operator, not a developer

What characterizes this profession is the posture. The agent operator does not code — or very little. They do not directly produce the deliverables. They orchestrate, supervise, validate. They give instructions, read results, decide what comes next. This is HITL — Human-in-the-Loop. A human in the loop, with veto rights and a supervision responsibility.

The developer writes code and debugs their own errors. The agent operator does something else: they debug someone else’s errors — a “someone” that does not exist, a statistical model that produces plausible text without understanding it. The operator’s judgment is the only independent source in the system. The agents and the copilot are the same model: same weights, same biases, same convergence. Only the human constitutes an outside viewpoint.

This observation changes everything about workstation design. The operator’s primary tool is not a code editor — it is a reader. Their critical task is not to write but to read. And to read with enough attention to catch what a statistical model can produce as false while presenting it with perfect confidence.

What the literature confirms

The literature review, conducted with parallel agents across approximately 160 identified sources, grounds the framing on several solid points.

The first is that interface design constitutes a control protocol, not an ornament. Thirty years of work on Ecological Interface Design (Vicente and Rasmussen, 1992) and Situation Awareness (Endsley, 1995) converge: the interface must make the domain’s constraints directly perceptible, without intermediate mental computation. Burns et al. (2008) empirically validated EID in a full-scope nuclear simulator with licensed operators — significantly superior results in scenarios without procedure. This is the value niche that interests me: unanticipated situations, outside procedure, where design becomes the only available control protocol. When orchestrating 22 agents in parallel and one of them behaves unexpectedly, there is no procedure. There is what the interface shows.

The second point concerns automation bias. Skitka et al. (1999) measured 41% omission errors with an automated system versus 3% without. Parasuraman and Manzey (2010), in a review cumulating over 1,500 citations, conclude that the bias persists among experts and resists simple training. The central assertion of the framing — “if the workstation is poorly designed, the operator validates without reading” — is directly supported by this convergence between automation bias and processing fluency.

Third point: the cost of supervision is quantifiable. Cummings and Guerlain (2007) empirically identify a saturation threshold at 70% operator occupancy rate. The fan-out model of Goodrich and Olsen (2003) formalizes supervision capacity: 12 to 15 entities under light monitoring, 3 to 5 under active control. The factor between the two is 3 to 5x. The 30-40% overhead figure that field experience shows falls within the documented range.

Finally, the trajectory from HITL to HOTL (Human-on-the-Loop) is theorized by three coexisting autonomy level frameworks: Sheridan-Verplank (1978), Parasuraman-Sheridan-Wickens (2000), and SAE J3016. Adjustable autonomy (Pynadath et al. 2002) is the dominant operational mechanism: agents dynamically transfer control to the human in key situations.

What contradicts — and what worries

The most uncomfortable finding for the framing comes from Reber, Schwarz, and Winkielman (2004): processing fluency is epistemically marked. What is fluent is perceived as true. Robins and Holmes (2008) confirm: the same content with higher aesthetic design is judged more credible in 67% of cases, before the reader has even read the text.

Translated for the workstation: if the markdown rendering is beautiful, the operator risks validating faster. Readability helps reading, but it can also encourage automatic validation. The framing now integrates deliberate friction as a countermeasure — highlighting risk zones, checkpoints before validation, enforced pauses. But the question of dosing between readability and friction remains entirely open.

Alarm fatigue is the second danger. In critical care, 72 to 99% of alarms are false alerts. The Joint Commission documented 80 deaths between 2009 and 2012 directly linked to this phenomenon. The mechanism is mundane: when everything sounds the same, nothing sounds anymore. Applied to the agent monitor: if every in-progress agent displays the same generic status, if every slight delay generates the same signal, the operator will end up ignoring the monitor. Design must hierarchize: background information, alert, alarm. Displaying everything at the same level amounts to displaying nothing.

The dual-screen raises a more subtle problem. Gallagher et al. (2021), in a systematic review of 18 studies published in Human Factors, conclude that no study demonstrates a strong performance increase with multiple screens. Users strongly prefer dual-monitor — but this preference does not translate into measurable objective gain. My conviction that “head left = control, head right = reflection” is consistent with Grudin’s (2001) model of focal/peripheral partitioning. But this specific thesis — one screen = one cognitive mode — has never been experimentally tested.

And then there is Bainbridge. His 1983 article in Automatica, with over 1,800 citations, remains the most devastating for anyone who talks about “gradually releasing control.” The more reliable the automation, the less capable the operator is of taking back control when it fails. Strauch (2017) confirms that the problem remains unresolved, 35 years later. Air France 447 is its most terrible manifestation. The nuance that partially saves the HOTL trajectory: Wickens (2018) shows that short cycles of manual/automatic alternation do not produce skill degradation. Work sessions with an orchestrator last 2 to 4 hours with regular active intervention. This is not passive monitoring — it is piloting. The question is: will it remain so when the system becomes more reliable?

The concept exists, fragmented and unnamed

The agent operator’s workstation already exists in the literature, but fragmented under five distinct terminologies. In human factors: Supervisory Control Station. In the LLM industry: Agent Control Plane. Industry players speak of Supervisory Intelligence (XMPro) or Agent HQ (GitHub). Frontend developers refer to Agentic UI. The academic field has no unified founding paper — it is an archipelago of contributions from aviation, robotics, nuclear, and LLMs, without synthesis.

Industry is ahead of research. GitHub Agent HQ, Microsoft Agent 365, Dify, LangSmith — functional operator interfaces exist before any academic theorization. And yet, a serious argument stands against a dedicated workstation: existing tools already cover many needs. Cline and Cursor provide granular approval, reviewable diffs, rollback, audit trail. The METR study (2025, RCT N=16) shows that AI tools lengthen time by 19%. Agarwal et al. (2026) find that autonomous agent gains are zero when an AI IDE is already present.

The nuance that maintains the framing rests on two points. First, Cline and Cursor are single-session: supervising N agents on N branches simultaneously — the parallel multi-agent orchestration pattern — has no native IDE interface. This is the only documented gap potentially justifying a dedicated layer. Second, all “You Ain’t Gonna Need It” literature concerns developers. An operator supervising agents on non-coding tasks — research, orchestration, content production — is out of scope for these studies.

Workstation design: two screens, two postures

The workstation is organized around two distinct postures, physically separated.

The left screen is the command screen. The operator monitors, launches, stops, reads metrics. The permanent question: “is the system working and what state is it in?” At the top, budget indicators — tokens remaining in the sliding window, requests remaining in the quota, remaining session context. In the center, the agent monitor: the state of the current deployment, dependencies, durations, agents exceeding expected time. At the bottom, the orchestrator terminal and shortcut buttons for frequent actions.

The right screen is the reflection screen. The operator reads deliverables, dialogues with the copilot, builds their thinking. The permanent question: “what did the system produce and what do I do with it?” A compass at the top — the planned tasks for the session and their progress — prevents scope drift. An integrated copilot tab places the AI assistant dialogue in the same visual field as the deliverables. An explorer tab provides clean markdown rendering. A separate thesis space tab protects the independence of judgment: the explorer shows what the system produces, the thesis space shows what the operator thinks.

Turn your head left: control. Turn your head right: reflection. The physical gesture marks the mental transition. This is already what the operator does daily for months. The workstation formalizes this habit; it does not invent it.

In short: a single multi-tab screen forces invisible cognitive transitions — the operator switches mental mode without a physical signal. Two dedicated screens (command on the left, reflection on the right) turn each transition into a bodily gesture: you turn your head, you switch mode. Switching cost drops because it becomes observable.

The gaps — what everyone is missing

The most striking gap is the complete absence of specific literature on human supervision of LLM agents. The 160 identified sources come from aviation, nuclear, robotics, medical, autonomous driving, and cognitive psychology. None directly addresses the case where a human operator supervises an LLM multi-agent orchestrator.

Transfer from these domains is plausible — the cognitive mechanisms at play (automation bias, switch cost, alarm fatigue, out-of-the-loop) are generic. But it is not validated. Supervising robots and supervising LLM agents differ on a fundamental point: robots execute physical actions in an observable world, LLM agents produce text in a semantic space. “Taking back control” does not mean the same thing.

The fan-out for orchestrated LLM systems does not exist. How many agents can an operator effectively supervise in parallel? Practitioner experience shows deployments of 12 to 41 agents, but there is no measurement of supervision quality as a function of number. The counterfactual “dedicated workstation vs. existing IDE” has never been measured either.

There is no synthesis paper establishing the equivalence between Supervisory Control Station (aviation), Agent Control Plane (LLM), and Agent Observability Dashboard (dev tooling). Writing this paper would be a real contribution.

What this study says about the profession

The agent operator is not a developer. They are not a project manager, a product manager, or a DevOps engineer either. This is a profession whose daily actions more closely resemble those of an air traffic controller or a nuclear control room operator than those of a software engineer.

They share with the air traffic controller the simultaneous supervision of autonomous entities, the need for permanent situational awareness, the ability to intervene quickly when something deviates. They share with the nuclear operator the management of systems whose internal functioning is partially opaque, the necessity of detecting anomalies before they cascade, the responsibility of a human gate on an automated process.

What distinguishes them from all these professions is the semantic nature of what they supervise. Agents do not move planes or control reactors. They produce text. The failure mode is not a physical accident but a confident hallucination, a silent drift, a plausible but false deliverable. The operator reads, and it is in reading that they exercise control. Their primary tool is their attention.

This is why workstation design is not a luxury. It is the control protocol itself. If the workstation is poorly designed — if the files are unreadable, if the metrics are hidden, if context switching is constant — the operator validates without reading. And an operator who validates without reading is a system without supervision. An open control loop, running and producing plausible output while nobody verifies whether it is true.

The profession already exists. The name does not yet. The dedicated workstation either. But the actions, the frictions, the biases, the risks — all of this is documented, in this study and in 160 sources that do not know they are talking about the same subject.