Towards Transparent Marine Autonomy: Visual Mission Composition for Teleoreactive AUVs
Christopher Thierauf, Charlotte Stieve, IEEE/MTS OCEANS 2026
Can a scientist plan a deep-sea robot mission themselves, without learning to program it?
This summarizes a recent publication, which you can read here. It builds directly on the cognitive teleoreactive framework I’ve been developing for AUV Sentry: that work gave Sentry the ability to reason about goals, and this work is about giving people a way to hand it those goals. Two open-source packages are released alongside it, which I’ll get to below.
Why We Need Something New
Mission-planning interfaces are the main way people and marine robots interact, yet very little of the thinking from cognitive science and human-robot interaction has made it into them. Most tools fall into one of two camps. Some, like Sentry’s existing Mission Controller (MC) and MOOS-IvP, are configured through text files that only experts can write. Others are visual, but don’t give the vehicle much autonomy: a human still turns intent into waypoints.
Planning systems that accept goals do exist, and they dramatically raise what the robot can do. But historically they’ve demanded that someone author a planning domain in a specialized language, which keeps them in the hands of specialists. So in practice, the people who actually want the data (the science users) can’t specify a mission themselves. They describe what they want to an expert operator, who translates it into a script, and they go back and forth until the launch deadline.
That setup causes a few problems:
- Miscommunication. The scientist has to trust that the operator understood them, and the operator has to interpret needs outside their own expertise.
- Opacity. If the scientist can’t see what the vehicle will do, they can’t judge whether it will do what they need.
- Miscalibrated trust. When autonomy is opaque, people either over-trust it and commit to missions it can’t complete, or under-trust it and hold it to minimal autonomy even when a more adaptable plan would work better.
There’s also a social cost. When the vehicle can’t communicate its own capability or intent, the operator becomes the go-between, and the client’s trust lands on the operator rather than the vehicle. Whether that operator feels able to speak up about a mission’s weaknesses is a question of psychological safety, which is something I’m increasingly interested in for teams that include autonomous agents.
The Key Idea: Interpretability Is Really Abstraction
The usual objection to planning systems is that domains are hard to author and outputs are hard to interpret. My argument in this paper is that this is mostly a problem of abstraction. Domain authoring is an engineering problem that the system builder can solve once. After that, the actual tasking a user needs to do is a small, tractable subset of the planner, and that subset can be represented as a handful of composable blocks.
So the interface separates what the operator decides from what the system reasons about, generates, and shows back to them. I think of this as “planner as a partner.”
The Interface
The whole thing is a web application, so there’s nothing to install, and it speaks the same HTTP API as the vehicle, so plans can be tested against a simulator before deployment. It’s built in Python with the Reflex framework and a FastAPI backend, and it can also run locally when there’s no internet at sea. Under the hood, every plan is just a Python file of calls to the mission executive’s API, which means saving, sharing, and archiving a plan is as simple as sharing that file.
It has three main parts:
- A geographic workspace. Users drag points and polygons onto a map, set parameters like altitude and speed (with presets for common cases like multibeam or camera surveys), and see the generated tracklines, optionally in 3D over imported bathymetry. Marking a zone as excluded turns it red, keeps the planner from routing through it, and sets up a geofence the vehicle enforces during the dive.
- A block-based plan constructor. Instead of writing symbolic goals, users stack blocks. A survey block asks for full coverage of a zone, a goto block asks the vehicle to visit a point, a plan block passes a general goal to the planner, an action block runs a behavior directly, and an order block lets the user require a sequence without hand-specifying everything in between. These map onto the levels-of-robot-autonomy framework, from near-teleoperation (action blocks) to supervisory control where the planner decomposes and optimizes. If a plan is infeasible, the planner says why: an unsatisfiable constraint, no valid way to achieve a goal, or not enough time or battery.
- A plan inspector. This is where transparency comes in. It’s built around the situation awareness-based agent transparency model, which asks what an agent needs to communicate: what it’s doing, why, and what happens next. A hierarchical task timeline answers the first, the step-by-step breakdown answers the second, and synchronized resource, battery, and depth charts answer the third. So before committing to anything, a user can answer questions like “will this mission run out of battery?” using performance data accumulated from prior missions.
The mission timeline and charting components came out of Charlotte Stieve’s internship with me in the Deep Submergence Lab. She built and released them as two open-source packages: reflex-apexcharts, which brings the ApexCharts library into Reflex for anyone building Python web apps, and reflex-mission-graphs, which packages the mission timeline so it can be dropped into other mission-planning interfaces.
How Does It Compare?
For a first evaluation, I used a comparative cognitive walkthrough: a structured method where you step through tasks and ask, at each step, whether an untrained user would try the right thing, notice the right control, connect it to their goal, and see that it worked. The modeled user is a typical Sentry science user, someone who knows exactly what data they want and is comfortable with maps and coordinates, but doesn’t know Sentry’s scripting tools or the physical constraints that make some paths infeasible. The baseline is the current workflow, which runs to roughly 27 steps on Sentry’s internal checklist, every one of them the operator’s.
I walked two scenarios in full and looked at three more as capability contrasts, scoring against my own interface whenever something was uncertain:
- A single survey. Both approaches have barriers here. In the visual interface, parameters like swath width are vehicle jargon, which presets help with. In MC, the user has to specify a path rather than an intent, and has to already know which commands exist.
- A multi-zone survey. The planner optimizes zone order unless told otherwise, which could surprise a user who expected their order to be kept, so I counted that against the interface.
- Exclusion zones. The visual interface supports keep-out zones the planner routes around and the vehicle enforces. MC has no equivalent; you just avoid drawing tracklines through the area.
- Mid-mission re-tasking. A new plan can be built from the vehicle’s current state and sent acoustically, as we demonstrated at sea in prior work. With MC, a change means aborting, recovering, and redoing the whole checklist.
- Non-survey goals. A plan block hands decomposition and ordering to the planner, where MC requires an expert to break the goal down by hand.
The main takeaway is that a user never has to touch a fluent, precondition, or any other symbolic planning concept to succeed. That shifts roles in a useful way: the scientist composes and checks the mission directly, and the operator becomes a safety approver who retains a veto. Each judgment ends up with the person who actually has the expertise to make it.
To be clear about what this does and doesn’t show: a cognitive walkthrough predicts where users will get stuck, but it involved no human participants and can’t measure workload, error rates, or whether trust is actually better calibrated. What it does show is a route toward calibrated trust that the current workflow doesn’t offer.
Why It Matters
The features here only exist because there’s a reasoning and planning system running onboard. Behavior trees, state machines, and linear command sequences can’t interpret goals, check feasibility, or re-plan mid-dive. But none of that power is useful if only its designer can drive it. By moving domain construction to the system builder, goal-level autonomy becomes something scientists, operators, and other non-experts can use directly. And while this was built for Sentry, the ideas aren’t specific to it; they apply anywhere an operator deploys an autonomous system on someone else’s behalf.
What’s Next
- A proper user study measuring workload (NASA-TLX), time to plan, and error rates.
- Exploring team dynamics: are operators more willing to discuss risk, and clients to revise objectives, when both are directly in the loop?
- At-sea data on re-planning events and intervention frequency to test the walkthrough’s predictions.
- Better ways to surface the planner’s self-assessments and failure recovery reasoning (see my ERAS work on failure) to everyday operators.
