August
2026
Annotation-Free Learning for Mobile Robot Navigation in Unstructured Environments
Authors:
Abstract:
Navigation in unstructured environments is a capability critical to many robotics applications such as forestry, construction, disaster response and defense. In these domains, robots have the potential to eliminate much of the dull, dirty and dangerous work that is currently performed by humans. Unfortunately, these environments pose a unique set of challenges for navigation that are not commonly found in controlled settings such as warehouses and urban driving.
We argue that the most significant of these challenges is building and utilizing a robust, yet sufficiently expressive representation of the robot's local environment for intelligent navigation decisions. Such a representation, in the context of unstructured environments, must encode a high level of geometric and semantic detail while using relatively modest computational resources, and must be able to do so in the presence of occlusion, sensor noise, and aggressive motion. Downstream navigation policies must leverage this representation to make correct navigation decisions despite the fact that unstructured environments are highly uncertain and complex, and often lack a straightforward mapping from semantics to decision-making. In practice, deployment of these systems requires extensive tuning of perception-planning interfaces, labeling of semantic images, etc. and is often brittle to changes in environment or embodiment. Unfortunately, this limits the widespread adoption of mobile robots, as existing systems are highly tuned to a given robot platform and environment, and are the result of months to years of dedicated effort from highly skilled teams of engineers.
This thesis will take a step towards eliminating this barrier by presenting a fully annotation-free system for robot navigation in unstructured environments. Importantly, this system is flexible, yet reliable in that it leverages large-scale, end-to-end, self-supervised learning scaffolded by a structured, environment-generic, geometric-semantic representation to enable high-speed navigation without and environment-specific annotations.
We first propose a voxel-based local mapping module enhanced with visual foundation model (VFM) features as an environment-generic, structured representation for navigation. We then leverage this representation to learn robust navigation policies from expert demonstration for aggressive driving in off-road terrain. Finally, we show that using self-supervised perceptual pretraining on our structured environment representation improves downstream policy capability and outperforms several state-of-the-art approaches for aggressive off-road driving. We validate our approach with extensive experimentation in multiple off-road environments using a full-scale autonomous all-terrain vehicle.
We argue that the most significant of these challenges is building and utilizing a robust, yet sufficiently expressive representation of the robot's local environment for intelligent navigation decisions. Such a representation, in the context of unstructured environments, must encode a high level of geometric and semantic detail while using relatively modest computational resources, and must be able to do so in the presence of occlusion, sensor noise, and aggressive motion. Downstream navigation policies must leverage this representation to make correct navigation decisions despite the fact that unstructured environments are highly uncertain and complex, and often lack a straightforward mapping from semantics to decision-making. In practice, deployment of these systems requires extensive tuning of perception-planning interfaces, labeling of semantic images, etc. and is often brittle to changes in environment or embodiment. Unfortunately, this limits the widespread adoption of mobile robots, as existing systems are highly tuned to a given robot platform and environment, and are the result of months to years of dedicated effort from highly skilled teams of engineers.
This thesis will take a step towards eliminating this barrier by presenting a fully annotation-free system for robot navigation in unstructured environments. Importantly, this system is flexible, yet reliable in that it leverages large-scale, end-to-end, self-supervised learning scaffolded by a structured, environment-generic, geometric-semantic representation to enable high-speed navigation without and environment-specific annotations.
We first propose a voxel-based local mapping module enhanced with visual foundation model (VFM) features as an environment-generic, structured representation for navigation. We then leverage this representation to learn robust navigation policies from expert demonstration for aggressive driving in off-road terrain. Finally, we show that using self-supervised perceptual pretraining on our structured environment representation improves downstream policy capability and outperforms several state-of-the-art approaches for aggressive off-road driving. We validate our approach with extensive experimentation in multiple off-road environments using a full-scale autonomous all-terrain vehicle.
Notes:
copied = false, 2000);
">
@phdthesis{Triest-2026-88358,
author = {Samuel Triest},
title = {Annotation-Free Learning for Mobile Robot Navigation in Unstructured Environments},
year = {2026},
month = {August},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-82},
}
author = {Samuel Triest},
title = {Annotation-Free Learning for Mobile Robot Navigation in Unstructured Environments},
year = {2026},
month = {August},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-82},
}