A physical UAV flies indoors while its perception system processes photorealistic maritime imagery.

A UAV approaching a ship needs to know its position and orientation relative to the vessel. Waves, changing viewpoints, lighting, and limited access to reliable navigation signals make this a demanding perception and control problem.

FDCL’s maritime autonomy research connects deep visual perception, geometric estimation, and experimental validation, in collaboration with the U.S. Naval Academy. It builds from shipboard visual–inertial flight experiments to transformer-based perception and mixed-reality testing of a complete autonomous flight system.

The deep-perception work led by Maneesha (“Maneesh”) Wickramasuriya forms a central part of his doctoral dissertation, Deep Transformer Network for Autonomous UAV Launch and Recovery in Ocean Environments.

The foundation: autonomous flight from a moving ship

Earlier work by Kanishke Gamagedara, Taeyoung Lee, and Murray Snyder developed the onboard hardware, estimation, and control software for autonomous launch and landing from a U.S. Naval Academy research vessel in Chesapeake Bay. The experiments tested both RTK-GPS-based relative positioning and vision-based flight.

A delayed Kalman filter fuses measurements according to when they were acquired: a delayed observation corrects a past estimate, which is then propagated to the current time. This makes the estimation system account for sensing and processing delays while the aircraft and ship continue moving.

An excerpt from the earlier shipboard visual–inertial flight experiments. Insets show the onboard camera view and visual features used for navigation. These experiments preceded the transformer-based perception system.

Learning to recognize the ship’s structure

The transformer approach learns recognizable ship structures from a single camera image. It predicts geometric keypoints on several parts of the vessel, recovers an individual pose from each part, and combines those estimates through Bayesian fusion. Using several structures reduces dependence on the visibility of any one ship component.

Collecting hundreds of thousands of labeled images at sea is impractical. Instead, the researchers reconstructed the vessel’s geometry and rendered approximately 435,000 synthetic training images, varying camera poses, textures, backgrounds, and lighting. The network learns ship structure across these variations rather than relying only on matching local image texture between successive frames.

Synthetic ship images with varied viewpoints, vessel textures, skies, and ocean backgrounds.
Synthetic training images vary appearance while preserving the vessel’s geometry, allowing automatic generation of images and corresponding pose labels.

From synthetic images to shipboard observations

The 2025 JGCD study evaluated the trained network on synthetic test images and real images collected during shipboard flight experiments. The real-image tests included normal, overexposed, and underexposed views.

Evaluation Mean position error Mean rotation error
Synthetic test images 0.204 m 0.91°
Real images: overexposed ship 0.112 m 1.8°
Real images: underexposed ship 0.089 m 1.1°
Real images: normal exposure 0.177 m 4.0°

For the three real-image datasets, position error was 0.66–0.97% of each dataset’s maximum range. These results demonstrate transfer from synthetic training to real maritime imagery under the tested conditions. They evaluate pose estimation from flight imagery; closed-loop flight using the transformer is addressed in the later study below.

Predicted ship keypoints and bounding boxes overlaid on real shipboard images under normal exposure. Predicted ship keypoints and bounding boxes under strong sunlight and overexposure.
Validation on real shipboard images: predicted keypoints and projected ship-part bounding boxes under normal exposure (top) and overexposure (bottom).

Mixed reality: connecting perception to a flying UAV

The ICUAS 2025 study introduced a vision-in-the-loop environment using 3D Gaussian Splatting. Images captured around the research vessel are used to construct a photorealistic scene that can be rendered from new viewpoints.

Four real photographs of the research vessel in the top row, with corresponding Gaussian-splatting renderings in the bottom row.
Real vessel images (top) and corresponding 3D Gaussian Splatting renderings (bottom), illustrating the maritime environment used for indoor verification.

The ICUAS 2026 study extends this environment to autonomous closed-loop flight with onboard perception and estimation:

  1. A motion-capture system measures the physical UAV’s pose so that the renderer can generate the corresponding maritime camera view.
  2. Those RGB images are streamed to the onboard computer, where the transformer estimates ship-relative position and orientation.
  3. A delayed Kalman filter combines the visual estimates with high-rate inertial measurements, and a geometric controller uses the resulting current-time estimate to command flight.

The experiments demonstrated autonomous takeoff, trajectory tracking, and landing in the indoor maritime emulation shown in the opening video. The setup exposes perception latency, asynchronous updates, and onboard computing constraints while retaining physical UAV dynamics. It provides an intermediate validation stage before at-sea deployment of the integrated transformer-based system.

Complementary sensing and geometric estimation

The same emphasis on vessel geometry also supports LiDAR-based pose estimation. A point transformer learns 40 ship keypoints from sparse 3D scans, providing a complementary route to relative localization when image appearance is difficult to interpret.

Related work on invariant Kalman filtering for relative dynamics develops the geometric estimation theory for motion between two systems. This complements the demonstrated delayed-filter flight architecture and supports further research on ship-relative sensor fusion.

Pose-estimation code · Doctoral defense announcement · Complementary LiDAR research