Ship-Relative UAV Pose Estimation with 3D LiDAR

Table of Contents

  1. Challenges in Shipboard Pose Estimation
  2. Dataset Generation and Collection
  3. Proposed Model Architecture
  4. Training Procedure
  5. Results and Evaluation
  6. Future Directions
  7. Related publication

Accurate ship-relative navigation is essential for autonomous landing and UAV operations at sea. GPS can be disrupted by jamming or spoofing, while camera-based perception is sensitive to lighting and visibility. LiDAR offers complementary measurements of three-dimensional ship geometry, with less dependence on ambient illumination.

Research by Karl Simon, Maneesha Wickramasuriya, Taeyoung Lee, and Murray Snyder develops a learning-based pipeline that uses 3D LiDAR scans to estimate the full six-degree-of-freedom (6DoF) pose of a UAV relative to a ship. The model was trained in simulation and validated with real-world data collected on the U.S. Naval Academy’s YP689 research vessel. The work complements FDCL’s vision-based maritime flight research, using geometric structure in 3D scans rather than monocular images.


Challenges in Shipboard Pose Estimation

Scans are often sparse and incomplete, especially at longer ranges, and occlusions can hide parts of the vessel. Classical registration algorithms such as ICP1 can converge to the wrong alignment when overlap is limited or the initial pose estimate is poor. Learned registration methods such as PointNetLK2 can also struggle with these partial views. The aim is therefore to recover recognizable ship structure from a single scan and use it to initialize alignment.

Sparse LiDAR returns misaligned with the ship model after ICP registration.
An example of incorrect ICP alignment with a sparse, partial scan. A learned geometric estimate provides an initialization for subsequent refinement.

Dataset Generation and Collection

Developing any deep learning model requires large amounts of training data, which is especially difficult to obtain in real maritime settings. To address this, we built a simulation environment in Gazebo using a CAD model of the YP689 training vessel. Within this environment, a virtual Ouster OS0-32 LiDAR was mounted on a UAV model to generate thousands of scans. These synthetic scans capture the ship geometry as viewed from varied positions and orientations representative of actual flight trajectories.

In addition to simulation data, we collected real-world validation data during experiments aboard the US Naval Academy’s YP689 research vessel in the Chesapeake Bay, MD. A UAV equipped with a real Ouster OS0-32 collected LiDAR data during full-length flight trajectories around the ship. High-accuracy ground-truth poses were also obtained using inertial integration3, as well as RTK GPS for benchmarking.


Proposed Model Architecture

The core of the system is a Point Transformer-based neural network adapted for ship-relative pose estimation. Forty canonical 3D keypoints are defined around the stern and superstructure, where the flight deck and nearby structures provide distinctive geometry for launch and recovery operations. The network predicts these keypoints in the LiDAR frame, even when only part of the vessel is visible.

Red canonical keypoints distributed around the stern and superstructure of the ship model.
Canonical ship keypoints define the geometric reference. Corresponding predictions in a LiDAR scan are aligned with this model to recover relative position and orientation.

  • Input: A single LiDAR scan (downsampled to 1024 points, normalized).
  • Feature extraction: Learns local and global scan features using self-attention layers.
  • Keypoint prediction: Network predicts ship keypoints in the LiDAR frame, using fixed CAD keypoint embeddings as decoder queries.
  • Pose estimation: Predicted keypoints are aligned with the known CAD-model keypoints using the closed-form Kabsch algorithm5 to recover rotation and translation.
  • Refinement: Finally, a lightweight registration further refines the alignment, where the model output serves as the initial guess.


Training Procedure

The model was trained using 20,000 synthetic LiDAR scans. To better reflect real-world conditions, scans were augmented with Gaussian sensor noise and occlusions from flight deck obstructions. For the training, we utilized an NVIDIA A100 GPU and took about 4 hours to complete 100 epochs on our dataset.

A key aspect of this experiment is that the network was trained only on simulated data, allowing a direct assessment of transfer to real scans. Applying the approach to another vessel would require its geometry and suitable training data; the reported results do not demonstrate generalization to unknown ship classes.


Results and Evaluation

The simulation-trained model was evaluated on 750 real-world LiDAR scans, covering UAV trajectories near the stern at ranges up to 10 m. The network’s keypoint predictions produce an initial pose, which can then be refined using GICP4:

Estimation stage Mean rotation error Mean translation error
Learned keypoints and geometric alignment 2.62° 0.357 m
After GICP refinement 0.62° 0.073 m (7.3 cm)

More than 96% of refined predictions had translation error below 20 cm. These are pose-estimation results from recorded shipboard flight data; they do not establish closed-loop autonomous landing with LiDAR.

These results show that a model trained entirely in simulation can generalize to real-world LiDAR scans collected during shipboard UAV operations. The learned stage supplies its own pose initialization for registration. Some scans still produce larger errors, including cases in which water returns distort the predicted geometry; GICP cannot always recover from an inaccurate initialization.

Future Directions

Future directions include incorporating sea spray and dynamic occlusions into simulation, integrating LiDAR with vision and inertial measurements, and optimizing the model for onboard computers. LiDAR’s reduced dependence on lighting does not establish all-weather performance: degraded visibility and adverse-weather operation require further evaluation. FDCL’s mixed-reality maritime flight environment provides a related approach for testing perception and control before deployment at sea.


K. Simon, M. Wickramasuriya, T. Lee, and M. Snyder, Light-Detection-and-Ranging–Based Unmanned Aerial Vehicle Pose Estimation in Ocean Environment with Deep Point Transformer Network, Journal of Guidance, Control, and Dynamics, 49(8), 2277–2290, 2026.

This work is based on my M.S. thesis Ship-Relative UAV Pose Estimation with 3D LiDAR (2025), advised by Prof. Taeyoung Lee. The source code is available at github.com/fdcl-gwu/point-transformer.

1 : P.J. Besl and Neil D. McKay. A method for registration of 3-d shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(2):239–256, 1992.

2 : Yasuhiro Aoki, Hunter Goforth, Arun Srivatsan Rangaprasad, and Simon Lucey. Pointnetlk: Robust & efficient point cloud registration using pointnet. pages 7156–7165, 06 2019.

3 : Kenny Chen, Ryan Nemiroff, and Brett T. Lopez. Direct lidar-inertial odom- etry: Lightweight lio with continuous-time motion correction. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 3983–3989, 2023.

4 : Aleksandr Segal, Dirk Hähnel, and Sebastian Thrun. Generalized-icp. 06 2009.

5 : Jim Lawrence, Javier Bernal, and Christoph Witzgall. A purely algebraic justifi- cation of the kabsch-umeyama algorithm. Journal of Research of the National Institute of Standards and Technology, 124, October 2019.