Background & Summary

Satellite payloads that include geostationary sensors, frequently detect long thin artificially created ship tracks in boundary layer clouds above large oceanic regions. These features are a result of the high concentrations of cloud condensation nuclei (CCN) that form from the nitrogen and sulfur oxides present in ship exhaust. In other words, the pollutants in ship exhaust can seed low-lying marine clouds by enhancing cloud droplet formation. However, because the clouds are water vapor limited, even in the marine boundary layer, cloud droplets, albeit more numerous, are typically smaller in clouds that have been influenced by ship effluents1. While suppression of droplet size may reduce the precipitation produced by marine clouds, the enhancement of droplet concentration from ship exhaust results in higher cloud albedo2,3. The microphysical process through which ship exhaust modulates cloud reflectivity is similar to a proposed climate adaptation strategy known as Marine Cloud Brightening (MCB)4, which is actively being studied as a potential solar climate intervention mechanism to slow down the impact of global warming, by reflecting more solar radiation back to space5,6,7,8,9. The study of ship tracks: satellite observable phenomena of natural MCB experiments, is thus critical in the development of MCB strategies.

The relatively bright tracks that may result from higher cloud albedos can persist for several hours after initial injection, before tracks fully disperse and become visibly indiscernible from other cloud features. However, it has been documented that a significant fraction of incoming solar energy reflected by tracks likely occurs over a longer lifetime (within 1-3 days after injection), with other studies indicating that clouds may take between a few hours to a few days to adjust and respond to aerosol injection5. In addition to determining the efficacy of MCB, the ability to quantify cloud responses to aerosol on varying timescales is crucial in furthering the scientific understanding of Aerosol-Cloud Interactions (ACI), poised as the largest source of uncertainty to global radiative forcing estimates10,11. Ship tracks, as natural experiments of aerosol injection, provide valuable insights into ACI by illustrating how aerosols influence cloud properties and behaviours12.

While expensive flight campaigns have historically enabled the bulk of high-resolution ACI measurement collection13,14, they are limited in data diversity; capturing spatio-temporal information from only small subsets of clouds that may cause difficulties in directly attributing aerosols to specific sources/times, and also presenting challenges for evaluating the spatio-temporal evolution of cloud parcels after injection. Thus, extracting similar measurements within satellite-measurable ship track regions could enhance the availability of data required for validation, though necessitate detailed and high-resolution representations of ship track regions which to the best of our knowledge, are not fully facilitated from current datasets.

The recent advent in field experiments of marine cloud aerosol injection15,16,17,18,19 to measure MCB efficacy, could directly benefit from quantitative comparison of realised cloud responses and brightening with those satellite-observed from ship tracks. Such comparisons, however, require a more complete understanding of track spreading. Ship track datasets which are currently available20,21,22 cannot be used to accurately measure track spreading, as they only provide partial information about coordinate label sets (i.e., a small set of spatial coordinates which tracks lie on). This is in contrast to the fully masked ship tracks presented in this paper, which provide gridded pixel locations covering all visible aspects of tracks.

While coordinate labels for tracks have shown promise in monitoring the density of ship tracks across the globe23,24, for more comprehensive analysis of the aerosol indirect effect and MCB efficacy, tracking the entire expanse of ship tracks is required25,26,27,28,29,30. For example, studies on how ship exhaust seeds existing clouds25, the formation and persistence of ship tracks in space and time31,32,33, when and where ship tracks are most apparent24, and their corresponding effects on cloud microphysics34, including insights into their overall climate forcing effects23, all necessitate the identification of fully masked ship tracks. Recent studies assessing the spreading behavior of tracks require complete track masks, especially of track edges, to compare changes in: strength, longevity, and spatial extent of cloud response and albedo, to those captured by new parameterizations of aerosol injection behavior33,35. In particular, ensuring the validity of such parameterizations against fully observed tracks would strengthen the ability to replicate cloud responses (and other aerosol indirect effects), thus having the potential to improve aerosol injection representation (that is currently poorly accounted for), in coupled climate models.

Despite requiring full tracks to measure track spreading rates via the calculation of changing plume widths (if visible) at different times after injection35, obtaining image masks of tracks is a challenging task. When visible, ship tracks appear distinct in certain satellite images (see Fig. 1), though they often resemble naturally forming cloud features with similar pixel intensities and textural attributes. As such, tracks cannot be separated from existing clouds through simple image manipulation or the determination of a pixel intensity threshold. Thus, although they can be distinct to human eyes, it is nontrivial to separate them from surrounding clouds. To this end, several semi-automated computational algorithms25,36,37 have been developed to extract pixels containing ship tracks from satellite imagery, using coordinate training data defining ship track locations (see e.g.34). Such algorithms in general utilise a thresholding approach to define regions of track vs no track, with thresholds and sizes of track regions chosen at fixed values, which may thus not necessarily cover all types of track lengths, features and trajectory shapes. Our proposed dataset aims to address this limitation by providing locations of tracks via pixel labels, allowing for more precise analysis and enhanced applicability. For instance, fully labelled ship tracks can be used directly as masks or rasters for other satellite data products beyond those which were used to create the labels, such as cloud liquid water path (LWP) as studied in37.

Fig. 1
Fig. 1
Full size image

Examples of ship tracks as long thin streaks in clouds that become wider and more diffuse, around many other cloud features that are not always distinguishable as ship tracks. Raw images with projected (a) vs swath (b) taken by MODIS on July 15 2007 at 2:20 am in the North Pacific ocean, North East of Japan. Atmospheric Corrected Reflectance data plotted at the 2.1 μm band.

While there have been numerous studies of ship tracks utilising satellite imagery26,33, existing publicly available datasets of ship tracks only partially capture their intricate features, in comparison to the dataset we propose in this paper that fully captures tracks. As previously mentioned, for instance, the datasets of20,21 provide a list of latitude/longitude coordinates as track labels, as opposed to satellite image masks presented in this paper, which identify pixels representing ship tracks. Specifically, the 2019 ship track data of20 consists of a list of latitude/longitude coordinates corresponding to points that lie on the track’s centreline that runs adjacent to its (longer) side edges. Similarly, the 2022a dataset in21 labels each track over the years 2003-2020 by its date, time, mean latitude, and mean longitude; computed over all the pixels that are machine-learned as representing ship tracks. The dataset does not contain the image masks corresponding to the machine-labelled pixels representing ship tracks, nor the hand-labelled training dataset used to train the machine learning algorithm described in23. On the other hand, the 2022b ship track dataset provided and utilised in24 provides both hand-labelled training38 and derived machine learned data22, in the form of centreline polygon masks representing latitude/longitude coordinates of a polygon encapsulating the central region of each ship track. From the references therein, the utilised training data leverages an unpublished ship track dataset, first mentioned in34, that highlights hand logged labels consisting of a latitude/longitude coordinate corresponding to the head, or most recently formed part of the track, and each coordinate representing a track’s turning point wherein it abruptly changes direction to form its quasi-linear structure. The centreline polygon masks are then obtained by connecting each of these hand-labelled coordinates by straight lines separated by a static width of 10 pixels (approximating the average ship track width of 9 km), in which the centreline corresponds to the line separating the polygon in half. Figure 2 visualises the differences in track labelling amongst the datasets that currently exist, and the dataset we present that aims to mask the entirety of track features.

Fig. 2
Fig. 2
Full size image

Example of current databases’ masking techniques against the masking technique used in our proposed dataset. Track observed on June 15 2006 at 6:45 pm.

The unpublished hand-labelled dataset of34, highlighting track head and turning point coordinates, has since been either directly utilised or added to in many subsequent ship track studies36,37,39. In particular, the dataset utilised in36 is the 2019 dataset20 that we study in this paper. To visualise these differences in a full satellite image, Fig. 3 shows the same image as Fig. 1 with now labelled tracks in the form of: centreline coordinates (as shown by the 2019 dataset of20), a mean coordinate (as given in the 2022a dataset of21), centreline polygon marks (as given in the 2022b dataset of38), and satellite masks consisting of all pixels that contain a track (our contribution).

Fig. 3
Fig. 3
Full size image

Swath MODIS image containing ship track labels from Watson-Parris et al. (shown in green)38, Song et al. (shown in red)21 and Toll et al. (shown in black)20 shown against our proposed data masks (shown in orange). Image taken July 15 2007 at 2:20 am in the North Pacific ocean, North East of Japan.

Curating a new dataset of ship track image masks during this specific time period is important for multiple reasons. First, both optimising the efficacy of MCB for maximum solar reflectivity and benchmarking varied observations of ACI, requires understanding of the atmospheric conditions leading to ship track formation and maximal persistence. However, formation of ship tracks itself is heavily dependent on the amount of sulphur contained in their corresponding emissions that can viably produce small enough droplet sizes when seeded, to be discernible from the surrounding clouds by satellite23. Notably, sulphur oxides have seen more stringent regulations over the past decade, particularly in specified emission control areas (ECAs), to minimise the negative impacts of airborne ship emissions. These regulations have led to a significant decline in observable ship tracks since 201023,24. As a result, there is strong evidence to suggest that studying ship tracks before the first restrictive policy in 2010 provides a more robust understanding of track formation, persistence and spreading, given their higher probability of formation, detection and discernibility.

Before 201023, show that the year 2006 provides the highest ship track density in the North Pacific and South East Atlantic regions, containing the majority of globally observed ship tracks, both inside and outside of ECAs. Further, tracks detected in the Northeast Pacific prior to 2009 produce the largest difference in cloud droplet number concentration between track and background since 200323, optimising their satellite discernibility and utility for studying MCB efficacy. For these reasons, we choose to look at tracks around the 2006 period to create a dense dataset of tracks in varied regions globally.

The MODIS observations were selected from NASA’s Earth Data repository1. The collection was influenced by those compiled in20, with the days selected intended to contain images where ship tracks were numerous and highly visible. The 2019 dataset of20 provides the dates and times of ship track instances, and a list of longitude and latitude coordinates along the centreline of each ship track.

The dataset provided in this paper can be used to build predictive models for automating the masking of unlabelled ship tracks from other images. Without an intermediate model to convert provided labels to image masks, building such predictive models is not straightforward using labels which consist of lists of latitude/longitude coordinates, centrelines, centreline polygon regions or mean coordinates, as provided in the publicly available datasets20,21,38. Training automated predictive models which label pixels in images as ship tracks typically requires training data which corresponds to the desired output format, namely, image masks which identify exactly which pixels in an image represent ship tracks. The dataset provided here contains such image masks, and can be used directly for training Machine Learning models which automatically generate new image masks for unlabelled images. With these masks, we additionally provide key track metadata including coordinates of their emission points, angles and centrelines, intended to provide further quantitative descriptors of each track for such downstream analyses.

Summary

We develop a hand-labelled dataset conveying full ship track trajectories in the form of image masks, using data from the MODIS (Moderate Resolution Imaging Spectroradiometer) instrument. To the best of our knowledge, our dataset40 providing full image masks representing ship tracks on images from MODIS is first-of-its-kind. The dataset is comprised of 300 separate MODIS observations over a span of time between June 15, 2006 and August 20, 2007: a period stretching more than a year with a total of 2,543 ship tracks present in the data.

The images were selected to cover a range of geographical locations, primarily over the North Pacific and South Atlantic oceans, where ship tracks are frequently observed. Further, the images are not evenly distributed in time and location but are concentrated in regions and seasonal periods where ship tracks are most likely to form. Specifically, the majority of images cover areas along known shipping lanes with high traffic, such as the North Pacific, the South Atlantic, the west coast of southern Africa, and the west coast of South America, under suitable atmospheric conditions for ship track formation. Studies have shown these conditions are primarily due to seasonal changes in the abundance of very low clouds, prevalent during warmer months (e.g., May-July in the Northern Hemisphere)41.

To summarise the data generation and validation process: we started by compiling the MODIS observations tagged in the 2019 dataset of20, as those containing ship tracks. Each of these observations were then validated by eye from a human labeller, and those with no obvious ship tracks discarded. Next, the observations were visualised in their raw form, i.e. without any projections, and pixels containing ship tracks masked in full, via their distinct features. This masking was done using the data’s non-projected form, to ensure that the masks would be able to be projected into every other format necessary. Since each projection may differ slightly, the most universal format was used for the foundational masking. Once the masks were conceived, they were validated by two other labellers, with ship track pixels not determined by the other two labellers discarded. Subsequently, we post-processed the hand-labelled masks to extract additional track statistics: including automatically computed track centrelines, track heads, and corresponding orientation angles. Finally, the dataset was complied, consisting of:

  • Raw (unprojected) directory: Containing unprojected raw image data and the corresponding masks/rasters.

  • Padded directory: Containing raw padded data and padded masks/rasters.

  • Projected directory: Containing projected image data onto a flat latitude/longitude grid using the plate Carrée projection with interpolated values being either the minimum (min) or interpolated (interp), and the corresponding projected masks/rasters.

  • Swath directory: Containing image data visualised as a MODIS swath and the corresponding swath-visualised masks/rasters.

  • Metadata: A JSON file providing additional track statistics, including the computed centrelines, track head coordinates, and orientation angles.

Footnote 1

The process of the dataset creation and validation is outlined in Fig. 4, where the masking described can be situated within the larger context of the outputted image data products, provided in the finalised dataset.

Fig. 4
Fig. 4
Full size image

The full process of our proposed dataset creation from initial masking to the compilation of the various datasets we provide.

We use the datasets of20,21,38 to validate our ship track masks and labelling methodology. More details about the validation process are provided in the Technical Validation section.

Methods

The first 300 MODIS observations identified in the 2019 dataset of20 were collected from NASA’s Earth Data repository. All pixels containing ship tracks in these images were subsequently hand labelled. These masked images are presumed to be similar to the training data utilised in23,24,26,36, though the masked images from those datasets are not publicly available.

Data Acquisition

The satellite data used in this dataset was acquired from NASA’s Earth Data repository and are MODIS data collected from the Aqua satellite. The MODIS instrument is one of several instruments carried by NASA’s Aqua and Terra satellites, providing large public quantities of data indirectly measuring atmospheric (aerosol) composition, cloud properties, atmospheric dynamics, ocean surface properties, and land features. For example, the MODIS-derived MYD06_L2 cloud product42 (MYD06_L2 specifies Level 2 data derived from cloud analysis from the Aqua platform), providing information on cloud properties such as cloud cover, top temperature and optical thickness, has been used to study cloud responses to aerosol injection from ship tracks and other pollution sources36,37,43,44. In this paper, the MODIS MYD06_L2 calibrated radiances product is demonstrated to produce ship track labels/image masks. MYD06_L2 provides provides calibrated and geolocated at-aperture radiance measurements across 36 different spectral bands ranging from the visible to the infrared spectrum, and captures satellite-observed ship track presence. Such solar reflectance data are obtained in a 2000 km swath with the 36 spectral bands ranging from 0.4 μm to 14.4 μm generating 288 high quality images daily. Within each of these images ship tracks can be identified through their unique shapes and radiance signatures, which appear in images via different radiances in each spectral band. Built off raw data collected from the MODIS instrument, the L2 Cloud package is a dataset released by NASA that uses post processing algorithms devised after the Aqua satellite was launched to enhance and augment the data that MODIS collects45. It is freely available for download under the MYD06-L2 directory in the Earth Data repository45.

From this package, the Atmospherically Corrected Reflectance (Atm-Corr-Refl) data-field, which is known for high observability of ship tracks23,25, was used for hand-labelling the ship tracks. Atm-Corr-Refl was calculated by NASA prior to its upload to the repository by post-processing of the raw MODIS data at the instrument’s 1 km resolution. The differences in cloud droplet sizes within the ship tracks when compared to surrounding clouds result in different reflectivites and thus distinct patterns in images produced using Atm-Corr-Refl, making it an attractive data product for identifying ship tracks. This distinctiveness is one reason for its use in previous datasets and its selection for this study.

Data Labelling

The image labels are masks of the Atm-Corr-Refl images. The images produced from Atm-Corr-Refl data were saved and manually labelled using GIMP46, an open-source photo editing software. Specifically, the Free Select Tool was used to trace, extract, and mask ship tracks from the surrounding image. Figures 1 and 5(a) show the raw Atmospherically Corrected Reflectance for a single MODIS observation.

Fig. 5
Fig. 5
Full size image

Raw observation (a) and masked ship track images (b) of the Atm-Corr-Refl data taken July 15 2007 at 2:20 am in the North Pacific ocean, North East of Japan. Tracks visible in the raw image are then outlined and masked in (b) with the inter-track pixels preserved.

Masking refers to creating an accompanying image raster of the raw image, where everything that is not a ship track is blank, and only the ship tracks themselves are preserved. To produce the image masks, we identified ship track pixels by their distinct features: hollow polygonal, thin, and elongated patterns that frequently showed cloud convection47. An emission point, or the head of a ship track, was defined as the thinner, more distinct starting point of the track relative to its elongated structure. This point was typically identified by an increase in brightness or albedo, indicating a higher concentration of smaller cloud droplets from ship exhaust. Tracks that were diffused fully into existing clouds, could not be distinguished from natural features, or were faint owing to lower reflectances, were ignored to avoid high occurrences of false positive observations (i.e., pixels erroneously labelled as ship tracks). For example, linear features along large cloud boundaries that lacked a clear emission point or a quasi-linear structure were excluded (see Fig. 1). Such features can arise from cirrus clouds26 at higher altitudes or from sea-ice boundaries at high latitudes, neither of which produce the reflectivity changes characteristic of ship tracks. By focusing on clear emission points, higher albedo, and hollow polygonal patterns, we minimised false positives and achieved greater accuracy even under challenging conditions. Furthermore, our labelling strategy leveraged attributes highlighted by26, whose algorithm successfully identifies ship tracks in cirrus-dense regions by emphasising these defining properties.

As emphasised, minimising false positives was a critical consideration for training high-quality machine learning models. We further ensured that tracks with visible though partially diffuse edges, as well as smaller or newly forming tracks, were labelled to capture a range of ship track types. Additionally, because many tracks form together in high-traffic shipping lanes48,49,50 and can curve in unison while crossing bands of weaker cloudiness47, we were careful to label multiple connected or curved structures simultaneously.

A common source of potential false positive tracks were linear features on the boundaries of large clouds. These often appeared near image edges or over regions with no clouds. If they lacked a well-defined head and blended into the ambient cloud field, we did not label them as ship tracks. Finally, if a track passed through a gap in cloud cover and then reappeared, we excluded the indeterminate gap region unless it exhibited a continuous structure.

All three labellers masked each image in two rounds. In the first round, obvious ship tracks were traced in GIMP using the Free Select Tool and placed on a black background. In the second round, each labeller inspected the data for faint or small ship tracks missed initially. A pixel was accepted as a track if at least two labellers labelled it. Although about 20–30% of tracks involved disagreement (mainly due to faint scenes), this consensus method reduced false positives, especially along track edges. This approach also helped capture both prominent and subtle ship tracks, ensuring minimal personal bias in distinguishing cloud from track. Figure 6 shows the labelling steps for one MODIS scene, including the final masked products.

Fig. 6
Fig. 6
Full size image

Raw Atm-Corr-Refl data, faded labelling, and masked labelling captured by the Aqua satellite at 9:15 pm on April 2nd, 2007, off the south coast of California.

While head identification was part of our manual labeling, we also augmented the dataset with automatically generated track statistics, namely the centreline and head of each labeled ship track. To do so, we deployed a skeletonisation algorithm to extract a one-pixel-wide centreline, then performed segment extraction and automated head detection using local wind data. This method provided a more detailed analysis of individual tracks in the masked images, particularly aiding the characterisation of fragmented, curved, or overlapping tracks.

Specifically, for each binary track mask, we first computed the skeleton via morphological thinning51, thereby preserving connectivity and topology in a one-pixel-wide centreline. The skeleton was then represented as a graph, with each pixel serving as a vertex and edges connecting adjacent pixels. We identified endpoints (pixels with a single neighbor) and junctions (pixels with more than two neighbors), and subsequently segmented the skeleton into independent paths by traversing from each endpoint through intermediate pixels (those with exactly two neighbors) until reaching another endpoint.

Due to the nature of labeled ship tracks, however, the skeletonisation process introduced fragmentation between single tracks due to noise, image artifacts, or overlapping cloud features, highlighting that continuous ship tracks could be broken into multiple identified segments. Thus, we merged segments to reconstruct the full track when the gap between an endpoint of one segment and an endpoint of an adjacent segment was small, and the angle at which the segments were directed was aligned, ensuring that separate segments were joined to become continuous-like track features.

Specifically, we merged two segments if the Euclidean distance between an endpoint ki ∈ {1, 2} from one segment i (denoted \({{\bf{e}}}_{{k}_{i}}\in {{\mathbb{R}}}^{2}\)) and an endpoint kj ∈ {1, 2} from another segment j (denoted \({{\bf{e}}}_{{k}_{j}}\in {{\mathbb{R}}}^{2}\)):

$$d({{\bf{e}}}_{{k}_{i}},{{\bf{e}}}_{{k}_{j}})=\parallel {{\bf{e}}}_{{k}_{i}}-{{\bf{e}}}_{{k}_{j}}{\parallel }_{2}$$

fell below a threshold of ten pixels, and if the cosine similarity between their local angles

$$\cos \theta =\cos ({\theta }_{{k}_{i}}-{\theta }_{{k}_{j}})$$

exceeded a threshold of 0.9. Here, each endpoint ek of a segment k is associated with a scalar angle θk ∈ [0, 2π], which defines the local orientation of the centreline at that endpoint, and was computed over a short segment of pixels lying on the segment’s centreline near the endpoint.

Consequently, once the segments were merged into a continuous track, we determined the head of each fully merged track by comparing its new endpoints relative to the prevailing wind direction θW at the boundary layer, obtained from the fifth generation European Centre for Medium–Range Weather Forecasts (ECMWF) reanalysis data52. Letting \({{\bf{e}}}_{1},{{\bf{e}}}_{2}\in {{\mathbb{R}}}^{2}\) denote the two endpoints of a track, we first computed their local angles θ1, θ2 ∈ [0, 2π]. To account for the advection process that aerosols undergo after injection, the endpoint whose direction was least aligned with the wind (i.e., that with the smallest cosine similarity with θW) was therefore selected as the track head:

$${{\bf{e}}}^{\ast }=\mathop{\arg \,\min }\limits_{i\in \{1,2\}}\,\cos ({\theta }_{i}-{\theta }_{W}).$$

This approach not only identifies the track head but also provides metadata statistics (e.g., head position, centreline geometry, and alignment angles) that facilitate further plume-spreading behaviour analyses.

Last, although this algorithm proved valuable in determining the majority of track heads (see Fig. 7), a small number of cases (≈5% of labelled tracks) such as very short or curved tracks occasionally resulted in misidentified heads (e.g., a single-pixel track may have its head placed mid-segment rather than visually at an endpoint). Since our primary goal was to generate robust ship track masks, these head annotations are provided as metadata track statistics and should be interpreted with caution or refined further if high-precision head positioning is required.

Fig. 7
Fig. 7
Full size image

Example output of the skeletonisation and head detection process on a ship track from an Atm-Corr-Refl image (July 17, 2006, 2:16 pm). Left: Projected track mask. Center: Merged skeleton segments with identified endpoints (blue). Right: Calculated track heads (red).

Data Records

The dataset includes four categories of image data (in .png format) as well as additional track statistics (metadata).

The first category is full-sized; these contain both raw and masked images plotted in a large 1354 × 2030 size with no padding.

The second category is padded: Here, the images are plotted in raw format and are centred and padded to a size of 690 by 1000 pixels, which maintains the proportions of the original data without altering the image resolution or coarse-graining. This padding allows for seamless use in image manipulation functions, such as image convolution, without loss of data.

The padded category contains both the masked and faded images, which are saved in the same folder for ease of use in applications that require consistent image dimensions.

The third category is projected. These images are resampled onto a uniform and rectangular latitude/longitude grid through the Plate Carrée map projection to accurately represent the latitude and longitude coordinates. This category allows for better analysis of location-based information as the data more accurately reflects the geospatial area it was collected over.

An issue with this projected approach is that some pixel values from the MODIS instrument are missing or corrupted. To correct for these, we created two further subsets of the projected category for visualisation purposes: interpolated and minimum, corresponding to missing values either being interpolated from the image or imputed using the minimum non-missing pixel intensity found in each image. We found that ship tracks can be highlighted more strongly in the interpolated images, though for climate analysis, the minimum version is recommended as it is less likely to generate false positives or represent swaths of missing data as a cloud feature. This ensures that the dataset is comprehensive and usable for various analyses.

The final category is swath: these are the images plotted using the actual latitude and longitude of each pixel, without resampling onto a uniform grid. These images accurately reflect the satellite’s swath path over the Earth’s surface, accounting for the curvature of the Earth. In contrast, the raw images are plotted directly from the data array without considering the geolocation of each pixel, resulting in rectangular images that may appear distorted or squeezed at the edges due to the Earth’s curvature not being taken into account. A swath example is shown in Fig. 1b.

The naming of the files within each directory follows the convention YearDay-Time-Category.png, which is inherited from the NASA Earth Data repository where the original files are hosted. Here, the Category is either Raw, Masked, or Faded, corresponding to the raw reflectance images, pixels of ship tracks, and faded images, respectively. As an example, a raw reflectance image taken on Day 166 in 2006 at 18:50 is named 2006166-1850-Raw.png. For further use of this dataset, images composed by other spectral reflectance bands 0.65 μm, 0.86 μm, 1.2 μm, and 1.6 μm can be easily extracted using the available masks.

The first class of images is the Atmospherically Corrected Reflectance data from the MODIS instrument. Since this data is initially attained in packed (raw) format, the images were first parsed to obtain the actual (unpacked) values of the reflectances. This conversion process is determined by the offset and scale factors that are provided through the L2 data package and are inherent to the raw (.hdf) format of the MODIS data.

In the raw images, the data is plotted directly from the data array without considering the geolocation of each pixel. This results in rectangular images of size 1354 x 2030 that may appear distorted or squeezed at the edges due to the Earth’s curvature not being taken into account.

The second class of images is the masked data. This is the data where only the pixels comprising a ship track are non-zero, with all other pixels set to zero. We note that the ship track pixels are unaltered and retain their original pixel values. These images are labelled in each directory appropriately.

Finally, the last class of images, saved only in the padded directory, are the faded images. This dataset was produced solely for visualisation purposes, in order to show the ship track labels in the context of the full image.

The use case for this data is for understanding and display of the masking, as is shown in Fig. 6. In these images, which serve as a step between the raw images and the masked images, all non-ship track pixels are reduced to 50% transparency, while the ship track pixels remain unaltered. This fading of the background or surrounding clouds allows for better visualisation of the ship tracks themselves for presentation or validation of correct masking. Just as with the masked images, these images have no alterations on the ship track pixels; only the non-ship track pixels are altered.

Additionally, we augment the dataset with automatically generated track statistics: the centreline and head for each labelled ship track as described previously. These metadata are stored in JSON format as nested dictionaries, with each track represented by keys including “head_lat”, “head_lon”, “head_angle”, and “wind_angle”. Here, “head_lat” and “head_lon” denote the spatial coordinates of the track head, while “head_angle” is defined as the average angle in degrees (with North as 0∘) between points along the centreline and the identified head. The “wind_angle” represents the average wind direction across the centreline points, differing by approximately 180∘ from the head angle to be consistent with the wind advecting the aerosol track from its head.

A summary of the dataset components, including types, categories, contents (classes) and intended usages is shown in Table 1.

Table 1 Summary of dataset components, dimensions, and contents.

Technical Validation

To ensure the validity of our labelled ship tracks, we compared our dataset40 with pre-existing ship track databases. Ideally, validation would involve verifying each ship track against the presence of a ship, as recorded by the Automatic Identification System (AIS). However, since AIS data from before 2009 is not available, our validation relies on quantitative comparisons with the other human and machine-labelled ship track databases discussed in this paper. In an ideal classification problem, validation would comprise of identifying:

  • True positives (TP): Correctly identified ship tracks.

  • False positives (FP): Incorrectly identified ship tracks.

  • False negatives (FN): Missed true ship tracks.

Defining a true negative in this study is not meaningful since all datasets only present labelled (and not unlabelled) ship tracks. Thus, due to the nature of the dataset, only false positives, true positives, and false negatives can be analysed.

The technical validation for this dataset is split into three parts. The first two parts utilise the hand-labelled 2019 dataset of20, as the source of the direct comparison with existing data. Due to large amounts of apparent false negative tracks, i.e. true ship tracks that were incorrectly unlabelled in the dataset of20, as demonstrated via comparisons between the 2022 datasets21,38 and seen in Fig. 1, the last two validation pieces subsequently invoke comparisons with the coordinate21 and polygon38 datasets. We opt not to show comparisons between our dataset40 and the machine labelled dataset of22, since it is directly derived from the raw 2022b hand-labelled dataset of38, which contains the track labels over the time period we study.

The images used to produce our dataset40 were selected using the corresponding dates and times of ship track labels present in the 2019 dataset20. Thus, the images used to produce our dataset were all presumed to have ship tracks present before they were hand-labelled/masked. However, among these images, a number were removed, as they had no discernible tracks. The validation was conducted after the masking was completed, by comparing the latitude/longitude coordinates along the track centrelines, as listed in20, with those masked and pre-filtered by the our dataset labels. This procedure was performed to validate that the correct features within each image were masked. To do so, the swath data was projected onto a latitude/longitude grid at 0.001 degree resolution, equivalent to a pixel at approximately 0.01 km spacing. Ship track pixels were then matched with sinusoidal projected coordinates identified by20. In doing so, we find that our labelled tracks recover 100% of the previously identified tracks in20. Further we find that the ship track pixels in our dataset contain 100% of the centreline coordinates of ship tracks in20. Although this analysis confirms that our labels are in agreement with those found previously, we found multiple discrepancies in that for several images analysed in the 2019 dataset20, both our labels and the labels of21 identify multiple ship tracks, while the labellers of20,38 only identify a subset of these.

A qualitative visual assessment of the similarity between unlabelled linear features and labelled ship tracks was thus preliminarily conducted. Figure 8 shows an example of three ship tracks that we40 and the dataset of21 label, in which only two are labelled in the datasets of20 and38.

Fig. 8
Fig. 8
Full size image

Raw Atm-Corr-Refl data at 6:45 pm on June 15th 2006 off the coast of Chile, plotted as a swath with all track labels from all existing datasets (a). Projected images, zoomed in on ship tracks. Tracks shown are projected and unlabelled (b) and with the validated ship tracks (c). Origin points of tracks tagged in20,38 are marked in red, while the unlabelled track (labelled both by the proposed40 and21 datasets) is marked in orange.

To better understand the accuracy of our dataset and quantify the apparent false negatives in20, we last utilise the datasets21,38 that consist of a central coordinate21, or a polygon region around the centreline38 of all labelled ship tracks in the time period we study. With each dataset (taken to be “ground-truth”), we compare the “ground-truth” labels and their locations with respect to our ship track pixel masks. Specifically, we use a buffer of 5 pixels (or 0.05 km) equivalent to an average of 1 standard deviation of a track’s half-width in line with24, to account for slight differences in labelling and determined locations, surrounding each of our track’s labelled pixels. In doing so, we study if at least one track coordinate (from “ground-truth” datasets20,21), or centreline polygon area coordinate (from “ground-truth” dataset38) lies in in this buffer. The number of “ground truth” tracks that do not lie in this zone, and the number of “ground-truth” tracks not belonging to one of these areas are also measured.

For each track, a true positive (denoted TP), i.e. a ship track labelled by the “ground-truth” dataset that is also labelled by us, is determined for every label from the “ground-truth” dataset (coordinates or centreline polygon) that can be matched within a track we label. A false positive (denoted FP) is determined if a track we label as a ship track, cannot be matched to a label from the “ground truth” dataset. A false negative (denoted FN) is determined if a label presented in the “ground truth” dataset cannot be matched to a track we label. By counting the number of true positives (#TP), false positives (#FP) and false negatives (#FN) between a dataset and the “ground-truth” dataset, standard comparative metrics, namely the precision, recall and F1 scores (a metric which combines both precision and recall into a single score), defined by

$$\,{\rm{Precision}}\,=\frac{\#TP}{\#TP+\#FP},\quad \,{\rm{Recall}}\,=\frac{\#TP}{\#TP+\#FN},\quad {F}_{1}=2\cdot \frac{{\rm{Precision}}\cdot {\rm{Recall}}}{{\rm{Precision}}+{\rm{Recall}}},$$

can be calculated.

In this context, the precision metric measures the proportion of true ship tracks among all tracks labelled as ship tracks. Recall, on the other hand, assesses labelling ability to find all true tracks in a given image. High recall measures how well labels capture true ship tracks, without leaving too many behind. The F1 score balances both the precision and recall to provide an overall performance measure. It indicates how well the labelling procedure correctly labels true ship tracks and ignores false tracks, thus providing a balanced view of the overall performance of a labelled dataset. Precision, recall, and the F1 score are all range between 0 and 1 (inclusive), with higher scores closer to 1 indicating better performance. We calculate these metrics between our proposed dataset and the three existing datasets20,21,38, and also between all pairs of these existing datasets. These metrics are shown in Table 2.

Table 2 Table showing the precision, recall and F1 scores between the proposed dataset and three existing ship track datasets.

The results of this analysis show a high level of variation between our dataset40 and the existing datasets of20,21,38 but also considerable disagreement between the presently existing datasets. In addition to providing the F1 score for our dataset40 against the three “ground-truth” datasets, Table 2 shows the F1 scores of the validation datasets against each other. It is shown here that the dataset providing the most agreement with our dataset is the dataset of21, with an F1 score of 58%. While the other F1 scores are lower at 34%, this is still shown to be higher than comparing the datasets of20,38 against the “ground-truth” dataset of21 at 23%. These results indicate that our dataset tends to label more tracks than20,38, but fewer than that of21, indicating a potentially higher number of false positives (i.e. many labelled false ship tracks) in the dataset of21, and a higher number false negatives (i.e. many unlabelled true ship tracks) in the datasets of20,38.

The exception to our analysis arises from the comparison between the 2019 dataset20 and the 2022b dataset38, which has the highest F1 score of 63% and similar precision and recall scores at the same value. This is not surprising, however, since both datasets are derived from the same unpublished ship track dataset, first mentioned in34, rendering the datasets20,38 non mutually exclusive. Further, while it remains unclear whether the large subset of unlabelled tracks in these datasets (against that of our data40 and21) were intentionally excluded or inadvertently missed during the original labelling process, we note that these datasets rely strictly on identifying a clear head for track confirmation34 as opposed to other identifying track features, which may additionally explain the disparities in the resulting precision and recall values.

Our overall validation process emphasises that while our dataset details tracks that are consistent with existing datasets, it also captures additional tracks missed by others. This highlights the need for careful consideration of both false positives and negatives in ship track labelling, and its potential impact in training a more robust classification algorithm for automating track identification.

Usage Notes

The images in this dataset was created and managed using Python and its PIL image module53. The additional track endpoints, heads and centrelines are provided as longitude/latitude coordinates and angles from North, that can be used in conjugation with the rasters, to identify specific track regions. This dataset primarily serves to be used as training data for supervised machine learning, or as a raster dataset for the analysis of ship track or surrounding climate behaviour in the given period. As such, the raw images and masked track images are the ones recommended to use, with the faded images only for use in visualisation or to reference where the ship tracks lay within the broader context of the image. Specific usage of our dataset40 per data directory is depicted in Table 3.

Table 3 Table showing the utility of each sub-dataset contained in our proposed full dataset for visualization, masking, validation and supervision in Machine Learning (ML) algorithms.