Introduction to a Machine Learning Model to Estimate Water Quality Parameters Part 2
The Story
Welcome to Part 2 of our specialized series on environmental data science: "Introduction to a Machine Learning Model to Estimate Water Quality Parameters Part 2."In this episode of the NASA Live Video Podcast, we move beyond the foundational concepts and dive deep into the practical deployment, training, and validation of machine learning algorithms for aquatic monitoring. Accurately assessing water quality across vast geographical scales requires a powerful combination of remote sensing data and predictive intelligence to turn raw satellite observations into real-time environmental insights.
We break down the technical workflows involved in modeling critical water quality parameters—such as Chlorophyll-a concentrations, Total Suspended Solids (TSS), Colored Dissolved Organic Matter (CDOM), and sea/lake surface temperatures. We discuss how regression and classification models utilize satellite-derived surface reflectance data, how to address overfitting, and the importance of ground-truth validation using in-situ data networks to ensure your model's accuracy.
Whether you are a data scientist, a hydrologist, an environmental engineer, or a space enthusiast curious about how artificial intelligence intersects with Earth observation, this second installment offers vital, technical insights. Subscribe to the NASA Live Video Podcast to catch up on Part 1 and stay connected with the absolute frontier of space exploration, remote sensing, and cutting-edge earth science!
Speaker 1: Hello everyone, Welcome back to this art set training on monitoring water quality in lakes and coastal regions using Stream. Today it's part two on introduction to a machine learning model to estimate water quality parameters based on satellite observations. We have a guest speaker today, doctor Ryan O'she and he will be talking about the machine learning model. Last week in part one, we saw background overview and demonstration of stream webtool and API, where Stream is satellite based TWOL for rapid evaluation of aquatic environment.
Speaker 1: We had a guest speaker mister William Wainwright. He showed us that Stream uses Landset eight and nine and Sentinel two A, B and C to obtain water quality parameters which include chlorophylea concentration, total suspended solids and seeky disk depths. The area covered our coastales and inland lakes in the US. The data are available at twenty to thirty meters special resolution, and we saw that stream API allows search and download of multiple images for an area of interest. We saw example of selection of water quality parameters using stream map and make time series of water quality parameters using qgis example shown here.
Speaker 1: We saw Pyramid lake and this map shows chlorophil a concentration. High resolution data that shows pattern and how chlorophyll a is distributed within the lake and in different parts that can be seen here. We downloaded multiple images of chlorophil a concentration as well as TSS into supeag bay and made time series using QGIS. Today now we will focus on introduction to machine learning model to estimate water quality parameters based on satellite observations. This is the model that's used to derive what quality parameters in stream.
Speaker 1: Just a note here there is a homework posted today on our website and it will be due on tenth of March and the certificate of completion will be ordered to those who attend all live sessions and complete the homework assignment before the given due date. So we'll start with today's session. Our objectives today are that by the end of part two, you'll be able to become familiar with a mixture density network or MDN based model used for deriving water quality parameters from satellite observations in stream.
Speaker 1: Recognize how to select and process satellite images for deriving water quality parameters in a water body of interest. Recognize how to access the md and model code for deriving water quality parameters in the water body of interest. Here's the outline for today. We will start with a very brief review of preprocessing satellite images for inputs to the MDN based model. Then we will have an overview of the md AND based model which is used in stream, and we'll have a demonstration of access and application of the md AND model code for deriving water quality parameters from satellite data.
Speaker 1: Our guestpicker dotor Osha will use examples to demonstrate how to use MDN code. He will use Sentinel three OGI images and Paysoci images to show water quality parameters er from this model in San Francisco Bay and Lake. Erie will have a brief review of preprocessing of satellite images. Now, satellite sensors measure top of atmosphere or TA radiance, and these radiances result from a combination of surface and atmospheric conditions, including the effects of clouds and aerosolt particles as seen in this figure.
Speaker 1: Toa radiances as a function of wavength are shown here and contribution is pretty large from atmosphere and also from surface, including this green curve that is contribution from water surface. So this water living reflectance that depends on BAX scattering and absorption of radiation due to materials in water that is optically active. As shown here that the water living reflectancies. They depend on what is going on in the water and these are the reflectancies they are used to estimate water quality parameters.
Speaker 1: To get water living remote sensing reflectancies or rrs at different wavelengths, we have to start with satellite observations of TOA radiances and apply atmospheric correction schemes to get rrs's. Now Here, there are a number of schemes or atmospheric correction algorithms listed. NASA has a couple NASA Ocean Biology Processing Group algorithm and NASA AKMA Verse algorithm which is used for stream. There is echo, light and polymer. These are some of the schemes generally used by the community to do atmospheric correction and get remote sensing reflectancies.
Speaker 1: So we would start with satellite images which is level one top of atmosphere reflectance and then find level two remote sensing reflectancies. Then they're used to estimate water quality parameters. So to summarize to get remote sensing reflectancies, will select a satellite and sensor and start with level one satellite data, then use appropriate atmospheric correction algorithm to derive level to water living reflectancies and then those can be input to the md and model that we will see in a few minutes.
Speaker 1: Here we're talking about NASA c Earth Atmosphere Data Analysis System or CEDS has Ocean Color Science Software or oCSS W that can help. That has algorithm for atmospheric correction and can be used to get level two remote sensing reflectancies. There are a couple of HORSET trainings they demonstrate how to use CIDs O c ss W to do atmospheric correction and get remote sensing reflectancies. So you can refer to these for more information. And c dos A is an open source software from NASA and can be downloaded and you can find information from NASA Earth Data and search with CDs and you can get more information or you can refer to these trainings for more details.
Speaker 1: We saw this table last week all the satellites and sensors which are used for water quality monitoring, and here are the sources where you can start with level one data. You can obtain level one satellite images from these satellites and so corresponding links are provided here. These are the data portals where you can download images from these satellites and sensors. With that, we'll start with today's session on overview and demonstration of MDN model for stream Just a note here, please put your questions in the questions box and we will letice them at the end of the webinar, and feel free to enter your questions as we go.
Speaker 1: We will try to get to all the questions during the Q and A session. After the webinar, the remainder of the questions will be answered in the Q and A document which will be posted on the training website about a week after the training. With that, I'll introduce our speaker for today, doctor Ryan O'she. Doctor Ryan O'she is a senior research scientist at NASA's Goddart Space Flight Center. Ryan received his PhD in Mechanical and Oceanographic engineering from the Massachusetts Institute of Technology and Woodshol Oceanographic Institution joint program in twenty twenty one while he was under a National Defense Science and Engineering Graduate fellowship.
Speaker 1: Doctor O'she's thesis, titled Computational Approaches for Submitter Ocean Color Remote Sensing, focused on the theoretical limitations, practical considerations, and potential future avenues for deploying lightweight hyperspectral cameras for ocean color remote sensing from low altitude platforms. Currently, Ryan is working on defining freshwater and ocean water products at NASA Garden. So we invite doctor O'sha to present about the md AND model used in Stream. Ryan.
Speaker 2: Everyone, my name is Ryan O'Shea. I'm a member of the Freshwater Sensing program at Science Systems and Applications, Inc. This program consists of myself, A Rin Sarenathen will Waynewright, and Brandon Smith. It was formerly headed by Nima Polophon, who's moved on to an SHQ. Today, I'll be presenting on monitoring water quality and lakes and coastal regions using Stream. Part two focuses on introduction to a machine learning model to ask me water quality parameters based on satellite observations.
Speaker 2: On the slide, I'm going to briefly take you through freshwater and littoral imaging from space, showing the mixture density network supported sensors in red or just supported by the MDN in blue, they're supported by stream and ingreen. They're coming student to stream. In two thousand, or actually the end of nineteen ninety nine, Modus was launched. Modus as a multi spectral instrument which can capture a kilometer scale resolution with a near one to two day temporal resolution. This satellite censors lasted for the last twenty six years, so it can produce time series and climatology of given regions.
Speaker 2: This is supported by the MDN VERS was also launched and I believe twenty twelve. It achieves about seven hundred and fifty meters spatial resolution and again one to two day spatial revisit. It has very similar capabilities as Modus, and we have models that support both of these and show generally pretty harmonized retrievals from both. We also support lance at eight and nine. These are the mos that you saw as part of Part one that are included on stream. These sensors have fewer spectral bands but much finer spatial resolution up to about twenty meters spatial scales that we can start to look at fine scale inland water bodies using these sensors.
Speaker 2: All supported on stream or Sentinel two A and two B with similar spectral resolution as landset and similar spatial resolution However, the costs for both of these sensors is temporal resolution, with Landset having a temporal resolution of about eight days, where Sentinel has a resolution of about five days. However, when you start to combine them together, you can improve the temporal resolution. So we already have Landset eight and Sentinel two supported on stream, but coming sooner stream will also be adding in Sentinel three A and Sentinel three B, for which we already have make sure density network models available.
Speaker 2: These have temporal resolutions of about one to two days, higher spectral resolutions which allow for improved product accuracy in these optically complex inland and coastal waters that we're sensing in, and spatial resolutions of about three hundred meters. So by combining the lands At eight, Sentinel two and Sentinel three imagery we can get higher temporal, spatial and spectral combinations.
Speaker 2: As additional sensors are launched, will also be adding these into stream. We also have mature densitly not network models available for commercial satellites like Planets super Dove, which achieves meter scale spatial resolution. We also support hyper spectral instruments, so we develop models for Hiko the hyperspectual imagery for the coastal ocean as well as PRISMA and EMIT. We're currently funded by EMMA as well as the PACE Science and Applications Team to further improve our products for hyperspectral imagery.
Speaker 2: The PACE model that we have available supports a wide range of products leveraging the full hyperspectral data, so not only does it support the standard color fol, TSS and c DOWN products, but also inherent optical properties like absorption due to phytoplankton, and PACE allows for one to two day temporal resolution and about a kilometer spatial resolution. We're also funded as part of the Lancet Science team to make models in advance of Landset Next, which should be launched around twenty thirty one.
Speaker 2: So this large set of sensors that we have developed models for are radiometrically capable for sensing in aquatic regions. The combined observations can get us towards daily observations if we're able to leverage estimates from all of these sensors combined. But they have different spatial and spectual cap abilities which allow us to sense in a wide range of spatial resolution region so we can sense an inland as well as coastal waters. The trick, however, is combining these estimates from these different sensors to get a better understanding of global inland and coastal waters.
Speaker 2: To do that, we need a robust workflow for producing seamless products over fresh and coastal estuaries. To do that, we plan to use machine learning model, and to train our machine learning model, we need a representative data set, so we've assembled a large data set of a variety of different institute products, including chlorophyll, a total suspended sediment, as well as color dissolved organic matter. This data set builds off of Moretz Lehman's Gloria database, which is cited in the bottom right.
Speaker 2: It has At this point, we have about greater than ten thousand globally distributed in situation measurements. These include those biogeocomical parameters I was talking about before, as well as inherent optical properties from inland and coastal waters. As you can see from the map, about twenty percent about eighty percent of this data set comes from North America as well as Europe, and twenty percent from the rest of the world. Shown here, we have histograms for chlorophyla, PC, TSS, and SDOM from that globally distributed insitue data set, and these histograms show the wide range of different optical conditions we have represented in the data.
Speaker 2: So if we just look at the chlorophyl a histogram in the top left, we have about six thousand measurements and they span about four orders of magnitude, so it represents anywhere from aligotrophic to utrophic waters quite well. We also have about a thousand measurements of PA see in the top right. PC stands for phychocyanin, which is another pigment, but this pigment specific to cyanobacteria biomass potentially toxic algae. We only have about a thousand measurements of those, which can be a little bit limiting during training.
Speaker 2: We also have five thousand measurements of toll suspended solids shown on the bottom left, as well as colored dissolved organic matter shown on the bottom right, proxies for nutrient and carbon availability. These again represent a wide range of different values, showing a wide range of different combinations of these different biogeochemical parameters representative in our data set. Okay, so we have this large globally distributed data set. We know we want to use a machine learning model to be able to represent and retrieve estimates in globally distributed inland and coastal waters, so we have to choose a machine learning algorithm.
Speaker 2: Well, the algorithm that our group chose was termed mixture density networks, and I'll quickly go over that algorithm first. As input, we take the remote sensing reflectance or the ocean color signal, basically just the color of the water at the wavelengths for any specific sensor. Here it's shown for either Hiko or Prisma, so it's hyperspectral. This ocean color signal then goes through band ratio and multi spectral algorithms that typically work well for say, chlorophyl estimation, and these are then input into the standard weights of a neural network after going through a normalization phase where they're scale between negative one and one.
Speaker 2: What separates mixture density networks from typical machine learning algorithms is the output layer. This output layer is a set of Gaussians which represent the probability distribution function for any given output product, for example, chlorophyl a.
Speaker 2: These Gaussians have a mean standard deviation and probability associated with them, and when we use a combination function, we can select the most likely value of chlorophyll for the input ocean color signal instead of the average value, which is what typical machine learning algorithms do. We can also simultaneously estimate all of the different products that we're interested in, including chlorophyla if I assign an tss, and seedam, as well as the absorbing IOPs like absorption due to phytoplankton, seed on on and non algo particles.
Speaker 2: And again this is for those hyper spectral instruments, so product availability depends on the spectral availability. So I'm visually just going to try to show you why mixture density networks are well suited to solve the non unique inverse problem. On the previous page, I showed that we have this sort of typical network structure. The mixture density networks differ because they output a set of gaussians, and these Gaussians can be used to predict let's say the corfy a concentration, when there could be a wide range of different coorfol a concentrations for any input water color spectrum.
Speaker 2: So let's say we take a water color spectrum like the one shown in this graphic, and we run it through our mixture density network model. What does the output of Gaussians actually look like? So we take this water Coli spectrum in we have let's say a measured in situe corfyle a concentration shown as this black dashed line. We have a prediction from the Ocean color three band or algorithm shown as this dashed blue line. And we have a prediction from a standard machine learning algorithm shown as the dashed red line.
Speaker 2: Again, these standard machine learning algorithms typically learned to estimate the average of all possible values of coorfal concentration instead of the most likely value. We can then overlay the Gaussians as shown here that our output from the mixture density network. Each of these Gaussians again has a mean, standard deviation and probability associated with it, So we can select the Gaussian with the highest probability, take its median value, and estimate that and for this particular highly eutrophic region, that ends up having a better estimate that matches more closely with the institute chlorophyl a concentration represented by that dash black line than either the standard machine learning algorithm or the ocean color three algorithm.
Speaker 2: So, in summary, mixture density networks allow us to better solve this non unique inverse problem where there could be many different combinations of chlorophyl a, PC, TSS, and STAM concentrations for the same ocean color signal. Okay, so we have this really large data set and we've picked out a machine learning algorithm that we think is going to work really well for this reason. We can then do a fifty to fifty split from our insitu data, use half of the data for training and half of the data for validation, and then compare the results between our mixture density network and operational algorithms.
Speaker 2: And that's what's shown here. So on the y axis, we have the estimated chlorophyl a concentration from our algorithm applied to in situ remote sensing respect reflectance data that's been resampled with the sensor spectral response function in this case MSI, and we compare that to estimates from standard operational algorithms like the Gillerson two band algorithm or the blend algorithm. As you can see from this graphic, mixture density networks are able to achieve lower error, lower bias, higher slope, and lower root means square log difference over a wide range of different chlorophyl values.
Speaker 2: But this is in the idealized scenario where we're estimating based on institute data, where the MP has seen half that data, the actual accuracy when we apply it to satellite instruments is reduced. So shown here we have MSI matchups, so their same day matchups where MSI the sentinel to satellite or the imager on the Sentinel two satellite overpassed while the insitue chlorophyla concentration was being measured, and on the y axis you can see the MDN retrieve chlorpyla from MSI. We have very similar aerror metrics, where MDSA is the error that we saw previously.
Speaker 2: Previously, it was about twenty eight percent when we were looking on institute data, but once we start looking on satellite retrieve data, when we had to remove that atmospheric signal, the accuracy drops down to about ninety four percent. So that means that the residuals from atmospheric correction substantially reduced the accuracy. However, what we can see from this plot is that the MDN matchups still well represent really wide ynamic range from zero point one to about one hundred milligrams per meter cubed of chlorophyla, So this means it represents anywhere from very illigotrophic waters to very eutrophic waters, so do a good job capturing the full bloom life cycle.
Speaker 2: Shown on this page, we have MDM derived products for Lake Erie and we can see that they're spatially consistent. So the products you see in the bottom left are actually from OLCIE instead of MSI, so OLGA has more bands. We expected to do a little bit better in terms of product accuracy as well as atmosphere correction for these regions, and we're using it here just to look at the spatial consistency of our mixture density networks when applied to atmospherically corrected data. In the top left you can see the chlorophyl a concentration.
Speaker 2: Next to it the toll suspended sediments, and the bottom left absorption due to seedam and the bomb right like a cyanin, which again is a proxy for potentially toxic cyano bacteria, specially in the Let's see if we look at the chlorophyl a map in the top left, you can see the Detroit River plum which has really low concentrations of chlorophyl a. In the western section of the image is mau Mi Bay. You can see there's higher concentrations there. In the southernmost section you can see really high concentrations in Sandusky Bay.
Speaker 2: And specially this lines up with what we've seen historically from this region. We can also look at matchups with in situe measurements that were taken, so I won't cover that here, and in general they line up quite well. Not only can we produce products with our algorithms, we can actually also produce uncertainties. And these uncertainties demonstrate the mall's confidence in its own estimates. So spatially they're very consistent, they don't have dramatic artifacts, and they line up with what we've seen previously.
Speaker 2: But what happens when we take a temporal time series of these products? So I did that leveraging again ULCI, which has this fine spatial, spectral and temporal resolution combination, and shown here you can see a time series of a harmful elgal bloom which starts in Malmi Bay and interacts with the Detroit River plume. So right now it's march as it turns into April, May and June. This bloom starts to develop here it's interacting with the Detroit River plume, and then we can see it sort of reset again.
Speaker 2: So temporally we can see strong consistency in our products and actually see some hydrological interactions in these regions. So I was focusing on chlorophyl a, but we can also look at phyic a sign in which again is this marker for potentially toxic cyanobacteria, And you can see really high PC in the Sandusky Bay and Malmy Bay in the months of July and August, suggesting there might be cyanobacteria there. So now I'm going to jump to how stakeholders or users like yourselves can apply our models to preprocessed satellite imagery.
Speaker 2: The tutorials can be found in the QR code linked below. If you have any issues with it, you can also email me at the blow email. But just to give you a quick overview of what's going to be included in this demonstration, first, we will import some packages as well as functions from those packages, including specific functions from the MDN itself. We will then visualize pre corrected using Acolyte Oulchi imagery and learn how to plot those with some of our own functions. Will then retrieve the RS from this pre processed atmospherically corrected tile.
Speaker 2: These functions can work with acolyte or LtGen or even polymer. Will then generate predictions from the R, so we'll generate predictions for say tss or chlorophyl a. Then we can also generate uncertainties from these rs and visually look at them and analyze the imagery. And finally we'll do the same thing but for L two gen corrected pace imagery so that we can see some of the more advanced products. All right, now we'll jump into that tutorial. So now we'll present demonstration of our Jupiter Hub notebook called Leveraging Mature Dencity Networks to generate biogeochemical parameter and IOP maps.
Speaker 2: This is one of three that we have available. The first two cover generating products from institu measurements and generating uncertainties, and they take a little bit of a deeper dive into how these MDMs work. But today I thought I would present on how to generate product and uncertainty maps. This is often what users are most interested in. If you have any questions on these, you can email me a run or a cash directly.
Speaker 2: So the first thing we're going to do in section one is import the required packages and environment settings so we can run this first block of code. This imports the Python packages needed for the notebook. They're pretty standard packages like numpos, pathlib, zip file, and map plotlib for plotting. And then there's some packages that we've written ourselves as part of the MDM package, which is installed via the instructions at the QR code that you saw during the presentation. I won't go through those instructions here, but if you have any issues again, please email me.
Speaker 2: These specific MDN functions I'll briefly go through here, and I'll dive into their exact uses a little bit more as we get into the actual code. These functions include get sensor bands, which when you provide it with the sensor name for example, OLGI will get you the required bands for the MDN sensor. Get tile data pulls the bands as well as RRS data from a given tile path so that we can get the rs from a preprocessed net CDF download example imagery we wrote specifically for the Jupiter Hub notebook.
Speaker 2: It pulls data that we have already atmospherically corrected, so that users can run this notebook without having to do that atmosphere correction themselves. We have fined RGB image which pulls the red, green, and blue bands and outputs that which we can then use as input to the display set RGB, which displays the satellite's RGB image so that we can look at the image itself identify the aquatic regions and cloudy pixels and land pixels. We can also get tiled gear graphic info. This pulls the extent like Latin longitude for the given image.
Speaker 2: We can then run the RRS data that we pulled using gettile data through map cube MDN full, and then from those predictions that it gives from the MDN, we can overlay the RGB image and the DM products to get final maps for both uncertainties as well as the products. And this will make more sense as I jump into the examples. The next block of code that we run right here just sets the display parameters for map plotlib. I won't go into that in detail. In section two, we just download an ULGA image that again I've already atmospherically corrected.
Speaker 2: If you were to do this yourself. You'll have to provide your own atmospherically corrected data and just switch out the tile paths below.
Speaker 2: So here we're setting the sensor to OULCI and the date location just allow us to pull the direct example imagery. I've already downloaded that tile locally, so and it's displayed at this tile path right here. After running that, we can then display the RGB image using displays at RGB, and that image looks like this, which I'll go through in a second. It just takes as input the tile path and sensor. You can change the size of the figure as well as the title which we pull from the sensor and location and date that we specify previously, and it outputs the image as an RGB image which we can use for future algorithms.
Speaker 2: So this can serve as the base map for future functions that we've written. So if we just like a look at the RGB image itself, you can see that this is an image from OLHI of San Francisco Bay on March sixteenth, twenty nineteen.
Speaker 2: Let's see, this is the Pacific Ocean. You can see it's pretty blue, meaning it's probably dominated mostly by chlorophyll, probably has low concentrations of kss and sediment. As it moves into the San Francisco Bay, you can see it turns maybe greener, maybe there's higher concentrations of chlorophyll there and higher phytoplankton biomass. As you move up to the northern section of San Francisco Bay, a turns sort of a yellowy color, suggesting maybe there's higher concentrations of suspended sediment there from the Sacramento River in San Juaquan Rivers.
Speaker 2: And if we look at the southern section of the bay, you can also see these little areas which are either bright green or bright orange or bright yellow. And these are the salt flats which are separated from the bay. So there's a wide range of different optically distinct regions represented here in this image. You can also see some clouds in the bottom left. Okay, so we have the tile data, we have the RGB image. Now we can pull out the remote sensing reflectance data with this block of code using get tile data with the tile path specified and the sensor specified to OLG, it'll pull out the bands and r rs required by our MD prediction function.
Speaker 2: We can also use get tile geographic info from the tile path to get the lat, long and extent for the image, and then we can plot the RS data for this specific band, just for the first band, just to check that we plot the RS correctly. Indeed, we can see the RS is pulled out for all the aquatic regions and there aren't any RS pixels for land or clouds. Now that we have the RS data, we can move on to generating predictions from the NDN using our MDN cube or map cube MDN full function. So I'll run that while I just go through this funk.
Speaker 2: You can see it running in the bottom. You can mute this output though here we've included it just so you can see that it is running. So from the MDN package we can pull map cube MDN full as well as get ARGs. These allow us to map the model predictions and uncertainties and to find the arguments that select the correct model. So, for example, in keywords, here we set the sensor to OLCI, the products to chlorophyl TSS, and c doom, and we set the model UID, which is the name of the model weights that we actually use.
Speaker 2: You don't have to change that for now. If you have questions about setting the model UID for specific sensors. You can again message me after we run get arguments, which pulls the correct arguments for this set of keyword arguments. We can then run map cube MDN full on the arguments and rrs that we've already pulled, as well as the bands. There's a few different that we can set, like land mask, which we set to false year as we just use Acolytes land mask. There's also figure subsample, which allows us to subsample the image so we could reduce the spatial resolution.
Speaker 2: Basically, this might be useful if you want to plot an MSI image quickly and it has too high spatial resolution and it takes a little too long on your computer to run. We also have scaler mode set to invert. This doesn't need to be changed, but just basically puts the model uncertainty values into the same scale as the model predictions. Block size is set to ten thousand. This is one that you might want to change as this controls the size of the or the number of the rs that are processed at a single time, so reducing it allows you to run this on a personal laptop and increasing it allows you to run things faster if you have more RAM available, and the uncertainty mode is set to composite, so instead of providing uncertainty bounds like low and high uncertainty estimates, it just produces a single model uncertainty.
Speaker 2: All right, it looks like this ran in the time I took to explain it, and you can see the outputs for each block of ten thousand rs. Again, you can mute this if necessary, but it looks like it finished running. We can then from the model predictions as to plot graphics for chlorophyll as well as for kss and for seed on, and then we can start to visually sort of analyze these products. So these plots are created using overlay RGBMDM products, where we take the RGB image as input the model predictions which were output by the previous algorithm that ran our mature density network model.
Speaker 2: We pull all the two D information and we can pull this LCE that corresponds to chlorophyl, and how we change it is we just change from chlorophyl to TSS or to seedom depending on the product that we want to plot. For this chlorophyll product, we have this string here that gets set that we can use as the title, and we can put prediction ticks from ten to the zero to ten to the two. So these are in the log scale and we can also set the figure size and that produces these plots with the RGB image background and the MDM products overlaid.
Speaker 2: Okay, so we saw darker waters in the Pacific and that corresponds to lower concentrations of chlorophyll. That checks out. We saw these really bright salt flats again which have high concentrations of chlorophyl. That also sort of makes sense. And we see a little bit of slightly higher concentrations of chlorophyl in the northern section of San Francisco Bay and the Sacramento and San Jake and rivers generally tracks with hope we saw visually in the RGB imagery. We can also look at the total suspended solids map that we produced from that same image.
Speaker 2: If we follow this from the Pacific Ocean, you can again see lower concentrations in the Pacific, which makes sense with how we understand the region. Higher concentrations seeming like you know, coming out of the San Francisco Bay, and as you move into the northern section you see really high concentrations which came from these two rivers, which again tracks with our understanding of the region,
Speaker 2: and we can do the same thing for SEAEDOM, where again you see lower concentrations in the Pacific Ocean and higher concentrations in the northern section of this bay. Not only can we get the products, but we can also visualize uncertainties again using this overlay RGB MDM products function. But now we just also set the image uncertainties using the output model uncertainties for the corresponding slice, for example, chlorophyll in this example, so we can run that for chlorophyl TSS as well as SEDOM, and we can start to see the model's confidence in its own predictions in the same space as the chlorophyl a predictions itself, so you can see the regions where there's higher uncertainties.
Speaker 2: Often you'll see some of these higher uncertainties in your cloud or maybe some of the atmosphere correction failed and in the really high concentration chlorophyl estimates in these salt ponds, which generally tracks with what we would expect for this region. Finally, we can move on to section six, where we're generating predictions from hyper spectral pace slash oci imagery. We can again just download some example imagery for a day where I've already corrected this imagery from an image of Lake Erie.
Speaker 2: Then we can display this RGB image from the pace image that we've again already atmospherically corrected. You can see it has darker waters in the eastern section, brighter waters in the western section in Sandusky Bay, which again tracks with our understanding of this region, and again lower brightness from the Detroit River plumb. We can then do all of the same functions or use all the same functions that we used previously, like get tile data, get the tile geographic info to get the r RS and lat long plot the RRS, and then use map cube MDN full on that RRS to get model predictions uncertainty as well as the slices.
Speaker 2: So I'll run that in the background here and again you can see it running with this process bar that'll run in the backround. But we can already look at some of the visual imagery that I've produced previously. In the top here you can see mgns made for chlorophyl A, PC, TSS and c DOM.
Speaker 2: So this hyperspectral algorithm allows us to pull out additional products like fecacyanin as well as more advanced products like IOPs, which you can see in the bottom, but spatially, I can just quickly go through what we expected to see, where you can see lower concentrations in the Detroit River plume and the eastern section of Erie, and higher concentrations in Malmi and Sandusky Bays. For PC, this is in May, so you'd expect there to be lower concentrations of pyc asyanin as these harmful albablims if not started yet, and you can again see that in most of this image, though there's a little bit of higher concentrations in Malmi Bay.
Speaker 2: You can see similar spatial distributions for TSS as well as SEEDM. We can also look at the more advanced products like absorption due to phytoplankton non algo particles as well as seed DOM. Shown here is mdnesimate at aph at four hundred and fifteen animeters four pace. You can also see a different way of length like four ninety five point fifty six twenty This one often corresponds to cyanobacteria and six hundred and seventy and we can also look spectually at absorption due to seedam or we've just plotted fourteen animeters four ninety five, sixty six fifty and six seventy five, and we did the same thing for non algal particles.
Speaker 2: So you can see that with this hyperspectral data we're able to represent a wide range of additional products. Thanks for listening to my demonstration. I'll be happy to field any questions you might have during the question answer period after a NTA provide some summary information. Thanks.
Speaker 1: Thank you very much Ryan for your excellent presentation and demonstration of the code, and also for sharing the code with everyone. That brings us to the close off today's session. To summarize, we started with an overview of preprocessing top of atmosphere radiances from satellite images. We would start with level one satellite data. We saw that there are several data portals where different satellite and sensor level one data are available. Then get atmospherically corrected level two remote sensing deflectancies from the data, and these remote sensing reflectancies can be used as inputs to the MDN models.
Speaker 1: Then OSHA provided an overview and demonstration of using md AND models to derive water quality parameters. Specifically, the models are trained using a large globally distributed data set of INC two water quality measurements and colocated hyperspectral measurements from a wide range of optically complex regions and sharing many of these samples with Gloria data. The MDN approach is well suited for addressing non unique inverse problems like ocean color remote sensing due to its output layer. The models retrieve key water quality parameters depending on spectral availability, including chlorophyl a, total suspended solids, colored dissolved organic matter, and phycosyanin, along with their associated uncertain Similar models were developed to support multiple multi and hyperspectral satellite images including MOTIS Weirs, Lancet eight and nine, ONLY, Sentinel two, MSI, Sentinel three, OLG, and base OCI.
Speaker 1: This brings us to the conclusion of this webinar series. To summarize, we started with an overview and demonstration of the satellite based Tool for Rapid Evaluation of Aquatic Environments or stream Webtool and API. In part one, we saw how to map water quality parameters including chlorophylate concentration, total suspended solids, and siky disk depth. The data are available from twenty twenty four to near real time in coastal estuaries and inland lakes across the US at twenty to thirty meters special resolution and these data are derived from Lancet eight and nine and Sentinel two A, two B and to C imagery.
Speaker 1: We saw how to search and download water quality products using the stream webtool and API, and then we examine time series of water quality parameters in the region of interest using QGIS. Today in part two, we had an overview and demonstration of the aquower's algorithm. So introduction to the md and models to derive water quality parameters and access an application of md AND based model code to retree water quality parameters. Finally, stream webtool and API and the MDEN code are all open source and you can download and use the code as we saw today from doctor Roche.
Speaker 1: There is one homework assignment and it's posted on the training webpage today. Answers must be submitted by Google Forms and the homework is due on tenth March. Certificate of completion will be awarded to those who attended all live webinar sessions. Would complete the homework assignment by the deadline and then you will receive a certificate via email approximately two months after completion of the course. Once again, we thank our guest speakers, William Wainwright and Ryan O'Shea for their presentations and demonstrations describing how what quality parameters are derived from remote sensing and then making these data available through stream website.
Speaker 1: Trainer's contact information is provided here and you can always contact our set. Our set website and our set YouTube link are given here. All the training materials this training and previous trainings material can be found here. For questions, comments, or to share how you have applied our trainings to your work or studies, email NASA dot r set at gmail dot com. Join our quarterly newsletter to stay up to date on our latest trainings and the steps are given here. Just send an email with no subject line to this email address and follow the instructions sent in response.
Speaker 1: Here are some useful resources. And with that we thank you all for attending today's session and we will go to the question and answer session. Now we'll start with the question and answer session, and once again we thank our guest speakers, William Wainwright and doctor Ryan O'she. We'll start with question one, how much does special resolution affect the processing time of the MTN? And you can mute and answer the question, all.
Speaker 2: Right, thanks. I'd say most of the variability comes from the number of aquatic pixels in the given image. So for the same area, MSI, if you're using it at the ten meter spatial resolution, would have almost you know, nine hundred times more pixels than say, OCHI at three hundred meter resolution and therefore take that much longer to process. But really it's a function of the number of available aquatic pixels.
Speaker 1: Great, thank you. Next question is what is the difference between PC and TSS. Yeah.
Speaker 2: Sure, PC is defined on slide twenty in the presentation. It stands for phytocyanin fyight. A sigin is a pigment that is specific to cyanobacterial biomass, which is a potentially toxic algae, and TSS stands for total suspended solids. These can be things like sand which has been suspended in the water column.
Speaker 1: Next question, is there any reason to select fifty to fifty training and validation data set? We usually considered eighty twenty or seventy thirty.
Speaker 2: Sure. We have a long response to this question in the Q and A document, but the short answers to ensure statistically robust and representative validation across the range of observed conditions, so we adopted this fifty to fifty split. We have a pretty large data set, so we're able to split it and get you know, robust validation metrics. Additionally, this is really are idealized aerometric I guess that we're reporting here. We have multiple other ways of testing these algorithms. One of them is a regional split where we'll split the data regionally, so we'll train on everything except for one region and test it in that region and then report the aerometrics there to get an out of training data set estimate that usually has like fifty uncertainty or residuals.
Speaker 2: And then the second way, which I showed here today was we'll also try to get matchups with the satellite sensor for the probably most realistic error metric of what you'll actually see from satellite imagery.
Speaker 1: Thank you, Rian. Question for is just to be sure stream uses day radiances for satellites as input to the md and model to derive water quality parameters.
Speaker 2: Sure, so stream actually has two different mDNS. One MDN retrieves the remote sensing reflectance or ocean color from the tay radiance, and then we actually have a second MDM that retrieves what our quality parameters from the remote sensor reflectants or ocean colors. So most of what we've published so far has been on that second MBM where we try to retrieve these water quality parameters, you know, things like corphyl ATSs SECU to step from the RS. But we are rapidly trying to get out a publication to retrieve the RS from the ta radiances as well, and we expect in the next I guess six months to have that MDN available online.
Speaker 2: But we'll keep you posting.
Speaker 1: Great thank you for that reply, and if you go back and see some trainings that our set did. In those trainings see dos orcss W have been used. That's also NASA software and soon a COOWERS will be available as well. Question five. The ms I scatter plots show a classic pteroctcit pattern of residual Sorry, there are stepatistical corrections that can account for this and provide better predictions. Has anyone tried that residuals are clearly a function of the X variable.
Speaker 2: So this isn't exactly my area of expertise, but somebody else in my group wrote a response which I will read here. The MSI scatterplot should not be interpreted in the classical regression sense. In this case, the x axis does not represent an independent predictor variable upon which the y axis depends. Mal predictions are generated from the input rs, whereas the x axis represents independent in situe measurements solely for validation, we acknowledge the presence of systematic bias and range dependent error behavior in certain portions of the dynamic range or the water quality variables.
Speaker 2: These trends are consistent with the headerosydastic characteristics of the data set and are an active area of ongoing mind refinement.
Speaker 1: Thank you. My question six is as what is temporal resolution is higher than lensor sentinel. Is there a way to use it as a support tool that works whenever there are clouds or glint effects and getting a higher amount of images.
Speaker 2: Yeah, So this is again another like active area of research. How do we harmonize between these different satellite sensors and combined together their outputs. So one thing we tried to do with these mixture density networks was apply them to the available satellites that we had access to. And I showed in one of the earlier slides all of the different satellites that we do have MDMs that we that support these different satellite sensors. These MDMs are trained on similar data sets, but you know, each of these sensors does have different spectralaracteristics.
Speaker 2: So what we end up seeing is that these mDNS generally produce pretty harmonized retrievals for the different products, including chlorophyll, between the different sensors. However, long term, we're trying to attempt to further improve harmonization between these sensors, specifically Landsat, Sentinel two and Sentinel three by leveraging even more advanced machine learning models. So I guess the short answer to that question is yes, you certainly can use our models to fill in and build time series, and we're trying to improve that in the future.
Speaker 2: Thank you.
Speaker 1: Question seven, Can planet data be used for MDN or is among the coming soon?
Speaker 2: Yeah? Actually, we do have an MDN available for chlorophyla and psyche de step from Planets super Dove. I provide a link in the Q and a text. However, it's not currently supported as part of the tutorials are GitHub. We do plan to addit soon, probably in the next month or two. As we are rapidly adding in these additional sensors that we have models for but aren't quite supported by the tutorials just yet. So feel free to email me and when we push the updates to our GitHub I can, I can send an email.
Speaker 1: Back wonderful, Thank you Rian. Our next question is have the md and results been compared with the detard products from Echolite.
Speaker 2: I'm not sure exactly what algorithm Accolte uses, but I guess the short answer to this question is no, we haven't done that direct comparison. We have done comparisons to other operationally available algorithms where we download the products from websites, like the Ocean Color V four algorithm, as well as cyan four regions where we have a lot of insitue data like Lake Erie and the Chesapeake Bay. And in general we've seen the mdn's outperform these algorithms across the wide dynamic range that's represented, particularly in Lake Erie.
Speaker 2: So for bluem formation in particular, they're they're quite robust.
Speaker 1: Next question is can this code be used anywhere? Do we have to change the input which depends on the zone that we study.
Speaker 2: Short answers, Yes, it can be used anywhere. The model only relies on r rests or the remote sensing reflectants the ocean color signal as input, so it can be used in any region. There aren't any other inputs.
Speaker 1: Question ten. Currently, I'm working on estimating TSS using surface reflectance data since high TSS in institute data are not available for training. My model predicts low TSS well used during extreme events, even though I know the actual TSS consultations should be much higher than the estimates. Could you please suggest how to address this issue.
Speaker 2: Let me think about that one for a second. In this case, or the TSS being estimated from the MDN and I guess how high is high for the TSS values. The short answer I guess is that if you did have high TSS values with RS available, you could potentially either retrain the MDN or apply a transfer learning approach where you take the MDN that we have in the industry retrain some of the later layers using this institute data. Yeah, I think that would probably be the best approach.
Speaker 1: So Ran, would you say that because this was trained MDN was trained with like a large range of values, we may perform better than other ritual schemes for DSS.
Speaker 2: I would expect to perform pretty well for the region or for the concentration ranges that we have available, but not necessarily outside of that. We do have limited data at the higher ends, so that might be a limitation for accuracy.
Speaker 1: Thank you. Question eleven, is it possible to find tune MD and using incito data?
Speaker 2: Yeah, that's what it builds off that last question. The short answer is yes, it is possible, particularly using transfer learning. I think we've tried that in our lab before. I do think we have some code for it. We don't currently support that via tutorials, but if you send me an email, we can set up a video conference or something and we can discuss how to do this yourself.
Speaker 1: Great. Thank you Ryan. Question twelve, the uncertainty map seem to look very similar to the water quality maps. High not wqu estimates equal high uncertainty. Does that mean that the model can be less reliable at higher values?
Speaker 2: I guess I'd say that the bulk value of uncertainty is higher for higher values, and we can use relative uncertainty or percentages to identify outliers in terms of prediction performance.
Speaker 1: Next question is can uranium contamination in surface water be estimated.
Speaker 2: I don't have any experience with mapping uranium contamination, and yeah, I don't think we've used that previously or our seen data from that.
Speaker 1: Question four is do you have any guidance or recommended sources resources for reprocessing the RS inputs? Specifically, I'm interested in MSI and only.
Speaker 2: Yes, so the I guess atmosphere corrections that from our experience we've seen work well in neutrophic waters is ACCOLAI and more oligotrophic waters is ltgener CDs. Also coming soon, we are producing our own in house aquaverse atmosphereic correction method. We are working on paper for that and should have that out in the next sort of six months, so stay turned.
Speaker 1: Next question is the simple dataset histogram shown earlier. What is the temporal range and the special extent of the data?
Speaker 2: Sure, it's similar to that of the Gloria database. We've taken Gloria database and sort of extended it with our own situ data as well. I'm not sure exactly on the temporal range, but much of the data is from the last twenty years and spacially most of the data from that map I showed. I don't know the slide number, but I do have that graphic where it shows about eighty percent of the data is from North America and Europe and maybe twenty percent is from the rest of the globe. We are trying to build that out.
Speaker 2: If if you have data set from underrepresented regions, we'd be happy to incorporate that into future versions of the MDN model, and theoretically that would improve the mdn's retrievals in those regions. So again, yeah, feel free to send me an email message if you're interested in that.
Speaker 1: Thank you. Next question, is there a new atmospheric correction embedded in the MDN. Do you suggest just starting with L two data.
Speaker 2: Yeah, that's correct, So the mdians that are available on the GitHub tutorial right now just start from the level to data, so i'd recommend something like L two gen or acolyte. Again, we are working on developing in house atmospheric correction models. We have some that work really well, but we have to get those publications out and then we can add those to the GitHub tutorials. We're planning to do this in advance of Ocean Optics, which is in I believe September, where we'll host a tutorial there.
Speaker 1: Next question, is it possible to apply these models to earlier versions of landsets, such as landset five or seven. What would be the limitations?
Speaker 2: Let me think about that for a second. I think we and somebody from my group can correct me if I'm wrong. I believe we have applied some of these models for SECI distep retrieval from landset five and them, but I don't think we've done that for corpyl A or TSS, so I'm not sure if we've explored that fully. I'm not sure how the SNR for these sensors would support retrieval of corpolate products, but that's definitely something we'd be interested in exploring.
Speaker 1: Thank you. Next question is are there any matchup studies for MD and base predicted, TSS and HL.
Speaker 2: Not published? However, we are doing that in house. We are funded as part of the PACE Science and Applications team. We are working on one such publication where we have matchups for corfy A and fykasionin across oligotrophic waters in the US that I don't know, hopefully will be out in the next year. And then also we have an in house validation data set from the PACE validation team, and those seem to do really well when compared to OCV four in more again utrophic inlandic coastal waters. So short answer is, yes, we've done these studies, but they're not published yet.
Speaker 1: Thank you. Next question, what is the accuracy of the md AND network?
Speaker 2: You know, that's a difficult question to answer because there's a few different accuracies that you can report, and for multiple different satellite sensors, so it really depends on the satellite sensors spectral capabilities as well as it's SNR the accuracy of the atmospheric correction, and you know, if you're applying with this algorithm to institute rs without any atmospheric correction, or if you're applying it to satellite imagery, if you are interested in the accuracy for specific sensors, we have a wide range of publications that are cited in the tutorials at the bottom.
Speaker 2: You can look through those, and in general we try to publish the accuracies in each of those different scenarios for the different sensors that we support.
Speaker 1: Thank you. Next question, is there a reason why you did not compare with other available machine learning algorithms?
Speaker 2: Yeah, I guess typically we try to compare with the sort of more operationally available algorithms, and we've generally focused on those. We'd be happy to compare against the other machine learning out agorithms if there's any specific one that you're interested in with any data so we might have.
Speaker 1: Thank you, Brian. Question twenty one is the same network used for all sensors.
Speaker 2: So for the multi spectral instruments, we typically try to use the same MDN network architecture with the same number of layers and neurons. However, when we move to hyper spectral sensors like EMIT, Prisma, and pace, we had to increase the number of neurons for the larger amount of inputs as well as outputs because now we were inputting hyper spectral data and we're outputting more complex products like absorption due to phytoplankton, which is spectually variable. So we have something like fifty inputs and seventy outputs.
Speaker 2: So the short answer is for the multi special instruments, yes, it's very similar. For hyper spectral we had to increase the number of neurons in the network. We use MDS for both. Sorry, yeah, yeah.
Speaker 1: So next question is I didn't understand that one MD in redevese rrs for iop ocean color properties, and the second MD and retrieves water quality parameters. Isn't it support supposed to be the same.
Speaker 2: Yeah, So for each sensor, there's they have different spectra that are supported. So for the multi spectral instruments, typically what we retrieve ischlorophy a, TSS, and c DOM, also pechey dis deep sometimes for the models that we support. But for the hyper spectral instruments, we started adding in the IOPs as well. So not only do we retrieve chlorophy a, TSS, and on, but we also simultaneously retrieve of absorption due to phytoplankton, color, dissolved organic matter, and non alcohol particles.
Speaker 2: Yeah. So it depends on the sensor which products we currently support if they do retree of them simultaneously.
Speaker 1: Next question is which bands are the network using for chlorophylla and the other parameters.
Speaker 2: Yeah, again, it depends very much on the sensor that we're talking about. Typically we try to keep the bands in the I believe it's four hundred and fifteen nanimeter to about seven hundred and fifteen animeter range that varies a little bit. Sensor to sensor. The exact way of lengths can be found in the tutorials using the get sensored bands function for the specific sensor, but we try to use all the bands in that range typically.
Speaker 1: Next question is is there a document that lists the water bodies where the India is trained and validated.
Speaker 2: There's an individual document. However, we do have again a wide set of publications available. These are cited at the bottom of our tutorials and in each one of those publications we go through very similar training where we do the fifty to fifty split. Sometimes we do a regional split as well, and then finally we try to do a matchup analysis for each one of the sensors, and in that we shall also talk about we do also talk about the sort of spatial distribution of that data.
Speaker 1: The next question is is there can it distinguish between oil spills and habs or oceans.
Speaker 2: We don't have experience using these MDMs for distinguishing between oil spills and habs. Okay.
Speaker 1: The next question is how to NASA Ocean Color and inland water Quality algorithms manage seasonal spectral shifts caused by monsoon driven trabid t seedom variations and changing solar geometry without introducing systematic bias in long term time series.
Speaker 2: Let me think about that one for a second. Yeah, I think this is a really good question. I think I'd have to think on it a little bit before providing an answer. I'll try to do that over the next week and get back to it.
Speaker 1: Great, Thank you so much. The next question is why was the NC two chlorophyl A data from USGS, which covers the United States not used.
Speaker 2: Let's see does that data set include r rs as well as colorful A. If so, we'd be happy to add it in. I'm not entirely sure if yeah, If it does include rs, we'd be happy to add it in when we retrain these models. If it is just coreful A, probably have pulled some of that data for validation, but we can't use it for training with the MDM since they require the RS as input.
Speaker 1: Question eight, does stream recommend regional recalibration of retrieval algorithms or does it rely on globally parameterized models with deviation correction layers.
Speaker 2: We do not currently recommend regional recalibration of these algorithms. Usually, what we would like to do is just keep adding in more and more data globally. So again, if you have in siture data with RS and the parameter of interest and you're willing to share it, please send me an email and we'd be happy on future model updates to include that data, which theoretically should increase the accuracy of the model for that given region. Yeah, the model is globally trained. We don't deviate at all based on the region, so the same model gets applied just to the RS no matter what region we're in.
Speaker 1: Thank you. Next question, Can the NDN models be used for volatile and non volatile compounds as part of water quality predictions in the case of nitrates, phosphorates, am monco, nitrates and other ions, could MDN be used if yes? In general in general practice, how many incituo observations are we looking here to train md IN models?
Speaker 2: Oh? I see the MDN can only be used for well, I guess the md and from the satellites with multi spectral or hyperspectual data and visible and urinef red range can only be used for products that change the ocean color signal if you have RS with different concentrating. I don't know specifically how these different compounds change the ocean color signal or if they do interact. If you do have a data set of RS with different concentrations of these different products and they do change it, then we could try to train an MDM to do that for how many institue observations?
Speaker 2: Again, that sort of depends, but I guess we've used in practice as little as a thousand samples to train our PC models, so that represented a wide range of different flectasign and concentrations pretty well, but I don't know if we've looked at an exact minimum.
Speaker 1: Next question, can mbian model run on wetland water quality monitoring?
Speaker 2: Let me think about that for one second. If the wetlands are optically deep, meaning that we can't see the bottom reflectance, then theoretically I believe they should work just as well for retrieving CHLOROPHYLA or TSS. However, if you can see the bottom reflectance, that is going to change the ocean color signal or the signal by these atmosphere correction algorithms, And we really haven't tested the model in these regions, thank you.
Speaker 1: Our next question is should input LTO data be gritted before using the MDN or does it not matter.
Speaker 2: One second? Now I think the answer is no, it doesn't matter. It doesn't have to be gritted. Yeah, the MDN just takes as input the remote sense and reflectance in basically an array. So as long as you can get it into the format similar to that that's applied in the tutorials, you should be good.
Speaker 1: Thank you so much. It looks like there's a lot of interest and we had a lot of great questions. We will be posting our Q and A documents on the training website in a week or so for both the sessions. Okay, I think there's one more question. How was the RS obtained for gloria and other data sets used where they downscale from any satellite data to the location of the sample site? If so, can I do similarly for my institute data sets, which only have glorophyll a.
Speaker 2: Sure, so the Gloria data set and we actually use it, you know, an augmented version of that, but it has very similar properties, has hyperspectral RRS and a few different in situ products, including chlorophyla. We take that hyperspectral RS and resample it with the sensors relative spectral response function and then we train the models using that resampled RS.
Speaker 2: You could do that with your own insert your data set if you have chlorophyla with RRS. If you only have chlorophyll a, you wouldn't be able to reach OT.
Speaker 1: I'm currently working on a coastal water quality study area. This area covers a small region and I want to use Stream, but Stream does not support this area with data. What can I do?
Speaker 2: You can email us. We are trying to support users that are interested in using our products. We do have some processing limitations that we're trying to work through, but if you email us, we can try to add a specific tile in that you might be interested in.
Speaker 1: That's fantastic, Thank you so much. Are these data exsible and Google attentionIn?
Speaker 2: No, we have not included these data and Google or if engineer. We do not currently have plans to.
Speaker 1: Hello, I'm from Mexico and I would like to use this tool in my country, but I saw it's only available in the US. Can be the possibility that we can use in Latin America? Sure?
Speaker 2: Again? Yeah, I feel free to email me if there's a specific site that you're interested in, and we can try to add that to Stream. Again. We do have some processing limitations, but if it's only one side should be able to do that if you're talking about stream, if you're talking about the mDNS themselves.
Speaker 1: We do have data.
Speaker 2: From again around the globe and we've used that in training.
Speaker 2: As long as you have RS data from satellite, using the tutorials that I showed here today, you should be able to produce your own product maps for that region. We're also as part of the land sat science team, we're also adding in data from some Latin American countries, including Mexico, and trying to validate our algorithms in these regions.
Speaker 1: This is one of the last questions we can take. In highly turbid Indonesian tropical river systems such as the Cupuas and Mahakam in Calimantan, where seed dom and suspended sediments are both elevated, how does the md AND handle spectral ambiguity between chlorophy a and si dom absorption in blue bends.
Speaker 2: Sure, so, we typically treat the MDN more as a black box. We don't really go into like the layers to see how it is span chlorophyll and seed dom. But you know the way it works is it's a empirical algorithms. So we have again a large data set with a wide range of chlorophy a, c DOM TSS values represented. So if we have similar chlorophyl a and seed dom combinations, we would expect it to perform relatively well. If we don't, it might not. But then you could again, if you have in such a data you're willing to share, you could email me and we could add it in to future trainings in the model.
Speaker 1: Great, thank you. I think we're almost at the end of today's session. Thank you all for attending this training, and once again we thank our guest speakers for their valuable time and information that they're provided. Also, thanks they are the team for their coordination and helping putting this training together. And we hope to see you at our next training and look forward to get suby results from you and feedback from you. Thank you,
Podbean