NASA ARSET Classification Methods for Land Cover P
The Story
Welcome to another highly technical and informative episode of the NASA Live Video Podcast: "NASA ARSET: Classification Methods for Land Cover."In this episode, we dive into the core engine of geospatial analysis—turning raw satellite imagery into meaningful, actionable maps. Mapping Earth's surface is essential for tracking urbanization, agricultural expansion, and ecosystem health, but achieving this requires robust, scientific image classification methods.
Through the framework of NASA’s Applied Remote Sensing Training (ARSET) program, we provide a comprehensive overview of the primary classification techniques used in remote sensing today. We break down the differences between pixel-based and object-based image analysis (OBIA), and contrast traditional methods like Supervised and Unsupervised classification with modern Machine Learning approaches—including Random Forest, Support Vector Machines (SVM), and Classification and Regression Trees (CART). Additionally, we discuss how to assess map accuracy to ensure your land cover datasets are reliable for decision-making.
Whether you are a GIS professional, an urban planner, an environmental scientist, or a space enthusiast curious about how computer algorithms read satellite data from orbit, this episode offers essential insights into modern data processing. Subscribe to the NASA Live Video Podcast to stay updated on the frontier of earth science, satellite imagery, and global exploration!
Speaker 1: Hello everyone, and welcome to Part one of the r SET training on visualizing land cover and land use change with NASA Satellite Imagery. My name is Justin Fain and I'm part of the r SET Ecological Conservation Team based out of the NASA AMES Research Center at the south end of San Francisco Bay. Before we get into the presentation, I'd like to take a moment to tell you about our set. Our SET is the Applied Remote Sensing Training Program. We provide free trainings on a variety of topics which are divided into thematic areas, which is what you see listed here on the right.
Speaker 1: We do our best to listen to feedback from our audience to make sure we're doing trainings on the topics you're most interested in. We cover remote sensing satellites, sensors, methods, and tools at various levels of complexity, from the fundamental's remote sensing to advanced trainings.
Speaker 1: Our SET trainings come in a few different formats. They can be online or in person, live and in stry, doctor led or asynchronous and self paste, but they're always provided free of charge. Many of our trainings are available with bilingual or multi lingual options. We use only open source software and data, and the trainings are targeted at a variety of levels of expertise. Let's take a quick look at what we'll be covering in this training and some of the things I hope I hope you will take away.
Speaker 1: These are the learning objectives for the entire two part training, of which this is part one. By the end of this training, you will be able to access NASA Earth Observation data HLS, in this case relevant to land cover and land use change mapping, convert NASA Earth observation data into distinct land covered land use classes using supervised and unsupervised machine learning classification methods in the our programming language. Recognize the role of classification methods as one part of a change monitoring strategy.
Speaker 1: Computer a chain matrix representing the change in land cover and land use between two dates, and create a map in our studio to visualize the differences in land cover and land use between two dates. We're going to cover a lot of background information today, and some of it may be seem boring or pointless, but please stick with me anyway. By the end of this first part, you're going to understand how to produce and analyze these maps. More importantly, you'll learn the principles which make these maps possible so that you can apply them to your own work.
Speaker 1: The prerequisites to get the most out of this training are the fundamentals of remote sensing course and some experience working with our and our studio. If you're comfortable with our but prefer to use a different ID, that's fine as well. Note that I'm not asking you to have downloaded any of the code or data for this training. I'm going to be jumping back and forth between slides and code later in a way that would be.
Speaker 2: Hard to follow.
Speaker 1: I've also gone ahead and prepared some of the parts of the code which take a while to run, so that we can keep the training moving along instead of silently watching progress bars. So I suggest that you follow along with presentation and then explore the code after the training is concluded. The code and data I'll be using in both sessions of this training will be available on the rset GitHub, which you can find from the training web page, along with these slides, a recording of this presentation and all other materials.
Speaker 1: If you're not familiar with cloning data from a GitHub repository, don't worry We've got you covered on that topic as well. We have a short video linked on the training web page which covers the cloning operation without too much extra detail. This is the first part of a two part training. The next session will be at the same time two days from now, on February twenty sixth. The homework for this training will open after the second session and is due two weeks later on March twelfth. You can find the homework on the training web page.
Speaker 1: We'll send out a certificate of completion to all who attend and both parts of this training and complete the homework on time. With logistics covered, we can begin Part one classification methods for land cover. As I said, my name is Justin Fain and I will be the sole trainer for both parts one and two. I'm a research scientist with NASA AMES and the Bay Area Environmental Research Institute based in Mountain View, California. We went over the objectives for the training, but here's a broad overview of what we're going to do in this section.
Speaker 1: In particular, please put any questions into the Q and A box. In this training, We've got a bunch of our set people watching the Q and A who are working to answer your questions while I'm presenting. We'll try to answer any questions you submit at the end of the training, but any questions we don't get to, either because of time or because they require us some extra time for us to fully answer will be answered in that Q and a document and uploaded to the training web page in a week or so. The questions you ask and the responses we get from you and the post training surveys help us decide on topics for future RSET trainings, so your contributions.
Speaker 2: Are greatly appreciated.
Speaker 1: Now we can get started. There are a few terms we need to define so that we can start with a shared vocabulary to discuss land cover and land use change. Land cover is a directly observable description of what's on the ground at a particular place on Earth. Land use is a derived class which is related to land cover but concerned with function. These two concepts are often grouped together and written as lc LU or less commonly in the.
Speaker 2: Reverse order LULC.
Speaker 1: Though they're related, land cover and land use aren't the same thing. Consider a forest, we can directly observe the trees, so the land cover would be forests or trees, but if the trees are a part of a national forest, the land use might be recreational. If instead those trees are being grown for timber production, then we might call the land use industrial forestry. We can track the change in land covered land use over time, which gives us the acronym lc LUC, commonly pronounced el Sae luck. The things which influence Elsie luck are often referred to as the drivers of change.
Speaker 1: We will be discussing classification methods, which you might expect to be complicated and math heavy, but the conceptual framework is really incredibly simple. Classification is the process of dividing things into groups called categories or classes. The classification models or classifiers we will explore are simply using math to do the sort of grouping tasks we do every day without even.
Speaker 2: Realizing that we're doing so.
Speaker 1: You put shoes on your feet because you classify them by their foot unction. You eat food and not stones because you classify them by their adibility. You decided to attend this training because you classified it as being valuable to you, and thanks for that.
Speaker 2: By the way.
Speaker 1: At the bottom, I have a random assortment of symbols in a set we could only broadly categorize things. From this set of things, we can explore some of the possible classifications we might want to apply. For example, this group represents plants. We may further categorize those plants based on other shared characteristics like edibility or growth habit. The uncategorized plant symbols here, which appear to represent some flowers and holly, might be sorted into a category for decorative plants. The remaining elements in our set of things make up the Knoll class.
Speaker 1: These elements aren't necessarily unrelated, but rather they're simply not captured by the categories we've chosen to focus on. I provide an example of classification in the previous slide, but that example was abstract and simplified, so it might not be immediately obvious how this applies to your real life and real work. Personally, I retain concepts way better when they're functionally embodied, which is to say, their concepts I can imagine myself using. So here, I have some questions and statements which I hope to help you imagine how you might relate these concepts outside of these training.
Speaker 1: Land cover and land use is concerned with the question what is that over there and describes the features on the landscape. The land use portion is specifically concerned with its function. Assigning a class is asking what kind of thing that is, as in which category should it be sorted into. A classifier assigns a class or kind of thing based on the property it shares with other things. Land Cover and land use change is keeping track of what class was assigned to a certain part of the earth at one time and comparing that to the same location at some other time.
Speaker 1: The drivers or driver of change are forces and conditions which caused the land cover land used class to change or in some cases, prevented it from changing. Natural areas that are protected by law from changing are a good example. So why is any of this interesting to study at all? Land cover and land use describes the way things are arranged on.
Speaker 2: The surface of the earth.
Speaker 1: These features and their arrangement changed do the and due to the influence of drivers, which may reveal information about the drivers themselves. Forest fires, urban growth erosion, landslides, and crop land expansion are just a few examples of the land cover and land use changes we see. When we talk about land cover and land use classes, we're dividing things based on shared characteristics that lets us reason about the relationships between classes, such as what makes some classes similar and others different.
Speaker 1: This also gives us the information we need to start examining the potential drivers of change. By sorting land cover into discrete classes, we can use repeated remote sensing observations to detect when there's a significant change in land cover. However, there are some important things we need to consider when we're setting up a classification for land cover and land use. Classification can only tell things apart when the spectral profiles their reflectance characteristics are sufficiently different.
Speaker 1: In the case of supervised machine learning models, the only possible classes are those appear in the training data. If you aren't careful, you can end up with a classifier that is consistent, confident, and completely wrong. If we have one model that's trained to classify trees and vegetation and another that's trained only to recognize vegetables from the market, we can get confident but obviously nonsense outputs. I'll explain what some of these models are doing internally, but this isn't going to be heavily focused on the math, because those parts are easy enough to look up.
Speaker 1: Instead, We're going to try to build an intuitive understanding of how classification works and why things happen internally. Before we get there, though, we need to circle back to discuss the drivers of land cover and land use change. There are many possible drivers of land cover and land use change, some of which I listed earlier. An exhaustive list would be impossible because of the sheer variety of things which may cause land cover and land use to change. Changes are often the result of indirect factors, which are distributed in space or time.
Speaker 1: A drought in one area might cause farmers in another area to change or intensify their production due to indirect market forces. Drivers may also be related, such that a single driver might cause multiple different effects. Establishing a new protected natural area would prevent urban expansion, but may also cause an increase in wetlands due to changes in management practices. We usually want to divide these drivers into neat categories, but the real world is much messier. Drivers of land cover and land use change are hard to attribute to a single source, as we see in the table here.
Speaker 1: The conclusions we draw about drivers from land cover land use change analysis must therefore be equally sensitive to these nuances. Let's go through a couple examples of land covered and land use change analysis in action.
Speaker 1: In this simplified map of land cover and land use for Sacramento, California, you can see a few clear signals, most notably the expansion of urban areas. If we combine these maps, which I will simulate here, we can create a map that shows where things changed and how they changed. We might even try to reason about the drivers at play, such as observing that the urban expansion is probably related to the fact that the population of Sacramento nearly doubled over the same time period. Now we can talk about classification methods, which are the ways in which we use models to determine things like landcover class from remote sensing data.
Speaker 1: Classification algorithms can be divided into two major types, supervised and unsupervised. Supervised methods require training data and can only identify the classes which are represented in that training data. Unsupervised methods don't use any training data and will classify the image into the number of classes the user specifies.
Speaker 1: When we talk about training data and the context of land cover classification, we're referring to points, lines, and polygons from which we can extract pixel values from the imagery. These values give us the profile of a representative spectral response curve, which describes the way that land cover reflects light at different wavelengths. In the image on the right, the color dots show an example of training data that I've prepared for this demonstration. Each color represents a different land cover class.
Speaker 1: Here we have two examples of spectral response curves for training data samples representing barren land at the top and forests below. We can overlay the two graphs to highlight their differences. Having multiple samples for each land cover class means that the model can build up an average representative spectral profile or curve. When the model is later given the spectral profile for a new location, it will attempt to pair it with the known land cover.
Speaker 2: It most closely matches.
Speaker 1: In this case, our hypothetical unknown sample resembles forest more closely than barren lands, so the model would assign it the LC class forest. Recall that this is a description of how training data is used in supervised classification, and we'll return to that in part two of this straining. For today, we're going to stick with unsupervised classification. We will look at K means clustering and depth, because I think it gives a good intuition about how classification works in general. We first have to project our data into feature space, which is an imagine canary space, where we place each pixel based on its values in each of the spectral bands of the image.
Speaker 1: To keep things simple, we'll imagine an image with only two bands A and B. These might be the red and near infrared bands if this were an actual image. We place a point for each pixel along the axes based on the value in each dimension. In the case of remote sensing imagery, this value is usually reflectant. Next, we select k random points. K is just some number greater than one and corresponds to the number of groups or classes we want to have in our classification. These points become the centers for each of the clusters we will be building.
Speaker 1: In this example, I've decided that K is four, so there are four centers.
Speaker 1: Then we measure the distance between each of the other points, and these centers. Whichever center is closest to any given point to time determines that points cluster. These clusters now represent four separate classes, which are grouped by the similarity of their spectral characteristics, so you might wonder which land cover types these clusters represent. Well, the problem is they don't really correspond to any specific land cover. Since we didn't provide labeled training data, the classification is only capable of clustering observations based on similarity.
Speaker 1: The k Menes classifier has no awareness of what kinds of land cover we expect in the image, and the decision to use a K of four was completely arbitrary. However, this doesn't mean that the result of Camine's classifications are completely meaningless. I showed the same image at the top of the training k Meane's clustering is what I was referring to when I mentioned being able to make a decent classifier without training data. The maps on the right are outputs of Kmine's clustering with K set to two, three, four, and five.
Speaker 1: Notice that some major features like water and vegetation are correctly classified, even though the classifier had no information about water or vegetation.
Speaker 2: We'll come back to.
Speaker 1: This image to discuss interpretation in more detail later. Before we jump into the code to generate k means clustering maps in R, I want to very quickly walk it through how I downloaded the data for this training using NASA's Earth Data Search. I provided a condensed version of the process for your reference here. So here's the Earth Data Search page where I've logged in and opened the project I'm using for this training. I found an area which I thought might be interesting and drew a rectangle with the spatial filter tool.
Speaker 1: With the rectangle tool selected, click and drag to cover the area of interest. After drawing the search rectangle, I entered HLS into the top bar and hit the big search button to submit the query. What's returned is a list of HLS products. We're specifically looking for the thirty meter daily surface reflectants. You can add all the imagery to your download request or continue to refine your search with the filters in the leftmost bar. When you're satisfied with your selection, click download all.
Speaker 1: Finally, click the download data button in the lower left corner to begin the download. Now we're going to open our studio and look at the code you need to generate the k means landcover classification map. I've been showing you throughout this presentation. As a reminder, you'll have access to the code, data, slides, and training materials via the training web page and our set GitHub. I highly recommend that you watch the demonstration and then explore the code after.
Speaker 2: The training is concluded.
Speaker 1: So with that, let's get into the code for part one on supervised classification with k Meade's clustering. We will begin by loading all of the libraries we'll be using in this training, including some that will only become relevant in part two. We should also set the random seed so that we always get the same results from the random pseudo random operations. I just use the date that the homework opens, but if you use that, all of your results should look like mine. For the sake of convenience, I've gone ahead and combined the HLS data into a single multiband raster for each year.
Speaker 1: This is what you'll find on the GitHub as well. We will now read in the imagery data for twenty seventeen and twenty twenty four using the rast function and specify which of the HLS bands correspond to the red, green, and blue wavelengths. In the case of HLS data, those are bands for three and two. Since band one is the coastal aerosol band, we can create a spatial raster collection to store both the twenty seventeen and twenty twenty four imagery data in the same object. Let's quickly inspect the spatial raster collection to see the structure of our data.
Speaker 1: The important things to notice are the collection has a length of two, which corresponds to our two years of data. Both years have the same number of rose and columns, and both years have the same number of layers or bands.
Speaker 2: We can create.
Speaker 1: RGB maps the data for both years and plot them side by side using the patchwork syntax.
Speaker 2: There we go.
Speaker 1: So the images might look strange because I've applied a histogram stretch to each raster for increased contrast in the plot, but that doesn't affect the underlying data values. It's very simple to implement K means clustering to our HLS data since we don't need to provide any training data for this unsupervised classification method. We simply need to specify the number of centers. I'll demonstrate this with the imagery from twenty seventeen for two, three, four, and five centers. These each take about a minute to finish running, so to avoid waiting, I've already processed them and saved the results.
Speaker 1: You will notice when you run this for yourself. The results of applying k means cluster to the twenty seventeen data isn't automatically shown as an image, but the result is a spat raster which we can use to generate in image later. Calling the resulting object directly gives us some additional information as well. We see here that the output has the same dimensions and projection information as the input, but now reduced to only one layer, which, in the case of k meines with five centers, has values ranging from one to five.
Speaker 1: To save some typing and make sure that our plots are consistent, I've written a plotting helper function here. We should also associate our numeric class labels with names so that we can more easily apply a categorical color palette. We're working with classes, so the integer class values are meant to represent discrete clusters. Naming them the letters A through Z means that we can avoid are trying to interpret our classes as though.
Speaker 2: They were numeric values.
Speaker 1: We only need five categories since we only went up to a K of five, but it's just as easy to make.
Speaker 2: The full set of twenty six.
Speaker 1: The plotting function here applies our value categories and makes a plot from our spatraster. It uses a standard philoesthetic and the void theme to remove all of the unnecessary plot elements. We can now create maps for each of our K means classification runs using our plotting helper function. Once again, we'll save these plots as variables so that we can put them together with the patchwork syntax. We can put them together with the patchwork syntax, which is a simple layout tool for multiple plot objects, and what we see below is the same array of K means maps I showed at the beginning of the train.
Speaker 1: Since the clusters are based on the similarity of reflectance characteristics, we can see that even in the case of K equals two, we're getting good detection of water class A,
Speaker 1: with everything that is not water being sorted into class B. At the opposite end of the spectrum, we have the K equals five case, where we might observe that water is seemingly being split into separate turbid class B and clear class E categories. As you can see, land cover classification by K means clustering is relatively simple and fast. It doesn't need any training data and only requires us to pick a suitable number of cluster centers. I want to reiterate that we don't need to provide any information about what kinds of land cover appear in our imagery.
Speaker 1: Since K means is so incredibly easy to use, it's often a good idea to try a variety of values for K and select the one which best classifies the land cover types you're interested in studying. So now that you've learned how to create and visualize a K means clustering model for LC classification, we're going to jump back to the slides and go over what we've learned today, what we're going to cover in part two, and then jump into the Q and A portion of this training.
Speaker 1: So what have we learned today? We've covered how to access HLS data with Earth Data Search, why land cover and land use classification, and land cover land use change analysis is important to our understanding of the Earth, what training data is and how it relates to supervised classification, How an unsupervised K means model classifies life and cover, and how to apply a K meins classification to remote sensing imagery. This is a lot of information for one session, so thanks for sticking with me In the next session.
Speaker 1: We're covering part two of this training, which includes a review of models for supervised classification, an explanation of random forest models, how to implement random forest for classification in our how to evaluate and interpret model results, and how to create maps to show land cover and land used change over time. A reminder about the homework. It will open following the second part of this training and is due on March twelfth. You'll find the homework on the training web page and you'll submit your answers via Google forms.
Speaker 1: You qualify for the Certificate of completion if you attend both live webinars and complete the homework assignment by the deadline.
Speaker 1: Of course, I've been your trainer, but none of this would be possible without the rest of the r SET team. You can contact me via email or direct any questions and comments to the rset Gmail address. You can also sign up for the rset newsletter to stay up to date on future trainings. Thank you very much, and with that we will break for a moment and move to the Q and A. Okay, hopefully everybody can hear me. Well, start with question one. Is it possible to use another programming language? Is it something tied to homework, can we submit in another language?
Speaker 1: Fortunately, I haven't designed the whole work to involve you running any code. I do give you some outputs of code that I ask you to interpret, but it is possible to do all of the things you'll see in this part and in part two in the coding language of your choice. Python is popular. A lot of people really like doing this sort of thing in Python. JavaScript is another one that gets used a lot because it's what's used in Google Earth Engine, which allows you to do some cloud computing, but there may be a cost associated with that.
Speaker 1: Where can I access the R code for today's presentation? I have both the ore markdown notebook document and the rendered version as an HTML document on the r set GitHub, along with all of the data that we'll be using in this part and the next one. Question three, is there any particular or special reason for using R as compared to say Python? No, Again, you can do this same sort of process in the coding language of your choice. The coding language of my choice just.
Speaker 2: Happens to be R.
Speaker 1: It's the one that I think in So when I'm trying to come up with these workflows, I normally write everything in R. And then go line by line and turn it into Python if I need to. From a practical perspective, Question four. From a practical perspective, what are the limitations of the k Meines clustering method What are alternative unsupervised classification methods? I have in the slides a list of some of the unsupervised classification methods as well as supervised classification methods that we will be going over in part two.
Speaker 1: But from a practical perspective, what are the limitations? That's a fantastic question because k meines doesn't rely on labeled training data. Because it's an unsupervised clustering approach, it is limited by the difference in the spectral profiles of the different classes that you're interested in. So it's very easy for a CA means model to do things like separating water from vegetation, which they're very spectrally different. However, in a scene, if you were interested in, say the difference between different crop types or different types of trees, K means wouldn't be a good fit for that use case because the vegetation, the crop types or trees are going to be more similar to each other than they are different from the other.
Speaker 3: Things in the scene.
Speaker 1: To question five, how sensitive is the satellite imagery and classification technology in twenty twenty six? Can we distinguish between perennial and annual crops or between two different types of crops using satellite imagery? Much like my answer to the last question, yes, possibly. We have an incredible data, very high resolution both spatially and spectrally, meaning the bands the slice of the electromagnetic spectrum are very thin, so you can get these very fine detailed differences in the spectral profiles.
Speaker 1: The problem there is usually that data is commercial, so it might be cost prohibitive, and the classification technology in twenty twenty six, we're going over some pretty easy and computationally cheap to run models, but there are fantastically complicated and very precise models using things like neural networks for classification. If there's interest, I would absolutely love to do a training on that. Question six, Please explain what supervised and unsupervised machine learning is. Supervised machine learning is using training data to inform the model as to what classes you expect to.
Speaker 2: See in the scene.
Speaker 1: Unsupervised classification, such as the clustering method I just demonstrated K means only relies on the contained within the spectral bands of your image, So k meines in a sense isn't aware of what is in the image. It's only grouping things based on how similar they are in their spectral profile, in this case using their reflectance. Will you send an email with the link for the assignment? The homework assignment, the homework, presentation, slides and recordings are all going to be on the web page for the training after part two in two days.
Speaker 1: The link is there.
Speaker 1: Can you show slowly how to download data from Earth Data Search? Also, how do I interact with the file? Do I load it up in our studio? I'm not sharing my screen. If you look at the presentation, slides and recording of this which should be posted, I do go over that, and you can go over that at your own pace. As far as how do you interact with the file? Early on in the code, which again will be made available to you, I use the RAST function to load the data into r The rast function is the primary function in the Terra package for r which allows us to interact with spatial data like the raster data RASP.
Speaker 1: If it's unclear, Rast is short for raster. Makes it very easy to remember in the Earth Data Search tool, what data set are you using for this training? Are their preferential data sets based on location function, etc. I'm using the harmonized LANDSAT sentinel data, so that is a combination of the Landsat archive and the European Space Agencies Sentinel to satellite. I chose this specifically because it has a fairly high resolution because of the sentinel data, and also a long period of record thanks to land SAT preferential data sets based on location function, et cetera.
Speaker 1: In general, you want the data set that's available.
Speaker 2: For your location.
Speaker 1: And for your time, for the time that you're interested in. In Part two, we're going to be going over how to calculate change between two dates. So you want to pick a sensor which has collected imagery for both of the dates you're interested in. If you needed to go back, say to nineteen eighty, you might be limited in what sensors are available. Unfortunately, satellites don't stay in orbit forever, so we don't have one consistent constant record. I have access to artgas pro can I accomplish the same K means on supervised classification in that tool.
Speaker 2: As I said at the top of the presentation.
Speaker 1: Our set likes to use open source software, so I haven't been using rjs pro. I would imagine that there are implementations of k meanes in art pro, but I can't give you any more specific information. How does self coding compared to the built in tools for classification programs such as QGIS or SNAP toolbox. Perhaps a little in the weeds here. Everything from r to QGIS to SNAP is doing the bulk of the calculations using the g doll c library, so they should be similar. And in fact, if you're using QGIS, if you go to the processing window, you can see the actual internal function calls that are happening to make the classification working QGIS, so that's very cool.
Speaker 1: If your area of interest only has vegetation and no water or large bear soil patches, could k meanes clustering be effective given the high level of spectral similarity overall? Would the difference thin be relatively enough to use Camines classification. Yes, you've caught onto something that's very important about Camines clustering. Because it's not given any information about the target classes. It's very sensitive to the other things in the image. If the only thing in your image.
Speaker 2: Is two types of vegetation.
Speaker 1: Let's say, even small differences will appear as separate classes. If there is anything in the background which is spectfully distinct from those two classes, that will get separated as its own cluster. So yeah, great intuition there, camines is sensitive to the rest of the image.
Speaker 2: Part two we will.
Speaker 1: Discuss some supervised classification methods, and those are far better at distinguishing small differences irrespective of the background. Question thirteen. If supervised and unsupervised classification methods allow us to map past and present land cover changes, how might combining random forest models with HLS data enable us to anticipate future ecological tipic points such as irreversible deforestation or wetland collapse. And what responsibility do scientists and policy makers bear in acting on such forecasts.
Speaker 1: That is a very heavy question. As far as the responsibility of scientists and policymakers, perhaps I can't speak to that. Supervised and unsupervised classification methods do allow us to track changes over time. We will get into random forest in Part two for land cover change, But as far as doing predictions doing forecasts, that's not something we're covering though. Random forest and other models are capable of doing that sort of forecasting, but you're going to rely heavily on domain expertise domain specific knowledge to detect things like tipping points.
Speaker 1: I would also say there has been some great work doing Markov Chain Monte Carlo for predicting future land cover change based on historical trends.
Speaker 1: Question fourteen is based on question five about the sensitivity of the satellite imagery. Are there practices that can extend the sensitivity using institute resources through algorithmic connections? As resolution of satellite imagery can be a major impediment to smaller scale features, it would get to know if these resources can be extended to these scales in some way. K means, as we've covered today, is an unsupervised model, which means that it doesn't have any awareness of anything like training data points institute data.
Speaker 1: In Part two, I have prepared some training data points using institute data that we can use to help with these predictions. You are always going to be fundamentally limited by the resolution, both in scale that is, the spatial resolution, the radiometric resolution, and the revisit time, also cloud cover and data availability, so that's always going to be a concern, But in part two you I think you'll get a little bit more satisfying information about how to use on the ground data. Getting an air with the raster code looks like the file pathways for your computer.
Speaker 1: It definitely is in the read me, I believe. I even stated that you're going to have to change that. Once you download the data. You can just change that path to wherever you download the data. If you successfully cloned everything from the gethub, it will be in that folder data rast directory, so look there. If you have trouble downloading the data using the gethub clone command, you might need to download the data directly from get hub in your browser. In the first session, people were having trouble with the get large file storage system that I used because the data is fairly large, even though I cut it down.
Speaker 1: Can you show the system that you use for searching tools. I believe that this question is referring to earth Data Search, which again is in the slides, and once the recording is up you can review that earth Data Search is a fantastic resource. Also, I have to give credit to the data archives, the dacks as they're called, which do all of the hard work of organizing the data so that we can start.
Speaker 1: Is there a way to run the GitHub r code on Google colab, Yes, quite simply. The the documents that I have provided in the GitHub are r m D. There are markdown files, so there's like a combination of some some prose, the written word and the code chunks. But you'll be able to see in there which which parts are code. Assuming you can run are on Google co lab, you should also be able to run this code. Do I have to process images before uploading to our studio?
Speaker 1: As as I mentioned, I did a bit of pre processing to trim these images down to make them a little smaller and more easily digestible in this training in general, No, you can use the rast function and the same workflow that I used with the spatial raster collection that methodology to read in any imagery and then combine the imagery into multipan data. I did some of that preprocessing ahead of time, so the data that you'll see on the gethub has had the multiple bands all combined into one file. Just makes it a little easier for you to download.
Speaker 1: Do you know if Kanine's clustering is available as an add on in qgis or RGIS. I will say again, I'm not super familiar with the RGIS tool set now, but qgis definitely has some classification plugins that I've used in the past. I cannot speak to how up to date they are. As with all open source code, it's it's a volunteer effort, so use caution, exercise caution in that way, but yes, there are there are some tools for QGIS. Question twenty, can you explain more about deciding centers in K means explain more?
Speaker 1: Yes, I guess the answer to that is I I could speak at length about deciding centers and K means. If you want to email me with a more specific question, I would be happy to talk more. But the short answer here is the centers are decided from the existing points that are acted into that abstract feature space. They're chosen at random. There are optimizations for choosing better centers.
Speaker 2: Which.
Speaker 1: If I did not leave that note in the PowerPoint, it will come up at the end of part two or in the middle of part two where we talk about the related ca nearest neighbor classification. So yeah, great question.
Speaker 1: Wouldn't Sentinel one be better? Sentinel two could have cloud cover affecting the image. I'm using the harmonized lansat Sentinel data, and as I mentioned in the Earth data part, the left sidebar has an array of filter options, and one of those options allows you to filter imagery by cloud cover percentage.
Speaker 2: So the imagery that I've.
Speaker 1: Gathered, I believe I set the upper limit on cloud cover to be something like twenty percent. So we're limited on the gates that fit that those criteria. But that's how I got around the cloud cover issue. As always in all applications, use the data that's available. If Sentinel one works better for your use case, then absolutely go ahead.
Speaker 3: With Sentinel one.
Speaker 1: Okay, I'm sorry about the silence. It took me a moment to read this next question. Given harmonized lands hat, Sentinel time series and random forest classification outputs, how does desk quantify and communicate change detection confidence so that transient spectral effects such as phonology, clouds, atmospheric scattering, sensor drift, et cetera, and classification uncertainly uncertainty are not mistaken for real land cover change. What formal protocols or decision thrustsholds prevent automated change products from being used to trigger policy enforcement or funding actions without independent ground or high reds that it's ground truth or high resolution imagry.
Speaker 1: I am personally and I do not believe it is the policy of our set in general to do any sort of recommendations on policy enforcement. As far as what I've shown today, these are sort of toy examples.
Speaker 2: I would call them.
Speaker 1: They're small, limited scope intentionally because the idea here is to allow you to better understand the how and why of land covered land use change mapping with supervised and unsupervised classification. But as to the rest of that question about protocol's decision thresholds, that is.
Speaker 2: Not my field.
Speaker 1: I'm a scientist, not a policy maker.
Speaker 2: Can I access and.
Speaker 1: Submit the homework by any other means? I don't agree with Google's terms of use. Unfortunately, we use Google forms for all of our homeworks, So sorry, sorry about that.
Speaker 2: Inconvenience.
Speaker 1: Oh yeah, question twenty four great, great question. It's one of those things that I deal with so regularly, I don't think to bring it up. What exactly does harmonize mean in the HLS data set? It is the harmonized LANDSAT Sentinel data set. So LANSAT has been a long running project. Sentinel has been around for fewer years, but the two sensor products tied together and brought into a common reference frame is what that harmonization process is called.
Speaker 2: So you get the.
Speaker 1: Benefits of the long period of record of LANSAT, and the higher resolution of Sentinel data makes.
Speaker 2: It particularly well suited to these sort of.
Speaker 1: Long term changes on the landscape. Is this kmines model the same for cities in rural areas? Since it's just based on an optical pixel, anything blue or dark could potentially be water. Wouldn't different models be better if they're different? How to do align labels? K means is best suited for individual scenes, that is, imagery taken one place at one time. You'll notice that, though the overall goals of this training are mapping change, we didn't use K means to apply to two different dates and calculate change that way.
Speaker 1: We'll get into that in part two. But K means is much better at a single scene because of the issue that you've identified there, the labels don't actually correspond to any particular land cover class, so it's very hard to do that sort of alignment. It's possible if you go back to the original imagery and figure out what different in the clustering is picking up on.
Speaker 2: That's possible to do.
Speaker 1: It's easier to do for simple differences things like water versus vegetation, but it gets far more complicated when the number of classes increases. What method would you use to validate the classified raster for a large area of interest, such as for a province or state? How can you produce metrics for error and classification? Is it possible to create boundaries for each class that will enable areas to be calculated for each class. I love questions like this because I get to say, tune into part two.
Speaker 1: In Part two, where we have the ground truth points, we have the training data that is what allows us to do this sort of large scale accuracy assessment as well as the quantification of the the degree and direction of land cover change.
Speaker 2: So Part two.
Speaker 1: Water is generally delineated successfully.
Speaker 2: This is true.
Speaker 1: How can we use satellite imagery to identify changes in soil moisture irrigation and precipitation. Does noise from vegetation affect reflectance in this case? There are methods using radar, not so much optical imagery for getting soil moisture but because soil moisture is necessarily more complicated than the surface characteristics, it's not really something that we can get at using this sort of optical imagery. I'm thinking if there is standing water, certainly you would see that that dark water characteristic as as part of the pixel.
Speaker 1: But you're generally talking about spatial scales that would make that complicated. You would need very high resolution data to be able to get any sort of like surface level soil moisture conditions, and it doesn't tell you anything.
Speaker 2: About what's below the surface.
Speaker 1: So interesting question, not really something that optical imagery is capable of. Unfortunately, be cool to have X ray vision.
Speaker 1: Question twenty seven would it makes sense to one calculate series of indices for two different dates, group each of the indices for each date into a single spat rast object, and perform the case means process in similar ways to the example, I'd be using this to monitor forest and forest by a risk. I'm sure about whether using the original bands would produce a better classification. So the way that camines operates is solely based on the distance, or rather the closeness of points in this feature space.
Speaker 1: If you calculated all of these indices in DVI, for example, is a simple band ratio, all of that variability should already be captured in the original bands. You can think of this as sort of like a form of dimensionality reduction, where including in DVII, since it is just a combination of the bands that already exist, shouldn't provide extra descriptive power, and it might at scale actually degrade the classification accuracy. But this is also for your particular application. I would definitely look into do a supervised classification instead of K means unsupervised, because it seems that you have very particular classes that you want to target in your forest use case.
Speaker 1: Can K means be used to determine the NDBI? K meanes would not be the appropriate tool for this, although if you're interested in how to calculate in DBI and other related metrics, there is an rset training that I also participated in on that subject, which you can find on the r set web page. Calculating indices in QJ, Yes,
Speaker 1: I believe question twenty nine is a repeat, so I'm going to skip that.
Speaker 2: Referred to my original answer.
Speaker 1: There, given a random forest model trained on imagery from one date to how do you quantify and mitigate temportal domain shift when applying it to imagery from another date with different phonology, sensor calibration, and land management practices. Specifically, which validation strategies, transfer learning or domain adaptation techniques and uncertainty propagation methods do you recommend to ensure detected changes reflect true land cover land use transitions rather than model Again, Part two is going to get a little into how we use the ground truth data to do accuracy assessment, and if you're super curious, all of that information is already up on the GitHub.
Speaker 1: At the end of the Part two document there is sort of an extended additional information section which goes into some of the metrics we use to assess the accuracy of the model rather than just hoping that our model is correct.
Speaker 2: Lately, wanted to use.
Speaker 1: Some sentinel data with cloud coverage made it unable to do so. Any way to deal with that, Yes and no, this is the HLS data. As I've mentioned, sentinel data in certain areas, as any any remote sensing data can be hampered by cloud coverage. In Earth data, as I mentioned, is filtering by cloud coverage percentage. Maximum cloud coverage can help The other option is to either build yourself or use a product that has been averaged over a series of collection dates, so that that you essentially remove the cloudy pixels.
Speaker 1: The issue there is if you're trying to assess changes that happen, you know, say in the matter of weeks, and then you're averaging your imagery over a matter of weeks, You're not going to be able to detect those those rapid changes in the landscape.
Speaker 2: Can do to.
Speaker 1: Install packaged tidy tarra, do you have any guidance? Not without seeing the specific error message. Generally, the fix there is to set your cran mirror that's the package archive for our set that to the cloud version. That can help, or downloading directly using the dev tools package if you're familiar with that. But yeah, without seeing the error message, I can't can't help anymore than in the general case.
Speaker 1: In remote sensing, when working with multi temporal satellite imagery to monitor deforestation across the large region, what preprocessing steps are necessary to ensure consistency between data sets, How can you quantify the accurate of change detection results, and what statistical methods might be applied to validate the findings. Multi temporal satellite imagery to monitor deforestation across Generally, what you're looking for in this case is going to be like level two or preferably level three products which have been been pre processed and standardized by the science teams for that particular sensor.
Speaker 1: A lot of that, that sort of work gets done behind the scenes. If you're asking about how you do sort of the harmonization process between say, you know, LANSAT and another sensor like Sentinel, that is a much more complicated question. That's not generally something that I would recommend anybody tries to do unless they're intimately familiar with the sensor itself. The issue primarily is that as you move from sensor to sensor or even through time, the band centers the centers of the wavelengths are different and can shift.
Speaker 1: So yeah, very complicated question that goes into the physics of light. A few years ago you gave a workshop that used the semi automatic Classification tool and qgis, Yes, great tool, it was very labor intensive. Now do you think that AI has advanced that unsupervised analysis using AI? Not saying that KAY means uses AI has gotten very efficient, rapid and accurate classifying land classes such that manually intensive analysis has become obsolete. No setting aside my personal thoughts on AI, the standard of science is reproducibility, and the thing that we should always strive for is something that we can justify.
Speaker 1: Writing the code yourself and understanding as I hope this training helps you to understand the why and the how of classification is an important part of being able to justify the outcomes. Just as a small example, we set the random seed at the beginning of our code, so I'm sure that the next time I run this code, I'm going to get exactly the same results. Currently, as AI and large language models work, now, you can't guarantee that you're going to get the same result every time, which is a problem for the scientific method.
Speaker 1: Reproducibility looks like the last question. Are there methods to determine the error rate and classification using k means or results verified only through the image result?
Speaker 1: K means is an unsupervised classification method, meaning it doesn't have training data, which means we don't actually have anything to compare it against. If you were interested in validating K means and you did have training data points, you could use those as ground truth to figure out to what degree the points that you've collected of, say class grass the grass class are grouped together in the same cluster in K means. Hopefully that makes sense. K means isn't aware of what's in the image. It's only grouping things based on similarity.
Speaker 1: But you could get an error rate assessment by looking at how tightly correlated the predicted clustered class is with the actual ground truth class. I hope that makes sense. I'm much more of a visual person. If you could see me, I'm swinging my hands around, which I hope would help make all of that make sense. Sorry, I'm seeing in the chat. Jeremy Hoffman. The specific error you're getting there is that the file cannot be opened.
Speaker 1: You definitely have get all installed selecting method.
Speaker 3: Uh.
Speaker 1: Yeah, it looks like you should check the filepath that That would be my first guess that filepath that's at the end of the error message there. If your file is not at that location. Uh, this is are telling you that it it can't find the file. H The other thing to check there would be your building. Git all doesn't seem doesn't seem to be able to recognize the TIFF format. If you have qgis or something like that, try opening it there. Git all should always install with the ability to open tiff file.
Speaker 1: So that's that's a strange one. And it looks like there are other solutions.
Speaker 3: In the in the chat.
Speaker 1: M, so thank you everybody. Looks like there aren't any more questions coming in again.
Speaker 2: The Q and A document is.
Speaker 1: Going to be posted within forty eight hours after I've had some time to go through and answer all of the rest of the questions that I need to answer in more detail. Thank you all again, and I hope to see you all in Part two, where we're going to cover a lot of the cool stuff. Not that this wasn't cool, but I'm really excited for Part two
Podbean