NASA ARSET Introduction to the GEDI Mission and its Derived Products
The Story
Welcome to this highly fascinating and technically rich episode of the NASA Live Video Podcast: "NASA ARSET: Introduction to the GEDI Mission and its Derived Products."In this episode, we venture into the heart of Earth's forests to explore how spaceborne lasers are revolutionizing our understanding of global carbon cycles and ecosystem structures. Operating from the International Space Station (ISS), NASA's GEDI (Global Ecosystem Dynamics Investigation) mission uses high-resolution laser ranging to provide the first high-quality measurements of forest canopy height, canopy vertical structure, and surface elevation globally.
Through the framework of NASA’s Applied Remote Sensing Training (ARSET) program, we introduce the foundational concepts of the GEDI mission and its core technology: full-waveform LiDAR. We break down the key derived data products available to the scientific community, explaining how researchers utilize GEDI data to estimate aboveground biomass, map complex forest structures, and improve biodiversity assessments and climate change modeling.
Whether you are an ecologist, a climate scientist, a forestry professional, a GIS analyst, or a space enthusiast eager to discover how NASA uses lasers from orbit to weigh the world's trees, this episode offers vital, foundational insights. Subscribe to the NASA Live Video Podcast to stay at the absolute forefront of space exploration, remote sensing data, and cutting-edge earth science!
Speaker 1: Welcome to our training series Spaceborne lightar for Monitoring vegetation, structure and biomass using Jedi. My name is Savannah Cooley. I'm a researcher at NASA Ames Research Center with the Bay Area Environmental Research Institute in Mountain View, California, and I'm a trainer with the RSET Ecological Conservation Team. There are three parts in this training. In Part two, today will further introduce the Jedi mission and its data products with the hands on exercise for accessing and analyzing data from Jedi's level to b HDF five files.
Speaker 1: Part one covered full waveform lightar foundations, and Part three will cover estimating biomass change with the OBI one application. Remember that the homework opens on the last day of the training on November six, twenty five. It is due two weeks later on November twentieth. A certificate of completion will be awarded to participants who attend all live sessions and complete the homework assignment before the given due date. Let's dive into Part two of the series, introduction to the Jedi mission and its derived products.
Speaker 1: We have one guest instructor for Part two. Stephanie Humenis is a research associate with the Satellite Needs Working Group Management Office based out of the National Space Science and Technology Center with NASA Marshall's Space Flight Center and the University of Alabama and Huntsville Earth System Science Center. Additional support for this training came from NASA earth Rise, an Earth Action program at Marshall's Space Flight Center. I would like to acknowledge our fellow colleagues at NASA earth Rise who supported the development of this training.
Speaker 1: There are three learning objectives for this session. By the end of the session, participants will be able to first identify the general characteristics, strengths, and limitations of the Jedi mission and its data. Second, identify Jedi data products and their characteristics for elevation, canopy, height, vegetation structure, and biomass estimation. Third, access and visualized Jedi elevation, height, vegetation, structure, and biomass data sets for an area of interest using openly available tools.
Speaker 1: If you have questions during today's training, you can put them in the Q and A box within WebEx at any time. At the end of the training, we will address as many questions as we can, and we will post these questions and written answers to the training page within a week after the training. With that, I will pass to Stephanie.
Speaker 2: All right, and with that, let's dive into Section one, which reviews what the JEDI mission is and how its data for ecosystem structure can be used. JEDI is a joint project between the University of Maryland and NASA Goddard Spaceflight Center and is a NASA Ventures instrument housed aboard the International Space Station. The sensor produces high resolution laser ranging observations of three dimensional structures on Earth's surface. The use of near infrared wavelength and the science team's waveform processing algorithms enable JEDI to optimally capture detail in the structure and density of vegetation.
Speaker 2: The mission was launched in twenty eighteen and began collecting data from April twenty nineteen through March twenty twenty three, but from March twenty twenty three to April twenty twenty four the instrument went into hibernation, but has since returned return of the JEDI and will collect data through this decade. So the purpose of this mission is to assess the current state of Earth's forest structure. From these measurements, models can be tested to observe how forest dynamics have changed historically and predict how they may change over time.
Speaker 2: From each JEDI sample or observation, we gain a better understanding of how ecosystems are structured dynamically. These novel perspectives of ecosystem structure can better inform the status of habitat quality and other biodiversity indicators when combining the data with other ecosystem information. JEDI not only observes forests, but it can summarize topography below vegetated canopies, track surface deformation, assess continental and coastal water resources, and even predict weather impacts.
Speaker 2: JEDI can inform human to landscape interactions, such as when managing e ecosystems, modeling fires, and assessing the carbon cycle with its data. So the next fuel Slides will offer a brief overview of possible applications of JEDI data. Biomass estimations are essential for quantifying carbon storage inflexes in terrestrial ecosystems. Information on biomass can inform models of the carbon cycle by constraining estimates of carbon sources and sinks. Improved models enable more accurate assessments of how forests contribute to climate regulation.
Speaker 2: Biomass can be modeled from jedi's data since it captures vegetation height and structure. Jedi's biomass estimations at the footprint and graded scales provide high resolution and globally consistent measurements to track global carbon budgets. Jedi's height metrics and its ability to measure structural complexity make it possible to assess dynamics in vegetation strata, ages, types, in competition, etc. As shown in the leftmost image. From this more in depth understanding of vegetation, professionals can better inform their efforts in conservation, restoration, habitat modeling, and fire management.
Speaker 2: An example data flow in the rightmost figure demonstrates how local measurements are matched to landscape scale measurements like those from JEDI or other remote sensing satellite based sources. Integrating data at this scale over time enables greater spatial extent to monitor fuels or vegetation regrowth or ecosystem disturbances. JEDI data provide detail measurements of the canopy, structure and ground elevation, which enhance land surface models used in weather prediction by improving the representation of vegetation, height, density and distribution.
Speaker 2: JEDI data can help refine simulations of surface energy balance, evapotransporation, and n and even wind dynamics. Each of these are diversely influenced by different ecosystem structures. These atmospheric effects are key processes influencing local and regional atmospheric conditions. LIGHTAR in general is commonly used to monitor water volume changes, wetland water levels, river reach slopes, stages, or discharge across inland water surfaces or coastal regions. JEDI can support these same hydrological analyzes with its sub kilometer sampling of elevation, including that of water and water under highly vegetated areas such as in mangrove forests.
Speaker 2: Accurate elevation data, including those beneath vegetation, are essential for mapping topography and detecting surface deformation, such as in landslides or subsidence. Understanding elevation and slope changes, especially below forest canopies, can be crucial for disaster risk management, such as when managing and responding to flooding, fire spread, or landslides, by capturing precise ground elevations, even in densely forested areas. JEDI enhances elevation modeling across diverse land cover types, so the mission specifications could potentially impact your ability to apply its datas products for your study area.
Speaker 2: This is a reference table for recognizing systematic considerations which may impact data availability. Given that JEDI is on the International Space Station, it only samples between fifty one point six degrees north and south, leaving out the far southern and northern study areas or regions of the globe, which are actually shaded in blue gray on the map below. It is important to recognize the sampling density of the sensor itself, where each of the eight beams lie six hundred meters apart and each sample or footprint is located sixty meters from the other within beam.
Speaker 2: Temporal availability is important as well, those samples may be available Spatially available records only fall between April twenty nineteen to March twenty twenty three, and then again from April twenty twenty four until now. Where data is available, the resolution and geolocation accuracy may be a big contender in how representative that data is of your desired application. Updated versioning and improved data products such as Fusing JEDI with other full coverage data sets like optical or SAR can improve accuracies, but may change resolutions.
Speaker 2: The twenty five meter footprints are valuable for their higher resolution and customizability, while the one kilometer or more grids and fused products may provide more reliable information for landscape scale studies. Spaceborne lightar is an incredible novel technology advancing the field of light ar, which has been a common and trusted remote sensing technique used to manage and monitor land and water across available light our systems, whether free and open sources or commercial. There are several major differences and or trade offs, as shown in the figure.
Speaker 2: Various methods for measuring vegetation with light ar or field studies occur at different spatial sampling patterns has shown on the bottom, or sampling resolutions as shown in the middle, and by the size of the beams, which are part in part dictated by the light our system itself, as you can visualize by the altitude of the sensors. As with many applications, some of the most trusted data are with field plots or human light observation. While highly accurate field surveys can be costly and time consuming, involving a lot of manual input.
Speaker 2: Because of this, these campaigns tend to cover small areas and are collected inconsistently, making it difficult to pair timely relevant data with other remote sensing products. Terrestrial light OAR systems or TLS are similarly reliable since they are placed on the ground in location. TLS can be costly, costly instruments and require specialized software, computation, storage, or skills to process and interpret the results, while quicker and incredibly detailed repeat coverage can also be costly.
Speaker 2: Airborne light our systems like NASA ELVIS, which is the land vegetation and ice system based out of the NASA Jet Propulsion Laboratory, can be used to calibrate and validate spaceborn light our missions like the ISATS and JEDI, because they cover large areas while collecting high density and accurate observations. Unmanned aerial vehicles like drones are similarly flexible. Greater spatial coverage can promote increased frequency of observation and coverage when funded, yet also require specialized software, computation, storage, and skills to process and interpret large amounts of data.
Speaker 2: Space based lighter though a great technological feat introduce trade off in accuracy and resolution compared to airborne UAV, terrestrial light our and field plot data. The density of the data may vary though repake coverage over a particular landscape. Globally, it is highly advantageous. While NASA works to improve accuracy, accessibility, and ease users skills to interpret the data, the ISAT missions and JEDI remain free and open source on public repositories.
Speaker 2: Many field inventory plots or other field data plots were used to derive jedimetrics based on alemetric equations that relate what is observed and measured on the ground with what JEDI and other remote sensors capture. The field inventory and analysis plots can also be used for local calibration or validating results. NASA has committed the past few decades to advancing space based light our systems. Here are the Earth observing satellites collecting light our perspectives of Earth from ISAT to ISAT, iiO to JEDI, and the future EDGE and STV missions.
Speaker 2: Earth observing lighter capabilities have undergone multiple iterations across time, starting with full waveform observations over predominantly water and ice landscapes. Full waveform LIGHTAR was then adapted to measure vegetation with increased density and resolution. A photon counting lighter for ice, water and terrestrial area with iat IW also greatly improve the spatial resolution. Features from each of these systems have contributed to the design of future missions like EDGE who will increase sampling density with additional beams and swath mapping for improved change detection capabilities and repeat coverage.
Speaker 2: Now let's dive into section two to overview Jedi's data products, including the waveform metrics and estimations and derived products for vegetation, structure, and biomass. When JEDI samples data, which occurus milis millions of times per week, several data products can be derived. The samples are collected as received waveforms, which are geolocated to Earth's surface and presented as the Level one B product. These waveforms capture the height structure of the surface features within the waveform.
Speaker 2: Several algorithms that are developed by the Jedi science team make sense of the waveform shape and its intensity based on its geographic location, the land cover type, and more to generate each of the other data products. These interpretations result in elevation and height data sets for bare earth elevation, top canopy height, and relative heights across the vertical profile as found in the Level two A product, which are footprints, Level three and the other derived products aggregate and reformat those elevation and height metrics at different resolutions and with certain quality checks applied.
Speaker 2: Additional measured variables for vegetation structure include canopy cover, plant area index, plant area volume density, and folio height diversity in the Level two B footprint data set. The waveform structural complexity index in Level four captures the complexity of the canopy within the entire lightar waveform return. Several of these metrics are total summaries, or they are calculated across the vertical profile along the height of the surface feature. From these different types of metrics, biomass can then be modeled either the density or the total biomass with associated uncertainty metrics that are included in jedi's Level four A footprint and Level four B or other derived biomass data products.
Speaker 2: The footprint level data sets are publicly available as HDF five files, But what is an HDF five file? These files organize data sets of multiple formats into hierarchies or groups along alongside their associated metadata. The root group stores all data within the file under which additional groups are housed. The groups direct to several stored data sets of metrics and metadata for each observation. An example of the fire files oh my god, Okay, let me repeat that. An example of the file structure is shown on the right.
Speaker 2: The level one B waveform, Level two A heights, Level two B structure, Level four A biomass, and level four C waveform Structural Complexity Index are accessible as HDF five formats. To date, opening and handling the HDF five files, including converting from HDF five to say CSV, shape file or in array, requires either using coding solutions in Python R or in Google Earth Engine, and there are some non coding oriented existing tools with great visualization capabilities with some limitations to data sets available over a certain study area with some size constraints.
Speaker 2: And by the way, the Google Earth Engine Jedi FI foot prints are formatted into monthly rosters, which is a bit more convenient to use when managing the HDF five file itself. Knowing this file structure will be really important for opening, subsetting and visualizing the data themselves. Each product has a unique setup, so the data dictionaries and specific path names become particularly important when working with a file. Beam xxxx indicates that the same groups of data sets are hosted eight times, meaning that they were collected and stored separately by for each beam.
Speaker 2: For example, if you want to find the latitude of the lowest mode of the waveform, you'll need to follow the path from the root group down to the data set. So if you're working with the root group, which would be for example, JEDI Level two A, you'll have to navigate to one of the beams, say Beam one thousand, then go to the geolocation group where the latitude lowest mode group is housed and all of the data values for each observation are there. If you wanted to get lat lowest mode for all of the beams, you'd have to select and go through this process eight times.
Speaker 2: Note that programs like RCGIS or QGIS cannot read the Jedi HDF five files directly, so they will need to be opened or transformed by alternative methods, which we can talk about later in this section. So all the other data products are provided in a more familiar GEOTIP format. This includes the Level three, Level four B and all the other gridded or derived products. These, however, may have predetermined quality filters and be at lower resolution, but could be more accessible and applicable to larger scale studies.
Speaker 2: Each product will include some combination of the listed conventions for identification and definition that could be with a particular time period, version number, spatial resolution or instrument as listed on the slide. For the rest of the section, we will share Jedi products reference tables with important information for choosing a Jedi data product to work with. For more detailed information on each of these products, I encourage you to review the linked reference flyers. So the waveforms, as we detailed in Part one are the basis for all of the other Jedi metrics and estimations.
Speaker 2: The original transmitted and received waveform data, which are shown as TX and RX and not publicly available, are geolocated to Earth's surface and made available as a Level one B product. We went over how to access this in Part one. Level two A is the next product level built from those geolocated waveforms. It includes data for the detected ground elevation, top canopy height, and relative heights along the vertical profiles that are zero to one hundred from interpreting the waveform itself with the Level two A processing algorithms.
Speaker 2: The Level three gridded surface metrics pictured on the bottom here offers a gridded approach to the elevation and height metrics in TIF image format rather than as footprints within the one kilometer by one kilometer grids globally, the JEDI data footprint counts are recorded within each grid to calculate the mean canopy height specifically with R one hundred and the mean ground elevation along with the standard deviations for each The calculated elevation and height variables are produced for a set of five different time periods of the mission in order to incorporate more of the footprint level data that's been accumulated as the mission continues.
Speaker 2: Now a lot of the other gridded products and derived products follow a similar method of aggregating over different time periods to generate those data sets. Level two B is also based on the geolocated waveform as well as the derived elevation and canopy top height from Level to A, but it mostly relies on the probability of complete laser penetration through the canopy and to the ground. In order to accurately calculate the structural metrics. The vegetation metrics from level to B are all built off one another along the vertical profile of the surface feature, canopy cover and plant area index are calculated from the vertical metrics.
Speaker 2: A total value for canopy cover fraction and the total plant area index can also be calculated. From those plant area index results. The vertical profile plant area of volume density is also calculated, and from that the foliage height diversity is derived. Each of these data sets offers a unique perspective of the vegetation characteristics, though they are closely linked algorithmically. The figure shows an example subset of the waveform structural complexity index prediction from level four C over the Eastern Amazon.
Speaker 2: Level four C provides an index for the structural complexity of the waveform and is estimated with models that are specified over four plant functional types. The index values for high and low complexity are available as a total for the entire footprint or as multiple values along the vertical profile like those in Level two B. The index combines the information from the vertical and horizontal canopy layers from the waveforms in relation to available airborne line or data into a single interpretable metric.
Speaker 2: The figure on the right shows above ground biomass density over the US with cloud free data from February twenty twenty three. The biomass level data sets are available in footprint and gridded formats, offering related information at multiple resolutions or aggregated mean over a one kilometer grid. The biomass is modeled in each case with environmental conditions in mind, such as leaf status whether it's leaf on or leaf off depending on the land cover type, the plant functional type, and geographic region considered respective to the algorithm that's applied to each footprint or guid.
Speaker 2: Since JEDI is a sampling mission and not an imaging or wall to wall mapping sensor, many workflows may need to combine JEDI with other data to fully map a landscape and measure change over time. An example of this is in the derived product put out by the JEDI mission. In many of the derived products put out by the JEDI mission that help increase the quality and density of the information. For example, a simultaneous spaceborne light our mission is SAT two collects photon point cloud data over vegetation in terrestrial areas as well that can be fused with JEDI for increasing the three D sampling density and temporal coverage of the observations, and it's provided as a fuse product at lower resolutions.
Speaker 2: For select pantropical forests. Shown on the leftmost map, height and biomass are calculated with a combination of JEDI and InSAR from the TANDEMX mission by the European Space Agency. These are tropically optimized wall to wall maps of height and biomass information. The figure on the bottom left presents canopy height from that TANNEMAX fusion over Gabon, Mexico, French Guiana and the Amazon Basin. Two other available products from the mission are highlighted here. The gridded version of all the Level two B vegetation structure metrics are presented at multiple resolutions and have been processed with stricter quality filters applied.
Speaker 2: For example, the mean foliage high diversity in the leftmost map was generated with JEDI shots acquired between April twenty nineteen to March twenty twenty three and aggregated in six kilometer grid cells. The Level four B country level summaries for above ground biomass are also completed at multiple time periods throughout the Jedi emission to help countries meet reporting standards. The figure on the right demonstrates country wide estimates of total above ground biomass in pedagrams created using the level four two point one version.
Speaker 2: This data set also includes biomass estimated by FAO for comparison. More details on how these metrics and data sets are calculated or derived will not be reviewed in this training, but there are additional trainings and resources in the NASA dacks. You can reference the data dictionaries, the algorithm theoretical based documents, or reference some additional trainings like the one we'll discuss at the end by NASA earth Rise for data preparation techniques for applied users and will that will dive into section three to discuss resources and tools for accessing, visualizing, and analyzing JEDI data products.
Speaker 2: So there are many considerations when choosing a data product to work with. Of course, as previously mentioned, ensuring the coverage and spatial and tempore resolution is primary. Your choice in using Jedi may depend on those mission specifications that we showed earlier. If you're looking to select data from specific timestamps. You may need to work with the footprint level samples since the gridded or other derived products aggregate annually or over larger time periods across the lifetime of the JEDI mission.
Speaker 2: Working with the footprint level data sets will be at twenty five meter resolution, which could be useful for localized applications or for use as calibration and validation or as training data themselves. However, these data sets are not represented analysis ready. This means that the user will need to deploy footprint level specific means of access, preparation and visualization strategies to work with the files, and will go over some open source tools and a demonstration with a Python notebook to exemplify this.
Speaker 2: If you are applying to a landscape, country, regional, or global scale application, using the gridded or derived products that are already quality filtered and aggregated over certain time periods could be really flexible options. These data sets are easy to work with since they're in GEOTIPH, but it's recommended to refer to the data preparation descriptions in order to better understand the assumptions that were made when generating those derived data sets and review whether those assumptions are optimal for your particular ecosystem's characteristics.
Speaker 2: So this reference table can be used to help you decide on how to handle your desired Jedi data product. So we will review the Earth Data Guy, tes Viz, slide rule, and Harmony Api in part two, while part three will review accessing Jedi data sets and Google Earth Engine, and the rest of the options are up here for your exploration. So the takeaway here is that not every product is available for as access or able to be processed across every platform. Some platforms like Earth Data Search are really only for data access and basic boundary overlap or temporal filtering, while platforms like tesviz, the Harmony API, slide rule client go beyond access and allow for more customized data processing, download, formatting, and visualization.
Speaker 2: For the greatest customization capabilities over large study areas, the user may need to deploy programming oriented approaches, especially when handling the footprint data. So platforms like and other packages in our Jedi Google Earth Engine or other Python workarounds like JEDIDB and other ones that will present in our demonstration today are good options that can enable you to have within platform data storage processing or link to other APIs or platforms. Any Jedi data product can be accessed via the Earth Data Search.
Speaker 2: It can be single download or bulk or cloud optimized methods. There are already many tutorials that exist to help you navigate the Earth Data system and interface that we will not replicate here today, and the same goes for coding related bulk access with the LP and ournl dax that provide command line and Python based guides to download from the catalog. With specific spatial and temporal subsetting linked here, you can acquire the footprints and download the HDF five files that roughly overlap with your location in time period from Earth Data Search, but further spatial clipping, variable subsetting, and any type of processing that will need to be done would have to be done on an alternative platform once you've acquired those files from the platform.
Speaker 2: So let's take a quick look at how the Jedi data products show up on the Earth Data Search interface. On the top left I typed Jedi into the search bar to get all the possible data product collections available to be filtered by the temporal and spatial extent specifications on the top left panel. There are the main options for temporal and spatial filtering. This is what controls the search results of the data products that can then be further opened or visualized. The geotips and HGF five products that will meet the filter criteria will appear and you can download the data.
Speaker 2: They might not be necessarily clipped or subset to the exact spatial extent we defined here with the polygon tool, so you input the time period, use the polygon tool to get a rough area of interest and as you can see, the entire swath, which includes the trajectories of all the eight of the beams are shown overlapping with the study area. The individual beams and samples are not visualized here. You can highlight the overlapping orbit to see its location and download the h five file for example.
Speaker 2: Other tutorials can go over AWS access for bulk download. But the level two B and Level four A, Level one B and Level four C waveform footprint products, so basically all the HDA five data sets show up in the same way
Speaker 2: where you're just seeing the outline and going back to the main results. Level three gridded elevation and height metric are actually visualized on the map where you can see the values and the spatial extents of each of those grids with a particular color legend and can be downloaded in a similar way. The Level four B biomass data values are also conveniently visualized with the color scheme and demonstrated with the grids. Some of the derived products are global TIFFs and will have associated p and gs for your convenience to visualize the data set, but other products the data will have to be downloaded to properly visualize.
Speaker 2: So for here we see the outline, the global outline of the data set, but there's no visualization. When we zoom back in we see that it just highlights that it overlaps. The Terrestrial Ecology Subsetting and Visualization Services platform is exclusively for the Level three gridded Height and Biomass and Level four A and Level four B biomass products. You can use this platform for acquiring data from conveniently pre processed locations or use the API tool for customization or subset by a user defined location.
Speaker 2: So let's walk through how to select a location for data processing of, for example, the gridded Level three product. Though you can acquire the prepared data sets or use the web service tool. We are going to demonstrate the user defined subsets tool. From here, you can define a location point or customize by manual input to grab generally overlapping data. This could pose an issue if you're looking to a subset by larger areas. Scrolling down the sensor and the data product can be selected. You can select multiple products at a time or type them in.
Speaker 2: You can also select specific variables if you desire. The next step is to input the desired time period to filter by, followed by optionally choosing to generate geotifs and choose between the projections. The data requests will be submitted and take some time to deliver depending on the demand of the request. The resulting subset data sets get sent to your email. When they are ready for download in multiple formats, the user is directed to a temporary page where the data in multiple formats can be downloaded.
Speaker 2: The order summary is also available with auxiliary maps and statistical summaries for modus and land cover and phonology, and other resources for visualization, as well as some statistics. In addition, the citation is generated and the location, subset and projection are summarized. If you wish to download the data. You can click the link and a local directory page pops up to save the files on your local computer. The same process can be applied to each of the available download links. Here you can see that each of the bands are available separately and in multiple formats like GEOTIP or as CSVS.
Speaker 2: Finally, the visualization tool can be accessed under the tab for order summary. From here you can run the visualization where each band from the Level three data set is visualized in a uniform way and a comparative table, so with the same color scheme you're getting each of those bands for the Level three product and statistics associated. Next, we have slide real client. It's a platform made in collaboration with the University of Washington and NASA Goddard's Face Flight Center. It's an open source framework for on demand processing of science data in the cloud with AWS.
Speaker 2: It also has access to ISAT two, Jedi, Lancelot, Arctic deem Rima, and a growing list of other data sets that are stored in AWSS three. This solution includes an API for customization and is well documented. Customizing the API may allow for more flexible use over larger areas, but this user interface tool that will go over today. Is also built for the ISAD data sets, which is a plus if you're looking to fuse these data sets in your application, but it is currently only developed for three of the footprint level data sets, which includes the Level one B waveform, Level two A heights, and Level four A biomass data.
Speaker 2: Let's take a look at how to interact with JEDI data using slide roll clients functionalities on the left panel, the user can select the data set they would like to visualize, process and download one at a time. On the map, use the polygon tool to define a relatively small study area to work with. This can be a limitation for users looking to use slide roll for larger studies. I can check the request parameters on the top left and using the advanced settings, additional parameters to what is predetermined for the selected data set can be applied.
Speaker 2: You can control timeout checks, select specific data variables, only specific beams or quality flags that you'd like to include in the final data set. And there's another parameter for data sets that we won't use. And as you can see, the request parameters are updated with these customizations. You click run to process the data. You may take a little while and then the user will be automatically directed to additional processing and visualization tools once the acquisition is run. There's a map that visualizes the beam transects and has advanced filter controls that allow you to specialize the visualization layers by individual beams, ground tracks, or select a particular orbital path.
Speaker 2: The map plotting can also be controlled where the data values for each of the shown attributes get displayed. When you hover over the data point, the user can interact with the map and retrieve the data values for every included JEDI variable. Next, the right most feature generates an elevation plotter for the data along the latitude. This is most relevant for the elevation and high products. There is also a three D viewer that displays the values along the latitude and longitude with a particular legend to demonstrate the range along the z access and then much of the parameters below are to control the interactive three D visualizations.
Speaker 2: Mostly Finally, there is a table tab that allows you to run an SQI querry to visualize the data sets that can be exported as a CSV. Okay, so now for our demonstration, we'll give an example of how to use the Harmony API, there's going to be a focus on accessing, downloading, visualizing, and analyzing the footprint level data. So really the question that we're trying to address is what do you do once you have the hd of five files in you know, whatever format that you selected, If you want to work with a larger area or specify with a particular shape file, you know, how can you do that?
Speaker 2: So when you download the h five files from Earth Data Access, they're not necessarily clipped directly to your spatial extent. So this solution will help you do that when you download the file. It also includes all of the variables that could be hundreds of variables, most of those you do not need. Maybe you want a handful about twenty two recommended So how can you down select those from the HDF five and then transform that data into a preferred format something that's maybe more compatible across multiple geospatial platforms like a shape file or a CSV.
Speaker 2: And just to know that Harmony API works with all of the waveform data. So that's level one B to A to B, level four A and four C.
Speaker 2: And so what you'll need is the NASA Earth Data Access account, the Google Drive account as well, when you'll need your credentials for the Earth Data Access ready to plug into the script and access to Google Collab. So this Python notebook will demonstrate how to download the data, select certain variables to subset that you'll want to analyze, generate some statistics and visualizations. We're going to use two example study areas. One of them is of Pinon Pine Forge around Albuquerque, New Mexico, as well as a paint Rock research forest in North Alabama, which is part of the geotrees Field Data Network and is largely managed by Alabama A and M University.
Speaker 2: So just as for your awareness the Google Collab notebook, if you do command or control question mark if you're looking to customize a notebook and plug in your own AOI or change the data products. So we'll be using an example with level two B vegetation structure. If you wanted to look at level one A or look at four A for biomass, you can command search for the question marks and follow the input prompts throughout the demonstration. Okay, so when you click the link, it will take you to this GitHub repot for our said too.
Speaker 2: This is where our script is held. The tutorial. There's a little read me here that gives an overview of what we'll be doing. So you're going to search and retrieve the Jedi footprints over particular areas of interest. You're going to filter them by specific metrics and other important variables which we'll talk a little bit about. We'll be able to visualize them with different figures, get some statistics as well as some three dimensional figures. You can export them as a CSV or a shape file, and everything is run in the Jupiter notebooks, so there's no set up on your local computer required.
Speaker 2: And just for a quick overview, this is all the steps that we'll be doing within the script. So we'll set up our environment by getting our libraries with PIP and stall import our core dependencies. Set up some directories, so that's using the temporary space which is housed here on the little folder tab that when you're connected to the run time, you're able to save all your files. Here. We'll get access to the NASA Harmony API capabilities. We'll set up our study area and I'll show you a couple ways to connect your study area data, establish your temporal period to filter Jedi by, and then we'll go ahead and execute the downloading the hd of five files with the Harmony API.
Speaker 2: But when you download those files, you're going to get a bunch of data that you probably don't want. There's in some of them hundreds of variables that you're probably not going to use, and we list the about twenty two recommended and not varies depending on the level product. So we're going to be looking at level two B the vegetation structure metrics, and we'll go ahead and subset the files greatly reduce how much data is being stored, which will really help with computation and when you're working within in other platforms.
Speaker 2: We'll take a quick look at those files to make sure that everything was good with them, and also just to demonstrate how to open the HDF five file in the Collab notebook, convert them to a geodata frame which is usable within the Jupiter Notebook in that temporary computational space, and then offer some options to export it to your drive as a shape file or a CSV. We'll talk a little bit about some quality filtering techniques that we're discussed in Part one, but if you would really like to go more in depth on preparing the data with more application specifications, in mind.
Speaker 2: The repo that this is housed in is a training that goes more in depth on that, and so there's several modules that talk a lot more about application specific. Quality filtering will just be applying a couple which includes getting rid of no data values and seeing how that changes the data set. We'll map those and take a look at what the data that looks like before and after that quality filtering, we'll do some basic statistics on counting the observations per year over the study area and generate our final filtered data set.
Speaker 2: We'll go ahead and plot and explore the data sets, so we'll have some violin plots that look at each variable from the vegetation structure metrics per year per study area, generate a pair plot that will show some relationships between those available data sets, and generate some three D visualizations. So when you're here, you have a couple options to access the file. So you can click this tab here and that will bring you to the script that you can run. So I'll open then in a new tab. Just for reference, you can go to the file itself here and click download.
Speaker 2: That's just a GitHub error that they're still fixing, but it's still a valid data script, so you can download it and then up load it to your Google Drive or Another neat trick is when you're open to this file, you can hit two colm right after GitHub and you'll see that the icon changes. Hit enter, and then you're taken to the script as well. So when you get to this area, you're probably going to need to connect here, so you hit connect and it'll set up. And just two important tabs here to note are this table of contents, which is essentially that outline that I just went over, which when you click somewhere it jumps you to the section where you can start executing code, and the folder here.
Speaker 2: So this when you don't have it connected, will not have these two folders. You'll just be connected to this temporary space as you can see here in this disc This is what you're computing under. But I'll show you when we start off, we connect to your drive, so here is your personal drive that you can and connect to data or upload data sets to, which will show. And then when you connect to the temporary space and we start running through the code, it will generate this folder here that only exists in this when you're working through this script live.
Speaker 2: So when you start, we have again a little overview and just a note here, so you can follow along with this script and run sell by sell. That's just kind of at your own pace. But you can also hit run all and follow the input prompts which will show in this tutorial in order to run through what's written in this script. If you want to do further customization and change, so we're working with level two B. If you want to go to level two A, level one B, level four A biomass, you will be prompted to input certain variables and or you can search, so you can do command F for a question mark and you'll see all of these areas here where you might want to customize.
Speaker 2: So, oh, I want to work with level two A instead, that's somewhere where I would type in the code to change that. If you want to input a different AOI and say you're working California or something like that, you would need to rename these And just for the sake of how this tutorial is written, if you want to customize it, it is written in dictionaries, so it's going to iterate over multiple aois, so just a reference on you'll have to modify some of the script if you want to look at just one area.
Speaker 2: So that's just a little thing to work around. So when you start, you can either hit run all follow the inputs, or go sell by Sell. Here, we're going to go sell by Sell. So to set up your collab, you're going to install, So I'm actually gonna head over to some pre run code here. It should take in total about thirty minutes to run the whole thing, depending on the size of your AOI, which isn't too bad. So when you do pip install, you get a bunch of these prompts. It may prompt you to restart the session.
Speaker 2: You don't need to do that. You can click cancel and as long as this green check mark comes up, you're good to go.
Speaker 2: And then from here you want to mount your drive. So when you mount your drive, a pop up will come up and it'll ask you to put in your Google Drive credentials, just like you're signing in to your email or whatnot, and then the folder will pop up here like I showed, where you can access all of your content. Now we're going to establish the directories. That is how we're going to save data and locate to our aois. So you can add more aois if you wanted. Again, be careful with some size. It's not unlimited.
Speaker 2: But if you're working with a product, you want to make sure that this variable is changed and make sure that your ais are set. So you can just click run for that. And I put some directions here on just clicking run, which means that it's just kind of automated. Everything is connected to the variables that you change, so you can just go ahead and click. And this sets the folders where everything is located. So you have paint Rock that's within that main directory and pinon. So we create this base path that we then change our directory to.
Speaker 2: So this is just telling our code, hey, we want to work out of this main folder, not anything else. So that helps us execute all of our code. And this may take a little while when you run the required packages.
Speaker 2: So we need all of these to plot, to use the widgets to map, to get Earth data access, to get Harmony, all that kind of stuff, and then we're ready for step two. So here this is where you're gonna need your credentials for Earth Data access. We're gonna basically set up the request for Harmony API, set up all of the requirements to make that call to Harmony to get the data. So first we need to input our credentials here and it'll prompt you with an enter to enter your your user name and your password this was also shown in part one.
Speaker 2: And then we go ahead and actually use that password to authenticate and connect to Harmony client. I might take a little moment, and then from here, this is where we define the Jedi product that we want to use. So when you hit this code, it'll ask you to enter something and so Harmony works with all the footprint level data. We're going to be working with Jedi two B, so you can either type it or copy it. So if you were plugging in another data set, you would copy any of these other names here and hit enter and there you go.
Speaker 2: So from here we want to make sure that the Harmony client is actually accessing the variables and all of the capabilities specific to the Jedi product that we just defined. So when you run that, it allows you to see that oh great. So for each beam this variable in RX processing called shot number, I'm able to access it. And so you see there's hundreds of different variables which we will show you how to choose from and get So now we go ahead and get the collection ID. This is what's going to be used in the Harmony API request.
Speaker 2: So really we're just telling it to give us these in a variable that we input later. I want to set up a variable that allows us to connect to the path, the area, the folder that is connected to our AOI that will use in the later functions. Again, when you change these above, these should change if you're plugging in your own AOI, and now we want to import our AOI data, So this is you're going to have to have them in geojason format, so that's quite important. What we're using here is the gethub a git hub geojason, so you can do this as well in your own repo.
Speaker 2: You would be able to upload wherever. It doesn't have to be in the same exact folder, but you could upload a GeoJSON. You go ahead and click it and get hub is this nice thing where you can visualize. But you want to click raw over here, which gives you the raw file and then you just copy this you here and paste it into your code here. So for here, we're just using what's openly available in our own rebel, so you can run through this tutorial, but you can change that yourself, or you can connect to a geodjason that is located in your own drive.
Speaker 2: So a new trick here is when you connect to drive and let's say I just pick a random I'll just pick a random file here. So say we are looking at this video and we want to copy the path here, So we click those three dots copy the path. That's what gives us this path right here. You would paste that, you would need to uncomment, so you would need to get rid of these hashtags in front. If you were to run with the plugging in your AOI to the drive instead, and that should work. Then you go over here and you want to convert those geojasons into a geodata frame.
Speaker 2: So that's basically just reading the file in this collab notebook. So when we do that, it makes it available to be mapped. And so we click this here too, and it'll display your AOI sites. So we're looking at pinon pine forage areas, so really yummy pine nuts here outside of Albuquerque, New Mexico. So that's quite a large area. And then we're looking at a very small area of a conservation forest here in northern Alabama.
Speaker 2: Not much has happened here, not much you know, in any logging fires and what have you. Super beautiful site. So those are what it's great we have those aois. They look great. And notice that, well this one was a rectangle, but this one had as a more complex polygon. So you're able to upload your GeoJSON that's more complex. But when you actually run the Harmony API, it takes a bounding box, so you may need to clip your data afterward in a different like within collab or whatever platform you're comfortable with.
Speaker 2: But when it makes the request to NASA the DAS, it needs just a really simple polygon. So this is what's taking the shape file that you input and making a box around that instead of using a complex polygon. Then we define these this variable that is just configuring the names of the sites that we're using for the AOI bounding boxes that we've created to the folder that it's housed in. So that's over here. These folders to help us with later functions. So now we've established AOI, So where are we here?
Speaker 2: So we've gone through and we've gotten Harmony. We've defined our Jedi product, we've defined our aois. Now we want to establish our temporal filter. So remember just a render here. This is when the data is available, and you're gonna get to input that yourself. So here, when you run this, it will allow you to hit enter. So I'm gonna try this over here just so I don't okay, So we can hit enter here and we can type in two January first hit enter, then we can do to the end of the year. Let's see that you don't have to use the zeros in front it.
Speaker 2: It takes regular single digit numbers as well, so you can have this enter or if you want just for something static, you can uncomment this and use a static variable. So now we're going to actually prepare for downloading the data with the Harmony and API and request it. So we're just going to define some tracking functions that you'll see in the code that's printed out. These are really helpful for understanding the size of the data. You know, how many files did you download, and you know, just other tracking information, how long did it take things like that, and so we'll run this download script.
Speaker 2: This can take a while sometimes if I did the whole time period here, the whole time period of all the data that's available as you can see for the most part, and it takes about ten minutes to run over these aois that can differ by the size of your AOI as well, so you don't need to know much about this code. But I did just want to show what the actual Harmony API request looks like. So remember, we define our concept ID, so that's what's specific to our level to be product. Then we have our bounding box and this is our dictionary of data day that we input for our two aois, and then we've defined our temporal range and so that's what's actually getting the file.
Speaker 2: So there's no variable subset. You can do variable subset with Harmony and I, you know, refer you to the documentation to do that. The reason why we chose this alternative is because there were really big limitations on the size of the AOI that you can use, and we wanted to enable a solution that would allow us to look at a larger area or look at multiple areas to compare over any time period that we'd like. There are still limitations here. You can't do the whole globe, you can't do a whole country that you would have to customize coding for.
Speaker 2: But it's a lot more flexible than having Harmony itself parsed through the HGA five files to get the specific variables. So instead we download the files in the temporary space and not onto your local folder and then move on to step three to actually subset the data to the variables we want. And so here this is what it looks like. It's a lot of coding, but it reminds you of what your temporal range is. It lets you know which AOI it's working on, where it's at. You know, is what's the percentage of processing that it goes through.
Speaker 2: And then it actually lists out all of the files that it downloaded successfully, and it does that for both of the aois, and then at the end you get a nice summary of how many aois were processed, how many raw files were there was gigabytes, and the average size per the files that you understand what's going on and where they're located. So you see, wow, the Pinon area we knew it was larger and had one hundred and six files. Paint Rock only had six files, even though we're looking at the whole range of years in twenty nineteen to twenty twenty five.
Speaker 2: So I wanted to show these really big differences, and we'll see that throughout the figures. You can have you know, Jedi. You could be wanting to use Jedi over an area, but it really depends on what your area looks like and how Jedi happens to collect over that region. So now we're going to move on to actually subsetting these variables. So we don't want, you know, six gigabytes of data, and we don't need all of those. We don't need hundreds of variables. So from here we're going to use this text file.
Speaker 2: This is located in the repo that I showed before. It's basically just a text file of all of the variables. It's like a dictionary of all of the variables that are available in the Level two products. So you can see hundreds and hundreds of variables, and we don't need all those. So we're going to use this text file that will just load in by running this code and see here, Great, it gets all the data in order to then plug in. Hey, I only want Level to BE data from this, so this will prompt us to put in and from here you can I'm working with a level to be, so I can copy this or you put in the product that you're working with, enter it here, great, and then it gets me only the variables that are available for that product.
Speaker 2: That's still many more than I might need depending on my application and what I'm doing, the research that I'm doing, and so from here we actually want to select those variables. So what's important about this list here and you can also visit the NASA Data dictionaries, is that we need this full filepath name. So even if I just want to get the leaf off flag, I can't just type leaf off. I need the folder or the group that it's in in the HDA five file. So this changes across the product. So that's why this step of visualizing and getting all of the data or looking at the NASUB Data dictionary is really important to get what that real total value is.
Speaker 2: So I want latitude of the highest return, well, it needs to be written as if it's from the geolocation folder. So from here this is going to ask you to input. You can either enter or you can uncomment here and create this list yourself, depending on the product. So something to notice, because we're working with Level two B, these products are specific to level to be for the most part, and these ones are just we recommend these. These are pretty commonly used if you're working with any Jedi data set that are important for quality filtering, for getting your location coordinates, for being able to understand the behavior or penetrative ability of each of the shots.
Speaker 2: So I highly recommend keeping these, although the prefixes might be different, so in like a different product, you know, maybe the degrade flag is in a different path folder here, might not be in geolocation, might be somewhere else, so you might need to double check that with when you're changing your folder to that. And so from here, I'm just I already wrote this out. So you just have the full path names separated by commas in a space, copy that and input it into your enter here. Or you can put a bunch of apostrophes like this and use this variable instead.
Speaker 2: So great, it loaded those in and it formatted it the way that it needs to be. And now we're going to define some functions and help us select the variable itself. So you just click run here to subset. Now you can also specify how many beams that you want, So again, subsetting the variables and by beams is something that's available in Harmony API, but it greatly reduces your computational ability, so we're doing it here post download. So a really common thing to do is to use only the power beams.
Speaker 2: So power means that the laser is that full power. Coverage means it was about reduced to about half of the amount of power, which means that maybe it wasn't able to penetrate denser forests all the way through. That depends on the AOI, and a lot of people are still researching that. So if you want to include all the beams and look at a comparison, and we'll show you a little bit of how to do that, that could be really really interesting. Or if you already know that you're working in a really dense tropical forest or something like that, you should probably just use the power beam.
Speaker 2: You can just select these four beams instead. But for this we want to look at all of them because I know that we don't have as dense forests in the US, and I would like the ability to compare them if I really wanted to. And also, when you separate those beams, you're greatly reducing the amount of data that you have available, and as you'll see, that can be a huge barrier to using Jedi. When you actually apply these quality filters, how much data is left. So now we're going to find some helper functions to actually download the HDA five file.
Speaker 2: So we'll just click run on these. They're just some functions that help us do stuff. And this is just basically opening the HDF five file, getting the data sets, putting them in the right place in the right order, reformatting them into the way that we want, and then we'll execute that function here. So you just click run on all of these, and what it'll look like is so I have it print out each of these things just to keep track of what's going on. So for paint Rock, AI are really small AOI we're making.
Speaker 2: We're creating a folder that's going to be our raw downloads, or where we're getting the raw downloads from is from this folder. But we're going to create a new folder that's for variables selected. And actually you could see that here. So we're in paint Rock Great, we created a folder here where we got all of our subset variables in. So as we do that, we're opening each of those files and it's you know, going through and finding each of the twenty two variables that we specified from each of the eight beams that we want and processes each file so it find defines these summaries of how many variables we're getting.
Speaker 2: But from this first file that we see here this from twenty to twenty two whatever, you see that there's no beams here, so it's skipping them and that's just because of the spatial clipping that we got. So even if you had the full file available, because of the spatial subset, it did not include all of the beams. So even though you're requesting for all the beams to be selected, that may not be the case depending on where they fall over the earth and over your shape all that you defined. So this is where we really get into and something that you know earth data.
Speaker 2: And until you convert the HDA five into something more readable and are able to look at this kind of information, you don't really know exactly how many shots within each of the beams are available over your aoy. So this is really useful for tracking what's actually going on. And then you realize that, oh, you know, I only I don't actually have coverage beams for this area. So we'll actually visualize this and make it a little easier to read. But this just goes through tracking, it defines the size and then you get to see, wow, we reduced that file ninety eight percent, which is great because we did not need all of that other data.
Speaker 2: So now we just do a quick check are they where we want them to be? It's just for security, and then we're going to go ahead and look at those HDA five files, so just something interesting to look at. We're not really going to work with them anymore since we're going to convert it. But if you want to look at the raw files, so let's take a look at a file in Pignon and we're looking at the raw files here, so you can choose from any of these that you want. You can say, okay, cool. So all of these beams are potentially available as well as these other groups.
Speaker 2: And these are all of the variables that are there from each of these beams in the raw data set. We don't really need all of them, but just to see great, they're all there. And then if you want to look at look at the variables selected, see what happens after you subset them. You get to see, oh, okay, great, So we actually only have these two beams available after we subset, and looks like we got all of the variables that we selected. So that's pretty good. And you can go ahead and look at the other files and see how these change specific to each file.
Speaker 2: So now we're gonna actually look at those variable characteristics across the HGA five file. So this is another example of opening up. So you hit run on that and you run this, and we're just going to look at first the raw download just to be able to see what's going on with this data set. And so something that we learned about in the lecture part of part two was that the HGF five files are holding data sets in each group. So if you have ancillary or the beam group that are of different types. And so when you look at this file here, and it's also in the Data dictionary that's an online page through the NASA, but you get to see the shape of the data set.
Speaker 2: Is it you know, one dimensional, two dimensional? How many points are in it? And what type of data is it? So is it integer? Is it float? Is it this or that? So if we want to look at okay, so let's look at this. We have cover here from this file, it looks like there's only one shot available and the cover is afloat. So this is just really to check what's going on if you needed to do some extra research with that and understand what's going on with each of those files. So these are again all the data sets, but if you want to look at something that you subset and see what's going on.
Speaker 2: Okay, So there's nine shots that are available here in this here you'll see that pa Z the Z profile. So we won't talk as much about this, and there's another training with earth Rise that talks a little more in detail about this, but this is just how it's a two dimensional data set of how the PAV values are stored along the vertical access so you know, along the heights you get a PIAI plant area index value. And so there's thirty different value possible values. Not all of them may be valid data for each of the shots, so it's a different format.
Speaker 2: It can't be opened the same way. We won't go over that in this training, but just to recognize that there are data sets of different sizes the same thing here, there's just different ways that they're formatted within the HD of five files. So if you were to customize, you need to pay attention to that and then we get a summary of those information. So here this is an optional step. If you're working on your disc is full and something's going on, you might want to go ahead and delete the raw downloads folder to get if you're not needing that and you just want to work with your final data set.
Speaker 2: So when you click run here and it's on false, it'll tell you, hey, set it to true in order to rerun and actually delete your files. So that means come over here and again there's a question mark to find that, and you'll just change that to true, hit run again, and then it will delete this entire folder. And if you do that again, if you needed that later, you're going to have to rerun the download to get those back. So now we're going to actually create the geodata frame with the subset of HDF five files.
Speaker 2: So this is what's going to help us start thinking about visualizing the data and being able to convert it to another data format. So we're going to define some functions that help us do this, lots of documentation throughout. If you're looking to customize, we try to make it easy and understandable. So now we're going to go ahead and actually execute converting those HTA five files for each AOI, and we're going to create a dictionary of geodata frames, so that's one geodata frame per AOI that we'll use to plot later.
Speaker 2: And again it's going to be using the data that was subset in order to do so. So when we do this, it goes through and it tells us what's going on with each of these again just to help us track. So here we have that it went for each beam. It's getting all of the records. So those are the individuals the sixteen shots here, thirty six shots, fifty eight shots in each of these beams for that file. What did it do to each of those data sets? So PEGAP datasy we don't really work it with this, but this is something that it truncated the information that might be something that needs to be investigated and fixed if you're using that data.
Speaker 2: But here it just lets you know, hey, I wasn't able to find this in these beams because those beams don't exist or how many records were actually processed. So for the most part, everything was successfully processed, but it might need to you know, be changed if you're working with these data sets. And you find that they're not correctly downloaded. And so here also the reference length that we're using, where did that go? The reference length is essentially just looking at lat lowest mode, so that's we know that we're looking at the latitude and longitude of each shot.
Speaker 2: We know that that shot exists. So if there's thirty six latitude points, then we're just going to assume that those are the thirty two shots that we want to be able to generate our geodata frame from.
Speaker 2: So that's just a bunch of tracking records for that, and then we're able to say, great, we have this available. We want to create a variable that allows us to access the individual geodata frame for each product or for each AOI for that product, so that you can look at the data sets. So great, we have the lat long Okay, everything looks great, it's in it's it's in order. You can see every every data set here some of the values and so this is what we're more familiar with.
Speaker 2: So now this is an optional step. So what we downloaded here was the so we had the raw file that was subset to the variables that we wanted, and then we converted that to a geodata frame, but no quality processing, no getting rid of, no data values, no look into the actual data sets themselves and their quality has been done at this point. So if you were looking at a study that was you know, comparing different quality filters that you wanted to apply, or looking at you know, validated the diferens in different validation of the results that you get, you might want to save the original data set that has had nothing done to it itself.
Speaker 2: So this is an option to do that here before you start applying different filters. And this might help you compare if you're if you're doing that, or maybe you found some literature that you think is really applicable to your ecosystem. You don't really care about the raw files you've already you know, you know, you're just going to apply the final filter data set and that's what you're going to work with, and you can skip this step. But when you do this, uh, it connects to your drive and it creates a folder for shape files or csvs where the data is actually saved.
Speaker 2: So you have that's just so it should look like this piinone with the product and then you have your shape file files all there and again that's the original nothing done to the data set, and now we're going to start thinking about exploring the quality filtering techniques. So we just click run on these cells to establish This just helps us visualize and then so we may have defined the beams by listing when we wanted to subset, but it's kind of hard to remember what each one of those numbers is corresponding to.
Speaker 2: You'd have to you know, memorize that. So instead we want to actually label them as power or coverage themselves. So we're going to define this function that allows us to create a column that tells us whether it's coverage or power instead of the individual beam name. So that'll help us generate some later figures. So now we have our original data set and we're really interested in seeing well some basic statistics of what's available when this could be make or break for your application of Jedi, and so we're going to define some of the functions that help us actually get this figure going plot it.
Speaker 2: So again you just click run on all of these. There's not much else to do. If you were changing some of the aois, you might want to change some of the colors, or if you added you might need to add some specifications to colors here and we go ahead and click run here. And so this is an example of what the output would look like. So here we have Wow, we have two hundred and fifty five total shots that are total observations that were available between twenty nineteen and twenty twenty five. Four pinon for Jedi level to be.
Speaker 2: And now we get a little summary chart here that tells US across the years, how many of those shots are associated with those U what was the peak month of how many shots they're worse? You see that can vary by what is that thousands across across each of these months per year, So it can be really important if you're looking at seasonal, seasonally specific applications. So I think I have this actually pulled up here. So what this does is it puts all the aois on one subfigure for each year, so you can do a comparison between them.
Speaker 2: And as you saw, paint Rock has very very few. It only has total of sixty four shots across all of these years. Now it's a small AOI. I don't know why it has that few, but this could be your point, you know, if you're looking, Hey, I want data from let's say twenty twenty two in April or May. For either of these, you can't do your study with data from that area because it's just not there. So that's what this code helps facilitate, is understanding. You know, even if they is potential, it really is variable.
Speaker 2: You just have to open the data in order to see how much is available for that time period. Or you see, you know, there's only nine ten shots here. There's no data from paint Rock in this time period, So yeah, it can be incredibly variable. So that's what this figure is doing, and then we get a nice other kind of descriptive comparison across those years to just help us understand. So if you're using it as trained data, this could be really really valuable for understanding the representativeness.
Speaker 2: Of course, this is total shots. This isn't even over individual land cover types. If you put a land cover mask, that could also greatly change the amount of available data
Speaker 2: depending on your application. So now we're going to make a bar plot that is similar to this, but now comparing beam types. Like I said, there could potentially be a very big difference between the power and the coverage beam. So when wee we just click run to find this function in order to generate our plot, and so we're going to get this monthly observations for one AOI. So we're just looking at payrock and we have color coded for each of the years, separated by power and coverage. So we see, okay, well, of all of the shots, they are all power, so maybe that's possibly more reliable shots.
Speaker 2: Really depends on the application, but this will help us better understand when we actually go and validate. And then for this one here, we have some coverage in twenty twenty one, but as you can see again, there's a very different distribution across the years, and of course we get a table that helps us visualize that. And I think I have yeah, and you can switch, So this widget here allows you to switch to the other. Ay, it might take a while depending on how much data there is that you're working with.
Speaker 2: And I kind of wanted to show jumping over here. Two. I only selected one year's worth of data from twenty twenty two. And again this is summarizing all of it, but if you just wanted to select, you can do some comparison there. So for pinon, we have a bunch of data that's distributed very differently across the coverage and power beams. This could really help you understand again if you're using it as training data, can help you understand the behavior of that information and summarize that across as a table.
Speaker 2: Okay, now we're going to look at us out a plot and see how many of these values are no data values.
Speaker 2: We want to get rid of those. So here we have before we separate. It's really difficult to see what's going on and just understand how many of those values are removed and what our information looks like across all the years, all the data for the AI when that no data is removed. And see, there's a lot more information that we're seeing here across the track of observations. So this is along each of the observations that get dealt. So see it's a lot again, there's only sixty four in paint Rock, but we understand the behavior a lot better.
Speaker 2: So now with that understanding of a little bit about the no data and how much is available, we want to go ahead and generate our filter data set. So again, if you don't have any data over the time period that you're looking at, then for the study area that you want, this might be the end of you know, you using this script where you'd have to change course if you want to use Jedi or if you have data and you want to make it higher quality, then you know, start applying these filters and see what happens.
Speaker 2: So we're going to remove the no data values. This will change when you're using a different product you're going to use you know, uh, this is for FH FH Foliage High Diversity plannary index and cover. This is what we're gonna be looking at in this level to be data set, but we also want to look at if you want to look at level two a elevation and height, you would change these here and get rid of those no data and then we have the quality flag and the degrade flag which we talked about in part one that will apply to remove to keep high quality flags and get rid of degraded flags.
Speaker 2: There are many other types of filtering that you can do that we won't cover in this training that you would add them to this section here if you wanted to apply those and compare. And now you see that the data sets reduced. So instead of sixty four, now we have forty two in paint rock and instead of two hundred fifty thousand, we have one hundred and twenty four thousand, So again that'll vary by the study area that you have. We can map the differences between that. So we have the original here and then we have the filtered data set, and we can see just how different those are.
Speaker 2: Again, the scale between the no data and the data values change from what's available, and we can check that for the different data products as well. Pai about a plannary index. Yeah, now just to show what happened after we filter, What does our distribution of the data across the months look like after we get rid of that data. So again that can be really variable. You just kind of have to see it to see what happens. Could stay largely the same paint rock, didn't have too many points that were taken out, but you'll see sometimes that could get rid of an entire month that you were hoping for by applying the filter, which can be really important.
Speaker 2: Okay, and then again an option to save it to shape file or CSV. This is your final filtered data set. This would be your analysis ready data would be saved in the same location with a different suffix specifying that it is filtered. So now we're just going to generate a few figures. We're going to make a violin plot that's gonna again help us summarize but get a little bit more information on some statistics of what's going on with each of these variables. So we have you know, the with the size of the violin plot is telling us how many of the data are there in that month that are contributing to these statistics across So as you can see, there's just different shapes that allow us to understand that.
Speaker 2: And I believe we can see a little bit more with pignone here. So you'll get a plot for each month to understand a little bit more about the distribution per you know, okay, what's going on here February not mentioned variability, So yeah, lots to look at, lots to interpret over your study area, and then you do that for each variable. So for level to be, we just looked at the foliage high diversity and you can do that for plant area index. So again you'll have the same time periods because all of these variables are built off of the same original shots, but they'll have different values and distributions.
Speaker 2: Now we'll generate a pairple This will include all of the variables that are included in that original subset that we copy and pasted everything, So we don't go over all of these. You might need to take an additional training or go to the algorithm theoretical based document for that, but this just helps you understand the behavior a little bit more. So we know plant area index against cover. We see that there's a really strong relationship here, and so this is just a nice way to start understanding a little bit more about the data sensitivity of the beam.
Speaker 2: How well is the the beam the shot able to penetrate through a canopy, and how does that relate to canopy cover. You get to see a little bit more about that information here. So that plots across all of the available data sets that are in your in your file that you subset, or you can select a few. And finally we'll go over some three dimensional plotting. So here we're going to make an interactive plot that allows us to play around with what's available. So these are our forty two shots. As we're looking at the filtered data set and how it looks, it's plotting the visual of the value of the variable itself.
Speaker 2: So right here we're looking at plant area index right, and the value of that plant area index is what's put on the z axis for each of the shots. So you can look at it in different ways. You get to see the tracks there and whatnot, so there's different ways to view. You can save it as a p ANDNG, you can pan, you can do different types of rotation across, make a video of yourself looking at it. Pretty cool, holo use interactive site and yeah, there's a lot more in that other one. We'll see if it loads.
Speaker 2: But then you can also create a static three D plots, so something that's a little bit easier to save and visualize. So here we have this example from Pignon that's you know, thousands of shots that we see in order to better understand each of these variables cover plant area index, foliage high diversity, get a sense for what our landscape looks like with all of these shots, and then again you see wow, you know, this might not be a lot of data. You'd have to think about what kind of value that plays in your study area.
Speaker 2: And again this is across all of the years, so it could be greatly reduced if you're just looking at a particular time period and not aggregating over and this is what it looks like when you have a lot of points. Okay, so thank you very much. If you have any issues, there is a discussion forum on the repo here having to do with these tutorials, and I hope you have fun exploring Jedi. So to conclude this part too, just wanted to leave with some recommendations on additional resources and how to stay updated.
Speaker 2: So, as we've mentioned before, a really great reference for understanding the basis and background of all of these data sets are there individual algorithm theoretical based documents as well as the individual user guides that are hosted by the docs and on Earth Data Catalog. The OURNL and lp doc sites are always processing updating products, uploading new versions and archiving older versions as well. So these are really the best sources of information and original and new data sets that Jedi and the Jedi Mission team are putting out Google Earth Engine are Jedi Jedi dB, which is a Python api that recently came out, as well as some of the other tools that we mentioned are continuously advancing their functionalities, so always revisiting their resources and their their main pages you know, will be updated with new functionalities to test and if you're interested in learning more application specific information on you know, once you have the data or you're looking at it, you finally mapped it with some of these example scripts or the tools that you've been able to access and run through, you know how to interpret the data, understand some quality checks, how to improve it or optimize it with other data sets.
Speaker 2: You can take the NASA earth Rise Jedi Applications focused training. It has another series of hands on tutorials UH in collaboration with some of the other collaborators who generated this training with our set that are specific for data preparation and analysis recommendations. And given that Jedi is a novel sensor it was launched in twenty nineteen, the research is still very active. So the literature is also a great place to go and keep up with in understanding new applications or recommendations of using Jedi, and the Jedi Mission Page conveniently hosts a Jedi related literature Zoto group that gets continuously updated with the latest literature.
Speaker 2: So thank you very much and I hope this training advances your use of Jedi.
Speaker 1: Thank you so much, Stephanie. Now I will provide a summary of Part two. Let's review the key concepts from Part two. First, we cover jedi's mission objectives and applications operating from the International Space Station with data from twenty nineteen to twenty twenty three and resumed in twenty twenty four. JETI quantifies three dimensional ecosystem structure for carbon cycle tracking, fire management, weather prediction, water resource monitoring beneath dense forests and topography mapping and vegetated areas.
Speaker 1: We explore jedi's data product hierarchy. Level one B products provide geolocated waveforms. Level two A extracts elevation and canopy height metrics. Level two B calculates canopy cover plant area index, plant area volume density and foliage height diversity. Level four A models footprint level biomass, while Level four C provides a structural complexity index, capturing three dimensional canopy complexity. Understanding file format it's crucial. Footprint level data sets at twenty five meter resolution use HDF five hierarchical format, requiring specialized tools or coding skills.
Speaker 1: Gridded products like L three and L four B use familiar GEOTIP format at one kilometer resolution, with quality filters already applied, making them more accessible. Data fusion products overcome Jedi's discrete sampling limitations in certain ways. The mission team developed fusion products combining Jedi with ISAT II for increased sampling density, Tandem x InSAR for wall to wall tropical maps, and landsat for continuous canopy height. Products leveraging Jedi's vertical precision alongside broader coverage sensors.
Speaker 1: Field inventory plots from networks like US Forest Service Inventory and Analysis established the aleometric equations relating ground measurements to Jedi observations. Future missions like Hedge will build on this with increased beam density and swath mapping for improved change detection. Finally, you learned about access tools. Earth Data Search provides downloads with spatial temporal filtering. Testfis offers subsetting for level three and level four B products. Slide Rule Client enables cloud based processing, Harmony API handles HDF five transformation.
Speaker 2: Coding.
Speaker 1: Platforms like our Jedi and Google Earth Engine provide custom analysis capabilities. Matching the right tool to your workflow is key to working efficiently with Jedi data. In Part three of this training, you will be able to utilize Jedi standard biomass products at both footprint and one kilometer scals. To map forest resources, You'll be able to access the open source obi wan API to generate estimates of biomass changing areas of interest. You'll also be able to identify how obi wan can be used to create baseline scenarios for forest carbon accounting projects.
Speaker 1: To receive a completion certificate for this training, you will need to submit the homework assignment, which opens on November six and closes on November twentyth.
Speaker 1: This is the contact information for our guest instructor, Stephanie Humanez, as well as for myself listed here. We're also including links to the r SET website and the r ST YouTube channel, where you can find many more trainings like this available for free online. Here is a list of references and additional resources to look into. Thank you so much for your participation in today's training.
Podbean