NASA ARSET_ Estimation of PM2.5 from AOD – Methodologies and Available Datasets
The Story
Welcome to this highly analytical and public health-focused episode of the NASA Live Video Podcast: "NASA ARSET: Estimation of PM2.5 from AOD – Methodologies and Available Datasets."In this episode, we tackle one of the most critical challenges in atmospheric science and environmental monitoring: tracking fine particulate matter (PM_{2.5}) from space. While ground-based air quality monitoring stations provide highly accurate tracking, their spatial coverage is heavily limited. To fill these global gaps, scientists rely on satellite-derived Aerosol Optical Depth (AOD) data to estimate ground-level air pollution and assess public health risks on a global scale.
Through the framework of NASA’s Applied Remote Sensing Training (ARSET) program, we break down the core methodologies used to translate columnar AOD values into accurate, surface-level PM_{2.5} measurements. We explore various quantitative approaches, ranging from standard empirical and statistical regressions to advanced chemical transport models and machine learning frameworks. Additionally, we provide a comprehensive overview of available open-access datasets—such as those from MODIS, VIIRS, and MAIAC—and discuss how to account for meteorological variables like planetary boundary layer height and relative humidity.
Whether you are an air quality manager, an epidemiologist, an atmospheric researcher, or a space enthusiast curious about how satellite optics measure the microscopic particles in the air we breathe, this episode delivers essential technical insights. Subscribe to the NASA Live Video Podcast to stay connected with the absolute frontier of space exploration, remote sensing data application, and cutting-edge earth science!
Speaker 1: Welcome to Part two of our aur set training on
Speaker 1: estimating surface PM two point five using satellite data and
Speaker 1: other information sources. In this part, we'll be learning about
Speaker 1: the estimation of PM two point five from AOD and
Speaker 1: discussing the methodologies and available data sets using these methodologies
Speaker 1: to create PM two point five estimates. The trainers for
Speaker 1: this part will be doctor Aaron von Donkolar, a research
Speaker 1: associate at the Washington University in Saint Louis, doctor Paulan Gupta,
Speaker 1: a research scientist in Godard Space Flight Center, and they'll
Speaker 1: be joined by doctor Jin Juncio, a assistant research scientist
Speaker 1: at Morgan State University. Our objectives for this training will
Speaker 1: be by the end of part two that participants will
Speaker 1: be able to explain the geophysical, hybrid and machine learning
Speaker 1: methodologies used to infer surface PM two point five from satellite,
Speaker 1: AOD information and other data sources. Second, we hope you'll
Speaker 1: be able to differentiate between these available data products for
Speaker 1: surface PM two point five from NASA and the Washington
Speaker 1: University in Saint Louis based on their methodologies and the
Speaker 1: strengths and weaknesses of the two data products. And finally,
Speaker 1: you'll be able to use NASA tools and the SATPM
Speaker 1: website to access these surface PM two point five data
Speaker 1: products for a region of interest and a time period
Speaker 1: of interest to you. Recall from Part one last week
Speaker 1: that PM two point five is a measure of the
Speaker 1: in situ mass concentrations of particles smaller than two point
Speaker 1: five microns and aerodynamic diameter, and that aerosol optical depth
Speaker 1: or AOD is an optical measurement of the atmospheric column
Speaker 1: aerosol loading, that is, the total presence of aerosols in
Speaker 1: the atmosphere from the surface up to the top of
Speaker 1: the atmosphere. AOD can be well related to PM two
Speaker 1: point five under certain conditions, but that relationship will do
Speaker 1: grade or break down if, for example, the aerosols are
Speaker 1: prevalent in the atmosphere well above the surface level, if
Speaker 1: coarser particles that has particles larger than two point five
Speaker 1: microns are dominant among the aerosols, if the aerosols are
Speaker 1: absorbing water, and if the humidity is variable and the
Speaker 1: particles absorb water under high humidity conditions. If the atmospheric
Speaker 1: or the surface conditions prevent reliable retrieval of AOD, for example,
Speaker 1: over highly reflective surfaces or under dense cloud cover, if
Speaker 1: the spatial resolution of the AOD does not capture the
Speaker 1: local PM two point five variability as measured by ground
Speaker 1: based PM two point five monitors, or if the timing
Speaker 1: of the PM two point five measurements does not well
Speaker 1: coincide with the satellite overpasses. As a reminder, if you
Speaker 1: have any questions during this part, please put them into
Speaker 1: the questions box and we'll address them all at the
Speaker 1: end of the webinar. You can put the questions in
Speaker 1: as you go. We'll collect those questions and those questions,
Speaker 1: along with the answers, will all be posted in a
Speaker 1: document which will be available on the training website about
Speaker 1: a week following today's training. So with that, I'll hand
Speaker 1: things over to doctor Aaron von Dongler to talk about
Speaker 1: geophysical and hybrid PM two point five estimation methods.
Speaker 2: Thanks Carl So. Last week we touched on using a
Speaker 2: simple linear regression between ground based observations of PN point
Speaker 2: five and satellite retrievals of aerostoptical depth as it means
Speaker 2: to use AOD to locally estimate P and point five concentrations.
Speaker 2: We saw that in certain cases this simple approach works
Speaker 2: really quite well, but in others there can be some
Speaker 2: severe limitations. For this part of the workshop, I want
Speaker 2: to focus on two other methods of using AOD to
Speaker 2: estimate P and T point five that specifically seek to
Speaker 2: either avoid the use of ground based observation altogether or
Speaker 2: at least ensure the information they provide me applied over
Speaker 2: a larger area. I wanted to start by showing the
Speaker 2: density of publicly available ground based P and twenty five
Speaker 2: observations around the world. Here of updated figure from Martin
Speaker 2: at Ull twenty nineteen showed the average population weighted distance
Speaker 2: to ground based monitors for countries around the world. You'll
Speaker 2: note that this average proximity can vary quite significantly, with
Speaker 2: populations in large regions in North America, Europe, and China
Speaker 2: is well a few others, typically residing within less than
Speaker 2: fifty kilometers in the nearest P and twenty five monitor.
Speaker 2: Populations within South Asia as a whole typically reside within
Speaker 2: seventy kilometers, Central Latin America within one hundred and thirty klomebers,
Speaker 2: and North Africa at least typically within a few hundred clombers.
Speaker 2: Even over a monitor dense region such as North America,
Speaker 2: the typical distance can vary significantly, from less than fifteen
Speaker 2: kilometers over western states to twenty five to fifty over
Speaker 2: more central ones. The net impact of all this with
Speaker 2: respect to estimates of P and twenty five from satellite
Speaker 2: is that any method that seeks to be globally applicable
Speaker 2: will need to either minimize its direct use of ground
Speaker 2: based observations or find an effective way to account for
Speaker 2: how the relationship used between AOD and P five changes
Speaker 2: with distance from the monitor locations used to constrain that relationship.
Speaker 2: One way to deal with this challenge is to remove
Speaker 2: the need for P and twoentty five bonders altogether. To
Speaker 2: do this, we need to revisit a relationship that was
Speaker 2: introduced to you last week that calculates AOD given P
Speaker 2: and twenty five concentration or set of assumptions, namely that
Speaker 2: skies are cloud free, the aerosols are well mixed, with
Speaker 2: none above the boundary layer, and last so that the
Speaker 2: aerosols present are optically homogeneous. That is similar, given these conditions,
Speaker 2: you can see that AOD is proportional to P and
Speaker 2: point five concentration itself as well as the boundary layer
Speaker 2: height and the extinction coefficient, which relates to how much
Speaker 2: the aerosol in question attenuates that it scatters or absorbs
Speaker 2: any passing light particle. Effective radius and density are inversely
Speaker 2: related with the potential for an additional effect from relative
Speaker 2: humidity if the aerosols hydrophilic. In order to broaden the
Speaker 2: applicability of this equation, we can relax a few of
Speaker 2: the initial assumptions and rearrange to solve for p and
Speaker 2: twoin five given the retrieval of AOD. Here, rather than
Speaker 2: assuming the aerosols are well mixed throughout the boundary layer,
Speaker 2: we're going to assume that we know what the fraction
Speaker 2: of AOD that occurs near the surface up to a
Speaker 2: heighth delta z and is associated with PAN twoin five.
Speaker 2: We also remove the assumption that there is no aerosol
Speaker 2: above the boundary layer and that all the aerosols are similar,
Speaker 2: instead just focusing on the qualities of those near surface aerosol.
Speaker 2: With the exception of this fractional term, the formula itself
Speaker 2: is quite similar, except that now the variables related to
Speaker 2: aerosol properties are no longer representative of the entire boundary layer,
Speaker 2: but again just to those aerosols that are near the
Speaker 2: surface within that delta z height. This equation shown again
Speaker 2: here forms the basis of the geophysical method to estimate
Speaker 2: PN twoiny five from ALI. It is theory based, physically driven,
Speaker 2: and completely independent of ground based P and twoin five observations.
Speaker 2: The challenge, however, is that it requires some assumptions or
Speaker 2: knowledge about the aerosols that are producing the retrieved AOD.
Speaker 2: This includes information on aerosol type which impacts a density,
Speaker 2: effective radius and extinction coefficient. It also needs an understanding
Speaker 2: as to how those aerosols respond to changes in relt
Speaker 2: of humidity. And lastly, it requires some knowledge with the
Speaker 2: vertical profile, that is, how much of the local airsol
Speaker 2: are located near the surface or again within that delta
Speaker 2: zt height. For simplicity, this large massive terms relating AODP
Speaker 2: and twenty five is typically reduced to a single scalar
Speaker 2: term called ATA. The challenge of the geophyscal method is
Speaker 2: how to best represent ata. This is where chemical transfer
Speaker 2: models play an essential role in most applications of the
Speaker 2: geophysical method. As a reminder, a CTM is a computer
Speaker 2: model that simulates atmosphere composition using the physical and chemical
Speaker 2: equations that govern the state of the atmosphere. It does
Speaker 2: this by inputting assimilated meteorology and natural and anthrogenic emissions
Speaker 2: into the model and allowing those equations to represent the
Speaker 2: atmosphere's chemical and dynamic response, as well as the transport
Speaker 2: and deposition of these individual chemical constituents or tracers. A
Speaker 2: CTM provides a complete representation of the atmosphere and in
Speaker 2: the context of geophyscal p point five, provides all the
Speaker 2: information we need to calculate ATA. Here we have an
Speaker 2: example of a CTM predicted ATA, in this case for
Speaker 2: the month of July twenty twenty three. The different colors
Speaker 2: are associated with the contributions of different aerosol types to ATA,
Speaker 2: with yellow fermeneral dust, blue for sea salt, green for
Speaker 2: organic aerosol, black for carbonaceous aerosol, and red for inorganics
Speaker 2: being sulfate, ammonium, and nitrate. The ATA values plotted here
Speaker 2: represents the relationship to instantaneous rather than twenty four hour
Speaker 2: Pene point five, and you can very clearly see a
Speaker 2: pulsing and intensity that circles the globe with a diurnal rhythm.
Speaker 2: This corresponds to the impact of solar heating on bound
Speaker 2: do lair heights, essentially diluting concentrations at the surface as
Speaker 2: the bound dolair increases with rising temperatures associated daylight hours,
Speaker 2: which in turn alters the vertical structure to decrease the
Speaker 2: value of ATA. Also visible are the regional differences in
Speaker 2: the particular aerosol sources and types driving the AOD to
Speaker 2: Pine point five relationship, such as dust emissions over the
Speaker 2: Sahara organic carbon VIEO biomass burning were parts of South
Speaker 2: America and Central Africa, or anthropogenic inergantic aerosol over China,
Speaker 2: Europe and North America. You can even see the impact
Speaker 2: of the long range transport of dust and biomass burning plumes,
Speaker 2: where ATA essentially drops to zero, implying that the aerosols
Speaker 2: almost exclusively aloft and therefore the relationship between AOD and
Speaker 2: P twointy five is effectively removed. As is pretty clear
Speaker 2: from this example, using a CTM to represent ATA has
Speaker 2: some definis advantages. You can see how the spatial and
Speaker 2: vertical variability impacting ATA is captured via the model's response
Speaker 2: to emissions. In meteorology, temporal variability is also captured which
Speaker 2: can have advantages when compensating for sellid overpast times and sampling.
Speaker 2: And also quite clear is the global applicability of this
Speaker 2: approach when compared to limiting yourself to justice locations that
Speaker 2: have dense ground based observations. Counter to these advantages, of course,
Speaker 2: CTM inherently has limitations to its resolution, whether by code,
Speaker 2: structure or resource limitations, and most global models are typically
Speaker 2: about it a resolution on the order of one hundred
Speaker 2: klometers or so. Any model is also, of course, only
Speaker 2: as good as the inputs that go into it, and
Speaker 2: regional uncertainty in emission, inventories or meteorology will impact the
Speaker 2: quality of the simulation. It should also be noted that
Speaker 2: while these simulations can be quite effective at representing general features,
Speaker 2: individual events are much more of a challenge and often
Speaker 2: limit it to the most effective applications of a simulated
Speaker 2: eight to two longer term means. And lastly, even at
Speaker 2: one hundred klumeter resolution, these simulations are quite computationally intensive
Speaker 2: to run, which can limit an individual's ability to run
Speaker 2: these simulations themselves and require a dependency on whatever simulation
Speaker 2: output is already available via other sources. All that said,
Speaker 2: a CTM based simulation of ATA is often an effective
Speaker 2: means to quantify how AOD relates to PM two point five,
Speaker 2: especially as the distance to the nearest ground based monitor increases.
Speaker 2: Here you can see an example over Eastern Asia, primarily China,
Speaker 2: for annual mean values during twenty twenty one. The left
Speaker 2: hand column shows AOD, whereas the right hand column shows
Speaker 2: the corresponding geophysical PM two point five. Black dots show
Speaker 2: the locations of the ground based monitors used for the
Speaker 2: comparisons in the lower row. In this example, you can
Speaker 2: see quite clearly that AOD is well related directly to
Speaker 2: P and TOY five concentrations over this region with an
Speaker 2: R squared of a round point four. Converting this AOD
Speaker 2: to PM two point five using a simulated ATA not
Speaker 2: only allows a direct interpretation of this AOD as PM
Speaker 2: twoint five, but additionally improves the R squared to approximately
Speaker 2: point five. This example nicely shows some of the strengths
Speaker 2: as well as from the challenges of this geophysical approach.
Speaker 2: You can see how major spatial features in PM twoint
Speaker 2: five are captured and even improved upon compared to AOD alone,
Speaker 2: and unlike AOD, the air quality implications can be more
Speaker 2: directly interpreted. On the side of challenges, there is potential
Speaker 2: for bias and scatter, which is difficult to evaluate without
Speaker 2: some level of ground based observations. As a result, purely
Speaker 2: geophysical estimates are typically best used for large spatial and
Speaker 2: temporal scales. Of course, ground based observations do not strictly
Speaker 2: have to be used for only evaluation, and can also
Speaker 2: be to understand any bias is present in the geophysical values.
Speaker 2: This idea forms the basis of what is known as
Speaker 2: hybrid PM tw point five or geophysical hybrid PM two
Speaker 2: point five. Hybrid PM two point five essentially builds on
Speaker 2: geophysical PM two point five using ground based observations. To
Speaker 2: give an example, here, i've shown the difference between ground
Speaker 2: based observations of PM two point five and coincidently sampled
Speaker 2: geophysical estimates. As you look at this figure, you'll begin
Speaker 2: to notice regional structures and features within these differences, such
Speaker 2: as a large scale underestimate of a few microgrants per
Speaker 2: cubic meter on the part of the geophysical values in
Speaker 2: the southeastern US, and a large scale overestimate of the
Speaker 2: same magnitude over the northeastern Great Lakes region. We don't
Speaker 2: see these large scale features over western regions, but rather
Speaker 2: we see a high degree of variability between seemingly neighboring
Speaker 2: monitor locations. Hybrid methods therefore, look at these features using
Speaker 2: statistical models and training to predict biases in geophysical PM
Speaker 2: two point five. In other words, these hybrid methods predict
Speaker 2: limitations associated with AOD retrievals and ATA simulations and their
Speaker 2: geophysical implications. So how do we do that well. Multiple
Speaker 2: statistical techniques can be used to make these associations, but
Speaker 2: perhaps the most intuitive approach uses multiple linear aggression. From
Speaker 2: last week's session and from other work, most of you
Speaker 2: will be quite familiar with single linear aggression, such as
Speaker 2: the example shown in the upper right. Here, a line
Speaker 2: of best fit is created between a single independent variable
Speaker 2: X one and some other desired dependant variable, why with
Speaker 2: a slope of beta one. One way to think of
Speaker 2: this result is that beta one shows the response of
Speaker 2: Y to changes in X one. Multiple linear aggression is
Speaker 2: an extension of single linear aggression, where as the name implies,
Speaker 2: multiple dependent variables are used. Rather than fitting a line
Speaker 2: of best fit, you're fitting a plane. But the concept
Speaker 2: is very similar and the result produces a series of
Speaker 2: predictor coefficients denoted by betas in this case, that can
Speaker 2: be interpreted as the linear response of your dependent variable
Speaker 2: to each independent predictor variable. For some systems, these beta
Speaker 2: predictor coefficients are a universal truth, that is, they don't
Speaker 2: change in space or time, but for our particular application
Speaker 2: estimating the bias in geophysical PM twoenty five, this is
Speaker 2: unlikely to be the case. As a result, an extension
Speaker 2: of multiple linear aggression called geographically weighted regression or GWR,
Speaker 2: is recommended. In effect, GWR allows the predictor coefficients from
Speaker 2: MLR to vary in space. In the equation shown here,
Speaker 2: this is denoted by the addition of the I and
Speaker 2: J subscripts, which are simply identifying different locations on a
Speaker 2: grid such as the one shown on the right here.
Speaker 2: In practice, the spatially unique coefficients are calculated by running
Speaker 2: a series of MLR regressions, one for each location on
Speaker 2: the grid, where the data points used for each regression
Speaker 2: are modified based on their local significance. This can either
Speaker 2: be through some sort of waiting scheme associated with the
Speaker 2: data points during regression, or their removal entirely if they
Speaker 2: are not relevant to the location at hand. It's worth
Speaker 2: noting also that while grids can be regularly spaced, such
Speaker 2: as I've shown here, they can also be irregular in nature.
Speaker 2: An irregular grid can be advantageous to capture changes in
Speaker 2: these relationships in regions where we either expect rapid relational
Speaker 2: changes in beta or have additional training capacity due to
Speaker 2: more ground based monitors, or alternatively reducing the computational costs
Speaker 2: associated with running on a fine grid over regions where
Speaker 2: neither of these conditions are expected. Okay, so what variables
Speaker 2: are going to be most effective at predicting the geophysical bias? Well,
Speaker 2: the ants would really tie back into one of two sources,
Speaker 2: either uncertainties in the retrieval of AOD itself or uncertainties
Speaker 2: in the relationship that's being used to relate AOD to
Speaker 2: PM two point five. Between this week and last, we've
Speaker 2: covered a lot of the sources of this potential uncertainty,
Speaker 2: which typically falls into four broad categories. Firstly, composition, Different
Speaker 2: aerosol types can be associated with different artical properties that
Speaker 2: may challenge AD retrievals, but probably more impactful in this
Speaker 2: context is the associations they can have with different emission sources,
Speaker 2: vertical profiles, and growth factors, all of which impact the
Speaker 2: AOD to PM two point five relationship. In this way,
Speaker 2: the presence of a particular aerosol component, such as mineral dust,
Speaker 2: for example, could be insightful due to uncertainties in the
Speaker 2: simulated representation of its emission transport, which in turn have
Speaker 2: the potential to bias ATA. Similarly, meteorological inputs also feedback
Speaker 2: on ATA within a chemical transport model, with local winds
Speaker 2: and temperature impacting several natural emission sources for example, or
Speaker 2: atmospheric conditions impacting the chemical production of aerosol. Land type
Speaker 2: information can also be related to emission type uncertainties, but
Speaker 2: primarily I tend to think of this sort of predictor
Speaker 2: as being associated with uncertainties in the AOD retrieval itself.
Speaker 2: The dominant assumptions in satellite aidor retrievals typically relate to
Speaker 2: the amount of light that is reflected off the Earth's surface,
Speaker 2: which naturally is strongly connected to the type of surface
Speaker 2: that is doing the reflecting. Urban and desert biases, for example,
Speaker 2: have both been documented and improved upon in AD retrievals
Speaker 2: over the years, but surface type remains a solid predictor
Speaker 2: of some of the challenge is facing an AOD retrieval,
Speaker 2: and any bias in AOD will directly carry forward into
Speaker 2: a bias in geophysical PM two point five. Lastly, elevation
Speaker 2: or more specifically, changes in elevation, can be an important
Speaker 2: predictor in geophysical bias. In this case, what is being
Speaker 2: represented is not so much of an error on the
Speaker 2: part of a model to simulate ATA, but rather changes
Speaker 2: in the AOD to PM point five relationship that are
Speaker 2: occurring at a resolution below that of our simulation. You'll
Speaker 2: recall that simulations are often run at fifty or one
Speaker 2: hundred klomebers, a mountain or valley can often occur at
Speaker 2: scales well below this, with obvious implications for how AOD
Speaker 2: relates to PM two point five. In essence, a predictor
Speaker 2: such as elevation changes is not so much accounting for
Speaker 2: errors in a model which could be correctly represent representing
Speaker 2: the average conditions at the simulated resolution, but still need
Speaker 2: adjustments to account for the impactive features that occur below
Speaker 2: the model's resolution. Bringing this all together, we end up
Speaker 2: with a GWR based equation such as you see here,
Speaker 2: where the difference between the ground based observations of P
Speaker 2: and point five and their coinstantly sampled geophysical estimates are
Speaker 2: locally regressed against a series of predicted variables associated with
Speaker 2: simulated composition neear logical parameters, elevation parameters, and land type parameters,
Speaker 2: and below you can see an example of the results
Speaker 2: that can be obtained. On the left we have the
Speaker 2: original biases that were observed against the ground based observations,
Speaker 2: and on the right we see the total GWR predicted
Speaker 2: bias in the geophysical PN point five. You can see
Speaker 2: other This approach captures the major features we discussed earlier,
Speaker 2: such as the overestimate in the Great Lakes region, but
Speaker 2: beyond that, it predicts that this feature extends further north
Speaker 2: and west into parts of the Canadian Prairies, a region
Speaker 2: with far less ground based constraints. One of the real
Speaker 2: strengths this GWR technique for predicting biases compared to some
Speaker 2: recent and arguably more advanced machine learning techniques, is a
Speaker 2: relative ease of understanding the impact of individual variables and
Speaker 2: how they shape the overall bias prediction. For example, I've
Speaker 2: plotted here both the predictor coefficients and that impact of
Speaker 2: the nitrate and percent urban variables on the overall predicted bias.
Speaker 2: As you can see in this case, nitrate large explains
Speaker 2: give shape to that Great Lakes Region over estimate, as
Speaker 2: well as its extension to under monitored regions. This opens
Speaker 2: up a question about what it is about nitrate that
Speaker 2: produces this bias, in this case potentially related to uncertainties
Speaker 2: in either the chemical reactions of all the local emissions.
Speaker 2: But knowing these connections can in turn direct developments and
Speaker 2: improvements within the CTM, which can then feed back into
Speaker 2: the original geophysical p in twenty five with potentially global benefits.
Speaker 2: You can also see the advantage of allowing predictor coefficients
Speaker 2: to vary in space. Again shown the left hand panels.
Speaker 2: Here is a very clear East West divide in the
Speaker 2: strength of association between urban content and predicted geophysical bias,
Speaker 2: with the coefficient approximately tripling over western North America compared
Speaker 2: to the east. Given the fine spatial scale at which
Speaker 2: present urban impacts hybrid PM twoint five, this may suggest
Speaker 2: something regionally unique in the way urban surfaces impact the
Speaker 2: AOD retrieval, or it could be a resolution impact on
Speaker 2: ATA where the urban environment impacts its relationship on a
Speaker 2: scale below that of a simulation more strongly in the west.
Speaker 2: In either event. Putting the impact of all these predicted
Speaker 2: variables together, we can demonstrate how hybrid PM two point
Speaker 2: five offers some significant gains compared to geophysical PM two
Speaker 2: point five alone. Here I've shown an overall global map
Speaker 2: of PM two point five provided by the hybrid methodology
Speaker 2: as as well as the improved agreement it offers, with
Speaker 2: an R squared increasing to zero point nine from approximate
Speaker 2: point seven, and that root means score difference decreasing by
Speaker 2: about a factor of two. One of the big strengths
Speaker 2: of this approach is that it builds on the fundamental
Speaker 2: strength of the geophysical method while simultaneously taking advantage of
Speaker 2: the constraint provided by ground based observations. There are, of course, limitations,
Speaker 2: and first and foremost, unlike purely geophysical estimates, we now
Speaker 2: require ground based observations, and the challenge becomes how to
Speaker 2: incorporate those measurements in a way that is most representative
Speaker 2: of your particular region of interest. Related to this, setting
Speaker 2: up a model such as GWR or alternatively, machine learning
Speaker 2: based algorithms like a convolutional know network or Ingredient boosting
Speaker 2: can be quite complicated, including but not limited to the
Speaker 2: selection of appropriate predictor variables. Lastly, evaluation of these models,
Speaker 2: like any statistical models, can be quite challenging, and to
Speaker 2: that point, appropriate cross validation is key. It's relatively easy
Speaker 2: to create a model which is very high agreement against
Speaker 2: the data points and locations that we're used to train it.
Speaker 2: It can be much more challenging to determine how well
Speaker 2: your model performs at other locations and times. With this
Speaker 2: in mind, I'm going to spend a few minutes touching
Speaker 2: on a few different methods of cross validation. Cross validation
Speaker 2: is a technique that's used to understand the uncertainty of
Speaker 2: statistical models. It involves the creation of a subset of
Speaker 2: data points that can be used for independent evaluation of
Speaker 2: that model. The idea being that they provide insight into
Speaker 2: how well a model will perform a part or away
Speaker 2: from the data that has been used to train it.
Speaker 2: Cross validation can be time consuming and realistically often has
Speaker 2: to balance its complexity, accuracy and representation. Above all, the
Speaker 2: cross validation data set must maintain what I would call
Speaker 2: true independence. The nature of which may depend on the
Speaker 2: questions being asked. For applications making use of the resulting model,
Speaker 2: The most basic form of cross validation is what is
Speaker 2: referred to as random cafold cross validation. Here, we randomly
Speaker 2: withhold the set fraction often ten percent of the data
Speaker 2: set from model training and use the remaining ninety percent
Speaker 2: for training purposes. We do this multiple times until we've
Speaker 2: withheld a sufficient amount of the original data set from
Speaker 2: at least one of the cross validation models, and use
Speaker 2: the cross validation models predictions of the withheld data points
Speaker 2: to evaluate the model quality. An advantage of this approach
Speaker 2: is that it's fairly simple to understand and implement. Is
Speaker 2: also perfectly suitable to many situations where the points in
Speaker 2: the data set being modeled are themselves fully independent of
Speaker 2: one another. For our particular application, those last point is
Speaker 2: often not the case, with ground based monitors often being
Speaker 2: clustered around urban centers and the concentration of PM two
Speaker 2: point five present at one station being very much connected
Speaker 2: to the concentration that we might expect at another. Buffered
Speaker 2: cross validation methods have been developed specifically to address this
Speaker 2: last point. The concept is that when a particular data
Speaker 2: point where station is withheld for cross validation. Not only
Speaker 2: is that point withheld for model training, but also any
Speaker 2: sites located within a given spatial distance. As with kfold
Speaker 2: cross validation, multiple models are trained and used to evaluate uncertainty,
Speaker 2: with the distinction being that the sites excluded from each
Speaker 2: cross validation model are not only the cross validation evaluation sites,
Speaker 2: but also those within the spatial buffer. The initial implementation
Speaker 2: of this approach was called buffered leave one out or
Speaker 2: blue cross validation. It essentially withheld one side at a
Speaker 2: time pluss buffer points running through each data point in
Speaker 2: the overall training data set. This is effective, but as
Speaker 2: you can imagine, is also computationally very expensive, as you
Speaker 2: need to train a unique cross validation model for each
Speaker 2: point in your data set. Buffer leave cluster out, and
Speaker 2: buffer leave isolated sites and cluster out methodologies were developed
Speaker 2: to combine the improved independence of the blue methodology with
Speaker 2: improved computational efficiency. Specifically, these methods withhold groups of nearly
Speaker 2: located sites plus or overlapping buffers out of each cross
Speaker 2: validation fold, in essence running multiple Blue like models simultaneously,
Speaker 2: thereby providing similar information. Is Blue both far less training
Speaker 2: of the individual models. The last type of cross validation
Speaker 2: that I wanted to highlight is temporal holdback, which is
Speaker 2: used to address a slightly different question. In this case,
Speaker 2: the focus of the model being developed is not to
Speaker 2: represent missing values at different locations, but rather at the
Speaker 2: same locations as the training data set, just for different
Speaker 2: time periods. As before, the goal here is to separate
Speaker 2: a set of data from the training of the model
Speaker 2: that is independent in the specific way that the model
Speaker 2: is to be used. In the case of a temporarily
Speaker 2: focused model, this requires the removal of specific blocks of
Speaker 2: time from the training process, which can then be used
Speaker 2: for subsequent validation. This approach does not inform both spatial
Speaker 2: uncertainties and should not be used in that way, but
Speaker 2: is quite insightful for specific temporal applications and can be
Speaker 2: combined with a spacely buffer analysis to provide a complete
Speaker 2: representation of model uncertainties. Okay, So, in summary, geophysical methods
Speaker 2: to estimate PM two point five avoid the use of
Speaker 2: ground based monitors by using the theory based physically driven
Speaker 2: relationship between total column aerosoloptical depth and PM two point five.
Speaker 2: CTMs are invaluable for this approach as they are able
Speaker 2: to represent all the necessary parameters and provide global coverage.
Speaker 2: Hybrid methods build on the geophysical methods using a statistical
Speaker 2: framework to predict the bias in geophysical PM two point five,
Speaker 2: and Lastly, effective cross validation of these or any models
Speaker 2: must consider how to ensure strong independence of the validation
Speaker 2: points that are being used. And now we're going to
Speaker 2: switch gears a little bit and focus specifically on the
Speaker 2: sat PM two point five data set produced at Washington
Speaker 2: University in Saint Louis by the Atmospheric Compositional Analysis Group
Speaker 2: or AGAC. The sat PM two point five data product
Speaker 2: is the culmination of fifteen to twenty years of research
Speaker 2: using aerosol optical depth to estimate near surface PM two
Speaker 2: point five concentrations. The goal of this long running project
Speaker 2: is to provide a consistent global, long term PM two
Speaker 2: point five data set, and its current iteration is observationally
Speaker 2: construc using both satellite AOD and ground based observations within
Speaker 2: the hybrid geophysical framework that we've been discussing is monthly
Speaker 2: with regular updates provided annually and at a fairly high
Speaker 2: resolution on a zero point zero one degree grid, which
Speaker 2: is about one kilometer. A key feature of this data
Speaker 2: set from its inception is that it is publicly available
Speaker 2: currently with its own dedicated website at SATPM dot org.
Speaker 2: On the right you can see its basic development structure,
Speaker 2: again following the hybrid geophysical framework, where AOD are converted
Speaker 2: to geophysical PM twointy five using a simulated representation of ATA,
Speaker 2: and these values are refined using a hybrid approach. At present,
Speaker 2: we do maintain a geographically weighted regression type algorithm such
Speaker 2: as we've been discussing, but primarily we encourage users to
Speaker 2: make use of a more current algorithm that makes use
Speaker 2: of a more modern machine learning method convolutional neural networks.
Speaker 2: I'm going to take a few minutes and address some
Speaker 2: of the unique challenges in setting up a full fledged
Speaker 2: hybrid geophysical PM two point five algorithm and how we've
Speaker 2: addressed these within the sat PM data set. Geophysical PM
Speaker 2: twoint five all starts with AOD, So the first questioning
Speaker 2: phase is which AOD retrieval is best suited for PM
Speaker 2: two point five estimation. The answer, of course can varya
Speaker 2: both space and time depending on the specific conditions. So
Speaker 2: to give you a better idea, of how AOD retrievals
Speaker 2: can vary. I've plotted here four separate AOD retrievals, all
Speaker 2: monthly means of available data for April twenty fourteen, taken
Speaker 2: from the same satellite platform. Two instruments represented MODUS and MISER,
Speaker 2: with the dark Target, Deep blue and MAIC retrievals being
Speaker 2: applied to the Modus instrument and the MISER instrument having
Speaker 2: its own specific retrieval algorithm. What will strike you is
Speaker 2: how in some ways these retrievals are also similar, but
Speaker 2: in other ways they look quite different. In mind, these
Speaker 2: are all good, high quality retrievals. Some differences are by design.
Speaker 2: Dark Target, for example, was never intended to retrieve AOD
Speaker 2: over deserts, and so data over such regions is simply
Speaker 2: missing from this retrieval. Deep Blues development initially focused on
Speaker 2: this very type of surface and well captures the features
Speaker 2: we might expect over such regions. MEEK specializes in complex terrain,
Speaker 2: using multiple overpasses to build up a representation of surface reflectants,
Speaker 2: something which MISER can do in a single overpass, but
Speaker 2: with far less spatial coverage each time, which results in
Speaker 2: almost stuttered look for MISER. Despite its high quality of
Speaker 2: retrieval owing to its reduced sampling of the monthly means
Speaker 2: shown in all cases, you can see data missing from
Speaker 2: the far North, where snow cover inhibits the use of
Speaker 2: any of these methods. So what can we do? We
Speaker 2: have multiple satellite instruments with multiple AOI retrievals, all which
Speaker 2: contain useful information, but not necessarily equally under all conditions.
Speaker 2: In our case, we turned to airnet. Aaronnet was briefly
Speaker 2: introduced last week, and for our purposes, all we really
Speaker 2: need to know is that aaronet is a ground based
Speaker 2: network of sun photometers that provides accurate long term measurements
Speaker 2: of AOD at different locations around the world. This allows
Speaker 2: the evaluation of satellite based AOD, but only at locations
Speaker 2: that have an AARNet site. So the question becomes, how
Speaker 2: can we extend this site specific information globally to understand
Speaker 2: the quality of a satellite retrieval anywhere in the world.
Speaker 2: What we've done is group AARNet locations together and perform
Speaker 2: separate month specific comparisons for each satellite instrument and or retrieval.
Speaker 2: These comparisons are grouped by land type DDI or normalized difference,
Speaker 2: vegetation index, and weighted by distance. We combine the findings
Speaker 2: of these comparisons based on local conditions around the world
Speaker 2: to produce a contents consistent definition of uncertainty for each
Speaker 2: AOD data source. You can see an example of this
Speaker 2: on the right, where I've shown the normalized root means
Speaker 2: square difference for Dark Target. Based on this approach. For
Speaker 2: those of you familiar with this retrieval, it shows a
Speaker 2: number of features we might expect, such as higher uncertainties
Speaker 2: over the Western US, where the surface tends to be brighter,
Speaker 2: presenting a challenge to some of dark Target's inherent assumptions,
Speaker 2: compared to the Eastern US, where richly vegetated surfaces align
Speaker 2: very well with those same assumptions. Once we've completed these
Speaker 2: comparisons for different AOD data sources, we need a way
Speaker 2: to bring this information together so that we locally rely
Speaker 2: on each data set relative to its accuracy at any
Speaker 2: given location. To do this, we employ the relationship shown
Speaker 2: in the lower right to represent each retrieval's overall uncertainty
Speaker 2: compared to the total uncertainty of all the other AOD
Speaker 2: data sources combined. Here are location's final AOD is determined
Speaker 2: as the weighted average of all available AOD data sources,
Speaker 2: with each one's weight based on a combination of its
Speaker 2: local normalized routing square difference. Bias is represented by the
Speaker 2: slope of the line of best fit and its availability.
Speaker 2: This equation can be evaluated anywhere in the world, which
Speaker 2: gives us the overall weighting factors for each retrieval shown here,
Speaker 2: which provide a measure of the relative local amount of
Speaker 2: uncertainty for each AOD retrieval. What's encouraging about these results
Speaker 2: is that these spatial features make intuitive sense based on
Speaker 2: an understanding of each AOD data source. You can see,
Speaker 2: for example, dark target contributing heavily over more vegetative regions,
Speaker 2: deep blue over many of the so called brighter surfaces,
Speaker 2: and may contributing most over some of the more complicated conditions.
Speaker 2: We've also brought in CTM simulated AOD as an additional
Speaker 2: AOD data source, which contributes predominantly over those regions where
Speaker 2: satellite based retrievals cannot find function, such as snow covered
Speaker 2: regions in the far North. Here, simulated AOD uncertainty follows
Speaker 2: a similar concept of comparisons against airnet, but rather it
Speaker 2: uses categories based on speciation and elevation, which are more
Speaker 2: globally connected to CTM uncertainty than land surface types in NDBI,
Speaker 2: and bringing this all together we can produce an overall
Speaker 2: combined AOD such as shown here, which takes advantage of
Speaker 2: each data set where it is most accurate. An additional
Speaker 2: challenge to bringing together multiple AOD data sets is that
Speaker 2: often there are different spatial scales or time periods. Here
Speaker 2: I have listed the satellites of retrievals as well as
Speaker 2: the model we currently use in our sat PM data set.
Speaker 2: You'll note that some of the earlier retrievals available from
Speaker 2: seaweeds using deep Blue are available from nineteen ninety eight
Speaker 2: to twenty ten at a resolution of fourteen kilometers, whereas
Speaker 2: the most recent editions include a veer's instrument starting in
Speaker 2: twenty eighteen at resolutions ranging from one to six kilometers. Gs.
Speaker 2: Chemical transfer model, for its part, is run at a
Speaker 2: resolution about fifty kilometers throughout this whole time period. One
Speaker 2: option to combine all these sources would be to simply
Speaker 2: interplate each source down to the finest one klometer resolution
Speaker 2: and combined, but this runs the risk of washing out
Speaker 2: fine scale features when relatively coarse AOD sources are weighted
Speaker 2: against truly fine ones. To address this possibility, we combine
Speaker 2: these data sets in a tiered fashion, first averaging all
Speaker 2: data sets onto the coursest fifty klounder grid. We then
Speaker 2: apply the ten kilometer relative variation in AOD within that
Speaker 2: fifty klomter grid according to the satellite based sources. And lastly,
Speaker 2: we apply the one clometer relative variation within the ten
Speaker 2: clumdra grid according to MAIC algorithm, which is the only
Speaker 2: satellite retrieval providing information at such a fine spatial scale. Finally,
Speaker 2: we linearly interprelate the simulated AOD to PM twenty five
Speaker 2: relationship down to the one cloonder grid which assumes smooth
Speaker 2: variation and ATA, and apply it to the one cloonder AOD.
Speaker 2: As I've stressed throughout this presentation, the geophysical method relies
Speaker 2: heavily on the inclusion of a chemical transport model. The
Speaker 2: Atmospheric Compositional Analysis Group is fortunate to be a proud
Speaker 2: part of the GSM community. GSM is an advanced chemical
Speaker 2: transport model with a specific mission to advance the understanding
Speaker 2: of human and natural influences on the environment through a
Speaker 2: comprehensive state of the science, readily accessible Global Model of
Speaker 2: Atmospheric Composition. It is continually developed and maintained by an
Speaker 2: international host of scientists, many of whom are shown in
Speaker 2: this photograph taken two years ago during the biannual GSM meeting,
Speaker 2: and from the map below, which shows locations of various
Speaker 2: GSM contributors for our particular purposes, representing the aerosol driven
Speaker 2: geophysical relationship between AOD and m PO point five. It
Speaker 2: is particularly relevant that GSM is open source and community driven.
Speaker 2: This allows us to include developments specific to our needs,
Speaker 2: such as recent developments and mineral dust emissions, or particular
Speaker 2: diagnostics that are well suited to ADA itself. GSM also
Speaker 2: boasts a detailed aerosol oxiden model representing the inorganic sulfate
Speaker 2: nitrate ammonium system. Carbonaceous aerosols and natural sources such as
Speaker 2: mineral dust and sea salts. Emissions are up to date
Speaker 2: and include meteorologically driven sources such as mineral dust and
Speaker 2: biogenic VOCs, as well as high resolution anthroogenic inventories such
Speaker 2: as SAIDs. Lastly, biomass spurning is informed using satellite observations
Speaker 2: using inventories such as GFAs and g FED. The whole
Speaker 2: system is driven using assimilated meteorology with variable resolution, but
Speaker 2: in our case we typically run at about fifty kilometers
Speaker 2: hybrid PM two point five, as you're aware, depends not
Speaker 2: only on satellite retrievals and CTM simulations, but requires high
Speaker 2: quality ground based observations. At AGAC, we actually maintain a
Speaker 2: database covering approximately fifteen thousand monitor locations around the world.
Speaker 2: The coverage of this database continues to grow, as you
Speaker 2: can see from the coverage map shown here, with different
Speaker 2: colors representing ground based monitors that were brought in during
Speaker 2: successive SAPPM releases. This database predominantly relies on government based
Speaker 2: networks and is limited to high quality monitors as opposed
Speaker 2: to low cost sensors which would complicate the usage in
Speaker 2: its context. Most of the data sources are publicly available,
Speaker 2: but there are exceptions which are obtained through direct contact.
Speaker 2: I've listed the various countries and sources of the monitors
Speaker 2: we use. This includes data from open AQ, which is
Speaker 2: an online nonprofit specifically devoted to the aggregation and harmonization
Speaker 2: of open access air quality data, but additionally, as I mentioned,
Speaker 2: direct access to many government environmental agencies. Akin to the
Speaker 2: challenge we face combining multiple AOD sources over a long
Speaker 2: time period, ground based observations are not temporarily complete. That is,
Speaker 2: there a very few ground based monitors that have been
Speaker 2: in continuous operation at the same location for the past
Speaker 2: twenty five years. As a result, changes to the distribution
Speaker 2: of monitors over time has the potential to weight the
Speaker 2: adjustments inferred from these monitors towards regions with the longest
Speaker 2: histories of monitors or worse introduced temporal changes in these
Speaker 2: adjustments that are based on changes to the distribution of
Speaker 2: available monitors over time rather than a meaningful representation of
Speaker 2: how the impact of bias itself has changed. To help
Speaker 2: limit this potential impact, we produce a temporarily complete PM
Speaker 2: two point five record at all ground based monitors prior
Speaker 2: to incorporating them into the hybrid framework. To complete this task,
Speaker 2: missing ground based values are inferred using relationships developed during
Speaker 2: overlapping periods with other available data sources. These other sources
Speaker 2: include not only geophysical and simulated PIN twoin five, but
Speaker 2: also other ground based observations from nearby or related stations.
Speaker 2: Multiple techniques are used to fill in this missing data.
Speaker 2: All are compared using a temporal holdback, and then combined
Speaker 2: based on the relative uncertainties. You can see the overall
Speaker 2: agreement between the inferred and observed P and twoint five
Speaker 2: and the lower scatterplots shown here, with a left hand
Speaker 2: panel showing the agreement when other ground based monitors are
Speaker 2: included within the inference calculations, compared to the right hand panel,
Speaker 2: which showed the agreement when based strictly on geophysical and
Speaker 2: simulated P and Twoiny five. A couple sample time series
Speaker 2: are shown above, with the red line showing the directly
Speaker 2: observed concentrations, the blue line showing the inferred values that
Speaker 2: are benefited from other local monitors, and the cyan color
Speaker 2: showing inferred values where no other local monitors were available.
Speaker 2: This representation of missing observation obviously adds some of its
Speaker 2: own certainties, but is overall very much a net benefit
Speaker 2: for during long term consistency and applying the insight gain
Speaker 2: from more recent heavily monitored time periods to earlier years.
Speaker 2: For the actual hybrid calculations, our most recent algorithms make
Speaker 2: use of a state of the science convolutional neural network
Speaker 2: or CNN. These types of machine learning algorithms are specifically
Speaker 2: designed to identify and learn from images, and in this
Speaker 2: way are highly effective for spatial tasks such as the
Speaker 2: relationship of geophysical PM twointy five with surrounding predictor variables.
Speaker 2: We train models separately for each month, which allows us
Speaker 2: to capture the temporal variability that we would expect relationships
Speaker 2: between the predictors and the geophysical bias. We additionally maintain
Speaker 2: an earlier GWR based algorithm, but this is a little
Speaker 2: bit more of a niche market. It's quite useful for
Speaker 2: a few particular applications, but in general we recommend the
Speaker 2: CNN based version. Six. Predictor variables are in line with
Speaker 2: the GWR discussion we had previously, including satellite based products,
Speaker 2: CTM simulations and related factors, emissions, meteorology, monitor location details,
Speaker 2: and geological type of information such as land type and elevation.
Speaker 2: We spend a lot of time developing cross validation strategies
Speaker 2: to help us understand the impacts of and ensure an
Speaker 2: effective model away from ground based monitor locations. These monitors
Speaker 2: provide an invaluable source of information and constraint, but they
Speaker 2: also run the risk of fooling us into misrepresenting or
Speaker 2: misunderstanding the qualities of our data set, especially as we
Speaker 2: get farther away from these training locations. On the left,
Speaker 2: here i've show on the effect on our squared of
Speaker 2: changes to buffer radius during our buffered cross validation analysis.
Speaker 2: In effect, this can be interpreted is how the agreement
Speaker 2: of hybrid PM two point five changes with distance from
Speaker 2: the nearest ground based monitor. For a series of differency
Speaker 2: and N models that were created during the development of
Speaker 2: our version six SAT PMD data set, the agreement of
Speaker 2: a purely geophysical PM twoenty five estimate shown in red,
Speaker 2: is independent of monitor distance, as shown the flat straight line. Interestingly,
Speaker 2: all the hybrid models we developed performed similarly at the
Speaker 2: monitor locations themselves if a buffer were included. However, models
Speaker 2: that excluded information from a CTM that did not include
Speaker 2: a physically and chemically meaningful process based understanding the atmosphere
Speaker 2: fell below the quality of a purely geophysical approach within
Speaker 2: about fifty kilometers. More advanced models that did include predictors
Speaker 2: based on CTM output fared better, but again within about
Speaker 2: one hundred two one hundred and fifty klometers underperformed the
Speaker 2: results provided by the geophysical PM twoenty five. Our CNN
Speaker 2: model has been optimized in a way that recognizes the
Speaker 2: relative strength of the purely geophysical estimates with increasing distance
Speaker 2: to the training sites. This allows improved agreement compared to
Speaker 2: the geophysical lessons alone, even at great distance from modern locations.
Speaker 2: Another way to look at this result is through what's
Speaker 2: called a shaply additive explanation or SHAP analysis. I'm not
Speaker 2: going to go through the details of this approach here,
Speaker 2: but only to say that a SHAP analysis is a
Speaker 2: method to rank the relative importance of individual predictor variables
Speaker 2: within a machine learning algorithm. I've shown such a result
Speaker 2: on the right, with the variables again ranked by their
Speaker 2: impact on the results of our hybrid CNN. As you
Speaker 2: can see, the process based predictors and the geophysical PM
Speaker 2: two pint five itself rank as the most influential predictors,
Speaker 2: again highlighting the limitations that would be present in a
Speaker 2: more purely statistical approach that didn't build upon a physically
Speaker 2: sound framework. And lastly, the results here these figures are
Speaker 2: briefly cycling through the global monthly hybrid geophysical PM two
Speaker 2: point five data set that were how to produce and
Speaker 2: please to make publicly available. The evaluations you see are
Speaker 2: against the highest possible standards, being both temporarily and spatially buffered.
Speaker 2: The sat PM two point five data set is particularly
Speaker 2: well suited for users that want to include the impact
Speaker 2: of features at a high space resolution and want to
Speaker 2: ensure consistency both over a long time period but also
Speaker 2: a distance from ground based constraints. So, in summary, AGAX
Speaker 2: sat PM two point five offers a global PM two
Speaker 2: point five data set that uses the geophysical hybrid methodology
Speaker 2: with multiple satellite based AOD products, an advanced chemical transfer model,
Speaker 2: and thousands of ground based observations, all within the state
Speaker 2: of the science convolutional neural network. We focus on consistency
Speaker 2: and quality, and we're proud to make this whole data
Speaker 2: set freely and publicly available. To the last section of
Speaker 2: my part of this talk, I'd like to give a
Speaker 2: brief walkthrough of the SATPM website and how to access
Speaker 2: the SATPIM data sets. For that, I'm going to switch
Speaker 2: over to a demo of the website itself. All right,
Speaker 2: So the easiest way to access the SATPM two point
Speaker 2: five data set is via its dedicated website, which can
Speaker 2: be found at SATPM dot org. The main homepage shown
Speaker 2: here gives a brief description of the data set with
Speaker 2: links to many of the associated information sources we've already
Speaker 2: covered today. One element I particularly like to draw your
Speaker 2: attention to is this form at the bottom of the page,
Speaker 2: which allows you to sign up for our mailing list.
Speaker 2: We primarily use this list to let our users know
Speaker 2: when an update has been released. Given the active nature
Speaker 2: of this data set, this happens at least annually is
Speaker 2: we extend sat PM two point five's time series forward
Speaker 2: in time, which we also coincide with expanded and updated
Speaker 2: ground based monitors and any other algorithm developments we've been
Speaker 2: working on. We really encourage our users to sign up
Speaker 2: so that they can stay informed about these updates and releases.
Speaker 2: The menus are fairly self explanatory, with access to our
Speaker 2: most recent data sets via the data access tab. Here,
Speaker 2: you'll find our latest global CNN based product V six
Speaker 2: GLO three, as well as the corresponding GWR based version
Speaker 2: V five GL six, and while not our focused today,
Speaker 2: there's also a North American regional CNN based product V
Speaker 2: six and AO one, which offers a compositional data set
Speaker 2: over North America following a similar methodology. Selecting any of
Speaker 2: these such as V six YLO three will then direct
Speaker 2: you to the product specific page. Here you'll find more
Speaker 2: detailed information related to this specific version, with a particular
Speaker 2: focus on any updates compared to the previous version. You
Speaker 2: also find reference to each data sets associated publication, which
Speaker 2: will provide you a detailed understanding of the data sets development.
Speaker 2: A little further down, you'll find details about the format
Speaker 2: and usage of the data itself, along with with access
Speaker 2: to the data set. For ease of access, we currently
Speaker 2: post this data set at both zero point zero one
Speaker 2: and zero point one degree resolution or about one in
Speaker 2: ten kilometer, and available from two separate repositories. This courser
Speaker 2: zero point one degree resolution is much easier for file
Speaker 2: size to work with and especially for the global files,
Speaker 2: but the zero point zero one degree files are therefo
Speaker 2: any projects that require it. For access, the first access
Speaker 2: point is via box folders using the links provided. To
Speaker 2: avoid downloading the whole data set, we've parsed the data
Speaker 2: by regions, say for South America, AF for Africa, GL
Speaker 2: for Global and A for North America, AS for Asia
Speaker 2: and EU for Europe. Within these you'll find subjectories for
Speaker 2: either monthly and or annual mean values, and then the
Speaker 2: net CDF formatted files themselves.
Speaker 3: Access is also.
Speaker 2: Available via the AWS Registry of Open Data via a
Speaker 2: dedicated S three bucket. Browsing the data via this repository
Speaker 2: is similar to the box folders, with subdirectories for region
Speaker 2: and timescale. Data can also be downloaded via the AWS
Speaker 2: command line interface, with specific instructions given here. Lastly, we
Speaker 2: recognize that often the questions being asked of our data
Speaker 2: set don't require access to the full high resolution data set.
Speaker 2: In fact, it can even be cumbersome to do so.
Speaker 2: To simplify such applications, we provide a select number of
Speaker 2: process data sets via the links here at the bottom
Speaker 2: of each version's page. These links will direct you to
Speaker 2: process national and regional summaries in a simple text based
Speaker 2: CSV format. These files can be downloaded and viewed within
Speaker 2: any number of programs, such as Excel, which I'm bringing
Speaker 2: onto the screen here. They contain national or subnational population
Speaker 2: and geographically weighted mean p and two point five concentrations
Speaker 2: by year covering the entire sat PM time series, as
Speaker 2: well as a breakdown of the percent of the population
Speaker 2: above certain thresholds. So if for example, you're interested in
Speaker 2: how exposures have changed over China for the past fifteen years,
Speaker 2: you could do a very quick searchdown to where China
Speaker 2: is being stored here the Chinese data that is, of course,
Speaker 2: select the last fifteen years or so, as well as
Speaker 2: the corresponding p twenty five concentrations, and end up by
Speaker 2: just simply plying that with an insert scatter plot you
Speaker 2: can see how concentrations have changed or China or any
Speaker 2: country you're interested in for the past little while. Of course,
Speaker 2: the exact structure of doing this would depend on your
Speaker 2: exact software of choice, but it's meant to be a
Speaker 2: very simple access point for answering these very questions quite quickly. Well,
Speaker 2: not our main focus, we do offer few tools and
Speaker 2: examples of accessing and reformatting the data, mainly provided by
Speaker 2: some of our users, which can help get you started
Speaker 2: on more advanced applications. The archive in Publications tab provide
Speaker 2: access to historic versions and to our group's publications related
Speaker 2: to sat PM two point five, and finally, the support
Speaker 2: team tab gives you a point of contact if you
Speaker 2: have any questions need to reach out, as well as
Speaker 2: specific contact for the members of a group responsible for
Speaker 2: supporting and overseeing this product. Lastly, if this SATPM product
Speaker 2: is of interest to you, I'll put in one final
Speaker 2: plug for joining our mailing list. We don't use extensively
Speaker 2: and so you won't be fluttered with emails, but it
Speaker 2: really does allow us to most effectively inform our users
Speaker 2: when new versions are released. And with that, I'll turn
Speaker 2: over to Pawan and Junhayung who will be discussing a
Speaker 2: direct machine learning approach for surface PM point five estimation.
Speaker 3: Thanks Eren for covering that. Now I'm going to talk
Speaker 3: about some of the direct machine learning approach to get
Speaker 3: the PM two point five using satellite data and model outputs.
Speaker 3: So before I start, I like to introduce three main
Speaker 3: categories of approaches used to estimate surface PM two point
Speaker 3: five from satellite data and how direct machine learning fit
Speaker 3: into that landscape. The first two approach are geophysical and
Speaker 3: hybrid approach, which we have covered in previous section by Eron.
Speaker 3: The third is direct mL approach. In this there is
Speaker 3: no city in the loop. Instead, machine learning models learn
Speaker 3: the mapping directly from the inputs such as satellite davids
Speaker 3: and reanalysis variable straight to surface PM two point five.
Speaker 3: Everything is learned end to end from the data. There
Speaker 3: are many examples in the community, and we'll explore one
Speaker 3: of them in here that is marred to CNN.
Speaker 2: Now.
Speaker 3: This slide shows how a machine learning based PM point
Speaker 3: four five product can support applications across three time horizon,
Speaker 3: the past, the present, and the future. On the left
Speaker 3: is the historical case answering question like what was PM
Speaker 3: two point five in twenty ten? Here we use complete
Speaker 3: retrospective archives, long term satellite data, reanalysis, and ground monitors.
Speaker 3: These products are ideal for exposure studies, trend analysis, and
Speaker 3: building decade long records. In the middle one is near
Speaker 3: real time what is PM to pint five right now?
Speaker 3: This relies streaming satellite data such as geostationary AEROSO, optical
Speaker 3: depth and short term metrilogical forecast. These rapid updates supports
Speaker 3: operational monitoring, wildfire response, dust storms, and public health alerts.
Speaker 3: And on the right is forecasting what will be PM
Speaker 3: to point five tomorrow. These models rely only on forecast
Speaker 3: metrilogy since feature satellite observations are ut available, and emission
Speaker 3: forecast forecasts can help with planning public alerts and proactive
Speaker 3: exposer mitigation strategies. Together, these three time horizons shows how
Speaker 3: machine learning can help us understand past air quality, monitor
Speaker 3: current conditions, and anticipate future. The typical workflow for machine
Speaker 3: learning based PM two point five estimation. We start with
Speaker 3: satellite aerosoloptical depth and reanalysis AEROSO product as the primary input.
Speaker 3: Then we bring in axillary variables such as humidity, boundary air,
Speaker 3: high temperature, windland cover, mostly from re analysis or forecasting
Speaker 3: models like AR five and MIREMERATU. These fields helps the
Speaker 3: model understand meteorological context. Finally, we use ground based PM
Speaker 3: to point five observation from networks like air now and
Speaker 3: open Aqui platform as target variables. All of these inputs
Speaker 3: feeds into machine learning models such as random forest ex
Speaker 3: you boost deep learning to produce surface PM to pint
Speaker 3: five estimates. Here I walked through the full pipeline we
Speaker 3: used to estimate surface PM to point five from satellite observation.
Speaker 3: We begin by acquiring data satellite aud and ground based
Speaker 3: PM to point five measurement. Next, we co located these
Speaker 3: data sets in time and space to ensure they represent
Speaker 3: the same atmospheric conditions. Then we build statistical or machine
Speaker 3: learning models that learn the relationship between aerosoloptical depth and
Speaker 3: surface PM too point five. Once train, the model estimate
Speaker 3: PM too point five across the satellite grade, providing continuous
Speaker 3: special coverage. As an optional final step, we convert those
Speaker 3: PM to point five values into air quality index categories
Speaker 3: to support public health communication. This is a reference like
Speaker 3: showing some example of machine learning use cases. Overall, this
Speaker 3: progression shows how machine learning is increasingly integrated into both
Speaker 3: observation driven and model driven PM to point five system
Speaker 3: improving accuracies across time scales. Here I'm highlighting the key
Speaker 3: strength and limitation of machine learning based PM to point
Speaker 3: five estimates. Machine learning approaches are powerful because they can
Speaker 3: capture nonlinear relationship between aerostoptical depth and PM to point five,
Speaker 3: incorporate diverse predictors or inputs, and perform well even when
Speaker 3: emission inventories are spares. They also reduce some systematic biases
Speaker 3: found in physics based models and allow fast inference one strand. However,
Speaker 3: there performance depends heavily on the density and quality of
Speaker 3: ground monitors. They may struggle with extreme events if those
Speaker 3: conditions are underrepresented in training. Database Generalization across region and
Speaker 3: time period is sometimes difficult in machine learning approaches. Okay, now,
Speaker 3: let's move on to the examples of machine learning. Deride
Speaker 3: PM to point five data sets called bias corrected MARA
Speaker 3: two PM two point five at our ly scale produced
Speaker 3: as a NASA's hay Cost project. Before we get to
Speaker 3: bias corrected MARAP to PM two point five, take a
Speaker 3: look at comparison of several major global PM too point
Speaker 3: five product. These data sets offer in different temporal coverage
Speaker 3: special resolution in the method used to generate PM two
Speaker 3: point five. For example, deepcamp and HAP combine the satellite
Speaker 3: ground and deep learning approach to provide high resolution hourly
Speaker 3: or daily estimates. Berkeley Earth relies primarily on the ground
Speaker 3: reasurement with interpolation, while the Washington University data sets which
Speaker 3: we just learned about, integrate satellite model and ground input
Speaker 3: at coursal temporal scale. Then, finally, MARA to CNN hay
Speaker 3: Cast PM two point five product, we will learn more
Speaker 3: on it incoming slide. The slide provides some more specific
Speaker 3: details on the MARA to CNN PM two point five
Speaker 3: data sets, although it says only available until twenty twenty four,
Speaker 3: but periodically continue to produce the data and currently as
Speaker 3: we will see, available until twenty twenty five. Temporal resolution
Speaker 3: is one hour with special resolution is about half a degree,
Speaker 3: which is original special resolution of MARA to data product.
Speaker 3: Each daily file contained global data in net city of
Speaker 3: format and the recently publication provides more details on the
Speaker 3: methodology and validation efforts. Each PM two point five estimation
Speaker 3: over each Mara to grid comes with a qaflag. This
Speaker 3: qaflag provides a simple grid level measure of confidence in
Speaker 3: the estimated PM two point five value. It determines by
Speaker 3: factors such as proximity to the surface, monitor line covered,
Speaker 3: type and latitude. Area with more monitoring or a stable
Speaker 3: surface condition generally receives a higher scores. The Quality Assurance
Speaker 3: flag range from one to four, where one represent low
Speaker 3: quality and four represent highest. For most quantitative analysis, we
Speaker 3: recommend using data with a flag value of three or four.
Speaker 3: The map on the right shows how this quality level
Speaker 3: vary across the globe. As you will notice that the
Speaker 3: flags are higher where the ground monitored density is high.
Speaker 3: Now we will get into some details on how these
Speaker 3: PM to point five products are derived. Marito provides PM
Speaker 3: to point five component at surface such as does sulfate, carbon, etc.
Speaker 3: And this equation is often used to combine these components
Speaker 3: to calculate total surface PM two point five. You'll notice
Speaker 3: that it is missing nitrate, which can be significant contributed
Speaker 3: in certain regions. When we will validate these across the globe,
Speaker 3: we found large biases and comparison with individual station, although
Speaker 3: it shows good consistency over larger areas. Here is the
Speaker 3: map showing its global performance against ground monitors in year
Speaker 3: twenty twenty. What we see is that bias pattern varies
Speaker 3: significantly by region and bi aerosol types. In places like California,
Speaker 3: Eastern China, and part of Middle East, marit tend to
Speaker 3: overestimate PM too point five. In contrast, regions such as
Speaker 3: the indogntetic plane and Southern and South America shows under estimation.
Speaker 3: These regions regionally varying biases highlight y machine learning correction
Speaker 3: is so important for producing PM too point five astimates.
Speaker 3: In this slide, we summarize the key mara to input
Speaker 3: be used to model PM too point five by capturing
Speaker 3: the major process that derives aerosol formation, transport and removal.
Speaker 3: On the aerosol side, we include black carbon, organic carbon, dust, sulfate, ESOTO,
Speaker 3: surface mass concentration as well as total extinction at five
Speaker 3: hundred fifteen nanometer that is our alsoloptical app These variables
Speaker 3: help represent both primary and second reparticle sources. On the
Speaker 3: meteorological site, we use surface pressure, humidity, temperature at multiple
Speaker 3: levels and ten meter wind component to capture boundary dynamics
Speaker 3: and atmospheric mixing. All inputs are taken at their native
Speaker 3: merit to spatial resolution and at hourly timing scale, ensuring
Speaker 3: consistent alignment across the data set. Here we describe how
Speaker 3: we use a reference grade PM two point five measurement
Speaker 3: collected from open eq platform as ground truth for training
Speaker 3: and validating our model. We begin by collecting measurement from
Speaker 3: more than five thousand monitoring station worldwide and collocating each
Speaker 3: site with the nearest Marato grid cell. We then temporarily
Speaker 3: match the ely PM to point failue with the crossmending
Speaker 3: Marato inputs. After collocation, we apply a strict quality control
Speaker 3: to remove outliers and invalid records, resulting in more than
Speaker 3: thirty million varied hourly observations. For model development, we used
Speaker 3: year twenty eighteen and nineteen data and for training, and
Speaker 3: for reserved the data from twenty twenty as an independent
Speaker 3: validation period to assess the performance. Here is a full
Speaker 3: workflow used to produce the bias corrected PM to point
Speaker 3: five data sets. As we described earlier, we start with
Speaker 3: especially and temporarily collocating marat meteorological and aerosol input with
Speaker 3: open EQ surface observation. Then we apply five different aerosol
Speaker 3: aware data processing strategies, each capturing PM two point five
Speaker 3: behavior from a slightly different perspective. The final product is
Speaker 3: a single bias corrected PM to point five estimate for
Speaker 3: every single grid cell every hour, with improved accuracies and
Speaker 3: regional consistency. Here is some more details on five day
Speaker 3: different strategies. Each strategy train its own independent convolution neural network,
Speaker 3: but they differ in how aerosol information is organized. For example,
Speaker 3: M one is our baseline model using the full data
Speaker 3: set with no studyfication. M two classified data by dominating
Speaker 3: aerosol species h each other, while M three assigned dominating
Speaker 3: species species by station for the data period. M four
Speaker 3: assign each mirror to grid cell in dominating species across
Speaker 3: the three years data sets, and M find find that
Speaker 3: by merging sulfate and cars aerosol as a single dominating species. Finally,
Speaker 3: the ensemble model, which is M six, combine output from
Speaker 3: all five approaches to produce a single more robust PM
Speaker 3: to pint five estimates. Here we show how the ensemble
Speaker 3: model M six improve PM to pint five estimates compared
Speaker 3: to the RAWMERA to data products. On on the left
Speaker 3: you see the training data which come which has data
Speaker 3: from year twenty twenty eight to twenty twenty nine. The
Speaker 3: red line shows MARA II and the blue line shows
Speaker 3: the M six output, and on the right is the
Speaker 3: validation period which is an INDEPENDENTATA set are coming from
Speaker 3: year twenty twenty two. So we perform this evaluation by
Speaker 3: binning observed PM two pointween two contiles and comparing model
Speaker 3: values across full concentration range. The blue line represent ensemble
Speaker 3: which closely track the oneine one line in both the
Speaker 3: training years and the independent twenty twenty validation journalists within
Speaker 3: the plus minus two sigma certainty weight. In contrast, the
Speaker 3: red line shows that the raw ERA to systematically overestimate
Speaker 3: low PM two point five values and understimate high concentration.
Speaker 3: These concentrations depend bias are clearly visible, and the results
Speaker 3: demonstrate how the ensemble approach using CNN substantially reviews these biases.
Speaker 3: Here we compare regional PM two point five biases in
Speaker 3: raw MARA to product with CNN corrected data. So as
Speaker 3: we have seen earlier. In the top panel, MARA to
Speaker 3: exhibit large regionally consistent error over estimation by more than
Speaker 3: twenty microgram per cubic meter in places like California, China,
Speaker 3: Medalist and understand entition by more than fifteen microgramperm meter
Speaker 3: over endopendetic plane in parts of South America. The bottom
Speaker 3: panel shows how CNN greatly reduce these biases, producing more
Speaker 3: localized and better calibrated pattern across diverse aerosol regimes. CNN
Speaker 3: correct PM two point five also has some limitation and
Speaker 3: it over predicts in eastern Southeast Asia at high concentrations.
Speaker 3: Here is another performance metric showing the performance of CNN
Speaker 3: versus RA MARA two PM two point five. The bar
Speaker 3: shows the medium index of agreement, which measures how well
Speaker 3: the estimations capture both the magnitude and pattern of observed
Speaker 3: pm too point five closer to one is better. Across
Speaker 3: every region, CNN performs sustentionally better. The improvement are especially
Speaker 3: strong in South Asia and Southeast Asia, where agreement nearly
Speaker 3: doubles compared to MARA two. Europe and the United States
Speaker 3: also shows clear gains. The key TAKEABA is that CNN
Speaker 3: more accurately represents both special and temporal variability a PM
Speaker 3: to point five, providing a more reliable global data sets
Speaker 3: than the raw MARA to PM two point five. So,
Speaker 3: in summary, Mara tow CNN hey cast PM too point
Speaker 3: five data sets provides long term global hourly data improve
Speaker 3: PM TOO point five estimation over Mara I bias corrected
Speaker 3: using ensembles, CNN approach and quality AWAID data sets. Now
Speaker 3: we'll look how to access these data sets. The slide
Speaker 3: kind of summarize the different ways in which a user
Speaker 3: can access the PM to point five data sets. The
Speaker 3: data ending page is the best starting point, offering an
Speaker 3: overview of product along with key links. The read we
Speaker 3: file provide detailed documentation on file structure and variable conditions.
Speaker 3: Users who prefer direct download can use the online archive
Speaker 3: or Earth Data search to retrieve files for their specific
Speaker 3: time period. For cloud based workflow. The data set is
Speaker 3: also available on AWS. Open Deep is another option which
Speaker 3: enables remote access through tools like Python, matlib or R
Speaker 3: and Giovanni for an easy browser based interface for subsetting,
Speaker 3: averaging and visualizing the data sets. Here is just screenshot
Speaker 3: of data lending page at NASA disk which hosts the
Speaker 3: pm to my five data sets from this project. Next
Speaker 3: we'll walk through a quick demo of how to access
Speaker 3: these data in NASA Giovanni tool. So now we are
Speaker 3: going to look at an exercise using NASA's Giovanni tool
Speaker 3: to access the merit to CNN PM to PINT five
Speaker 3: data sets and see a couple of example how we
Speaker 3: can access and map and do the time stage type
Speaker 3: of analysis online. So in order to search Giovanni page,
Speaker 3: I would just go to Google and search Giovanni. NASA
Speaker 3: Giovanni will be much better and you will see the
Speaker 3: first link which is NASA Giovanni. I'll just click on that.
Speaker 3: When you first time click on that link, you will
Speaker 3: notice it will ask for a username and passwords. So
Speaker 3: if you do not have Earth data log in user
Speaker 3: name and password, you might want to register and create.
Speaker 3: It's a free If you already have it, you should
Speaker 3: be able to log in. My computer already had so
Speaker 3: it already logged in. And if you do not log
Speaker 3: into the Earth Data log in, then you will have
Speaker 3: limited capabilities from the Giovanni. You wan't to be able
Speaker 3: to do everything. So this is homepage of Giovanni. It's
Speaker 3: a tool which actually allows visualization and analysis of many
Speaker 3: many NASA's data sets. As you can see on the
Speaker 3: left side is the list of all the different platforms,
Speaker 3: variable disciplines the data sets are available. We if you
Speaker 3: actually want to explore, you can click each of these
Speaker 3: taps like measurements and you will see a lot of
Speaker 3: different types of measurements which are listed here. There's like
Speaker 3: five hundred or more. You can also select by platform
Speaker 3: or instruments such as like Modi, is Mara or whatever
Speaker 3: you want. You can also subsidt based on the resolution
Speaker 3: and other things. And then if you go on the
Speaker 3: top that's where the analysis tools are available. So if
Speaker 3: you go to the select part, it will give you
Speaker 3: the list of different type of analysis you can do.
Speaker 3: You can do maps, comparison, time series, histograms, and so
Speaker 3: on and so forth. And this is the selection of
Speaker 3: date starting an end date. And then finally you have
Speaker 3: a geographical region selection and the geographical reasons selection. There
Speaker 3: are various ways we can explore that. First, so since
Speaker 3: we are going to look the MARA to c in
Speaker 3: NPM two point five, instead of going through this whole list,
Speaker 3: I will just try to search here, I'll just say
Speaker 3: CNN and hopefully that should pop up. So here is
Speaker 3: as soon as I do the CNN and do the search.
Speaker 3: The parameter which I'm looking for is called bias corrected
Speaker 3: surface total PM too point five, mass concentration all quality level.
Speaker 3: And then you needs and source and resolutions. And as
Speaker 3: I mentioned earlier, the data is actually now available until
Speaker 3: end of last year twenty twenty five, so we process
Speaker 3: the data frequently. So first thing we are going to
Speaker 3: look we selected the data. Now, first thing we are
Speaker 3: going to do a time averaged map. Okay, so I'm
Speaker 3: going to just select a map type and then I
Speaker 3: will select a date in twenty twenty five last year.
Speaker 3: And for some reason, I'm going to pick August, let's
Speaker 3: say August eight, and that is my starting and then
Speaker 3: same date, I'm going to pick August eighty five, so
Speaker 3: I'm just for quick analysis. I'm just picking one day,
Speaker 3: so it has since this is hourly data, so the
Speaker 3: total data will be twenty four timestamp. And then I
Speaker 3: will select the geographical area. Now, in the geographical area,
Speaker 3: I can actually select a box rectangular box around any
Speaker 3: geographic region, or I can actually select specific countries like
Speaker 3: or many other US states and many other things. So
Speaker 3: for simplicity, we're going to just select a simple box
Speaker 3: over continental northern America. I'm going to select something which
Speaker 3: covers both Canada and US to show a specific case here.
Speaker 3: So and you can if you're walking through with me,
Speaker 3: you can select any region of the world you like
Speaker 3: to select. Once you do that, so now we have
Speaker 3: selected the type of analysis we want to do, which
Speaker 3: is time abridge map. We have selected the dat range,
Speaker 3: we have selected the region of interest, and we have
Speaker 3: selected the variable which we want to plot. Once you've
Speaker 3: done all of that, you just click on that bottom
Speaker 3: right corner plot data. Once you do that, the system
Speaker 3: is going to go and get the data for that
Speaker 3: particular day, all twenty four time of stems and then
Speaker 3: start making plot. So here is the plot. It's very quick.
Speaker 3: You have a color scale here on the right, and
Speaker 3: then you have a map. Now what you see here
Speaker 3: these high values in Canada is actually there huge fire
Speaker 3: during this period. So that is why I selected that
Speaker 3: specific month and day to demonstrate the capture of smoke
Speaker 3: loom within these data sets. So we can make this
Speaker 3: plot look a little bit better. So to do that,
Speaker 3: I will click on the option and under the options
Speaker 3: again there is option I will change the color scales
Speaker 3: to a little bit more, to seventy five, and then
Speaker 3: this palette I can actually select another palette. I'm going
Speaker 3: to select something which is a little bit more. I'm
Speaker 3: going to select this orange from yellow to oranges. And
Speaker 3: then also I'm going to apply some smoothing so that
Speaker 3: the map looks a little bit better, and then the
Speaker 3: projection is fine, and just say replot. Once I do,
Speaker 3: you can see that the color scale has changed the values.
Speaker 3: Since I've only selected the maxim value up to fifty
Speaker 3: of the high values are masked. So let's go back
Speaker 3: and actually revise that and maybe go to hundreds and
Speaker 3: do the replotting again. And you can use the lab
Speaker 3: ramathic scale also if you like to see the plots
Speaker 3: and difference. So now we can see really nice smoke
Speaker 3: prooms from these fires showing high concentration of PM two
Speaker 3: point five here. So this is very simple to use tool.
Speaker 3: You can do this for any parts of the world.
Speaker 3: Now let's try another. We will use the same day
Speaker 3: and we will just do a quick time series over
Speaker 3: this region and see how the PM two point five
Speaker 3: is evolving throughout the day. So to do that, I'll
Speaker 3: just click again bottom right corner which says back to
Speaker 3: the data selection. So once I do that, it goes
Speaker 3: back to my previous selection. Now I can change the
Speaker 3: plot type. I'm going to do time series area averaged, okay,
Speaker 3: and I will keep the day same because I want
Speaker 3: to see the hourly evaluation of the PM two point five.
Speaker 3: But I will just focus on this small area where
Speaker 3: we saw very high concentration of PM too point five.
Speaker 3: Just the small area because we want to see the
Speaker 3: time step evaluation. Right, So again I changed the plot
Speaker 3: type to time series date. I kept the same from
Speaker 3: zero to twenty three hours and the area I have changed. Again,
Speaker 3: I'll say plot the data and since it is averaging
Speaker 3: very small area, so it should be very fast if
Speaker 3: you're doing this over longer a period of time. Many
Speaker 3: many years or larger area, then the tool can take
Speaker 3: much longer time. So now you can see a dinal
Speaker 3: variability on August eight, twenty twenty five, and all the
Speaker 3: times are in UTC, so make sure that you remember
Speaker 3: that you can convert that into local time if you want.
Speaker 3: But you can see the concentrations are very high at
Speaker 3: the start of the day and they started going down actually,
Speaker 3: so again I would suggest you to actually play with
Speaker 3: this tool. It's really nice. I can show you some
Speaker 3: other example which I have run earlier. This is another
Speaker 3: day where we have done the times. There is from
Speaker 3: July first to August July thirty first, and this is
Speaker 3: a different part of the US actually, but it does
Speaker 3: show some increasing trend and again this increase is actually
Speaker 3: resulting from the transport of the smoke in that part
Speaker 3: of the region. This is another one from August. You
Speaker 3: can see from earlier August there were a lot more
Speaker 3: PM to point five concentration and August and this again
Speaker 3: influenced by the those wildfires which we have seen. So
Speaker 3: it's a really nice and easy tool. I would strongly
Speaker 3: recommend you play with different types of analysis you have here.
Speaker 3: You can do the histogram how the PM too point
Speaker 3: five actually looked on that particular area. So it will
Speaker 3: basically take all the data in that box which we
Speaker 3: have selected for that particular day and do a quick
Speaker 3: histogram to understand what kind of values we see. So
Speaker 3: you can see very often the values are very low
Speaker 3: most of the time, but it does shows longer tail
Speaker 3: with values ranging all the two hundred hundred twenty which
Speaker 3: we saw earlier in the map. Also, so you can
Speaker 3: actually do a lot of good things. You can also
Speaker 3: download this as a PNG or net city of files,
Speaker 3: and then there are other options. So I hope this
Speaker 3: tool is useful and next week CALL will actually go
Speaker 3: through this data sets a little bit more and provide
Speaker 3: you more ways in which you can actually analyze this
Speaker 3: data along with the ground data and the satellite based
Speaker 3: PM to point five station. Thank you back to the call.
Speaker 1: So thank you very much doctor Goop and doctor von
Speaker 1: Donklar for those overviews of the data products and their
Speaker 1: associated methodologies. To summarize what we discussed in this part,
Speaker 1: the two data sets we discussed are the Washington University
Speaker 1: in Saint Louis SAT PM two point five data set
Speaker 1: and the NASA bias corrected Merit two PM two point
Speaker 1: five data set. For the SATPM data set, the methodology
Speaker 1: it employs is a hybrid approach where a geophysical estimate
Speaker 1: of PM two point five from model simulations and aerosolo
Speaker 1: optical depth is corrected for biases using ground based measurements
Speaker 1: of PM two point five on a global scale. The
Speaker 1: key strengths of this data set are its fine spatial resolution.
Speaker 1: Gridded products are available at point zero one degree roughly
Speaker 1: one kilometer spatial resolution, as well as having a long
Speaker 1: record going back to nineteen ninety eight and continuing till
Speaker 1: twenty twenty four. The data set has consistent performance even
Speaker 1: far from the location of ground based monitors, as assessed
Speaker 1: through robust spatial cross validation. A possible limitation to consider
Speaker 1: when using this data set are the monthly temporal resolution
Speaker 1: of the products, so this would be unable to resolve
Speaker 1: variability at daily or hourly scales. The data are free
Speaker 1: and openly accessible via the SATPM website. For the bias
Speaker 1: corrected Merit II data set, the methodology used by that
Speaker 1: data set is an ensemble approach where a series of
Speaker 1: convolutional neural network models are developed and calibrated for different
Speaker 1: dominant aerosol conditions, and then a final ensemble model is
Speaker 1: used to select the appropriate model depending on the dominant
Speaker 1: conditions for any location. Key strengths of this data set
Speaker 1: are its high temporal resolution. It's available at hourly temporal resolution.
Speaker 1: It also has a rather long data record from twenty
Speaker 1: to twenty twenty four, and this data set represents PM
Speaker 1: two point five spatial and ten p poral variability much
Speaker 1: better compared to the Merit II Global reanalysis data set
Speaker 1: on which it was based. A possible limitation to consider
Speaker 1: when using this data set is its relatively coarse spatial resolution.
Speaker 1: It's available on the original MERRITI grid, which is a
Speaker 1: point five zero point six two five degree or roughly
Speaker 1: fifty kilometer spatial resolution grid. These data are also free
Speaker 1: and openly accessible via NASA Earth Data and for example,
Speaker 1: you can use tools like Giovanni to do online analysis
Speaker 1: with this data product. In Part three, we'll be looking
Speaker 1: at a case study where we will download and compare
Speaker 1: the sat PM the bias correct Merit two data sets
Speaker 1: to ground based PM two point five measurements for specific
Speaker 1: region and time period of interest. This will demonstrate both
Speaker 1: how to access these data sets. And how to evaluate
Speaker 1: them in a real world setting for practical applications. To
Speaker 1: be sure that you're ready for Part three, we're asking
Speaker 1: you to prepare it advance three different things. First, ensure
Speaker 1: that you have access to Google Collab. This is accessible
Speaker 1: to anyone with a Google account such as a Gmail
Speaker 1: address if you have that. If not, those are free
Speaker 1: to sign up for, and accessing a Gmail address or
Speaker 1: having a Google account will enable you to access Google
Speaker 1: Collab and use it during the training. Second, you'll need
Speaker 1: a account with NASA Earth Data. Again, this is a
Speaker 1: free account. You can sign up for one if you
Speaker 1: don't already have one, and for the training, you should
Speaker 1: have your username and password ready to input into the
Speaker 1: code we're using to allow you to access the NASA data.
Speaker 1: And finally, we ask you to have an open aq
Speaker 1: account to access an open aq API key, which will
Speaker 1: also be using to download data from that global data
Speaker 1: aggregation service that we discussed in part one. If you
Speaker 1: don't have an account, again, that's free to sign up for,
Speaker 1: so you can go to the open Aq website and
Speaker 1: register for an account and they will provide you with
Speaker 1: an API key as a reminder, the homework for this
Speaker 1: training will be issued after the last part of the training,
Speaker 1: which will be Part three next week, and in order
Speaker 1: to receive a certificate for a completion of the training,
Speaker 1: you should attend all three parts of the training through
Speaker 1: the webinars and also complete the homework assignment by the deadline,
Speaker 1: which will be two weeks after the homework is issued.
Speaker 1: This is the contact information for myself and the trainers
Speaker 1: who contribute to this part of the training, as well
Speaker 1: as for the RSET program in general. And here are
Speaker 1: the references included in the material. So thank you all
Speaker 1: very much for your attention, and we'll move on to
Speaker 1: the question and answer portion of the training. All right,
Speaker 1: So as we transition to the Q and A, I
Speaker 1: just want to remind everyone we'll try to answer as
Speaker 1: many questions as possible in the rest of this session.
Speaker 1: We'll be recording all our answers in a document, including
Speaker 1: the questions we don't get to in the live session,
Speaker 1: and we'll be posting that document to the to the
Speaker 1: training website as soon as we've gone through and completed
Speaker 1: answering all the questions. So let me just check. Okay,
Speaker 1: here we go. So the first question is related given
Speaker 1: the non unique relationship between satellite observed AOD and PMT
Speaker 1: one five, Can conditional density estimation models such as mixed
Speaker 1: density networks represent which of the uncertainty threshod exceed the
Speaker 1: probability early warning and so forth? And what validation framework
Speaker 1: is needed to demonstrate spatial transferability to monitor sparse regions.
Speaker 1: We had a couple of questions similar to this, so
Speaker 1: I'll just answer sort of generally here.
Speaker 3: Uh.
Speaker 1: You know, the techniques we've talked about in this training
Speaker 1: are are two of the techniques, but there are you know,
Speaker 1: other techniques such as the one you mentioned, can also
Speaker 1: be applicable, and they may have their own strengths and
Speaker 1: weaknesses which would have to be determined for a given application.
Speaker 1: But in general, the validation framework needed to demonstrate the
Speaker 1: transferability of any methodology, including the methodologies we talked about today.
Speaker 1: Were the methodologies which which Aaron von Donkel are covered
Speaker 1: in his portion. Uh. Those include buffer leave, cluster out
Speaker 1: cross validation to assess spatial generalizability and temporal hold back
Speaker 1: to general to assess the temporal uh generalizability or transferability.
Speaker 1: So I'll just use again, I'll just use that as
Speaker 1: a general opportunity general opportunity to answer questions about UH validation,
Speaker 1: and similarly for a question too. This question is specifically
Speaker 1: about how models trained in a country can be transferred
Speaker 1: to Korea, but I'll use this as an opportunity to
Speaker 1: ask kind of the general question of generalizability and transferability
Speaker 1: of models, including the models presented here. So in this case,
Speaker 1: maybe i'll I'll ask Aaron and then maybe pol On
Speaker 1: to comment on this on the more general question of transferability.
Speaker 1: So Aaron, let's start with you.
Speaker 2: Fuir, So, I mean, in general I would not advise
Speaker 2: using a statistically trained model that was trying to want
Speaker 2: region to go to another. It really depends on the structure.
Speaker 2: Using One of the strengths you know when I discussed
Speaker 2: the geophysical method was it is really designed to try
Speaker 2: to be as globally obflitable as possible. But really, in general,
Speaker 2: you don't want to train a model one addict another,
Speaker 2: meaning in either event. Really, as you know, Carl, answer
Speaker 2: the first question, I mean, good robust validation of any
Speaker 2: model that you want to use is really helpful. In
Speaker 2: such cases. You want to make sure that if you're
Speaker 2: trying to extend the model beyond where it's been trained
Speaker 2: you have some sort of independent data sets to try
Speaker 2: to validate that.
Speaker 1: Okay, thank you, Paul. On anything to add to that or.
Speaker 3: No, I think that is true. Only thing I would
Speaker 3: add is it really depends. There are many things which
Speaker 3: you can consider. One is what kind of PM two
Speaker 3: pint five concentrations and aerosols types are in different regions.
Speaker 3: So let's say you trend a model over US and
Speaker 3: you want to apply over Europe. In some cases it
Speaker 3: will work. But then, as Aaron says, if you have
Speaker 3: if you do not have any ground measurement to validate,
Speaker 3: then it is going to be somebody's guests.
Speaker 1: Right.
Speaker 3: We don't know what is the quality of that. So
Speaker 3: depending on specific situation type of ASO concentration range, meteorological condition,
Speaker 3: geographical terrain, and other kind of look constraint, the model
Speaker 3: may or may not work. So it's it's a really
Speaker 3: case by case determination whether it should, whether it will
Speaker 3: work or not.
Speaker 1: Okay, thank you. So moving on to question three, was
Speaker 1: the minimum exploitable spatial and temporal resolution for estimating PM
Speaker 1: two point five at the scale of a Sephlian city
Speaker 1: like Waga, dougu or Kaya using mayak or Veer's products,
Speaker 1: So the Mayak, the Modus Mayak, and the Veers MAAC
Speaker 1: products specifically are available at a one kilometer spatial resolution,
Speaker 1: so that would be the maximum possible, the best possible
Speaker 1: spatial resolution, and these would be these are both polar
Speaker 1: orbiting satellites, so those would be available at a one
Speaker 1: day temporal resolution. So if you restricted yourself to just
Speaker 1: those sources of information, that would be the the best
Speaker 1: possible spatial and temporal resolution you can get. As discussed
Speaker 1: in the training, there are ways to bring in other
Speaker 1: types of information to get for the biace correcting Mara
Speaker 1: two product, for example, to get the higher tempore resolution.
Speaker 1: So moving on to question four, I think we basically
Speaker 1: answered this. This is kind of a similar to question one,
Speaker 1: So we'll just to sort of reiterate yes, it may
Speaker 1: be applicable and robust validation strategy would be needed to
Speaker 1: assess that. Question five. Are the AODPM two point five
Speaker 1: conversion coefficients already calibrated for West Africa or the the hell?
Speaker 1: Or is it necessary to estimate them locally? Yeah, I'll
Speaker 1: just kind of quickly We talked about that a little
Speaker 1: bit before as well, in the sort of general question
Speaker 1: about generalizability. So, yes, these data sets, in the data
Speaker 1: sets we discussed today include localized calibrations that would make
Speaker 1: these locally relevant. But if you want to create your
Speaker 1: own data set you would have to perform those local
Speaker 1: calibrations yourself. Anything else to add on that from Aaron
Speaker 1: or on?
Speaker 2: No, not, not really, I mean that's yeah. I mean
Speaker 2: one of the reasons we precee these, that's just to
Speaker 2: avoid having to your own calibrations. But there are cases
Speaker 2: string that aren't covered by what we do. In those cases, yes,
Speaker 2: some local calibration can be can be helpful.
Speaker 1: Well I don't have anything, okay, okay, thank you. So
Speaker 1: question six, how can surface PM two point five be
Speaker 1: distinguished from elevated smoke dust or transported aerosol layers when
Speaker 1: they produce similar AOD values? This is a good question.
Speaker 1: We alluded to this a little bit in Part one,
Speaker 1: and then in Part two you kind of saw the
Speaker 1: specific strategies that the two data sets we talked about use. So,
Speaker 1: in particular, the SATPM product uses the geos Kim simulation
Speaker 1: to simulate the vertical profiles of aerosols, and MERIT two
Speaker 1: uses the go Kart simulation. The Aerosol Simulation module, and
Speaker 1: there are also observational based methods to look at, especially
Speaker 1: the vertical distribution of aerosol from either ground or space
Speaker 1: based light ar observations which can provide the vertical distribution
Speaker 1: of atmospheric particles. Question seven, how is the correlation between
Speaker 1: AOD and maybe PM two point five at shorter temporal
Speaker 1: scales such as days or weeks. I'd say there's no
Speaker 1: sort of general answer to this question. It can depend
Speaker 1: on a lot of factors. For example, you know, averaging
Speaker 1: over longer periods of time could average out some noise
Speaker 1: in the signal, But then you're also looking at changing
Speaker 1: meteorological conditions, and as we discussed in part one, those
Speaker 1: meteorological conditions such as humidity of mounted boundary layer height
Speaker 1: can have a big impact on the AOD to PM
Speaker 1: two point five relationships. So looking at a longer period
Speaker 1: of time or a shorter period of time may may
Speaker 1: not help. Yeah, I'll maybe leave it there. Question eight,
Speaker 1: how can I obtain MYCAOD data to estimate PM two
Speaker 1: point five? We put some links there for you to
Speaker 1: access those data. They're available through NASA's Earth Data site,
Speaker 1: and once we post this document, to the website, you'll
Speaker 1: be able to access those links directly, but in general,
Speaker 1: if you go to the NASA Earth Data website and
Speaker 1: search for the data product, you should be able to
Speaker 1: find information on how to download it. To the method
Speaker 1: question nine, to the methodologies in estimating pm slash aerosol
Speaker 1: apply to all other air pollutants, So that's a good question.
Speaker 1: While this specific methodology wouldn't directly apply, there are certainly
Speaker 1: ways to modify the methodology or use similar methodologies for
Speaker 1: other pollutants. Aaron, maybe just briefly comment on your group's
Speaker 1: work there, sure so.
Speaker 2: I mean, as Carl mentioned, a lot of the concepts
Speaker 2: become the same, a lot of the details become slightly different.
Speaker 2: So N two is when our group has worked with
Speaker 2: a little bit. You know, the idea is very similar.
Speaker 2: You have a satellite retrieval of total column N two,
Speaker 2: you can relate that to near service N two using
Speaker 2: a chemical transfer model, and you can use a statistical
Speaker 2: framework to further refine those estimates. So by broad strokes,
Speaker 2: it can be very similar. But a lot of the
Speaker 2: details are often quite different and the unique challenges that
Speaker 2: come with that, so there can be a lot of
Speaker 2: work to just to copy that those concepts over to
Speaker 2: a different species, but the concepts to broadly hold.
Speaker 1: Okay, thank you. Question ten. Can MLR model for PM
Speaker 1: two point five bias correction be applied in regions with
Speaker 1: limited ground based monitoring stations like North Africa? How do
Speaker 1: you handle missing data? Again? I think, Aaron, if you
Speaker 1: want to comment on this, but we've sort of touched
Speaker 1: on this already with some of the previous questions about generalizability.
Speaker 1: But if you have anything else to add or Pollen,
Speaker 1: if you have anything else to add.
Speaker 2: Not dramatically beyond what was covered in the presentation. I
Speaker 2: think this question might have come in a bit earlier
Speaker 2: on in the presentation, So hopefully we've covered some of
Speaker 2: this during our talk. But I mean probably a lot
Speaker 2: of the methods that we're covering here were designed specifically
Speaker 2: to be more globally applicable trying to handle missing data,
Speaker 2: fill in missing data, relate those relationships or feel those
Speaker 2: relationships over broader regions as well, So they're inherently trying
Speaker 2: to account for a lot of those challenges. But again,
Speaker 2: each region is slightly unique.
Speaker 1: Okay, thank you. Question eleven, which is recommended for bright
Speaker 1: sparsely vegetated surfaces. I interpreted this as a question about
Speaker 1: the AOD products, and in general, we tend to think
Speaker 1: the deep blue AOD algorithm is better for those types
Speaker 1: of surfaces. And I believe, Aaron, you did cover that
Speaker 1: in slides twenty eight through thirty four the presentation, where
Speaker 1: there's some examples of comparing the performance of the different
Speaker 1: algorithms in different regions. I believe this question did come in,
Speaker 1: you know, before you got to those slides, So hopefully
Speaker 1: by referring to those slides, this person can get an
Speaker 1: answer to their question question twelve. Okay, this is pretty specific.
Speaker 1: When evaluating interurban environmental exposure inequality using the version six
Speaker 1: Global product TATPM product in dense megacities, we notice the
Speaker 1: gridded product heavily flattened spatial gradients, i e. Very low
Speaker 1: spatial coefficient of variation across districts compared to localized ground
Speaker 1: monitors that show very large microclimatic and seasonal peaks. How
Speaker 1: do you recommend researchers to account for or bias correct
Speaker 1: this localized spatial smoothing when conducting fine scale subdistrict exposure
Speaker 1: and Genie coefficient analyses. So, Aaron, I think that's a
Speaker 1: question for you.
Speaker 2: Sure. So, I mean there can be a number of
Speaker 2: factors affecting the scatter that you're seeing between the local monitors.
Speaker 2: I think Carl showed a slide last week that that
Speaker 2: really highlighted the amount of variability within monitors within a
Speaker 2: given city, how much monitors can vary even within a
Speaker 2: one kilometer distance. So it's worth braining mind that the
Speaker 2: estimates we're providing our average over one kilometer. Beyond that,
Speaker 2: of course, it is challenging to capture these fine scale features.
Speaker 2: One of the things within the soldering my presentation is
Speaker 2: a lot of different resolutions of data sets to go
Speaker 2: into it.
Speaker 3: And so.
Speaker 2: As much as these statistcal models are attempting to account
Speaker 2: for this fine scale of variability, there are a lot
Speaker 2: of resolutions going on, and so we tend to think
Speaker 2: of it as the one klumeter product may not fully
Speaker 2: capture those one those subparometer features, even one klmeter features fully,
Speaker 2: and so we typically recommend averaging over slightly larger areas. Dramatically,
Speaker 2: you know, up to ten colmers you can see some
Speaker 2: pretty dramatic improvement, although in my experience a lot of
Speaker 2: health studies, people tend to move around, and so I
Speaker 2: tend to think this is not necessarily for most applications
Speaker 2: a severe limitation, but that may depend on the particular
Speaker 2: study you have in mind.
Speaker 1: Okay, thank you. Question thirteen. When we build local machine
Speaker 1: learning models or downscale or validate satellite estimates with limited
Speaker 1: ground data, our leave one station out spatial cross validation
Speaker 1: yields a near zero spatial R squared even if temporal
Speaker 1: validation is high. What are your recommended best practices for
Speaker 1: validating spatial models and establishing local confidence in satellite derive
Speaker 1: PM two point five when ground stations are so severely limited.
Speaker 1: So I would actually say that you're probably doing pretty
Speaker 1: much the best you can. So the leave one station
Speaker 1: out is probably a good cross validation method for you,
Speaker 1: especially considering a condition of of sparse ground monitoring stations.
Speaker 1: If you had more of a dense network, then probably
Speaker 1: a buffered leave one station out or a buffered clustered
Speaker 1: leave a cluster out basically method would be probably more recommended.
Speaker 1: But in the case of not so dense ground measurements
Speaker 1: that there's probably not much statistical difference between these different
Speaker 1: cross validation techniques so yeah, I would say, there's there's
Speaker 1: no really no simple answer here. You're I think the
Speaker 1: believe one station out strategy is probably a good strategy
Speaker 1: for your case. And the fact that may just be
Speaker 1: that the methodology you're looking at is not able to
Speaker 1: give you a good spatial generalization. There may just be
Speaker 1: too much variability, or you're you're you're missing an information
Speaker 1: on which would explain that variability. And then my my
Speaker 1: other recommendation here would be, in such situations where ground
Speaker 1: monitoring is very sparse, it may actually be better to
Speaker 1: use a pre existing product like the saven PM product
Speaker 1: or the bias corrected Barriti product that we talked about today,
Speaker 1: and use your limited ground based monitoring information purely for
Speaker 1: evaluation of the quality of that existing product. Oh and sorry,
Speaker 1: I see polland stands up.
Speaker 3: Go ahead, Well, no, go ahead, finish your question. Answer.
Speaker 3: I wanted to just make an announcement.
Speaker 1: Okay, yeah, so yeah, I think that's basically what I
Speaker 1: want to say. So I would say the most efficient
Speaker 1: use of such sparse data is it would really be
Speaker 1: instead of trying to develop your own model, use that
Speaker 1: data to evaluate the models that are already out there
Speaker 1: and see which one is the most appropriate for your setting. Okay,
Speaker 1: go ahead, Polan, Yeah.
Speaker 3: I just want to, I think, put a node out there.
Speaker 3: I know several of you are trying to access the
Speaker 3: merit to CNN using geo Wanni and you're not able
Speaker 3: to find the data. And I was not aware, but
Speaker 3: I just found out that Giovanni team is moving that
Speaker 3: data sets to the cloud. So for time being, they
Speaker 3: have actually taken it off, so it will take about
Speaker 3: ten days for them to move it. So after July
Speaker 3: twenty fourth, the data should be available back in the
Speaker 3: Giovanna So we apologize for this inconvenience. I think we
Speaker 3: didn't plant ahead in time, but hopefully the data sets
Speaker 3: will be available soon back in the system and you
Speaker 3: will be accessed.
Speaker 1: Okay, yeah, so thank you for that announcement, Pollen. We'll
Speaker 1: maybe see I don't think that should affect your the
Speaker 1: homework that we're assigning for this. We're not using Giovanni
Speaker 1: specifically for the homework assignment, but maybe we'll just double
Speaker 1: check that that won't affect people's ability to complete the
Speaker 1: homework assignment. But yeah, thank you for that announcement. All right,
Speaker 1: So question fourteen, what proportions of nitrate dominated urb in
Speaker 1: PM two point five are formed from local emissions generated
Speaker 1: and surrounding regionals or transported across national boundaries, and how
Speaker 1: do these change under different meteorological and pollution conditions. So
Speaker 1: that's a very complicated question. You may be able to
Speaker 1: use something like a chemical transport model to get an
Speaker 1: answer to that question. It's really a little bit beyond
Speaker 1: the scope of today's training, so we won't go into
Speaker 1: more detail than that, but feel free to maybe contact
Speaker 1: us or first explore that training that we put a
Speaker 1: link to, and if that doesn't answer your question, you
Speaker 1: might need to contact us for further discussions. Question fifteen,
Speaker 1: how well does the vertical scaling of GEOSCM capture highly
Speaker 1: localized boundary layer dynamics and low altitude micro level emission
Speaker 1: sources in extremely dense megacities And are their future efforts
Speaker 1: to integrate local high resolution meteorology to improve this? So, Aaron,
Speaker 1: I think you spoke to this a little bit in
Speaker 1: your previous answers, but if you want to add anything
Speaker 1: onto this, go ahead, sure.
Speaker 2: I mean, so it's a good question, and I mean
Speaker 2: I like the idea of trying to integrate higher resolution
Speaker 2: meteorology to improve local conditions. There has been some work
Speaker 2: done on like ship emissions representing subgrid variation stuff. I'm
Speaker 2: not aware of any of that specifically looks at megacities
Speaker 2: for now. Really, what we're relying on is more of
Speaker 2: the statistical models to capture those any shortcomings within just
Speaker 2: came to capture those really high resolution features that maybe
Speaker 2: may come about because of those features such as megacities.
Speaker 2: It's it's pretty good, we think, but it's certainly an
Speaker 2: actual physically based model would be even better if we
Speaker 2: could to pull it off.
Speaker 1: Okay, thank you. Question sixteen. For developing countries which cannot
Speaker 1: afford ground measurements and they have a little to no
Speaker 1: ground stations and there's limited technical capacity for advanced machine learning,
Speaker 1: what would you recommend as the ideal data and method
Speaker 1: to analyze air pollution there? So, I think that's, you know,
Speaker 1: partial part of the motivation for this training is the
Speaker 1: fact that both the SATPM and the bias corrected MERRIT
Speaker 1: two data sets are global, provide global coverage, and are
Speaker 1: publicly freely available everywhere in the world. The reason, but
Speaker 1: you know, the reasoning behind that is exactly to provide
Speaker 1: this kind of information to those who have more limited resources.
Speaker 1: So using these data sets that are already available to you,
Speaker 1: and wherever possible, comparing them to whatever local information you
Speaker 1: do have to establish their data quality would be kind
Speaker 1: of my recommendation for the best place to start your analysis. Okay,
Speaker 1: So for question seventeen, given random cafolds, spatially buffered and
Speaker 1: temporal holdout cross byalidation, which of these gives the most
Speaker 1: reliable estimate of error when the model is applied to
Speaker 1: unmonitored locations in future data? So I think the recommendation
Speaker 1: of best practice here is to use both a spatially
Speaker 1: buffered cross validation to assess the spatial generalizability, that is,
Speaker 1: the applicability to unmonitored locations, and a temporal holdout strategy
Speaker 1: to assess the temporal generalizability, so the applicability to future data.
Speaker 1: Question eighteen, will the retrievals be biased over airnet locations?
Speaker 1: I'm not sure if this refers to the aerosoloptical depth
Speaker 1: retrievals or the PM two point five data sets, but
Speaker 1: I think in either case there really shouldn't be any
Speaker 1: kind of bias specific to the airnet stations. In either
Speaker 1: of these methodologies. Question nineteen for deep investigations about the
Speaker 1: problems and possible solutions for Shalian countries and the Sahelian
Speaker 1: Alliance countries, what kind of partnership is possible for NASA
Speaker 1: R SET if the University of Thomas Sunkar University. So
Speaker 1: two points here. First, our set is a training program.
Speaker 1: We don't really do our own research and investigations. That's
Speaker 1: sort of something that's done separately by others at our set.
Speaker 1: But with that being said, you know, the people who
Speaker 1: participate in this training, myself and doctor von Donkler and
Speaker 1: doctor Gupta are involved in research around the world. So
Speaker 1: if you have specific questions or you know specific research
Speaker 1: topics which you think you want to collaborate on, you
Speaker 1: can reach out to us directly about that. Question twenty.
Speaker 1: Since process based constraints are interpolated smoothly from the course
Speaker 1: fifty kilometer model, what limitations or what are the limitations
Speaker 1: of using the downscaled one kilometer product for subgrid, micro
Speaker 1: environment or local buffer analyses in a highly complex urban
Speaker 1: area is Again I guess this is a question for
Speaker 1: Aaron and anything to add, you know, beyond what you've
Speaker 1: already said basically.
Speaker 2: No, I mean, it's a great question, as I mentioned before,
Speaker 2: I mean, actually one thing, I'm really glad to hear
Speaker 2: that you're thinking this way, whoever asked us questions, is
Speaker 2: a great way to think about the limitations of the model,
Speaker 2: and it really highlights why we talk about the value
Speaker 2: you know, as much as possible average to a larger area.
Speaker 2: So as much as we do have stysical models and
Speaker 2: things trying to capture that fine scale resolution down to
Speaker 2: one kilometer, it's quite reasonable to assume that those really
Speaker 2: really fine scale features will not be fully represented and
Speaker 2: they'll higher uncertainties. So it's a really good question we
Speaker 2: are relying on the statistical models to capture that really
Speaker 2: fine scale and how things will vary over that really
Speaker 2: really fine scale at least, how that relationship will over
Speaker 2: fine scale.
Speaker 1: Okay, thank you. Question twenty one, Again, this was a
Speaker 1: pretty detailed specific question and a bit beyond the scope
Speaker 1: of this training. I think this is sort of a
Speaker 1: whole research topic in and of itself, so we can't
Speaker 1: really provide a simple answer to it, unfortunately. So question
Speaker 1: twenty two, to what extent does visitor traffic inside and
Speaker 1: closed museum environments contribute to indoor PM two point five. Again,
Speaker 1: indoor PM concentrations is not really the be on the
Speaker 1: scope of this training, and I don't think it's really
Speaker 1: in any of our expertise unfortunately, so we can't really
Speaker 1: comment on that in a very informed way. So let's
Speaker 1: go to question twenty three next, scroll down a little bit.
Speaker 1: There we go, due to the lack of ground based
Speaker 1: measurements in Africa, would you expect there to be significant
Speaker 1: uncertainties in the derived PM two point five concentrations there
Speaker 1: from these data sets? Would you expect similar performance or
Speaker 1: performance to be similar, better or worse than merit to
Speaker 1: over Africa? So I think, yeah, maybe to Aaron and
Speaker 1: then to poll on about regionally specific differences in uncertainties
Speaker 1: and the products.
Speaker 2: Sure, so yeah, sorry, As the question you know quite
Speaker 2: rightly implies Africa is a certainly going to be a
Speaker 2: more uncertain region. There are some really there's some challenging
Speaker 2: conditions for models to capture both from major dust sources,
Speaker 2: major byomass bringing sources, and both these methods and painals
Speaker 2: weak because I think we are trying our best over
Speaker 2: those regions to speak to some of the strengths of
Speaker 2: the set PM data set. We really have that observational
Speaker 2: constraint from satellite, very just hardwired in what we're doing.
Speaker 2: The methods are designed in our by design using sparse
Speaker 2: ground based measurements as much as we have available to
Speaker 2: try and calibrate that better. It's definitely a region of
Speaker 2: more uncertainties, but I would still argue that what we're
Speaker 2: producing is probably one of the better estimates that is
Speaker 2: available out there.
Speaker 1: Okay, I just want to.
Speaker 3: Add some the Africa, as Aaron mentioned, is always very
Speaker 3: interesting reason and there's a lake up ground measurement. We
Speaker 3: all know about it, right, so it's very very hard
Speaker 3: to actually assess what kind of uncertaintives we might have.
Speaker 3: But I just want to let people know that there
Speaker 3: have been a lot of efforts going on in Africa
Speaker 3: in deploying low cost census although we know they are
Speaker 3: not best to use for this kind of validation, but
Speaker 3: they do provide useful information, and there are several groups
Speaker 3: in Africa and from other parts of the world who
Speaker 3: are actually trying and deploying and collecting the data. So
Speaker 3: we are hoping to actually collaborate with these groups to
Speaker 3: validate our products in future and provide more accurate assessment
Speaker 3: as we move forward in processing more data.
Speaker 1: Okay, thank you. So question twenty four was a question
Speaker 1: about actually like future funding opportunities, which is again not
Speaker 1: really something we're qualified to speak to, but I encourage
Speaker 1: you to take a look at the nspire's website, which
Speaker 1: is a place where all NASA funding opportunities are posted
Speaker 1: for you to look at. So question twenty five, in
Speaker 1: my own PM two point five machine learning framework SHAP
Speaker 1: analysis ranked my separate satellite AOD retrieval as the weakest
Speaker 1: predictor compared to continuous meteorological and spatial features. Have you
Speaker 1: observed similar trends where the model downweights physical AOD in
Speaker 1: favor of net data proxies. So I'll just speak maybe
Speaker 1: to my own experience, I would say that this is
Speaker 1: certainly possible in specific conditions that especially where PM two
Speaker 1: point five might be highly meteorology dependent. There's also the
Speaker 1: possibility that there are multiple layers of aerosols in the
Speaker 1: atmosphere and so the total AOD is really not reflective
Speaker 1: of the surface concentration very well. And also especially if
Speaker 1: you smoke, if you focus your analysis on a very
Speaker 1: small region where the meteorology is similar, and the PM
Speaker 1: two point five is also similar. Statistical methods might just
Speaker 1: be able to pick out on similar patterns and meteorology
Speaker 1: and AOD and made the connection between them much more
Speaker 1: easily than if you tried the same methodology on a
Speaker 1: larger region or even on a global scale, where the
Speaker 1: relationship between AOD and PM two point five is is
Speaker 1: much less consistent, and therefore other information sources like AOD
Speaker 1: would have a bigger role to play. Paul On or
Speaker 1: Aaron if you want to add on to that.
Speaker 2: No, no, I think you covered it. Sorry, go ahead,
Speaker 2: paul On.
Speaker 3: No, No, I was just saying, good thing.
Speaker 2: Well, you covered it very nicely, Carl.
Speaker 1: Okay, thanks, Okay. Question twenty six, how do we ensure
Speaker 1: the data is used accurately? Okay, this is the pretty
Speaker 1: broad question, but at a simple level, I'd say, you know,
Speaker 1: the this course, for one, is trying to convey to
Speaker 1: you basically what the limitations and the capabilities are of
Speaker 1: the data. So keeping those in mind as you use
Speaker 1: the data will help you to use it accurately. And
Speaker 1: as I think we've said several times, wherever possible, performing
Speaker 1: your own local assessments with locally collected data is really
Speaker 1: the best way to establish the quality of these data
Speaker 1: products for your specific region of interest. So anything to
Speaker 1: add on to that, I'll ho or eron.
Speaker 3: I can add one thing about specifically related to the
Speaker 3: statistical modeling or machine learning method. One thing which is
Speaker 3: very very critical whenever we do the statistical or machine
Speaker 3: learning method is the data distribution. So if you have
Speaker 3: several input parameter and you're trying to map with the
Speaker 3: output PM two point five, do check how the data
Speaker 3: has distributed. Because machine learning algorithm and statistical methods often
Speaker 3: try to learn what is around the mean and the
Speaker 3: extreme values become actually has less impact on the relationship
Speaker 3: which are which is formed, and often you will notice
Speaker 3: that these methods either underestimate the very high values because
Speaker 3: there are very few data points in your training and
Speaker 3: overestimate the low values because that's again very small values
Speaker 3: in your OWLD data set. So the data distribution, checking
Speaker 3: and adjusting your validation data versus the training data can
Speaker 3: be an important aspect of preparing your data before you
Speaker 3: start modeling. Whether you use a multiple linear regression or
Speaker 3: linear regression or machine learning method, it applies to all
Speaker 3: of those.
Speaker 1: Yeah, think that's a yeah, that's a really good point
Speaker 1: of the extreme values, especially if you have an a
Speaker 1: special interest in the extreme values as opposed to the mean,
Speaker 1: then that's that's an important caveat to keep in mind. Okay, yeah,
Speaker 1: let's maybe move on to the next question. Let's scroll
Speaker 1: down a little bit to question twenty seven. Will the
Speaker 1: mL based estimate capture events such as dust storms and
Speaker 1: smoke from fires? So, powin, do you want to speak
Speaker 1: to this a little bit?
Speaker 3: Sorry I was muted. I think it will capture the
Speaker 3: mara to as written here in the response, does assimilate
Speaker 3: satellite based aerosol optical depth and the fire emissions. And
Speaker 3: if that event is captured by the satellite is specifically
Speaker 3: from the modies, then the merit should have it captured
Speaker 3: and the machine learning based bias corrections should actually capture it.
Speaker 3: In some cases it might if the machine learning training
Speaker 3: didn't have dust or a smoke type of cases, then
Speaker 3: it might miss. So there are certain cases where smoke
Speaker 3: is very thick or dust plume is very thick, and
Speaker 3: the satellite sometimes get confused whether it is dust or
Speaker 3: cloud and mask it as a cloud and it does
Speaker 3: not get propagated into the assimilation stream and that can
Speaker 3: actually be not seen in the MARI too. So I
Speaker 3: would say it veries case by case, but in general
Speaker 3: it should be able to capture qualitatively. Quantitatively things can
Speaker 3: be very different, and depending on which event and how
Speaker 3: far we are looking from the location of the source,
Speaker 3: the quantitative numbers can be different. Qualitative lids should be captured.
Speaker 1: Okay, thank you. Question twenty eight, what are the main
Speaker 1: challenges in making PM two point five products like the
Speaker 1: Questionington University SAT PM two point five or the Meritus
Speaker 1: and then hay Cass product from retrospective processing to near
Speaker 1: real time to delivery. So this is a question about
Speaker 1: like the data latency and how maybe how that might
Speaker 1: be reduced in the future. So I guess yeah, Aaron, Aaron,
Speaker 1: go ahead, sure.
Speaker 2: So, as I wrote, I mean, that's a great question.
Speaker 2: We'd love to be able to get data sets, data
Speaker 2: sets out even faster. It's unforced, its often not practical.
Speaker 2: There is you know, as I've written here, as you
Speaker 2: gather from the presentation, is an awful lot of different
Speaker 2: data sets that come together to make these final PM
Speaker 2: twoint five data sets possible. From biomass burning emission inventories,
Speaker 2: which seems to be brought you brought into the kepital
Speaker 2: trust models models needs to be run. One very practical
Speaker 2: example of a delay that we face is to get
Speaker 2: the highest quality ground based monitors. Often governmental agencies will
Speaker 2: perform their own internal checks for usually about six months.
Speaker 2: So really for next year's release, you know which we're
Speaker 2: playing later this year, we're just starting to work with
Speaker 2: the ground based monitors now, which are the highest reference
Speaker 2: grade model that as well as our own internal validations
Speaker 2: take time. It's really multifaceted why it takes us so long,
Speaker 2: but I really feel that the time gives us an
Speaker 2: opportunity to reach the highest quality data product that we can.
Speaker 1: Okay, thank you, I on it.
Speaker 3: I think it applies to the matra TO or any
Speaker 3: other satellite product in the same ways because they all
Speaker 3: depends on some ancile data sets. Specifically in case of
Speaker 3: matter TO, it assimilates to so many data sets, so
Speaker 3: to get all those data prepared that for assimilation some time,
Speaker 3: and MERI to is re analysis is always I think
Speaker 3: several weeks behind because it's a re analysis. The idea
Speaker 3: is to retrospective analysis. So I think that is why
Speaker 3: we want to produce the best availabled data sets, and
Speaker 3: there are some real time and forecasting more data products
Speaker 3: which have a different way of processing.
Speaker 1: Okay, thank you, so question twenty nine this refers to
Speaker 1: a specific slide where we're talking about the index of
Speaker 1: agreement and asking about some of the other metrics. So
Speaker 1: some of the other slides did have different metrics, especially bias.
Speaker 1: And then also the paper describing the data set itself,
Speaker 1: which we'll link here obviously has a much more detailed
Speaker 1: description of the performance. Okay. Question thirty is this at
Speaker 1: PM data set available for North Africa? Yes, it is,
Speaker 1: and was the spatial and temporal resolution it's point zero
Speaker 1: one degree spatial end monthly ten or a resolution. We
Speaker 1: covered that in the presentations. Question thirty one, okay. Another
Speaker 1: regional question. In Indonesia, the days with the worst air pollution,
Speaker 1: for example hayes from fires are often the same days
Speaker 1: covered by clouds, so these days get removed from the
Speaker 1: data the AOD data. Specifically, how do we fix PM
Speaker 1: two point five estimates so they don't end up too
Speaker 1: low because of this? Aaron, maybe I think you answered
Speaker 1: this question here, so.
Speaker 2: Sure, I mean, I'll speak to how we address that
Speaker 2: within the SATPM data set, and maybe Poland has more
Speaker 2: to add after within sat PM. This is we account
Speaker 2: for these sampling bias in a few different ways. First off,
Speaker 2: the ground based monitors themselves we okay you said. First off,
Speaker 2: the sat AOD we directed to a monthly mean based
Speaker 2: upon sampling adjustment factors provided by the CTM, so again
Speaker 2: using that chepical transt model to infer how the days
Speaker 2: we observe relate to the days we haven't observed, and
Speaker 2: then providing a more representative overall monthly mean. Subsequent to that,
Speaker 2: the ground based monitors themselves will sample irrespective of that
Speaker 2: cloud cover. So when we apply those machine learning models
Speaker 2: or GBR models as a subsequent calibration to the geophysical estimates,
Speaker 2: that inherently also corrects for any sampling biases that may
Speaker 2: be from limitations to the satellite retrievals themselves. So there's
Speaker 2: a few ways in which this is hopefully addressed. None
Speaker 2: of them are obviously perfect, but there should not be
Speaker 2: inherently a sampling bias due to cloud covered days in
Speaker 2: our product.
Speaker 3: Okay, yeah, for Merit too, I think as it is
Speaker 3: a model produced data, so the data should be available
Speaker 3: even if there is our cloud. Now, in terms of
Speaker 3: the bias, I think the bias can arise is because
Speaker 3: if the satellite a woods are missing, then the model
Speaker 3: is relying completely on its own capability. I don't think
Speaker 3: we do any specific correction to actually account for that
Speaker 3: sampling bias there, but we haven't seen any clear evident
Speaker 3: in all validation studies that it should it should have
Speaker 3: a biases.
Speaker 1: Okay, thank you, So question two question thirty two, rather
Speaker 1: as a follow on, I think asking about which of
Speaker 1: these satellites would work best for Indonesia where there are
Speaker 1: cloud covered So just in terms of the this is
Speaker 1: maybe a little bit more general, but a geostationary satellite
Speaker 1: like Kimawari offering more frequent observations would give you more
Speaker 1: opportunities for cloud free observations. However, you know, any of
Speaker 1: these instruments would be I think similarly affected by the
Speaker 1: presence of clouds. And I also just want to note
Speaker 1: we've gone a little bit over our allotted time. If
Speaker 1: it's all right with with Aaron and Pallen, we might
Speaker 1: keep going for maybe ten more minutes and answer some
Speaker 1: more questions and then everything anything. We don't get to
Speaker 1: live we can answer in this document offline. Yeah, so
Speaker 1: let's keep going for a little bit longer. Question thirty four,
Speaker 1: you mentioned that multiliner regression is used to predict you
Speaker 1: a physical bias in PM two point five. Could you
Speaker 1: specify which independent variables or predictors are typically used in
Speaker 1: the set pm MLR model. So, I think to save time,
Speaker 1: we'll just sort of refer you to the publication where
Speaker 1: these are listed, and then in general they fall into
Speaker 1: the categories of physical and process based parameters, emissions, meteorology,
Speaker 1: and site location. Believe Aaron covered this in the slides
Speaker 1: as well. So question thirty five, within the workflow of
Speaker 1: the six ensemble predictive model developed for the Maritu sine
Speaker 1: n Heyksta, is that what is the scientific rationale for
Speaker 1: training five distinct convolitional neural network strategies rather than relying
Speaker 1: on a single baseline strategy basically the M one model. So,
Speaker 1: poland if you want to, yeah, speak to that freequently.
Speaker 3: Sure. Sure. So. I think, as we have discussed and
Speaker 3: seen in many cases, that the relationship between input parameter
Speaker 3: and the output target variable PM too point five is
Speaker 3: not linear and it varies as a function of aerosols type, meteorology,
Speaker 3: range of Pm two point five and many other parameters.
Speaker 3: And our goal was to what best we can do right,
Speaker 3: So in order to actually get the best of the
Speaker 3: data sets, we divided the data into different nominating aerosol
Speaker 3: component time so that we can get a better sense
Speaker 3: of how the input to output mapping can be done
Speaker 3: with more reliability. And that is why we actually try
Speaker 3: to play with different schemes and then we came up
Speaker 3: with these five different model which provide us an optimized solution.
Speaker 3: It may not be the best, but it provided an
Speaker 3: optimal solution which worked globally. And then the ensemble kind
Speaker 3: of bring them together in a more unified consistent data
Speaker 3: sets across the globe. So to use the strength of
Speaker 3: narrow too and their aerosol component, use those different approach
Speaker 3: to get an optimized global Team nine five data sient.
Speaker 1: All right, thank you. Question thirty six Quick clarification Slide
Speaker 1: sixty three notes that the mL generalization across region and
Speaker 1: time isn't guaranteed well. Slide seventy eight cosmeritive product robust
Speaker 1: onscing regions in years, given that it's trained on twenty
Speaker 1: eighteen to twenty nineteen dow date on twenty twenty. Could
Speaker 1: you help me understand what a bust onceing readings and
Speaker 1: years is based on Powen if you want to, Yeah.
Speaker 3: So, I think we continue to evaluate this data, although
Speaker 3: in those years which we have mentioned in the those
Speaker 3: are based on the published work which we did, but
Speaker 3: we continue to evaluate and I think we haven't seen
Speaker 3: any degression in the model performance over even beyond twenty
Speaker 3: twenty one and twenty twenty four, which we have done recently,
Speaker 3: which is not part of this paper. So I think
Speaker 3: what we can say is robust over time now over
Speaker 3: the region. The problem is the same as we have
Speaker 3: been discussing in earlier question is that if they're not
Speaker 3: ground stations available over certain regions, then it's very very
Speaker 3: hard for us to actually even predict what kind of
Speaker 3: uncertainties or error might be there. And that is why
Speaker 3: we try to provide a quality assurance flat with the
Speaker 3: data sets which basically rely on the ground density of
Speaker 3: the ground measurements and few other parameters which we are
Speaker 3: important in obour training algorithm, and that gives you an
Speaker 3: idea about whether a product is reliable in certain revision
Speaker 3: or not. So I think the quality assurance is one
Speaker 3: way to kind of see whether the product is robost
Speaker 3: or not.
Speaker 1: Okay, thank you. Question thirty seven, how do do you
Speaker 1: phittic physical models handle sudden extreme local emissions like crop
Speaker 1: burning or wildfires, especially when satellite AAD observations are missing
Speaker 1: due to cloud cover, Aaron, if you want to add
Speaker 1: anything on that.
Speaker 2: Sure mean this probably ties a little bit back into
Speaker 2: the earlier question why it takes so long to put
Speaker 2: these estimates together each time. So within the geophyscal model,
Speaker 2: it's obviously driven by a chemical transport model, and so
Speaker 2: in our case, and we use daily biomass burning emissions,
Speaker 2: and so those would capture within those daily inventories. They
Speaker 2: specific events like a sudden residual burning or wildfires, And
Speaker 2: those emission inventories themselves aren't necessarily dependent on the same limiting,
Speaker 2: don't have the same dependence sort of limitations, don't have
Speaker 2: the same limitations as the AOD retrieval. So a number
Speaker 2: of those daily biomass burning inventories can work as fine
Speaker 2: even when there is some cloud cover. Now that's not
Speaker 2: to say they're perfect, of course, I know crop resume
Speaker 2: burning has received a little bit more attention in a
Speaker 2: few papers I've read recently, thing that can be a
Speaker 2: higher uncertainty. But for the geophysical lestments, they're trying to
Speaker 2: within those daily inventories. And one of the reasons we
Speaker 2: of course enhance the geophysical lestments using the hybrid approach
Speaker 2: is to try to account for any misunderstandings or underrepresentations
Speaker 2: that occur within those inventories. So that's kind of how
Speaker 2: we we account for it.
Speaker 1: Okay, thank you. And question thirty eight another question about
Speaker 1: Indonesia due to the persistent cloud cover, which for the
Speaker 1: machine learning methods, which meteorological variables or predictors are the
Speaker 1: most crucial to include so the model doesn't become biased
Speaker 1: in a human tropical tropical region that experiences extreme humidity
Speaker 1: fluctuations during hotspot seasons. I guess Fallen, if you want
Speaker 1: to offer suggestions there or this might just be something
Speaker 1: that would need to be studied, but go ahead, go ahead, Yeah,
Speaker 1: let me.
Speaker 3: Just read the tropical system cloud co throat the US
Speaker 3: most moder Yeah. This this is a real problem, I guess.
Speaker 3: And it is not only over Indonesia.
Speaker 1: Uh.
Speaker 3: This is if you look across the ITCZ, which is
Speaker 3: has a consistent clouder cover even over South America or Africa.
Speaker 3: Part of that, and satellite the current passive satellite observation
Speaker 3: of aerosols will always have that problem because if there
Speaker 3: are cloud, the signal is so dominated by the cloud
Speaker 3: we cannot really get anything any information about aerosols. So
Speaker 3: I don't think there is any observational solution there. We
Speaker 3: will have to rely on the model for that, and
Speaker 3: more we can improve on the model, I think the
Speaker 3: better answer we can get. The other option is rely
Speaker 3: on ground measurement and active satellite based observation, which are
Speaker 3: a few of them right now in orbit and the
Speaker 3: few morees coming in the future which should be able
Speaker 3: to provide some more details or aerosol information in presence
Speaker 3: of cloud, and once they get assimilated into models, we
Speaker 3: will have a better idea what is happening in those reasons.
Speaker 1: Okay, thank you, And then I also saw another question
Speaker 1: just came in which was also referring to Indonesia. I
Speaker 1: think you have probably answered answered this in your response
Speaker 1: as well. About the general challenges so well, I think we'll,
Speaker 1: you know, we'll read through these and make sure that
Speaker 1: we've captured all of them. But I think for the
Speaker 1: sake of time, we're gonna I think call it call
Speaker 1: an end to the Q and A here. So thank
Speaker 1: you very much doctor Owen Dunkelar and doctor Kupta for
Speaker 1: contributing your time and your expertise to this training. And
Speaker 1: thank you very much to our attendees for your interest
Speaker 1: and all these very interesting and insightful questions about how
Speaker 1: to use these data sets appropriately. As a reminder, the
Speaker 1: next and final part of the training will be next week,
Speaker 1: and to prepare yourself with that training, you should test
Speaker 1: out that you're able to access Google Collab the sorry
Speaker 1: the have an Earth Data login password for accessing the
Speaker 1: NASA Earth data and also have an account with open
Speaker 1: Aq to access the open aq data set via their API,
Speaker 1: and there'll be some links to those. The slides will
Speaker 1: have links to those when they're posts the training website.
Speaker 1: So yeah, again, thank you everyone, and we will see
Speaker 1: you next week.
Podbean