conventions for anomaly data#600
Conversation
…ussion with David
|
Suggested revised text for section 7.5: An "anomaly" is the difference between a physical quantity and its statistical norm. For example, a commonly-used anomaly is the temperature at a specific time and place minus the long-term global average temperature. CF offers two conventions for describing anomaly data. In Section 7.5.3, "Temporal anomalies using anomaly standard names", we describe a simple convention that depends on special standard names. It can only be used for simple temporal anomalies, and is insufficient for some use-cases. In the remainder of this section and the following two (Section 7.5.1, "Anomalies with respect to a norm data variable" and Section 7.5.2, "Anomalies with respect to a norm metadata variable"), we describe two more general conventions for anomalies that can handle more complex use-cases. The generalized definition of an anomaly value A of some physical quantity q is the difference P - N between a particular value P of q and a normal value or norm N of q. N is some statistic calculated from the values of q that lie within specified ranges of one or more of its coordinates. P can be, but is not necessarily, one of the set of values from which N is calculated. In the same way, a data variable A containing anomalies with respect to a norm is notionally the difference between a data variable P containing the original data and a data variable N containing the statistical norm. P is typically not present in the dataset, and N is usually absent as well. The three variables have matching dimensions. P has all the same dimensions and coordinate variables as A, but N shares only a subset of the dimensions and coordinate variables of A. The other dimensions of N are the ones over which the norm is calculated. The commonest kind of anomaly A is a "temporal anomaly": the difference between the value P of a quantity and the mean N of the same quantity over some range of time coordinates, usually called the "climatological normal", the "climate normal", or the "climatology". N is most often either a time-mean over a continuous period of multiple years or a climatological time-mean (Section 7.4, "Climatological Statistics"). The time coordinate of the anomaly may or may not lie within the range of times from which N is calculated. N has all the same dimensions and coordinate variables as A except for time. The general convention described here is for anomalies with respect to a statistical norm calculated from any single dimension or combination of dimensions; moreover, the norm statistic does not have to be a mean. For example, anomalies might be calculated (as a function of longitude) with respect to the zonal mean, or (as a function of horizontal location) with respect to the minimum value in the area. In these examples, the norm is the zonal mean or the area minimum, respectively. When temporal anomalies are described following the convention of this section, more information can be recorded about the norm than when following the convention of Section 7.5.3, "Temporal anomalies using anomaly standard names". [cont'd] |
|
Under this convention, a data variable is described as an anomaly by including in its
Each name must be the name of an axis (i.e., a dimension and its corresponding coordinate variable, or a scalar coordinate variable) of the anomaly data variable A. We call these the "anomaly axes". They must have standard_name attributes. Usually there is only one anomaly axis, and usually it is a spatiotemporal axis. For instance, for a data variable containing anomalies with respect to the zonal mean, name identifies longitude as the anomaly axis e.g. " For each anomaly axis (although usually there is only one), N has a coordinate variable or scalar coordinate variable that indicates the range(s) of coordinates over which the statistic N was calculated from the variation of P. We call these (scalar) coordinate variables the "norm coordinate variables". For instance, if N contains zonal means, it has a norm coordinate variable for longitude that indicates the range of longitudes over which the zonal mean was calculated. There are two alternatives for norm. In both cases, norm is an ancillary variable of the anomaly data variable (Section 3.4, "Ancillary Data"). [cont'd] |
|
If the data variable containing the norm N (the "norm data variable") is present in the dataset, it can be named as norm in the The second alternative (Section 7.5.2, "Anomalies with respect to a norm metadata variable") is where norm in For example, the second method must be used for a timeseries of hourly mean anomalies with respect to a climatological hourly mean diurnal cycle, because the norm axis is multivalued and each anomaly value is relative to a different norm value (the one for the appropriate hour). Note that the anomaly time dimension may be larger than the norm climatological time dimension: in this example, the anomaly timeseries may be several days long, while the norm time dimension spans only one climatological day. [cont'd] |
|
7.5.1. Anomalies with respect to a norm data variable In this case, the word norm in The norm data variable must have all the same axes (dimension and coordinate variable or scalar coordinate variables) as the anomaly data variable, except for the anomaly axes. For each anomaly axis, the norm data variable must have either a scalar coordinate variable (named in its A norm coordinate variable cannot have a dimension greater than 1. Any norm coordinate dimensions must be included among the dimensions of the anomaly data variable as well as the norm data variable; likewise, any scalar norm coordinate variables must also be named in the coordinates attribute of the anomaly data variable. The norm data variable must have a [cont'd] |
|
A norm data variable for a climatological statistic (in the sense of Section 7.4, "Climatological Statistics") has a norm coordinate variable that must have
Example 7.15 shows how the In Example 7.15, the norm coordinate variable of time has just one element. If the climatological time axis is multivalued, a norm metadata variable is required (Section 7.5.2, "Anomalies with respect to a norm metadata variable"). [cont'd] |
|
Example 7.15. Distinguishing temporal anomalies with different kinds of norm The anomaly data variable [cont'd] |
|
The Another possibility is that the daily anomalies are calculated with respect to the 30-year July climatological mean, contained in Equivalently, the [cont'd] |
|
7.5.2. Anomalies with respect to a norm metadata variable In this case, the word norm in For each anomaly axis, the norm metadata variable must have a coordinate variable with the same [cont'd] |
|
In the case where N has a multivalued climatological time axis (such as those illustrated in Examples 7.9, 7.10, and 7.11), the norm metadata variable has, as its sole dimension, the corresponding climatological time dimension. In this case, the norm metadata variable has more than one element; in all others, it has only a single element. Regardless, it is a "dummy" variable whose data values are immaterial. In all cases, norm coordinate variables must have boundary variables that indicate the coordinate ranges over which N was calculated from P. The norm metadata variable must have a [cont'd] Note: at the end of the first sentence of the first paragraph, I changed "dimension of the anomaly coordinate variable of time" to "corresponding climatological time dimension" because I found "anomaly coordinate variable of time" very hard to parse; it sounds to me like it's referring to the time dimension of the anomaly variable, not to the time dimension of the norm metadata variable. I believe that what I've written has captured the concept correctly, but let me know if I've missed something. |
|
The use of norm metadata variables is illustrated by Examples 7.16, 7.17 and 7.18. The treatment of multivalued climatological time is described and illustrated after Example 7.18. Example 7.16. Temporal anomalies with a climatological norm metadata variable This example shows how the metadata of Example 7.15 can be recorded using a norm metadata variable. The data would be the same as in that example. A data variable If the daily anomalies are calculated with respect to the 30-year July climatological mean: As in Example 7.15, the metadata of [cont'd] |
|
Example 7.17. Anomalies with respect to a zonal mean The [cont'd] |
|
Example 7.18. Anomalies with respect to the minimum within a horizontal area The data variable [cont'd] |
|
If the norm has a multivalued climatological time axis, further information must be provided to describe the correspondence between elements of the anomaly and elements of the norm. For example, in the case of an anomaly relative to a monthly climatology, all of the January anomaly values will be relative to the average value for January, all the February anomalies will be relative to the average value for February, and so on. The mapping between the time axis of the anomaly variable and the climatological time axis of the norm is recorded by an auxiliary coordinate variable named by the The climatological time axis of In the following example, by "timestep In an abstract sense, the norm metadata variable indicates that Example 7.19 illustrates this convention, using the example described above. [cont'd] Comment: I had a hard time understanding this section, I think because the way it was originally written presumed proficiency with compression by gathering, which I've never used. I've rewritten it pretty extensively to try to factor that out into a separate piece of the explanation, so please check that I didn't lose anything. |
|
Example 7.19. An anomaly data variable whose norm has a multivalued climatological time coordinate variable The anomaly data variable Element 0 of |
|
7.5.3. Temporal anomalies using anomaly standard names Several CF standard names ending with Example 7.20. An anomaly data variable with a reference epoch The data variable In this example, |
|
This convention, using Example 7.21. Ambiguity in interpreting an anomaly data variable with a reference epoch The standard name One possibility is that Another possibility is that the daily anomalies are calculated with respect to the 30-year climatological mean for July: Without additional information about N, these possibilities (and others) cannot be distinguished using this convention. (Note that in all cases, |
|
|
||
| The second alternative (<<anomalies-norm-metadata>>) is where __norm__ in **`cell_methods`** identifies a "norm metadata variable" instead of the norm data variable. | ||
| A norm metadata variable contains information about the norm axes, but no data of its own. | ||
| This method can be used regardless of whether the norm data variable is present in the dataset as well. |
There was a problem hiding this comment.
"present in the dataset as well" sounds weird to me. I would prefer "also present in the dataset."
There was a problem hiding this comment.
OK, changed.
| The norm coordinate variable must have the same **`standard_name`** as the anomaly coordinate variable. | ||
| It must also have boundary variables to indicate the coordinate range over which __N__ was calculated from __P__. | ||
|
|
||
| Norm coordinate variables cannot have a dimension greater than 1, and these dimensions must be included among the dimensions of the anomaly data variable as well being dimensions of the norm data variable. |
There was a problem hiding this comment.
Typo: "as well being" -> "as well as being"
There was a problem hiding this comment.
Thanks, fixed.
| If either the anomaly data variable or the norm data variable has a **`standard_name`** attribute, it must __not__ be a standard name ending in **`_anomaly`**, and if they both have **`standard_name`** attributes, they must contain the same standard name. | ||
| The anomaly data variable must name the norm data variable in its **`ancillary_variables`** attribute (<<ancillary-data>>), as well as in **`cell_methods`**, in order to indicate the link between them. | ||
|
|
||
| The __norm__ data variable __N__ must have all the same axes as the anomaly data variable __A__, __except__ for the anomaly axes. |
There was a problem hiding this comment.
I think 'norm' should not be italicized here.
There was a problem hiding this comment.
Yes, quite right. I got trigger-happy with italics.
| Norm coordinate variables cannot have a dimension greater than 1, and these dimensions must be included among the dimensions of the anomaly data variable as well being dimensions of the norm data variable. | ||
| Likewise, any scalar norm coordinate variables must be named in the **`coordinates`** attribute of the anomaly data variable as well as the norm data variable. | ||
|
|
||
| The __norm__ data variable must have a **`cell_methods`** attribute with an entry for the norm coordinate variable (or the combination of them if more than one) to indicate how __N__ was computed from the variation of __P__. |
There was a problem hiding this comment.
I think norm should not be italicized here, either.
| Likewise, any scalar norm coordinate variables must be named in the **`coordinates`** attribute of the anomaly data variable as well as the norm data variable. | ||
|
|
||
| The __norm__ data variable must have a **`cell_methods`** attribute with an entry for the norm coordinate variable (or the combination of them if more than one) to indicate how __N__ was computed from the variation of __P__. | ||
| For instance, the __norm__ data variable for anomalies with respect to a time-mean must have a norm coordinate variable for time, and a **`cell_methods`** attribute containing an entry naming this variable. |
There was a problem hiding this comment.
Nor should norm be italicized here, IMO.
There was a problem hiding this comment.
I agree. I have checked (maybe you did too) and there are no other offending italics.
| The norm data variable is not necessarily present in the file, and even if it is present, this approach does not provide any link between the anomaly variable and the norm variable. | ||
| Hence, interpretation of the anomaly can be unclear. | ||
| Example 7.21 illustrates this ambiguity as it manifests in Example 7.20. | ||
| Such ambiguities can be resolved by the conventions of <<anomalies-norm-data>> and and <<anomalies-norm-metadata>>. |
| time:bounds="time_bounds"; | ||
| time:calendar="standard"; | ||
| double time_bounds(time,two); | ||
| double climatological_time; // norm coordinate variable |
There was a problem hiding this comment.
| double climatological_time; // norm coordinate variable | |
| double climatological_time; // norm and anomaly coordinate variable |
There was a problem hiding this comment.
Not done, as discussed
| climatological_time:calendar="standard"; | ||
| double climatological_time_bounds(two); | ||
| ---- | ||
| If the daily anomalies are calculated with respect to the 30-year July climatological mean: |
There was a problem hiding this comment.
Sorry - I un-apologise :) The problem is in the CDL below (this comment is on the caption for that)
|
Yes!
|
| dimensions: | ||
| time=6; | ||
| climatological_time=12; | ||
| variables: | ||
| float delta_tas(time,latitude,longitude); // anomaly data variable | ||
| delta_tas:standard_name="air_temperature"; | ||
| delta_tas:units="degC"; | ||
| delta_tas:units_metadata="temperature: difference"; | ||
| delta_tas:cell_methods="time: maximum time: anomaly_wrt climatological_tas"; | ||
| delta_tas:ancillary_variables="climatological_tas_metadata"; | ||
| int climatological_tas(time); // norm metadata variable | ||
| climatological_tas:coordinates="month_indices"; | ||
| climatological_tas:cell_methods="climatological_time: mean within years | ||
| climatological_time: mean over years"; | ||
| int month_indices(time); | ||
| month_indices:compress="climatological_time"; | ||
| double time(time); // anomaly coordinate variable | ||
| time:standard_name="time"; | ||
| time:units="days since 2023-06-01"; | ||
| time:bounds="time_bounds"; | ||
| time:calendar="standard"; | ||
| double time_bounds(time,two); | ||
| double climatological_time(climatological_time); | ||
| climatological_time:standard_name="time"; | ||
| climatological_time:units="days since 1990-01-01"; | ||
| climatological_time:bounds="climatological_time_bounds"; | ||
| climatological_time:calendar="standard"; | ||
| double climatological_time_bounds(climatological_time,two); | ||
| data: | ||
| time=15, 45, 76, 381, 411, 442; | ||
| // 2023-06-16, 2023-07-16, 2023-08-16, 2024-06-16, 2024-07-16, 2024-08-16 | ||
| time_bounds=0,30, 30,61, 61,92, 366,396, 396,427, 427,458; | ||
| // beginning and end of Jun, Jul and Aug of 2023 and 2024 | ||
| climatological_time=15, 45, ... 349; // 1990-01-16, 1990-02-15 ... 1990-12-16 | ||
| climatological_time_bounds=0,10623, 31,10651, ... 334,10957; | ||
| // 1990-01-01,2019-02-01, 1990-02-01,2019-03-01 ... 1990-12-01,2020-01-01 | ||
| month_indices=5, 6, 7, 5, 6, 7; |
There was a problem hiding this comment.
I'm somewhat lost with this example.
- I presume that
climatological_tasshould beclimatological_tas_metadata, as per the caption - The compression by gathering applies to the data variable, not the
climatological_timecoordinate variable - The compressed variable's dimension is also the name of a coordinate variable (
time), which causes problems when the variable is uncompressed to size 12
There was a problem hiding this comment.
OK - I've absorbed some of the new rules on how this works, and it's not really compression by gathering ... but it looks like it!
There was a problem hiding this comment.
I've made suggestions in compress->select above, which negates most of this comment thread, but:
- I presume that
climatological_tasshould beclimatological_tas_metadata, as per the caption
still stands, I think.
There was a problem hiding this comment.
I support using "select" instead of "compress". I see the original motivation for re-using the functionality (which is very similar), but I agree that it's better to use different names to distinguish that they have different meanings.
There was a problem hiding this comment.
Yes, fine, thanks. Also I have corrected climatological_tas_metadata. Thanks for noticing.
| The climatological time axis of __N__ (i.e. its dimension and coordinate variable) must be included in the file, although __N__ itself need not be present. | ||
| The auxiliary coordinate variable has a **`compress`** attribute naming the climatological time dimension, in order to make the link between them. | ||
| This method of mapping between axes is equivalent to <<compression-by-gathering>>. | ||
| Using this method means that the norm metadata variable can refer to a subset of elements of the climatological time axis if only some of them are relevant, and it can refer repeatedly to elements of climatological time axis if there is more than one anomaly time referring to a given climatological time. |
There was a problem hiding this comment.
I don't think that this is how compression by gathering works in chapter 8 works. Compression by gathering applies to the data of the variable that carries the compressed dimension, and also references the index variable. This seems to a brand new mechanism that allows you to cherry-pick elements from an existing coordinates.
Edit (18:18Z): I have suggested an alternative, that just entails not using "compress" (in favour of "select"), and spelling out how the selection occurs.
There was a problem hiding this comment.
I have changed compress to select and removed the sentence comparing the mechanism to compression by gathering. The way it works is described in the previous paragraph, which I have reordered thus:
The norm metadata variable for a multivalued climatological time axis has the time dimension of the anomaly data variable as its sole dimension, and is thus itself multivalued, but its values are arbitrary and meaningless. The mapping between the anomaly time axis of the anomaly variable and the climatological time axis of the norm is recorded by an auxiliary coordinate variable of integer type named by the
coordinatesattribute of the norm metadata variable. The auxiliary coordinate variable is one-dimensional and has the same time dimension as the anomaly data variable. The climatological time axis of N (i.e. its dimension and coordinate variable) must be included in the file, although N itself need not be present. The auxiliary coordinate variable has aselectattribute naming the climatological time dimension, in order to make the link between them. The value of element i of the auxiliary coordinate variable is the index (numbering from 0) along the climatological time dimension of the norm corresponding to element i of the anomaly time dimension. Using this method means that the norm metadata variable can refer to a subset of elements of the climatological time axis if only some of them are relevant, and it can refer repeatedly to elements of climatological time axis if there is more than one anomaly time referring to a given climatological time.
| Size-one dimensions of the norm metadata variable must also be dimensions of the anomaly data variable. | ||
| The norm metadata variable must have no coordinate variables or scalar coordinate variables other than the norm coordinate variables, which are of size one. | ||
| Therefore the norm metadata variable has only one element. | ||
| It must have a **`_FillValue`** attribute, and its single element must be equal to the **`_FillValue`**, to indicate that it contains no meaningful data. |
There was a problem hiding this comment.
I don't think we should mandata a _FillValue attribute. The test is that it contains only missing values, and shouldn't get into the mine field of how the missing values are encoded.
There was a problem hiding this comment.
OK. I have therefore removed the recommendation that it "not have any of the attributes of <<attribute-appendix>> other than cell_methods, coordinates and _FillValue" because it's too complicated to rephrase this including several attributes which might be relevant to missing data. I've left it as "[The norm metadata variables] does not need attributes describing the norm quantity (standard name, units, etc.), because they must be the same as for the anomaly data variable," and "its single element must indicate missing data."
| | C | ||
| | <<compression-by-gathering>>, <<reduced-horizontal-grid>> | ||
| | Records dimensions which have been compressed by gathering. | ||
| | <<reduced-horizontal-grid>>, <<anomalies-norm-metadata>>, <<compression-by-gathering>> |
There was a problem hiding this comment.
As per the compress->select suggestion
| | <<reduced-horizontal-grid>>, <<anomalies-norm-metadata>>, <<compression-by-gathering>> | |
| | <<reduced-horizontal-grid>>, <<compression-by-gathering>> |
There was a problem hiding this comment.
Also inserted a new entry in the table:
| **`select`**
| S
| C
| <<anomalies-norm-metadata>>
| Identifies a dimension to which the values in this variable are indices.
| The norm metadata variable also has the same time dimension as the anomaly data variable (rather than being a scalar), but as usual, its values are arbitrary and meaningless. | ||
|
|
||
| The climatological time axis of __N__ (i.e. its dimension and coordinate variable) must be included in the file, although __N__ itself need not be present. | ||
| The auxiliary coordinate variable has a **`compress`** attribute naming the climatological time dimension, in order to make the link between them. |
There was a problem hiding this comment.
| The auxiliary coordinate variable has a **`compress`** attribute naming the climatological time dimension, in order to make the link between them. | |
| The auxiliary coordinate variable has a **`select`** attribute naming the climatological time dimension, in order to make the link between them. |
There was a problem hiding this comment.
OK. select is a good name for it. It is not quite the same as compress anyway, because it allows repeated indices.
|
|
||
| The climatological time axis of __N__ (i.e. its dimension and coordinate variable) must be included in the file, although __N__ itself need not be present. | ||
| The auxiliary coordinate variable has a **`compress`** attribute naming the climatological time dimension, in order to make the link between them. | ||
| This method of mapping between axes is equivalent to <<compression-by-gathering>>. |
There was a problem hiding this comment.
| This method of mapping between axes is equivalent to <<compression-by-gathering>>. | |
| The auxiliary coordinate variable implies the existence of another auxiliary coordinate variable of the same size, not in the file, whose data values are selected from the climatological time coordinate variable according to the climatological time axis index positions given in the data array. |
There was a problem hiding this comment.
I'm sorry, I'm too tired to understand this sentence, so I've done only the deletion. The function of the aux coord var is explained in the previous paragraph.
| climatological_tas:cell_methods="climatological_time: mean within years | ||
| climatological_time: mean over years"; | ||
| int month_indices(time); | ||
| month_indices:compress="climatological_time"; |
There was a problem hiding this comment.
As part of the compress->select suggestion:
| month_indices:compress="climatological_time"; | |
| month_indices:select="climatological_time"; |
| See also the **`add_offset`** attribute. | ||
| In cases where there is a strong constraint on dataset size, it is allowed to pack the coordinate variables (using add_offset and/or scale_factor), but this is not recommended in general. | ||
|
|
||
| | **`source`** |
There was a problem hiding this comment.
As per the compress->select suggestion
| | **`select`** | |
| | S | |
| | C | |
| | <<anomalies-norm-metadata>> | |
| | Identifies other dimension to which the values in this variable are indices, used for selection of climatological time coordinates for norm metadata variables. | |
| | **`source`** |
Co-authored-by: David Hassell <davidhassell@users.noreply.github.com>
See issue #582 for discussion of these changes.
Release checklist
history.adocup to date?