Showing posts with label moving averages. Show all posts
Showing posts with label moving averages. Show all posts

Monday, February 5, 2024

Why bother with seasonal adjustment?

Why bother with seasonal adjustment, extreme adjustment, and smoothing?  Because often the unprocessed data are not very clear.

The chart below shows the original (unadjusted), seasonally and extreme-adjusted, and the smoothed data for Turkey's total vehicle sales.

Notice the extreme seasonality.   

If you just looked at the original data, you might conclude from December 2023's (seasonal) spike that vehicle sales were doing very well.   But in fact, on a seasonally adjusted basis, vehicle sales are down from earlier this year, which is consistent with the sharp monetary tightening by the Bank of Türkiye over the last six months.   In fact, the BOT started raising interest rates in July, which is also when the seasonally adjusted sales peaked.

The smoothed version of the seasonally adjusted data has turned down, suggesting that this is a downturn which will last more than a couple of months.  The smoothed data peaked in July at 105,000 units and have fallen steadily since to 103,000 units in December.

 

Click to see clearer image

Here's what the data look like over a shorter time period:

Click to see clearer image


I should point out that the seasonal and extreme adjustment are done by programs I've written, based on algorithms created by the US Bureau of the Census.  I make no manual adjustments; it's all automated.  However, I can choose the length and type of the smoothing.  In this case, I used a 3-month centred linear moving average of a 9-month centred linear moving average with weights calculated by the Bureau of the Census for the X-11 seasonal adjustment program.  This is called a 3 x 9 moving average.  "Linear" refers to the pattern of weights, which is the same for each month in the central span of the time series.  You can also use quadratic or quartic moving averages.  The moving averages are "centred", because a conventional moving average lags the turning points in the underlying data. If you use different moving averages for different time series, because they have different "spikiness", you would get spurious variations in cyclical turning points between the series, not derived from the underlying data, but a function of the length of the moving average you've chosen. 

Quite technical stuff, but I thought you might be interested. 

Thursday, June 2, 2022

US average May PMIs down

 As usual, focus on the average of the two national purchasing manager indices, the green line.  This reduces the statistical 'noise' or the random measurement fluctuations around the underlying trend. As often happens, over the month the two surveys moved in opposite directions, this time, the ISM up and the PMI down.  But the average continues to slide.

Note that I have extreme-adjusted both series, and that has attenuated the sharp 2-month downward spike due to Covid in early 2020.




Sunday, May 3, 2020

Wreckage

The ISM—the original and longest-running US survey of manufacturing conditions—fell fast in April.  My extreme adjustment algorithm reduced the extent of the fall as you can see in the chart below.  The program in effect assumes that the downward spike in April will be at least partially reversed in May.  That may not be true.  If the May number is also so weak, my extreme-adjustment  algorithm will accept the May number unchanged.   During the SARS outbreak in 2003, the downward spike was brief, and the extreme-adjustment program rightly corrected it.  This time round, especially in the US, the decline will probly not be that short.

Unadjusted, the ISM fell to a point last seen during the GFC, in 2008/2009.



As usual, the average of the ISM and the PMI produces a time series with less random fluctuation.  That's the green line in the chart below.  Both component series have been extreme adjusted, so the comments applicable above to the ISM also applies to this.

 

Monday, August 5, 2019

Contradictory stories

Let's suppose that you ask a random 1 in 100 people in your city what they feel about, I dunno, elephants.  Now let's suppose you ask a new random sample a month later what they think about elephants.  You'll get a different answer, even if the opinions of the underlying population haven't changed at all.  This is random fluctuation.   It has nothing to do with fluctuations in the underlying reality.  It's an artifact of the estimation process.

There are ways of reducing the influence of random variations.  First, you can fit some kind of moving average.  The assumption is that the random error each month is independent of the error for the next or the previous month.  Therefore it should, over time, average out to zero.  In principle, therefore, to minimise the random error, you should take a longer moving average.  But that would also hide the variability of the underlying data too, which would mean you might miss turning points until it long after they've happened.

The second way is a variant of fitting a moving average.  It's called extreme adjustment.  What this does is to fit a moving average, then estimate the error term, i.e., the deviation between the moving average and the original data.  Any observation where the error term is larger than 2 standard deviations away from zero is replaced with a value closer to the assumed underlying time series, i.e., the moving average.  Other observations are however left unchanged.  This works well for smooth time series with occasional large fluctuations, but "spiky" time series have such large "error terms" (random fluctuations) that larger fluctuations are simply assumed to be normal.

The third way is to use time series with different sample universes.  For example, you can get a very good idea of the performance of the US economy by adding together the volume of retail sales, the volume of industrial production, and total non-agricultural employment.  The assumption here is the error terms of each time series are not correlated with each other, which means that their average will be closer to zero.  We will have a better insight into the underlying economic movements from their average than from each series on its own.

Here's an excellent example of random fluctuations apparently producing quite different stories about what is actually happening.  It comes from two headlines on the same news page of TradingEconomics.com.

  • The Commonwealth Bank of Australia Services PMI was revised higher to 52.3 in July 2019 from a preliminary estimate of 51.9 and compared to the previous month's final figure of 52.6. Output and new orders rose at a slower pace, despite a solid increase in new export business. 

  • The AIG Australian Performance of Services Index plunged 8.3 points from the previous month to 43.9 in July 2019, pointing to the steepest month of contraction in the service sector since November 2014. All five activity indexes in the Australian PSI were negative in July: sales (-8.0 points to 45.1); employment (-3.8 points to 43.8); new orders (-12.7 points to 44.1); supplier deliveries (-12.3 points to 40.5); and finished stocks (-3.4 points to 46.2). 

Which is the better guide to what's happening?  How can we tell, since we don't know the underlying reality—we can only infer it from the data?  We should prolly use both.  So let's add them together and see what we get.  The chart below shows the AIG services PMI vs the Commonwealth Bank services PMI together with an average of those two, and a 3 month by 9 month moving average of the average.  You still with me?



The CBA survey is the smoother of the two (presumably because of a larger sample size).  But it doesn't go back as far as the AIG survey.  Note that there's a good relationship between the CBA and AIG surveys, with the AIG survey much more volatile ("spiky")  than the CBA one, though turning points are not identical.  Perhaps this is because of overlapping sample spaces, since some survey correspondents may be in both surveys.  Note that the average of the two is less "spiky" than the AIG's alone.  The moving average of the average of the two surveys says that there is a slow upturn taking place, despite the falls in July in both series.  It's probably our best guide to what is actually happening.  Except ....  a different moving average might give a different result (see appendix below)

I use all these methods to try and divine what's happening in economies.  Composite and diffusion indices or even just simple averages using many time series as components will more truly reflect the underlying reality than a single series.  Extreme adjustment or  moving averages make it easier to see trends.  But the lesson is also not to just look at the latest observation.  It doesn't matter if July's value is down if June's value is up by a larger amount, and vice versa.  The media often just tell you the latest data release, when what you need to know is the trend over recent months or years.  Statistics can be very misleading if you don't know their context.


APPENDIX: CENTRED VS TRAILING MOVING AVERAGES

The simplest kind of moving average is just the sum of the last N periods divided by N.  (There are more complex moving averages which use a pattern of weights which resemble a quadratic or quartic, but we won't go there today!)  Such a moving average lags the putative turning point by half the span of N.  This is useful for technical analysis in markets, where we might use, say, a 30-day moving average and a 120-day moving average and compare the two.  But for economic time series we want to see clearly when we have entered or exited a recession or a boom.   For that we use a centred moving average.  This is particularly important if we are using moving averages of different lengths (spans) because the underlying time series have different levels of "spikiness".  We might perhaps wish to compare an unsmoothed series with one to which we have fitted a moving average.

With a 9-term moving average, for example, we shift it backwards by 4 months to centre it.   You can see the effect of that in the chart below, which shows the US ISM index and two 9 month moving averages, one trailing and one centred.  Now, note that at the end of the centred moving average, we have "lost" observations.  One way of filling these gaps is to estimate the most recent moving average values using  shorter moving averages.  For example to pad a 9-term centred moving average we could use a 7-term moving average for the 4th observation from the end, 5-term for the 3rd, 3-term for the second, and the actual data point for the last observation.  But for technical reasons this proves unsatisfactory.  Instead we use a pattern of weights.  For example, the most recent value for a 9-term centred moving average is  35% times the last observation, 25% the second last, 21.2% the third last, 15.4% the 4th last, 9.6% the fifth last, 3.8%% the sixth last, 0% for the seventh last, 0% for the eighth last and 0% for the ninth last.  The second most recent value uses a different set of weights, and so on.  This produces an approximation of the 9-term moving average as if we had an additional 4 data points.  However, it is subject to revision: as a new data point becomes available, the moving average weights shift.   







Wednesday, June 26, 2019

Fooled by noise

I talked here about a post by Tamino who writes Open Mind, which discussed signal and noise, and how easy it was to get mislead by short-term random changes in the data of a time series.   Such considerations are why I often show extreme-adjusted or smoothed data, because these processes allow us to disentangle random fluctuations from the underlying trend.

He has written another post about this issue called Fooled by Noise.  I've summarised it below.


The chart below shows NASA's calculation of the global temperature anomaly. Observe how much the data fluctuate.  Hard to detect trends!  But some denialists maintain that global temperatures have started dropping again, over the last couple of years.  Of course, we can see similar downward spikes right through the last 140 years, only to be reversed later.  So how can we get a better handle on the data?



This chart shows the same data, except it's an annual average.  By taking a 12-month average, the random fluctuations are reduced.  Though the "spikiness" of the data has been reduced (i.e., the random fluctuations have been reduced) there are still some fluctuations, plus also some obvious longer waves or cycles.



The chart below shows even more smoothing.  It plots the 5 year averages of the monthly anomalies.  Note, now, how clear the trends are.  Global temperatures fell from circa 1880 to 1910, then rose steadily until the early 1940s, before levelling off until the late 70s.  The rise then resumed with a vengeance, and the most recent 5 year period shows a sharp jump.  The stabilisation of global temperatures from 1945 to 1975 despite rising levels of CO₂ is attributed to the equally rapid rise in the emissions of sulphur dioxide.  When acid rain was brought under control by scrubbing the smoke from factory and power stations flues of SO₂ in the late 70s, the underlying trend due to CO₂ was quickly revealed.  Note that you can't blame the most recent spike on El Niño.  The equally strong El Niño in 1998 barely tweaks the 5 year average.


The decadal averages are even clearer.  Once again El Niño is undetectable in the data.



You can apply the same analysis to the extent of Arctic sea ice.  Here is the annual chart, and below that is the chart of the decadal averages.



The moral of the story is that when denialists say "Arctic sea ice is recovering" or "global temperatures have stopped rising" what they are citing as evidence is unsmoothed data, still embodying random fluctuations.  We can only be sure global temperatures have peaked or the extent of Arctic sea ice has troughed when the decadal averages have turned, because these averages dramatically reduce random data errors, so we can see the true trends.  And they haven't: global temperatures are still rising and Arctic sea ice is still declining.  Nor would it be logical for them to do so,  except temporarily after a massive volcanic eruption, when sulphur dioxide is released into the atmosphere, lowering global temperatures. And as you can see from the decadal averages, previous eruptions cannot be seen in the data, though they do appear in the annual data, e.g., Mt Pinatubo in 1991.


Tuesday, March 21, 2017

Extremes

Why would the number of extremely hot days rise by a much larger percentage than the rise in average temperatures?

It's to do with the shape of the statistical distribution.  Let's suppose that the average temperature for a season or a year is X, and that the observations are normally distributed.  That means that the number of times  the temperature is X-10, X-9, X-8 ...... X+8, X+9, X+10 (just as an example)  starts out with very few times it's X-10 (or X+10), a few more times when it's X-9 (or X+9), still more times it's X-8 (or X+8), in a shape which is reminiscent of a bell.  The normal distribution is often called the bell curve because of this.

There are other distributions which are skewed to one end of the curve or the other, which means that they are asymmetrical.  For example, wealth is very asymmetrically distributed: there are far fewer rich people who own most of the wealth, and lots of poor people who own little.  But observations for weather tend to be symmetrically distributed, but because of global warming, they have drifted over time.  The distribution is still bell-shaped but the mean  (the average) has shifted higher.  And that means that the "tails" of the bell shape have also shifted.  So whereas in the old days you had say 1 very hot day 3 days a year, 1 % of the time, now you have 2 or 3 or 4 times as many very hot days.  Similarly, in the old days you used to get 1 very cold day 3 days a year, now you get none.  The graphic below demonstrates this principle very nicely.


Source


Lots of people think that a rise of 1 degree C in average temperatures is small, because they are used to the shifts in temperature from day to night, and from winter to summer, which are obviously much larger.  But what global warming means is that we will have far more hot days in summer, with temperatures reaching and exceeding new highs more often and by more.  We have just seen this in our Australian summer, and they've just seen this in the US, where February heat has been at July levels,  If the mean, and its whole distribution shifts still higher the number of super hot days will double or treble or quadruple, with devastating consequences for us and for our world.

Saturday, November 19, 2016

2016 another record

It seems clear that 2016 is going to be another record year for global temperatures, the third in a row.  The chart below, from Better Nature, plots the 12 month moving average of the monthly anomaly, plus a 132 month (11 year) moving average, plus a linear OLSQ trend from 1970 to 2016.  The dots show the annual (January-December) average.  Obviously, there isn't yet one for 2016, because we don't have data for November or December.

The 11 year moving average is very close to the linear trend, suggesting that an 11 year moving average is a pretty good approximation.  And, the rise over the last 4 years is larger than the rise leading up to the last big El Niño in 1998, so if global temperatures don't fall as they did after the last El Niño, then the slope of the 11 year moving average will steepen, hinting that the trend also has shifted.  If they do fall, of course the denialists will be back with their rubbish "temperatures haven't risen for the last X years."  But the thing to watch is the 11 year moving average.

I'll keep you posted.

(Source)

Monday, February 1, 2016

Signal versus noise

Everybody who is in financial markets knows about this problem.  Is the latest low US GDP data point the beginning of a new slower trend or is it a blip?  Is the stock market still going up--recent trends have been down, so is this a change in longer-term trends or what?  What we normally do to estimate trend in economics or shares is to use a moving average.  And you look at the fluctuations around the moving average.  Is there a pattern of rising highs and lows, or is each new high and each new low below the previous one?

What struck me about the charts in this article was how it's impossible to determine the trend in the unsmoothed data, but how clear the trend is in the smoothed data.  The chart shows the temperature record for central England.  Since 1945 it's been adjusted for the urban heat island effect, i.e., it has been reduced to compensate for the fact that cities are much warmer than the surrounding countryside.

Here's the chart of the unsmoothed data, with the December 2015 anomaly highlighted in the red circles:



Here's a simple 30 year moving average of the data:



Note that the moving average is only available up to 2001, because traditionally you centre the moving average at the midpoint of its span.  So the last observation on the chart above is the average from January 1986 to December 2015.  This is a problem with moving averages--you "lose" data at the beginning and end of your underlying data series.  There are some mathematical techniques you can use to get both smoother and more up to date moving averages.  One of these is the LOWESS moving average. (in these circumstances,  I use a Henderson curve, which is a bit less sophisticated and easier to calculate.)



You can see why those who don't understand statistics will be flummoxed by the big random day-to-day and month-to-month fluctuations. But the thing to focus on is the longer-term trends, and these are clearly shown by the moving averages.