Sunday, May 14, 2023
Just how big is China?
Sunday, November 7, 2021
The pandemic's true death toll
From The Economist
How many people have died because of the covid-19 pandemic? The answer depends both on the data available, and on how you define “because”. Many people who die while infected with SARS-CoV-2 are never tested for it, and do not enter the official totals. Conversely, some people whose deaths have been attributed to covid-19 had other ailments that might have ended their lives on a similar timeframe anyway. And what about people who died of preventable causes during the pandemic, because hospitals full of covid-19 patients could not treat them? If such cases count, they must be offset by deaths that did not occur but would have in normal times, such as those caused by flu or air pollution.
Rather than trying to distinguish between types of deaths, The Economist’s approach is to count all of them. The standard method of tracking changes in total mortality is “excess deaths”. This number is the gap between how many people died in a given region during a given time period, regardless of cause, and how many deaths would have been expected if a particular circumstance (such as a natural disaster or disease outbreak) had not occurred. Although the official number of deaths caused by covid-19 is now 5m, our single best estimate is that the actual toll is 17.1m people. We find that there is a 95% chance that the true value lies between 10.5m and 19.7m additional deaths.
Monday, August 5, 2019
Contradictory stories
There are ways of reducing the influence of random variations. First, you can fit some kind of moving average. The assumption is that the random error each month is independent of the error for the next or the previous month. Therefore it should, over time, average out to zero. In principle, therefore, to minimise the random error, you should take a longer moving average. But that would also hide the variability of the underlying data too, which would mean you might miss turning points until it long after they've happened.
The second way is a variant of fitting a moving average. It's called extreme adjustment. What this does is to fit a moving average, then estimate the error term, i.e., the deviation between the moving average and the original data. Any observation where the error term is larger than 2 standard deviations away from zero is replaced with a value closer to the assumed underlying time series, i.e., the moving average. Other observations are however left unchanged. This works well for smooth time series with occasional large fluctuations, but "spiky" time series have such large "error terms" (random fluctuations) that larger fluctuations are simply assumed to be normal.
The third way is to use time series with different sample universes. For example, you can get a very good idea of the performance of the US economy by adding together the volume of retail sales, the volume of industrial production, and total non-agricultural employment. The assumption here is the error terms of each time series are not correlated with each other, which means that their average will be closer to zero. We will have a better insight into the underlying economic movements from their average than from each series on its own.
Here's an excellent example of random fluctuations apparently producing quite different stories about what is actually happening. It comes from two headlines on the same news page of TradingEconomics.com.
- The Commonwealth Bank of Australia Services PMI was revised higher to 52.3 in July 2019 from a preliminary estimate of 51.9 and compared to the previous month's final figure of 52.6. Output and new orders rose at a slower pace, despite a solid increase in new export business.
- The AIG Australian Performance of Services Index plunged 8.3 points from the previous month to 43.9 in July 2019, pointing to the steepest month of contraction in the service sector since November 2014. All five activity indexes in the Australian PSI were negative in July: sales (-8.0 points to 45.1); employment (-3.8 points to 43.8); new orders (-12.7 points to 44.1); supplier deliveries (-12.3 points to 40.5); and finished stocks (-3.4 points to 46.2).
Which is the better guide to what's happening? How can we tell, since we don't know the underlying reality—we can only infer it from the data? We should prolly use both. So let's add them together and see what we get. The chart below shows the AIG services PMI vs the Commonwealth Bank services PMI together with an average of those two, and a 3 month by 9 month moving average of the average. You still with me?
The CBA survey is the smoother of the two (presumably because of a larger sample size). But it doesn't go back as far as the AIG survey. Note that there's a good relationship between the CBA and AIG surveys, with the AIG survey much more volatile ("spiky") than the CBA one, though turning points are not identical. Perhaps this is because of overlapping sample spaces, since some survey correspondents may be in both surveys. Note that the average of the two is less "spiky" than the AIG's alone. The moving average of the average of the two surveys says that there is a slow upturn taking place, despite the falls in July in both series. It's probably our best guide to what is actually happening. Except .... a different moving average might give a different result (see appendix below)
I use all these methods to try and divine what's happening in economies. Composite and diffusion indices or even just simple averages using many time series as components will more truly reflect the underlying reality than a single series. Extreme adjustment or moving averages make it easier to see trends. But the lesson is also not to just look at the latest observation. It doesn't matter if July's value is down if June's value is up by a larger amount, and vice versa. The media often just tell you the latest data release, when what you need to know is the trend over recent months or years. Statistics can be very misleading if you don't know their context.
APPENDIX: CENTRED VS TRAILING MOVING AVERAGES
The simplest kind of moving average is just the sum of the last N periods divided by N. (There are more complex moving averages which use a pattern of weights which resemble a quadratic or quartic, but we won't go there today!) Such a moving average lags the putative turning point by half the span of N. This is useful for technical analysis in markets, where we might use, say, a 30-day moving average and a 120-day moving average and compare the two. But for economic time series we want to see clearly when we have entered or exited a recession or a boom. For that we use a centred moving average. This is particularly important if we are using moving averages of different lengths (spans) because the underlying time series have different levels of "spikiness". We might perhaps wish to compare an unsmoothed series with one to which we have fitted a moving average.
With a 9-term moving average, for example, we shift it backwards by 4 months to centre it. You can see the effect of that in the chart below, which shows the US ISM index and two 9 month moving averages, one trailing and one centred. Now, note that at the end of the centred moving average, we have "lost" observations. One way of filling these gaps is to estimate the most recent moving average values using shorter moving averages. For example to pad a 9-term centred moving average we could use a 7-term moving average for the 4th observation from the end, 5-term for the 3rd, 3-term for the second, and the actual data point for the last observation. But for technical reasons this proves unsatisfactory. Instead we use a pattern of weights. For example, the most recent value for a 9-term centred moving average is 35% times the last observation, 25% the second last, 21.2% the third last, 15.4% the 4th last, 9.6% the fifth last, 3.8%% the sixth last, 0% for the seventh last, 0% for the eighth last and 0% for the ninth last. The second most recent value uses a different set of weights, and so on. This produces an approximation of the 9-term moving average as if we had an additional 4 data points. However, it is subject to revision: as a new data point becomes available, the moving average weights shift.
Wednesday, June 26, 2019
Fooled by noise
He has written another post about this issue called Fooled by Noise. I've summarised it below.
The chart below shows NASA's calculation of the global temperature anomaly. Observe how much the data fluctuate. Hard to detect trends! But some denialists maintain that global temperatures have started dropping again, over the last couple of years. Of course, we can see similar downward spikes right through the last 140 years, only to be reversed later. So how can we get a better handle on the data?
This chart shows the same data, except it's an annual average. By taking a 12-month average, the random fluctuations are reduced. Though the "spikiness" of the data has been reduced (i.e., the random fluctuations have been reduced) there are still some fluctuations, plus also some obvious longer waves or cycles.
The chart below shows even more smoothing. It plots the 5 year averages of the monthly anomalies. Note, now, how clear the trends are. Global temperatures fell from circa 1880 to 1910, then rose steadily until the early 1940s, before levelling off until the late 70s. The rise then resumed with a vengeance, and the most recent 5 year period shows a sharp jump. The stabilisation of global temperatures from 1945 to 1975 despite rising levels of CO₂ is attributed to the equally rapid rise in the emissions of sulphur dioxide. When acid rain was brought under control by scrubbing the smoke from factory and power stations flues of SO₂ in the late 70s, the underlying trend due to CO₂ was quickly revealed. Note that you can't blame the most recent spike on El Niño. The equally strong El Niño in 1998 barely tweaks the 5 year average.
The decadal averages are even clearer. Once again El Niño is undetectable in the data.
You can apply the same analysis to the extent of Arctic sea ice. Here is the annual chart, and below that is the chart of the decadal averages.
The moral of the story is that when denialists say "Arctic sea ice is recovering" or "global temperatures have stopped rising" what they are citing as evidence is unsmoothed data, still embodying random fluctuations. We can only be sure global temperatures have peaked or the extent of Arctic sea ice has troughed when the decadal averages have turned, because these averages dramatically reduce random data errors, so we can see the true trends. And they haven't: global temperatures are still rising and Arctic sea ice is still declining. Nor would it be logical for them to do so, except temporarily after a massive volcanic eruption, when sulphur dioxide is released into the atmosphere, lowering global temperatures. And as you can see from the decadal averages, previous eruptions cannot be seen in the data, though they do appear in the annual data, e.g., Mt Pinatubo in 1991.
Tuesday, December 12, 2017
Record highs far outpacing record lows
So if you do see a pattern of ever higher highs, or ever lower lows, whether it's in market indices, or share prices, or climate records, it means there is a trend. A trend doesn't mean it's higher/lower every year, it means that over time new highs/lows keep on being attained. One year may be lower than the previous year without meaning that an upward trend has changed. On the other hand, if successive observations do show a new pattern of falling highs and lower lows, then it becomes more and more likely that a new trend is in place.
![]() |
| Source: Climate Central |
Daily record highs are vastly outpacing daily record lows in the U.S. We will always have warm years and cold years, but in a world without global warming, those warm and cold years would balance over time. However, that’s not what we are seeing. According to the 2017 U.S. Climate Science Special Report, after a rigorous reanalysis of GHCN stations back to 1930, 15 of the last 20 years had more daily record highs than daily record lows. The number of daily record highs outpaced daily record lows more than 4 to 1 in 1998, 2012, and 2016.
A first look at the data from NOAA/NCEI indicates that 2017 continues the warming trend, as daily record highs are beating daily record lows by a 3.5-to-1 margin so far. Below are some preliminary 2017 stats through the end of November. Visit the NOAA Daily Weather Records tool to get the daily updates on these numbers:
- Monthly record highs have outnumbered monthly record lows at a rate of 9.7 to 1
- All-time record highs have outnumbered all-time record lows 8.7 to 1
- Record high minimum temperatures have outnumbered record low minimums 4.6 to 1
As the concentration of greenhouse gases increases in the atmosphere, the ratio of record highs to record lows will likely increase even further. With no change in current emissions trends, model projections indicate that record highs could outpace record lows by 15 to 1 by the end of the century.
[Read more here]
The evidence that climate change is happening right now, and that we don't have to wait for decades to see its effects, is mounting. Records aren't just falling for temperatures, but also for severity and length of droughts, rainfall events, and hurricanes. Almost everybody, even if they are concerned about global warming, thinks it's something which will only affect our grandchildren. But it's happening right now. By the time our grandchildren are adult, it will be incomparably worse, unless we move aggressively and rapidly to de-carbonise our economies. Fortunately, technology and the costs of renewables are helping facilitate this transformation, but we need to speed the process up, and also address areas where such technological progress is not being made: air travel, cement, iron & steel, land clearing, shipping.
Monday, January 2, 2017
Averages
![]() |
| (Source) |
However, when you aggregate the demand from millions of households, it becomes reasonably predictable from day to day and month to month. The chart below shows Californian electricity demand on a hot day in 1999 (it doesn't show net demand, that is, demand net of power generated by solar panels, which produces the "peaking duck" curve)
![]() |
| [Source] |
There are still fluctuations of course, with daytime demand peaking at twice the level that demand falls to at its low point in the wee hours of the morning, but the curves are much smoother. On hot days, conveniently, the demand peak partially coincides with a likely supply peak from solar. With a few hours of storage, solar could provide for all daytime demand, at least for places between latitudes 35 or 40 N and S.
Now the key point here is that if the grid had to be built to supply the peak power demands of each individual household, electricity would be extremely expensive. That's why "going off the grid" is too expensive to do, no matter how tempted we are by the outrageous electricity prices and network fees charged by our utilities. An individual house or small business would need far too much capacity to make sure it never ran out of power. For the grid as a whole, far less spare capacity is needed, because individual peaks and troughs are averaged out. This spare capacity is often provided by gas peaking plants, which can start up quickly and turn off quickly whereas most "baseload" power (coal and nuclear) can't be easily scaled up or down.
The same averaging process is true when you are looking at individual sources of supply too, including the variable supply from renewable generating sources. Take a look at the chart below (from this website) It shows the wind output of various Australian wind farms expressed as a percentage of total capacity. The black line shows total Australian wind farm output, also as a percentage of total capacity. See how individual wind farms can fluctuate from nearly 100% of capacity down to 0% in the space of a few hours, but how the total of all wind farms across the country (the black line) varies between 30% and 50%. On average, over time, Oz wind farms produce 30 to 35% of nameplate capacity.
![]() |
| (Source) |
This is all relevant because when we're working out how much backup or storage we need, there are some who maintain that each wind or solar farm needs 100% backup from a conventional fossil fuel generator. And of course, this makes wind and solar appear impossibly expensive (that's the point, I suspect). But we don't actually need that much back up or storage. Here're the views of the CSIRO (Commonwealth Scientific and Industrial Research Organisation), an Australian independent government-funded research body:
CSIRO Energy chief economist Paul Graham says that additional storage is not needed for up to 40 to 50 per cent wind and solar penetration. That’s because the grid can rely on existing back-up ( built to meet peaks in demand and for when coal and gas “baseload” generators trip or need to be repaired).
Beyond those levels, storage needs to be part of the equation. But again, not as much as many would think. But as the back-up generators gradually exit the grid, they can be replaced by various storage types, until storage then becomes the principal form of back-up and grid security on the grid.The CSIRO modelling showed that at very high levels of wind and solar, a maximum of half a day’s average demand was needed for storage. In some areas of the grid, only around three hours might be needed.
This is an important point, because some renewable critics say that about a week’s worth of storage is needed, and multiples of wind and solar capacity required for back up. These would be the same people that argue that climate science is a hoax, but it is a view that has more traction than it should.
Graham says the CSIRO modelling indicated that at those very high levels, about 0.8GW of back-up was required for about every GW of wind and solar capacity. This is around the same amount of back up capacity currently needed by centralised power plants to meet peak demand and outages.
[Read more here; my emphasis]
What's true for Australia is likely true for most other geographies where power is generated from a variety of sources, except as I discuss here, countries in high latitudes where the majority of renewable power will come from wind. Though note that the CSIRO's forecasts assume that more than half of Australia's power will be wind generated by 2050.
Lazard have recently updated their LCOEs for different generation sources, and for the first time have provided an estimate of solar with 10 hours of storage. They come up with a cost of US$92/MWh. Average coal (in the US, ignoring the implicit cost of pollution, i.e., without a carbon tax) is $100/MWh. Combined cycle natural gas is cheaper, averaging US$63/MWh, but as the chart below shows, natural gas prices fluctuate widely. And if--when--fracking is curtailed because of its environmental costs, the price is likely to rise further. So in future gas may not be cheaper than solar plus batteries, even without a carbon tax.
![]() |
| [Source] |
You might notice that though output from an generator fuelled by fossil fuels is stable, apart from outages, its costs are not stable; while on the other hand output from renewable energy sources is variable, but its costs are fixed, because its "fuel" is free. A sort of Heisenberg uncertainty principle for electricity generation. But utilities and users value price stability as much as they value supply reliability, Even if gas is cheaper, they might still find the price and supply stability promised by renewables plus storage as very attractive.
Lazard has made a much more detailed costing of storage than I was able to. What it shows is that for most of the world, even when you add in the necessary storage, solar can provide us with all our electricity needs. And when I say most of the world, I mean India, China, Africa, the southern US and Europe, Australia and South America., as you can see from the chart below. These regions can all be powered from solar plus storage. Right now. As cheaply as or more cheaply than coal. And with more stable costings than gas. Even without a carbon tax. These locations might also have wind energy, and that's a bonus, because the output from wind plus solar is on average less variable than the output of either alone. But they can do it on solar. Right now.
Global temperatures are rising by 0.2 deg C per decade on average. If it takes us 2 decades to switch our entire generation fleet to renewables, and to electrify most transport, temperatures will still have risen by another 0.4 deg C. The 1 deg C rise we've already had has caused enough problems. We can't afford to wait, and now we have no reason to either. Switching to renewables is now affordable. And it won't cost us the earth.
Monday, February 1, 2016
Signal versus noise
What struck me about the charts in this article was how it's impossible to determine the trend in the unsmoothed data, but how clear the trend is in the smoothed data. The chart shows the temperature record for central England. Since 1945 it's been adjusted for the urban heat island effect, i.e., it has been reduced to compensate for the fact that cities are much warmer than the surrounding countryside.
Here's the chart of the unsmoothed data, with the December 2015 anomaly highlighted in the red circles:
Here's a simple 30 year moving average of the data:
Note that the moving average is only available up to 2001, because traditionally you centre the moving average at the midpoint of its span. So the last observation on the chart above is the average from January 1986 to December 2015. This is a problem with moving averages--you "lose" data at the beginning and end of your underlying data series. There are some mathematical techniques you can use to get both smoother and more up to date moving averages. One of these is the LOWESS moving average. (in these circumstances, I use a Henderson curve, which is a bit less sophisticated and easier to calculate.)
You can see why those who don't understand statistics will be flummoxed by the big random day-to-day and month-to-month fluctuations. But the thing to focus on is the longer-term trends, and these are clearly shown by the moving averages.
Wednesday, July 23, 2014
Global Warming Stats
But you can also look at the mean, the moving average itself, to get an idea of what's happening. In the chart below (via NOAA), I've plotted the 60 month (five year) moving average to June. After record global temperature in May and June I suspected that the 5 year average would be high, and it is. In fact it's equal hottest to the five years to June 2007. And it's clear that there isn't much of a "pause" in the inexorable rise in temperatures.
So what about the long pause from the 40s to the end of the 70s? My guess is that that was due to the releases of aerosols (sulphur dioxide) as industrialisation proceeded apace. Aerosols reduce global warming. When the impacts of acid rain became clearer over time, sulphur dioxide emissions were slashed, but carbon dioxide emissions kept on rising, so global temps resumed their rise. But much of the world's industrial production growth in the last decade or so has come in China and India, where aerosol emissions are prodigious (look at the images of Chinese skies). This has been one factor (I suspect: others are ENSO and the sunspot cycle) ) keeping the world cooler just as it was (I suspect) between 1940 and 1970. Which means that as China and India clean up their air, by cutting CO2 and SO2 emissions, global temps will resume their uptrend even as CO2 emissions peak and start to decline. A terrifying prospect.


















