(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 1 — #1
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
Springer Texts in StatisticsRobert H. ShumwayDavid S. Stoff erTime Series Analysis and Its ApplicationsWith R Examples Fourth Edition (cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 2 — #2
(cid:105)
(cid:105)
Robert H. Shumway
David S. Stoffer
Time Series Analysis and
Its Applications
With R Examples
Fourth Edition
live free or bark
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page v — #3
(cid:105)
(cid:105)
Preface to the Fourth Edition
The fourth edition follows the general layout of the third edition but includes some
modernization of topics as well as the coverage of additional topics. The preface to
the third edition—which follows—still applies, so we concentrate on the differences
between the two editions here. As in the third edition, R code for each example
is given in the text, even if the code is excruciatingly long. Most of the examples
with seemingly endless coding are in the latter chapters. The R package for the text,
astsa, is still supported and details may be found in Appendix R. A number of data
sets have been updated. For example, the global temperature deviation series have
been updated to 2015 and are included in the newest version of the package; the
corresponding examples and problems have been updated accordingly.
Chapter 1 of this edition is similar to the previous edition, but we have included
the definition of trend stationarity and the the concept of prewhitening when using
cross-correlation. The New York Stock Exchange data set, which focused on an old
financial crisis, was replaced with a more current series of the Dow Jones Indus-
trial Average, which focuses on a newer financial crisis. In Chapter 2, we rewrote
some of the regression review, changed the smoothing examples from the mortality
data example to the Southern Oscillation Index and finding El Niño. We also ex-
panded the discussion of lagged regression to Chapter 3 to include the possibility of
autocorrelated errors.
In Chapter 3, we removed normality from definition of ARMA models; while the
assumption is not necessary for the definition, it is essential for inference and pre-
diction. We added a section on regression with ARMA errors and the corresponding
problems; this section was previously in Chapter 5. Some of the examples have been
modified and we added some examples in the seasonal ARMA section.
In Chapter 4, we improved and added some examples. The idea of modulated
series is discussed using the classic star magnitude data set. We moved some of the
filtering section forward for easier access to information when needed. We removed
the reliance on spec.pgram (from the stats package) to mvspec (from the astsa
package) so we can avoid having to spend pages explaining the quirks of spec.pgram,
which tended to take over the narrative. The section on wavelets was removed because
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page vi — #4
(cid:105)
(cid:105)
vi
Preface to the Fourth Edition
there are so many accessible texts available. The spectral representation theorems are
discussed in a little more detail using examples based on simple harmonic processes.
The general layout of Chapter 5 and of Chapter 7 is the same, although we have
revised some of the examples. As previously mentioned, we moved regression with
ARMA errors to Chapter 3.
Chapter 6 sees the biggest change in this edition. We have added a section on
smoothing splines, and a section on hidden Markov models and switching autore-
gressions. The Bayesian section is completely rewritten and is on linear Gaussian
state space models only. The nonlinear material in the previous edition is removed
because it was old, and the newer material is in Douc, Moulines, and Stoffer (2014).
Many of the examples have been rewritten to make the chapter more accessible. Our
goal was to be able to have a course on state space models based primarily on the
material in Chapter 6.
The Appendices are similar, with some minor changes to Appendix A and Ap-
pendix B. We added material to Appendix C, including a discussion of Riemann–
Stieltjes and stochastic integration, a proof of the fact that the spectra of autoregressive
processes are dense in the space of spectral densities, and a proof of the fact that spec-
tra are approximately the eigenvalues of the covariance matrix of a stationary process.
We tweaked, rewrote, improved, or revised some of the exercises, but the overall
ordering and coverage is roughly the same. And, of course, we moved regression with
ARMA errors problems to Chapter 3 and removed the Chapter 4 wavelet problems.
The exercises for Chapter 6 have been updated accordingly to reflect the new and
improved version of the chapter.
Davis, CA
Pittsburgh, PA
September 2016
Robert H. Shumway
David S. Stoffer
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page vii — #5
(cid:105)
(cid:105)
Preface to the Third Edition
The goals of this book are to develop an appreciation for the richness and versatility
of modern time series analysis as a tool for analyzing data, and still maintain a
commitment to theoretical integrity, as exemplified by the seminal works of Brillinger
(1975) and Hannan (1970) and the texts by Brockwell and Davis (1991) and Fuller
(1995). The advent of inexpensive powerful computing has provided both real data
and new software that can take one considerably beyond the fitting of simple time
domain models, such as have been elegantly described in the landmark work of Box
and Jenkins (1970). This book is designed to be useful as a text for courses in time
series on several different levels and as a reference work for practitioners facing the
analysis of time-correlated data in the physical, biological, and social sciences.
We have used earlier versions of the text at both the undergraduate and gradu-
ate levels over the past decade. Our experience is that an undergraduate course can
be accessible to students with a background in regression analysis and may include
Section 1.1–Section 1.5, Section 2.1–Section 2.3, the results and numerical parts of
Section 3.1–Section 3.9, and briefly the results and numerical parts of Section 4.1–
Publicité
Section 4.4. At the advanced undergraduate or master’s level, where the students
have some mathematical statistics background, more detailed coverage of the same
sections, with the inclusion of extra topics from Chapter 5 or Chapter 6 can be used
as a one-semester course. Often, the extra topics are chosen by the students according
to their interests. Finally, a two-semester upper-level graduate course for mathemat-
ics, statistics, and engineering graduate students can be crafted by adding selected
theoretical appendices. For the upper-level graduate course, we should mention that
we are striving for a broader but less rigorous level of coverage than that which is
attained by Brockwell and Davis (1991), the classic entry at this level.
The major difference between this third edition of the text and the second edition
is that we provide R code for almost all of the numerical examples. An R package
called astsa is provided for use with the text; see Section R.2 for details. R code
is provided simply to enhance the exposition by making the numerical examples
reproducible.
We have tried, where possible, to keep the problem sets in order so that an
instructor may have an easy time moving from the second edition to the third edition.
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page viii — #6
(cid:105)
(cid:105)
viii
Preface to the Third Edition
However, some of the old problems have been revised and there are some new
problems. Also, some of the data sets have been updated. We added one section
in Chapter 5 on unit roots and enhanced some of the presentations throughout the
text. The exposition on state-space modeling, ARMAX models, and (multivariate)
regression with autocorrelated errors in Chapter 6 have been expanded. In this edition,
we use standard R functions as much as possible, but we use our own scripts (included
in astsa) when we feel it is necessary to avoid problems with a particular R function;
these problems are discussed in detail on the website for the text under R Issues.
We thank John Kimmel, Executive Editor, Springer Statistics, for his guidance
in the preparation and production of this edition of the text. We are grateful to Don
Percival, University of Washington, for numerous suggestions that led to substantial
improvement to the presentation in the second edition, and consequently in this
edition. We thank Doug Wiens, University of Alberta, for help with some of the R
code in Chapter 4 and Chapter 7, and for his many suggestions for improvement of
the exposition. We are grateful for the continued help and advice of Pierre Duchesne,
University of Montreal, and Alexander Aue, University of California, Davis. We also
thank the many students and other readers who took the time to mention typographical
errors and other corrections to the first and second editions. Finally, work on the this
edition was supported by the National Science Foundation while one of us (D.S.S.)
was working at the Foundation under the Intergovernmental Personnel Act.
Davis, CA
Pittsburgh, PA
September 2010
Robert H. Shumway
David S. Stoffer
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page ix — #7
(cid:105)
(cid:105)
Contents
Preface to the Fourth Edition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
v
Preface to the Third Edition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
vii
1
2
3
Characteristics of Time Series . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.1 The Nature of Time Series Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.2 Time Series Statistical Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.3 Measures of Dependence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.4 Stationary Time Series . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.5 Estimation of Correlation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1.6 Vector-Valued and Multidimensional Series . . . . . . . . . . . . . . . . . . . . .
Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
Time Series Regression and Exploratory Data Analysis . . . . . . . . . . . . .
2.1 Classical Regression in the Time Series Context . . . . . . . . . . . . . . . . .
2.2 Exploratory Data Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
2.3 Smoothing in the Time Series Context . . . . . . . . . . . . . . . . . . . . . . . . .
Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
1
2
8
14
19
26
33
38
47
47
56
67
72
77
ARIMA Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
77
3.1 Autoregressive Moving Average Models . . . . . . . . . . . . . . . . . . . . . . .
90
3.2 Difference Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3.3 Autocorrelation and Partial Autocorrelation . . . . . . . . . . . . . . . . . . . . .
96
3.4 Forecasting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
3.5 Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115
3.6 Integrated Models for Nonstationary Data . . . . . . . . . . . . . . . . . . . . . . 133
3.7 Building ARIMA Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
3.8 Regression with Autocorrelated Errors . . . . . . . . . . . . . . . . . . . . . . . . 145
3.9 Multiplicative Seasonal ARIMA Models . . . . . . . . . . . . . . . . . . . . . . . 148
Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page x — #8
(cid:105)
(cid:105)
x
4
5
6
7
Contents
Publicité
Spectral Analysis and Filtering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
4.1 Cyclical Behavior and Periodicity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
4.2 The Spectral Density . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174
4.3 Periodogram and Discrete Fourier Transform . . . . . . . . . . . . . . . . . . . 181
4.4 Nonparametric Spectral Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191
4.5 Parametric Spectral Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205
4.6 Multiple Series and Cross-Spectra . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208
4.7 Linear Filters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213
4.8 Lagged Regression Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 218
4.9 Signal Extraction and Optimum Filtering . . . . . . . . . . . . . . . . . . . . . . . 223
4.10 Spectral Analysis of Multidimensional Series . . . . . . . . . . . . . . . . . . . 227
Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 230
Additional Time Domain Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 241
5.1 Long Memory ARMA and Fractional Differencing . . . . . . . . . . . . . . 241
5.2 Unit Root Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250
5.3 GARCH Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 253
5.4 Threshold Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261
5.5 Lagged Regression and Transfer Function Modeling . . . . . . . . . . . . . 265
5.6 Multivariate ARMAX Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271
Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 284
State Space Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 287
6.1 Linear Gaussian Model
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 288
6.2 Filtering, Smoothing, and Forecasting . . . . . . . . . . . . . . . . . . . . . . . . . 292
6.3 Maximum Likelihood Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 302
6.4 Missing Data Modifications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 310
6.5 Structural Models: Signal Extraction and Forecasting . . . . . . . . . . . . 315
6.6 State-Space Models with Correlated Errors . . . . . . . . . . . . . . . . . . . . . 319
6.6.1 ARMAX Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 320
6.6.2 Multivariate Regression with Autocorrelated Errors . . . . . . . 322
6.7 Bootstrapping State Space Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 325
6.8 Smoothing Splines and the Kalman Smoother . . . . . . . . . . . . . . . . . . . 331
6.9 Hidden Markov Models and Switching Autoregression . . . . . . . . . . . 334
6.10 Dynamic Linear Models with Switching . . . . . . . . . . . . . . . . . . . . . . . 345
6.11 Stochastic Volatility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357
6.12 Bayesian Analysis of State Space Models . . . . . . . . . . . . . . . . . . . . . . 365
Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375
Statistical Methods in the Frequency Domain . . . . . . . . . . . . . . . . . . . . . 383
7.1
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383
7.2 Spectral Matrices and Likelihood Functions . . . . . . . . . . . . . . . . . . . . 386
. . . . . . . . . . . . . . . . . . . . . . . 388
7.3 Regression for Jointly Stationary Series
7.4 Regression with Deterministic Inputs
. . . . . . . . . . . . . . . . . . . . . . . . . 397
7.5 Random Coefficient Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 405
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page xi — #9
(cid:105)
(cid:105)
Contents
xi
7.6 Analysis of Designed Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 407
7.7 Discriminant and Cluster Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . 421
7.8 Principal Components and Factor Analysis . . . . . . . . . . . . . . . . . . . . . 437
7.9 The Spectral Envelope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 453
Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 464
Appendix A Large Sample Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 471
A.1 Convergence Modes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 471
A.2 Central Limit Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 478
A.3 The Mean and Autocorrelation Functions . . . . . . . . . . . . . . . . . . . . . . 482
Appendix B Time Domain Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491
B.1 Hilbert Spaces and the Projection Theorem . . . . . . . . . . . . . . . . . . . . . 491
B.2 Causal Conditions for ARMA Models . . . . . . . . . . . . . . . . . . . . . . . . . 495
B.3 Large Sample Distribution of the AR Conditional Least Squares
Estimators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 497
B.4 The Wold Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 500
Appendix C Spectral Domain Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503
C.1 Spectral Representation Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503
C.2 Large Sample Distribution of the Smoothed Periodogram . . . . . . . . . 507
C.3 The Complex Multivariate Normal Distribution . . . . . . . . . . . . . . . . . 517
C.4 Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 522
C.4.1 Riemann–Stieltjes Integration . . . . . . . . . . . . . . . . . . . . . . . . . . 522
C.4.2 Stochastic Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 524
C.5 Spectral Analysis as Principal Component Analysis . . . . . . . . . . . . . . 525
C.6 Parametric Spectral Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 529
Appendix R R Supplement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531
R.1 First Things First . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531
R.2 astsa . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531
R.3 Getting Started . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 532
R.4 Time Series Primer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 536
R.4.1 Graphics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 539
References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 541
Index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 553
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page xii — #10
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 1 — #11
(cid:105)
(cid:105)
Chapter 1
Characteristics of Time Series
The analysis of experimental data that have been observed at different points in time
leads to new and unique problems in statistical modeling and inference. The obvi-
ous correlation introduced by the sampling of adjacent points in time can severely
restrict the applicability of the many conventional statistical methods traditionally
dependent on the assumption that these adjacent observations are independent and
identically distributed. The systematic approach by which one goes about answer-
ing the mathematical and statistical questions posed by these time correlations is
commonly referred to as time series analysis.
The impact of time series analysis on scientific applications can be partially doc-
umented by producing an abbreviated listing of the diverse fields in which important
time series problems may arise. For example, many familiar time series occur in the
field of economics, where we are continually exposed to daily stock market quota-
tions or monthly unemployment figures. Social scientists follow population series,
such as birthrates or school enrollments. An epidemiologist might be interested in
the number of influenza cases observed over some time period. In medicine, blood
pressure measurements traced over time could be useful for evaluating drugs used
Publicité
in treating hypertension. Functional magnetic resonance imaging of brain-wave time
series patterns might be used to study how the brain reacts to certain stimuli under
various experimental conditions.
In our view, the first step in any time series investigation always involves careful
examination of the recorded data plotted over time. This scrutiny often suggests
the method of analysis as well as statistics that will be of use in summarizing the
information in the data. Before looking more closely at the particular statistical
methods, it is appropriate to mention that two separate, but not necessarily mutually
exclusive, approaches to time series analysis exist, commonly identified as the time
domain approach and the frequency domain approach. The time domain approach
views the investigation of lagged relationships as most important (e.g., how does
what happened today affect what will happen tomorrow), whereas the frequency
domain approach views the investigation of cycles as most important (e.g., what is
the economic cycle through periods of expansion and recession). We will explore
both types of approaches in the following sections.
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 2 — #12
(cid:105)
(cid:105)
2
1 Characteristics of Time Series
Fig. 1.1. Johnson & Johnson quarterly earnings per share, 84 quarters, 1960-I to 1980-IV.
1.1 The Nature of Time Series Data
Some of the problems and questions of interest to the prospective time series analyst
can best be exposed by considering real experimental data taken from different subject
areas. The following cases illustrate some of the common kinds of experimental time
series data as well as some of the statistical questions that might be asked about such
data.
Example 1.1 Johnson & Johnson Quarterly Earnings
Figure 1.1 shows quarterly earnings per share for the U.S. company Johnson &
Johnson, furnished by Professor Paul Griffin (personal communication) of the
Graduate School of Management, University of California, Davis. There are 84
quarters (21 years) measured from the first quarter of 1960 to the last quarter of
1980. Modeling such series begins by observing the primary patterns in the time
history. In this case, note the gradually increasing underlying trend and the rather
regular variation superimposed on the trend that seems to repeat over quarters.
Methods for analyzing data such as these are explored in Chapter 2 and Chapter 6.
To plot the data using the R statistical package, type the following:1.1
library(astsa)
SEE THE FOOTNOTE
plot(jj, type="o", ylab="Quarterly Earnings per Share")
Example 1.2 Global Warming
Consider the global temperature series record shown in Figure 1.2. The data are the
global mean land–ocean temperature index from 1880 to 2015, with the base period
1951-1980. In particular, the data are deviations, measured in degrees centigrade,
from the 1951-1980 average, and are an update of Hansen et al. (2006). We note an
apparent upward trend in the series during the latter part of the twentieth century
that has been used as an argument for the global warming hypothesis. Note also
the leveling off at about 1935 and then another rather sharp upward trend at about
1.1 Throughout the text, we assume that the R package for the book, astsa, has been installed and loaded.
See Section R.2 for further details.
(cid:105)
(cid:105)
(cid:105)
(cid:105)
TimeQuarterly Earnings per Share19601965197019751980051015llllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllll(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 3 — #13
(cid:105)
(cid:105)
1.1 The Nature of Time Series Data
3
Fig. 1.2. Yearly average global temperature deviations (1880–2015) in degrees centigrade.
Fig. 1.3. Speech recording of the syllable aaa · · · hhh sampled at 10,000 points per second
with n = 1020 points.
1970. The question of interest for global warming proponents and opponents is
whether the overall trend is natural or whether it is caused by some human-induced
interface. Problem 2.8 examines 634 years of glacial sediment data that might be
taken as a long-term temperature proxy. Such percentage changes in temperature
do not seem to be unusual over a time period of 100 years. Again, the question of
trend is of more interest than particular periodicities. The R code for this example
is similar to the code in Example 1.1:
plot(globtemp, type="o", ylab="Global Temperature Deviations")
Example 1.3 Speech Data
Figure 1.3 shows a small .1 second (1000 point) sample of recorded speech for
the phrase aaa · · · hhh, and we note the repetitive nature of the signal and the
rather regular periodicities. One current problem of great interest is computer
recognition of speech, which would require converting this particular signal into
the recorded phrase aaa · · · hhh. Spectral analysis can be used in this context to
produce a signature of this phrase that can be compared with signatures of various
(cid:105)
(cid:105)
(cid:105)
(cid:105)
TimeGlobal Temperature Deviations18801900192019401960198020002020−0.40.00.40.8llllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllTimespeech0200400600800100001000200030004000(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 4 — #14
(cid:105)
(cid:105)
4
1 Characteristics of Time Series
Fig. 1.4. The daily returns of the Dow Jones Industrial Average (DJIA) from April 20, 2006 to
April 20, 2016.
library syllables to look for a match. One can immediately notice the rather regular
repetition of small wavelets. The separation between the packets is known as the
pitch period and represents the response of the vocal tract filter to a periodic
sequence of pulses stimulated by the opening and closing of the glottis. In R, you
can reproduce Figure 1.3 using plot(speech).
Example 1.4 Dow Jones Industrial Average
As an example of financial time series data, Figure 1.4 shows the daily returns
(or percent change) of the Dow Jones Industrial Average (DJIA) from April 20,
2006 to April 20, 2016. It is easy to spot the financial crisis of 2008 in the figure.
The data shown in Figure 1.4 are typical of return data. The mean of the series
appears to be stable with an average return of approximately zero, however, highly
volatile (variable) periods tend to be clustered together. A problem in the analysis
of these type of financial data is to forecast the volatility of future returns. Models
such as ARCH and GARCH models (Engle, 1982; Bollerslev, 1986) and stochastic
volatility models (Harvey, Ruiz and Shephard, 1994) have been developed to handle
these problems. We will discuss these models and the analysis of financial data in
Chapter 5 and Chapter 6. The data were obtained using the Technical Trading Rules
(TTR) package to download the data from YahooTM and then plot it. We then used
the fact that if xt is the actual value of the DJIA and rt = (xt −xt−1)/xt−1 is the return,
then 1 + rt = xt /xt−1 and log(1 + rt ) = log(xt /xt−1) = log(xt ) − log(xt−1) ≈ rt .1.2
The data set is also available in astsa, but xts must be loaded.
library(TTR)
djia
library(xts)
djiar = diff(log(djia$Close))[-1] # approximate returns
plot(djiar, main="DJIA Returns", type="n")
lines(djiar)
Publicité
= getYahooData("^DJI", start=20060420, end=20160420, freq="daily")
1.2 log(1 + p) = p − p
2 + p
expansion are negligible.
2
3
3 − · · · for −1 < p ≤ 1. If p is near zero, the higher-order terms in the
(cid:105)
(cid:105)
(cid:105)
(cid:105)
Apr 21 2006Apr 01 2008Apr 01 2010Apr 02 2012Apr 01 2014Mar 31 2016−0.050.000.050.10DJIA Returns(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 5 — #15
(cid:105)
(cid:105)
1.1 The Nature of Time Series Data
5
Fig. 1.5. Monthly SOI and Recruitment (estimated new fish), 1950-1987.
Example 1.5 El Niño and Fish Population
We may also be interested in analyzing several time series at once. Figure 1.5
shows monthly values of an environmental series called the Southern Oscillation
Index (SOI) and associated Recruitment (number of new fish) furnished byDr. Roy
Mendelssohn of the Pacific Environmental Fisheries Group (personal communica-
tion). Both series are for a period of 453 months ranging over the years 1950–1987.
The SOI measures changes in air pressure, related to sea surface temperatures in
the central Pacific Ocean. The central Pacific warms every three to seven years due
to the El Niño effect, which has been blamed for various global extreme weather
events. Both series in Figure 1.5 exhibit repetitive behavior, with regularly repeating
cycles that are easily visible. This periodic behavior is of interest because under-
lying processes of interest may be regular and the rate or frequency of oscillation
characterizing the behavior of the underlying series would help to identify them.
The series show two basic oscillations types, an obvious annual cycle (hot in the
summer, cold in the winter), and a slower frequency that seems to repeat about
every 4 years. The study of the kinds of cycles and their strengths is the subject of
Chapter 4. The two series are also related; it is easy to imagine the fish population is
dependent on the ocean temperature. This possibility suggests trying some version
of regression analysis as a procedure for relating the two series. Transfer function
(cid:105)
(cid:105)
(cid:105)
(cid:105)
Southern Oscillation Index1950196019701980−1.0−0.50.00.51.0RecruitmentTime1950196019701980020406080100(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 6 — #16
(cid:105)
(cid:105)
6
1 Characteristics of Time Series
Fig. 1.6. fMRI data from various locations in the cortex, thalamus, and cerebellum; n = 128
points, one observation taken every 2 seconds.
modeling, as considered in Chapter 5, can also be applied in this case. The following
R code will reproduce Figure 1.5:
par(mfrow = c(2,1)) # set up the graphics
plot(soi, ylab="", xlab="", main="Southern Oscillation Index")
plot(rec, ylab="", xlab="", main="Recruitment")
Example 1.6 fMRI Imaging
A fundamental problem in classical statistics occurs when we are given a collection
of independent series or vectors of series, generated under varying experimental
conditions or treatment configurations. Such a set of series is shown in Figure 1.6,
where we observe data collected from various locations in the brain via functional
magnetic resonance imaging (fMRI). In this example, five subjects were given pe-
riodic brushing on the hand. The stimulus was applied for 32 seconds and then
stopped for 32 seconds; thus, the signal period is 64 seconds. The sampling rate
was one observation every 2 seconds for 256 seconds (n = 128). For this example,
we averaged the results over subjects (these were evoked responses, and all subjects
were in phase). The series shown in Figure 1.6 are consecutive measures of blood
oxygenation-level dependent (bold) signal intensity, which measures areas of acti-
vation in the brain. Notice that the periodicities appear strongly in the motor cortex
(cid:105)
(cid:105)
(cid:105)
(cid:105)
CortexBOLD020406080100120−0.6−0.20.20.6CortexBOLD020406080100120−0.6−0.20.20.6Thalamus & CerebellumBOLD020406080100120−0.6−0.20.20.6Thalamus & CerebellumBOLD020406080100120−0.6−0.20.20.6Time (1 pt = 2 sec)(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 7 — #17
(cid:105)
(cid:105)
1.1 The Nature of Time Series Data
7
Fig. 1.7. Arrival phases from an earthquake (top) and explosion (bottom) at 40 points per
second.
series and less strongly in the thalamus and cerebellum. The fact that one has series
from different areas of the brain suggests testing whether the areas are responding
differently to the brush stimulus. Analysis of variance techniques accomplish this in
classical statistics, and we show in Chapter 7 how these classical techniques extend
to the time series case, leading to a spectral analysis of variance. The following R
commands can be used to plot the data:
par(mfrow=c(2,1))
ts.plot(fmri1[,2:5], col=1:4, ylab="BOLD", main="Cortex")
ts.plot(fmri1[,6:9], col=1:4, ylab="BOLD", main="Thalamus & Cerebellum")
Example 1.7 Earthquakes and Explosions
As a final example, the series in Figure 1.7 represent two phases or arrivals along
the surface, denoted by P (t = 1, . . . , 1024) and S (t = 1025, . . . , 2048), at a seismic
recording station. The recording instruments in Scandinavia are observing earth-
quakes and mining explosions with one of each shown in Figure 1.7. The general
problem of interest is in distinguishing or discriminating between waveforms gen-
erated by earthquakes and those generated by explosions. Features that may be
important are the rough amplitude ratios of the first phase P to the second phase
S, which tend to be smaller for earthquakes than for explosions. In the case of the
(cid:105)
(cid:105)
(cid:105)
(cid:105)
EarthquakeEQ50500100015002000−0.40.00.20.4ExplosionEXP60500100015002000−0.40.00.20.4Time(cid:105)
(cid:105)
“tsa4_trimmed” — 2017/12/8 — 15:01 — page 8 — #18
(cid:105)
(cid:105)
8
1 Characteristics of Time Series
two events in Figure 1.7, the ratio of maximum amplitudes appears to be somewhat
less than .5 for the earthquake and about 1 for the explosion. Otherwise, note a
subtle difference exists in the periodic nature of the S phase for the earthquake. We
can again think about spectral analysis of variance for testing the equality of the
periodic components of earthquakes and explosions. We would also like to be able
to classify future P and S components from events of unknown origin, leading to
the time series discriminant analysis developed in Chapter 7.
To plot the data as in this example, use the following commands in R:
par(mfrow=c(2,1))
plot(EQ5, main="Earthquake")
plot(EXP6, mai...