Time Series Analysis and Its Applications with R Examples (Fourth Edition)

Springer
Page 1 sur 568Lecteur de document UniversityLib

Time Series Analysis and Its Applications with R Examples (Fourth Edition)

Statistics and Data Analysis · notes

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 1 — #1

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

Springer Texts in StatisticsRobert H. ShumwayDavid S. Stoff erTime Series Analysis and Its ApplicationsWith R Examples Fourth Edition (cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 2 — #2

(cid:105)

(cid:105)

Robert H. Shumway

David S. Stoffer

Time Series Analysis and

Its Applications

With R Examples

Fourth Edition

live free or bark

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page v — #3

(cid:105)

(cid:105)

Preface to the Fourth Edition

The fourth edition follows the general layout of the third edition but includes some

modernization of topics as well as the coverage of additional topics. The preface to

the third edition—which follows—still applies, so we concentrate on the differences

between the two editions here. As in the third edition, R code for each example

is given in the text, even if the code is excruciatingly long. Most of the examples

with seemingly endless coding are in the latter chapters. The R package for the text,

astsa, is still supported and details may be found in Appendix R. A number of data

sets have been updated. For example, the global temperature deviation series have

been updated to 2015 and are included in the newest version of the package; the

corresponding examples and problems have been updated accordingly.

Chapter 1 of this edition is similar to the previous edition, but we have included

the definition of trend stationarity and the the concept of prewhitening when using

cross-correlation. The New York Stock Exchange data set, which focused on an old

financial crisis, was replaced with a more current series of the Dow Jones Indus-

trial Average, which focuses on a newer financial crisis. In Chapter 2, we rewrote

some of the regression review, changed the smoothing examples from the mortality

data example to the Southern Oscillation Index and finding El Niño. We also ex-

panded the discussion of lagged regression to Chapter 3 to include the possibility of

autocorrelated errors.

In Chapter 3, we removed normality from definition of ARMA models; while the

assumption is not necessary for the definition, it is essential for inference and pre-

diction. We added a section on regression with ARMA errors and the corresponding

problems; this section was previously in Chapter 5. Some of the examples have been

modified and we added some examples in the seasonal ARMA section.

In Chapter 4, we improved and added some examples. The idea of modulated

series is discussed using the classic star magnitude data set. We moved some of the

filtering section forward for easier access to information when needed. We removed

the reliance on spec.pgram (from the stats package) to mvspec (from the astsa

package) so we can avoid having to spend pages explaining the quirks of spec.pgram,

which tended to take over the narrative. The section on wavelets was removed because

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page vi — #4

(cid:105)

(cid:105)

vi

Preface to the Fourth Edition

there are so many accessible texts available. The spectral representation theorems are

discussed in a little more detail using examples based on simple harmonic processes.

The general layout of Chapter 5 and of Chapter 7 is the same, although we have

revised some of the examples. As previously mentioned, we moved regression with

ARMA errors to Chapter 3.

Chapter 6 sees the biggest change in this edition. We have added a section on

smoothing splines, and a section on hidden Markov models and switching autore-

gressions. The Bayesian section is completely rewritten and is on linear Gaussian

state space models only. The nonlinear material in the previous edition is removed

because it was old, and the newer material is in Douc, Moulines, and Stoffer (2014).

Many of the examples have been rewritten to make the chapter more accessible. Our

goal was to be able to have a course on state space models based primarily on the

material in Chapter 6.

The Appendices are similar, with some minor changes to Appendix A and Ap-

pendix B. We added material to Appendix C, including a discussion of Riemann–

Stieltjes and stochastic integration, a proof of the fact that the spectra of autoregressive

processes are dense in the space of spectral densities, and a proof of the fact that spec-

tra are approximately the eigenvalues of the covariance matrix of a stationary process.

We tweaked, rewrote, improved, or revised some of the exercises, but the overall

ordering and coverage is roughly the same. And, of course, we moved regression with

ARMA errors problems to Chapter 3 and removed the Chapter 4 wavelet problems.

The exercises for Chapter 6 have been updated accordingly to reflect the new and

improved version of the chapter.

Davis, CA

Pittsburgh, PA

September 2016

Robert H. Shumway

David S. Stoffer

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page vii — #5

(cid:105)

(cid:105)

Preface to the Third Edition

The goals of this book are to develop an appreciation for the richness and versatility

of modern time series analysis as a tool for analyzing data, and still maintain a

commitment to theoretical integrity, as exemplified by the seminal works of Brillinger

(1975) and Hannan (1970) and the texts by Brockwell and Davis (1991) and Fuller

(1995). The advent of inexpensive powerful computing has provided both real data

and new software that can take one considerably beyond the fitting of simple time

domain models, such as have been elegantly described in the landmark work of Box

and Jenkins (1970). This book is designed to be useful as a text for courses in time

series on several different levels and as a reference work for practitioners facing the

analysis of time-correlated data in the physical, biological, and social sciences.

We have used earlier versions of the text at both the undergraduate and gradu-

ate levels over the past decade. Our experience is that an undergraduate course can

be accessible to students with a background in regression analysis and may include

Section 1.1–Section 1.5, Section 2.1–Section 2.3, the results and numerical parts of

Section 3.1–Section 3.9, and briefly the results and numerical parts of Section 4.1–

Publicité

Section 4.4. At the advanced undergraduate or master’s level, where the students

have some mathematical statistics background, more detailed coverage of the same

sections, with the inclusion of extra topics from Chapter 5 or Chapter 6 can be used

as a one-semester course. Often, the extra topics are chosen by the students according

to their interests. Finally, a two-semester upper-level graduate course for mathemat-

ics, statistics, and engineering graduate students can be crafted by adding selected

theoretical appendices. For the upper-level graduate course, we should mention that

we are striving for a broader but less rigorous level of coverage than that which is

attained by Brockwell and Davis (1991), the classic entry at this level.

The major difference between this third edition of the text and the second edition

is that we provide R code for almost all of the numerical examples. An R package

called astsa is provided for use with the text; see Section R.2 for details. R code

is provided simply to enhance the exposition by making the numerical examples

reproducible.

We have tried, where possible, to keep the problem sets in order so that an

instructor may have an easy time moving from the second edition to the third edition.

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page viii — #6

(cid:105)

(cid:105)

viii

Preface to the Third Edition

However, some of the old problems have been revised and there are some new

problems. Also, some of the data sets have been updated. We added one section

in Chapter 5 on unit roots and enhanced some of the presentations throughout the

text. The exposition on state-space modeling, ARMAX models, and (multivariate)

regression with autocorrelated errors in Chapter 6 have been expanded. In this edition,

we use standard R functions as much as possible, but we use our own scripts (included

in astsa) when we feel it is necessary to avoid problems with a particular R function;

these problems are discussed in detail on the website for the text under R Issues.

We thank John Kimmel, Executive Editor, Springer Statistics, for his guidance

in the preparation and production of this edition of the text. We are grateful to Don

Percival, University of Washington, for numerous suggestions that led to substantial

improvement to the presentation in the second edition, and consequently in this

edition. We thank Doug Wiens, University of Alberta, for help with some of the R

code in Chapter 4 and Chapter 7, and for his many suggestions for improvement of

the exposition. We are grateful for the continued help and advice of Pierre Duchesne,

University of Montreal, and Alexander Aue, University of California, Davis. We also

thank the many students and other readers who took the time to mention typographical

errors and other corrections to the first and second editions. Finally, work on the this

edition was supported by the National Science Foundation while one of us (D.S.S.)

was working at the Foundation under the Intergovernmental Personnel Act.

Davis, CA

Pittsburgh, PA

September 2010

Robert H. Shumway

David S. Stoffer

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page ix — #7

(cid:105)

(cid:105)

Contents

Preface to the Fourth Edition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

v

Preface to the Third Edition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

vii

1

2

3

Characteristics of Time Series . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1.1 The Nature of Time Series Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1.2 Time Series Statistical Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1.3 Measures of Dependence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1.4 Stationary Time Series . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1.5 Estimation of Correlation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1.6 Vector-Valued and Multidimensional Series . . . . . . . . . . . . . . . . . . . . .

Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

Time Series Regression and Exploratory Data Analysis . . . . . . . . . . . . .

2.1 Classical Regression in the Time Series Context . . . . . . . . . . . . . . . . .

2.2 Exploratory Data Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

2.3 Smoothing in the Time Series Context . . . . . . . . . . . . . . . . . . . . . . . . .

Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

1

2

8

14

19

26

33

38

47

47

56

67

72

77

ARIMA Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

77

3.1 Autoregressive Moving Average Models . . . . . . . . . . . . . . . . . . . . . . .

90

3.2 Difference Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

3.3 Autocorrelation and Partial Autocorrelation . . . . . . . . . . . . . . . . . . . . .

96

3.4 Forecasting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102

3.5 Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115

3.6 Integrated Models for Nonstationary Data . . . . . . . . . . . . . . . . . . . . . . 133

3.7 Building ARIMA Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137

3.8 Regression with Autocorrelated Errors . . . . . . . . . . . . . . . . . . . . . . . . 145

3.9 Multiplicative Seasonal ARIMA Models . . . . . . . . . . . . . . . . . . . . . . . 148

Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page x — #8

(cid:105)

(cid:105)

x

4

5

6

7

Contents

Publicité

Spectral Analysis and Filtering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167

4.1 Cyclical Behavior and Periodicity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168

4.2 The Spectral Density . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174

4.3 Periodogram and Discrete Fourier Transform . . . . . . . . . . . . . . . . . . . 181

4.4 Nonparametric Spectral Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191

4.5 Parametric Spectral Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205

4.6 Multiple Series and Cross-Spectra . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208

4.7 Linear Filters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213

4.8 Lagged Regression Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 218

4.9 Signal Extraction and Optimum Filtering . . . . . . . . . . . . . . . . . . . . . . . 223

4.10 Spectral Analysis of Multidimensional Series . . . . . . . . . . . . . . . . . . . 227

Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 230

Additional Time Domain Topics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 241

5.1 Long Memory ARMA and Fractional Differencing . . . . . . . . . . . . . . 241

5.2 Unit Root Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250

5.3 GARCH Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 253

5.4 Threshold Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 261

5.5 Lagged Regression and Transfer Function Modeling . . . . . . . . . . . . . 265

5.6 Multivariate ARMAX Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 271

Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 284

State Space Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 287

6.1 Linear Gaussian Model

. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 288

6.2 Filtering, Smoothing, and Forecasting . . . . . . . . . . . . . . . . . . . . . . . . . 292

6.3 Maximum Likelihood Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 302

6.4 Missing Data Modifications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 310

6.5 Structural Models: Signal Extraction and Forecasting . . . . . . . . . . . . 315

6.6 State-Space Models with Correlated Errors . . . . . . . . . . . . . . . . . . . . . 319

6.6.1 ARMAX Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 320

6.6.2 Multivariate Regression with Autocorrelated Errors . . . . . . . 322

6.7 Bootstrapping State Space Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 325

6.8 Smoothing Splines and the Kalman Smoother . . . . . . . . . . . . . . . . . . . 331

6.9 Hidden Markov Models and Switching Autoregression . . . . . . . . . . . 334

6.10 Dynamic Linear Models with Switching . . . . . . . . . . . . . . . . . . . . . . . 345

6.11 Stochastic Volatility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 357

6.12 Bayesian Analysis of State Space Models . . . . . . . . . . . . . . . . . . . . . . 365

Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 375

Statistical Methods in the Frequency Domain . . . . . . . . . . . . . . . . . . . . . 383

7.1

Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 383

7.2 Spectral Matrices and Likelihood Functions . . . . . . . . . . . . . . . . . . . . 386

. . . . . . . . . . . . . . . . . . . . . . . 388

7.3 Regression for Jointly Stationary Series

7.4 Regression with Deterministic Inputs

. . . . . . . . . . . . . . . . . . . . . . . . . 397

7.5 Random Coefficient Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 405

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page xi — #9

(cid:105)

(cid:105)

Contents

xi

7.6 Analysis of Designed Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 407

7.7 Discriminant and Cluster Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . 421

7.8 Principal Components and Factor Analysis . . . . . . . . . . . . . . . . . . . . . 437

7.9 The Spectral Envelope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 453

Problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 464

Appendix A Large Sample Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 471

A.1 Convergence Modes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 471

A.2 Central Limit Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 478

A.3 The Mean and Autocorrelation Functions . . . . . . . . . . . . . . . . . . . . . . 482

Appendix B Time Domain Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491

B.1 Hilbert Spaces and the Projection Theorem . . . . . . . . . . . . . . . . . . . . . 491

B.2 Causal Conditions for ARMA Models . . . . . . . . . . . . . . . . . . . . . . . . . 495

B.3 Large Sample Distribution of the AR Conditional Least Squares

Estimators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 497

B.4 The Wold Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 500

Appendix C Spectral Domain Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503

C.1 Spectral Representation Theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . 503

C.2 Large Sample Distribution of the Smoothed Periodogram . . . . . . . . . 507

C.3 The Complex Multivariate Normal Distribution . . . . . . . . . . . . . . . . . 517

C.4 Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 522

C.4.1 Riemann–Stieltjes Integration . . . . . . . . . . . . . . . . . . . . . . . . . . 522

C.4.2 Stochastic Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 524

C.5 Spectral Analysis as Principal Component Analysis . . . . . . . . . . . . . . 525

C.6 Parametric Spectral Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 529

Appendix R R Supplement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531

R.1 First Things First . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531

R.2 astsa . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 531

R.3 Getting Started . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 532

R.4 Time Series Primer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 536

R.4.1 Graphics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 539

References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 541

Index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 553

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page xii — #10

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 1 — #11

(cid:105)

(cid:105)

Chapter 1

Characteristics of Time Series

The analysis of experimental data that have been observed at different points in time

leads to new and unique problems in statistical modeling and inference. The obvi-

ous correlation introduced by the sampling of adjacent points in time can severely

restrict the applicability of the many conventional statistical methods traditionally

dependent on the assumption that these adjacent observations are independent and

identically distributed. The systematic approach by which one goes about answer-

ing the mathematical and statistical questions posed by these time correlations is

commonly referred to as time series analysis.

The impact of time series analysis on scientific applications can be partially doc-

umented by producing an abbreviated listing of the diverse fields in which important

time series problems may arise. For example, many familiar time series occur in the

field of economics, where we are continually exposed to daily stock market quota-

tions or monthly unemployment figures. Social scientists follow population series,

such as birthrates or school enrollments. An epidemiologist might be interested in

the number of influenza cases observed over some time period. In medicine, blood

pressure measurements traced over time could be useful for evaluating drugs used

Publicité

in treating hypertension. Functional magnetic resonance imaging of brain-wave time

series patterns might be used to study how the brain reacts to certain stimuli under

various experimental conditions.

In our view, the first step in any time series investigation always involves careful

examination of the recorded data plotted over time. This scrutiny often suggests

the method of analysis as well as statistics that will be of use in summarizing the

information in the data. Before looking more closely at the particular statistical

methods, it is appropriate to mention that two separate, but not necessarily mutually

exclusive, approaches to time series analysis exist, commonly identified as the time

domain approach and the frequency domain approach. The time domain approach

views the investigation of lagged relationships as most important (e.g., how does

what happened today affect what will happen tomorrow), whereas the frequency

domain approach views the investigation of cycles as most important (e.g., what is

the economic cycle through periods of expansion and recession). We will explore

both types of approaches in the following sections.

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 2 — #12

(cid:105)

(cid:105)

2

1 Characteristics of Time Series

Fig. 1.1. Johnson & Johnson quarterly earnings per share, 84 quarters, 1960-I to 1980-IV.

1.1 The Nature of Time Series Data

Some of the problems and questions of interest to the prospective time series analyst

can best be exposed by considering real experimental data taken from different subject

areas. The following cases illustrate some of the common kinds of experimental time

series data as well as some of the statistical questions that might be asked about such

data.

Example 1.1 Johnson & Johnson Quarterly Earnings

Figure 1.1 shows quarterly earnings per share for the U.S. company Johnson &

Johnson, furnished by Professor Paul Griffin (personal communication) of the

Graduate School of Management, University of California, Davis. There are 84

quarters (21 years) measured from the first quarter of 1960 to the last quarter of

1980. Modeling such series begins by observing the primary patterns in the time

history. In this case, note the gradually increasing underlying trend and the rather

regular variation superimposed on the trend that seems to repeat over quarters.

Methods for analyzing data such as these are explored in Chapter 2 and Chapter 6.

To plot the data using the R statistical package, type the following:1.1

library(astsa)

SEE THE FOOTNOTE

plot(jj, type="o", ylab="Quarterly Earnings per Share")

Example 1.2 Global Warming

Consider the global temperature series record shown in Figure 1.2. The data are the

global mean land–ocean temperature index from 1880 to 2015, with the base period

1951-1980. In particular, the data are deviations, measured in degrees centigrade,

from the 1951-1980 average, and are an update of Hansen et al. (2006). We note an

apparent upward trend in the series during the latter part of the twentieth century

that has been used as an argument for the global warming hypothesis. Note also

the leveling off at about 1935 and then another rather sharp upward trend at about

1.1 Throughout the text, we assume that the R package for the book, astsa, has been installed and loaded.

See Section R.2 for further details.

(cid:105)

(cid:105)

(cid:105)

(cid:105)

TimeQuarterly Earnings per Share19601965197019751980051015llllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllll(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 3 — #13

(cid:105)

(cid:105)

1.1 The Nature of Time Series Data

3

Fig. 1.2. Yearly average global temperature deviations (1880–2015) in degrees centigrade.

Fig. 1.3. Speech recording of the syllable aaa · · · hhh sampled at 10,000 points per second

with n = 1020 points.

1970. The question of interest for global warming proponents and opponents is

whether the overall trend is natural or whether it is caused by some human-induced

interface. Problem 2.8 examines 634 years of glacial sediment data that might be

taken as a long-term temperature proxy. Such percentage changes in temperature

do not seem to be unusual over a time period of 100 years. Again, the question of

trend is of more interest than particular periodicities. The R code for this example

is similar to the code in Example 1.1:

plot(globtemp, type="o", ylab="Global Temperature Deviations")

Example 1.3 Speech Data

Figure 1.3 shows a small .1 second (1000 point) sample of recorded speech for

the phrase aaa · · · hhh, and we note the repetitive nature of the signal and the

rather regular periodicities. One current problem of great interest is computer

recognition of speech, which would require converting this particular signal into

the recorded phrase aaa · · · hhh. Spectral analysis can be used in this context to

produce a signature of this phrase that can be compared with signatures of various

(cid:105)

(cid:105)

(cid:105)

(cid:105)

TimeGlobal Temperature Deviations18801900192019401960198020002020−0.40.00.40.8llllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllllTimespeech0200400600800100001000200030004000(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 4 — #14

(cid:105)

(cid:105)

4

1 Characteristics of Time Series

Fig. 1.4. The daily returns of the Dow Jones Industrial Average (DJIA) from April 20, 2006 to

April 20, 2016.

library syllables to look for a match. One can immediately notice the rather regular

repetition of small wavelets. The separation between the packets is known as the

pitch period and represents the response of the vocal tract filter to a periodic

sequence of pulses stimulated by the opening and closing of the glottis. In R, you

can reproduce Figure 1.3 using plot(speech).

Example 1.4 Dow Jones Industrial Average

As an example of financial time series data, Figure 1.4 shows the daily returns

(or percent change) of the Dow Jones Industrial Average (DJIA) from April 20,

2006 to April 20, 2016. It is easy to spot the financial crisis of 2008 in the figure.

The data shown in Figure 1.4 are typical of return data. The mean of the series

appears to be stable with an average return of approximately zero, however, highly

volatile (variable) periods tend to be clustered together. A problem in the analysis

of these type of financial data is to forecast the volatility of future returns. Models

such as ARCH and GARCH models (Engle, 1982; Bollerslev, 1986) and stochastic

volatility models (Harvey, Ruiz and Shephard, 1994) have been developed to handle

these problems. We will discuss these models and the analysis of financial data in

Chapter 5 and Chapter 6. The data were obtained using the Technical Trading Rules

(TTR) package to download the data from YahooTM and then plot it. We then used

the fact that if xt is the actual value of the DJIA and rt = (xt −xt−1)/xt−1 is the return,

then 1 + rt = xt /xt−1 and log(1 + rt ) = log(xt /xt−1) = log(xt ) − log(xt−1) ≈ rt .1.2

The data set is also available in astsa, but xts must be loaded.

library(TTR)

djia

library(xts)

djiar = diff(log(djia$Close))[-1] # approximate returns

plot(djiar, main="DJIA Returns", type="n")

lines(djiar)

Publicité

= getYahooData("^DJI", start=20060420, end=20160420, freq="daily")

1.2 log(1 + p) = p − p

2 + p

expansion are negligible.

2

3

3 − · · · for −1 < p ≤ 1. If p is near zero, the higher-order terms in the

(cid:105)

(cid:105)

(cid:105)

(cid:105)

Apr 21 2006Apr 01 2008Apr 01 2010Apr 02 2012Apr 01 2014Mar 31 2016−0.050.000.050.10DJIA Returns(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 5 — #15

(cid:105)

(cid:105)

1.1 The Nature of Time Series Data

5

Fig. 1.5. Monthly SOI and Recruitment (estimated new fish), 1950-1987.

Example 1.5 El Niño and Fish Population

We may also be interested in analyzing several time series at once. Figure 1.5

shows monthly values of an environmental series called the Southern Oscillation

Index (SOI) and associated Recruitment (number of new fish) furnished byDr. Roy

Mendelssohn of the Pacific Environmental Fisheries Group (personal communica-

tion). Both series are for a period of 453 months ranging over the years 1950–1987.

The SOI measures changes in air pressure, related to sea surface temperatures in

the central Pacific Ocean. The central Pacific warms every three to seven years due

to the El Niño effect, which has been blamed for various global extreme weather

events. Both series in Figure 1.5 exhibit repetitive behavior, with regularly repeating

cycles that are easily visible. This periodic behavior is of interest because under-

lying processes of interest may be regular and the rate or frequency of oscillation

characterizing the behavior of the underlying series would help to identify them.

The series show two basic oscillations types, an obvious annual cycle (hot in the

summer, cold in the winter), and a slower frequency that seems to repeat about

every 4 years. The study of the kinds of cycles and their strengths is the subject of

Chapter 4. The two series are also related; it is easy to imagine the fish population is

dependent on the ocean temperature. This possibility suggests trying some version

of regression analysis as a procedure for relating the two series. Transfer function

(cid:105)

(cid:105)

(cid:105)

(cid:105)

Southern Oscillation Index1950196019701980−1.0−0.50.00.51.0RecruitmentTime1950196019701980020406080100(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 6 — #16

(cid:105)

(cid:105)

6

1 Characteristics of Time Series

Fig. 1.6. fMRI data from various locations in the cortex, thalamus, and cerebellum; n = 128

points, one observation taken every 2 seconds.

modeling, as considered in Chapter 5, can also be applied in this case. The following

R code will reproduce Figure 1.5:

par(mfrow = c(2,1)) # set up the graphics

plot(soi, ylab="", xlab="", main="Southern Oscillation Index")

plot(rec, ylab="", xlab="", main="Recruitment")

Example 1.6 fMRI Imaging

A fundamental problem in classical statistics occurs when we are given a collection

of independent series or vectors of series, generated under varying experimental

conditions or treatment configurations. Such a set of series is shown in Figure 1.6,

where we observe data collected from various locations in the brain via functional

magnetic resonance imaging (fMRI). In this example, five subjects were given pe-

riodic brushing on the hand. The stimulus was applied for 32 seconds and then

stopped for 32 seconds; thus, the signal period is 64 seconds. The sampling rate

was one observation every 2 seconds for 256 seconds (n = 128). For this example,

we averaged the results over subjects (these were evoked responses, and all subjects

were in phase). The series shown in Figure 1.6 are consecutive measures of blood

oxygenation-level dependent (bold) signal intensity, which measures areas of acti-

vation in the brain. Notice that the periodicities appear strongly in the motor cortex

(cid:105)

(cid:105)

(cid:105)

(cid:105)

CortexBOLD020406080100120−0.6−0.20.20.6CortexBOLD020406080100120−0.6−0.20.20.6Thalamus & CerebellumBOLD020406080100120−0.6−0.20.20.6Thalamus & CerebellumBOLD020406080100120−0.6−0.20.20.6Time (1 pt = 2 sec)(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 7 — #17

(cid:105)

(cid:105)

1.1 The Nature of Time Series Data

7

Fig. 1.7. Arrival phases from an earthquake (top) and explosion (bottom) at 40 points per

second.

series and less strongly in the thalamus and cerebellum. The fact that one has series

from different areas of the brain suggests testing whether the areas are responding

differently to the brush stimulus. Analysis of variance techniques accomplish this in

classical statistics, and we show in Chapter 7 how these classical techniques extend

to the time series case, leading to a spectral analysis of variance. The following R

commands can be used to plot the data:

par(mfrow=c(2,1))

ts.plot(fmri1[,2:5], col=1:4, ylab="BOLD", main="Cortex")

ts.plot(fmri1[,6:9], col=1:4, ylab="BOLD", main="Thalamus & Cerebellum")

Example 1.7 Earthquakes and Explosions

As a final example, the series in Figure 1.7 represent two phases or arrivals along

the surface, denoted by P (t = 1, . . . , 1024) and S (t = 1025, . . . , 2048), at a seismic

recording station. The recording instruments in Scandinavia are observing earth-

quakes and mining explosions with one of each shown in Figure 1.7. The general

problem of interest is in distinguishing or discriminating between waveforms gen-

erated by earthquakes and those generated by explosions. Features that may be

important are the rough amplitude ratios of the first phase P to the second phase

S, which tend to be smaller for earthquakes than for explosions. In the case of the

(cid:105)

(cid:105)

(cid:105)

(cid:105)

EarthquakeEQ50500100015002000−0.40.00.20.4ExplosionEXP60500100015002000−0.40.00.20.4Time(cid:105)

(cid:105)

“tsa4_trimmed” — 2017/12/8 — 15:01 — page 8 — #18

(cid:105)

(cid:105)

8

1 Characteristics of Time Series

two events in Figure 1.7, the ratio of maximum amplitudes appears to be somewhat

less than .5 for the earthquake and about 1 for the explosion. Otherwise, note a

subtle difference exists in the periodic nature of the S phase for the earthquake. We

can again think about spectral analysis of variance for testing the equality of the

periodic components of earthquakes and explosions. We would also like to be able

to classify future P and S components from events of unknown origin, leading to

the time series discriminant analysis developed in Chapter 7.

To plot the data as in this example, use the following commands in R:

par(mfrow=c(2,1))

plot(EQ5, main="Earthquake")

plot(EXP6, mai...