Selection of Best ARIMA Model for Forecasting Average Daily Share Price Index of Pharmaceutical Companies in Bangladesh: A Case Study on Square Pharmaceutical Ltd

Table of contents

1. Introduction

tock exchange plays a vital role in the national economy of Bangladesh. Stock market is an essential part of the capital market. The economy of a country largely depends on capital market. In the capital market the investors invest the money to get the profit. The investors buy the security bond of different company on the priority basis. They choose the security bond of different company on the basis of the different factors. Some of the significant factors are Company's information analysis & prediction, dividend declaration, etc. A large amount of investors has no knowledge about the market analysis and proper prediction of the future prices of different types of shares available in the market. So, most of the time they spend the money to buy security bond of different companies on the basis of wrong and thumb idea, without any idea about data analysis and prediction.

For this reason there are extreme ups and downs in the daily share price indices, sometimes rise very quickly and fall sharply. In this situation, the market condition becomes unpredictable. Hence, a large amount of investors loss their capital in this unstable capital market. As a result the general investors do not find interest to invest the money in the capital market. Then there arises a crisis in the capital market which creates problem and hampers the national economic growth.

Therefore, if it is possible to provide a better model for the share market which can enable the investors to predict the prices in advance, it would help the investors as well as keep stability of the national economy. This study is an effort towards that direction.

2. II.

Literature Review Contreras et al. (2003) used ARIMA models to predict next day electricity prices; they have found two ARIMA models to predict hourly prices in the electricity markets of Spain & California. The Spanish model needs 5 hours to predict future prices as opposed to the 2 hours needed by the Californian model. Kumar et al. (2004) used ARIMA model to forecast daily maximum surface ozone concentrations in Brunei Darussalam. They have found that ARIMA (1, 0, 1) was suitable for the surface O 3 data collected at the airport in Brunei Darussalam. Tsitsika et al. (2007) used ARIMA model to forecast pelagic fish production. The final model selected were of the form ARIMA (1, 0, 1) & ARIMA (0, 1, 1). From the above mentioned studies it is clear that ARIMA can be used to forecast. In very few of them the authors tried to find out best ARIMA model, but in most of the articles the authors used ARIMA to forecast. The present study is designed to select the best ARIMA model to forecast average daily price index of listed companies in Dhaka Stock Exchange.

3. III.

4. Objectives of the Study

Share price index is a time series data. One of the important objectives of the time series analysis is to study the past behavior of the available data and then forecast with fitting a suitable model with the help of econometric or statistical techniques. Thus, the specific objectives of this study are as follows: 1. To check whether the selected time series data is stationary or not. If not, the data are to be transformed into stationary using suitable transformation. 2. To select the best ARIMA model using some selection criteria. Then ARIMA techniques are applied to fit and forecast the average daily share price indices of DSE data for the Square Pharmaceuticals Limited (SPL) Company. 3. Finally, to draw a conclusion for forecasting the average daily share price indices of the selected company efficiently.

IV.

5. Data and Methodology

The ADSPI data recorded against SPL have been collected from Dhaka Stock Exchange (DSE) for the year 2011. Thus we obtained a total of 236 observations against all working days from Square Pharmaceuticals limited. The stepwise methodology used in this study is outlined below:

Firstly, the data is presented graphically to check whether the data series is stationary or not. For this purpose, the statistics like Ljung-Box-Pierce Qstatistic (1978) based on auto correlation; Dickey-Fuller test (DF) (1979), Augmented Dickey-Fuller (ADF) test (1982) based on unit root process have been applied.

To select the best ARIMA (p, d, q) type of models fitted for the company, their goodness of fit have been compared using following criteria; a) The Akaike Information Criteria (AIC) b) The Corrected Akaike Information Criteria (AICc) c) Schwartz Information Criteria (SIC) d) Mean Absolute Percent Error (MAPE) e) Root Mean Square Error (RMSE) and f) Absolute Mean Error (AME)

A brief description about the criteria for the selection of best ARIMA model is given below: a) Akaike Information Criterion (AIC) AIC is an important and leading statistics by which we can determine the order of an autoregressive model Mr. Akaike developed this statistics. According to his name this statistics is known as Akaike Information Criterion (AIC). The AIC takes into account both how well the model fits the observed series and the number of parameters to be used in the fit. AIC due to Akaike (1969) is defined as ( )

1 2 1 2 + + ? ? ? ? ? ? + ? = ? p In N AIC

Where the parameter bears the usual meaning. Akaike also mention that the minimum AIC criterion produced a selected model, which is hopefully closer to the best possible choice.

6. b) Corrected Akaike Information Criterion

Sometimes the AIC does not provide the efficient order of model selection, which asymptotic efficiency is more desirable criterion. Shibata in 1976 shown that AIC criterion is not consistent too. Thus Hurvich and Tsai (1989) provide a criterion of AIC for bias. The correlation is of particular use when the sample size is small or when the number of fitted parameter is a moderate to a large fraction of sample size. The criterion is defined as

N P N P N AIC c 2 1 1 ln 2 + ? + + ? = ? i.e, ( )( ) ( )2

7. C

Thus AIC c is the sum of AIC and an additional non-stochastic penalty term 2(p+1) (p+2) / (N-p+2), where the parameter bears the usual meaning.

8. c) Schwaetz Information Criteria

In 1978 Schwaetz discussed a criterion denoted by SIC which help in deciding the order of auto regression. Initially he developed this criterion for taking decisions about the regress subset. Later Engel et. al, in 1992 use this criterion as a tool for determining the order of auto regression and they defined this criterion as below

N p N N p SIC 2 2 1 ? ? ? ? ? ? ? ? = ?

Where, the parameters bear the usual meaning. Schwartz also shows that this criterion is better than AIC. The model with minimum SIC assumes to describe the data series adequately. The minimum value of this criterion is desirable for the adequacy of a model. The criteria mentioned above are compared for correct determination of the order of auto regression and the degree of differencing and this criterion is computed only for estimation period. But for the selection of an ARIMA model, which adequately describes the data series, the values of the following criteria are compared for three periods viz, estimation period, validation period and total period. The criteria used in this study are as follows:

The mean of the absolute deviation of predicted and observed values is called absolute mean error and is defined as

? ? ? ? = ? = pred obs T I AME 1

This criterion is used for the comparison of the models in three periods.

The square root of the sum of square of the deviation of the predicted values from the observed value dividing by their number of observation is known as the root mean square error. The root mean square error is defined as ( )

2 1 1 ? = ? ? ? ? = T I pred obs

9. RMSE

Where, T is the number of periods. This criterion is used for the comparison of the models in three periods.

The mean of the sum of absolute deviation of predicted and observed value dividing by the observed value is called mean absolute error. For comparison we have multiplied by 100, which is called mean absolute percent error and which is defined as

100 1 1 × ? ? ? ? ? = ? = T t obs pred obs

10. MAPE

Where, the parameters bear the usual meaning.

From the above discussion it is clear that the smaller error better the forecasting performance of the observed variables and if the model variable perform well, so will the model as a whole do too.

For the data series a separate ARIMA model has been used. For that purpose, a general concept of ARIMA (p, d, and q) model is discussed below:

ARIMA models are, in theory, the most general class of models for forecasting a time series that can be stationeries by transformations such as differencing and logging. If we have to difference a time series d times to make it stationary and then apply the ARMA (p, q) model to it, we can say that the original time series is ARIMA (p, d, q), that is it is an autoregressive integrated moving average time series, where p denotes the number of autoregressive terms, d denotes the time series have to be differenced before it becomes stationary and q denotes the number of moving average terms. Thus an ARIMA (2,1,2) time series has to be differenced once (d=1) before it becomes stationary and the stationary time series can be modeled as an ARMA (2,2) process that is it has two AR and two MA terms. Of course if d=0 then ARIMA (p, d=0,q) = ARMA (p, q). A most general ARIMA model constitutes three types of process named as autoregressive (AR) process, differencing to strip of the integration (I) and moving average (MA) process. The goodness of fit with respect to every criterion are examined and the model which satisfies most of the criterion, is considered as the best one.

In an autoregressive process each value in a series is linear function of the preceding value. Thus in the first order autoregressive process only the single preceding value is used as a function of current value. In the second order autoregressive process two preceding values are used as a function of the current value and so on. The first order autoregressive is denoted by AR (1), the second order autoregressive is denoted by AR (2) and up to the p th order autoregressive is denoted by AR (p). The model ( 1) is known as AR (1) model. But if we consider the model

t t t t u + ? ? + ? ? + ? = ? ? ? 2 2 1 1 (2) Where ( ) 2 , 0 ~u t IN u ?

The model ( 2) is known as AR (2) model.

In general we can write Differencing is a comparatively simple operation that involves calculating consecutive changes in the values of the data series. Differencing is used when the mean of a series is changing over time to time. A consciousness that is homogeneously non-stationary can be transform into stationary by differencing. Differencing is not dealing with non-stationary variance. To difference a series once (d=1) we have to calculate the period to period change, to difference a series twice (d=2) we have to calculate the period to period changes in the first difference series and so on for further differences.

t p t p t t t u + ? + + ? + ? + = ? ? ? ? ? ? ? ? ......

In Statistics, a moving average or rolling average is one of a family of similar techniques used to analyze time series data. It is applied in finance and especially in technical analysis. It can also be used as a generic smoothing operation, in which case the raw data need not be a time series.

A moving average series can be calculated for any time series. In finance it is most often applied to stock prices, returns or trading volumes. Moving averages are used to smooth out short-term fluctuations, thus highlighting longer-term trends or cycles. The threshold between short-term and long-term depends on the application, and the parameters of the moving average will be set accordingly.

Mathematically, each of these moving averages is an example of a convolution. These averages are also similar to the low-pass filters used in signal processing.

1 1 ? ? + + ? = ? t t t u u (4)

Where ? is constant and u is the white noise error term i.e., u~N ( ) 2 , 0 ? . Here Y at time t is equal to a constant plus a moving average of the current and past error terms. In this case, we say that Y follows a first order moving average or MA (1) process. But if Y follows the expression

2 2 1 1 ? ? ? + ? + + ? = ? t t t t u u u (5)

Then we say that Y follows a second order moving average or MA (2) process. In general,

q t q t t t t u u u u ? ? ? ? + + ? + ? + + ? = ? ...... .......... 2 2 1 1 (6)

Then we say that Y follows a q th order moving average or MA (q) process.

In short, a moving average process is simply a linear combination of white noise error terms. In order to identify the tentative ARIMA model for the ADSPI of SPL, the steps described by Box and Jenkins have been followed. For this purpose the data

11. Characteristics of a good ARIMA model

12. Selection of ARIMA models for ADSPI of SPL data series

In moving average process, each value is determined by the average of the current disturbance and one or more previous disturbances. Suppose the model Y as follows:

are partitioned into two stages. The first stage is known as the estimation stage and second is known as the validation stage. The sample of observations 1 to 226 has been used in estimation stage and the rest has been used for testing the validity of model.

Ten ARIMA models with tentatively selected various values of p, d and q are estimated by using computer software SHAZAM versions 8.0 for windows. The ten tentatively selected models are ARIMA (1,1,1), ARIMA (1,1,2), ARIMA (2,1,1), ARIMA (2,1,2), ARIMA (1,1,3), ARIMA (2,1,3), ARIMA (3,1,1), ARIMA (3,1,2), ARIMA (3,1,3) and ARIMA (1,1,4). Among the models only five comparatively well performed models are displayed in the table -1c. Table-1c discloses that ARIMA with p=2, d =1 and q=2 process has maximum number of lowest values of all the selected criteria AIC, AICc, SIC, and AME, RMSE, MAPE in the three periods i.e., estimation period, validation period and total period Hence, ARIMA (2,1,2) model has been selected for forecasting the ADSPI of SPL data series.

The fitted ARIMA (2, 1, and 2) model selected for SPL data series is given by (1-0.6636*B-0.60032*B 2 ) SPL -t = 0.0040125+ (1-0.44229*B-0.489*B 2 ) a t (0.2925) (0.1737) (0.3036) (0.1971) (Values in the parenthesis are corresponding t-values and '*' means statistical significance p<0.01) v.

13. Results and Discussion

The major findings of the study are as follows: 1. The upward trends of plots of the data series are visualized although the overall trends are not smooth.

14. VI. Conclusion

This study made the best endeavor to develop the best ARIMA model to efficiently forecasting the Average Daily Share Price Indices (ADSPI) of the Square Pharmaceuticals Limited (SPL), because if it is possible to provide a better model for the share market which can enable the investors to predict the prices in advance, it would help the investors as well as stability of the national economy. The empirical analysis indicated that the ARIMA (2,1,2) model is best for forecasting the Average Daily Share Price Indices (ADSPI) of the Square Pharmaceuticals Limited (SPL) data series so far the diagnostic criteria are concerned. Finally, the Average Daily Share Price Indies (ADSPI) for Square Pharmaceuticals Limited (SPL) data series is forecasted up to February, 2012 by using the selected model.

Figure 1. C?
Criteria used for testing the validity of model a) Absolute Mean Error (AME) b) Root Mean Square Error (RMSE) c) Mean Absolute Percent Error (MAPE) d) Absolute Mean Error (AME) e) Root Mean Square Error (RMSE) Mean Absolute Percent Error (MAPE) f) Auto Regressive (AR) ProcessLet us suppose that the variable t
Figure 2.
Datta (2011) used ARIMA model in forecasting
inflation in the Bangladesh Economy. He showed that
ARIMA (1, 0, 1) model fits the inflation data of
Bangladesh satisfactorily. Al-Zeaud (2011) used ARIMA
model in modeling &forecasting volatility. The result
shows that best ARIMA models at 95% confidence
interval for banks sector is ARIMA (2, 0, and 2) model.
Uko et al. (2012) examined the relative predictive power
of ARIMA, VAR & ECM models in forecasting inflation in
Nigeria. The result shows that ARIMA is a good predictor
of inflation in Nigeria & serves as a benchmark model in
inflation forecasting.
Figure 3. Table 1 (
1
4. The Dickey-Fuller unit root test statistic and the
Ljung-Box-pierce Q-Statistic also indicate that the
Average Daily Share Price Indices (ADSPI) of SPL
data series is non-stationary. The computed
absolute values of the ?-statistic for SPL is found as
Figure 4. Table 1 (
1
Transformation Difference=1
Validation of diagnostic criteria for the model
Criteria Period ARIMA ARIMA ARIMA ARIMA ARIMA
(1,1,1) (1,1,2) (1,1,3) (2,1,1) (2,1,2)
AIC Estimation -6.4781 -6.5012* -6.4729 -6.4797 -6.4679
AICc Estimation -6.4275 -6.4506* -6.4223 -6.3780 -6.3662
SIC Estimation -6.4339 -6.4423* -6.3993 -6.4208 -6.3943
Estimation 0.00098364 0.0010903 0.00096091 0.00093692 0.0008849*
AME Validation 0.00096132* 0.0011192 0.0010414 0.0014722 0.0024209
Total 0.00060532 0.0006709 0.00059133 0.00057657 0.0005446*
Figure 5. Table 1 (
1
model for ADSPI of SPL data series
Future Date Lower Forecast Upper Actual Error
1

Appendix A

Appendix A.1 Appendix

Figure 1 (a) : The ACF and PACF plots of original data for average daily share price indices of SPL data series

Appendix B

  1. Concentrations in Brunei Darussalam-An ARIMA Modeling Approach, 54 p. .
  2. Forecasting Exchange Rates of Bangladesh using ANN & ARIMA models: A comparative study. A K Azad , M Mahsin . International Journal of Advanced Engineering Science & Technologies 2011. (10) p. .
  3. Inflation Forecasts with ARIMA, Vector Autoregressive & Error Correction Models in Nigeria. A Uko , E Nkoro . Finance & Administrative Science 2012. 50 p. 71. (European Journal of Economics)
  4. Modeling & forecasting pelagic fish production using univariate and multivariate ARIMA models, E V Tsitsika , C Maravelias , J Haralatous . 2007. 73 p. .
  5. Modeling &Forecasting Volatility using ARIMA model. H A Al-Zeaud . Finance & Administrative Science 2011. 35 p. . (European Journal of Economics)
  6. ARIMA models to predict Next Day Electricity Prices. J Contreras , R Espinola , F J Nogales , A J Conejo . IFEE Transactions on power system 2003. 18 (3) p. .
  7. ARIMA Forecasting of Inflation in the Bangladesh Economy. K Datta . The IUP Journal of Bank Management 2011. X (4) p. .
  8. Forecasting Daily Maximum Surface Ozone, K Kumar , A K Yadav , M P Singh , H Hassan , V K Jain . 2004.
  9. Next Day Stock market Forecasting: An Application of ANN & ARIMA. N Merh , V P Saxena , K R Pardasani . The IUP Journal of Applied Science 2011. 17 (1) p. .
  10. Forecasting incidence of hemorrhagic fever with renal syndrome in China using ARIMA model. Q Liv , X Liu , B Jiang , W Yang . Biomed Central 2011. p. .
Notes
1
© 2013 Global Journals Inc. (US)
Date: 2013-01-15