Learn how IMPLAN forecasts national industry Employment, Output and income for 2025 to 2034 using COVID-adjusted ARIMA models.
OVERVIEW
IMPLAN has always provided its users with a vast amount of data for use in various types of economic analysis, with data going all the way back to 2001. Now for the first time IMPLAN is going into the future. Using our own historical data, IMPLAN has produced Industry-specific forecasts at the national level. These forecasts include key economic indicators such as Employment, Output, and Labor Income. They cover 2025 to 2034 at the national (U.S.) level. In this article we will provide a high-level overview of the forecasting technique utilized in developing IMPLAN’s forecasted data, as well as more specific methods employed to ensure consistency with IMPLAN’s existing data products.
WHAT IS AN ARIMA MODEL?
Autoregressive Integrated Moving-Average (ARIMA) is a modeling technique for analyzing and forecasting time-series data. More specifically, ARIMA models are used to generate estimates of future values given a set of past values. ARIMA modeling is a robust method by which to project time-series data into the future. ARIMA models are specified with three key parameters (p, d, and q), each of which correspond to components of the name “ARIMA” (Noble, n.d.):
- p = Autoregressive component (AR): variable of interest is regressed on its own prior values, p will tell us the number of lagged predictors included in the model.
- d = Integrated component (I): how many times the data has been differenced for it to become stationary (i.e., time-series data has a constant mean and variance).
- q= Moving average component (MA): models the error of the forecast based on previous error terms.
ARIMA models can be referred to by their order as ARIMA (p,d,q). For example, an ARIMA (1,1,1) has an autoregressive order of 1, has been differenced 1 time, and uses the forecast error from the period t-1. An autoregressive order of 1 simply means that we believe that the value in the period t is influenced by the preceding period t-1, so the previous year’s value is used to help predict the current year value. An autoregressive order of 2 would mean that we believe that the value in the period t is influenced by the periods t-1 and t-2, so the values from the previous two years are used to help predict the current year value.
Differencing means that instead of using the raw values of a time-series to forecast future values, we analyze the changes in the data from one period to the next. The purpose of differencing is to make sure that the time-series is stationary, meaning that it has a constant mean and variance over time, which is important for ARIMA modeling. A 1st difference would indicate that we are analyzing the growth from period t-1 to t. A 2nd difference means that we are analyzing how growth (1st difference) increased or decreased from the preceding period. For example, say that we have a time-series for Employment in an Industry from 2019-2024, our 1st and 2nd differences would be as follows:
Year |
Industry Employment |
1st Difference (d=1) |
2nd Difference (d=2) |
|---|---|---|---|
2019 |
50 |
n/a |
n/a |
2020 |
70 |
20 = 70 - 50 |
n/a |
2021 |
100 |
30 = 100 - 70 |
10 = 30 – 20 |
2022 |
150 |
50 = 150 - 100 |
20 = 50 – 30 |
2023 |
230 |
80 = 230 - 150 |
30 = 80 – 50 |
2024 |
320 |
90 = 320 - 230 |
10 = 90 - 80 |
The moving average (MA) term (q) tells us the number of past forecast errors (residuals) that are included in the ARIMA model. An error term helps the model to capture random shocks that cannot be predicted based on previous values and are not captured in the AR (p) component. An error term is another way to quantify how far off a predicted value is from the actual observed value in a time-series. When an ARIMA model has MA term of 1, this means that we are using the error from the previous period, t-1, to help refine the estimated value in the current period, t. This previous error term being included in the ARIMA will help the model to smooth out random fluctuations over time and produce better forecasts as a result.
When deciding which of the various possible ARIMA orders is best to use for forecasting a particular time series, model selection criterion (MSC) such as the Akaike Information Criterion (AIC) or Bayesian Information Criterion (BIC) are utilized to assess the tradeoff between model fit and model complexity. The AIC is designed to find the model that is the best at predicting future values in the time-series (favoring the most accurate), while the BIC aims to find the true underlying model (more harshly penalizing model complexity). The AIC would be a preferable MSC when looking to forecast only a few periods ahead, whereas the BIC would be more commonly used in the case of longer-term forecasts (Noble, n.d.).
When using either of these MSC to assess what the best ARIMA model is, the idea is that we would prefer to select the model with the lowest AIC or BIC. Regardless of the absolute value of the MSC, the model with the BIC or AIC closest to zero or most negative would be deemed the best fit. For example, a model with a BIC of -100 would be preferable to a model with a BIC of +1 (all else equal).
For more information on the mechanics of ARIMA models refer to the following article written by Joshua Noble from IBM - Introducing ARIMA Models.
HOW DOES IMPLAN BUILD ITS INDUSTRY FORECASTS?
In collaboration with Daniel Jerrett, PhD at SmartMacro Labs LLC, IMPLAN developed a method to estimate Industry-specific COVID Adjusted ARIMA models for the following economic indicators:
- Wage and Salary Employment
- Proprietor Employment
- Output
- Employee Compensation
- Proprietor Income
- Taxes On Production and Imports Net of Subsidies
- Other Property Income
The data being used to train each of these models represents an annual time-series of the variable in question for a particular Industry from 2001-2024 (e.g., annual Employee Compensation in Industry 1 from 2001-2024). Each of the economic indicators for an Industry will have different ARIMA model because they utilize a different set of data when estimating their parameters (i.e., different time-series). These models are subsequently used to forecast the same data point 10 periods (years) out, 2025-2034.
Prior to being used to train the ARIMA models, each time-series was transformed from their levels (the default data point) to stabilize variance and normalize the time-series. In other words, the transformations aim to make fluctuations consistent over time and reduce skewness in the data. Econometric models such as ARIMA generally work best (thus generating better forecasts) when the data is normally distributed (Chan, 2025). The two transformations that were performed on the IMPLAN data are natural logarithm and Yeo-Johnson. Both transformations help to accomplish the goal of making a time-series distribution more normalized but are used for different variables depending on the nature of the data.
Natural logarithm was applied when estimating the ARIMA models for which the underlying time-series data can only take on a positive value: Wage and Salary Employment, Proprietor Employment, Output, and Employee Compensation. Whenever the variable being analyzed can take on a positive or negative value, the natural logarithm is not a valid transformation, so each had a Yeo-Johnson transformation – which can handle negative values – applied: Proprietor Income, Taxes on Production and Imports, and Other Property Income. Note that a backwards transformation is applied for both the natural logarithm and Yeo-Johnson, so that the forecasted values are presented in levels.
As mentioned previously, the integrated component (I) denotes how many times the data has been differenced for it to become stationary. To prevent unrealistic explosive growth, the time-series data used to train the ARIMA models estimated by IMPLAN can only be differenced a maximum of one time (d = 1). Another potential parameter that could be used in cases of an ARIMA specification that has d = 1 is drift. Drift is in effect a constant term that allows the model to generate forecasts that follow a linear trend (either increasing or decreasing) over time, meaning there is a non-zero change from one year to the next. However, this drift term is only included in ARIMA (0,1,0) models when drift meets an economic significance threshold.
The final component of the estimated ARIMA models is the COVID intervention variable. This variable serves as an indicator for if the year of a particular data point was during the height of the COVID-19 Pandemic (2020 or 2021) or any other year. The rationale for including this intervention variable was to isolate the economic shocks that were evident in IMPLAN data during these years from the longer-term trend. These shocks stemmed from business interruptions, increased government subsidies, changes in Employment, etc. ARIMA models rely on past data to predict future values, and these COVID shocks often led to a divergence from historical patterns. If we were not isolating the data for the COVID years, an ARIMA model could potentially pick up on abnormal short-term trends from these shocks and weigh them too heavily when forecasting into the future. Including the COVID intervention variable is telling the model that the economic shocks in 2020 and 2021 were short-term exceptions and that we do not want to treat them as a part of the “normal” pattern.
Lastly, the Bayesian Information Criterion (BIC) was the chosen metric by which the ARIMA models were specified. The BIC was chosen because it penalizes complexity to prevent overfitting the model with too many parameters, making it a better fit to identify the true model for our 10-period forecast horizon.
As will be described in more depth in the subsequent section, there were certain instances where the ARIMA specifications for an Industry led to “flat line” forecasts where, for a particular year and industry, the forecasted value is the same as the previous year’s value. This is an artifact of the ARIMA model failing to identify a consistent trend or pattern within the time-series being used to train the model. For these select sectors, additional steps were taken to generate year-over-year growth rates in line with those of the more aggregate (2-digit NAICS) industry to which the specific industry belongs.
WHAT HAPPENS WHEN A FORECAST IS FLAT?
How does IMPLAN set growth rates for flat forecasts?
- Employment and Output — Use the 2-digit NAICS Bureau of Labor Statistics projections as the target 10-year growth rate.
- Other variables: 4-digit NAICS first — If at least two other IMPLAN Industries in the same 4-digit NAICS have good (non-flat) forecasts, use their aggregate growth rate as the target.
- Then 3-digit NAICS — If not, apply the same rule at the 3-digit NAICS level.
- Then 2-digit NAICS — If still not, apply the same rule at the 2-digit NAICS level.
The time-series data used to estimate each Industry and indicator specific ARIMA model extends back to 2001, yielding a relatively small sample, which, for certain industries, can make it difficult for an ARIMA process to distinguish between complex trends vs. random fluctuations. In this section, we will go through the steps taken to account for instances of the ARIMA model failing to identify a long-term trend using historical IMPLAN data, thus leading to projected values that do not change over time. In the models mentioned below that show flatline forecasts over multiple years, IMPLAN implemented a hierarchical system (based on data availability) to set target 10-year growth rates for the affected sectors outside of the forecast function itself. These growth rates will be based on aggregate growth derived from:
-
2-Digit NAICS Bureau of Labor Statistics projections
- Because of data availability, these are implemented only for Employment and Output. Output and Employment projections are available at the 2-digit NAICS level.
- Aggregated IMPLAN Industry forecasts bridged to North American Industry Classification System (NAICS) codes (4 to 2 Digit, with preference for more detailed NAICS) projection results.
- This approach was used for all other variables.
- Ex) Forecasted Employee Compensation in Industry 1 has the flat line behavior for the entire 10-year forecast.
- We will look for the 4-digit NAICS that Industry 1 is a part of. If there are at least two other IMPLAN Industries (say Industries 2 and 3) that fall into this 4-digit NAICS code that are not flat lines, we will use the aggregate growth rate in those as the target growth rate for Industry 1.
- If there are not at least two IMPLAN Industries with good ARIMA model estimates in that 4-digit NAICS, we then move to Industry 1’s 3-digit NAICS with the same criteria and method.
- If there are not at least two IMPLAN Industries with good ARIMA model estimates in that 3-digit NAICS, we then move to Industry 1’s 2-digit NAICS with the same criteria and method.
These methods apply to Industries where the BIC selected one of the following ARIMA specifications:
| Specification | What it means | How IMPLAN handles it |
|---|---|---|
| ARIMA (0,1,0) without drift | A random walk: the simplest integrated model (d = 1). Predicts the last observed value (Yt) for all future periods (Yt+h). | Overwrite all forecast years with aggregate growth rates. |
| ARIMA (0,0,0) | White noise. | Overwrite all forecast years with aggregate growth rates. |
| ARIMA (0,0,1) | A Moving Average model of order 1 (MA(1)) on stationary, non-differenced data: the current (mean) value is influenced by random noise and the t-1 error. | Overwrite all forecast years with aggregate growth rates. |
| ARIMA (0,0,2) | A Moving Average model of order 2 (MA(2)): random noise plus the t-1 and t-2 errors. | Overwrite all forecast years with aggregate growth rates. |
| ARIMA (0,0,3) | A Moving Average model of order 3 (MA(3)): random noise plus the t-1, t-2 and t-3 errors. | Overwrite all forecast years with aggregate growth rates. |
| ARIMA (0,1,1) | Smooths recent data but produces a horizontal line after the first forecast period. | If the model predicts a change every year, keep its forecasts. If it changes only in year 1, keep year 1 and apply aggregate growth to years 2-10. |
| ARIMA (0,1,2) | Smooths recent data but produces a horizontal line after the second forecast period. | If the model predicts a change every year, keep its forecasts. If it changes only in years 1 and 2, keep those and apply aggregate growth to years 3-10. |
A final thing to note about this projection process is the relationship between Employment and Income. In addition to the steps listed above, IMPLAN established caps for Employee Compensation per Wage and Salary Worker and Proprietor Income per Proprietor based on historical highs or lows of these values. These caps serve to prevent Employee Compensation and Proprietor Income from growing at a rate much faster or slower than the forecasted Employment.
HOW DOES IMPLAN PREVENT NEGATIVE INTERMEDIATE INPUTS?
The Leontief Production Function (LPF) tells us that:
Leontief Production Function
Output = Intermediate Inputs + Value Added
Value Added = Employee Compensation + Proprietor Income + Taxes on Production and Imports + Other Property Income
Intermediate Inputs, or purchases of non-durable goods and services that are used to produce other goods and services rather than for final consumption, cannot be less than zero (i.e., an Industry cannot be paid to receive Intermediate Inputs). From this we can discern that when subtracting Value Added from Output, that the resulting value (Intermediate Inputs) must be greater than or equal to zero.
However, during our ARIMA modeling and projecting select sectors process, we estimated future values of Output and the components of Value Added separately. This led to instances where, for a particular Industry in a particular year, Output was less than Value Added.
To ensure that our Intermediate Inputs are not negative, we identified every Industry and Year combination where adding the components of Value Added resulted in a value greater than the forecasted Output for that same Industry and Year. For each of these forecast years, Other Property Income was adjusted downward until Value Added was less than Output.
REFERENCES
Noble, J. (n.d.). What are ARIMA models? IBM.
Chan, B. (2025). Log, Box-Cox, and Yeo-Johnson: Transform Skewed Data the Right Way? Medium.
ADDITIONAL RESOURCES
BLS Industry Output and Employment Projections