Why your demand forecast is 40% wrong
The median mid-market demand forecast runs at 30-50% MAPE at the SKU-week level. 15% is achievable. Here's where the error actually lives, and what it takes to close the gap without blowing the budget.
Your demand forecast is probably 30 to 50% wrong, measured as MAPE at the SKU-week level. This is not an accusation. It's the median accuracy we see when we open the hood on forecasting systems at mid-market Australian wholesalers, retailers, and distributors. And it's expensive. A 40% MAPE on a business doing $80m in inventory cycles through in wasted working capital, expedite freight, stockouts, and overstock writedowns at somewhere between 2% and 5% of revenue every year.
The good news: 15% MAPE is achievable for the majority of SKUs at the majority of mid-market businesses with work that costs less than most people assume. The work isn't glamorous, and it isn't just a better model. It's mostly richer signals, better hierarchy, and tighter coupling between the forecast and the replenishment workflow it feeds.
Why most forecasts are bad
The typical mid-market forecast in 2026 is some combination of:
- A simple time-series model. Usually exponential smoothing, ARIMA, or Prophet. Applied at the SKU level. Fit on 2-3 years of sales history.
- A manual override layer. Planners adjust the output weekly based on knowledge the model doesn't have. Promotions. Seasonal peaks. New product launches. Known stockouts in the training data.
- A replenishment system that takes the final forecast as gospel. Or worse, that's disconnected from the forecast and uses its own reorder points.
Each piece is defensible. The combination is fragile, because the model's errors, the planner's overrides, and the replenishment system's logic don't know about each other, and they compound.
The three failure modes this produces, in order of how much money they cost:
Failure mode 1: the model ignores information the business has. A Prophet model on sales history knows nothing about next Tuesday's promotion, the cyclone forecast for North Queensland, or the fact that two SKUs were out of stock for three weeks in June. Planners know. But planner overrides are informal, poorly documented, and inconsistent. The result is a forecast that's systematically wrong in predictable ways, and nobody can tell you exactly why.
Failure mode 2: SKU-level forecasting ignores hierarchy. A 12-pack of a product and a single unit of the same product are forecast independently. When you aggregate across a category, the independence assumption falls apart, and the aggregate forecast is worse than the parts. Hierarchical reconciliation is a solved problem in the academic literature and an unsolved problem in most mid-market deployments.
Failure mode 3: the feedback loop back to the forecast is broken. The replenishment system placed an order based on last week's forecast. This week, actual sales came in different. Nothing feeds back. The forecast for next week doesn't know about the gap, and the cycle repeats. In a well-designed system, actuals would update the forecast within 24 hours.
What pulls MAPE from 40% to 15%
Four moves, in rough order of impact:
1. Add the signals your current forecast ignores. Promotions, price changes, weather (for the categories where it matters), competitor stockout data if available, holiday calendars that aren't just "December gets a bump." This usually takes a 40% MAPE forecast to 25-30% with no model change. Most of the win is in the data, not the algorithm.
2. Forecast at the hierarchy that matches how the business decides. Reconcile top-down (category-level forecasts) with bottom-up (SKU-level forecasts) using formal hierarchical methods. Not spreadsheet roll-ups. The two layers will disagree, and the reconciliation is where you pick up 3-5 percentage points of accuracy on anything that isn't a perfectly stable fast mover.
3. Move from weekly to daily forecasts on the top 20% of SKUs. The 80/20 curve is savage in retail and wholesale. 80% of the revenue is on 20% of the SKUs, and the top 20% are the ones where a daily forecast beats a weekly one because demand is volatile enough to matter.
4. Close the feedback loop. Every day's actuals update every forecast. Every promotion outcome retrains the promotion-lift model. Every planner override is logged and reviewed weekly. This is less about the model and more about operating discipline, which is why businesses that don't invest in the weekly review never achieve the accuracy their model is capable of.
A 15% MAPE forecast that the replenishment system uses in near-real-time outperforms a 10% MAPE forecast that ships in a weekly batch. Accuracy matters less than latency once you're past a certain threshold. Most businesses are still well below that threshold.
What's worth forecasting and what isn't
Not every SKU needs a sophisticated forecast. The 80/20 curve means most of the accuracy gains come from getting the top 20% of SKUs right, and the long tail can be handled with simpler methods.
- Top 20% of SKUs (revenue-weighted). Full forecasting pipeline, daily refresh, planner oversight, hierarchical reconciliation. Worth every dollar.
- Middle 60%. Weekly forecast, basic signals (price, promotion, seasonality), automated refresh, planner review in aggregate.
- Long tail 20%. Reorder-point logic with service-level targets. Don't try to forecast. The variance is too high for any model to outperform a well-chosen safety stock.
Treating every SKU the same is the single most common design mistake we see. It produces a forecast that's bad everywhere rather than good where it matters.
Organisational fixes that beat algorithmic ones
- One team owns the forecast. Not planning AND supply chain AND sales AND finance. One team. Which team is less important than the singularity.
- Forecasts are reviewed weekly against a specific benchmark. MAPE. Bias. Lift vs no-change baseline. If nobody is looking at those numbers every Monday, they'll drift in silence.
- Planner overrides are logged with reasons and reviewed. Most planner overrides are wrong. Most planners think theirs are right. Logging forces the conversation.
- The replenishment system trusts the forecast. If it doesn't, fix the forecast or fix the replenishment system. Don't live with two disagreeing sources of truth.
What it takes to ship
A proper demand forecasting engagement for a mid-market Australian wholesaler, retailer, or distributor runs 8-16 weeks for the build, with a typical payback period of 3-9 months based on inventory reduction, stockout reduction, and expedite-freight reduction combined. The first engagement is usually paid back faster than the sponsor expects, which is why forecasting investments tend to chain into predictive-replenishment work the following year.
The honest risk: if the organisation isn't ready to trust the forecast (planners keep overriding, replenishment keeps using its own reorder points, sales keeps asking for "their" version), the technology investment stops delivering after month three. Organisational readiness beats model sophistication every time.
Frequently asked
Got one of these problems in front of you?
Beyond Data runs engagements that put the ideas in this insight into practice.
More insights
Predictive maintenance: when it pays, when it doesn't
PM isn't a universal upgrade on time-based maintenance. It pays when assets are expensive, failures are gradual, and data exists. When any of those is missing, the economics collapse.
Shadow AI: 45% of your AI adoption is invisible to IT
Half your business is using AI for real work, on personal accounts, outside policy. Banning it makes the shadow harder to see. Here's the operating model that channels demand and blocks the genuinely dangerous fraction at the technology layer.