Efficiency through machine learning
Hundreds of thousands of businesses lose hundreds of thousands of dollars because they produce a large amount of surplus product that is not wanted by customers, or vice versa. This is often due to the inability to correctly predict customer demand, and this problem exists in many different industries. Modern problems require modern solutions! One of them is specialized demand forecasting systems based on machine learning technologies. We talked to Almir Davletov, a leading developer in the field, about the capabilities and peculiarities of one such system.
The huge world of modern technologies cannot be developed and serviced without a huge number of modern specialists. At what point did you realize that you would link your career with the development of IT technologies?
As a child, perhaps. I was madly fond of programming, I learned how to write code on my own in the 7th grade when I discovered a computer at school. On school computers Flash (then Macromedia, now Adobe) was installed, and there were folders with source codes of some small projects, like mini-games and interactive banners, I opened them and started to dig through them. From then on I've been in IT, you could say.
Neural networks are the advanced technology of modern mankind, which is often considered by the average person as a prototype of real artificial intelligence. As an experienced specialist, please tell me what is the difference between the "neural network", "machine learning" and "artificial intelligence" concepts?
Neural network is a type of machine learning algorithm that simulates the work of the human brain to solve certain tasks.
Machine learning is a subsection of artificial intelligence, with a primary focus on developing algorithms that can "learn" from data.
Artificial intelligence is a broader term that encompasses the creation of algorithms that can perform tasks that require intelligent analysis. It includes machine learning, neural networks, and more.
You are a specialist in machine learning technologies. What interesting projects have you been involved in in this role?
In many, of course, there have been quite a few projects. Among the interesting ones: there was a project with a chain of coffee shops in the States. For them, I was developing a model to predict the demand for baked goods sales in order to minimize unnecessary costs and increase sales revenues. They had a rule: no selling yesterday's baked goods, only fresh ones. And that rule was costing them money: either because of the surplus that had to be thrown away, or because of lost profits due to lack of baked goods.
A large coffee shop chain with many branches is a complex structure in itself. Your development had to help build a forecast both in the parent establishments and in the branches. How was the process of implementing the program in the work of the establishments going? Was this process simultaneous for all coffee shops, or was it split into several stages? If so, was there any difference between the implementation of the product in the main outlets and different branches of the franchise?
There was no concept of a main establishment or branch, it was the owner's network. The implementation of the final solution took place at a later stage, and we started the development without direct communication with the coffee shops. We only needed to know historical sales data (most recently for the past week) and surplus data from all the coffee shops to start training the model.
Machine learning technology was used for this project. How exactly was the process of putting together the algorithms for this development? What input data was used in this process?
We used standard algorithms for solving prediction problems, in our case they were ARIMA and SARIMA. I won't go into technical details, except to say that these are models that can predict time series and take into account seasonal fluctuations. From the input data we took everything that could directly or indirectly affect the result: how many pastries were ordered for each day of the week, how many were sold/thrown away, what the weather was like, what city/location the coffee shop was located in, demographics of the location, seasonality (holidays, weekends), etc. We used all the data that could directly or indirectly affect the result. After collecting historical data for the last couple of years and dividing it into training and validation data, we began training the model and observing which data affected the predictions positively and which did not affect them at all, cutting out the less useful ones, until the model learned to predict sales more or less accurately on the validation data.
It turns out that the program has "learned" to predict how many people will come today for a cup of coffee and how many for fresh baked goods? How is such a prediction made?
Coffee is a separate story, we worked with baked goods. Coffee can stay fresh longer than a morning croissant, so the whole focus was just on baked goods. We made predictions for the entire next week, pre-training the model on the sales data from the previous week, while exploring ways to optimize the algorithm. After a week, we looked at the accuracy of the predictions, made notes, made a plan for the week, and did it all over again.
When developing your product, did you pay attention to the variety of baked goods and coffee produced by the establishments? After all, different customers want different products. Can your program predict which products will be bought in a certain period of time? And how can you reach such accuracy?
Naturally, we made predictions for each type of pastry separately, that was the main difficulty. Today all croissants are sold, tomorrow all brownies are. Accuracy, in this case, is difficult to achieve, as trends change, something loses popularity over time and something gains, hindering our models from accurate prediction of the result. Factors such as the economic situation, events that change the order of things, etc. also greatly affect the quality of predictions. At such times, it is important not just to retrain the model, but also to revise the algorithms that are used for prediction so that accuracy is not lost. When everything around you changes, it is natural that the model will also change.
The system you have created has shown impressive results. However, as a rule, in such large projects, the results do not appear immediately. Tell us, when did you realize that your hard work and the program you created began to bear fruit? Were there positive changes in the work of the coffee shops right away?
We were competing with senior coffee shop managers who purchased pastries based on their feelings and experiences. They didn't rely on our predictions, and didn't even see them until the model started consistently predicting sales better than a human being. If I'm not mistaken, this took about 3 months from the start of development. That's when we started to introduce the model to coffee shops and familiarize managers with the new tool. Obviously, there was resistance at the beginning, but over time managers quickly appreciated the opportunity to take care of more important things than planning bake sales.
Demand for coffee often depends on many external factors. It changes, for example, with the changing seasons of the year, and with the weather in general. It is quite natural that on a cold fall morning, a person will be more than happy to get a cup of hot coffee. However, on a hot summer afternoon, it is rather difficult to imagine such a picture. In addition, there are other external factors such as the appearance of a major competitor, for example. Tell me how the system reacts to such changes, how do the forecasting algorithms adjust to these types of factors?
These factors are the weakest point of prediction systems, their occurrence cannot be predicted nor their effect, naturally. However, if you wisely approach the development of machine learning models in general, if you continue to study the data and to experiment, you can keep the accuracy of predictions under a certain control.
Since the development concerns the production in coffee shops of a large number of different products, how many predictions of purchases in one coffee shop per day are we talking about? How great is the accuracy of such forecasts?
Everything depends a lot on the size of the coffee house, location, and seasonality. Unfortunately, I can't give you exact figures, but as a rule, the larger the volume, the more accurate the predictions. If the model predicted the sale of 15 buns and sold 10, it is a 33% surplus. And if it predicted 50 and sold 40, that's only a 20% surplus. We achieved a 25% increase in sales and reduced the surplus by 70%, it was a huge success. Of course, the model did better in some coffee shops and worse in others, but the overall situation improved significantly, and customers were happy with the result. However, later, due to the lockdown during the pandemic, most of the coffee shops switched to customer service via delivery, but that's another story.
As a specialist who has invested a lot of effort in the development of demand forecasting technologies, can you tell us how much demand there is for such technologies today? Such developments look like a very promising product to be used in many sectors of the economy, including not only in the service sector. In your opinion, in which industries is further introduction of forecasting technologies feasible?
They are indeed very much in demand. Predictive technologies are already actively used in retail, manufacturing, financial services, energy, healthcare, and a host of other areas. With the growing amount of data and improved algorithms, these technologies will become increasingly efficient and therefore will gain more and more popularity, expanding the scope of applications.
As a specialist with an impressive track record of developing various solutions, what is your experience with the demand forecasting system in your career? What other large and interesting projects have you been involved in?
This experience definitely holds an important place in my career. I have learned to better understand the dynamics of data changes and how external factors can influence the results of predictions. I still apply this knowledge to this day. As for other projects, I optimized, for instance, internal logistics at a metallurgical plant (not much machine learning there, though), helped build D2C solutions for large manufacturers of carbonated drinks, and analyzed the effectiveness of advertising campaigns.
What aspects should an aspiring specialist who has decided to connect his/her activity with demand forecasting technologies pay attention to?
From the obvious — to study the fundamentals, of course, without those there is no way. Then — data quality, without accurate data it will be difficult to identify patterns. High granularity of data can also help you understand the dynamics better. And then be patient, think outside the box, and test all the hypotheses that come to mind. One day, one of them may turn out to be correct and become the basis for a future decision.
(Disclaimer: Devdiscourse's journalists were not involved in the production of this article. The facts and opinions appearing in the article do not reflect the views of Devdiscourse and Devdiscourse does not claim any responsibility for the same.)
Google News