Self-service Analytics, or Your Next Step Towards a Data-driven Company


Editor’s note: Marina explains why self-service analytics has gained much traction and shares how to leverage its capabilities to the fullest. If you consider integrating self-service analytics into your analytics environment or improving the existing solution, don’t hesitate to turn to ScienceSoft’s data analytics consulting services for professional assistance.

Nowadays, with data being widely recognized as a valuable asset, I see how companies are willing to make data analysis easily available for more business users. And the market offers them a solution for that – self-service data analytics. The term speaks for itself: when employed, self-service analytics allows business users to perform queries and generate reports on their own.

I bet this ambitious possibility raises some questions. Does that mean that traditional data analytics, when you have to request reports from the IT department to obtain insights, is a relic of times past? Can every employee access any corporate data now? How can people with no understanding of data modeling conduct effective analysis anyway? Keep on reading for the answers and see how to make self-service analytics work for you, not against you.

self-service analytics

Top 3 reasons why I think self-service analytics is worth investing in

Data analytics is a well-known facilitator of data-driven decision-making for the business. However, there are times when traditional data analytics cannot satisfy urgent business requirements promptly. And here comes self-service analytics. Below you can see the self-service analytics benefits that make it a perfect amendment for the traditional data analysis:

1. Faster decision-making

With self-service analytics, business users don’t need to wait for the reports to be done for them: they can run queries and get whatever data they need to make timely decisions as fast as self-service analytics software allows. For example, in one of ScienceSoft’s projects, self-service analytics software enabled the customer to modify the existing reports according to the current business needs and conduct analysis just by pressing certain buttons on their user application home screen.

2. Self-sufficiency for business users + enhanced productivity for data analysts

Our customers particularly praise self-service analytics for the ability to make ad-hoc reporting and analytics accessible for employees with no technical background.

Additionally, as more employees obtain independence in running queries and conducting data analysis, data scientists and skilled analysts can shift their focus from simple analytics tasks towards their core and more complicated ones.

3. Data democratization

Self-service analytics facilitates data literacy and the spread of data-driven culture by granting access to data to a larger number of employees. Surely, it doesn’t mean that any employee gains free access to critical business data as access should be regulated by data governance policies. However, you should remember that the chosen security procedures may affect the performance of the analytics solution (e.g., it can take too long for the system to produce the required reports). To avoid such a negative outcome, I advise paying particular attention to tuning user access control. You may see this approach realized in one of our projects, where we set up an access model so that it wouldn’t slow down the overall self-service analytics solution. As a result, the customer could democratize data across their corporation with no risks of data breaches and solution performance issues.

Schedule a free demo!

ScienceSoft’s team will leverage top self-service software to present your sample data in the form of immersive reports and dashboards.

3 steps for self-service analytics success

To ensure that self-service analytics brings the desired results, I advise you to focus on:

1. Informed choice of self-service analytics tools

When choosing self-service software, you should think about such aspects as data integration capabilities, advanced analytics capabilities, speed of reporting, visualization capabilities, data security, and much more. Taking into account that there are so many factors to consider, configuring the right technology stack for your self-service solution is much easier with professional help. With experienced consultants, you will be able to opt for software that satisfies both your immediate and long-term business goals.

Many customers I work with consider Microsoft Power BI to be a great facilitator of self-service data analysis. You are welcome to read my article about Power BI pros and cons, which may help you assess the feasibility of this self-service analytics tool for your company.

2. User adoption

Remember that by granting business users access to self-service business intelligence, you don’t automatically give these users the skillset they need for leveraging its potential. To help employees embrace new capabilities, I advise you to:

  • Choose software with a simple and intuitive user interface, so that non-technical users could master it. In case you wonder what an intuitive user interface may look like, feel free to watch our Power BI demo.  
  • Follow the agile adoption approach: don’t rush your employees into harnessing the full breadth of your new software. Let them feel and appreciate the first tangible outcomes and gradually encourage them to explore software further.
  • Conduct proper training for end-users to ensure that self-service analytics software is in full compliance with the skills they possess.

3. Data management procedures

Keep in mind that delivered reports are only as accurate as the data depicted in them, so you’ll have to set up proper data management procedures. Regardless of the fact that you employ self-service analytics software, the role of a data analyst is still crucial. Only a professional data analyst can perform such activities as data cleaning, preparing data sets that are further used by end-users, conducting advanced data analysis and monitoring system performance to ensure its high effectiveness.

So, how to let your business users own the data?

Although self-service analytics empowers more business users to make better decisions at the speed of business, reaching this level of data maturity and scaling your corporate analytical culture can be tough. You need to equip business users with the properly chosen self-service tools, the right level of access to data based on their business roles, and the guidance they need. More than often, achieving it is impossible without professional help. If you don’t know how to start your transition to a truly data-driven company, or you’ve encountered some problems within the existing self-service analytics solution, I am here to offer my assistance.


Are you striving for informed decision-making? We will convert your historical and real-time data into actionable insights and set up forecasting.

7 Major Big Data Challenges and Ways to Solve Them


Before going to battle, each general needs to study his opponents: how big their army is, what their weapons are, how many battles they’ve had and what primary tactics they use. This knowledge can enable the general to craft the right strategy and be ready for battle.

Just like that, before going big data, each decision maker has to know what they are dealing with. Here, our big data consultants cover 7 major big data challenges and offer their solutions. Using this ‘insider info’, you will be able to tame the scary big data creatures without letting them defeat you in the battle for building a data-driven business.

Big data challenges

Challenge #1: Insufficient understanding and acceptance of big data

Oftentimes, companies fail to know even the basics: what big data actually is, what its benefits are, what infrastructure is needed, etc. Without a clear understanding, a big data adoption project risks to be doomed to failure. Companies may waste lots of time and resources on things they don’t even know how to use.

And if employees don’t understand big data’s value and/or don’t want to change the existing processes for the sake of its adoption, they can resist it and impede the company’s progress.

Solution:

Big data, being a huge change for a company, should be accepted by top management first and then down the ladder. To ensure big data understanding and acceptance at all levels, IT departments need to organize numerous trainings and workshops.

To see to big data acceptance even more, the implementation and use of the new big data solution need to be monitored and controlled. However, top management should not overdo with control because it may have an adverse effect.

Challenge #2: Confusing variety of big data technologies

Variety of big data technologies

It can be easy to get lost in the variety of big data technologies now available on the market. Do you need Spark or would the speeds of Hadoop MapReduce be enough? Is it better to store data in Cassandra or HBase? Finding the answers can be tricky. And it’s even easier to choose poorly, if you are exploring the ocean of technological opportunities without a clear view of what you need.

Solution:

If you are new to the world of big data, trying to seek professional help would be the right way to go. You could hire an expert or turn to a vendor for big data consulting. In both cases, with joint efforts, you’ll be able to work out a strategy and, based on that, choose the needed technology stack.

Read more:

Challenge #3: Paying loads of money

Big data on-premises vs. in-cloud costs

Big data adoption projects entail lots of expenses. If you opt for an on-premises solution, you’ll have to mind the costs of new hardware, new hires (administrators and developers), electricity and so on. Plus: although the needed frameworks are open-source, you’ll still need to pay for the development, setup, configuration and maintenance of new software.

If you decide on a cloud-based big data solution, you’ll still need to hire staff (as above) and pay for cloud services, big data solution development as well as setup and maintenance of needed frameworks.

Moreover, in both cases, you’ll need to allow for future expansions to avoid big data growth getting out of hand and costing you a fortune.

Solution:

The particular salvation of your company’s wallet will depend on your company’s specific technological needs and business goals. For instance, companies who want flexibility benefit from cloud. While companies with extremely harsh security requirements go on-premises.

There are also hybrid solutions when parts of data are stored and processed in cloud and parts – on-premises, which can also be cost-effective. And resorting to data lakes or algorithm optimizations (if done properly) can also save money:

  1. Data lakes can provide cheap storage opportunities for the data you don’t need to analyze at the moment.
  2. Optimized algorithms, in their turn, can reduce computing power consumption by 5 to 100 times. Or even more.

All in all, the key to solving this challenge is properly analyzing your needs and choosing a corresponding course of action.

Challenge #4: Complexity of managing data quality

Data from diverse sources

Sooner or later, you’ll run into the problem of data integration, since the data you need to analyze comes from diverse sources in a variety of different formats. For instance, ecommerce companies need to analyze data from website logs, call-centers, competitors’ website ‘scans’ and social media. Data formats will obviously differ, and matching them can be problematic. For example, your solution has to know that skis named SALOMON QST 92 17/18, Salomon QST 92 2017-18 and Salomon QST 92 Skis 2018 are the same thing, while companies ScienceSoft and Sciencesoft are not.

Unreliable data

Nobody is hiding the fact that big data isn’t 100% accurate. And all in all, it’s not that critical. But it doesn’t mean that you shouldn’t at all control how reliable your data is. Not only can it contain wrong information, but also duplicate itself, as well as contain contradictions. And it’s unlikely that data of extremely inferior quality can bring any useful insights or shiny opportunities to your precision-demanding business tasks.

Solution:

Big data quality

There is a whole bunch of techniques dedicated to cleansing data. But first things first. Your big data needs to have a proper model. Only after creating that, you can go ahead and do other things, like:

  • Compare data to the single point of truth (for instance, compare variants of addresses to their spellings in the postal system database).
  • Match records and merge them, if they relate to the same entity.

But mind that big data is never 100% accurate. You have to know it and deal with it, which is something this article on big data quality can help you with.

Challenge #5: Dangerous big data security holes

Security challenges of big data are quite a vast issue that deserves a whole other article dedicated to the topic. But let’s look at the problem on a larger scale.

Quite often, big data adoption projects put security off till later stages. And, frankly speaking, this is not too much of a smart move. Big data technologies do evolve, but their security features are still neglected, since it’s hoped that security will be granted on the application level. And what do we get? Both times (with technology advancement and project implementation) big data security just gets cast aside.

Solution:

The precaution against your possible big data security challenges is putting security first. It is particularly important at the stage of designing your solution’s architecture. Because if you don’t get along with big data security from the very start, it’ll bite you when you least expect it.

Challenge #6: Tricky process of converting big data into valuable insights

Valuable insights with big data

Here’s an example: your super-cool big data analytics looks at what item pairs people buy (say, a needle and thread) solely based on your historical data about customer behavior. Meanwhile, on Instagram, a certain soccer player posts his new look, and the two characteristic things he’s wearing are white Nike sneakers and a beige cap. He looks good in them, and people who see that want to look this way too. Thus, they rush to buy a similar pair of sneakers and a similar cap. But in your store, you have only the sneakers. As a result, you lose revenue and maybe some loyal customers.

Solution:

The reason that you failed to have the needed items in stock is that your big data tool doesn’t analyze data from social networks or competitor’s web stores. While your rival’s big data among other things does note trends in social media in near-real time. And their shop has both items and even offers a 15% discount if you buy both.

The idea here is that you need to create a proper system of factors and data sources, whose analysis will bring the needed insights, and ensure that nothing falls out of scope. Such a system should often include external sources, even if it may be difficult to obtain and analyze external data.

Challenge #7: Troubles of upscaling

The most typical feature of big data is its dramatic ability to grow. And one of the most serious challenges of big data is associated exactly with this.

Your solution’s design may be thought through and adjusted to upscaling with no extra efforts. But the real problem isn’t the actual process of introducing new processing and storing capacities. It lies in the complexity of scaling up so, that your system’s performance doesn’t decline, and you stay within budget.

Solution:

The first and foremost precaution for challenges like this is a decent architecture of your big data solution. As long as your big data solution can boast such a thing, less problems are likely to occur later. Another highly important thing to do is designing your big data algorithms while keeping future upscaling in mind.

But besides that, you also need to plan for your system’s maintenance and support so that any changes related to data growth are properly attended to. And on top of that, holding systematic performance audits can help identify weak spots and timely address them.

Win or lose?

As you could have noticed, most of the reviewed challenges can be foreseen and dealt with, if your big data solution has a decent, well-organized and thought-through architecture. And this means that companies should undertake a systematic approach to it. But besides that, companies should:

  • Hold workshops for employees to ensure big data adoption.
  • Carefully select technology stack.
  • Mind costs and plan for future upscaling.
  • Remember that data isn’t 100% accurate but still manage its quality.
  • Dig deep and wide for actionable insights.
  • Never neglect big data security.

If your company follows these tips, it has a fair chance to defeat the Scary Seven.


Big data is another step to your business success. We will help you to adopt an advanced approach to big data to unleash its full potential.

4 Retail Data Analytics Trends to Win More Sales


The Christmas season is over, and you’ve heroically lived through it. The beginning of the year is a great time to look around and check the latest retail trends. We’ve scanned the expectations and selected the ideas that are likely to add value. The compilation has another unifier: all the initiatives should be supported with data analytics. Let’s have a close look at the four chosen trends.

Retail data analytics

1. Going omnichannel

Retail started to move beyond brick-and-mortar stores long ago. However, many retailers still consider making the first step in this direction. If you are among those, you need a proper retail data analytics solution.

Let’s say, a consumer electronics retailer runs both physical stores and an online one. The retailer’s brick-and-mortar stores are showing brilliant sales, while their online store is terribly lagging behind. Naturally, the retailer is not happy with the format that does not bring sales. However, is their e-store really useless, and they are just wasting money to keep it running?

After scrutinizing customer big data for both channels, the retailer may find out the following: 75% of website visitors are surfing the online catalog to compare product features and finish their purchases in one of the physical stores. In this case, if the retailer abandons an online store, they may also lose the majority of customers who would prefer shopping with another retailer in a way that they find convenient.

2. Creating unique customer experience

Customers make purchases in brick-and-mortar and online stores, take part in loyalty programs, create shopping lists and place orders in the apps. In a word, they interact with a retailer in different ways and expect them to offer a personalized approach in return. As a retailer, you must understand how precious customer data is, so you should be striving to repay your customers with targeted marketing campaigns and product offers.

Let’s assume, you are a drugstore retailer. You collect and analyze customer data to understand their behavior and preferences. Your analytical system knows that Customer A used to visit your store once a month to buy 3 packs of nappies, a washing powder of brand X and a dishwashing liquid of brand Y. But this month, Customer A didn’t appear. To encourage him or her to visit your store, you may then send them a 5% coupon for their favorite washing powder brand.

As a real-life example of data analytics in action, we’ve chosen Nordstrom with their initiative Reserve Online & Try In Store. Their app users can select items and book them for trying on in a particular store and at a convenient time. Nordstrom has gone further: the retailer can recognize if a customer is passing by their store and kindly invite them to come in. How can anyone stay indifferent when they receive the message: Hello from Nordstrom. It looks like you are nearby, and your item is ready to try! To crown it all, once a customer is in the store, he or she will find their name on the door of a fitting room. Here is the personalized approach in action.

3. Dynamic pricing to stay competitive

Once competitive intelligence was a challenge for brick-and-mortar retailers. Naturally, the rivals were unwilling to share any information, and price monitoring at the competitor’s stores was time-consuming, prone to mistakes and exhausting. In the era of ecommerce, retailers can benefit from new approaches to competitive intelligence. By definition, online stores are publicly available, and a lot of information such as product details, promo offers, prices and category hierarchy is always at hand. This all made dynamic pricing possible, when the system is able to scan competitors’ prices in real time, run a complex analysis very quickly and change a retailer’s prices automatically based on the defined rules. Now, if an ecommerce retailer wants to be 5% cheaper than their competitors are, with big data analytics they are able to do this.

Brick-and-mortar retailers can also rely on big data analytics, though some time ago they experienced some implementation limitations, as store employees had to replace price tags manually, which was time-consuming. With the invention of electronic price tags, dynamic pricing has become available to brick-and-mortar retailers, too. In no time, the system can change prices based on analyzing the sales, stock, competitors’ prices, customer demand, shelf life, etc.

4. Building effective relationships with suppliers

An industry benchmark, Walmart laid the rules on how to collaborate with the suppliers. And this policy can add $1 billion to Walmart’s revenue. The retailer has implemented a scoring system to assess the suppliers according to On-Time, In-Full principle. In simple words, either delivery is late or early, the supplier should pay a fine. The same scenario works if the supplier delivers timely, but the quantity of goods is wrong, their quality is poor, or packaging is damaged.

If you are to follow Walmart’s best practice, you can also think about tuning your data analytics system to differentiate between strategic and non-strategic suppliers, critical and non-critical product categories. Additionally, you can set different thresholds. For instance, the critical number of troublesome deliveries may be 10%.

With such a scoring system, a retailer will easily identify whether a supplier is reliable or not. Besides, this approach will contribute to a more efficient inventory management. So, a retailer’s persistent headache caused by over-stocks and out-of-stocks should finally subside.

To sum it up

Retailers’ aspirations don’t change significantly with the time. Satisfying customer needs better, outperforming competitors and building up effective relationships with suppliers are at the top of every retailer’s wish list. Despite the retail industry is constantly developing and there appear new sophisticated ways to solve daily and strategic challenges, this does not guarantee that every retailer reaches the desired result. However, supporting the initiatives with data analytics can make the difference. 


Are you striving for informed decision-making? We will convert your historical and real-time data into actionable insights and set up forecasting.

Demand Forecasting Using Data Science


From our consulting practice, we know that even the companies that have put significant effort into demand forecasting can still go the extra mile and improve the accuracy of their predictions. So, if you’re one of the companies who want reliable demand forecasts on their radars, this is the right page for you.

Though a 100% precision is impossible to achieve, we believe data science can get you closer to it, and we’ll show how. Our data scientists have chosen the most prominent demand forecasting methods based on both traditional and contemporary data science to show you how they work and what their strengths and limitations are. We hope that our overview will help you opt for the right method, which is one of the essential steps to creating a powerful demand forecasting solution.

Demand forecasting using data science

Traditional data science: The ARIMA model

A well-known traditional data science method is the autoregressive integrated moving average (ARIMA) model. As the name suggests, its main parameters are autoregressive order (AR), integration order (I) and moving average order (MA).

The AR parameter identifies how the values of the previous period influence the values of the current period. For example, tomorrow the sales for SKU X will be high if the sales for SKU X were high during the last three days.

The I parameter defines how the difference in the values of the previous period influence the value in the current period: tomorrow the sales for SKU X will be the same if the difference in sales for SKU X was minimal during the last three days.

The MA parameter identifies the model’s error based on all the observed errors in its forecasts.

Strengths of the ARIMA model

  • ARIMA works well when the forecast horizon is short-term and when the number of demand-influencing factors is limited.

Limitations of the ARIMA model

  • ARIMA is unlikely to produce accurate long-term forecasts as it doesn’t store insights for long time periods.
  • ARIMA assumes that your data doesn’t show any trend or seasonal fluctuations, while these conditions are sure not to be met in real life.
  • ARIMA requires extensive feature engineering efforts to capture root causes of data fluctuations and that is a lengthy and labor-intensive process. For example, a data scientist should mark particular days of the month as weekends for ARIMA to take into account this factor. Otherwise, it won’t recognize the impact of a particular day on sales.
  • The model can be time-consuming as every SKU or subcategory requires separate tuning.
  • It can only handle numerical data, such as sales values. This means that you can’t take into account such factors as weather, store type, store location and promotion influence.
  • It fails to capture non-linear dependencies, and that’s the kind of dependencies that is most frequent. For example, with 5% off promotion, toys from Frozen witnessed a 3% increase in sales. If the discount becomes twice higher – 10%, this doesn’t mean that the company should expect a double increase in sales to 6%. Besides, if they run a 5% promotion for Barbie dolls, their sales can increase by 9% as promotion influences various categories differently.

Contemporary data science: Deep neural networks

Since there are so many limitations to traditional data science, it’s natural that there are other, more reliable approaches, namely contemporary data science. There’s no better candidate to represent contemporary data science than a deep neural network (DNN). Recent research papers show that DNNs outperform all the other forecasting approaches in terms of effectiveness and accuracy of predictions. To usher you into the promising world of deep learning, our data scientists composed a 5-minute introduction to DNNs that comprises both the theory part and the practical example.

What are DNNs made of?

Deep neural network architecture

Here’s the architecture of a standard DNN. To read this scheme, you should know just 2 terms – a neuron and a weight. Neurons (also called ‘nodes’) are the main building blocks of a neural network. They are organized in layers to transmit the data along the net, from its input layer all the way to the output one.

As to the weights, you can regard them as coefficients applied to the values produced by the neurons of the previous layer. Weights are of extreme importance as they transform the data along its way through a DNN, thus influencing the output. The more layers a DNN has or the more neurons each layer contains, the more weights appear.

What data can DNNs analyze?

DNNs can deal equally well with numerical and categorical values. In the case with numerical values, you give the network all needed figures. And in case with categorical values, you’ll need to use ‘0-1’ language. It usually works like this: if you want to input a particular day of the week (say, Wednesday), you should have seven neurons, and you’ll give 1 to the third neuron (which will mean Wednesday) and zeroes to all the rest.

The vast diversity of data that a DNN is able to ingest and analyze allows considering multiple factors that can influence demand, thus improving the accuracy of forecasts. The factors can be internal, such as store location, store type and promotion influence, and external ones – weather, changes in GDP, inflation rate, average income rate, etc.

And now, a practical example. Say, you are a manufacturer who uses deep neural networks to forecast weekly demand for their finished goods. Then, you may choose the following diverse factors and data for analysis.

Factors to analyze What each factor reflects Number of neurons for the input layer
8 previous weeks’ sales figures Latest trends 8
Weeks of the year Seasonality 52 (according to the number of weeks in a year)
SKUs Patterns specific to each SKU 119 (according to the number of SKUs in your product portfolio)
Promotion The influence of promotion 1 (Yes or No)
    Total number of input neurons: 180

In addition to showing the diversity of data, the table also draws the connection between the business and technical aspects of the demand forecasting task. Here, you can see how factors are finally converted into neurons. This information will be useful for understanding the sections that follow.

Where does DNN intelligence come from?

There are two ways for a DNN to get intelligence, and they peacefully coexist. Firstly, this intelligence comes from data scientists who set the network’s hyperparameters and choose most suitable activation functions. Secondly, to put its weights right, a DNN learns from its mistakes.

Activation functions

Each neuron has an activation function at its core. The functions are diverse and each of them takes a different approach to converting the values they take in. Therefore, different activation functions can reveal various complex linear and non-linear dependencies. To ensure the accuracy of demand forecasts and not to miss or misinterpret exponential growth or decline, surges and temporary falls, waves, and other patterns that data shows, data scientists carefully choose the best set of activation functions for each case.

Hyperparameters

There are dozens of hyperparameters, but we’d like to focus on a more down-to-earth one, such as the number of hidden layers required. Choosing this parameter right is critical for making a DNN able to identify complex dependencies. The more layers, the more complex dependencies a DNN can recognize. Each business task, and consequently, each DNN architecture designed to solve this task, requires an individual approach to the number of its hidden layers.

Suppose in our example, data scientists decided that the neural network requires 3 hidden layers. They also came up with the coefficients that change the number of neurons in the hidden layers (these coefficients are always applied to the number of neurons in the input layer). Here are their findings:

Layer Coefficient Number of neurons in the layer
Input layer   180
Hidden layer 1 1.5 270
Hidden layer 2 1 180
Hidden layer 3 0.5 90
Output layer   1
    Total number of neurons in the network: 721

Usually, data scientists create several neural networks and test which one shows better performance and higher accuracy of predictions.

Weights

To work properly, a DNN should learn which of its actions is right and which one is wrong. Let’s look at how the network learns to set the weights right. At this stage, regard it as a toddler who learns from their personal experience and with some supervision of their parents.

The network takes the inputs from your training data set. This data set is, in fact, your historical sales data broken down to SKU and store level, which may also contain store attributes, prices, promotions, etc. Then, the network lets this data pass through its layers. And, at first, it applies random weights to it and uses predefined activation functions.

However, the network doesn’t stop when it produces an output – a weekly demand for SKU X. Instead, it uses loss function to calculate to which extent the output the network got differs from the one that your historical data shows. Then, the network triggers optimization algorithms to reassign the weights and starts the whole process from the very beginning. The network repeats this as many times (can be thousands and millions) as needed to minimize the mistake and produce an optimal demand.

To let you understand the scale of it all: the number of weights that a neural network tunes can reach hundreds of thousands. In our example, we’ll deal with 113,490 weights. No serious math is required to get this figure. You should just multiply the number of neurons in one layer by the number of neurons in the layer that follows and sum it all up: 180×270 + 270×180 + 180×90 + 90×1 = 113,490. Impressive, right?

Demand forecasting challenges that DNNs overcome

New product introduction

Challenge: Historical data is either limited or doesn’t exist at all.

Solution: A DNN allows clustering SKUs to find lookalikes (for instance, based on their prices, product attributes or appearance) and use their sales histories to bootstrap forecasting.

The thing is that you have all the historical data for the lookalikes because they are your tried-and-tested SKUs. So, you can take their weekly sales data and use it as a training data set to estimate the demand for a new product. As discussed earlier, you can also add external data to increase the accuracy of demand predictions – for example, social media data.

Another scenario here could be: a DNN is tuned to cluster new products according to their performance. This helps to predict how a newly launched product will perform based on its behavior at the earliest stages compared to the behavior of other new product launches.

Complex seasonality

Challenge: For some products (like skis for the winter or sunbathing suits for the summer), the seasonality is obvious, while for others, the patterns are not so easy to spot. If you are looking for multiple seasonal periods or high-frequency seasonality, you need something more efficient than trivial methods.

Solution: Just like with new product introductions, the task of identifying complex seasonality can be solved with the help of clustering. A DNN sifts through hundreds and thousands of sales patterns of each SKU to find similar ones. If particular SKUs belong to the same cluster, they are likely to show the same sales patterns in the future.

Weighing the pros and cons of DNNs

Now that we know how a DNN works, we can consider the upsides and downsides of this method.

Strengths of DNNs

Compared to traditional data science approaches, DNNs can:

  • Consider multiple factors based on diverse data (both external and internal, numerical and categorical), thus increasing the accuracy of forecasts.
  • Capture complex dependencies in data (both linear and non-linear) thanks to multiple activation functions embedded into the neurons and cleverly set weights.
  • Successfully solve typical demand forecasting challenges, such as new product introductions and complex seasonality.

Limitations of DNNs

Although DNNs are the smartest data science method for demand forecasting, they still have some limitations:

  • DNNs don’t choose analysis factors on their own. If a data scientist disregards some factor, a DNN won’t know of its influence on the demand.
  • DNNs are greedy for data to learn from. The size of the training data set should not be less than the number of weights. And, as we have already discussed, you can easily end up with hundreds of thousands of weights. Correspondingly, you’ll need as many data records.
  • If a DNN is trained incorrectly, it can fail to distinguish erroneous data from the meaningful signals. As a result, such a network can produce accurate forecasts on the training data but bring up distorted outputs while dealing with new incoming data. This problem is called overfitting, and data scientists can fight it using a dropout technique.
  • Non-technical audience tends to perceive DNNs as ‘magic boxes’ that produce ungrounded figures. You should put some effort into making your account managers trust DNNs.
  • DNNs still can’t take into account force majeure, like natural disasters, government decisions, etc.

So, where does your heart lie?

From our consulting experience, we see that contemporary data science in most cases outperforms traditional methods, especially when it comes to identifying non-linear dependencies in data. However, this doesn’t mean that traditional data science methods should be completely disregarded. They still can be considered for producing short-term forecasts. For example, recently we successfully delivered sales forecasting for an FMCG manufacturer, where we applied linear regression, ARIMA, median forecasting, and zero forecasting.


Bringing data science on board is promising, yet difficult. We’ll solve all the challenges and let you enjoy the advantages that data science offers.

Hire Data Scientists Efficiently with Our 3 Tips


Optimized supply chains, improved production efficiency, personalized customer experience, and boosted sales effectiveness are just some of the gains that our customers pursue when they turn to data science consulting services. Facing the growing demand for data science talent, we, at ScienceSoft, decided to cover a burning topic of how to hire data scientists. Here, we answer 3 main questions: what skills a data scientist should possess, how to assess those skills, and where to search for the right person.

how to hire a data scientist

What data scientist do you need?

To narrow the initial list of candidates down and make the shortlisting pipeline more efficient, we recommend that you clearly define a needed data scientist’s profile. With a vast variety of skills that a data scientist is expected to possess (including the endless list of big data technologies and machine learning algorithms), you can never find a data science unicorn who handles all these things with the same mastery.

So, you can devise an ideal data scientist’s profile on your own or find the appropriate option among the existing classifications. For example, ScienceSoft adheres to a classification that recognizes 2 data scientists types: analysts and technicians.

How to assess data scientists’ skills?

The approach to skills assessment depends on which of the 3 scenarios listed below your company favors:

  1. Growing in-house data science capabilities (this scenario also covers team augmentation).
  2. Resorting to data science consulting services (when you hire an external consultant for knowledge transfer to boost the development of your internal data science capabilities).
  3. Outsourcing data science (when you don’t plan to develop in-house data science capabilities).

Approach 1. When you search to grow in-house data science capabilities.

  • Check the candidates’ CVs.
  • Challenge a candidate with a test to validate their skills.
  • (Optional) Organize an in-house data challenge (for example, the way Airbnb does).

Approach 2. When you search for a consulting/outsourcing partner.

  • Check a candidate company’s competence and experience: study their portfolio of implemented projects, check the attained partnerships and certificates.
  • Ask to deliver a proof of concept (for complex projects).

Where to find a data scientist?

Now, when you know whom to chase, let’s discuss where to search for. Job sites, recruitment agencies, and professional networks like LinkedIn is the triad that easily comes to mind. However, considering the shortage of data scientists, these traditional resources may turn out to be insufficient. In addition, these channels are mainly tuned to hiring data scientists for growing in-house teams. If you consider data science consulting or outsourcing rather than team augmentation, ScienceSoft recommends turning your attention to three extra sources:

  • Tech communities like GitHub and Stack Overflow – you’ll find the profiles of data scientists there.
  • Listings, like this one featuring the best data science consultancies.
  • Homepages of data science consulting and outsourcing companies where you can check the service and project portfolio of a certain vendor.

Let the effective search for data scientists begin!

Now, you know what these fantastic data scientists are and where to find them. We hope that our tips will help you make your hunt for data scientists efficient and fast, and your data science-powered projects a true success.


Bringing data science on board is promising, yet difficult. We’ll solve all the challenges and let you enjoy the advantages that data science offers.

Ecommerce Business Intelligence: Features, Gains, Costs


SAP BusinessObjects Business Intelligence

Oracle Business Intelligence

Basic ecommerce BI capabilities

Advanced visualization capabilities

Native: with other SAP products.

Via SAP Translator Program or RFC interface: with third-party software.

Native: with other Oracle products.

Via Oracle connectors: with third-party software.

Native: with Microsoft products.

Via APIs: with third-party software.

Seamless integration with any required business solution (including legacy software) and third-party systems.

Compliance with global data security standards.

Compliance with global data security standards.

Compliance with global data security standards.

Compliance with all required global and regional ecommerce data protection regulations.

Pricing 50 users/ 3 years of use

Upon request to the vendor

Expect to pay initial setup fees + configuration, customization, and integration fees + maintenance and support fees.

~$370,000 depending on the use of data integration functionality.

NB: Updates and maintenance costs as well as analytics server administrator rights fees are not included.

$200,000 with data-sharing rights for all users.

NB: may require updates and maintenance costs.

Upfront investments of $50,000$1,000,000, depending on solution complexity.

Unlimited number of users, integrations, and any required advanced capabilities included.

No additional fees.

Making Use of the Alliance


Having implemented a business intelligence solution long ago, you keep monitoring recent trends to understand what the market can offer. While doing your research, you must have noticed that business intelligence often goes side by side with data analytics. Does this mean that you can enrich your existing BI solution by following promising data analytics trends? Is the synergy of BI and DA possible? Let’s find this out.

BI and data analytics

Business intelligence and data analytics – two sides of the same coin          

A quick search on the internet is enough to understand that these two terms are used inconsistently: different vendors, service providers and other market players keep to their internal definitions. That is why, some sources consider business intelligence and data analytics two different concepts, while the others use these terms interchangeably. We define them as follows:

Business intelligence (BI) is a technology-based process of analyzing data and presenting actionable insights to help business users make informed decisions. The implementation of BI includes three main stages:

  • Developing a data warehouse
  • Designing online analytical processing (OLAP) cubes
  • Visualizing data.

Data analytics is a catchall term that encompasses business intelligence as well as advanced approaches and methods to collecting, processing and analyzing data sets to identify trends, dependencies and correlations. The term is broad and is applicable to both business and science. Data analytics includes:

  • Data mining
  • Predictive and prescriptive analytics
  • Big data analytics, etc.

How to achieve the synergy of BI and DA

Usually invisible to end business users, data analytics uses complicated algorithms and statistical approaches to provide extra insights, which then can enrich habitual reports. Here, we share some illustrative examples of how business intelligence and data analytics can work together.

  1. Cohort analysis allows considering online store visitors not as a whole, but broken down into different user groups that show similar behavior patterns. Such groups may become a dimension for the OLAP cube. Business decision makers can compare them by sales, profit, the number of orders per month, etc. to design personalized marketing activities.
  2. Regression analysis allows identifying the relationship between variables. The dependency (or the lack of dependency) between them can provide companies with extra insights, as opposed to historical data alone. For instance, it is interesting to look at the total number of complaints and top-10 complaints. But with regression analysis, you may also find out whether the wait time and the number of complaints are connected.
  3. Time series analysis is applied to historical data to create forecasts. Let’s say, you want to predict sales. For this, you need to have sales figures for several previous years, split by month. Based on this data, an analytical system will identify past trends, monthly growth/decline rates, repeating patterns, if any, and will make the best possible estimate for the future.

Data analytics trends are definitely worth attention

Let’s look at some data analytics trends and find out how they can help enrich your existing BI solution. By the way, you may find the same trend in both BI and data analytics lists (do you remember the inconsistency of the terms we mentioned above’). The implementation of some initiatives will cause little pain, while others may require significant changes in the technology stack, approaches and methods.

1. Machine-learning based artificial intelligence

Let’s talk here about the problem of customer churn. Traditional BI solutions help you understand how many customers left you last week/quarter/month. Looking at the churn rates, you naturally start thinking about how to return these customers. Still, the moment is gone – the customers have already switched to your competitors, and now you’ll have to do your best to win them back.

With machine learning-based AI, businesses can identify high-risk customer segments well in advance. The analytical system can assess customers’ activities across all channels and signal if their behavior looks like they are going to leave. For example, a customer contacts the support center more frequently than the average customer does, or they start using your services less often, or their average spend significantly decreases. Of course, the set of symptoms will be specific to each industry. And you need to identify the ones that are essential for your business, score each symptom and let your analytical system learn. As a result, your system will inform you about a possible churn well in advance for you to take actions, such as targeted marketing campaigns.

2. Predictive analytics

Having timely and accurate reports that depict historical data is great. However, many businesses may find this insufficient. In fact, companies also need to understand what is likely to happen in the future to take preventive actions today. Here, predictive analytics comes to rescue.

Imagine a manufacturer of outdoor clothing who is planning their range for the next winter season. They may look at past sales split by category and plan production volumes accordingly. But will it be insightful? Alternatively, they may apply a time series analysis that we have described above.

As fashion is fast-paced, the manufacturer needs even more insights to forecast customer demand, decide on the winter range and plan the production volume for every item. For example, the producer can additionally analyze weather forecasts (the colder the winter, the less three-quarter coats should be in the range) and the trends that are becoming popular on social media.

3. Big data

If your business is on the brink of an important change that will require collecting, processing and analyzing big data (such as installing sensors to your machinery to foster preventive maintenance or launching an e-store in addition to brick-and-mortar ones, etc.), your analytical system should also be able to handle this new challenge. A traditional BI solution should be extended. Big data requires a dedicated technology stack, such as Apache Hadoop, Apache Hive, Apache Spark, etc. To get valuable insights from big data sets, you have a wide variety of data analytics methods and techniques at your disposal, for instance, pattern matching.

On a final note

If you have a BI solution implemented, this does not mean that you have hit the ceiling. The market is always developing and offering new business intelligence services. There are always ways to improve the existing solution.

While checking the recent trends, don’t limit yourselves to business intelligence only. Check out the list available for data analytics, as well. For some businesses, traditional business intelligence may suffice, but some companies may find it reasonable to enrich the existing solution with data analytics to get more insights.


Are you striving for informed decision-making? We will convert your historical and real-time data into actionable insights and set up forecasting.

How Big Data Influences Your IoT Solution


The number of Internet connected devices is projected to triple by 2025. Correspondingly, IoT is joining the line of important big data sources. This makes data practitioners turn their attention to IoT big data.

IoT big data nature

The nature of IoT big data

IoT big data is distinctly different from other big data types. To form a clear picture, imagine a network of sensors that continuously generate data. In manufacturing, for example, it can be the temperature values of a particular machinery part, as well as vibration, lubrication, humidity, pressure and more. So, IoT big data is machine-generated, not created by humans. And it mainly represents the flow of numbers, not chunks of text.

Now, imagine that each sensor produces 5 measurements per second and, overall, you have 1,000 sensors installed. And this high-volume data is incessantly flowing (by the way, such data has a special name – streaming data). Definitely, pure data collection is not your ultimate goal – you need valuable insights, some of which as close as possible to real-time. If the pressure starts suddenly plunging to the critical level, you won’t be happy to know about this only in a couple of hours. By that time, your maintenance team might have already been trying to repair a broken machinery unit.

Besides, IoT data is location and time specific. While examples can be numerous, here we’ll mention only a couple: location data is critical to understand which of the sensors communicates the readings that are likely to signal an upcoming failure, while a timestamp is essential to identify a particular pattern that is likely to cause a machinery breakdown. For instance, every ten seconds a temperature value increases by 5 F still without surpassing a threshold, which leads to increasing pressure by 1,000 Pa for one minute.

Storage, preprocessing and analysis of IoT big data

Of course, it’s your business objectives that always lay the foundation for the solution’s architecture. Still, the nature of IoT big data leaves its mark on data storage, preprocessing and analysis. So, let’s take a closer look at the specific features of each process.

IoT big data storage

As you’ll have to deal with high volumes of quickly arriving structured and unstructured data in different formats, a traditional data warehouse will not meet your requirements – you need a data lake and a big data warehouse. A data lake may be split into several zones such as a landing zone (for raw data in their original format), a staging zone (for the data after a basic cleaning and filtering applied and for raw data from other data sources), as well as analytics sandbox (for data science and exploratory activities). A big data warehouse is required to extract the data from a data lake, transform it and store in a more organized way.

IoT big data preprocessing

It’s important to decide whether you would like to store raw or already preprocessed data. In fact, answering this question right is one of the challenges connected to IoT big data. Let’s return to our example with a sensor that communicates 5 temperature values per second. One option is to store all 5 readings, while the other is to store only one value such as their average/median/mode per aggregation period of one second. To clearly visualize what difference such an approach makes to the required storage capacity, you should multiply the overall number of sensors by their expected running time and then by their reading frequency.  

If you belong to 70% of the organizations that value managing data in real time, and a part of your plan is getting real-time insights, it’s still possible to have real-time alerts without sending all the readings to the data storage. For example, your system is able to ingest the whole flow of data, and you’ve set critical thresholds or deviations that trigger instant alerts. Still, only some filtered or compressed data is sent to the data storage.

Ways to avoid data losses

It’s also necessary to think in advance what if the flow of readings stops for some reason, let’s say due to a temporary failure of a sensor or a loss of its connection with the gateway.

Here, two approaches are possible:

  • Using robust algorithms that are reliable to data omissions.
  • Using redundant sensors, for example, having several sensors to measure the same parameter. On the one hand, this increases reliability: if one sensor fails, the others will continue sending their readings. On the other hand, this approach requires more complicated analytics, as the sensors may generate slightly different values what should be processed by analytical algorithms.

IoT big data analysis

IoT big data demands two types of analytics: batch and streaming. Batch analytics is inherent in all big data types, and IoT big data is not an exception. It is widely used to run a complex analysis on the captured data to identify trends, correlations, patterns and dependencies. Batch analytics involves sophisticated algorithms and statistical models applied to historical data.

Streaming analytics perfectly covers all the specifics of IoT big data. It is designed to deal with high-speed flows of data generated within small time intervals and to provide near real-time insights. For different systems, this ‘real-time’ parameter will vary. In some cases, it can be measured in milliseconds, while in others – in several minutes. To get insights as fast as possible, the captured data can be analyzed at the system’s edge or even in a data streaming processor.

To sum it up

By nature, IoT big data is machine-generated, high-volume, streaming, location and time specific. Big data consulting practice proves how important it is to have these features considered prior to designing and developing an IoT solution. We are sure that you don’t want to run out of storage space in just a couple of months, or miss real-time insights just because your solution does not support streaming analytics, or face any other problem that undermines the robustness of your IoT solution. To avoid this, it’s necessary to clearly identify your short-term and long-term business requirements, as well as carefully choose an optimal big data architecture and technology stack from multiple options.


Big data is another step to your business success. We will help you to adopt an advanced approach to big data to unleash its full potential.

twins or just strangers with similar looks?


Apache Cassandra and Apache HBase are much like two strangers whom you meet in the street and think to be twins. You don’t really know them, but their similar height, clothes and hairstyles make you see no differences between them. However, after having a closer look, you realize that these two looked identical only at a distance.

Having numerous similarities, like being NoSQL wide-column stores and descending from BigTable, Cassandra and HBase do differ. For instance, HBase doesn’t have a query language, which means that you’ll have to work with JRuby-based HBase shell and involve extra technologies like Apache Hive, Apache Drill or something of the kind. While Cassandra can boast its own CQL (Cassandra Query Language), which Cassandra specialists find most helpful.

Cassandra vs. HBase

1. Data model

HBase

HBase data model

Here we have a table that consists of cells organized by row keys and column families. Sometimes, a column family (CF) has a number of column qualifiers to help better organize data within a CF.

A cell contains a value and a timestamp. And a column is a collection of cells under a common column qualifier and a common CF.

Within a table, data is partitioned by 1-column row key in lexicographical order, where topically related data is stored close together to maximize performance. The design of the row key is crucial and has to be thoroughly thought through in the algorithm written by the developer to ensure efficient data lookups.

Cassandra

Cassandra data model

Here we have a column family that consists of columns organized by row keys. A column contains a name/key, a value and a timestamp. In addition to a usual column, Cassandra also has super columns containing two or more subcolumns. Such units are grouped into super column families (although these are rarely used).

In the cluster, data is partitioned by a multi-column primary key that gets a hash value and is sent to the node whose token is numerically bigger than the hash value. Besides that, the data is also written to an additional number of nodes that depends on the replication factor set by Cassandra practitioners. The choice of additional nodes may depend on their physical location in the cluster.

HBase vs. Cassandra (data model comparison)

The terms are almost the same, but their meanings are different. Starting with a column: Cassandra’s column is more like a cell in HBase. A column family in Cassandra is more like an HBase table. And the column qualifier in HBase reminds of a super column in Cassandra, but the latter contains at least 2 subcolumns, while the former – only one.

Besides, Cassandra allows for a primary key to contain multiple columns and HBase, unlike Cassandra, has only 1-column row key and lays the burden of row key design on the developer. Also, Cassandra’s primary key consist of a partition key and clustering columns, where the partition key also can contain multiple columns.

Despite these ‘conflicts,’ the meaning of both data models is pretty much the same. They have no joins, which is why they group topically related data together. They both can have no value in a certain cell/column, which takes up no storage space. They both need to have column families specified while schema design and can’t change them afterwards, while allowing for columns’ or column qualifiers’ flexibility at any time. But, most importantly, they both are good for storing big data.

2. Architecture

Cassandra has a masterless architecture, while HBase has a master-based one. This is the same architectural difference as between Cassandra and HDFS.

This means that HBase has a single point of failure, while Cassandra doesn’t. An HBase client does communicate directly with the slave-server without contacting the master, which gives the cluster some working time after the master goes down. But, this can hardly compete with the always-available Cassandra cluster. So, if you can’t afford any downtimes, Cassandra is your choice.

However, to ensure availability, Cassandra replicates and duplicates data, which leads to data consistency problems. This makes Cassandra a bad choice if your solution depends heavily on data consistency, unlike the strongly consistent HBase. Because the latter writes data only to one place and always knows where to find it (data replication is done ‘externally’ in HDFS).

Besides, Cassandra’s architecture supports both data management and storage, while HBase’s architecture is designed for data management only. By its nature, HBase relies heavily on other technologies, such as HDFS for storage, Apache Zookeeper for server status management and metadata. And again, it needs extra technologies to run queries.

3. Performance

Cassandra’s and HBase’s on-server write paths are very much alike. There’re only slight differences: names for data structures and the fact that, unlike Cassandra, HBase doesn’t write to the log and cache simultaneously (it makes writes slower).

On the higher architectural level, HBase has even more disadvantages:

  1. Before getting to the needed server, the client has to ‘ask’ Zookeeper which server has the hbase:meta table containing info about all tables’ locations in the cluster. Then, the client asks the meta-table-holding server ‘who’ stores the actual table it needs to write to. And only after that the client writes the data to the needed place. If such writes (and also reads) are frequent, this info is of course cached. But if a table region is moved to another server, the client needs to do the full round again. While Cassandra’s data distribution and partitioning based on consistent hashing is much cleverer and quicker than that.
  2. As soon as the in-HBase write path ends (cached data gets flushed to the disk), HDFS also needs time to physically store the data.

Cassandra vs. HBase write

Moreover, the actual measurements of Cassandra’s write performance (in a 32-node cluster, almost 326,500 operations per second versus HBase’s 297,000) also prove that Cassandra is better at writes than HBase.

If you need lots of fast and consistent reads (random access to data and scans), then you can opt for HBase. It writes only on one server, so there is no need to compare different nodes’ data versions. HBase servers also don’t have too many data structures to check before finding your data. You may think that HBase’s read is inefficient since the data is actually stored in HDFS, and HBase needs to get it out of there every time. But HBase has a block cache that has all frequently accessed HDFS data, plus bloom filters with all other data’s approximate ‘addresses,’ which speeds up data retrieval. Essentially, HBase and HDFS’s index system is multi-layered, which is much more efficient than Cassandra’s indexes (check out our article on Cassandra performance to find out more about reads).

If you’ve read that Cassandra is also very good at reads, you may be bewildered by the conclusion that HBase is better. Especially if you saw this benchmarking experience where Cassandra handles 129,000 reads per second against HBase’s just 8,000 (in a 32-node cluster). The thing is, these reads are targeted (based on known primary keys) and, chances are, they are also quite inconsistent. So, Cassandra’s huge numbers fade, if we’re speaking about scans and consistency.

4. Security

Like all NoSQL databases, HBase and Cassandra have their security issues (the main one being that securing data spoils performance making the system heavy and inflexible). But it’s safe to say that both databases have some features to ensure data security: authentication and authorization in both and inter-node + client-to-node encryption in Cassandra. HBase, in its turn, provides the much-needed means for secure communication with other technologies it relies upon.

A bit more detail:

Both Cassandra and HBase provide not just database-wide access control but allow a certain level of granularity. Cassandra enables row-level access and HBase goes as deep as cell-level. Cassandra defines user roles and sets conditions for these roles which later determine whether a user can see particular data or not. While HBase has an inverse ‘move.’ Its administrators assign a visibility label to data sets and then ‘tell’ users and user groups what labels they can see.

5. Application areas

Judging by how Cassandra and HBase organize their data models, they are both really good with time-series data: sensor readings in IoT systems, website visits and customer behavior, stock exchange data, etc. They both store and read such values nicely. Besides that, scalability is the property they both have: Cassandra – linear, HBase – linear and modular ones.

However, when it comes to scanning huge volumes of data to find a small number of results, due to having no data duplications, HBase is better. For instance, this reason applies to HBase’s ability to handle text analysis (based on web pages, social network posts, dictionaries and so on). Plus, HBase can do well with data management platforms and basic data analysis (counting, summing and such; due to its coprocessors in Java).

Cassandra is good for huge volumes of data ingestion, since it’s an efficient write-oriented database. With it, you’ll build a reliable and available data store. In addition, Cassandra enables you to create data centers in different countries and keep them running in sync. Besides, if you couple Cassandra with Spark, you can also achieve good scan performance.

But the main difference between applying Cassandra and HBase in real projects is this. Cassandra is good for ‘always-on’ web or mobile apps and projects with complex and/or real-time analytics. But if there’s no rush for analysis results (for instance, doing data lake experiments or creating machine learning models), HBase may be a good choice. Especially if you’ve already invested in Hadoop infrastructure and skill set.

Cassandra vs. HBase – a recap

Cassandra is a ‘self-sufficient’ technology for data storage and management, while HBase is not. The latter was intended as a tool for random data input/output for HDFS, which is why all its data is stored there. Besides, HBase uses Zookeeper as a server status manager and the ‘guru’ that knows where all metadata is (to avoid immediate cluster failures, when the metadata-containing master goes down). Consequently, HBase’s complex interdependent system is more difficult to configure, secure and maintain.

Cassandra is good at writes, whereas HBase is good at intensive reads. Cassandra’s weak spot is data consistency, while HBase’s pain is data availability, although both try to mitigate the adverse consequences of these problems. Also, both don’t stand frequent data deletes and updates.

So, Cassandra and HBase are definitely not twins but just two strangers with a similar hairstyle. To choose between the two, you should thoroughly analyze your tasks. And then, try to find a way to strengthen the database’s weak spots without affecting its performance.


Need professional advice on big data and dedicated technologies? Get it from ScienceSoft, big data expertise since 2013. 

How to translate a corporate strategy into KPIs


Imagine a spine-chilling scenario: a corporate strategy that was supposed to bring a company to success, lies on a shelf forgotten by its creators and collects dust. At first glance, this scenario is unlikely, as a strategy is the cornerstone of any business. However, BI consulting practitioners have the opposite opinion: this is exactly what happens if a corporate strategy is not supported with the right KPIs.

A strategy may be brilliant, still its execution is likely to fail if the team lacks understanding of how their daily efforts influence the final target. Besides, if a company does not reflect its strategy in KPIs, it is easy to lose focus and step aside, especially at the very beginning when the progress is not obvious. Without right KPIs, a company may lack focus in its actions and consistency in its messages and activities.

Corporate strategy in KPIs

Defining KPIs metrics: examples of doing it right and wrong

Let’s take a look at a great real-life example. In 2015, Walmart was brave to declare their intentions publicly through posting their 3-year growth plan on their website. Together with a simple and straightforward goal of turning omnichannel and delivering a seamless shopping experience at scale, Walmart indicated the target of $45-60B new sales to measure the success. Besides, Walmart indicated 5 growth areas and listed relevant KPIs for each.

Thanks to KPIs, Walmart ensured that everybody spoke the same language, understood the priorities, and knew how to measure the progress. For example, Delivering value goes with price leadership and private brands KPIs; the strategic objective of Providing convenience aims at e-commerce, online grocery and smaller formats; key geographies are narrowed down to the North America and China.

Unlike Walmart, companies usually leave their targets for internal use only. However, it does not mean that they fail to set the right KPIs. For instance, Procter and Gamble also published their strategic objectives, albeit with no targets indicated on their website. At the same time, it is clear that the company did a great job of defining KPIs, such as operating total shareholder return for the Value creation strategy, the number of product categories under focus for Portfolio transformation, etc.

However, even large and well-known companies can make mistakes. Wells Fargo’s notorious case is an example. The bank selected cross-selling as one of their strategic initiatives and developed a relevant motivation scheme with a catchy slogan Eight is great. The idea was to incentivize the employees to reach a target of 8 products sold per customer. Years later, it turned out that Wells Fargo chose a wrong KPI. The KPI motivated to boost sales, even if it did not increase the revenue.  Besides, it was widely unrealistic, and employees opened 2 million fake accounts to reach their targets.

How business intelligence helps

Let’s find out how business intelligence can help while defining and executing a company’s strategy.

BI for defining a strategy

In a market economy with numerous players, companies tend to choose strategies that will strengthen their competitive advantages. Naturally, the first step is to identify these. At such an important stage, what matters is a fact-based opinion, not a gut feeling. To get much-needed insights, companies should analyze both internal and external data. Market research findings coupled with historical data reveal trends and opportunities to seize. For example, a FMCG manufacturer would be able to choose the markets to expand, customers’ behavior at those markets, the best-fitting product portfolio to offer (say, organic food).

BI for setting KPIs

Once a strategy is approved, the next stage is to define KPIs that will be challenging yet attainable. To do this, business analysts go for historical data analysis and forecasting. Besides, business analysts should develop a hierarchy of non-conflicting KPIs: company-based, departmental and individual. With a hierarchy of KPIs, everybody will focus on their piece of work, but all team members will be working towards a common goal. For instance, the NASA brilliantly put this principle into practice: when John F. Kennedy asked the NASA’s janitor about his job, the latter replied, ‘I am helping to put a man to the moon.’

BI for progress monitoring

When KPIs are defined and communicated to the team, performance monitoring starts. Comparing targets vs. facts and tracking the progress is essential to understand how the company advanced in executing its strategy. To make this monitoring possible, business analysts should develop a set of reports and dashboards, identify how often end users need each report, and what data (input and output) should be there. You are welcome to check our BI demo to see how such dashboards may look like.

To sum it up

Any company should support its corporate strategy with the right KPIs and constantly track the progress. Otherwise, a strategy execution is likely to fail. At all stages of strategic management, business intelligence consulting can bring the synergy. For instance, a company can benefit from data analysis services that are helpful to define a corporate strategy, set the right KPIs and monitor the progress.


We offer BI consulting services to answer your business questions and make your analytics insightful, reliable and timely.