In 2013 IBS, which back then seemed to be creating Data Division, asked me to do a brain dump (exclusively based on my experience with corporate oil and gas clients) regarding the problem area of Big Data and Data in general. I stumbled upon it after 7 years and found it amusing. Some things are obvious. Some turned out to be not entirely accurate, but... 7 years have passed.
I wrote it in English and thought of translating it into Russian. Maybe something is still relevant today? (I will translate the bullet points, but I'll leave the tables in English out of laziness. Green means good, red means dangerous, blue means a dream).
I will add minimal comments from 'today' italic, to make it understandable and distinguishable.
So, DATA! We need data...
The Data Division is like the Blood Division because data can be compared to blood flowing through the veins and arteries of a business organism. However, while the blood is the same, the organisms are different, and thus productization is very challenging, but it also represents an opportunity for development.
There are people who notice the data right in front of them - these are We.
And there are people who, unfortunately, do not see the data at all. This, again, unfortunately, includes our Clients!

So, business axioms...
- We sell to businesses, not IT (may all IT specialists forgive me right away) as we solve global issues, and also aim for more money.
- All business problems concentrate around thematic industry verticals and require adequate specialization.
- Attempts to prove the value of 'data' or, even more difficult, the value of 'data management' to businesses – this is eternal suffering and pain. Essentially, it's like going to someone who feels pretty well and saying: 'Dude, we are going to treat your blood, and, dude, it's going to be expensive!'
- My 'wet dream' is to sell 'data extraction' and 'analytics' within a SaaS model to small and medium businesses, which have dived into 123 cloud services with cool interfaces: project management, helpdesk, accounting, CRM, payroll, time reporting, marketing, ... you name it, and have buried themselves in data. Youcalc and Successfactors (there probably aren’t such anymore) are good!
- Look for people who love to tinker with crunch data. They are rare and strange (like coffee grounds readers), but key to business. A poet, for example, may be quite good at understanding correlation.
- Engineers are needed! They are necessary to turn the problems that Crunchers extract from data into solutions. The success or failure of the solution entirely depends on them.
- Development opensource projects represent enormous value and allow for the 'construction' of complex solutions practically 'from scratch'.
- But… one must not forget that Hadoop is a library, and Lucene is also a library, and the gap between a library and an industrial product is significant!
- The solutions built will need to be substantially adapted, therefore modularity and integrability are key aspects.
- Agile (God help us) is a key technique in interaction with the client and hypothesis testing , of which there will be many.Outsourcing any coding and UI is especially possible and necessary. All business analytics and specifications
- of the backend should be retained and regarded as a core competency. inside Decision-makers in business must be constantly 'informed' about
- the necessity of working correctly with data and the ongoing search for new ways to analyze it. The combination of technical and business competencies of our employees will help elevate the status of the entire organization as a whole. — is an endless source of inspiration (
- Internet there weren't so many kittens back then) regarding approaches to corporate data management despite the fact that tasks and scales differ significantly.Technological postulates…

There is enormous developmental potential in
- simplifying how data is presented to people. One could call it 'iPhonization'. Despite BI vendors claiming they directly
- bring analytics to end-users , (and they are certainly moving in this direction) – a breakthrough has still not occurred. People simply do not understandmultidimensional data well. A user interface presenting more or less complex loosely structured data in
- a faceted format – also presents an endless number of problems. Conclusion: the flatter it is – the better. A platform built on automatic data extraction from sources (which are not always intended for such extraction) is heavily dependent on the sources, the stability of connectors, and infrastructure. Any failure to deliver results will always be blamed on the platform (the messenger).
- The platform built on automated data extraction from sources (which are not always intended for such extraction) is significantly dependent on these sources, the stability of the connectors, and the infrastructure. Any failure to deliver results will always lead to blame being placed on the platform (the messenger). Trust – capital of this kind of platforms. Capital that is difficult to earn and easy to lose.
- From a business perspective, there's no difference between analyzing Big Data and Just Data. Often behind simple calculations like 2x2 lies the potential for millions of dollars. A good example is the data on the lifespan of infrastructure elements on the Norwegian shelf. When all the future maintenance dates for all equipment were plotted on a single timeline and it was discovered that in N years a veritable shelf Armageddon was approaching — one very wealthy individual got up from his seat, politely excused himself, and left the room saying, "Excuse me, I don’t have much time, I need to prepare a fleet..."
- Excel, and essentially a clear and concise tabular representation of data has enormous power and a great future. I believe in beautiful tables (and still do) and that's it!
- The main bow of all this "analytics" is decision-making automation. That's where the juiciest opportunities lie, but also the highest risks, because big opportunities come with big risks, hence the opportunities and the risks... 🙂 Managing well drilling, for example...
- If "integrability" is the key feature, then the data should de facto be represented as a service. REST It's great, but one must not forget about optimization performance., which is often sacrificed for integrability, as computational power continues to grow.
- Master data is what needs to be localized, extracted, standardized, before addressing any business issues. Master data is small, but the problems with it are big! As the semantics brothers say – 50% of all the world's problems stem from people calling the same things by different names, and the other 50% come from them calling different things the same name.
- Any encapsulation at the storage level limits the openness of the solution and leads to SILO-ification. It's fine if you're a large vendor, otherwise it’s not so great. (Here, we are certainly not talking about the block level or AWS S3, which had already been around for 6 years back then, but about files).
- Relational modeling is no longer our friend. RDF and key-value are cool! We have seen magical transformations of relational databases with models from 2000 tables to 15 tables, and none of the users lost anything.
- The Internet works because there is URL as the unified means of addressing. The importance of URLs, or rather URI for an enterprise's informational resources is difficult to overestimate.
- Text mining and NLP are popular. On the Internet. But even in the corporate sector, significant success can be achieved by extracting structured data from unstructured corporate data.
- Synergy between structured data and information extracted from unstructured data, i.e. files – an analytical goldmine.
- When extracting data, don't forget about rights and copyrights..
- A data extraction company should establish adepartment of hackers, in the good sense of the word. Inspired by the tough battle against the Yellow Pages' security systems from search bots.
- Before working with data, it is necessary to "see" it in its entirety. This is hard to explain. I think of table forms. Some might think of graphical representations, but any graph is already an interpretation. One way or another… "see"!
- Reiterating the issue of user "trust" in the frontend. Trust in the connectors/data generation processes, trust in the data, trust in the decisions made..
Source: habr.com
