Saturday, September 7, 2013

Gap Maps - Great Way to Measure Product Operations Maturity

Often, people connected with a product wonder where their product is in comparison to other similar products or competitors, so to speak. The management gurus figured out a way to bridge that by creating “Gap Maps”.

The great thing about gap maps is that they can be used for comparing any entity. The entity can be products, sports teams, presidential candidates, tools, countries, choices…..anything at all.

The way you do it is to find two determinant attributes – two very distinguishable and principal attributes that define the set of entities. To explain this further, I take my favorite example – if I want to buy a new car, I could have mileage per gallon and MTBR as two determinant attributes. I could also have “smell of new car” as one of the determinant attribute. Though it is very endearing thing, the new car’s smell is certainly not very good attribute to go by, especially when you are putting tons of money down on a car.

Moral of the story – the gap maps are just a tool. How reliable it would be depends entirely on your choice of attributes. It is classic case of GIGO (Garbage in, garbage out) – if you chose the right attributes, you will get good gap map, bad choices will lead to bad gap map.

Why we need to get a good gap map? Simple reason is that a good map unravels so many stories about your product and similar products (read competition). A bad gap map may sweep under the rug many of the flaws and may lead you to believe that you have a winner product.

Now, this is one part of the story. Hold this strand somewhere in your L2 cache while I do a context switch.
Let me start another thread in this story and see if I can bring multiple the threads together to weave a single story, at some point in this blog.

We had gone live with a product earlier this year. It is one of our Data Systems pipeline that carries data (aka events) from hundreds of thousands of serving hosts back to our own “Deep Blue” backend system which then takes this petabytes of data and makes sense out of it.

Before we went live, we did something called Operational Readiness Certification or ORC. ORC is a long laundry list of hundreds of questions – some require subjective answers, some need numbers and yet others get Boolean type answers.

A good example of question asked in ORC is – “Do you have a BCP” – the answer is pretty Boolean – Yes or No. (Ok, there will always be the *cautious* types that will start their answer with “It depends…..” LOL!)
So this pipeline had passed ORC with flying colors and everyone was happy. However, when we started ramping up the volume, we found three conspicuous challenges on this pipeline:-

  • Backlog Catch up Rate
  • Reprocessing
  • Data Discrepancies

Each of these three was causing us to throw tons of manual cycles at it with no light blinking at the end of tunnel. There was no ready metrics that we could take to our product and engineering partners and tell them where this product was in terms of production readiness and where we would want it to be.

It was a perplexing problem and luckily for us our Dev team is very brilliant who didn’t really need us to put lot of data behind these three issues before they would even pick them up for resolution.

So what was the problem statement? In a very high level, bulleted version, it would look something like this:-

  • ORC is great but somewhat subjective and boolean
  • ORC is also a matrix of blockers, failures, exceptions, action items. Once you are past those, nothing more comes out of ORC.
  • After ORC is done, what is the next step? All products are at the same level.
  • No way to classify/score maturity of a product compared to its peers
  • No method of comparing different properties
  • Comparison of similar systems could help
  • ORC doesn’t allow time series trending

As I said above, we are lucky to have great developers at Yahoo and within a quarter, we almost solved all the three problems. However, lack of a good measurement method or absence of a tool that we could use to evaluate and compare our product with similar and very mature product in that space frustrated us.

Now let me bring in the third string in this story. I was reading a book on product management where the authors discuss gap maps and it dawned on me that if I have something similar, we could use it to compare different properties at Yahoo that we support.

Bringing all different strands together, I decided to look at gap maps, used for decades in the product management industry, to evaluate and compare our different data systems pipelines.  The first and foremost challenge was to figure out correct determinant attributes. The challenge was that there are so many great attributes that could be used to peel this onion. I took a different approach and decided to create two uber attributes with any number of sub-attributes.


Sample attributes and sub-attributes

Performance and Operability were the two overarching attributes and each had many sub-attributes.While selecting attributes and sub-attributes, it should be kept in mind it is not necessary that different categories of properties have similar attributes. For example, a property like Yahoo Frontpage or Facebook main landing page may have different attributes than a backend data warehouse system. Decide your attributes and sub-attributes carefully and diligently. This will be time well invested upfront in the whole exercise.Once the sub-attributes have been chosen, you could use a simple spreadsheet to compute the values for the sub-attributes and aggregate them to arrive at a value for the parent attribute. Please see the figure below. I have also given some guidance for scoring them, but you should create your own guidance.


Scoring spreadsheet sample

The point to keep in mind is that this guidance should not change from property to property (or product to product) in the same category. So if you are comparing different data warehouse systems or massive analytic systems, the guidance should remain the same. However, like I said above, different categories may have different attributes, some similar some totally dissimilar and there guidance for scoring may also be vastly different from the above scoring model.

Once you get the values from this type of spreadsheet for main two determinant attributes, use those values to plot on the simple X,Y Axes graphs.

The names of the product are somewhat fictitious and so is the sample data (I could use the disclaimer “The characters and story in this movie are fictional and any resemblance to people, living or dead, is merely coincidental…” J)


Sample Scores for POM

 Once you have the values for X and Y Axes, plotting of the graphs is fairly simple.


POM Graph

Now that you have the graph which gives where different similar products fall, you may ask the question, now what?

Well, to start with, you (developers, Service engineers, product managers, managers – in short all connected with the product) can understand where your product is compared to its peers.

The first endeavor should be to get your product to the first quadrant (positive quadrant). Once it is in the positive quadrant, the next endeavor should be to continually move it to in the north east direction.

It also gives you an understanding why EDW needs tons of manpower to support it and why CMS data warehouse needs half an FTE (full time employee) to support it.

Further, EDW folks can talk to CMS folks to get a handle on what are the different things CMS team did to get to where they are.

Finally, if EDW team starts working on the betterment of the product, they can use two snapshot of this graph – the first one in the present time and the next one a quarter or two later to evaluate progress (hopefully) that the product has made on the two determinant attributes.

Love to get feedback.

Monday, September 2, 2013

Development, SE and SRE teams - why all of them are critical?

Have you ever wondered what happens when you have a very motivated team that is very wrong for a given job? The team works extremely hard, slogging hours and clocking tons of time, killing themselves over weekends and finally comes up dissatisfied with what they have (non)achieved. In other words, square peg in round hole. Sounds familiar? Read on.

This case study is for service engineering team in any company of 1000 people or more. In the industry, service engineering team typically sets up all the framework to take the developer’s code to production. This involves setting up robust CI/CD (Continuous Integration/Continuous Delivery), monitoring, automating some parts of Dev team provided engineering and QA tests, post-deployment smoke tests and full blown monitoring framework for all alerts around the code.  

SE and SRE Teams
Most organizations have SE and SRE teams in a combined single team. They may chose to call it SE or SRE. But make no mistake, this team has two different and very conspicuous flavors – engineering and operations. Like I said earlier, most teams have both the flavors built into one single team. That implies that the same team will have folks with great engineering competence and those with operations bias as well. However, in some companies, especially the larger ones, the two flavors may be two very different teams working separately under different leadership for common goals as stated above. 

In some places, the engineering focused team is called DevOps and the operations team is called SRE. Yet other companies name the engineering focused team as Service Engineering team and operations team as just operations team.

In Yahoo, we have engineering biased team as Service Engineering team (SE team) and operations focused team as Service Reliability Engineering team (SRE team).

When the teams are created ab initio, the work is segregated and defined for each team. The two teams are also seeded differently – engineering focused team will have more people who can code, understand the code, get into innards of the code base and file bugs when we hit issues due to buggy code. They almost tell the developers “here is where your code is throwing an exception, please fix this part”. So they “can read” and “understand” the code but since they do not “own” the code, they hand off the bug resolution to developers.

Then we have the SRE team which is our first and second line of defense and mostly attends to all the alerts, provides first level investigation and triaging and either resolves them or escalates them to service engineering team. In a very matured SRE team, we expect 85-90% alerts being handled and resolved by SRE team. The 15-10% alerts that are escalated to SE team are mostly resolved by SEs. You can expect 1-2% of those alerts being escalated to development team.

If you look at the work, SRE work is totally interrupt driven, SE work is partially interrupt driven and large part is planned work. The developer’s work, on the other hand, should be largely planned so he or she can totally focus on the new features, new products and enhancements etc.

It is possible that over a period, the teams may morph into something better – at that point, we say “oh, this team has really matured into a fantastic SRE or SE or development team”. That is mostly possible if the teams are doing the type of work that they have been designed to do and in the manner (interrupt or plan driven) they have been conceptualized to do. So this is good scenario, eh?

When things start going awry….
What happens when the scenario doesn’t turn out to be as good as we wanted it to be or as favorable to each team as we would have wished for? Well, then we have a challenge….

If we are not continuously monitoring our teams for type of skill sets that seed these teams, the type of work that falls in their laps, it is very possible that the fiber of the team(s) may undergo a mutation – typically, for worse.

Imagine a scenario, we have attrition due to any number of reasons in SRE team. It could be leadership or lack of it, lack of good management, gaps in people’s expectation, company not doing well, work load being pure killer…any number of reasons. And as a organization, we fail to see this coming, even after it happens, we do not backfill attrition immediately or fast enough, the workload on remaining people will continue to increase since attrition of people doesn’t necessarily translate into reduction of workload. So, the remaining workforce comes under resources crunch and work overload. This then starts a downward spiral that if not arrested well and fast enough can pretty much cause annihilation of SRE team. Once SRE team reduces without proportional reduction in the workload, we start spilling workload to SE team. SE team now suddenly discovers, much to its angst and disappointment that it is the de-facto team doing SRE work while expectation around SE work have not diminished at all. So SE team starts focusing on totally operational work and bends backward to make the Site stay up.

During this time, SE team has also changed their work routine as follows so they can accommodate the operations workload that has been thrust on them:-
  •           Stops going to development team’s daily scrum meetings since SE was up fighting operations and incident late last night, during weekends and long weekends
  •           Doesn’t have time to do code review with dev team
  •           Doesn’t have time to build monitoring new feature that got pushed last evening through CI/CD pipeline
  •           Backlog of the SE related work starts building up
Not many people realize this but during this time, the SE team also moved away from being largely driven by planned work through Kanban/Scrum to interrupt driven work.

Slowly but steadily the SE team becomes the new SRE team. Management and leadership don’t really mind it – they cut the cost down in operating expenses by dismantling a full team that was called SRE team and in their minds and words, they have made the SE team very “efficient”.

The management is so focused on dollars that it misses the deep, dense forest for few “shinning” trees.

The whole change takes place over at least a year, it can’t happen over few months.

The management is applauded and rewarded for cost cutting and the “success stories” are told and exchanged with other teams. No one yet understands the deeper damage this cost cutting and change has done to the SE team. By the time people and management will realize this, the current leadership of management would have long moved to different role, different company, and different team to continue the good work there.

Now let us bring our focus back to poor SE team that has been forced to morph into an SRE team. There are some very brilliant engineers in the SE team who are not happy with the current state of affairs and are waiting for management to put them out of their misery by recreating the SRE team. When they realize that management is not even thinking on those lines – of recreating SRE team, reinvesting in the SRE people - they make their mind and jump ship. Jump ship to different team, different company, anywhere but their current team.

Now the next phase kicks in, the SE team starts seeing the same attrition that SRE team saw – reasons may be the same or different – but good and brilliant team members are first to leave. The immediate management of SE team is suddenly running around like headless chicken, trying to do firefighting by getting enough resources so Site can be kept up. In their rush to get any and every resource they can find, they naturally lower the hiring bar.

Once they lower the hiring bar, the once brilliant and industry recognized SE team starts hiring sub-standard material for that team. Please be aware that the new hires are not bad or incompetent people. When I say "sub-standard" it doesn't mean that they are incompetent. All I am saying is that they are not the right fit for the SE team. They may be pretty hard working but are ill-suited for the SE team which was until then seeded by brilliant engineers who fully understood how to take very "developer" code and change it into a very "production ready" code.

Now, this is where it gets interesting. In the first phase, we lost SRE team. In the second phase, we forced the SE team to become SRE and SE team combined into one. In the third phase, we made the SE team to change totally into SRE team. In the fourth phase, we started losing SE team. In the fifth and final phase, we hired and seeded SE team with different level of people.

…here is the kicker, the SE team that started operating like an operations team is at its lowest morale since the workload is totally interrupt driven, they have by now lost the respect of their development team. At this point, development team pretty much agrees that SE team doesn't do any level of investigation on any issue before chucking it over the fence to them.

…And the disease spreads to Development team as well….
So now, this is perhaps that last stage, the development team changes its workload from being totally plan driven workload to substantially interrupt driven workload.  At this stage, a development team that was entirely focused on new products, new features is now forced to change its focus to part new features/product and part sustenance of existing products or site up. A development team that was able to bring to market at least two big products a year is now struggling to bring even one big product in beta phase.

The slow development team frustrates their management since their management is trying to catch up or stay ahead of competition. As a result, they are continuously changing roadmap and plan of record.

In the past, the development team used to complete a big product in 4-6 months, which used to allow their management to do course corrections rapidly. Now, the course corrections have to happen at the same pace as earlier but imagine that a development which is moving a very slow pace, same team is now forced to absorbed course corrections while their beta version is also not put out. This results in what development team largely sees as "scope creep" on the given project. This frustrates the developers and they understand this as “directionless” development team management.

Now the maladies and issues of SRE and SE team have become contagious and started hitting development team as well. Developers, frustrated by their management’s constant change of direction, start leaving causing a drain even bigger than the SRE and SE team attrition.

At this point the whole organization is paralyzed by these issues and starts slowing down to the extent, that at some point, it comes to a grinding halt. At this juncture, all the teams are focused on keeping the site up.

How do we rebuild from here?
First thing we need to do is to baseline our team to estimate the "damage" - to ascertain how much our team's bias has changed.

To measure where an SE or SRE or Development team stands, we can plot all the dimensions required for a good team - SRE, SE (with engineering bias) or Dev team on X axis and then measure them on a scale of 0 through 10. Zero on the scale shows total absence of the specific dimension and 10 shows reasonably high proficiency in that dimension.

For the purpose of this article, I will focus only on Service Engineering team.

Let us also assume that hypothetically a very mature Service Engineering team in the industry will perhaps be at 7.5 and above.

There is a distinct difference between Service Engineering team as well as Service Reliability Engineering (SRE) team. Both are equally important for a company. At the risk of being shouted down by many, I would say that for a company, SRE team is more important than SE team. I draw the metaphor of a hospital. Every hospital has an Emergency Room or ER. Some places also call it Trauma center. This place receives people who are in dire need of immediate first aid else they would die. They receive all the accidents, heart attack, gun violence related people… They are America’s life line and the first line of response. Without these wonderful folks we would have hundreds of thousands of additional causalities in US every year.

Then there is a medical system that does more of diagnostic and preventive medication. These are also the people for whom every day is a “Monday” – they don’t have holidays,

These are the guys who save us daily.

Similarly, SRE team is our first line of defense. These are the guys who receive the alerts and respond to them immediately. Depending on how egregious an alert is, and how critical a property is, their response may vary. For example, in case of a DoS attack, they may start and IRC, collaborate with other companies having the same challenge and do everything to repel the attack. Actually, discussion on DoS will need a separate book by itself. LOL!

SE team is the medical system that has much more time than the ER or SRE team and therefore, can focus on addressing the causes as opposed to just the symptoms.
Engineering-Operations Graphs or EO-factor

I call these graph Engineering-Operations graph or “EO factor” for short. This is like pH factor that determines if a solution is acidic or basic. The pH factor of pure water is 7 and is considered neutral. pH factor above 7 is considered alkaline or basic and as pH factor increases beyond 7, the basicity or alkalinity of the solution increases. Similarly, as it goes below 7, the acidity of solution increases.

EO factor should be read in the same way, as we move away from the “Operations Line” the team becomes more biased and focused on a given track – operations or engineering.  As you move north of this line, the team tends to become very engineering focused. In the same vein, if you move south of this line, the team has more of operations competence.

This graph can be used by any team with different dimensions. For example, an SE or SRE team working on Big Data systems will have somewhat different dimensions than a team working on Mail or a team working on company landing home page. It is entirely up to the managers to figure out correct dimensions to measure their teams and use those dimensions to then chart out the future course. 

EO Factor
Build the new team...
Once you have ascertained the baseline, take remedial steps to rebuild it - slowly and steadily.

Seed it with right skill set, keep the right workload type, see to it that it comes in right manner - interrupt or planned, keep irrigating and feeding it with right talent. And above all, watch it carefully as shown below.

Watch the team's focus...

Let us take a scenario where I have EO factor for a team and the team exists for a long period. You are interested in seeing how that team either stays in the fiber that you created it for or if, over a period, it has changed its bias. A time series trending will be great graph to have for this type of trending. 

EO Factor Trending
This time series helps us to understand how our team is trending over a period. Remember, no side is good or bad in this graph. You want to keep a good watch on your team’s tendency to shift its bias from Operations to Engineering or vice versa. Depending on what bias the team was intended to be seeded with, you have very compelling motivation to keep the bias in the same side. If you are not watchful and very mindful of this, the teams always have tendencies to move from one focus to the other.

We should ensure that we seed an SE team so that it has an EO factor of about 6 or above. We have also got to ensure that this EO factor always stays above that for that SE team. Similarly, we need to ensure that our SRE team has EO factor below 5 or 5.5 so it keeps operations as their bias.

An SRE team can be seeded with very brilliant engineering focused people as well. The challenge is that most likely the brilliant people do not like operations work and may again start rolling the same juggernaut which led to the attrition of SRE in the first place.

Conclusion
We need to be always aware and cognizant of type of talent pool in every team and make sure that right team is staffed for doing the right and intended job. This has to be consistently and continually monitored else we get to a point where we need to make radical changes. Making changes and having those changes make impact is multi-year project and therefore, painfully slow.  Hence, a stitch in time may act as a preventer for nine few years later.

Sunday, June 16, 2013

Data Service Engineers - the new hybrid!

Until recently, at least 3-4 years ago, there were two distinct skill sets in IT market – Database Administrators (DBAs) and Service Engineers (SEs). In some companies, Service Engineers are also referred to as Service Reliability Engineers (SREs). Service Engineers are essential operations folks that have many duties including pushing code to production, bringing up services, deploying new hosts, triaging incidents, monitoring alerts dashboards etc – in short the whole 9-yards for operations.

Every company that has online operations and IT staff maintaining those online operations needs to have SEs or SREs. For the purpose of this blog, I am going to restrict to using the single term SE as opposed to use both SE and SRE.

DBAs in an IT shop in any company are responsible for managing databases. Typically these databases are relational databases like – Oracle, MySQL, MS SQL Server or DB2
.
Both the set of folks – DBAs as well as SEs – work very cohesively to form a compact team. Their duties are also largely segregated. SEs maintain the applications and front end of any application while DBAs tend to the backend databases.

However, in past few years or so (and this could be different in different companies – some are further along in this journey and maturity) – we are seeing a blurring of the lines between the two skill sets. The blurring is largely caused by the advent of all databases that are non-relational data stores. Examples of such data stores are NoSQL (Key Value Stores) – Cassandra, Dynamo DB, MongoDB; Columnar databases like SenSage, Sybase IQ, C-Store, EXASOL, MonetDB, Vertica;  flat file based distributed systems like Hadoop.

The wholesale adoption of these data stores by applications have somewhat destroyed the silos that DBAs had to maintain to keep the data integral. For example, to protect a relational database, the DBAs needed to always maintain the access controls etc very rigid. The more “black box” a relational database was, the more rigid and higher were the silo walls. For example, there are MySQL databases administered by SEs whereas Oracle or DB2 databases or for that matter, even MS SQL Server and Sybase databases are always managed by DBAs.

Since non-relational data stores (and to certain extent even in case of MySQL) are not as rigid in structures as Oracle, there was no deep knowledge of internals of the data store needed to manage these databases. So essentially, the small and nimble IT teams started looking at saving money on dedicated DBA hires and started resorting to SEs managing not only the Applications and Frontend part of the system but also the backend systems comprising of NoSQL and/or MySQL.

This has led to the mushrooming of a new breed of SEs – I refer to them often as Data SEs or DSEs. These are going to be the new breed of operations folks who will manage the Big Data stack in companies. The DSEs will not only have deep knowledge of all front end tools and technologies like apache, tomcat, Perl, PHP, Python, Github, HTML5, CSS, Jenkin and myriad other languages/tools but also be very skilled on managing and working through the technologies like Hadoop, HBase, Hive, STORM, Spark, Shark, Cassandra, MongoDB, MySQL.


As more and more applications move to adopt MySQL/NoSQL data stores in favor of hardcore and very expensive relational stores like Oracle, the lines between DBAs and SEs will vaporize and emergence of Data SEs become an inevitable reality.

Sunday, February 17, 2013

Is Oracle Dying......


Oracle Corporation was founded in Jun 1977 by Larry Ellison, Bob Miner, Ed Oates. Over the years, it has risen to become almost indisputable leader of the Relational Database Management System (RDBMS) market with 44% (Source: IDC 2009) – at least, for now, though, no one is sure how long that numero uno position will last. There were heady days of 1996-2008 or so when Oracle ruled the world of RDBMS. It was unchallenged crown king that could do no wrong. Hundreds of thousands of Database engineers, architects, administrators spoke of Oracle as if it was actually the famed “Oracle of Delphi”. Conference passes to Oracle Open World were so coveted that it was distributed to star employees in any company using Oracle Products. 
However, after 2008, the downward spiral was very perceptible to the database communities. The hush hush talks could now be heard loud and clear. Only that Oracle was perhaps hearing it but not listening. It continued to maintain the arrogance of a star past its prime - denying that it was aging, claiming that the talent would always trump the age.
I think the Oracle Goliath had forgotten that for every arrogant Goliath, there is a David that is bound to introduce it to its nemesis.  But my guess is that this downward spiral perhaps set into motion long before 2008 or so when rest of world started noticing it or at least it became very perceptible.
Time machine
Let us trace Oracle's journey through its very meager beginnings and how it lost its course along the way. The chronological sequence of this journey could be roughly as I have shown below:-
1977  SDL (Oracle's predecessor) founded
1978   Oracle Version 1 developed 
1979   First commercial SQL RDBMS
1983   Oracle Version 3, built on the C programming language, is the first RDBMS to run on mainframes, minicomputers, and PCs, VMS Based database
1984   first RDBMS to offer read-consistency
1985   Released of Oracle Version 5, one of the first relational database systems to operate in client/server environments
1986   Oracle goes public on the NASDAQ exchange
1987   Becomes world’s largest database company, Oracle gets into building enterprise applications, introduces UNIX-based Oracle applications
1988   Oracle Version 6 debuts with several major advances: Row-level locking, Hot backup, PL/SQL
1989   Oracle provides DB support online transaction processing (OLTP) and moves Oracle moves its headquarters Redwood Shores, California, campus.
1990   Launches Oracle Applications Release 8
1992   Launched Oracle 7, offers full applications implementation methodology
1993   Client/server environments enhancements
1994   Oracle earns the industry’s first independent security evaluations
1995   Offers the first 64-bit RDBMS
1996   Releases feature rich 7.3, Oracle to manage any type of data—text, video, maps, sound, or images, moves towards an open standards-based, web-enabled architecture
1998   With Oracle8 Database and Oracle Applications 10.7, Oracle is the first enterprise computing company to embrace the Java
1999   Offers its first DBMS with XML support
2000   Oracle ships Oracle E-Business Suite Release 11i
2001   Oracle9i Database adds Oracle Real Application Clusters,  becomes the first to complete 3 terabyte TPC-H world record
2002  Offers the first database to pass 15 industry standard security evaluations
2003  Oracle debuts Oracle Database 10g, more robust clustering software
2004  Declares Oracle “the Information Company” and spreads into many other areas
2005  Oracle completes the acquisition of applications rival PeopleSoft , releases its first free database, Oracle Database 10g Express Edition (XE)
2006   Declares a 30-year commitment to open standards computing with Unbreakable Linux—giving customers
2008   HP Oracle Database Machine/Exadata storage
2009   Gets into too many things - including BEA products, launch of Oracle Fusion Middleware, 11g advance Oracle
2010   Oracle acquires Sun Microsystems, announces Sun based Exadata/Exalogic machines
2011   Keeps adding bells and whistles to same Exadata/Exalogic machines
2012   Announces initiative focused on Cloud

Rise of Oracle
Most of the engineers in software industry were not even born when in late seventies, it struck young Larry Ellison, after reading the 1970 paper written by Dr Edgar F. Codd on relational database management systems (RDBMS) named "A Relational Model of Data for Large Shared Data Banks, that a software could be designed that could follow the principles of relational databases.  His belief was reinforced when he read an article, published in the IBM Research Journal, by Ed Oates of IBM about the IBM System R database. System R itself was based on Codd's theories. In 1977, Ellison co-founded Oracle Corporation with Bob Miner and Ed Oates under the name Software Development Laboratories (SDL) and in 1979, SDL was rechristened as Relational Software, Inc. only to change its name again in 1982 to Oracle Systems Corporation. In 1995, Oracle Systems Corporation changed its name to Oracle Corporation. From 1979 through 1992, Oracle primarily focused its attention on making its flagship product, Oracle RDBMS, strong. Oracle was getting complacent after version 5 and then it came out with version 6 – this was huge fiasco product and it was nightmare for customer support and Oracle support. Corporate customers were threatening to pull off Oracle. Version 6 was quickly followed up by version 7 which saved the day for Oracle. 7.3.4 turned out to be very stable product. Version 8i, 9i and 10g added to Oracle RDBMS core competence. These versions by themselves attracted customers to Oracle.  If everything was so good at that point in time and continues to be good then why do I particularly feel that Oracle could be dying as a company?
Lack of Level 5 leadership
Oracle has been led by Larry Ellison all these years - since its inception. Larry is a level 4 leader – wish he was level 5! Under his leadership, Oracle has always focused on “what” should be done and a very little on “how” will it be done. Level 5 CEOs first focus on “who” and then on “what” and “how”.  People part of the equation remains very flaky, to say the least, with Oracle. It has been notoriously uncaring about exodus of top talent. Many ex-Oracle top performers have gone on to form companies, rise to be C staff levels, unleash innovations but Oracle didn’t really do anything specific to stop the fleeing top talent.  
Also, like many other celebrity CEOs, Ellison is getting too distracted by things that he and his company should not be focusing on – example, Oracle’s American Cup sponsorship, Ellison’s many prime properties, Ellison’s unflinching support for former ousted HP CEO and great friend Mark Hurd, Ellison’s purchasing Lanai Island and somewhat ill advised potshots at erst-while partner turned arch rival, HP. All these have direct impact on Oracle’s future – why? Because all these are issues that distract the CEO. Similar distraction proved a debacle for Lee Iacocca – once he turned around Chrysler, he focused more on politics, image building, helping White House with many initiatives which distracted him from his duties as a CEO. And Chrysler slid back into the mess that it had barely recovered from. Mark’s hiring into Oracle forced Ellison to send Charles Phillips off. Charles was a great executive and leader recognized for his talent in and outside Oracle. Letting a great leader go in favor of a friend whose moral ethics are somewhat doubtful can never go well with the employees.
Also, Oracle doesn't have conversations like “what can we do to stop you from leaving” with most of their top talent attritions.
5 Phases of a perilous corporation
Any company going through the general growth, if not managed in a disciplined manner, can hurtle itself into peril. Jim Collins brings this out very succinctly in his book “How the Mighty Fall – any why some companies never give in”. The 5 stages of this journey from greatness to perish are very perceptible when they happen.
The Path to Destruction
So if it is not 2008, when do I think Oracle started slipping? I suspect Oracle’s downward spiral started after 2001-2002 (or at least sometime during that period). It could not come to terms with the ever high stock price of more than $45 and started developing unreasonable greed.


Perhaps under some implicit or explicit mandate from Uncle Larry, the sales people were sent marching to see how much more they could milk out of their unsuspecting and totally Oracle dependent customers. And perhaps the sales people came back with the message that customers would not mind paying more for the crown jewel product – core RDBMS as well as Oracle ERP Suite – 11i. Oracle (read, Larry Ellison) could not stand competition – especially those then started looking at how to kill rivals – hostile and non-hostile acquisitions of rival JD Edwards, PeopleSoft and Siebel.
Every growing company reaches a point where growth starts flattening – happened with Apple, happened with Google and will happen with next big shinning company as well – Oracle was not particularly immune to it so in an attempt to offset the flattening growth of its flagship core database product, Oracle started developing another front that it could open - this was business of  application servers - an exploding market back in the day.

An application server is software that helps developers write and deploy specific applications. The market has exploded past decade or so since many application server vendors are trying to build dynamic applications for the mobile devices. The market could perhaps be as lucrative as the core database market.

Oracle was very late entrant into this market but it quickly acquired BEA Software (leader in the space) and started competing neck to neck with IBM WebSphere. Within Oracle, Application Server business is viewed as “third business” besides core RDBMS and ERP.
Oracle built its business by dominating the database market, providing the central repositories of crucial information that businesses must maintain and use to complete transactions. This has given it an unrivaled position of power when dealing with customers. Capitalizing on such an edge, Oracle’s sales representatives have earned a fearsome reputation as hard-line negotiators determined to squeeze customers – and they are squeezing where it hurts the customers most – at their licensing and support costs.
However, like it had opened a third front by getting into Application Servers market, it has since then opened many more such fronts via its acquisition spree - Oracle moved well beyond the database and into business software, buying up the important products that companies use to keep track of their technology infrastructure, employees, sales, inventory and customers.
Undisciplined growth
In their pursuit to keep up with their YoY growth, Oracle has descended into a very undisciplined growth. There was also very unreasonable desire to grow into every domain. While growing via acquisitions, Oracle Executive Management has forgotten that it is not simply enough to acquire good companies, it takes good and dedicated diligence to grow them into great companies. Some of the companies Oracle acquired are as under:-
2013                      
Feb-13  Acme Packet  Networking hardware for telecom service providers
2012                      
Dec-12  Eloqua   Marketing Automation platform for managing sales/leads
Dec-12  DataRaker    Cloud based Analytic platform
Nov-12 Instantis   Cloud and premises-based Project Portfolio Management 
Sep-12  SelectMinds   Cloud-based social talent sourcing 
Jul-12    Xsigo Systems   Provider of network virtualization technology 
Jul-12    Skire  Solutions provider for managing capital projects/facilities
Jul-12    Involver   Social media development platform
Jun-12   Collective Intellect  Cloud-based social intelligence solutions
May-12 Vitrue   Social Marketing Platform provider
Mar-12 ClearTrial Cloud-based Clinical Trial Operations and Analytics products
Feb-12  Taleo   Talent Management Software
2011                      
Oct-11   RightNow Technologies    Cloud-based CRM
Oct-11   Endeca E-commerce & Business Intelligence
Sep-11  GoAHead    Service Availability and Management Software
Jul-11    InQuira Service Knowledge Management Software
Jul-11    Ksplice Rebootless Linux kernel updates
Jun-11   FatWire Software  Web Content and Web Experience Management 
Jun-11   Pillar Data Systems         Storage systems
Apr-11   Datanomic          Data Quality Software
Feb-11   Ndevr - Select IP only/Environmental Reporting/BI
2010                      
Nov-10 Art Technology Group   Ecommerce software vendor
May-10 Pre-Paid Software    Payment Solutions
May-10 Market2Lead  Applications
May-10 Secerno  Data protection hardware and software
Apr-10  Phase Forward Applications for life sciences cos/healthcare providers
Feb-10  AmberPoint  Service-Oriented Architecture (SOA) management
Feb-10  Convergin   Telecom Service Broker
Jan-10   Sun Microsystems Computer servers, storage, networks, Java, MySQL database, software, and services
Jan-10   Silver Creek Systems  Product Data Quality Solutions
2009                      
Oct-09   SOPHOI  IP management for Media & Entertainment Industry
Sep-09  HyperRoll   Financials, software and IT services
Jun-09   Conformia  Product Lifecycle Management
May-09 Virtual Iron Software Server Virtualization Management Software
Mar-09 Relsys International Drug Safety and Risk Management
2008                      
Oct-08   Haley (RuleBurst Holdings)  Natural Language Business Rules / Policy Automation
Oct-08   Advanced Visual Technology  Retail Space Planning
Oct-08   Primavera  Project Portfolio Management
Jun-08   Skywire Software    Document Management
May-08 AdminServer  Insurance Policy Administration
Jan-08   BEA Systems   Enterprise Software
2007                      
Dec-07  Moniforce   Real User Experience Monitoring
Sep-07  Bridgestream    Enterprise Role Management software
Jul-07    Bharosa, Inc  Online Identity Theft and Fraud Detection
May-07 Agile Software Corporation   Product Lifecycle Management
Apr-07  Lodestar Corporation  Utilities Application Software
Mar-07 Hyperion Corporation   Enterprise Performance Management
Mar-07  Tangosol Inc   Datagrid Software
2006                      
Nov-06 Stellent Inc. Universal Content Management, Digital Rights Management
Nov-06 SPL WorldGroup  Utility Billing and Customer Service Systems
Oct-06   Sunopsis   ETL, Data Integration
Oct-06   MetaSolv OSS service activation
Jun-06   Demantra  Demand-Driven Planning Solution
Jun-06   Telephony@Work   Leading IP-based Contact Center Solution
Apr-06  Portal Software  Billing/Revenue Management solutions
Feb-06  HotSip   Communications infrastructure solutions
Feb-06  Sleepycat Software  Open-source db software for embedded applications
Jan-06   360Commerce   Retail Industry Solutions
Jan-06   Siebel Systems  Customer relationship management
2005                      
Dec-05   Temposoft    Workforce Management Applications organization
Nov-05 OctetString   Virtual Directory Solutions
Nov-05 Thor Technologies     Enterprise-wide User Provisioning Solutions.
Oct-05   Innobase  Discrete Transactional Open Source Database Technology
Sep-05  G-Log    Transportation Management Solutions
Aug-05  i-flex      Banking Industry Solutions
Jul-05    Context Media  Enterprise Content Integration
Jul-05    ProfitLogic           Retail Industry Solutions
Jun-05   TimesTen      Real-time Enterprise Solutions
Jun-05   TripleHop   Context-sensitive Enterprise Search
Apr-05  Retek    Retail Industry Solutions
Mar-05 Oblix      Identity Management Solutions
Jan-05   PeopleSoft     Enterprise Software
2004                      
Jun-04   Collaxa  Business process management
May-04 Phaos    Identity management
Jan-04   SiteWorks Solutions   Clinical trials management
2003                      
Jun-03   Reliaty  Enterprise data protection
Jun-03   FileFish Enterprise content management
2002                      
Jun-02   Steltor  Enterprise calendaring system
Jan-02   NetForce   Adverse event reporting system
Jan-02   Indicast      Voice portals
Jan-02   TopLink   Object-relation mapping technology
1999                      
Jun-99   Thinking Machines Corporation datamining technology
1995                      
Aug-95  IRI Software  OLAP products
1994                      
Oct-94   Rdb Division of Digital Equipment Corporation   Relational database

The early acquisitions show Oracle focus on growing its databases market but acquisitions of past few years including very surprising $5 Billion acquisition of Sun MicroSystems do not give us good sense of where Oracle’s focus is. The strategic theme in Oracle’s acquisition spree is missing and seems more like reactions of leadership focusing only on “growth”. Take a look at spread of Oracle into sectors and even a layman would agree that it is stretching itself far too thin.
If people outside of Oracle can’t understand why Oracle acquired Sun Microsystems, the confusion is equally evident inside Oracle as well. No one can put a figure on if Oracle acquired Sun for hardware market entry point or MySQL or Sun Solaris OS or was it a combination of all these and then some.
Oracle has come out with an integrated ERP product suite – Fusion. The sales teams do not fully comprehend how to package Fusion compared to Oracle 12 version. As such Fusion itself is at least four years too late. In its attempt to create a unified platform for ERP software, it has managed to successfully scare customer who want just a small focused set of modules – like AR and GL or Manufacturing.

There was Steve Jobs who made the famous statement that “…we tell customers what they want…”. Larry Ellison can make the same claim – but to be successful at doing that, you have to be a visionary and not be distracted so hopelessly as Ellison currently is. And, customers seem to the last thing on most of Oracle’s moves. For example, some of Sun’s largest former customers consist of the large Wall Street players, and they were miffed and pushed back last year when Oracle wanted restrict their choices around the Sun technology. Oracle ultimately gave in to their defiance, reaffirming deals that would let Hewlett-Packard and Dell offer prized Sun software on their hardware.

“Customers will always gripe about giving too much control to any one company,” said Israel Hernandez, director of software research at Barclays Capital.

In 2010, Oracle hired Mark V. Hurd, the former chief executive of H.P., as a co-president. Analysts viewed the hiring as a positive outcome for Oracle as it looks to expand. However, Mr. Hurd’s arrival was quickly followed by departure of one of Oracle long-timer, Charles Phillips.

Oracle customers are worried about putting all their eggs in one basket. Every company that they tend to do business is being bought by Oracle – much to customers’ dislike. And for hosts of Oracle’s Annual Open World program, San Francisco officials must wonder if the city could survive the demands of an Oracle four times its current size. A look at its portfolio will tell you more about scary reach and disappointing and unfocused spread that Oracle has now – 110 product lines spread across 14 different domains.

DATABASE
    DataScaler (October 2010)
    e-Test (acquired from Empirix) (March 2008)
    Innobase (October 2005)
    Moniforce (December 2007)
    mValent (February 2009)
    Secerno (May 2010)
    Sleepycat (February 2006)
    TimesTen (June 2005)
    TripleHop (June 2005)
               
MIDDLEWARE
    AmberPoint (February 2010)
    BEA (January 2008)
    Bharosa (July 2007)
    Bridgestream (September 2007)
    Captovation (January 2008)
    ClearApp (September 2008)
    Context Media (July 2005)
    Datanomic (April 2011)
    FatWire (June 2011)
    HyperRoll (September 2009)
    GoldenGate (July 2009)
    Java (April 2009)
    Oblix (March 2005)
    OctetString (November 2005)
    Passlogix (October 2010)
    Sigma Dynamics (August 2006)
    Silver Creek Systems (January 2010)
    Stellent (November 2006)
    Sunopsis (October 2006)
    Tacit Software (November 2008)
    Tangosol (March 2007)
    Thor Technologies (November 2005)
               
APPLICATIONS
    AppForge (April 2007)
    Collective Intellect (June 2012)
    Eloqua (December 2012)
    Haley (October 2008)
    InQuira (July 2011)
    Interlace Systems (October 2007)
    Involver (July 2012)
    LogicalApps (October 2007)
    Market2Lead (May 2010)
    Ndevr (February 2011)
    RightNow (October 2011)
    SelectMinds (September 2012)
    Taleo (February 2012)
    TempoSoft (December 2005)
    Vitrue (May 2012)

PRODUCT LINES
    Agile (May 2007)
    ATG (November 2010)
    Endeca (October 2011)
    Hyperion (March 2007)
    PeopleSoft (January 2005)
    Primavera (October 2008)
    Siebel (January 2006)
    Telephony@Work (June 2006)
  
IMPLEMENTATION AND INTEGRATION TOOLS
    Global Knowledge Software (GKS) (July 2008)

SERVERS, STORAGE, AND NETWORKING
    Ksplice (July 2011)
    Pillar Data Systems (June 2011
    Sun (April 2009)
    Xsigo Systems (July 2012)
    Virtual Iron (May 2009)
               
INDUSTRY SOLUTIONS
               
COMMUNICATIONS AND MEDIA
    Acme Packet (February 2013) (pending)
    Convergin (February 2010)
    eServGlobal's Universal Service Platform (USP) (May 2010)
    GoAhead (September 2011)
    HotSip (February 2006)
    MetaSolv Software (October 2006)
    Net4Call (April 2006)
    Netsure Telecom Limited     (September 2007)
    Portal Software (April 2006)
    Sophoi (October 2009)
  
ENGINEERING AND CONSTRUCTION
    Instantis (November 2012)
    Primavera (October 2008)
    Skire (July 2012)

FINANCIAL SERVICES
    i-flex (August 2005)

HEALTH SCIENCES
    ClearTrial (March 2012)
    Phase Forward (April 2010)
    Relsys (March 2009)

INDUSTRIAL MANUFACTURING
    Agile (May 2007)
    Conformia Software (June 2009)
    Demantra (June 2006)
    G-Log (September 2005)

INSURANCE
    AdminServer (May 2008)
    Skywire Software (June 2008)

RETAIL
    360Commerce (January 2006)
    Advanced Visual Technology (AVT) (October 2008)
    ProfitLogic (July 2005)
    Retek (April 2005)

UTILITIES
    DataRaker (December 2012)
    SPL WorldGroup (November 2006)
    LODESTAR (April 2007)

Failure to accept reality
It is also felt that Oracle executive management is out of touch with reality. The typical strategy is to first make fun of competitors, then ridicule them and finally scare the wits out of the customers who were even thinking of adopting competitors’ products.
They did this for Sun, HP, NetApp, EMC, VMWare, Salesforce, Microsoft (for MS SQL Server). Most of the times, customers can see through this and continue their adaptation of new products from customers.
Finally, Oracle sees the “writing on the wall” and realizes that it has to do something quickly – one in that mode, it joins the race by acquiring a competitor or second best guy in the race. PeopleSoft, Siebel, Sun (MySQL) and hosts of cloud acquisitions are cases in point.
Most recent examples are Oracle’s taking potshots for two consecutive years in Oracle Open World 2010 and 2011 at Salesforce.com and then when it couldn’t wean away customers from Salesforce.com or slow down the ramp up, it launched its own versions of cloud offerings in 2012 Open World.

Grasping for straws
Good news first, Oracle has not yet reached this stage yet – in this stage, very perceptible symptoms are – changing CEOs and executive staff in quick rotation and changing the product directions every so often. However, there is bound to be a moment, not in very distant future, when we will find that people will become so weary of Oracle products that Ellison will be either dislodged by a hostile board or will leave on his own. He has essentially no succession plan in place except bunch of execs like Thomas Kurian or Mark Hurd who can stake their claim to the crown. Thomas is well respected within the company but lacks charisma and chutzpah of Ellison. Mark may not be as respected but has good experience of cutting costs – like he did at HP.
Death Knell
In this stage, the company will slowly vanish into irrelevance or acquired/merge into another competitor or go belly up. For the sake of hundreds of thousands of professionals using, preaching and earning their bread from Oracle Technologies, I just hope Oracle never reaches that stage.
Will it be able to recover from this downward spiral?
Oracle can arrest this dance towards its vanishing into oblivion – question that really begs for an answer is – will it have the honesty to first admit and then stop this march?
First of all, Oracle should focus and determine its core strength and then focus on building up on those. There is no prudence demonstrated in draining money on acquisitions and then selling those companies at markdown, or worst, writing off the charge as a loss.
It is about time Oracle give up its greed on squeezing more money out of its customer and first create products and value that customers will willingly play obscene amount of money for.