Friday, December 17, 2010

Salesforce.com's $212 Million Acquisition of Heorku - A Sparkling Gem In Radiant Future Of Cloud And PaaS

I met James Lindenbaum, a founder of Heroku, in early 2009, at the Under The Radar conference in Mountain View. We had a long conversation on cloud as a great platform for Ruby, why Ruby on Rails is a better framework than PHP, and viability of PaaS as a business model. He also explained to me why he chose to work on Heroku at Y Combinator. I was sold on their future, on that day, and kept in touch with them since then. The last week, Salesforce.com acquired Heroku for $212 million. That's one successful exit, which is good news in many different dimensions.

PaaS is a viable business model

PaaS is not easy. It takes time, laser sharp focus, and hard work to build something that the developers would use and pay for. A few companies have tried and many have failed. But, it is refreshing to see the platform and the ecosystem that Heroku has built since its inception. Heroku did not raise a lot of money, kept the cost low, and attracted customers early on. I was told (by Byron, I think) that an average cost for Heroku to run a free Ruby app for a month was $1. They considered it as marketing cost to get new customers and convert the free customers to paying ones, as they outgrew their needs. I cannot overpraise this brilliant execution model. I hope to see more and more entrepreneurs being inspired from - simplicity, elegance, and execution of Heroku's model - to help the developers deploy, run, and scale their applications on the cloud. In the last few years, we have seen a great deal of innovation in dynamic programming languages, access algorithms, and NoSQL persistence stores. They all require a PaaS that the developers can rely on - without worrying about the underlying nuts and bolts - and focus on what they are good at - building great applications. If anyone had the slightest doubt on viability of PaaS as a business model, this acquisition is a proof point that PaaS is indeed the future. Heroku is just the beginning and I am hoping for more and more horizontal as well as vertical PaaS that the entrepreneurs will aspire to build.

Superangels and incubators do work

There have been many debates on viability of the investing approach of the superangels and the incubators, where people are questioning, whether the approach of thin slicing the investment, by investing into tens and hundreds of companies, would yield similar returns, as compared to return on traditional venture capital investment. I also blogged about the imminent change in the VC climate, and decided to watch their returns. The numbers are in with Heroku. It's a first proof point that a superangel or an incubator approach, structurally, does not limit the return on the investment. I believe in investors investing in right people solving the right problems. If you ever meet James and hear him passionately talk about Ruby, the Heroku platform, and the developer community, you will quickly find out why they were successful. Hats off to YC on finding this "jewel". No such thing as too little investment, or too many companies.

Ruby goes enterprise

I know many large ISVs that have been experimenting with Ruby for a while, but typically these efforts are confined to a few small projects. It's good to see that Ruby, now, has a shot of getting much broader adoption. This would mean more developers learning Ruby, cranking out great enterprise gems, embracing Git, and hopefully open source some of their work. I have had many religious discussions, with a few cloud thought leaders and bloggers in the past few months, regarding the boundaries of PaaS. The boundaries have always been blurry - somewhere between SaaS and IaaS - but, I don't care. My heart is at delivering the applications off the cloud that scales, delivers compelling experiences, and leverages economies of scale and network effects. To me, PaaS is means to an end and not the end. I am hoping that an acquisition of a PaaS vendor by a successful SaaS vendor will make Ruby more attractive to enterprise ISVs and non-Ruby developers.

I have no specific insights into what Salesforce.com will do with Heroku, but I hope, they make a good home for Heroku, where they flourish and continue to do great work on Ruby and PaaS. This is what a cloud and Ruby enthusiast would wish for.

Monday, November 15, 2010

10 Business Books In 2010

These are the 10 business books published in 2010, that I would recommend you to read. Originally, I wrote this on Quora, in response to "What are must read business books of 2010?". Yes, I have read all of them, and no, they are not in any specific order.

1) What the Dog Saw by Malcolm Gladwell

I am a big fan of Malcolm Gladwell and his style. This is a compilation of his "The New Yorker" stories. Even though the articles are available on his website, this book makes it a great read.

2) Cognitive Surplus by Clay Shirky

The next time someone asks you how come people have so much time to blog, answer questions on Quora, or contribute to Wikipedia, ask them to read this book.

3) The Big Short by Michael Lewis

Want to know all about CDO and subprime mortgage and still be entertained? This is the book. Michael Lewis has great storytelling skills that makes serious and complex topics fun to read. I like this book as much as I liked Moneyball - http://amzn.to/b9YPx9

4) Open Leadership by Charlene Li

If you liked Groundswell - http://amzn.to/c8faH5 - you will like this as well. If you are interested in organizational transformation through social media, this will make a great read. Social media adoption can certainly make the leaders more credible, open, and transparent. Being a social media freak and an enterprise 2.0 strategist, I loved this book.

5) Engage by Brian Solis

This book is Seth Godin meet Social Media. It's a must-read if you are a marketer, trying to understand the impact of social media on your brand and working on engaging your customers using social media. Brian Solis has a fluid style with a lot of relevant examples.

6) The New Polymath by Vinnie Mirchandani

Vinnie is a great enterprise software analyst and a prolific blogger. I closely follow his work. This is an upbeat book that will excite the technologists as well as the business folks. If you think you have a stretch goal and want to change the world, this book will further stretch your stretch goals, and will give you a reason and purpose, to get out of bed every morning and run for it.

7) Rework by Jason Fried

I have followed 37Signals and Jason's blog. This book puts everything together with illustrations and a simple style making it easy to read, just like 37Signals. If you are itching to be an entrepreneur, this might make you take that leap. If you're starting out and want inspiration and design principles, this is the book. All design is re-design and so is this book.

8) The Facebook Effect by David Kirkpatrick

Some watch the movie, I prefer to read a book. The book is more accurate than the movie. Well, duh. David is a great writer, and he used the access that he had to Zuckerberg and Facebook, to produce a great book. It's quite insightful.

9) Gamestorming by Dave Gray

I love XPLANE. They do a great job and now they are part of Dachis group where I am expecting them to do even better. It's incredibly difficult to take complex concepts and simplify to communicate to any audience. The book outlines great approaches to accomplish the simplicity and facilitate learning, discovery, and decision making.

10) Delivering Happiness by Tony Hsieh

Zappos is a great company. I have learned a lot from its culture and from Tony's management style. This is a must-read, if you believe you want to excel in serving your customers and have your entire team live by those values.

And this is the first 2011 book that you may want to read:


Knowing Umair, this will be a great book.

Tuesday, November 9, 2010

Challenging Stonebraker’s Assertions On Data Warehouses - Part 2

Check out the Part 1 if you haven’t already read it to better understand the context and my disclaimer. This is the Part 2 covering the assertions from 6 to 10.

Assertion 6: Appliances should be "software only."

“In my 40 years of experience as a computer science professional in the DBMS field, I have yet to see a specialized hardware architecture—a so-called database machine—that wins.”

This is a black swan effect; just because someone hasn’t seen an event occur in his or her lifetime, it doesn’t mean that it won’t happen. This statement could also be re-written as “In my 40 years of experience, I have yet to see a social network that is used by 500 million people.” You get the point. I am the first one who would vote in favor of commodity hardware against a specialized hardware, but there are very specific reasons why the specialized hardware makes sense in some cases.

“In other words, one can buy general purpose CPU cycles from the major chip vendors or specialized CPU cycles from a database machine vendor.”

Specialized machines don’t necessarily mean specialized CPU cycles. I hope the word “CPU cycle” is used as metaphor and not to indicate its literal meaning.

“Since the volume of the general purpose vendors are 10,000 or 100,000 times the volume of the specialized vendors, their prices are an order of magnitude under those of the specialized vendor.”

This isn’t true. The vendors who make general-purpose hardware also make specialized hardware, and no, it’s not an order of magnitude expensive.

“To be a price- performance winner, the specialized vendor must be at least a factor of 20-30 faster.”

It’s a wrong assumption that BI vendors use specialized hardware just because of the performance reasons. The “specialized” in many cases for an appliance is simply a specialized configuration. The appliance vendors also leverage their relationship with the hardware vendors to fine tune the configuration based on their requirements, negotiate a hefty discount, and execute a joint go-to-market strategy.

The enterprise software follows value-based pricing and not cost-based pricing. The price difference between a commodity and a specialized appliance is not just the difference of the cost of hardware that it runs on.

“However, every decade several vendors try (and fail).”

Not sure what is the success criteria behind this assertion to declare someone a winner or a failure. Acquisitions of Netezza, Greenplum, and Kickfire are recent examples of how well the appliance companies have performed. The incumbent appliance vendors are doing great, too.

“Put differently, I think database appliances are a packaging exercise”

The appliances are far more than a packaging exercise. Other than making sure that the software appliance works on the selected hardware, commoditized or otherwise, they provide a black box lifecycle management approach to the customers. The upfront cost of an appliance is a small fraction of the overall money that the customers would end up spending during the entire lifecycle of an appliance and the related BI efforts. The customers do welcome an approach where they are responsible for managing one appliance against five different systems at ten different levels with fifteen different technology stack versions.

Assertion 7: Hybrid workloads are not optimized by "one-size fits all."

Yes, I agree, but that’s not the point. It’s difficult to optimize hybrid workloads for a row or a column store, but it is not as difficult, if it’s a hybrid store.

“Put differently, two specialized systems can each be a factor of 50 faster than the single "one size fits all" system in solution 1.”

Once again, I agree, but it does not apply to all the situations. As I discussed earlier, the performance is not the only criteria that matters in the BI world. In fact, I would argue the opposite. Just because the OLTP and OLAP systems are orthogonal, the vendors compromised everything else to gain the performance. Now that’s changing. Let’s take an example of an operational report. This is the kind of report that only has the value if consumed in realtime. For such reports, the users can’t wait until the data is extracted out of the OLTP system, cleaned up, and transferred into the OLAP system. Yes, it could be 50 times faster, but completely useless, since you missed the boat.

The hybrid systems, the once that combine OLTP and OLAP, are fairly new, but they promise to solve a very specific problem, which is real real-time. While the hybrid systems evolve, the computational capabilities of OLTP and OLAP systems have started to change as well. I now see OLAP systems supporting write-backs with a reasonable throughput and OLTP systems with good BI style query performance, all of these achieved through modern hardware and clever use of architectural components.

Let’s not forget what optimization really is. It means desired functionality at reasonable performance. A real-time report, that takes 10 seconds to run could be far more valuable than a report that runs under ten milliseconds, three days later.

“A factor of 50 is nothing to sneeze at.”

Yes, point taken. :-)

Assertion 8: Essentially all data warehouse installations want high availability (HA).

No, they don’t. This is like saying all the customers want five 9 SLA on the cloud. I don’t underestimate the business criticality of a DW if it goes down, but not all the DW are being used 24x7 and are mission critical. One size doesn’t fit all. And, if your DW is not required to be highly available, you need to ask yourself, whether it is fair for you to pay for the HA architectural cost, if you don’t want it. Tiered SLAs are not new, and tiered HA is not a terrible idea.

Let’s talk about the DWs that do require to be highly available.

“Moreover, there is no reason to write a DBMS log if this is going to be the recovery tactic. As such, a source of run-time overhead can be avoided.”

I am a little confused how this is worded. Which logs are we referring to - the source systems or the target systems? The source systems are beyond the control of a BI vendor. There are newer approaches to design an OLTP system without a log, but that’s not up for discussion for this assertion. If the assertion is referring to the logs of the target system, how does that become a run-time overhead? Traditional DW systems are a read-only system at runtime. They don’t write logs back to the system. If he is referring to the logs while the data is being moved to DW, that’s not really run-time, unless we are referring to it as a hot-transfer.

There is one more approach, NoSQL, where eventual consistency is achieved over a period of time and the concept of a “corrupted system” is going away. Incomplete data is an expected behavior and people should plan for it. That’s the norm, regardless of a system being HA or not. Recently Netflix moved some of its applications to the cloud, where they have designed a background data fixer to deal with data inconsistencies.

HA is not black and white, and there are way more approaches, beyond the logs, to accomplish to achieve desired outcome.

Assertion 9: DBMSs should support online reprovisioning.

“Hardly anybody wants to take the required amount of down time to dump and reload the DBMS. Likewise, it is a DBA hassle to do so. A much better solution is for the DBMS to support reprovisioning, without going offline. Few systems have this capability today, but vendors should be encouraged to move quickly to provide this feature.”

I agree. I would add one thing. The vendors, even today, have a trouble supporting offline provisioning to cater to the increasing load. On-line reprovisioning is not trivial, since in many cases, it requires to re-architect their systems. The vendors typically get away with this, since the most customers don’t do capacity planning in real-time. Unfortunately, traditional BI systems are not commodity where the customers can plug-in more blades when they want and take them out when they don’t.

This is the fundamental premise behind why cloud makes it a great BI platform to address such re-provisioning issues with elastic computing. Read my post “The Future Of BI In The Cloud”, if you are inclined to understand how horizontal scale-out systems can help.

Assertion 10: Virtualization often has performance problems in a DBMS world.

This assertion, and the one before this, made me write the post “The Future Of BI In The Cloud”. I would not repeat what I wrote there, but I will quickly highlight what is relevant.

“Until better and cheaper networking makes remote I/O as fast as local I/O at a reasonable cost, one should be very careful about virtualizing DBMS software.”

Virtualizing I/O is not a solution for large DW with complex queries. However, as I wrote in the post, a good solution is not to make the remote I/O faster, but rather tap into the innovation of software-only SSD block I/O that are local.

“Of course, the benefits of a virtualized environment are not insignificant, and they may outweigh the performance hit. My only point is to note that virtualizing I/O is not cheap.”

This is what a disruption initially looks like. You start seeing good enough value in an approach, for certain types of solutions, that seems expensive for other set of solutions. Over a period of time, rapid innovation and economies of scale remove this price barrier. I think that’s where the virtualization stands, today. The organizations have started to use the cloud for IaaS and SaaS for a variety of solutions including good enough self-service BI and performance optimization solutions. I expect to see more and more innovation in this area where traditional large DW will be able to get enough value out of the cloud, even after paying the virtualization overhead.

Wednesday, November 3, 2010

Bottom Of The Pyramid – Nokia’s Second Act

The two-third of world’s 4.6 billon mobile users live in the emerging markets. Millions of these users live below the poverty line and are part of the bottom of the pyramid (BOP). Nokia is the market leader in these emerging markets, at least for now, with 34% market share. It’s clear from rapidly declining Nokia’s marketshare and an appointment of new CEO, Stephen Elop, that Nokia needs a second act. I believe the BOP is what could be the next big thing for Nokia.

The recent NYTimes story highlights a Nokia’s service, to supply commodity data to the farmers in India, using a text message. So far, 6.3 million people have signed up for this service. Nokia is planning to roll out this service, Life tools, in Nigeria as well. This is part of their Ovi mobile business.

I have written before on impact of cloud computing and mobile on the bottom of the pyramid and the importance of public policy innovation in emerging markets. The BOP is one of the biggest opportunities that Nokia currently has. Nokia has been losing marketshare in the smartphone category, and it is going to get increasingly difficult for Nokia to compete with Apple, Google, RIM, and now Microsoft. However, the very same vendors will find it equally difficult to move down the chain to compete with Nokia in the emerging markets.

One of the biggest business challenges to cater to the BOP is not a desire to market or a product to offer, but it is the lack of direct access to these consumers. The people at the BOP are incredibly difficult to reach. I have seen many go-to-market plans fail because it is either impossible or prohibitively expensive to market to these consumers. One of the biggest assets Nokia has is the relationship, the channel, with the people at the BOP. Now is the time to focus and leverage that channel by providing them with the content and the services that could be served on these phones via a strong platform, built for the BOP, and a vibrant ecosystem built around it.

My two cents: exit from the Smartphone category and double down the investment to serve the people at the bottom of the pyramid.

Nokia, that could be your second act.

Thursday, October 28, 2010

Challenging Stonebraker’s Assertions On Data Warehouses - Part 1

I have tremendous respect for Michael Stonebraker. He is an apt visionary. What I like the most about him is his drive and passion to commercialize the academic concepts. ACM recently published his article “My Top 10 Assertions About Data Warehouses." If you haven’t read it, I would encourage you to read it.

I agree with some of his assertions and disagree with a few. I am grounded in reality, but I do have a progressive viewpoint on this topic. This is my attempt to bring an alternate perspective to the rapidly changing BI world that I am seeing. I hope the readers take it as constructive criticism. This post has been sitting in my draft folder for a while. I finally managed to publish it. This is Part 1 covering the assertions 1 to 5. The Part 2 with the rest of the assertions will follow in a few days.

“Please note that I have a financial interest in several database companies, and may be biased in a number of different ways.”

I appreciate Stonebraker’s disclaimer. I do believe that his view is skewed to what he has seen and has invested into. I don’t believe there is anything wrong with it. I like when people put money where their mouth is.

As you might know, I work for SAP, but this is my independent blog and these are my views and not those of SAP’s. I also try hard not to have SAP product or strategy references on this blog to maintain my neutral perspective and avoid any possible conflict of interest.

Assertion 1: Star and snowflake schemas are a good idea in the data warehouse world.

This reads like an incomplete statement. The star and snowflake schemas are a good idea because they have been proven to perform well in the data warehouse world with row and column stores. However, there are emergent NoSQL based data warehouse architectures I have started to see that are far from a star or a snowflake. They are in fact schemaless.

“Star and Snowflake schemas are clean, simple, easy to parallelize, and usually result in very high-performance database management system (DBMS) applications.”

The following statement contradicts the statement above.

“However, you will often come up with a design having a large number of attributes in the fact table; 40 attributes are routine and 200 are not uncommon. Current data warehouse administrators usually stand on their heads to make "fat" fact tables perform on current relational database management systems (RDBMSs).”

There are a couple of problems with this assertion:
  1. The schema is not simple; 200 attributes, fact tables, and complex joins. What exactly is simple?
  2. Efficient parallelization of a query is based on many factors, beyond the schema. How the data is stored and partitioned, performance of a database engine, and hardware configuration are a few to name.
"If you are a data warehouse designer and come up with something other than a snowflake schema, you should probably rethink your design.”

Really?

The requirement, that the schema has to be perfect upfront, has introduced most of the problems in the BI world. I call it the design time latency. This is the time it takes after a business user decides what report/information to request and by the time she gets it (mostly the wrong one.) The problem is that you can only report based what you have in your DW and what’s tuned.

This is why the schemaless approach seems more promising as it can cut down the design time latency by allowing the business users to explore the data and run ad hoc queries without locking down on a specific structure.

Assertion 2: Column stores will dominate the data warehouse market over time, replacing row stores.

This assertion assumes that there are only two ways of organizing data, either in a row store or in a column store. This is not true. Look at my NoSQL explanation above and also in my post “The Future Of BI In The Cloud”, for an alternate storage approach.

This assertion also assumes that the access performance is tightly dependent on how the data is stored. While this is true in the most cases, many vendors are challenging this assumption by introducing an acceleration layer on top of the storage layer. This approach makes is feasible to achieve consistent query performance, by clever acceleration architecture, that acts as an access layer, and does not depend on how data is stored and organized.

“Since fact tables are getting fatter over time as business analysts want access to more and more information, this architectural difference will become increasingly significant. Even when "skinny" fact tables occur or where many attributes are read, a column store is still likely to be advantageous because of its superior compression ability."

I don’t agree with the solution that we should have fatter fact tables when business analysts want more information. Even if this is true, how will column store be advantageous when the data grows beyond the limit where compression isn’t that useful?

“For these reasons, over time, column stores will clearly win”

Even if it is only about rows versus columns, the column store may not be a clear commercial winner in the marketplace. Runtime performance is just one of many factors that the customers consider while investing in DW and business intelligence.

“Note that almost all traditional RDBMSs are row stores, including Oracle, SQLServer, Postgres, MySQL, and DB2.”

Exactly!

The row stores, with optimization and acceleration, have demonstrated reasonably good performance to stay competitive. Not that I favor one over the other, but not all row-based DW are that large or growing rapidly, and have serious performance issues, warranting a switch from a row to a column.

This leads me to my last issue with this assertion. What about a hybrid store – row and column? Many vendors are trying to figure this one out and if they are successful, this could change the BI outlook. I will wait and watch.

Assertion 3: The vast majority of data warehouses are not candidates for mainmemory or flash memory.

I am assuming that he is referring to the volatile flash memory and not flash memory as storage. Though, the SSD block storage have huge potential in the BI world.

“It will take a long time before main memory or flash memory becomes cheap enough to handle most warehouse problems.”

Not all DW are growing at the same speed. One size does not fit all. Even if I agree that the price won’t go down significantly, at the current price point, main memory and flash memory can speed up many DW without breaking the bank.

The cost of DW, and especially the cost of flash memory, is a small fraction of the overall cost; hardware, license, maintenance, and people. If the added cost of flash memory makes business more agile, reduces maintenance cost, and allows the companies to make faster decisions based on smarter insights, it’s worth it. The upfront capital cost is not the only deciding factor for BI systems.

“As such, non-disk technology should only be considered for temporary tables, very "hot" data elements, or very small data warehouses.”

This is easier said than done. The customers will spend significant more time and energy, on a complicated architecture, to isolate the hot elements and running them on a different software/hardware configuration.

Assertion 4: Massively parallel processor (MPP) systems will be omnipresent in this market.

Yes, MPP is the future. No disagreements. The assertion is not about on-premise or the cloud, but I truly believe that cloud is the future for MPP. There are other BI issues that need to be addressed before cloud makes it a good BI platform for a massive scale DW, but the cloud will beat any other platform when it comes to MPP with computational elasticity.

Assertion 5: "No knobs" is the only thing that makes any sense.

“In other words, look for "no knobs" as the only way to cut down DBA costs.”

I agree that “no knobs” is what the customers should thrive for to simplify and streamline their DW administration, but I don’t expect these knobs to significantly drive down the overall operational cost, or even the cost just associated with the DBAs. Not all the DBAs have a full time job to manage and tune the DW. The DW deployments go through a cycle where the tasks include schema design, requirements gathering, ETL design etc. Tuning or using the “knobs” is just one of many tasks that the DBAs perform. I absolutely agree that the no-knobs would certainly take some burden off the shoulders of a DBA, but I disagree that it would result into significant DBA cost-savings.

For a fairly large deployment, there is significant cost associated with the number of IT layers
that are responsible to channel the reports to the business users. There is an opportunity to invest into the right kind of architecture, technology-stack for the DW, and the tools on top of that to help increase the ratio of Business users to the BI IT. This should also help speed up the decision-making process based on the insights gained from the data. Isn’t that the purpose to have a DW to begin with? I see the self-service BI as the only way to make IT scale. Instead of cutting the DBA cost, I would rather focus on scaling the BI IT with the same budget and a broader coverage amongst the business users in an organization.

Monday, October 25, 2010

The Future Of BI In The Cloud



Actual numbers vary based on whom you ask, but the general consensus is that the Business Intelligence (BI) and Analytics in the cloud is a fast growing market. IDC expects a compounded annual growth rate (CAGR) of 22.4% through 2013. This growth is primarily driven by two kinds of SaaS applications. The first kind is a purpose-specific analytics-driven application for business processes such as financial planning, cost optimization, inventory analysis etc. The second kind is a self-service horizontal analytics application/tool that allows the customers and ISVs to analyze data and create, embed, and share analysis and visualizations.

The category that is still nascent and would require significant work is the traditional general-purpose BI on large data warehouses (DW) in the cloud. For the most enterprises, not only all the DW are on-premise, but the majority of the business systems that feed data into these DW are on-premise as well. If these enterprises were to adopt BI in the cloud, it would mean moving all the data, warehouses, and the associated processes such as ETL in the cloud. But then, the biggest opportunities to innovate in the cloud exist to innovate the outside of it. I see significant potential to build black-box appliance style systems that sit on-premise and encapsulate the on-premise complexity – ETL, lifecycle management, and integration - in moving the data to the cloud.

Assuming that the enterprises succeed in moving data to the cloud, I see a couple of challenges, if treated as opportunities, will spur the most BI innovation in the cloud.

Traditional OLAP data warehouses don’t translate well into the cloud:

The majority of on-premise data warehouses run on some flavor of a relational or a columnar database. The most BI tools use SQL to access data from these DW. These databases are not inherently designed to run natively on the cloud. On top of that, the optimizations performed on these DW such as sharding, indices, compression etc. don’t translate well into the cloud either since cloud is a horizontally elastic scale-out platform and not a vertically integrated, scale-up, system.

The organizations are rethinking their persistence as well as access languages and algorithms options, while moving their data to the cloud. Recently, Netflix started moving their systems into the cloud. It’s not a BI system, but it has the similar characteristics such as high volume of read-only data, a few index-based look-ups etc. The new system uses S3 and SimpleDB instead of Oracle (on-premise). During this transition, Netflix picked availability over consistency. Eventual consistency is certainly an option that BI vendors should consider in the cloud. I have also started seeing DW in the cloud that uses HDFS, Dynamo, and Cassandra. Not all the relational and columnar DW systems will translate well into NoSQL, but I cannot overemphasize the importance of re-evaluating persistence store and access options when you decide to move your data into the cloud.

Hive, a DW infrastructure built on top of Hadoop, is a MapReduce meet SQL approach. Facebook has a 15 petabytes of data in their DW running Hive to support their BI needs. There are a very few companies that would require such a scale, but the best thing about this approach is that you can grow linearly, technologically as well as economically.

The cloud does not make it a good platform for I/O intensive applications such as BI:

One of the major issues with the large data warehouses is, well, the data itself. Any kind of complex query typically involves an intensive I/O computation. But, the I/O virtualization on the cloud, simply does not work for large data sets. The remote I/O, due to its latency, is not a viable option. The block I/O is a popular approach for I/O intensive applications. Amazon EC2 does have block I/O for each instance, but it obviously can’t hold all the data and it’s still a disk-based approach.

For BI in the cloud to be successful, what we really need is ability for scale-out block I/O, just like scale-out computing. Good news is that there is at least one company, Solidfire, that I know, working on it. I met Dave, the founder, at the Structure conference reception. He explained to me what he is up to. Solidfire has a software solution that uses solid state drives (SSD) as scale-out block I/O. I see huge potential in how this can be used for BI applications.

When you put all the pieces together, it makes sense. The data is distributed across the cloud on a number of SSDs that is available to the processors as block I/O. You run some flavor of NoSQL to store and access this data that leverages modern algorithms and more importantly horizontally elastic cloud platform. What you get is commodity and blazingly fast BI at a fraction of cost with pay-as-you-go subscription model.
Now, that’s what I call the future of BI in the cloud.

Friday, October 15, 2010

Can A Product Manager Be Effective Without Product Design Skills?

I am very passionate about the topic of design and design-thinking. When I saw this question on Quora, I decided to post my answer. Following is directly from my answer to this question on Quora:

The answer is "Definitely not."

It's not about the product design by itself, but it's about applying core and transferable product design skills to product management. Let's break it down:

1) Understanding users: Good product designers have great user research, observation, and listening skills to put themselves into the shoes of a user and understand the real, mostly unspoken and latent, needs of the end users.

2) Being self-critical: If you are a trained designer, you would stay away from self-referential design, which is a root cause for many failed products. Good product designers are self-critical about their approach and the deliverables and are always open to feedback to iterate on their design.

3) Working with designers: If you are a designer, you have great empathy for fellow designers. I have seen products fail, simply because, the product managers can't work with the designers and don't share the same mindset.

4) A "maker" mentality: The designers are makers. They make things. The product managers typically don't, the engineers do. For a product manager, it's incredibly important to have a "maker" mentality. They should continuously be making and refining, by themselves or with the help of the engineers. The product managers, who believe that their responsibility ends when they are done gathering the requirements are likely to fail, miserably in most cases.

5) A "T-shaped" product manager: If you're a product manager, the vertical line of the "T" is your core PM skills. However, successful product managers go beyond their core skills, the horizontal line in the letter "T", to learn more about product design, engineering etc. This ensures that they have a holistic perspective of the product. That leads me to my last point.

6) General Manager: viable, feasible, and desirable: A good product from a vendor's perspective is commercially viable, technologically feasible, and desirable by the end users. Many product managers stop at the business needs, but they truly need to go beyond that to work with the engineering to make it technologically feasible, and have a design mindset to work with the designers to make it desirable by the end users. The product managers should thrive for a "general manager" mindset, of which, product design is a core element.