Hook

Their other posts in the index, biggest breakout first.
Databricks is the hardest company I've tried to explain yet. This is part seven in my series explaining top private tech companies. For decades, companies stored their data in one of two places. A data warehouse, which is perfect for structured data like spreadsheets and order forms. It's fast and reliable, but expensive and rigid. Or a data lake, built for messier data like videos and images. It's cheap and flexible, but slow and unreliable. Most large companies used both. Two separate systems, two teams managing the whole mess. In 2013, seven researchers from Berkeley built a way to process massive data sets by distributing the work across thousands of computers at the same time. They called it Apache Spark. On top of Spark, they built something called the Lakehouse, a hybrid of the two that combined the best of both. Fast, reliable, cheap, and flexible. Only possible because it was built on Spark. Then they made a strange decision. They open-sourced both Apache Spark and the Lakehouse, giving it away entirely for free. They bet that if the world built on top of their technology, the adoption would create a moat. It totally worked. Over the next decade, millions of developers across thousands of big companies built their entire data infrastructure on Spark. Databricks instead sold proprietary features and managed platform service on top of this open-source foundation. In 2022, when ChatGPT was released, Databricks was pretty perfectly positioned. Their customers could now build AI directly on top of the data they already had, using Databricks for retrieval, model training, and autonomous agents. And then pay Databricks like a utility bill. The more AI workloads they ran, the higher the utility bill. Databricks' existing infrastructure just became infinitely more valuable because of AI. The revenue growth is wild. From $200 million in 2020 to a $5.4 billion run rate in 2026. Roughly 27x in six years, with the curve steepening every year after ChatGPT. Over that period, Databricks raised $20 billion from Andreessen, Thrive, NVIDIA, and others, with a current valuation of $134 billion. An IPO is on the horizon. If data really is the new oil, Databricks own the pipeline, and they built it by letting the oil flow freely.