This is part two in our six-part series exploring alternatives to Apache Parquet. In our first post, we examined Parquet’s dominance in columnar storage. Today we explore Apache Arrow, not as a replacement for Parquet, but as its essential in-memory counterpart.
In our previous post, we established how Parquet revolutionised analytical storage through efficient columnar compression and encoding. But there’s a fundamental trade-off: the very optimisations that make Parquet storage-efficient (dictionary encoding, bit-packing, run-length encoding) also make it computationally expensive to process. Every query requires decompressing and decoding data before computation can begin. Apache Arrow flips this equation, optimising for computational speed over storage compactness.