Format | Apache Arrow
arrow.apache.org · 375 words · saved by 1 readers
Arrow Format
Apache Arrow Overview Apache Arrow is a multi-language toolbox for building high performance applications that process and transport large data sets. It is designed to both improve the performance of analytical algorithms and the efficiency of moving data from one system or programming language to another. A critical component of Apache Arrow is its in-memory columnar format, a standardized, language-agnostic specification for representing structured, table-like datasets in-memory. This data format has a rich data type system (included nested and user-defined data types) designed to…
saved by
related reading
- Databases in 2025: A Year in Review // Blog // Andy Pavlo - Carnegie Mellon Universitycs.cmu.edu
- Lance Format and LanceDB: Columnar Storage for the Embedding Age | Andrea Bozzo | Blogandreabozzo.github.io
- NYSRGnotes.ekzhang.com
- Pruning for Icebergsnowflake.com
- Databricks + Tabular | Databricks Blogdatabricks.com
- Schema Registry in Kafka: Avro, JSON and Protobuf | by Ismael Sánchez Chaves | .Net Programming | Mediummedium.com
- From monolith to Lakebase to LTAP: rethinking the database from storage up | Databricks Blogdatabricks.com
- Introducing Husky, Datadog’s third-generation event store | Datadogdatadoghq.com
- Using MongoDB with Pandas, NumPy, and PyArrow Data Engineeringanalyticsvidhya.com
- homepages.cwi.nl/~boncz/lsde/papers/p215-dageville-snowflake.pdfhomepages.cwi.nl
- Snowflake key concepts and architecture | Snowflake Documentationdocs.snowflake.com
- Grafana Tempo 1.5 release: New metrics features with OpenTelemetry, Parquet support, and the path to 2.0grafana.com