Lance Format and LanceDB: Columnar Storage for the Embedding Age | Andrea Bozzo | Blog
Lance is a columnar format built for ML workloads — fast random access, native vector indexing, and zero-copy Arrow integration. This article walks through the format, LanceDB, and wiring it into a live NATS stream.
Columnar storage for the embedding age Table of Contents Lance Format and LanceDB: Columnar Storage for the Embedding Age # I have been spending time lately looking at storage formats from the perspective of ML infrastructure rather than analytics. Most of my day-to-day work sits in the Arrow/DataFusion ecosystem, so when I wanted to add a vector search layer to a small event streaming project, my first instinct was to ask: is there a format that speaks Arrow natively, handles embeddings without a separate vector database, and doesn't require me to run a separate server? Lance turned out to be
Explore this link on the map →saved by
related reading
- Building a web search engine from scratch in two months with 3 billion neural embeddingsblog.wilsonl.in
- Not All Vector Databases Are Made Equal | Towards Data Sciencetowardsdatascience.com
- Investing in Pinecone | Andreessen Horowitza16z.com
- From monolith to Lakebase to LTAP: rethinking the database from storage up | Databricks Blogdatabricks.com
- Chroma raises $18M seed round | Chromatrychroma.com
- What Is a Vector Database? | IBMibm.com
- Build and Scale a Powerful Query Engine with LlamaIndex & Rayanyscale.com
- How We Store and Search 30 Billion Facesclearview.ai
- Databases in 2024: A Year in Review // Blog // Andy Pavlo - Carnegie Mellon Universitycs.cmu.edu
- From prototype to production: Vector databases in generative AI applications - Stack Overflowstackoverflow.blog
- Announcing ScaNN: Efficient Vector Similarity Searchai.googleblog.com
- Vector databases explained | Lantern Bloglantern.dev