LanceDB
by LanceDB Inc.
The multimodal lakehouse for AI — vector, full-text, and hybrid search at billion-row scale
LanceDB is an open-source, AI-native multimodal lakehouse that unifies vector, full-text, and hybrid search with data curation, feature engineering, versioning, and high-throughput training-data access. It is built for ML and AI engineering teams running production retrieval and training pipelines at scale.
LanceDB is an AI-native data platform built on the open-source Lance columnar format that unifies data management across the entire ML lifecycle — from raw multimodal files (text, images, video, audio) through feature engineering to model training, retrieval, and deployment. It provides vector, full-text, and hybrid search with SQL filtering, automatic data versioning (branching, tagging, and rollback), deduplication and curation, Python-UDF feature engineering with schema evolution that avoids table rewrites, and high-throughput training-data access (up to ~70% Model FLOPS Utilization with fast random access and no egress bottlenecks). LanceDB scales to 100B+ rows per table and 100K+ queries per second on object storage, and is used by teams at Netflix, Runway, Midjourney, ByteDance, Uber, Character.AI, WeRide, and CodeRabbit. The company raised a $30M Series A led by Theory Ventures in June 2025 (with CRV, Y Combinator, and Databricks Ventures participating) and entered 2026 with Lance-native SQL retrieval via DuckDB, multi-bucket object storage, and continued open-source momentum. It is offered as free Apache-2.0 open source, a usage-based LanceDB Cloud (public beta), and an annual-commit LanceDB Enterprise tier.
At a Glance
- Category
- Data & Analytics
- Pricing
- Free, Usage-based, Contact for pricing
- Target Market
- CTOs, Data Scientists, Enterprise Developers, AI Engineers
- Founded
- 2021
- Headquarters
- San Francisco, USA
Key Features
- ✓Multimodal search
Vector, full-text, and hybrid search with SQL filtering across text, images, video, and audio.
- ✓Lance open format
Built on the open-source Lance columnar format designed for fast random access and no-egress training workflows.
- ✓Data curation & versioning
Deduplication, edge-case discovery, and automatic branching, tagging, and rollback of datasets.
- ✓Feature engineering
Python UDFs with automatic updates and schema evolution without rewriting tables.
- ✓Training-data access
High Model FLOPS Utilization with fast random access and no egress bottlenecks for model training.
- ✓Scale
Handles 100B+ rows per table and 100K+ queries per second on object storage.
Capabilities
Use Cases
- •Retrieval for RAG and agents
Serve low-latency vector and hybrid retrieval for generative-AI and agent applications.
- •Training-data lakehouse
Curate, version, and stream multimodal datasets directly into model-training pipelines.
- •Multimodal data curation
Deduplicate and organize embeddings, images, documents, and video for production ML.
Ideal For
Best For
- ✓Multimodal vector and hybrid search for RAG
- ✓Curating and versioning training datasets at scale
- ✓Feature engineering across text, image, and video
Integrations
Deployment
Market Analysis
Pros
- ✓Open-source and self-hostable with no lock-in
- ✓Unifies the full ML data lifecycle in one format
- ✓Proven at scale with users like Netflix, ByteDance, and Uber
Cons
- ✗Younger than incumbent vector databases
- ✗Cloud and Enterprise tiers still maturing (Cloud in public beta)
Pricing
Open Source
$0
- ✓Apache 2.0 license
- ✓Self-hosted
- ✓All core features
LanceDB Cloud
Usage-based
- ✓Managed service
- ✓Public beta
- ✓Pay-as-you-go
LanceDB Enterprise
Contact for pricing
- ✓Petabyte-scale distributed lakehouse
- ✓Private deployment
- ✓Annual commit via AWS Marketplace
Apache-2.0 open source is free and self-hosted; LanceDB Cloud is usage-based (public beta); Enterprise is an annual commit via AWS Marketplace or custom negotiation.
Sources
This page was written from 2 sources.
- 1.lancedb.com — lancedb.comvendor
- 2.lancedb.com — series a fundingvendor
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
QueryStory
Agentic data platform that turns plain-English questions into auditable, decision-ready business narratives
turbopuffer
Vector and full-text search built object-storage-first, at roughly a tenth the cost of RAM-resident vector databases
Ellis
AI-native operations platform that reconciles private credit fund data and runs the close, reporting and monitoring
Monte Carlo
Data and AI agent observability for enterprises running production AI