CROWDB is a distributed storage platform for Iceberg tables, AI datasets, and S3 objects. Each access model keeps its own semantics while sharing one storage core for distributed state, placement, protection, streaming, and recovery.
- Iceberg: native catalog and FileIO, implemented.
- S3: core HTTP object operations, implemented.
- Dataset: native access and direct GPU delivery, in design.
Storage bottlenecks move—from disks to CPUs, networks, and data movement—but system boundaries tend to stay. A change that crosses a metadata service, an object gateway, and a separate storage engine can become a negotiation between systems rather than an improvement to one data path.
CROWDB owns enough of that path to change it when workloads and hardware change. S3 objects, Iceberg tables, and planned AI datasets are native access models built over shared infrastructure, not conventions layered on top of one another. The goal is not to claim novelty for Paxos, WALs, trees, or erasure coding; it is to make their contracts agree on durability, placement, bounded buffers, and recovery.
Read the full motivation in Why we’re building CROWDB.
The 0.1.0 Linux amd64 container is available for public evaluation. Use disposable data. Production use and on-disk upgrade compatibility are not supported, and Dataset and direct GPU delivery are not available yet. The container's PyIceberg write and pandas example is in the quick start.
The repository contains the storage implementation, tests, and source design documents. Start with:
- Contributing for the development environment and workflow.
- Documentation index for subsystem designs.
- Backlog for current delivery scope.
- Single-node container for image development and packaging.
The code in this project was written with AI assistance. The architecture, naming, module boundaries, and trade-offs remain human choices. AI is the compiler. The intent is mine.
Licensed under the Apache License 2.0.