← All repos

Daft

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

ai-engineeringai-pipelinearrowartificial-intelligencebig-datadata-engineeringdistributeddistributed-computingdistributed-systemsembeddingsetlhuggingfaceicebergmachine-learningmultimodalparquetpythonrayrust
Browse cluster: SQL databases and distributed systems
4,532commits
188contributors
10languages

Tech stack & purpose

Daft is a high-performance data engine built for processing AI and multimodal workloads including images, audio, video, and structured data at scale. The project is developed by Eventual-Inc and is written in Python and Rust, built on technologies like Apache Arrow, Ray, and supporting formats such as Parquet and Iceberg. The repository includes benchmarks demonstrating Daft's performance advantages over Ray Data and Spark across tasks like audio transcription, document embedding, image classification, and video object detection, along with examples of native extensions and Kubernetes deployment tooling.

Languages

Rust
46.5%
Python
35.9%
Jupyter Notebook
16.4%
TypeScript
1.1%
Shell
0.1%
Makefile
0.0%
CSS
0.0%
Go Template
0.0%
Dockerfile
0.0%
JavaScript
0.0%

Contributors (top 30 of 188)