DuckDB & Arrow in Action : Scalable Pipelines and Vectorized Queries in Python and Rust
Overview
Reactive Publishing
Modern data engineering is shifting toward lightweight, high-performance in-memory analytics. DuckDB & Arrow in Action provides a practical, hands-on guide to combining the fast SQL query processing of DuckDB with Apache Arrow's zero-copy memory format.
Whether you are building real-time data pipelines or optimizing local analytical processing, this book delivers clear patterns for high-throughput data manipulation without the overhead of massive distributed clusters.
What You Will Learn:
Core Fundamentals: Master DuckDB's columnar execution engine and the memory layout of Apache Arrow.
Zero-Copy Workflows: Pass large datasets seamlessly between DuckDB, PyArrow, and Rust without serialization overhead.
Vectorized Processing: Write high-performance SQL and programmatic queries that leverage SIMD hardware acceleration.
Pipeline Integration: Build efficient ETL and ELT data pipelines in Python and Rust for production environments.
Performance Tuning: Optimize memory usage, indexing, and query execution plans for large-scale datasets.
Designed for software engineers, data engineers, and data scientists, this book equips you with the tools to build faster, cost-effective data infrastructure using modern open-source technologies.
This item is Non-Returnable
Customers Also Bought
Details
- ISBN-13: 9798191549248
- ISBN-10: 9798191549248
- Publisher: Independently Published
- Publish Date: August 2026
- Dimensions: 9 x 6 x 1.54 inches
- Shipping Weight: 1.63 pounds
- Page Count: 622
Related Categories
