Using 512gb of ram for a 22gb remote file does feel a bit weird for a benchmark but maybe they couldn’t get a large number of cores without lots of memory?
Most cloud providers start with a 2:1 ratio of memory in GiB to CPU cores and go up from there. Databases also are the most common workload for large-memory systems because they benefit so much from large buffer caches.
Brilliant deep dive into asynchronous I/O and execution thread models. Essential reading for high-performance data engineering.
Using 512gb of ram for a 22gb remote file does feel a bit weird for a benchmark but maybe they couldn’t get a large number of cores without lots of memory?
Most cloud providers start with a 2:1 ratio of memory in GiB to CPU cores and go up from there. Databases also are the most common workload for large-memory systems because they benefit so much from large buffer caches.
Deep dive into asynchronous I/O architectures like this is pure engineering gold for high-performance data processing. Excellent breakdown.
I wonder how this would work in trying to parallelize the worker threads (multiple duckdb instances) coordinating them via Quack.
Ducks all the way down!
DuckDB is trending towards becoming a query engine, specifically the fastest analytical query engine. This is very good.
Do they have SSL updates yet? Signing is great, but using https means not fighting firewalls to start a job
This is such a long waited feature!