Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Cluster-Duck

Cluster-Duck is an experimental federated query coordinator for DuckDB. My example uses three AWS EC2 workers each with their independent DuckDB database installed. The coordinator runs on worker 1, sends SQL to the workers through DuckDB's Quack protocol, starts independent statements concurrently and displays their results and timings.

It is not a true DuckDB cluster. There is no distributed optimiser, automatic sharding, replication, failover, cross-node transaction manager or transparent distributed join. DuckDB stores and processes each worker's data; Quack carries SQL and results; Python performs coordination.The name cluster-duck is just a jokey take on the better known, slightly ruder expression. I give full permission for the folks at DuckDB to steal it when they eventually get round to producing a proper DuckDB cluster.

What is included

  • python-reference/infra/aws/cluster-duck-3-node.yaml: the three-node AWS environment, Quack services, Systems Manager access and automatic shutdown.
  • python-reference/infra/aws/related_cluster_sql.py: the coordinator program installed and run on worker 1.
  • python-reference/examples/aws_queries/: example SQL files for the three related datasets.
  • python-reference/src/cluster_duck/: the tested coordination and SQL API.
  • native-extension/: the separate C++ DuckDB extension scaffold.
  • docs/DEMO_OUTPUT.md: concise output from the AWS test.

Deploy from AWS CloudShell

The template creates three t4g.nano instances by default. Review AWS pricing before deployment; EC2, EBS and public IPv4 charges may apply.

Open AWS CloudShell in the region you want to use, upload the repository or template, then deploy it:

aws cloudformation deploy \
  --region us-east-2 \
  --stack-name cluster-duck-test-v2 \
  --template-file python-reference/infra/aws/cluster-duck-3-node.yaml \
  --capabilities CAPABILITY_NAMED_IAM

Run on the coordinator

In the EC2 console, select worker 1 and open Connect → Session Manager. All Cluster-Duck commands are run in that session:

sudo -i
cluster-duck ready
cluster-duck-sql

Submit any number of independent read-only statements:

cluster-duck-sql \
  --query "worker-1=SELECT sale_status, COUNT(*) FROM sales GROUP BY sale_status" \
  --query "worker-2=SELECT country, COUNT(*) FROM customers GROUP BY country" \
  --query "worker-3=SELECT category, COUNT(*) FROM products GROUP BY category"

Longer statements can be supplied with repeatable --query-file WORKER=PATH arguments. Mutating SQL requires --allow-write. Statements in one invocation run concurrently, so dependent operations belong in separate invocations. Each ad-hoc query-N result can print the exact SQL assigned to that label when the command includes --show-sql.

Documentation

See the coordinator documentation, the AWS deployment notes and the measured demonstration output.

Quack is beta software. See DuckDB's Quack documentation.

About

Experimental federated query coordinator for DuckDB using Quack

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages