Building ยท October 7, 2026
I built a CLI that turns several Kaggle accounts into one GPU cluster
I built a CLI that turns several Kaggle accounts into one GPU cluster.
Compute Pool is open source, a Rust CLI with a Python training runtime. You log in to each account once (up to 8 slots), and one command launches the whole cluster.
From there you can run independent jobs on every GPU, like sweeps or batch inference; data-parallel training with an SSH all-reduce; and pipeline-parallel training, with one model split across nodes.
The hard part was the network. Kaggle sessions don't accept inbound connections, so each worker opens an outbound Cloudflare tunnel and the master connects back over SSH, with host keys pinned. A health check pings every node every 30 seconds, and a node that misses three drops out until it answers again.
Here's what I've tested: LoRA fine-tuning of an 8B Llama on two- and three-node clusters, with checkpoint and resume, and automatic reconnect when a node drops mid-run. The all-reduce example matches a single-node run to floating-point precision.
Here's what I measured: the link runs at about 16 MB/s with a 70 ms round trip, so pipeline training is network-bound and the GPUs mostly wait. It's in the README, and I didn't bury it.
Kaggle's policy is one account per person, so read the risk section before you use it.
The repo is here: iam-saiteja/Compute-Pool. Try it, break it, open an issue.
Written by Thanniru Sai Teja. More writing.