Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Virtual SLURM HPC cluster in a Docker Compose

Details

External ID
45982280
Source
HN
Company
—
Product
Virtual SLURM HPC cluster in a Docker Compose
Website domain
github.com
Launched
Nov. 19, 2025
Cohort
—
Upvotes
57
Upvotes percentile
0.8482532751091703
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

I'm the main developer behind vHPC, a SLURM HPC cluster in a docker compose.As part of my job, I'm working on a software solution that needs to interact with one of the largest Italian HPC clusters (Cineca Leonardo, 270 PFLOPS). Of course developing on the production system was out of question, as it would have led to unbearably long feedback loops. I thus started looking around for existing containerised solutions, which were always lacking some key ingredient in order to suitably mock our target system (accounting, MPI, out of date software, ...).I thus decided that it was worth it to make my own virtual cluster from scratch, learning a thing or two about SLURM in the process. Even though it satisfies the particular needs of the project I'm working on, I tried to keep vHPC as simple and versatile as possible.I proposed the company to open source it, and as of this morning (CET) vHPC is FLOSS for others to use and tweak. I am around to answer any question.

Enrichment

Theme
interactive simulations and creative experiments
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
local slurm hpc cluster in docker
Manually corrected
False

Could you build this?

Partial The basic Docker Compose and container setup is achievable, but accurately emulating a multi-node SLURM HPC environment with realistic networking, munge auth, and cgroup resource tracking requires deep Linux and HPC admin knowledge.

What it would actually take: The stack comprises containerized SLURM daemons (slurmctld, slurmd, slurmdbd), MySQL for accounting, MUNGE for inter-node cryptographic authentication, and shared NFS mounts. The hard part is managing container systemd/init systems, multi-node cgroups for CPU/memory isolation, and emulating MPI networking without bare-metal hardware. Requires seasoned Linux systems administration and HPC cluster engineering expertise.

Discussion

17 comments analyzed.

Competitors mentioned: SGE (Sun Grid Engine), LSF (Load Sharing Facility), OpenOnDemand, AWS ParallelCluster, NVIDIA Base Command Manager

Concerns raised: SLURM uses 4 ports per job, limiting simultaneous jobs to thousands before controller TCP port exhaustion, Long queuing times and 2FA requirements on shared reference clusters, SLURM poorly suited for automated testing workflows with large job counts, Command-line tools feel outdated with very small default column widths, Array jobs and fire-and-forget submission are janky workarounds for large workloads

Feature requests: Modern replacement to SLURM with better handling of large job counts, Better support for programmatic job submission and output handling, Improved command-line interface with sensible defaults, Native support for nested jobs without workarounds

Competitors

Other products that read as similar to this one — 29 launches clear the similarity bar, closest 8 shown.

Attention rank: #6 of 30 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 2 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.