Autonomous recovery for distributed training jobs
Details
- External ID
- 46812909
- Source
- HN
- Company
- —
- Product
- Autonomous recovery for distributed training jobs
- Website domain
- tensorpool.dev
- Launched
- Jan. 29, 2026
- Cohort
- —
- Upvotes
- 12
- Upvotes percentile
- 0.5586297760210803
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi HN! We’re TensorPool. We help companies access and optimize large scale compute for training foundation models.The ProblemIt’s been almost a year since we’ve finished YC, and we’ve just crossed 100,000 multinode training GPU hours run on our platform.On those training runs, we’ve seen countless 3am job crashes because of issues like an Xid error from a flaky GPU or an S3 timeout that corrupted a checkpoint save. By the time you wake up and notice, you've lost 8+ hours of compute. You scramble to diagnose the issue, manually restart from the last checkpoint, and hope it doesn't happen again. Rinse and repeat.For training runs that take days to weeks, this constant babysitting is exhausting and expensive. The research iteration cycles lost can also make or break a model release (especially for short reservations).What We BuiltThis agent monitors your training jobs and autonomously recovers them when things go wrong. It works with Kubernetes, Slurm, and TensorPool Jobs.We originally built the TensorPool Agent as an internal tool to help us debug failures with our own customers. Over time, we realized its performance was so good that we could automate the entire triage process. We're now releasing a public beta for people to use.Best case: The TensorPool Agent detects the failure, diagnoses the root cause, fixes it, and restarts your job from the last checkpoint – all while you sleep ;)Worst case: If the TensorPool agent can't fix the issue automatically, it delivers a preliminary RCA and a list of actions it attempted, giving you a head start on debugging.How It Works1) Registration – You provide credentials to your job scheduler via our dashboard. Perms are granted on a whitelist basis; you explicitly control what actions the agent can take.2) Monitoring – The agent continuously monitors your job for failure conditions.3) Recovery – On failure, the agent analyzes logs and attempts to diagnose the issue. If successful, it restarts the job from the last checkpoint and resumes monitoring. If not, you get an alert with full context.Target Failure ModesThe agent is specifically designed for runtime errors that occur deep into training, like:- CUDA OOM: Memory leaks, gradient explosions- Xid errors: GPU hardware faults (Xid 79, 63, 48, etc.)- Distributed communication failures: NCCL timeouts, rank failures- Storage I/O errors: Checkpoint corruption- Network issues: S3 request timeouts on mounted object storage
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- recovery for distributed training jobs
- Manually corrected
- False
Could you build this?
No Autonomous recovery for multi-node distributed AI training requires deep systems programming, NCCL/InfiniBand network diagnostics, and GPU hardware cluster orchestration at scale.
What it would actually take: Requires an agent monitoring low-level GPU metrics, NVLink/NCCL error logs, and InfiniBand status across hundreds of bare-metal nodes. The system must integrate directly with distributed frameworks (PyTorch DDP, Megatron-LM) to execute rapid zero-bubble elastic checkpointing and live node-swapping without killing entire jobs. This requires deep systems engineering, kernel-level debugging, and access to expensive multi-node GPU clusters.
Discussion
3 comments analyzed.
Concerns raised: Silent failures hard to detect (NCCL hangs, gradient explosions without crashes), Relies on explicit error logs rather than behavioral signals, Difficulty distinguishing job running normally vs. running but degraded
Feature requests: Automatic detection of silent failures and hung processes, Business metrics-based alerting instead of just raw measurements, Recovery/restart mechanisms for long-running training jobs
Competitors
Other products that read as similar to this one — 116 launches clear the similarity bar, closest 8 shown.
Attention rank: #57 of 117 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 86 days after the earliest competitor.
- First autonomous ML and AI engineering Agent · hn · 2026-01-22 · 5 upvotes · similarity 0.47
- OpenTiger · hn · 2026-02-22 · 11 upvotes · similarity 0.43
- Agent framework that generates its own topology and evolves at runtime · hn · 2026-02-11 · 107 upvotes · similarity 0.41
- SF Tensor - Infrastructure for the Era of Large-Scale AI Training ⚡ · yc · 2025-11-04 · 23 upvotes · similarity 0.39
- Mini-AGI · hn · 2026-09-21 · 277 upvotes · similarity 0.38
- Superserve · hn · 2026-07-21 · 9 upvotes · similarity 0.38
- Recurse · hn · 2026-09-25 · 8 upvotes · similarity 0.38
- Ontological Directed Synthesis Network · ph · 2026-09-22 · 1 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.