Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

A package manager for agent skills with built-in evals

Details

External ID
46900933
Source
HN
Company
—
Product
A package manager for agent skills with built-in evals
Website domain
tessl.io
Launched
Feb. 5, 2026
Cohort
—
Upvotes
7
Upvotes percentile
0.38207547169811323
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I'm Guy, the founder behind Snyk — now building Tessl, a package manager for agent skills.We’ve recently witnessed that most teams still treat skills as static artifacts: markdown files, created or copied from repo to repo.This approach offers a strong initial boost, but quickly creates debt:- Skills are duplicated, and updates never roll out. - Poor quality skills go unseen, misguiding agents instead of helping. - Skill knowledge grows stale, and don’t keep up with the systems and practices they describe.Without a way to evaluate skills, teams have no clear way to understand how good a skill actually is, or if it degraded over time.Our belief is that evaluations are the foundation for having quality skills.With that in mind, I’m glad to announce that Tessl Registry contains review evals for over 2,000 skills, and you can request an evaluation for any public skill.Super excited to be launching this — keen to get your feedback, and looking forward to the many more enhancements in the queue!

Enrichment

Theme
AI agent frameworks and developer tools
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
package manager for ai agent capabilities
Manually corrected
False

Could you build this?

No Building an enterprise-grade package manager and evaluation infrastructure for agent skills requires complex automated sandbox environments, rigorous statistical benchmarking against nondeterministic models, and package registry architecture.

What it would actually take: A real platform requires a distributed registry with semantic versioning and dependency resolution, combined with isolated execution sandboxes (Firecracker microVMs or secure Docker containers) to run agent evals safely. The difficult core is designing an evaluation harness that executes diverse LLMs against agent skills, handles nondeterministic tool-calling evaluations, computes reliable benchmarks across model updates, and manages multi-tenant registry security. It demands deep domain expertise in compiler/package manager architecture, LLM evals, and secure multi-tenant infrastructure.

Discussion

2 comments analyzed.

Concerns raised: Model changes cause prompt/skill regressions over time, Difficulty tracking skill quality and drift across versions

Feature requests: CI integration to pin skill versions and fail builds on eval score drops, Skill quality visibility over time with regression detection

Competitors

Other products that read as similar to this one — 197 launches clear the similarity bar, closest 8 shown.

Attention rank: #124 of 198 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 99 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.