The Hessian of tall-skinny networks is easy to invert
Details
- External ID
- 46638894
- Source
- HN
- Company
- —
- Product
- The Hessian of tall-skinny networks is easy to invert
- Website domain
- github.com
- Launched
- Jan. 15, 2026
- Cohort
- —
- Upvotes
- 31
- Upvotes percentile
- 0.7200263504611331
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
It turns out the inverse of the Hessian of a deep net is easy to apply to a vector. Doing this naively takes cubically many operations in the number of layers (so impractical), but it's possible to do this in time linear in the number of layers (so very practical)!This is possible because the Hessian of a deep net has a matrix polynomial structure that factorizes nicely. The Hessian-inverse-product algorithm that takes advantage of this is similar to running backprop on a dual version of the deep net. It echoes an old idea of Pearlmutter's for computing Hessian-vector products.Maybe this idea is useful as a preconditioner for stochastic gradient descent?
Enrichment
- Theme
- scientific computing and research algorithms
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- neural network optimization research
- Manually corrected
- False
Could you build this?
No This is a mathematical machine learning research project deriving and proving linear-time Hessian inversion algorithms for deep neural networks, representing novel academic computer science research.
What it would actually take: Implementing this requires deriving exact second-order gradient formulations for specific neural network topologies and implementing custom matrix factorizations (e.g. specialized Woodbury matrix identity variants or backpropagation through Kronecker-factored approximations). It requires a research scientist with a PhD-level background in theoretical deep learning, optimization theory, and high-performance linear algebra (C++/CUDA).
Discussion
20 comments analyzed.
Competitors mentioned: BFGS, Conjugate Gradient (CG), SGD-type optimizers, Levenberg-Marquardt, Jacobian-free Newton-Krylov methods
Concerns raised: Situation dependent - no free lunch theorem applies, Hessian often not invertible without regularization, Computational and memory overhead vs. gradient descent, Requires model architecture and tricks to be worthwhile on larger models
Feature requests: Training runs and experimental results, Comparison with Conjugate Gradient methods, Better educational resources explaining Hessians and Jacobians
Competitors
Other products that read as similar to this one — 127 launches clear the similarity bar, closest 8 shown.
Attention rank: #25 of 128 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 68 days after the earliest competitor.
- Neural Inverse Cloud · hn · 2026-07-02 · 7 upvotes · similarity 0.45
- genpark-interior-point-primal-dual-barrier-skill · github · 2026-09-28 · 7 upvotes · similarity 0.43
- Deep learning without gradient descent, 500 layers, no skip connections · hn · 2026-01-07 · 5 upvotes · similarity 0.43
- genpark-graph-convolutional-network-gcn-layer-skill · github · 2026-09-28 · 7 upvotes · similarity 0.42
- GitHub · hn · 2026-01-16 · 6 upvotes · similarity 0.41
- genpark-conjugate-gradient-krylov-subspace-solver-skill · github · 2026-09-09 · 8 upvotes · similarity 0.40
- genpark-conjugate-gradient-krylov-subspace-solver-skill · github · 2026-09-09 · 8 upvotes · similarity 0.40
- Ontological Directed Synthesis Network · ph · 2026-09-22 · 1 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.