C discrete event SIM w stackful coroutines runs 45x faster than SimPy
Details
- External ID
- 46872818
- Source
- HN
- Company
- —
- Product
- C discrete event SIM w stackful coroutines runs 45x faster than SimPy
- Website domain
- github.com
- Launched
- Feb. 3, 2026
- Cohort
- —
- Upvotes
- 69
- Upvotes percentile
- 0.8578167115902965
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi all,I have built Cimba, a multithreaded discrete event simulation library in C.Cimba uses POSIX pthread multithreading for parallel execution of multiple simulation trials, while coroutines provide concurrency inside each simulated trial universe. The simulated processes are based on asymmetric stackful coroutines with the context switching hand-coded in assembly.The stackful coroutines make it natural to express agentic behavior by conceptually placing oneself "inside" that process and describing what it does. A process can run in an infinite loop or just act as a one-shot customer passing through the system, yielding and resuming execution from any level of its call stack, acting both as an active agent and a passive object as needed. This is inspired by my own experience programming in Simula67, many moons ago, where I found the coroutines more important than the deservedly famous object-orientation.Cimba turned out to run really fast. In a simple benchmark, 100 trials of an M/M/1 queue run for one million time units each, it ran 45 times faster than an equivalent model built in SimPy + Python multiprocessing. The running time was reduced by 97.8 % vs the SimPy model. Cimba even processed more simulated events per second on a single CPU core than SimPy could do on all 64 cores.The speed is not only due to the efficient coroutines. Other parts are also designed for speed, such as a hash-heap event queue (binary heap plus Fibonacci hash map), fast random number generators and distributions, memory pools for frequently used object types, and so on.The initial implementation supports the AMD64/x86-64 architecture for Linux and Windows. I plan to target Apple Silicon next, then probably ARM.I believe this may interest the HN community. I would appreciate your views on both the API and the code. Any thoughts on future target architectures to consider?Docs: https://cimba.readthedocs.io/en/latest/Repo: https://github.com/ambonvik/cimba
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- discrete event simulation engine
- Manually corrected
- False
Could you build this?
No Implementing a high-performance discrete event simulation engine in C using POSIX pthreads and custom asymmetric stackful coroutines requires low-level systems programming and assembly/context-switching expertise.
What it would actually take: The engine requires low-level C systems programming with custom user-space context switching (like makecontext/swapcontext or custom assembly routines to swap CPU register states and maintain separate call stacks). The hard part is synchronization: designing lock-free event priority queues (e.g., calendar queues or pairing heaps) that operate deterministically across multithreaded simulation boundaries without lock contention. This demands deep expertise in operating system internals, CPU memory models, and systems performance optimization.
Discussion
18 comments analyzed.
Competitors mentioned: Bunki, SimPy, Mojo, getcontext/setcontext, swapcontext
Concerns raised: Stack overflow and function return safety with fixed-size stacks, Context switch overhead is time-consuming part of performance, Sanitizers (ASan/UBSan/valgrind) compatibility unclear, Brittleness of stackful coroutine implementations, Intel CET special handling may be needed
Feature requests: Support for additional CPU architectures beyond x86-64, Optimize context switches further with compiler attributes like __attribute__((preserve_none)), Growing/dynamic stack support instead of fixed-size stacks, Better sanitizer support for debugging user code
Competitors
Other products that read as similar to this one — 46 launches clear the similarity bar, closest 8 shown.
Attention rank: #10 of 47 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 97 days after the earliest competitor.
- Writing a C++20M:N Scheduler from Scratch (EBR, Work-Stealing) · hn · 2026-02-17 · 18 upvotes · similarity 0.42
- Forkrun · hn · 2026-03-27 · 151 upvotes · similarity 0.41
- Runloom · hn · 2026-07-10 · 47 upvotes · similarity 0.41
- CCo · hn · 2026-08-03 · 5 upvotes · similarity 0.40
- Cicada · hn · 2026-01-30 · 57 upvotes · similarity 0.40
- RunMat · hn · 2025-12-02 · 21 upvotes · similarity 0.38
- The Type 1 Civilization Toolkit · ph · 2026-09-24 · 1 upvotes · similarity 0.38
- Coderive · hn · 2025-12-22 · 8 upvotes · similarity 0.38
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.