Open source ยท August 29, 2026
26.4K stars, and a little piece of my code is now part of colibri
26.4K+ stars. 2.9K+ forks. And somehow, a little piece of my code is now part of it.
colibri is an open-source inference engine for running large language models locally, including models that are normally far beyond what you'd expect to run on consumer hardware.
A few weeks ago, I was going through its DeepSeek V4 indexer and found a small performance bottleneck that was happening over and over again: 12 dynamic malloc/free calls on every batch.
So I tried a simpler approach. Instead of allocating and freeing memory every time, I introduced a persistent 32-byte-aligned scratch arena.
The numbers surprised me: 0.1539 ms โ 0.0006 ms per batch call. That's a 99.6% reduction in allocation overhead per step.
The PR got merged into dev. And today, I saw it in the colibri v1.9.0 release. PR #1179. colibri v1.9.0.
It's a small contribution. But seeing my code go from "I think I can improve this" to actually becoming part of a project feels pretty good.
Written by Thanniru Sai Teja. More writing.