Open source ยท August 29, 2026

26.4K stars, and a little piece of my code is now part of colibri

26.4K+ stars. 2.9K+ forks. And somehow, a little piece of my code is now part of it.

colibri is an open-source inference engine for running large language models locally, including models that are normally far beyond what you'd expect to run on consumer hardware.

A few weeks ago, I was going through its DeepSeek V4 indexer and found a small performance bottleneck that was happening over and over again: 12 dynamic malloc/free calls on every batch.

So I tried a simpler approach. Instead of allocating and freeing memory every time, I introduced a persistent 32-byte-aligned scratch arena.

The numbers surprised me: 0.1539 ms โ†’ 0.0006 ms per batch call. That's a 99.6% reduction in allocation overhead per step.

The PR got merged into dev. And today, I saw it in the colibri v1.9.0 release. PR #1179. colibri v1.9.0.

It's a small contribution. But seeing my code go from "I think I can improve this" to actually becoming part of a project feels pretty good.

Written by Thanniru Sai Teja. More writing.