# Thanniru Sai Teja > AI Systems Researcher and Full-Stack Engineer. 4th-year B.Tech CSE (Data Science) student at BVRIT Narsapur, Telangana, India (2023-2027, CGPA 8.85/10). Builds attention-free sequence architectures and memory systems on consumer hardware, and ships performance patches upstream to curl, nvm, colibri, llmfit and Soup. Open to SWE and ML roles. Core idea: most AI research assumes a room full of expensive servers. Sai Teja builds the same architectures to run on consumer hardware like a single 4GB GPU, because a good idea should prove itself under real hardware constraints first. ## Contact - Email: iamsaitejathanniru@gmail.com - Phone: +91 9133708990 - Location: Greater Hyderabad Area, Narsapur, India - LinkedIn: [linkedin.com/in/thannirusaiteja](https://linkedin.com/in/thannirusaiteja) - GitHub: [github.com/iam-saiteja](https://github.com/iam-saiteja) - [Portfolio](https://thannirusaiteja.dev/): the human-facing site, with a contact form - [Shell](https://thannirusaiteja.dev/shell): a terminal over the same content. Commands: help, whoami, ls, cat , recall , stats, activity, contact ## Open source - [llmfit](https://github.com/AlexsJones/llmfit) (Rust, 38k stars): 46% faster model search. Short-circuiting field checks in the TUI's apply_filters(), replacing ~2,000 heap allocations per keypress. ~950 µs → 510 µs per filter pass. [PR #934](https://github.com/AlexsJones/llmfit/pull/934) - [nvm](https://github.com/nvm-sh/nvm) (Shell, 95k stars): 141x faster nvm_tree_contains_path. Replaced repeated dirname calls with an in-memory POSIX parent walk. ~350 process forks down to 0. 16,274 ms → 115 ms for 50 calls. [PR #3912](https://github.com/nvm-sh/nvm/pull/3912), [v0.40.8](https://github.com/nvm-sh/nvm/releases/tag/v0.40.8) - [Soup](https://github.com/MakazhanAlpamys/Soup) (Python, 8.4k stars): 30% faster dataset validation. A single-pass validator with hashable row signatures and no intermediate allocations. 0.1397 s → 0.0971 s on 20,000 rows. [PR #771](https://github.com/MakazhanAlpamys/Soup/pull/771), [v0.75.0](https://github.com/MakazhanAlpamys/Soup/releases/tag/v0.75.0) - [curl](https://github.com/curl/curl) (C, 43k stars): 36% faster Base64 decode. Rewrote the decode lookup table and quantum loop. Benchmarked by maintainer Daniel Stenberg over 10M operations and committed upstream. 353.1 ns → 226.43 ns per decode. [commit e7ea5d1](https://github.com/curl/curl/commit/e7ea5d10b31fcdde4c69c867793b5a396b975849), [PR #22903](https://github.com/curl/curl/pull/22903), [perf history](https://curl.se/perf/index.html#b64dec) - [colibri](https://github.com/JustVugg/colibri) (C, 41k stars): 99.6% less allocation overhead. Swapped 12 malloc/free calls per batch in the DeepSeek V4 indexer for a persistent 32-byte-aligned scratch arena. 0.1539 ms → 0.0006 ms per batch call. [PR #1179](https://github.com/JustVugg/colibri/pull/1179), [v1.9.0](https://github.com/JustVugg/colibri/releases/tag/v1.9.0) - [Windows-MCP](https://github.com/CursorTouch/Windows-MCP) (Python, 8.5k stars): 52% faster coordinate resolution. Replaced an N+1 lookup in MultiEdit and MultiSelect with bulk resolution: 37.45% faster at 3,000 resolutions, 52.09% at 35,000. My first open-source contribution. [PR #219](https://github.com/CursorTouch/Windows-MCP/pull/219) ## Research - [HGDM](https://github.com/iam-saiteja/HGDM-Hierarchical-Gated-Delta-Memory): Hierarchical Gated Delta Memory. 100% attention-free byte-level sequence model. 1.006B parameters pushed to 1.46B training tokens on a 4GB VRAM GPU. O(1) inference memory across 20x context scaling. Built with Nitro (custom Triton fused CUDA scan kernel). Features Hierarchical Temporal Decimation and Factored Bilinear State Highways. - [NCM](https://github.com/iam-saiteja/NCM): Native Cognitive Memory. Tensor-based episodic memory with four-dimensional retrieval: semantic, emotional, state-conditioned, and temporal. The same query can recall different memories depending on the system's state. Validated across 18 empirical experiments. Now testing contradiction-aware retrieval and scaling to 100k memories. - [Small-model RAG study](https://github.com/saiharish587/RAG-Research): LFM2-350M vs LFM2-700M, paper under review. 60,000 controlled evaluations on 600 HotpotQA questions across five RAG architectures. The 350M model beat the 700M one on Token F1 (11.07% vs 7.27% with Naive RAG). Even with oracle evidence the 350M model reached only 6.22% Token F1, so the open question is how well small models use context. Co-authored with Sai Harish Puchala. - ZS-ISAB (not public yet): Linear-complexity attention for TabPFN. Scales TabPFN's attention from quadratic to linear with seeded anchor selection, an online softmax accumulator, and a zero-shot attention mask, enabling 500K+ row inference on 4GB VRAM. In iteration toward a TMLR submission; running a 91-dataset TabZilla benchmark against tree-based baselines. - NOUS (not public yet): AI-native OS in Rust, architecture phase. A bare-metal operating system for x86_64 where memory is addressed by semantic embeddings rather than pointers, and the resident kernel is a recurrent neural model. Early design work: UEFI framebuffer dashboard, capability-gated intent syscalls, and NCM as the memory core. - [HGDM preprint on Zenodo](https://zenodo.org/records/20286026): records/20286026, accepted into Advanced Research in Artificial Intelligence Systems (ARAIS). ## Engineering - [Compute Pool](https://github.com/iam-saiteja/Compute-Pool): Several Kaggle accounts, one GPU cluster. A Rust CLI with a Python training runtime that pools up to 8 Kaggle accounts into one cluster: independent jobs, data-parallel training with an SSH all-reduce, and pipeline-parallel training. Workers open outbound Cloudflare tunnels and the master connects back over SSH with pinned host keys. Tested with LoRA fine-tuning of an 8B Llama on two- and three-node clusters, with checkpoint, resume and auto-reconnect. - [Event Ticket Booking System](https://github.com/iam-saiteja/Event-Ticket-Booking-System): Event & ticketing platform, 300+ users. End-to-end ticketing platform with real-time booking, QR code ticket generation, PDF export, and a Cloud Functions backend. Served 300+ real users across campus events. Handles concurrent bookings with Firestore transactions, instant QR verification, and automated email receipt delivery. - STUD Recruitment Platform (private): Full-stack recruitment portal, built in 3 days. Built and shipped a production recruitment platform for Stud Entertainments in 3 days flat, with admin tracking, automated interview scheduling, and auth. Application tracking, a notification system, and multi-role dashboards for candidates and admins. - [Orbius AI](https://github.com/saiharish587/Orbius-AI): AI development environment, built in a 2-day buildathon. Built at the NxtWave × OpenAI Academy state-level buildathon. Multi-Agent Mode lets GPT, Gemini, Perplexity and other models work on the same problem together. Led development: 17 commits and 12,000+ lines in two days, presented to the judges in Hyderabad. - [FormatForge](https://github.com/iam-saiteja/FormatForge): AI product image pipeline. Upload one photo, get platform-ready assets for Amazon, Flipkart, Instagram, Spotify with natural language in-place AI re-editing. Powered by Gemini 2.5 Flash Image. Live demo at formatforge.streamlit.app. - [Chaff App](https://github.com/iam-saiteja/Chaff): Agricultural marketplace platform. Web-based farmer-to-manufacturer stubble marketplace turning agricultural waste into valuable industrial resources. Connects farmers directly with straw/chaff buyers to prevent crop burning while monetizing waste. ## Experience - Open-Source Performance Contributor, curl, nvm, colibri, llmfit, Soup, Windows-MCP, Remote (Apr 2026 - Present): curl: 36% faster Base64 decode (353.1 ns → 226.43 ns), benchmarked and committed by the maintainer. nvm: ~141x faster nvm_tree_contains_path() with ~350 process forks down to 0, released in v0.40.8. colibri: 99.6% less allocation overhead in the DeepSeek V4 indexer, released in v1.9.0. llmfit, Soup and Windows-MCP: 30-52% faster hot paths, all merged upstream. - Technical Lead, Stud Entertainments, Hyderabad, India (Nov 2024 - Jun 2025): Built & shipped a full-stack recruitment platform in 3 days flat using TypeScript, React & Firebase. Built user-friendly admin dashboards, application tracking, automatic interview scheduling, auth, and notification systems. Conducted technical interviews, hired & led an engineering team, establishing delivery workflows that eliminated production delays. Developed event management automation tools improving operational efficiency by 25% through workflow optimization. Developed machine learning models related to film & media workflow acceleration. - Independent AI Systems Researcher, Self-Directed Research Track, Remote / Hyderabad (2023 - Present): Designed HGDM: attention-free 1B parameter sequence architecture with O(1) inference memory footprint on a single 4GB VRAM GPU. Author of Zenodo published preprint (records/20286026), accepted into Advanced Research in Artificial Intelligence Systems (ARAIS). Architected Native Cognitive Memory (NCM) episodic memory system validated across 18 empirical experiments. Co-authored a 60,000-evaluation small-model RAG study (LFM2-350M vs 700M), now under review. Developed ZS-ISAB (scaling TabPFN attention from quadratic to linear) and PAS-Offload LLM inference engine. - Core Member & Technical Lead, Chalana Chithram Club, BVRIT Narsapur (2023 - Present): Technical: Built college event platforms serving 300+ active users, ticketing portals, and live festival dashboards. Creative & Leadership: Directed 'Our Story' short film; managed complete technical setup and coordination for university film festival. ## Skills - Programming Languages: Python, TypeScript, JavaScript, C, C++, Rust, Java, SQL, Shell - Cloud & DevOps: AWS, Google Cloud, IBM Cloud, Azure, Docker, Linux (Ubuntu), Git, Firebase - AI/ML & Data Science: PyTorch, Triton / CUDA, Scikit-learn, Transformers, Generative AI, LLMs, Prompt Engineering, Data Mining, Predictive Analytics, NumPy, Matplotlib, Power BI - Web & Backend: React, Node.js, REST APIs, HTML5 / CSS3, Tailwind CSS, PostgreSQL, MySQL, Firestore - Security & Governance: Cybersecurity, Cloud Security, Network Security, Ethical Hacking, Threat Analysis, AI Ethics & Governance - Languages Spoken: Telugu (Native), English (Highly Proficient), Hindi (Proficient) ## Certifications - [Data Structures & Algorithms](https://smartinterviews.in/certificate/59ba20c7), Smart Interviews (2025) - [Python Essentials 1 & 2](https://www.credly.com/badges/2aa46df9-1837-4e4e-b934-ac31a2127078/linked_in_profile), Cisco (2023) - [Google AI Essentials](https://www.coursera.org/account/accomplishments/verify/CSETQHEXL9WC), Google (2024) - [Applied Data Science with Python](https://www.credly.com/badges/a060f074-5aee-4413-a349-97108b1da71e), IBM (2024) - [Career Essentials in Generative AI](https://www.linkedin.com/learning/certificates/aaca8a5d82966153eee25dbcd3660c6a82907757efcdd990739bf8d2a274c35a), Microsoft (2024) - [Career Essentials in Data Analysis](https://www.linkedin.com/learning/certificates/e9aea547bdad1fef70206793b6b8b747fc7cbb83837e37b3eb87260f08efd0ea), Microsoft (2024) - [Google Cybersecurity Specialization](https://www.coursera.org/account/accomplishments/specialization/NP9J5GQIZ816), Google (2023) - [Ethical Hacker Nanodegree](https://www.udacity.com/certificate/e/ea9f6164-f938-11ef-9e6e-4f7e68af4e8d), Udacity (2025) ## Writing - [I built a CLI that turns several Kaggle accounts into one GPU cluster](https://thannirusaiteja.dev/writing/compute-pool) (2026-10-07, Building): Compute Pool is a Rust CLI with a Python training runtime. One command launches the whole cluster. - [One of nvm's shell functions just got ~140x faster](https://thannirusaiteja.dev/writing/nvm-141x) (2026-09-28, Open source): nvm_tree_contains_path() went from 16,274 ms to 115 ms for 50 calls, and from ~350 process forks to 0. Merged and released in nvm v0.40.8. - [30% faster dataset validation in Soup](https://thannirusaiteja.dev/writing/soup-validator) (2026-09-15, Open source): People training models on laptops just got 30% faster dataset validation on Soup. Shipped in v0.75.0. - [One small part of curl is now 36% faster](https://thannirusaiteja.dev/writing/curl-base64) (2026-09-11, Open source): I optimized curl's Base64 decode lookup table and rewrote the loop. Bagder benchmarked it, committed it himself, and changed his mind about the approach. - [More bots than humans have been visiting my portfolio](https://thannirusaiteja.dev/writing/portfolio-bots) (2026-09-06, Building): It crossed 17K+ pageviews, so I checked who was visiting. Roughly 89% automated. - [Thank you, Sublime HQ. You made my childhood memorable](https://thannirusaiteja.dev/writing/sublime-text) (2026-09-04, Personal): My first laptop had 4GB RAM and no GPU. Sublime Text is where I learned Python, HTML, JavaScript, PHP and OpenCV. - [26.4K stars, and a little piece of my code is now part of colibri](https://thannirusaiteja.dev/writing/colibri-arena) (2026-08-29, Open source): A persistent scratch arena in colibri's DeepSeek V4 indexer: 0.1539 ms → 0.0006 ms per batch call. Released in v1.9.0. - [A better system around a small model beat a bigger model](https://thannirusaiteja.dev/writing/small-models-rag) (2026-08-27, Research): 60,000 controlled evaluations on 600 HotpotQA questions: LFM2-350M consistently outperformed LFM2-700M on Token F1. Paper under review. - [What if attention isn't all you need?](https://thannirusaiteja.dev/writing/attention-isnt-all) (2026-08-14, Research): Hybrid models are asking how much of sequence modeling needs attention. HGDM is my attention-free experiment: 1.006B parameters, 1.46B training tokens. - [My first open-source contribution](https://thannirusaiteja.dev/writing/windows-mcp) (2026-08-13, Open source): An N+1 lookup in Windows-MCP's hot path, replaced with bulk coordinate resolution: 37-52% faster. - [AI memory needs to be more than a vector database](https://thannirusaiteja.dev/writing/ncm-memory) (2026-08-12, Research): Why should two memories be the same just because they mean the same thing? That question became NCM. - [Two days, one buildathon, and an idea that resonated later](https://thannirusaiteja.dev/writing/orbius-buildathon) (2026-08-03, Building): At the NxtWave × OpenAI Academy buildathon we built Orbius AI, with a multi-agent mode where several models work on one problem. Then llm-council arrived. - [Rebuilt my portfolio from scratch](https://thannirusaiteja.dev/writing/portfolio-rebuild) (2026-08-02, Building): React, TypeScript, Vite and Cloudflare Pages. Live GitHub stats, a research section, and a small Q&A widget. - [50+ Rejections, 20 Years Old, and the Space Between Building and Prepping](https://thannirusaiteja.dev/writing/50-rejections) (2026-08-01, Personal): Yesterday was my birthday. Ironically, I got an email from Google. Rejected. That pushes me past 50 off-campus rejections... ## Optional - [Full text of every essay](https://thannirusaiteja.dev/llms-full.txt)