IndexBuild log
Projects
One year, one direction: down the LLM serving stack. Statuses here are honest — repos, benchmarks, and merged PRs land on this page when they exist, not before.
01
In progress · Jul – Sep 2026
Relay — LLM serving gateway
A Go gateway streaming OpenAI-compatible traffic across vLLM replicas on GKE. The question that makes it interesting: does prefix-hash session affinity beat round-robin on cache hit rate and time-to-first-token? Answered with OpenTelemetry traces, Grafana dashboards, and load replays of real multi-turn conversations from LMSYS-Chat-1M — then written up, including why you might just use llm-d or the GKE Inference Gateway instead.
Go · Kubernetes (GKE) · vLLM · OpenTelemetry · Grafana
02
Planned · Nov 2026 – Mar 2027
Ember — an inference engine from scratch
A from-scratch LLM inference engine for 1–3B open-weight models: paged KV-cache block management, prefix sharing, and chunked-prefill continuous batching, verified token-for-token against the HuggingFace reference. The end state is a benchmark against vLLM at matched batch sizes and a gap analysis of exactly what each optimization buys — and what still separates a student engine from the real thing.
Python · PyTorch · FastAPI · Docker