IndexBuild log

Projects

One year, one direction: down the LLM serving stack. Statuses here are honest — repos, benchmarks, and merged PRs land on this page when they exist, not before.

  • 01

    In progress · Jul – Sep 2026

    Relay — LLM serving gateway

    A Go gateway streaming OpenAI-compatible traffic across vLLM replicas on GKE. The question that makes it interesting: does prefix-hash session affinity beat round-robin on cache hit rate and time-to-first-token? Answered with OpenTelemetry traces, Grafana dashboards, and load replays of real multi-turn conversations from LMSYS-Chat-1M — then written up, including why you might just use llm-d or the GKE Inference Gateway instead.

    Go · Kubernetes (GKE) · vLLM · OpenTelemetry · Grafana

  • 02

    Planned · Nov 2026 – Mar 2027

    Ember — an inference engine from scratch

    A from-scratch LLM inference engine for 1–3B open-weight models: paged KV-cache block management, prefix sharing, and chunked-prefill continuous batching, verified token-for-token against the HuggingFace reference. The end state is a benchmark against vLLM at matched batch sizes and a gap analysis of exactly what each optimization buys — and what still separates a student engine from the real thing.

    Python · PyTorch · FastAPI · Docker