LLM infrastructure, mostly around LiteLLM. I like problems where the answer comes from measuring the wire rather than reading the docs, and most of what I publish started as something that was broken in my own cluster.
litellm-mysubs serves Claude Max,
ChatGPT Plus and Google Antigravity subscriptions as ordinary LiteLLM models, so any
OpenAI-compatible client can use them. One line in callbacks: installs it. On
PyPI.
crowdcompute is a demand-validation MVP for community-funded European AI compute: 300 people, one 8xH200 cluster.
cobol4i migrates IBM i ILE COBOL to Java through deterministic gates first, and only then lets a model touch what is left.
ggml-org/llama.cpp#24710 (merged), making tensor-split regex patterns static so they are not recompiled per tensor.
BerriAI/litellm#42467, where a provider reporting a model name the price map does not carry silently bills the turn at zero. The issue has the measurements, including the two wrong theories I had to walk back first.
I also filed BerriAI/litellm#42172, where an Anthropic OAuth token from a Claude Code client reaches third-party deployments and replaces their configured key.
Measure before claiming. Most of my issue reports and commit messages carry the numbers that led to the conclusion, and when the conclusion turns out wrong I correct it in the same thread rather than quietly moving on. Tests are there to fail when the code breaks, not to make coverage look good.
