Invariance
Letting a language model restyle a live product is easy to demo and hard to ship, because the failure mode is an unreadable page in front of a paying tenant. Invariance is the version that can actually go to production: the model proposes, and a set of invariants decides.
The LLM is confined to a proposal stage bounded by brand and accessibility invariants, which caught 100% of violations before publish. Nothing reaches a user without clearing contrast and brand constraints first.
The control plane runs on Neon Postgres with Drizzle, keeping schema as the source of truth for manifests, versioned per-tenant themes, rollback, audit trails, and prompts.
Each theme ships as an immutable, content-addressed Cloudflare R2 artifact behind the CDN, with a short-TTL KV pointer on the hot path. Source is never rewritten at runtime, so a bad theme is a pointer flip away from being undone.
BriefCase
LLMs can read a contract clause and reason about its risk, but running one over every clause in a deal is expensive, and legal text is exactly the kind of data that cannot leave a client's infrastructure. The question our four-person team asked: can training-time reasoning from a locally hosted open-weight teacher substitute for in-domain pretraining in a small encoder?
We built the experiment on CUAD, mapping its 41 clause types onto a 10-class risk taxonomy and generating teacher rationales for all 10.5k training clauses. The design is a 2x2 crossing encoder type, DistilBERT against Legal-BERT, with training variant, vanilla against rationales generated by a locally hosted Llama-3.1-8B-Instruct teacher. Rationales are concatenated at fine-tuning time and stripped at inference, so every model classifies from the clause alone and the CoT variants cost nothing extra to serve.
The best student reached 0.896 macro-F1. But the result that matters is the gap underneath it: every fine-tuned student beat its own teacher's direct zero-shot inference by roughly 47 macro-F1 points, 0.87 and above against the teacher's 0.40. A 110M-parameter encoder outclassed the 8B model it learned from, which says the bottleneck is teacher competence on the task, not the distillation framework.
CoT's own effect was small and honest: +0.26 macro-F1 points on DistilBERT and +0.33 on Legal-BERT, both inside the single-seed noise band, while in-domain pretraining held a robust 2-point gap that CoT never closed. So the answer to the original question, at 8B teacher scale, is no. We report it as a negative result rather than rounding it up.
The whole pipeline is on-premise and reproducible end to end in 2.5 GPU-hours on a single A100, at $0 in external API fees.
GARCH BTC
GARCH models are the standard tool for volatility in traditional finance and underused in crypto, where the volatility clustering they were designed for is at its most extreme. This project applies them to Bitcoin, pairing the classical models with hybrid and Bayesian-tuned variants.
I implemented GJR-GARCH from scratch with Numba for performance, then benchmarked it against GARCH(1,1) and EGARCH baselines on MSE, AIC, and log-likelihood over Bitcoin log returns, with Optuna driving hyperparameter search and LSTM and Transformer hybrids layered on top.