Local-First AI Foundations
Get a real model running on your own machine and understand what you're actually looking at — weights, quantization, context windows, and where the memory goes.
- Run and serve local models with Ollama and llama.cpp
- Read a model card and predict whether it fits your VRAM
- Choose a quantization level on evidence, not vibes
Capstone A local assistant with a documented hardware profile and measured tokens/sec.