Local AI Lab
A home AI environment on consumer hardware for running LLMs, small language models, speech models and embeddings locally — to learn what these workloads really cost when there is no API to hide behind.
The hardware
GPU node
Ryzen 7 5700X, RTX 4060, 32 GB
Ollama, Whisper, embeddings
Always-on node
Raspberry Pi 5, 8 GB
MQTT, Piper TTS, monitoring
Questions it is built to answer
- What fits in 8 GB of VRAM, and what does it cost in latency?
- When is a small model good enough for the task?
- What should the system do when the GPU box is switched off?