Local AI Lab

Role
Design & build
Type
Personal
Focus
AI infrastructure

A home AI environment on consumer hardware for running LLMs, small language models, speech models and embeddings locally — to learn what these workloads really cost when there is no API to hide behind.

The hardware

GPU node

Ryzen 7 5700X, RTX 4060, 32 GB

Ollama, Whisper, embeddings

Always-on node

Raspberry Pi 5, 8 GB

MQTT, Piper TTS, monitoring

Questions it is built to answer

  • What fits in 8 GB of VRAM, and what does it cost in latency?
  • When is a small model good enough for the task?
  • What should the system do when the GPU box is switched off?

Stack

  • Ollama
  • SLMs
  • Embeddings
  • Docker
  • Wyoming
  • LAN + TLS