Local AI Coaching

Run AI Models
On Your Own Computer

Set up Ollama, pick the right model for your hardware, and build private AI workflows that work without internet — no cloud subscription, no data leaving your machine.

Book a 60-Minute Session — $150
Zoom · 60 minutes · Tell me your hardware when you book
🖥️ Runs local models daily — M-series Mac + custom Linux rig
🏢 Ex-Amazon, 6 years Senior TAM, 12+ AWS certs
☁️ Knows cloud vs local tradeoffs from both sides
👥 30,000+ students on Udemy

Who This Is For

What We Can Work On

Ollama Setup & Model Management

Install, configure, pull the right models, set up the local API, and wire it to Open WebUI or your own app. From zero to running in one session.

Model Selection for Your Hardware

Which model to run at which quantization level for your RAM and GPU. Stop guessing — get the right model for your actual hardware.

LM Studio & Open WebUI

Set up a ChatGPT-like UI on your own machine. LM Studio for a GUI-first experience, Open WebUI for a full-featured local chat interface.

Privacy Workflows

Build document Q&A, summarization, and analysis pipelines that run entirely on your hardware — for contracts, patient notes, financial records.

Local Coding Assistant

Set up Continue.dev or Cursor with a local Ollama backend so your code never leaves your machine. Offline coding assistance with no subscription.

Cloud vs Local Decision

Honest assessment of when local AI makes sense vs when cloud is better. Don't waste money on GPU RAM you don't need for your use case.

What Your Hardware Can Run

Quick reference — we'll calibrate to your exact setup in the session.

Hardware
Best Models
Use Case Fit
Mac M-series 16GB

Llama 3.2 3B, Mistral 7B (Q4), Phi-3 Mini

Chat, summarization, simple coding tasks

Mac M-series 32GB

Llama 3.1 8B (full), Llama 3.1 70B (Q4), Qwen 2.5 32B

Strong coding, document analysis, complex reasoning

Mac M-series 64–128GB

Llama 3.1 70B (full quality), DeepSeek, Qwen 72B

Production-quality local inference, replaces cloud for most tasks

Windows/Linux + NVIDIA GPU

Depends on VRAM — RTX 3080 (10GB) fits 8B models in VRAM

Fast GPU inference; CPU fallback for larger models

Top Local Models Right Now

Llama 3.1 8B
8GB+ RAM
Best all-around small model. Fast, smart, great for coding and chat.
Llama 3.1 70B
32GB+ RAM (Q4)
Near-GPT-4 quality locally. The jump from 8B is significant.
Qwen 2.5 32B
24GB+ RAM
Excellent coding and reasoning. Strong multilingual support.
Mistral 7B
8GB+ RAM
Fast and efficient. Good for structured tasks and instruction following.
Phi-3 Mini
4GB+ RAM
Surprisingly capable for its size. Runs on almost anything.
DeepSeek Coder
16GB+ RAM
Specialized for code. Competes with GPT-4 on many coding benchmarks.

Pricing

$150 / 60-minute session
  • 1:1 Zoom — screen share your machine, we set it up live
  • Hardware assessment and model recommendations
  • Ollama or LM Studio setup and configuration
  • Local API wiring to your tools (Open WebUI, Continue.dev, etc.)
  • Privacy workflow design for your specific use case
  • Follow-up notes and model list after the session
Book Your Session →

Frequently Asked Questions

Can I run good AI models on a MacBook?
Yes — Apple Silicon Macs (M1 through M4) are excellent for local AI. The unified memory means the GPU and CPU share the same RAM pool. A 32GB Mac can run Llama 70B at Q4 quantization at usable speeds. A 64GB or 128GB Mac runs it comfortably. Even a 16GB Mac runs strong 7–8B models well.
Is local AI actually private?
Yes — when you run a model with Ollama or LM Studio, your prompts never leave your machine. No API calls, no logging, no training on your data. This matters for legal documents, patient data, confidential business files, and proprietary code you can't send to a cloud provider.
What's the difference between Ollama and LM Studio?
Ollama is CLI-first and exposes an OpenAI-compatible API — great for developers who want to wire local models into their apps or tools. LM Studio is GUI-first with a built-in chat UI — better for non-developers who want a desktop app experience similar to ChatGPT but running locally. Many people use both.
Do I need a GPU?
Not on a Mac — Apple Silicon handles inference efficiently on the integrated GPU in unified memory. On Windows or Linux, an NVIDIA GPU (RTX 3080 or better) dramatically speeds things up, but smaller models (7–8B) run on CPU-only setups at usable speeds. We'll set realistic expectations based on your exact hardware.
How does local AI compare to ChatGPT or Claude?
Honestly: for most tasks, GPT-4o and Claude Sonnet are still stronger than what you can run locally on consumer hardware. The tradeoff is privacy, cost, and offline capability. Local models are catching up fast — Llama 3.1 70B on a 64GB Mac is genuinely impressive. The session will help you figure out whether local is good enough for your use case or whether you should stay on cloud.

Related Pages

Your data. Your hardware. Your AI.

60 minutes to go from zero to a working local AI setup you actually use.

Book a Session — $150 →