Run private AI agents on your own GPU instance. Your code, your prompts, your data — never leave your AWS account.
Hermes Desktop + Ollama + Terraform automation. One-command start/stop with GPU auto-stop — pay only when it's running.
The Problem
ChatGPT, Claude, and API-based tools are great — but every request goes to someone else's servers. For some teams, that's a dealbreaker.
Data leaves your control
Code, prompts, internal docs — sent to third-party APIs. A compliance problem for FinTech, HealthTech, or anyone with an NDA.
Subscriptions stack up
$20/mo per seat × 5 people × multiple tools = $200+/month, forever, with no ownership.
Rate limits & downtime
API throttling, outages, model deprecations — all outside your control, all on someone else's schedule.
Your laptop gets hot
Running local models on your own machine drains battery, heats up your laptop, and competes with your actual work.
What's Inside
A cold, quiet laptop — and a powerful AI agent running somewhere else, on hardware you control.
A clean desktop interface to interact with your private models — chat, agents, and tool use, all running on your own infrastructure.
Run open-source models (Llama, Qwen, DeepSeek, and more) locally on the GPU instance. Pull and switch models with one command.
Runs on g4dn.xlarge by default — enough for 7B–13B models at solid speed. Lives entirely inside your own AWS account.
Infrastructure as code — spin up, tear down, or rebuild your AI node in minutes. Includes GPU auto-stop to avoid idle costs.
The economics
vs $200+/month in stacked AI subscriptions — and that's before anyone hits a rate limit or a model gets deprecated.
GPU auto-stop means you pay AWS only while the instance is actually running. Idle nights and weekends cost close to nothing.
// Your data never leaves your AWS account. No per-seat pricing. No rate limits.
Choose Your Path
Start with the free repo. Skip the setup time with the Blueprint. Or have it deployed for you.
The full stack, open source. For developers comfortable with Terraform and AWS who want to deploy it themselves.
Everything from the free repo, plus the parts that take hours to get right yourself — pre-configured and tested.
I deploy it in your AWS account, configure it for your team's needs, and walk you through using it.
Data Sovereignty
Every request stays inside your own AWS account — on a GPU instance you control. No third-party API sees your code, your prompts, or your internal data.
For custom deployments, I use the same OIDC keyless access as AWS Vibe Deploy — I never see your AWS credentials. See how that works →
Custom Deployment
"Fill this in. I'll get back within 2 hours."