AWS VibeDeploy
Private AI Infrastructure

Your own AI infrastructure.
Inside your AWS.

Run private AI agents on your own GPU instance. Your code, your prompts, your data — never leave your AWS account.

Hermes Desktop + Ollama + Terraform automation. One-command start/stop with GPU auto-stop — pay only when it's running.

Open source repo Data never leaves AWS ~$58/mo for a team of 5 GPU auto-stop included

The Problem

Every prompt you send
leaves your perimeter.

ChatGPT, Claude, and API-based tools are great — but every request goes to someone else's servers. For some teams, that's a dealbreaker.

Data leaves your control

Code, prompts, internal docs — sent to third-party APIs. A compliance problem for FinTech, HealthTech, or anyone with an NDA.

Subscriptions stack up

$20/mo per seat × 5 people × multiple tools = $200+/month, forever, with no ownership.

Rate limits & downtime

API throttling, outages, model deprecations — all outside your control, all on someone else's schedule.

Your laptop gets hot

Running local models on your own machine drains battery, heats up your laptop, and competes with your actual work.

What's Inside

A private inference node.
Inside your AWS account.

A cold, quiet laptop — and a powerful AI agent running somewhere else, on hardware you control.

🖥️

Hermes Desktop

A clean desktop interface to interact with your private models — chat, agents, and tool use, all running on your own infrastructure.

🦙

Ollama

Run open-source models (Llama, Qwen, DeepSeek, and more) locally on the GPU instance. Pull and switch models with one command.

AWS EC2 GPU Instance

Runs on g4dn.xlarge by default — enough for 7B–13B models at solid speed. Lives entirely inside your own AWS account.

📐

Terraform Automation

Infrastructure as code — spin up, tear down, or rebuild your AI node in minutes. Includes GPU auto-stop to avoid idle costs.

The economics

~$58/month
for a team of 5

vs $200+/month in stacked AI subscriptions — and that's before anyone hits a rate limit or a model gets deprecated.

GPU auto-stop means you pay AWS only while the instance is actually running. Idle nights and weekends cost close to nothing.

5× AI subscriptions$200+/mo
Private AI Circuit~$58/mo
Monthly savings ~$142+

// Your data never leaves your AWS account. No per-seat pricing. No rate limits.

Choose Your Path

From free to fully managed.

Start with the free repo. Skip the setup time with the Blueprint. Or have it deployed for you.

Open Source
Free
self-deploy · GitHub repo

The full stack, open source. For developers comfortable with Terraform and AWS who want to deploy it themselves.

  • Full Terraform source code
  • Setup documentation
  • Hermes Desktop + Ollama configs
  • Community support via GitHub Issues
View on GitHub
MOST POPULAR
Deployment Blueprint
$49
one-time · self-deploy, faster

Everything from the free repo, plus the parts that take hours to get right yourself — pre-configured and tested.

  • Everything in Open Source
  • Elastic IP pre-configured
  • GPU auto-stop script
  • Model pull selector
  • Step-by-step deployment guide
  • Tested on g4dn.xlarge
Get the Blueprint — $49
Custom Deployment
$800+
scoped after conversation

I deploy it in your AWS account, configure it for your team's needs, and walk you through using it.

  • Everything in Blueprint
  • Deployed in your AWS account
  • Configured for your team size & models
  • OIDC keyless access (same as Vibe Deploy)
  • Walkthrough for your team
  • 1 week of support
Discuss My Setup →

Data Sovereignty

Your prompts.
Your AWS. Your rules.

Every request stays inside your own AWS account — on a GPU instance you control. No third-party API sees your code, your prompts, or your internal data.

For custom deployments, I use the same OIDC keyless access as AWS Vibe Deploy — I never see your AWS credentials. See how that works →

Your team
↓ prompts, code, agents
Your AWS account (GPU)
OpenAI / Anthropic API never sees it

Custom Deployment

Get it deployed for you

"Fill this in. I'll get back within 2 hours."

// Response within 1–2 hours.

Team size

Primary use case

// Or skip the wait: GitHub repo (free) · Blueprint ($49)