Skip to content

Commit 1f777c0

Browse files
Add use cases, cost management, and privacy/compliance guides
- 20 verified real-world use cases from HN, blogs, and community reports - Cost management with horror stories, optimization strategies, and 97% reduction case study - Privacy & compliance covering GDPR, SOC 2, HIPAA, air-gapped deployments, and government responses
1 parent dd46cc1 commit 1f777c0

4 files changed

Lines changed: 926 additions & 0 deletions

File tree

‎docs/guides/cost-management.md‎

Lines changed: 314 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,314 @@
1+
---
2+
sidebar_position: 12
3+
title: Cost Management
4+
description: Control OpenClaw spending — horror stories, cost breakdowns, optimization strategies, monitoring tools, and a documented 97% cost reduction
5+
---
6+
7+
# Cost Management
8+
9+
OpenClaw's default configuration sends everything to one model, includes full conversation history, and heartbeats fire every 30 minutes. Without tuning, **active use can cost $300-750/month**. This guide covers what drives costs, how to control them, and real strategies that have achieved up to **97% cost reductions**.
10+
11+
:::danger Real Horror Stories
12+
These are documented incidents, not hypotheticals:
13+
- **$3,600/month** — Federico Viticci's "Navi" assistant consumed 180M Anthropic tokens in one month
14+
- **$20 overnight** — Benjamin De Kraker's heartbeat drained his entire balance checking "is it daytime yet?" every 30 minutes at $0.75/check
15+
- **$200/day** — An automated task stuck in an infinite loop, churning tokens while accomplishing nothing
16+
- **$8 per 30 min** — Processing Moltbook browsing content, adding up to $380+/day
17+
:::
18+
19+
---
20+
21+
## Where Does the Money Go?
22+
23+
### Cost Breakdown by Component
24+
25+
```mermaid
26+
pie title Typical Token Consumption
27+
"Heartbeat checks" : 35
28+
"Chat/conversation" : 25
29+
"Skill execution" : 20
30+
"Context overhead" : 15
31+
"Sub-agent spawning" : 5
32+
```
33+
34+
### Heartbeat Costs
35+
36+
The heartbeat is the biggest silent cost driver. Each heartbeat sends the **entire context window** to the API:
37+
38+
| Model | Cost per Heartbeat (~120K tokens) | Monthly (48/day) |
39+
|-------|----------------------------------|-------------------|
40+
| Claude Opus 4.5 | ~$0.75 | ~$1,080 |
41+
| Claude Sonnet 4.5 | ~$0.45 | ~$648 |
42+
| Claude Haiku | ~$0.15 | ~$216 |
43+
| Gemini 2.5 Flash | ~$0.05 | ~$72 |
44+
| Gemini 2.5 Flash-Lite | ~$0.01 | ~$14 |
45+
| Local (Ollama) | $0.00 | **$0** |
46+
47+
> "Using Opus for a heartbeat is like hiring a lawyer to check your mailbox."
48+
49+
### Context Accumulation
50+
51+
Every request sends the full conversation history. One user's session grew to **1.4 MB / 100K tokens** of pure overhead. System prompts alone are 5,000-10,000 tokens and are resent with every API call.
52+
53+
### What Causes Runaway Costs
54+
55+
| Cause | Impact | Example |
56+
|-------|--------|---------|
57+
| **Infinite loops** | Tokens burn on repeated failures | Tool validation fails, error appended to context, model retries endlessly |
58+
| **Context snowball** | Cost per request grows over time | Session at 56-58% of 400K window means 200K+ cached tokens per query |
59+
| **Wrong model for task** | 60x overpayment | Opus for heartbeat vs. Flash-Lite |
60+
| **Thinking mode** | 10-50x token explosion | Extended reasoning on simple tasks |
61+
| **No spending limits** | Unpredictable bills | Overnight drain with no caps configured |
62+
| **Large tool outputs** | Massive context per loop | 1M file paths sent with full history on every iteration |
63+
64+
---
65+
66+
## Optimization Strategies
67+
68+
### 1. Model Routing (Biggest Impact)
69+
70+
Use different models for different tasks. The cost difference between Opus and Flash-Lite is **60x**.
71+
72+
| Task Type | Recommended Model | Cost/1M Tokens | Notes |
73+
|-----------|------------------|----------------|-------|
74+
| Heartbeat checks | Gemini 2.5 Flash-Lite | **$0.50** | No quality difference for simple checks |
75+
| Simple tasks | Gemini 3 Flash | **$0.15** | Quick lookups, status checks |
76+
| Long context | Kimi K2.5 | **$0.50** | Large document handling |
77+
| Coding | Gemini 3 Pro | **$1.25** | Mid-tier reasoning |
78+
| Quick responses | Claude Haiku | **$1.00/$5.00** | 90% of Sonnet quality at 1/3 price |
79+
| Mid-tier general | Claude Sonnet | **$3.00/$15.00** | Good balance |
80+
| Complex reasoning | Claude Opus 4.5 | **$5.00/$25.00** | Reserve for complex tasks only |
81+
| Budget alternative | DeepSeek V3.2 | **$0.53** | 60x cheaper than Opus |
82+
83+
Configure hybrid routing:
84+
85+
```yaml title="~/.openclaw/config.yml"
86+
brain:
87+
provider: "anthropic"
88+
model: "claude-sonnet-4"
89+
90+
# Use local model for heartbeat
91+
heartbeat_override:
92+
provider: "local"
93+
local:
94+
endpoint: "http://localhost:11434"
95+
model: "qwen3:14b"
96+
type: "ollama"
97+
```
98+
99+
### 2. Context Window Management
100+
101+
| Strategy | How | Impact |
102+
|----------|-----|--------|
103+
| **Start fresh sessions** | Use `/new` or `/reset` regularly | Prevents context snowball |
104+
| **Compact context** | Use `/compact` to summarize old messages | Frees window space |
105+
| **Limit context size** | `openclaw config set max_context_tokens 8000` | Per-session cost dropped from $0.40 to $0.05 |
106+
| **Limit response size** | `openclaw config set max_tokens_per_request 4000` | Prevents verbose responses |
107+
| **Daily auto-reset** | Default 4:00 AM local time | Creates fresh session automatically |
108+
109+
### 3. Heartbeat Tuning
110+
111+
Configure active hours to stop overnight token drain:
112+
113+
```json title="~/.openclaw/openclaw.json"
114+
{
115+
"heartbeat": {
116+
"every": "1h",
117+
"activeHours": {
118+
"start": "08:00",
119+
"end": "22:00"
120+
}
121+
}
122+
}
123+
```
124+
125+
| Setting | Cost Impact |
126+
|---------|-------------|
127+
| Increase interval from 30m to 1h | **-50%** heartbeat cost |
128+
| Increase interval from 30m to 2h | **-75%** heartbeat cost |
129+
| Set active hours (14h/day vs 24h) | **-42%** heartbeat cost |
130+
| Use local model for heartbeat | **-100%** heartbeat cost |
131+
| Set to `0m` | Disables heartbeats entirely |
132+
133+
### 4. Disable Thinking Mode for Routine Tasks
134+
135+
Thinking/reasoning mode can increase token usage by **10-50x**. Disable it for routine tasks:
136+
137+
```json5 title="~/.config/openclaw/config.json5"
138+
{
139+
"thinking": "disabled"
140+
}
141+
```
142+
143+
### 5. Leverage Prompt Caching
144+
145+
Anthropic prompt caching charges cache hits at only **10% of normal cost**:
146+
147+
- OpenClaw automatically applies `cacheRetention: "short"` (5-minute cache)
148+
- System prompts (5K-10K tokens) only charged once during cache validity
149+
- High-frequency users save **30-50%** with caching
150+
- Extend with `cacheRetention: "long"` in model config
151+
152+
### 6. Use Existing Subscriptions (Zero Additional Cost)
153+
154+
If you already pay for a subscription, connect it to OpenClaw with **no additional API costs**:
155+
156+
| Subscription | Monthly Cost | How to Connect |
157+
|-------------|-------------|----------------|
158+
| Claude Pro | $20/month | Authenticate via claude-cli |
159+
| Claude Max | $100/month | Same method, higher limits |
160+
| ChatGPT Plus | $20/month | Connect API key |
161+
| Google One AI Premium | $20/month | Connect API key |
162+
163+
During onboarding, select "Anthropic" as your provider, then choose "Claude Code CLI" as the auth method to use your existing subscription.
164+
165+
---
166+
167+
## Budget Controls
168+
169+
### Set Spending Limits
170+
171+
```bash
172+
# OpenClaw built-in limits
173+
openclaw config set daily_budget_usd 5.00
174+
openclaw config set monthly_budget_usd 50.00
175+
```
176+
177+
### Provider-Level Alerts
178+
179+
| Provider | How | Recommendation |
180+
|----------|-----|---------------|
181+
| **Anthropic** | Dashboard spending alerts | Set alerts at 50%, 75%, 90% of budget |
182+
| **OpenAI** | Hard monthly spending caps | Configure directly in dashboard |
183+
| **OpenRouter** | Built-in spending tracking | Monitor across all routed models |
184+
185+
### Model Failover Chain
186+
187+
```json title="~/.openclaw/openclaw.json"
188+
{
189+
"models": {
190+
"primary": "anthropic/claude-sonnet-4",
191+
"fallback": "anthropic/claude-haiku",
192+
"complex": "anthropic/claude-opus-4.5"
193+
}
194+
}
195+
```
196+
197+
---
198+
199+
## Real Cost Data
200+
201+
### Monthly Spending by Usage Pattern
202+
203+
| Usage Pattern | Monthly Cost | Description |
204+
|--------------|-------------|-------------|
205+
| **Free / Minimal** | $0-8 | Local models (Ollama) or existing subscription |
206+
| **Light / Casual** | $10-30 | Few commands per day, basic automation |
207+
| **Moderate** | $30-70 | Regular file tasks, research, moderate chat |
208+
| **Power User** | $70-200 | Constant automation, long sessions |
209+
| **Heavy / Always-On** | $300-750 | Full proactive assistant, multiple integrations |
210+
| **Enterprise / Opus** | $500-5,000 | Full-time Opus agent, heavy tool usage |
211+
212+
> The median: **90% of users spend under $12/day**, and the average developer spends roughly **$100-200/month**.
213+
214+
### Additional Costs
215+
216+
| Component | Cost | Notes |
217+
|-----------|------|-------|
218+
| OpenClaw software | **$0** | MIT open source |
219+
| Local hardware | **$0** | Your own machine |
220+
| VPS hosting | **$5-20/month** | Optional, for always-on |
221+
| API keys | Variable | See usage patterns above |
222+
223+
---
224+
225+
## The $1,200 to $36 Case Study
226+
227+
A documented case achieved a **97% cost reduction** by combining five changes:
228+
229+
| Change | Effect |
230+
|--------|--------|
231+
| 1. Local search (QMDskill with BM25 + vector) | Token usage dropped 95% |
232+
| 2. Session context reduction (50KB to 8KB) | Per-session cost: $0.40 to $0.05 |
233+
| 3. Free web search via Exa AI | Eliminated paid search tokens |
234+
| 4. Automatic model routing | Right model for each task |
235+
| 5. Local heartbeat handling | Eliminated recurring paid heartbeats |
236+
237+
These reductions **compound multiplicatively** — $1,200/month became $36/month.
238+
239+
---
240+
241+
## Monitoring Tools
242+
243+
### Built-in
244+
245+
| Command | What It Shows |
246+
|---------|--------------|
247+
| `/status` | Agent reachability, context usage, current model |
248+
| `/usage full` | Token usage for every response |
249+
250+
### Third-Party Dashboards
251+
252+
| Tool | Features |
253+
|------|----------|
254+
| [**Clawalytics**](https://www.clawalytics.com/) | Real-time spending dashboard, per-agent breakdowns, daily cost charts, suspicious activity alerts |
255+
| [**ClawWatcher**](https://news.ycombinator.com/item?id=46954200) | Real-time token usage, cost per model, skills/actions tracking |
256+
| [**claw-dash**](https://github.com/openclaw/openclaw/discussions/8277) | Sessions, tokens (24h), costs, model info, cron jobs, system health gauges |
257+
| [**openclaw-dashboard**](https://github.com/tugcantopaloglu/openclaw-dashboard) | Browser notifications for usage limits, cost analysis by model |
258+
| [**OpenClaw Cost Calculator**](https://calculator.vlvt.sh/) | Pre-deployment cost estimation |
259+
260+
### OpenRouter Auto-Routing
261+
262+
For centralized model routing, OpenRouter can automatically select the most cost-effective model per prompt using `openrouter/openrouter/auto` — single API key for all providers with built-in cost tracking.
263+
264+
---
265+
266+
## Quick-Start: Cut Costs Today
267+
268+
If you're spending too much right now, apply these changes immediately:
269+
270+
1. **Switch heartbeat to a cheap model** — biggest single savings
271+
2. **Set active hours** — stop overnight token drain
272+
3. **Set daily and monthly budget limits** — prevent surprises
273+
4. **Use `/compact` regularly** — prevent context snowball
274+
5. **Disable thinking mode** for routine tasks
275+
276+
```yaml title="~/.openclaw/config.yml — Minimum cost configuration"
277+
brain:
278+
provider: "anthropic"
279+
model: "claude-sonnet-4"
280+
281+
heartbeat_override:
282+
provider: "local"
283+
local:
284+
endpoint: "http://localhost:11434"
285+
model: "llama3.1:8b"
286+
type: "ollama"
287+
```
288+
289+
```json title="~/.openclaw/openclaw.json — Budget limits"
290+
{
291+
"heartbeat": {
292+
"every": "1h",
293+
"activeHours": {
294+
"start": "08:00",
295+
"end": "22:00"
296+
}
297+
}
298+
}
299+
```
300+
301+
```bash
302+
openclaw config set daily_budget_usd 5.00
303+
openclaw config set monthly_budget_usd 50.00
304+
```
305+
306+
---
307+
308+
## See Also
309+
310+
- [Local Models](/guides/local-models) — Zero-cost local inference
311+
- [Cloud GPU & Self-Hosted Models](/guides/cloud-gpu-models) — Cost comparison of cloud GPU options
312+
- [Heartbeat System](/architecture/heartbeat) — How heartbeat scheduling works
313+
- [Configuration Reference](/reference/configuration) — All cost-related settings
314+
- [Real-World Use Cases](/guides/use-cases) — Cost data from real deployments

0 commit comments

Comments
 (0)