|
| 1 | +--- |
| 2 | +sidebar_position: 12 |
| 3 | +title: Cost Management |
| 4 | +description: Control OpenClaw spending — horror stories, cost breakdowns, optimization strategies, monitoring tools, and a documented 97% cost reduction |
| 5 | +--- |
| 6 | + |
| 7 | +# Cost Management |
| 8 | + |
| 9 | +OpenClaw's default configuration sends everything to one model, includes full conversation history, and heartbeats fire every 30 minutes. Without tuning, **active use can cost $300-750/month**. This guide covers what drives costs, how to control them, and real strategies that have achieved up to **97% cost reductions**. |
| 10 | + |
| 11 | +:::danger Real Horror Stories |
| 12 | +These are documented incidents, not hypotheticals: |
| 13 | +- **$3,600/month** — Federico Viticci's "Navi" assistant consumed 180M Anthropic tokens in one month |
| 14 | +- **$20 overnight** — Benjamin De Kraker's heartbeat drained his entire balance checking "is it daytime yet?" every 30 minutes at $0.75/check |
| 15 | +- **$200/day** — An automated task stuck in an infinite loop, churning tokens while accomplishing nothing |
| 16 | +- **$8 per 30 min** — Processing Moltbook browsing content, adding up to $380+/day |
| 17 | +::: |
| 18 | + |
| 19 | +--- |
| 20 | + |
| 21 | +## Where Does the Money Go? |
| 22 | + |
| 23 | +### Cost Breakdown by Component |
| 24 | + |
| 25 | +```mermaid |
| 26 | +pie title Typical Token Consumption |
| 27 | + "Heartbeat checks" : 35 |
| 28 | + "Chat/conversation" : 25 |
| 29 | + "Skill execution" : 20 |
| 30 | + "Context overhead" : 15 |
| 31 | + "Sub-agent spawning" : 5 |
| 32 | +``` |
| 33 | + |
| 34 | +### Heartbeat Costs |
| 35 | + |
| 36 | +The heartbeat is the biggest silent cost driver. Each heartbeat sends the **entire context window** to the API: |
| 37 | + |
| 38 | +| Model | Cost per Heartbeat (~120K tokens) | Monthly (48/day) | |
| 39 | +|-------|----------------------------------|-------------------| |
| 40 | +| Claude Opus 4.5 | ~$0.75 | ~$1,080 | |
| 41 | +| Claude Sonnet 4.5 | ~$0.45 | ~$648 | |
| 42 | +| Claude Haiku | ~$0.15 | ~$216 | |
| 43 | +| Gemini 2.5 Flash | ~$0.05 | ~$72 | |
| 44 | +| Gemini 2.5 Flash-Lite | ~$0.01 | ~$14 | |
| 45 | +| Local (Ollama) | $0.00 | **$0** | |
| 46 | + |
| 47 | +> "Using Opus for a heartbeat is like hiring a lawyer to check your mailbox." |
| 48 | +
|
| 49 | +### Context Accumulation |
| 50 | + |
| 51 | +Every request sends the full conversation history. One user's session grew to **1.4 MB / 100K tokens** of pure overhead. System prompts alone are 5,000-10,000 tokens and are resent with every API call. |
| 52 | + |
| 53 | +### What Causes Runaway Costs |
| 54 | + |
| 55 | +| Cause | Impact | Example | |
| 56 | +|-------|--------|---------| |
| 57 | +| **Infinite loops** | Tokens burn on repeated failures | Tool validation fails, error appended to context, model retries endlessly | |
| 58 | +| **Context snowball** | Cost per request grows over time | Session at 56-58% of 400K window means 200K+ cached tokens per query | |
| 59 | +| **Wrong model for task** | 60x overpayment | Opus for heartbeat vs. Flash-Lite | |
| 60 | +| **Thinking mode** | 10-50x token explosion | Extended reasoning on simple tasks | |
| 61 | +| **No spending limits** | Unpredictable bills | Overnight drain with no caps configured | |
| 62 | +| **Large tool outputs** | Massive context per loop | 1M file paths sent with full history on every iteration | |
| 63 | + |
| 64 | +--- |
| 65 | + |
| 66 | +## Optimization Strategies |
| 67 | + |
| 68 | +### 1. Model Routing (Biggest Impact) |
| 69 | + |
| 70 | +Use different models for different tasks. The cost difference between Opus and Flash-Lite is **60x**. |
| 71 | + |
| 72 | +| Task Type | Recommended Model | Cost/1M Tokens | Notes | |
| 73 | +|-----------|------------------|----------------|-------| |
| 74 | +| Heartbeat checks | Gemini 2.5 Flash-Lite | **$0.50** | No quality difference for simple checks | |
| 75 | +| Simple tasks | Gemini 3 Flash | **$0.15** | Quick lookups, status checks | |
| 76 | +| Long context | Kimi K2.5 | **$0.50** | Large document handling | |
| 77 | +| Coding | Gemini 3 Pro | **$1.25** | Mid-tier reasoning | |
| 78 | +| Quick responses | Claude Haiku | **$1.00/$5.00** | 90% of Sonnet quality at 1/3 price | |
| 79 | +| Mid-tier general | Claude Sonnet | **$3.00/$15.00** | Good balance | |
| 80 | +| Complex reasoning | Claude Opus 4.5 | **$5.00/$25.00** | Reserve for complex tasks only | |
| 81 | +| Budget alternative | DeepSeek V3.2 | **$0.53** | 60x cheaper than Opus | |
| 82 | + |
| 83 | +Configure hybrid routing: |
| 84 | + |
| 85 | +```yaml title="~/.openclaw/config.yml" |
| 86 | +brain: |
| 87 | + provider: "anthropic" |
| 88 | + model: "claude-sonnet-4" |
| 89 | + |
| 90 | + # Use local model for heartbeat |
| 91 | + heartbeat_override: |
| 92 | + provider: "local" |
| 93 | + local: |
| 94 | + endpoint: "http://localhost:11434" |
| 95 | + model: "qwen3:14b" |
| 96 | + type: "ollama" |
| 97 | +``` |
| 98 | +
|
| 99 | +### 2. Context Window Management |
| 100 | +
|
| 101 | +| Strategy | How | Impact | |
| 102 | +|----------|-----|--------| |
| 103 | +| **Start fresh sessions** | Use `/new` or `/reset` regularly | Prevents context snowball | |
| 104 | +| **Compact context** | Use `/compact` to summarize old messages | Frees window space | |
| 105 | +| **Limit context size** | `openclaw config set max_context_tokens 8000` | Per-session cost dropped from $0.40 to $0.05 | |
| 106 | +| **Limit response size** | `openclaw config set max_tokens_per_request 4000` | Prevents verbose responses | |
| 107 | +| **Daily auto-reset** | Default 4:00 AM local time | Creates fresh session automatically | |
| 108 | + |
| 109 | +### 3. Heartbeat Tuning |
| 110 | + |
| 111 | +Configure active hours to stop overnight token drain: |
| 112 | + |
| 113 | +```json title="~/.openclaw/openclaw.json" |
| 114 | +{ |
| 115 | + "heartbeat": { |
| 116 | + "every": "1h", |
| 117 | + "activeHours": { |
| 118 | + "start": "08:00", |
| 119 | + "end": "22:00" |
| 120 | + } |
| 121 | + } |
| 122 | +} |
| 123 | +``` |
| 124 | + |
| 125 | +| Setting | Cost Impact | |
| 126 | +|---------|-------------| |
| 127 | +| Increase interval from 30m to 1h | **-50%** heartbeat cost | |
| 128 | +| Increase interval from 30m to 2h | **-75%** heartbeat cost | |
| 129 | +| Set active hours (14h/day vs 24h) | **-42%** heartbeat cost | |
| 130 | +| Use local model for heartbeat | **-100%** heartbeat cost | |
| 131 | +| Set to `0m` | Disables heartbeats entirely | |
| 132 | + |
| 133 | +### 4. Disable Thinking Mode for Routine Tasks |
| 134 | + |
| 135 | +Thinking/reasoning mode can increase token usage by **10-50x**. Disable it for routine tasks: |
| 136 | + |
| 137 | +```json5 title="~/.config/openclaw/config.json5" |
| 138 | +{ |
| 139 | + "thinking": "disabled" |
| 140 | +} |
| 141 | +``` |
| 142 | + |
| 143 | +### 5. Leverage Prompt Caching |
| 144 | + |
| 145 | +Anthropic prompt caching charges cache hits at only **10% of normal cost**: |
| 146 | + |
| 147 | +- OpenClaw automatically applies `cacheRetention: "short"` (5-minute cache) |
| 148 | +- System prompts (5K-10K tokens) only charged once during cache validity |
| 149 | +- High-frequency users save **30-50%** with caching |
| 150 | +- Extend with `cacheRetention: "long"` in model config |
| 151 | + |
| 152 | +### 6. Use Existing Subscriptions (Zero Additional Cost) |
| 153 | + |
| 154 | +If you already pay for a subscription, connect it to OpenClaw with **no additional API costs**: |
| 155 | + |
| 156 | +| Subscription | Monthly Cost | How to Connect | |
| 157 | +|-------------|-------------|----------------| |
| 158 | +| Claude Pro | $20/month | Authenticate via claude-cli | |
| 159 | +| Claude Max | $100/month | Same method, higher limits | |
| 160 | +| ChatGPT Plus | $20/month | Connect API key | |
| 161 | +| Google One AI Premium | $20/month | Connect API key | |
| 162 | + |
| 163 | +During onboarding, select "Anthropic" as your provider, then choose "Claude Code CLI" as the auth method to use your existing subscription. |
| 164 | + |
| 165 | +--- |
| 166 | + |
| 167 | +## Budget Controls |
| 168 | + |
| 169 | +### Set Spending Limits |
| 170 | + |
| 171 | +```bash |
| 172 | +# OpenClaw built-in limits |
| 173 | +openclaw config set daily_budget_usd 5.00 |
| 174 | +openclaw config set monthly_budget_usd 50.00 |
| 175 | +``` |
| 176 | + |
| 177 | +### Provider-Level Alerts |
| 178 | + |
| 179 | +| Provider | How | Recommendation | |
| 180 | +|----------|-----|---------------| |
| 181 | +| **Anthropic** | Dashboard spending alerts | Set alerts at 50%, 75%, 90% of budget | |
| 182 | +| **OpenAI** | Hard monthly spending caps | Configure directly in dashboard | |
| 183 | +| **OpenRouter** | Built-in spending tracking | Monitor across all routed models | |
| 184 | + |
| 185 | +### Model Failover Chain |
| 186 | + |
| 187 | +```json title="~/.openclaw/openclaw.json" |
| 188 | +{ |
| 189 | + "models": { |
| 190 | + "primary": "anthropic/claude-sonnet-4", |
| 191 | + "fallback": "anthropic/claude-haiku", |
| 192 | + "complex": "anthropic/claude-opus-4.5" |
| 193 | + } |
| 194 | +} |
| 195 | +``` |
| 196 | + |
| 197 | +--- |
| 198 | + |
| 199 | +## Real Cost Data |
| 200 | + |
| 201 | +### Monthly Spending by Usage Pattern |
| 202 | + |
| 203 | +| Usage Pattern | Monthly Cost | Description | |
| 204 | +|--------------|-------------|-------------| |
| 205 | +| **Free / Minimal** | $0-8 | Local models (Ollama) or existing subscription | |
| 206 | +| **Light / Casual** | $10-30 | Few commands per day, basic automation | |
| 207 | +| **Moderate** | $30-70 | Regular file tasks, research, moderate chat | |
| 208 | +| **Power User** | $70-200 | Constant automation, long sessions | |
| 209 | +| **Heavy / Always-On** | $300-750 | Full proactive assistant, multiple integrations | |
| 210 | +| **Enterprise / Opus** | $500-5,000 | Full-time Opus agent, heavy tool usage | |
| 211 | + |
| 212 | +> The median: **90% of users spend under $12/day**, and the average developer spends roughly **$100-200/month**. |
| 213 | + |
| 214 | +### Additional Costs |
| 215 | + |
| 216 | +| Component | Cost | Notes | |
| 217 | +|-----------|------|-------| |
| 218 | +| OpenClaw software | **$0** | MIT open source | |
| 219 | +| Local hardware | **$0** | Your own machine | |
| 220 | +| VPS hosting | **$5-20/month** | Optional, for always-on | |
| 221 | +| API keys | Variable | See usage patterns above | |
| 222 | + |
| 223 | +--- |
| 224 | + |
| 225 | +## The $1,200 to $36 Case Study |
| 226 | + |
| 227 | +A documented case achieved a **97% cost reduction** by combining five changes: |
| 228 | + |
| 229 | +| Change | Effect | |
| 230 | +|--------|--------| |
| 231 | +| 1. Local search (QMDskill with BM25 + vector) | Token usage dropped 95% | |
| 232 | +| 2. Session context reduction (50KB to 8KB) | Per-session cost: $0.40 to $0.05 | |
| 233 | +| 3. Free web search via Exa AI | Eliminated paid search tokens | |
| 234 | +| 4. Automatic model routing | Right model for each task | |
| 235 | +| 5. Local heartbeat handling | Eliminated recurring paid heartbeats | |
| 236 | + |
| 237 | +These reductions **compound multiplicatively** — $1,200/month became $36/month. |
| 238 | + |
| 239 | +--- |
| 240 | + |
| 241 | +## Monitoring Tools |
| 242 | + |
| 243 | +### Built-in |
| 244 | + |
| 245 | +| Command | What It Shows | |
| 246 | +|---------|--------------| |
| 247 | +| `/status` | Agent reachability, context usage, current model | |
| 248 | +| `/usage full` | Token usage for every response | |
| 249 | + |
| 250 | +### Third-Party Dashboards |
| 251 | + |
| 252 | +| Tool | Features | |
| 253 | +|------|----------| |
| 254 | +| [**Clawalytics**](https://www.clawalytics.com/) | Real-time spending dashboard, per-agent breakdowns, daily cost charts, suspicious activity alerts | |
| 255 | +| [**ClawWatcher**](https://news.ycombinator.com/item?id=46954200) | Real-time token usage, cost per model, skills/actions tracking | |
| 256 | +| [**claw-dash**](https://github.com/openclaw/openclaw/discussions/8277) | Sessions, tokens (24h), costs, model info, cron jobs, system health gauges | |
| 257 | +| [**openclaw-dashboard**](https://github.com/tugcantopaloglu/openclaw-dashboard) | Browser notifications for usage limits, cost analysis by model | |
| 258 | +| [**OpenClaw Cost Calculator**](https://calculator.vlvt.sh/) | Pre-deployment cost estimation | |
| 259 | + |
| 260 | +### OpenRouter Auto-Routing |
| 261 | + |
| 262 | +For centralized model routing, OpenRouter can automatically select the most cost-effective model per prompt using `openrouter/openrouter/auto` — single API key for all providers with built-in cost tracking. |
| 263 | + |
| 264 | +--- |
| 265 | + |
| 266 | +## Quick-Start: Cut Costs Today |
| 267 | + |
| 268 | +If you're spending too much right now, apply these changes immediately: |
| 269 | + |
| 270 | +1. **Switch heartbeat to a cheap model** — biggest single savings |
| 271 | +2. **Set active hours** — stop overnight token drain |
| 272 | +3. **Set daily and monthly budget limits** — prevent surprises |
| 273 | +4. **Use `/compact` regularly** — prevent context snowball |
| 274 | +5. **Disable thinking mode** for routine tasks |
| 275 | + |
| 276 | +```yaml title="~/.openclaw/config.yml — Minimum cost configuration" |
| 277 | +brain: |
| 278 | + provider: "anthropic" |
| 279 | + model: "claude-sonnet-4" |
| 280 | +
|
| 281 | + heartbeat_override: |
| 282 | + provider: "local" |
| 283 | + local: |
| 284 | + endpoint: "http://localhost:11434" |
| 285 | + model: "llama3.1:8b" |
| 286 | + type: "ollama" |
| 287 | +``` |
| 288 | + |
| 289 | +```json title="~/.openclaw/openclaw.json — Budget limits" |
| 290 | +{ |
| 291 | + "heartbeat": { |
| 292 | + "every": "1h", |
| 293 | + "activeHours": { |
| 294 | + "start": "08:00", |
| 295 | + "end": "22:00" |
| 296 | + } |
| 297 | + } |
| 298 | +} |
| 299 | +``` |
| 300 | + |
| 301 | +```bash |
| 302 | +openclaw config set daily_budget_usd 5.00 |
| 303 | +openclaw config set monthly_budget_usd 50.00 |
| 304 | +``` |
| 305 | + |
| 306 | +--- |
| 307 | + |
| 308 | +## See Also |
| 309 | + |
| 310 | +- [Local Models](/guides/local-models) — Zero-cost local inference |
| 311 | +- [Cloud GPU & Self-Hosted Models](/guides/cloud-gpu-models) — Cost comparison of cloud GPU options |
| 312 | +- [Heartbeat System](/architecture/heartbeat) — How heartbeat scheduling works |
| 313 | +- [Configuration Reference](/reference/configuration) — All cost-related settings |
| 314 | +- [Real-World Use Cases](/guides/use-cases) — Cost data from real deployments |
0 commit comments