How Going Back to the Stone Age Can Cut AI Costs

I taught my AI sub-agents to talk like cavemen. Alert investigations got 23% cheaper, 20% faster with no impact on quality

How Going Back to the Stone Age Can Cut AI Costs
AI expensive. Caveman optimize

AI expensive. Caveman optimize

One of the agents which I built is used for debugging an alert. Debugging an alert requires running parallel hypothesis investigations and going through multiple signals like logs, metrics, traces. This can cause context-pollution and force context compaction if I only use one agent. So how I orchestrate it is via one coordinator agent that delegates to a bunch of specialist sub-agents.

While building this workflow, I stumbled upon caveman mode. Turns out this actually works!

One agent bad. Many agent good.

Not always, but for alert debugging multi-agent workflows make a big difference. The coordinator agent reads the alert, makes a small plan, and fans out questions in parallel. Something like "what errors show up in order-worker logs between 13:30 and 14:30?" goes to the logs agent, a metrics question goes to the metrics agent, and so on.

                  alert
                    │
             ┌──────▼──────┐
             │ coordinator │ ──▶ root cause for the human
             └──────┬──────┘
        ┌───────────┼───────────┐
     ┌──▼──┐     ┌──▼────┐   ┌──▼───┐
     │logs │     │metrics│   │traces│   ...
     └─────┘     └───────┘   └──────┘

Each sub-agent runs its queries, chews through thousands of log lines or metric time series, and sends back a text summary. The raw data never goes back up to the coordinator agent. The coordinator only ever sees the plan and the summaries, so its context stays clean.

That fixed the context problem, but created another one.

Agent talk too much

Now everything hangs on that one summary, and I pay for each word in it twice. First as the sub-agent's output tokens, which are the expensive kind (5x the input price on Opus 5).
Then again as the coordinator's input, on every turn after that, because the summary sits in its context until the investigation ends.

It also slows everything down. The coordinator can't move on until the slowest sub-agent finishes typing.

So what were my sub-agents typing? I asked the logs agent why some pods were restarting, and it came back with a title. Then a "Root Cause" heading. Then a timeline, as a markdown table. Then an "Error volume" section that repeated half the timeline. A "Mechanism" section. A section on co-firing alerts that ended with "not confirmed here" and finally "Suggested next steps for remediation owner".

400 words, beautifully formatted, for another LLM. Which read it, ignored the formatting, and moved on.

Caveman speak

So I told the sub-agents to talk like cavemen. This is the block that goes into every sub-agent prompt:

COMMUNICATION STYLE (CRITICAL — every response):
Reader = coordinator agent, not human. It already has alert context and your question. Never restate them.
Terse. Technical substance exact. Only fluff die.
Drop: articles, filler, pleasantries, hedging, narration of queries run, tool-limitation notes that do not change conclusion, recommendations/next steps (coordinator decides).
No headers, tables, bold, rules, emoji. Write nothing between tool calls.
Format: one bullet per finding:
- <UTC time> <signal>: <value> (baseline <value>) — <evidence>
State each value once. No timeline that repeats bullets.
Quote each error string once, shortest unique part, in backticks.
Keep ruled-out hypotheses, one bullet each: `- ruled out <X>: <evidence>`.
Keep evidence links as markdown links.
Last line: `Cause: <a> → <b> → <c>` or `Cause: unknown — missing <data>`.
Target ≤150 words.

The first line is the one I like most. "Be concise" means nothing until the model knows who it's writing for. Once it knows the reader is another agent that already has the alert and the question, it stops restating them.

I also kept one thing on purpose: ruled-out hypotheses. "The DB is fine" sounds like fluff, but it's evidence. Drop it, and the coordinator sends another agent to go check the DB. That costs way more than the eight words you saved.

And the coordinator? Its answer goes to a human and that human deserves full sentences. So I do not enable caveman mode here.

Caveman measure

I used my real production prompts and the same model we run in production (Claude Opus 5.5). I debugged a demo incident with a known answer: a ConfigMap change bumps the Kafka batch size of order-worker from 500 to 5000, the JVM runs out of heap, and the pods go into an OOMKill loop. To keep things honest I threw in two red herrings: a node with DiskPressure, and a service spitting out webhook 429s at the same time.

The logs, metrics and traces tools returned the same canned data every run, so the only thing changing was the prompt. Then a separate model (Claude Opus 5.5) graded every summary and every final answer against the real root cause.

Caveman show log

Here is one run of the logs agent, the middle one of three on each side. Same question, same data. "Why are order-worker containers restarting between 13:30 and 14:30 UTC? Check error logs, k8s events, and any config changes."

Without caveman - 1,331 tokens.

I'll investigate the order-worker restarts across logs and events.
# order-worker Restart Investigation — 2026-10-05 13:30–14:30 UTC

## Root Cause: ConfigMap change (max.poll.records 500 → 5000) → JVM heap exhaustion → OOMKill loop

### Timeline
| Time (UTC) | Event |
|---|---|
| **13:58:12** | **ConfigMap `orders/order-worker-config` updated by `ci-bot` via ArgoCD sync, commit `3be0c71` "tune throughput"** — resourceVersion 88121 → 88340. `KAFKA_MAX_POLL_RECORDS=5000` (previously **500**, a 10× increase) |
| 14:01:40 | First large batch: "Fetched 5000 records from orders-events-3 in 412ms" |
| 14:03:58 | WARN — GC overhead: 4120ms pause (G1 Full GC), heap **1498M/1536M** (97% full) |
| 14:04:49 | ERROR — `java.lang.OutOfMemoryError: Java heap space` at `OrderBatchProcessor.enrich(OrderBatchProcessor.java:142)` (pod `order-worker-6d9f-2hx7k`) |
| 14:04:51 | Event `OOMKilling` — killed process 1 (java), anon-rss 2.0GB, **exit code 137**, limit 2Gi |
| 14:05:30 | Restart: "Starting OrderWorker v4.8.1 (JVM **-Xmx1536m**)" — unchanged heap config |
| 14:10:30/33 | Second OOM + OOMKilled (pod `order-worker-6d9f-9kq2m`) |
| 14:12:02 | WARN — Consumer group `order-worker` rebalancing; member left (session timeout) |
| 14:16:38/40 | Third OOM + OOMKilled (pod `2hx7k`) |
| 14:17:02 | Event `BackOff` — "Back-off restarting failed container order-worker" → CrashLoopBackOff |
| 14:22:00 | Alert fires (7 restarts in 15m) |

### Error volume vs. baseline
- Window 13:30–14:30: **402 ERROR**, 1510 WARN (88,410 INFO)
- Prior day same window baseline: **3 ERROR**, 41 WARN → **~134× increase in errors**
- Error histogram (10m buckets): 13:50 = 0, 14:00 = 61, 14:10 = 158, 14:20 = 183 — errors start only *after* the 13:58 config change, and are still climbing.

### Mechanism
`OrderBatchProcessor.enrich()` (line 142) holds the entire polled batch in memory. 10× more records per poll exceeds the 1536m JVM heap inside a 2Gi container limit → OutOfMemoryError → cgroup OOMKill (137) → restart → re-consume same offsets → repeat. Each crash also triggers consumer-group rebalances, amplifying the loop.

### Co-firing alerts — assessment
- **Node ip-10-2-3-4 DiskPressure (14:05)**: event explicitly lists affected pods as `logging-agent, metrics-agent, search-indexer` — **no order-worker pods**. Appears **unrelated/coincidental**; order-worker deaths are exit-137 memory OOMKills, not eviction.
- **payments-webhook 429 rate**: no evidence in order-worker logs; likely a downstream symptom of the restart/retry storm or unrelated. Not confirmed here.

### Suggested next steps for remediation owner
1. Roll back ConfigMap to `KAFKA_MAX_POLL_RECORDS=500` (revert commit `3be0c71`), or
2. Raise container memory limit + `-Xmx` if the 5000 batch size is desired.

With caveman - 687 tokens:

- 13:58:12 config change: ConfigMap `order-worker-config` rv 88121→88340, `KAFKA_MAX_POLL_RECORDS=5000 (previous 500)` — pushed by ci-bot via argocd sync, commit 3be0c71 "tune throughput"
- 14:01:40 first large fetch: `Fetched 5000 records from orders-events-3 in 412ms`
- 14:03:58 GC pressure: `GC overhead: 4120ms pause (G1 Full GC), heap 1498M/1536M` (JVM -Xmx1536m, container limit 2Gi)
- 14:04:49 first ERROR: `java.lang.OutOfMemoryError: Java heap space` at `OrderBatchProcessor.enrich(OrderBatchProcessor.java:142)`; repeats 14:10:30, 14:16:38
- 14:04:51 / 14:10:33 / 14:16:40 OOMKilling events: `exit code 137 (OOMKilled), limit 2Gi`, pods order-worker-6d9f-2hx7k and -9kq2m; 14:17:02 `Back-off restarting failed container`
- ERROR volume 402 in 13:30–14:30 (baseline 3 prev day same window); WARN 1510 (baseline 41); errors start 14:00 bucket 61 → 158 → 183
- 14:12:02 secondary effect: `Consumer group order-worker rebalancing; member left (session timeout)`
- ruled out node DiskPressure: event at 14:05:12 on node/ip-10-2-3-4 lists pods logging-agent, metrics-agent, search-indexer — no order-worker pods

Cause: ConfigMap max.poll.records 500→5000 at 13:58 → 10x batch heap in OrderBatchProcessor.enrich exceeds -Xmx1536m/2Gi limit → OOMKilled exit 137 restart loop + CrashLoopBackOff

Do you see it kept the important stuff, and dropped the fluff?

Timestamps, error strings, commit hash, baselines, ruled-out hypotheses, and full causal chain are all still there. Title, table, repeated timeline, essay, and the redundant info: that's all gone.

Numbers no lie

The log above is one run of one agent. The table averages everything: the sub-agent rows cover 9 runs per side across the logs, metrics and traces agents, and the investigation rows cover 3 full alert investigations per side. The logs agent writes the longest caveman summaries, so the average comes out a bit lower than the example.

Tokens

Without caveman With caveman Change
Summary sent to coordinator 1,355 tokens 552 tokens −59%
Sub-agent LLM time 27.3 s 17.4 s −36%
Last turn of the sub-agent 16.9 s 6.8 s −59%
Cost $0.087 $0.064 −26%

Alert investigation

Without caveman With caveman Change
Cost $0.90 $0.70 −23%
Time 175 s 140 s −20%

Accuracy

Without caveman With caveman Change
Root cause correct 3/3 3/3 same
Blamed a red herring 0/3 0/3 same
Final answer score (out of 10) 8.7 8.7 same
Claims not backed by data, per sub-agent summary 3.0 1.3 −56%

Cost and speed got better, which was expected. What surprised me was the accuracy.

Agents hallucination reduced with shorter summaries. The reason was pretty clear. All the made-up stuff lived in the "Mechanism" essay, the guesses about co-firing alerts, the remediation advice - which are no longer part of the output.

And the engineer reading the final answer can't tell the difference. The coordinator still writes a full report in plain English:

13:58:12 — ci-bot applied an ArgoCD sync (commit 3be0c71, "tune throughput") that updated ConfigMap orders/order-worker-config, changing KAFKA_MAX_POLL_RECORDS from 500 to 5000. The worker hot-reloaded this value, so there was no pod rollout — the change took effect silently on running pods with no canary or restart gate.

Caveman learn

A few things I'd tell anyone trying this:

  • Tell the model who's reading. An agent writing for an agent can skip everything the reader already knows.
  • The fluff is in the formatting more than in the grammar. Headers and tables cost more than "the" and "a".
  • Do not compress the evidence. Timestamps, error strings and "ruled out X" all stay.
  • Leave the human-facing answer alone.

You try now

Want to see the coordinator and its cavemen dig through a real alert? Try the Oodle playground or in your own account.

Unga bunga 🪨