My AI Portfolio Reviewer, Part 2: How Fast an AI System Gets Expensive — A $5 Bill, 104K Tokens, and the Many Ways to Cut Token Costs

✍️ Writing · August 2026

claude-apillm-cost-optimizationgcpbigquerydata-engineeringrag

A $5 monthly Claude bill, a 104K-token spike, and everything I changed to bring it down. Part 2 of a 3-part series.

One paid call, traced end to end: the web_search tool (4 searches/stock, response content excluded, the 104K→37K token fix), into a model split by task (claude-opus-5 struck through, replaced by claude-sonnet-5 and claude-haiku-4-5 at effort: low), logged one row per call into the claude_usage BigQuery table, summed into a real cost of $0.76 per cycle, down from ~$5.

📖 Preface

Part 1 was about the system — the pipeline, the staging, why Claude only ever steps in once, for one real judgment call. This one is about the money side of that: what it actually costs to run that judgment call every month, what surprised me about the bill, what I changed, and what I learned along the way — some of it good, some of it not. Part 3 is still on hold, waiting on real hardware.

Same rule as before: everything here comes from real numbers, pulled straight from a table I built to track this. Let’s get into it. 💸

🎯 Intro

This didn’t start after the pipeline went live — it started while I was still building it. I was just curious what running Claude through this would actually cost, so before wiring anything into production I added some credits and started testing.

First mix-up: I assumed the API cost and my Claude subscription were the same thing. They’re not. The subscription is a flat monthly plan for using Claude the assistant — the API is billed separately, pay-as-you-go, and it’s what this pipeline actually runs on.

So I added $20 in API credits to test with. It was gone after a handful of test runs. I added another $10, thinking that would buy me more room to experiment — it lasted two or three calls. That’s when the real cost analysis started: not curiosity anymore, but “I need to know exactly where this money is going before I let this run unattended every month.”

💰 The bill showed up

Once I actually sat down to measure it, the picture got clear fast. Stage B is the one step in this whole pipeline where Claude actually looks at a stock and forms an opinion — the only place that spends real money. In its early version, it was running on Opus (the strongest, priciest model available), with the reasoning effort turned up to “medium”, and no limit yet on how much it could search the web for each stock. That setup worked out to about $5 for a full monthly cycle — roughly 4 to 6 stocks getting a Stage B review.

On paper, $5 isn’t much. But watching $20, then $10, disappear in a handful of calls during testing told a different story — this wasn’t a task that should be burning money that fast. So I built a claude_usage BigQuery table: one row per API call, logging model, input tokens, output tokens, and cost_usd, tagged with the run_id and stock symbol.

That’s the actual run_id=2026-08-20 cycle, straight out of the table — not a mockup:

Call Model Input tokens Output tokens Cost
IRCTC Sonnet 5 78,981 1,998 $0.178
ITC Sonnet 5 79,968 2,861 $0.189
JYOTIRES Sonnet 5 75,273 2,466 $0.175
KSOLVES Sonnet 5 51,484 1,509 $0.118
PIIND Sonnet 5 44,177 1,161 $0.100
gsr_lookup Haiku 4.5 9,867 138 $0.011
notes Haiku 4.5 410 108 $0.001
Full cycle — 7 calls $0.77

🔍 How the whole thing is tracked

The mechanism is simple:

That table is what everything below is built on. Every number past this point is a logged row I queried directly, not a memory or an estimate — which is also, not incidentally, what made it possible to write this article honestly instead of from recollection.

📊 First finding: input tokens dominate, and they can spike hard

Real per-token pricing (checked against the Claude API pricing page):

Model Input ($/M tokens) Output ($/M tokens)
Opus 5 $5 $25
Sonnet 5 $2 $10
Haiku 4.5 $1 $5

That felt like the fix. It wasn’t the fix — it was one blocklist entry. More on that below.

🎛️ The tuning decisions, in order

Put simply: from nearly $5 a cycle on Opus with the effort dial turned up, down to a real, logged $0.77–$1.24 per cycle today — roughly 3-6x cheaper, moving cycle to cycle with how much that month’s web searches happen to pull back.

The real end-to-end email footer, from an actual production run (run_id=2026-08-20, 5 stocks evaluated):

This run cost ~$0.76 (~₹66) in Claude API usage

That’s a cent under the $0.77 I get summing all 7 logged rows myself — the email’s own running total doesn’t include the free-standing GSR lookup call, just the Stage B and notes calls. A small gap, but worth being precise about rather than rounding it away.

The received monthly email, showing the real cost banner: this run cost $0.76 (₹66) in Claude API usage.

⚠️ Two things the tuning didn’t fix

1. New outliers keep appearing. The domain blocklist only stops a source after it’s already cost money once — so new ones just keep slipping through:

This isn’t a solved problem. It’s a queue of not-yet-caught ones.

2. A retry silently doubles the bill.

The pipeline degrades gracefully by retrying, and “gracefully” has a price tag nobody sees unless they’re looking at the log directly — not the email footer from a single pass.

🚫 Two “obvious” optimizations that made things worse

Not every lever pulls the direction you’d expect. Two are worth sharing because they’re the kind of thing that looks right on paper.

Switching Stage B fully to Haiku. Half of Sonnet’s sticker price — should be cheaper, right? Tested it directly:

Narrowing to an allowlist of trusted domains. For two thinly-covered small-cap stocks, I tried restricting search to a few trusted finance domains instead of blocking a few bad ones.

One more gotcha, if you’re using the web_search tool: moneycontrol.com and economictimes.indiatimes.com actively block Anthropic’s search crawler and get rejected outright by the API if you list them in either allowed_domains or blocked_domains. Found that one the hard way.

🐛 The bug that got through anyway

All of the above still didn’t prevent a real production incident:

Takeaway for anyone building an LLM-in-the-loop pipeline: exception handling that degrades gracefully is good practice, but it also means a real bug can hide behind a technically-successful run. Trigger it for real, periodically, and actually read what came back — don’t just check the HTTP status code.

🧠 Knowing when not to call the model

The pipeline’s most recent addition is an “Equity Allocation Parity” diagnostic:

That’s the real skill this project reinforced for me, more than any single tuning trick above: being a data engineer who leverages AI well isn’t about routing everything through a model. It’s architecting so the model only ever sits in the one place actual judgment is required, instrumenting that one place properly, and being willing to rip out an “obvious” optimization the moment the real numbers say it made things worse.

Claude didn’t make this pipeline disciplined. The staging, the cost table, the willingness to reverse a decision that looked right on paper — all of that had to exist first. Claude just got to be the one piece of judgment none of it could replace.

🔮 Where retrieval might go next

The domain blocklist is a reactive fix — it only stops a source after it’s already cost money once (see the outliers above). That’s because web search’s token cost is a function of whatever content the tool happens to pull back on a given day, and that’s fundamentally unbounded. On top of that, Stage B starts from zero every single cycle: it searches the open web for a stock it may well have evaluated before, with no memory of what it concluded last time, or why.

Retrieval-augmented generation (RAG) is the direction that could fix both problems at once — and it’s specifically the bounded-context property of RAG that matters here, not just “give the model memory”:

This is deliberately framed as a direction, not a result — nothing above is built yet. But the pieces are already close: BigQuery’s vector search and Vertex AI Vector Search are both already in reach of this stack, and Stage B’s own historical output is exactly the kind of grounded, first-party data a RAG system should be built on, rather than reaching for the open web for something the pipeline already figured out once. When it’s actually built, it’s getting its own write-up — not a footnote here.


Part 1 covers the system this all sits inside of — the goal, the architecture, and the reasoning behind the staging in the first place. Part 3 (on hold until there’s something real to show) covers putting the results on a physical screen. The RAG direction above will get its own dedicated article once it’s actually built.