optimAIzr docs
Find where your LLM spend is wasted, then verify and apply the savings. Everything runs on your machine: no account, no upload.
npx optimaizr profile- Step 1Install
One global package. Node 20.11 or later, no runtime deps.
- Step 2Profile
optimaizr profilereads your Claude Code and Codex sessions. - Step 3Follow the lead
It names the biggest bottleneck and the command to run next.
See your profile in one command
If you use Claude Code or Codex there is nothing to set up. optimAIzr reads their session transcripts straight from disk and answers three questions: how much you use AI, where you waste the most, and what to look at next.
optimaizr profile AI usage
Spend $42.18 $46.90/month at this rate
Calls 1,842
Tokens 18.4M 18.1M in / 312K out
Optimization
Flagged calls 214 11.6% of calls
Potential waste $13.72 in this window
Potential savings $10.47/mo $127.40/year
Biggest opportunity
! Oversized context
Small jobs are inheriting a whole session's context.
Next step
optimaizr simulate oversized-inputIllustrative figures. Yours come from your own usage.
Agents, apps, or an export
Reads ~/.claude/projects and ~/.codex/sessions. Anthropic and OpenAI agent spend land in one report.
Wrap an Anthropic or OpenAI client. Calls keep their signature, and recording can never break a request.
import optimaizr from "optimaizr";
const claude = optimaizr.wrap(new Anthropic(), { service: "checkout-api" });
const openai = optimaizr.wrap(new OpenAI(), { service: "checkout-api" });Label call sites to get per-route advice:
import { withRoute } from "optimaizr";
await withRoute("summarise-ticket", () => client.messages.create({ ... }));Already have a usage export? Column names are matched loosely, a row's own cost is trusted over our arithmetic, and unpriceable rows are flagged, never dropped.
optimaizr import usage-export.csv --service billingTwenty commands, all free
optimaizr profile- Usage, waste and your biggest bottleneck, on one screen
optimaizr audit- What you spend, what is recoverable, and why
optimaizr scan- The full spend report, with next steps
optimaizr live- Watch calls as they land and surface fixes
optimaizr waste- Just the opportunities
optimaizr tokens- Token analytics and the priciest calls
optimaizr why [path...]- Drill down: provider, model, project, workload
optimaizr show <rule>- The requests a recommendation touches
optimaizr guide- Which model for which job, with the arithmetic
optimaizr recommend- Ranked actions with impact and confidence
optimaizr simulate <rule>- What the change would cost, arithmetically
optimaizr verify <rule>- Prove a fix against your quality bar
optimaizr apply <rule>- Get the exact change, once verified
optimaizr import <file>- Load a CSV or JSON usage export
optimaizr limit- Record a Claude session-limit hit, so --plan learns it
optimaizr card- Your last 30 days as an image to post, totals only
optimaizr report- Write a self-contained HTML dashboard
optimaizr providers- What can be read, and from where
optimaizr privacy- What is collected, stored and sent
optimaizr metrics- How much has been analysed, and found
Flags
--why- Show every calculation and assumption
--budget N- Monthly cap in USD: the day it runs out, and what buys days back
--plan pro|max5|max20- Read usage as a Claude subscription: sessions and your limit
--days N- Only the last N days
--project STR- Filter by project path
--source all|agents|sdk- Which data to read
--tz utc|local|<IANA>- Which midnight day buckets use (default UTC)
--json- Machine-readable output
live also takes --backfill N, --window N, --min-usd N and --no-prompt. With --json it never prompts.
When you answer live
Y says what it actually did. On a Claude Code model swap it edits your settings and tells you how to undo it. When no lever exists it says so, records the decision, and hands you optimaizr verify, rather than claiming to have changed a call that was billed before it ever reached the screen.
It declines when the change is too blunt. Under 80% of your traffic affected, a global model change would downgrade the rest, so it refuses, prints the share, and points at the per-agent override instead.
From a hunch to a shipped change
profileWhere am I wasting the most?whyWhere does the money go?recommendWhat can I change?simulateWhat would that save?verifyWould the output still be good?applyThe exact change, once verified
simulate is arithmetic on your recorded tokens and says nothing about quality. verify replays your traffic and scores it. apply never edits your code: it writes the change to optimaizr-change.md for you to review.
Every figure says what it is
Provider-reported usage, priced at the rate in force. A fact.
A pattern derived from measured data by a stated rule.
A projection of how a different model or setting would behave. Could be wrong.
Confidence is about the money: how far to trust the dollar figure (low, medium or high). Impact is about the output: a saving can be arithmetically certain and still be a bad idea, which is what verify is for. Add --why to any report to see each calculation and assumption.
When two findings touch the same call, the headline counts that call once, at its largest claim. The total is a floor, and can sit below the sum of the findings.
A cap that runs out, or a plan that stops you
A company cap. If you get a fixed amount per month, reset on the 1st, give it the cap and it answers in dates: when the cap runs out at this pace, and how many days the fixes buy back.
Monthly cap $300.00 resets Oct 1 (UTC)
Used this month $246.40 ████████████████░░░░ 82% · day 22 of 30
At this rate Sep 27 cap reached 4 days before reset
With the fixes below Sep 28 +1 day · estimatedoptimaizr live --budget 300 warns once each at 50%, 80%, 95% and 100%. The month is cut in the --tz zone, UTC by default. It counts calls on this machine only, so usage elsewhere brings the real cap sooner.
Claude Pro or Max. On a flat plan the dollars are what the same work would cost on the API, which is still the right measure of a session: a larger model drains one faster in roughly the proportion it costs more.
API-equivalent $443.55/mo 22x what you pay
5-hour sessions 18 in the last 30 days
Waste per session 10% fix it and you'd hit the limit ~11% later
Your session limit ~$26.90 learned from 3 recorded hits
This session $12.40 since 20:00, resets 01:00 · ~46% of itAnthropic does not publish the limit, so it is learned: run optimaizr limit when Claude says you have hit it (--at 15:10 after the fact, undo to take one back). live --plan pro then warns at 80% and 95% of it. Sessions are rebuilt from timestamps, five hours from the top of the hour.
ChatGPT Plus or Pro. Nothing to set: Codex records OpenAI’s own limit meter in its session files, so profile shows your plan, the share of the 5-hour and weekly windows used and when each resets, next to the waste. live warns at 80% and 95% of each.
ChatGPT Plus $20.00/mo what you pay
API-equivalent $117.72/mo 5.9x what you pay
5-hour window 82% used · resets 01:49
Weekly window 38% used · resets Sep 25, 22:49Set either once in config: { "budget": 300 } or { "plan": "pro" }. ChatGPT plans need no setting.
Cheaper, and still good
Anyone can tell you a cheaper model is cheaper. verify checks it still does the job, on your own traffic, before you change anything.
PASS 40 samples replayed
cost/call $0.0121 -> $0.0034
projected $61.40/mo saved
this check cost you $0.38- Replay, not simulation. Recorded requests are re-sent under the change and compared with the production response.
- Deterministic checks first. JSON parses, required fields exist, the same tool is called, length stays in budget.
- Then a pairwise judge, each pair judged twice with positions swapped, so position bias cannot pick a winner.
- A failure is a result. If the cheaper option is worse, you are told to keep what you have.
Verification needs prompts, so turn on capture: opt-in, sampled, redacted, local.
optimaizr.wrap(client, { service: "checkout-api", capture: { rate: 0.02 } });Fixes that change behaviour without being a request rewrite (giving small jobs their own context, say) cannot be replayed. apply <rule> --accept-risk records your sign-off instead.
Thirteen detectors
cache-churnA prefix that changes between calls, so caching never pays offrepeat-tool-callsFiles re-read inside one session, re-billed on every later callmodel-fitMechanical turns running on a model priced for reasoningreasoning-effortHeavy deliberation that produced almost no outputoversized-tool-outputTool results large enough to distort a whole sessionerror-loopsThe same failing command retried unchangedprompt-bloatAn oversized system prompt paid for on every callrepeated-contextThe same prefix re-sent across separate sessionsoversized-inputSmall jobs inheriting a whole session's contextoversized-outputResponses far longer than the medianspend-concentrationA few workloads driving most of the billcost-spikeA sudden increase that call volume does not explainpricing-changeRates changing under youFindings are safe to apply (pure waste removal) or need verification (could change output). Rate increases you have to absorb are reported separately and never counted as savings.
Where naive cost tools go wrong
- Transcripts double-count. Claude Code stamps every content block with the same usage; summing rows inflates spend about 1.9x. optimAIzr keeps one per response.
- Images are priced by pixels. Real dimensions come from the image header, not from the base64 length.
- Prices are dated. Each call is costed at the rate in force that day, including cache reads, cache writes and batch discounts.
- Per-request charges count. Web search is billed per search and priced as its own line.
- “At this rate” is the recent rate. Past 30 days of history, monthly figures project the last 30, not an average that lets April dilute September.
- Days need a timezone. Buckets default to UTC; pass
--tz localor any IANA zone.
Your quality bar, your prices
Read from optimaizr.config.json or .optimaizr.json in the current directory, then ~/.optimaizr/config.json.
{
"qualityBar": {
"sampleSize": 40,
"checks": [
{ "type": "json-parses" },
{ "type": "contains", "value": "summary" },
{ "type": "max-chars", "value": 2000 }
],
"judge": {
"model": "claude-opus-5",
"criteria": "Which response is more accurate and complete?",
"minWinRate": 0.45
}
}
}Prices are data. Add a model, a price change or a whole provider with no code change:
[
{
"id": "acme-turbo-1",
"label": "Turbo 1",
"provider": "acme",
"tier": "fast",
"rates": [{ "from": "2026-01-01", "inputPerM": 0.5, "outputPerM": 1.5 }],
"contextTokens": 128000,
"maxOutputTokens": 8000,
"capabilities": ["tools"],
"bestFor": "Cheap bulk work."
}
]Gemini is built but off until its rate cards are verified. Enable it with OPTIMAIZR_GEMINI=1.
Nothing leaves your machine
- Analysis is entirely local. No account, no telemetry, no upload.
- Data lives in
~/.optimaizr/as plain JSONL. Delete the folder to remove everything. - The only network call is
verify, to your own provider with your own key, which is never stored or logged. - Prompt contents are never stored unless you turn on capture.
- HTML reports are self-contained and fetch nothing when opened.
Run optimaizr privacy for the full answer, or read the privacy policy. verify needs the optional @anthropic-ai/sdk or openai package, and reads ANTHROPIC_API_KEY or OPENAI_API_KEY.