Shipping a public LLM endpoint without a surprise bill
The Scope tool on this site calls a frontier model, from an unauthenticated public route, with no login. Here is every guardrail holding that up.
Answers“how to add an llm feature without runaway api costs”
The homepage of this site has a tool that takes a description of a problem and returns an architecture for it. It calls a frontier model. It has no login, no paywall and no email gate, and the URL is public. Anyone reading this can open a terminal and POST to it right now.
That is, on paper, the worst possible cost profile: an unauthenticated endpoint that spends real money per request, published on a marketing page whose entire job is to attract strangers. The standard advice — put it behind a login — would have killed the feature, because the feature is the proof of work. A visitor who has to sign up to see whether I can architect will not sign up.
It has been up without incident. That is not luck, and it is not because the traffic is small. It is five specific controls, each covering a different failure mode, plus one design decision that matters more than the other five combined and which almost nobody builds first.
Here is the whole arrangement, in the order a request meets it.
- 01
Parse & cap
Six length caps on public input
400 / 413
- 02
Config check
No key, or no limiter
503
- 03
Rate limit
5/hour/IP, sliding
429
- 04
Model call
Schema-bound, 4k cap
refused / failed
- 05
Stream result
NDJSON, then log
What one request is actually allowed to cost
Before the controls, the budget they are enforcing. Every number below is a constant in the code rather than an intention, which is the only kind of budget that holds.
Fixed. Scoping rules, plus the constraints that keep the output from turning into brochure copy.
Interpolated from the hand-written scopes. The input line I am least willing to cut — they are what hold the register on the cheaper model.
Hard-capped at the route boundary. A longer body is a 413, not a silent truncation.
max_tokens. Output is the expensive half of most pricing, so this is the number that actually bounds the bill.
Knowing that number is the point of the exercise. Multiply the worst case by your rate limit and you have the real ceiling on what one abusive IP can spend in an hour — for this endpoint, five requests, so about 26,000 tokens. That is a number I can look at and stop worrying about. "We have a rate limit" is not.
1. A hard per-IP rate limit, treated as required
Five requests per hour per IP, sliding window, backed by Redis.
Three details matter more than the number.
Sliding, not fixed. A fixed window lets someone burst the full allowance at the end of one window and again at the start of the next, so the effective rate at the seam is double the stated limit. With a five-per-hour limit that is the difference between five and ten, which is survivable; with a five-per-minute limit on an expensive model it is not. Sliding windows cost slightly more Redis work and remove the seam entirely.
The limiter is mandatory in production. If the rate limit store is not configured, the route refuses to serve at all:
if (process.env.NODE_ENV === "production" && !isRateLimitConfigured()) {
return json({ error: "unconfigured" }, 503);
}
Fail closed
- Missing limiter config in production → 503, no model call
- The feature is loudly off, and the fallback covers the visitor
- One misconfiguration costs a degraded page for as long as nobody notices
- Recovery is setting an env var
Fail open
- Missing limiter config → serve anyway, uncapped
- The feature keeps working perfectly, which is why nobody notices
- One misconfiguration is an uncapped public endpoint on a paid API
- Recovery is a conversation with your provider about the bill
VerdictFail closed. A missing limiter is a deployment error, not a degraded mode.
The instinct to degrade gracefully is a good instinct that is wrong here. It is right for a recommendation widget and wrong for a spend gate, and the way to tell them apart is to ask what the failure costs. A missing recommendation costs a worse page. A missing spend gate costs money at whatever rate the internet decides.
The IP never lands in storage. It is hashed with SHA-256 and a server-side salt before it reaches Redis or Postgres. Correlating requests from one source does not require holding the address, and not holding it means the table is not personal data I have to manage, disclose, or delete on request.
2. A bounded output, always
max_tokens is 4,000. Output tokens are the expensive half of most model
pricing — often three to five times the input rate — and an unbounded
generation is an unbounded line item.
Setting it deliberately is also a design forcing function, which is the part worth stealing. If the response does not fit in the cap, the shape of what you are asking for is wrong, and that is much better discovered while writing the prompt than while reading an invoice. When I first wrote this schema it asked for five sections; two of them were never rendered by the panel. Generating fields nothing displays is pure cost, and the token cap is what made that obvious.
3. Structured outputs instead of parse-and-pray
The response is constrained to a JSON Schema at the API level, not requested in prose and parsed afterwards.
The cost argument for this is not obvious until you have shipped without it. Asking a model to "reply in JSON" produces valid JSON almost always, and the almost is the whole problem: every malformed response is either a retry — which is a second full-price call, at the worst possible moment, under a latency budget a visitor is watching — or a failure the visitor sees. Schema-constrained output turns a probabilistic contract into an enforced one, and retries stop being a budget line rather than becoming a smaller one.
There is a second-order benefit that took me longer to appreciate. Because the schema is a real object in the codebase, the render shape is the source of truth and the generated shape extends it:
export type GeneratedScope = ScopeResult & {
issue: string;
};
Change what the panel renders and this file fails to compile. Without that, a field quietly stops being displayed, keeps being generated, and you pay for it on every request for months.
4. Reasoning effort held at the floor
The paid path runs at effort: "low". The free path holds thinkingLevel at
its minimum.
Reasoning tokens are billed, and they are not free latency either. For this workload — one bounded generation against a fixed schema, with a visitor watching a cursor blink — high effort buys very little and costs on both axes. This is worth calibrating per feature rather than defaulting to maximum: the reflex to turn everything up is expensive, and for constrained generation it is frequently not even better.
Two implementation notes, both of which cost me time:
Thinking is left on, not disabled, on the paid path. With thinking disabled
the model can leak <thinking> tags into the output, which then has to be
stripped — and lowering effort had already bought back most of the latency, so
disabling it was paying a correctness risk for nothing.
On Gemini, thinking cannot be turned off at all any more. Sending
thinkingBudget: 0 on the 3.x models is a hard 400 with invalid argument,
not a warning and not a silent clamp. The floor is a low thinking level. If you
are porting a config forward from an older Gemini integration, this is the line
that breaks.
5. A provider swap that costs one environment variable
Two implementations sit behind one interface: a paid path on Claude Opus 5 and a free-tier path on Gemini Flash.
Selection is by credential, not by a mode flag: whichever key is present wins,
ANTHROPIC_API_KEY wins if both are set, and SCOPE_PROVIDER pins one
explicitly when you want to override that order.
This is a cost control and not merely an architecture preference. It means that if spend becomes a problem, the response is an environment variable change with no redeploy, rather than a refactor negotiated under pressure at the exact moment you are least able to think clearly. Having the cheap path built and tested before you need it is the difference between a switch and a project.
It does not abstract perfectly, and pretending otherwise would be the dishonest
version of this post. The schema had to be translated: Gemini rejects
additionalProperties outright, wants type names as an uppercase enum, and
generates fields in the order it is given them, so the adapter carries a
translation function and a propertyOrdering array. One definition of the
scope shape, two dialects of it. I wrote about how that abstraction is put
together separately.
The decision that matters more than all five
Every control above is a way of failing. Rate limited. Quota exhausted. Provider down. Response refused. No key configured. Five controls means five new ways for a visitor to hit a broken feature — and the more aggressively you tune them, the more often that happens.
So the actual invariant is this: every failure path resolves to a real result rather than an error.
Behind the live model sit thirteen hand-written scopes, composed by me, one per bottleneck the studio actually sells against. The client-side transport is written so that no path throws:
| What went wrong | Signal | What the visitor sees |
| --- | --- | --- |
| Over the hourly limit | 429 | A hand-written scope |
| No key, or no limiter in prod | 503 | A hand-written scope |
| Model declined | refused event | A hand-written scope |
| Provider down, timeout, dropped stream | failed event | A hand-written scope |
| Nothing | result event | A generated scope |
Note what is not in the right-hand column: a spinner that never resolves, a stack trace, a toast saying "something went wrong", an empty panel. There is no on-screen tell at all, which is a deliberate and slightly uncomfortable choice — I will come back to it.
This inverts the usual relationship between cost control and product quality. Normally the two fight: every cap you add is a worse experience at the boundary, so caps get set generously, so they stop being caps. With a credible fallback the fight ends, because tripping a control is no longer a visible failure. Five an hour is a limit I can set precisely because the sixth request still returns something real.
It also means the honest description of the feature is "an interactive scoping tool that usually runs live", not "an AI-powered tool". I am fine with that. The widget is the site's proof of work; it is never allowed to look broken, and a static-but-correct answer serves that better than a live one that sometimes is not there.
The short version
If you are putting an LLM behind a public endpoint:
- Work out the worst-case token cost of one request, multiply by your rate limit, and look at the number. If you cannot produce that number, you do not have a cost control, you have a hope.
- Rate limit per IP on a sliding window, and make the limiter mandatory rather than best-effort. Fail closed.
- Hash the IP before it reaches storage. You need to correlate, not to identify.
- Cap output tokens explicitly. Unbounded generation is an unbounded bill, and the cap doubles as a design check on what you are asking for.
- Constrain the response with a real schema so malformed output does not become paid retries — and keep the schema tied to your render shape so dead fields fail the build instead of billing quietly.
- Match reasoning effort to the actual task instead of defaulting to maximum.
- Build the cheap provider path before you need it, behind the same interface.
- Then make every one of those controls invisible by giving the feature something real to fall back to.
The last one is what lets the first seven be strict. Build it first, not last — in this system the fallback library predates the API call by four phases, and that ordering is the only reason the controls above could be set where they are.
