Shipping a public LLM endpoint without a surprise bill

The Scope tool on this site calls a frontier model, from an unauthenticated public route, with no login. Here is every guardrail holding that up.

Anadi11 min read

Answers“how to add an llm feature without runaway api costs”

The homepage of this site has a tool that takes a description of a problem and returns an architecture for it. It calls a frontier model. It has no login, no paywall and no email gate, and the URL is public. Anyone reading this can open a terminal and POST to it right now.

That is, on paper, the worst possible cost profile: an unauthenticated endpoint that spends real money per request, published on a marketing page whose entire job is to attract strangers. The standard advice — put it behind a login — would have killed the feature, because the feature is the proof of work. A visitor who has to sign up to see whether I can architect will not sign up.

It has been up without incident. That is not luck, and it is not because the traffic is small. It is five specific controls, each covering a different failure mode, plus one design decision that matters more than the other five combined and which almost nobody builds first.

Here is the whole arrangement, in the order a request meets it.

Fig. 01The request path
  1. 01

    Parse & cap

    Six length caps on public input

    400 / 413

  2. 02

    Config check

    No key, or no limiter

    503

  3. 03

    Rate limit

    5/hour/IP, sliding

    429

  4. 04

    Model call

    Schema-bound, 4k cap

    refused / failed

  5. 05

    Stream result

    NDJSON, then log

Every gate is cheap, synchronous, and sits before the paid call. The dashed branches are the codes a tripped gate returns; every one of them resolves to a hand-written scope on the client rather than an error state, which is what licenses the rest to be strict.

What one request is actually allowed to cost

Before the controls, the budget they are enforcing. Every number below is a constant in the code rather than an intention, which is the only kind of budget that holds.

Fig. 02Token budget for a single scope
Instructions and voice rules~700 tok

Fixed. Scoping rules, plus the constraints that keep the output from turning into brochure copy.

Two worked examples~360 tok

Interpolated from the hand-written scopes. The input line I am least willing to cut — they are what hold the register on the cheaper model.

Visitor's bottleneck≤600 chars

Hard-capped at the route boundary. A longer body is a 413, not a silent truncation.

Output≤4,000 tok

max_tokens. Output is the expensive half of most pricing, so this is the number that actually bounds the bill.

Worst case per request≈5,200 tokens
The input side is fixed by me and the output side is capped. The only part a visitor controls is 600 characters — about three per cent of the worst-case request.

Knowing that number is the point of the exercise. Multiply the worst case by your rate limit and you have the real ceiling on what one abusive IP can spend in an hour — for this endpoint, five requests, so about 26,000 tokens. That is a number I can look at and stop worrying about. "We have a rate limit" is not.

1. A hard per-IP rate limit, treated as required

Five requests per hour per IP, sliding window, backed by Redis.

Three details matter more than the number.

Sliding, not fixed. A fixed window lets someone burst the full allowance at the end of one window and again at the start of the next, so the effective rate at the seam is double the stated limit. With a five-per-hour limit that is the difference between five and ten, which is survivable; with a five-per-minute limit on an expensive model it is not. Sliding windows cost slightly more Redis work and remove the seam entirely.

The limiter is mandatory in production. If the rate limit store is not configured, the route refuses to serve at all:

if (process.env.NODE_ENV === "production" && !isRateLimitConfigured()) {
  return json({ error: "unconfigured" }, 503);
}
Fig. 03The choice that gets made wrong
What this route does

Fail closed

  • Missing limiter config in production → 503, no model call
  • The feature is loudly off, and the fallback covers the visitor
  • One misconfiguration costs a degraded page for as long as nobody notices
  • Recovery is setting an env var
The tempting default

Fail open

  • Missing limiter config → serve anyway, uncapped
  • The feature keeps working perfectly, which is why nobody notices
  • One misconfiguration is an uncapped public endpoint on a paid API
  • Recovery is a conversation with your provider about the bill

VerdictFail closed. A missing limiter is a deployment error, not a degraded mode.

Both branches are defensible in a design review. Only one of them survives someone rotating an Upstash token on a Friday afternoon and not noticing the app came back up without it.

The instinct to degrade gracefully is a good instinct that is wrong here. It is right for a recommendation widget and wrong for a spend gate, and the way to tell them apart is to ask what the failure costs. A missing recommendation costs a worse page. A missing spend gate costs money at whatever rate the internet decides.

The IP never lands in storage. It is hashed with SHA-256 and a server-side salt before it reaches Redis or Postgres. Correlating requests from one source does not require holding the address, and not holding it means the table is not personal data I have to manage, disclose, or delete on request.

2. A bounded output, always

max_tokens is 4,000. Output tokens are the expensive half of most model pricing — often three to five times the input rate — and an unbounded generation is an unbounded line item.

Setting it deliberately is also a design forcing function, which is the part worth stealing. If the response does not fit in the cap, the shape of what you are asking for is wrong, and that is much better discovered while writing the prompt than while reading an invoice. When I first wrote this schema it asked for five sections; two of them were never rendered by the panel. Generating fields nothing displays is pure cost, and the token cap is what made that obvious.

3. Structured outputs instead of parse-and-pray

The response is constrained to a JSON Schema at the API level, not requested in prose and parsed afterwards.

The cost argument for this is not obvious until you have shipped without it. Asking a model to "reply in JSON" produces valid JSON almost always, and the almost is the whole problem: every malformed response is either a retry — which is a second full-price call, at the worst possible moment, under a latency budget a visitor is watching — or a failure the visitor sees. Schema-constrained output turns a probabilistic contract into an enforced one, and retries stop being a budget line rather than becoming a smaller one.

There is a second-order benefit that took me longer to appreciate. Because the schema is a real object in the codebase, the render shape is the source of truth and the generated shape extends it:

export type GeneratedScope = ScopeResult & {
  issue: string;
};

Change what the panel renders and this file fails to compile. Without that, a field quietly stops being displayed, keeps being generated, and you pay for it on every request for months.

4. Reasoning effort held at the floor

The paid path runs at effort: "low". The free path holds thinkingLevel at its minimum.

Reasoning tokens are billed, and they are not free latency either. For this workload — one bounded generation against a fixed schema, with a visitor watching a cursor blink — high effort buys very little and costs on both axes. This is worth calibrating per feature rather than defaulting to maximum: the reflex to turn everything up is expensive, and for constrained generation it is frequently not even better.

Two implementation notes, both of which cost me time:

Thinking is left on, not disabled, on the paid path. With thinking disabled the model can leak <thinking> tags into the output, which then has to be stripped — and lowering effort had already bought back most of the latency, so disabling it was paying a correctness risk for nothing.

On Gemini, thinking cannot be turned off at all any more. Sending thinkingBudget: 0 on the 3.x models is a hard 400 with invalid argument, not a warning and not a silent clamp. The floor is a low thinking level. If you are porting a config forward from an older Gemini integration, this is the line that breaks.

5. A provider swap that costs one environment variable

Two implementations sit behind one interface: a paid path on Claude Opus 5 and a free-tier path on Gemini Flash.

Fig. 04The provider seam
POST /api/scoperoute handler
lib/scope/prompt.tsone prompt, both providers
lib/scope/provider.tsselects by credential
anthropic → claude-opus-5 (paid)gemini → gemini-3.6-flash (free tier)
lib/scope/errors.tsScopeRefusedError · ScopeUnavailableError
The route never learns which model answered. Both adapters export the same three symbols and throw the same two error classes, so the discrimination the route does — refused versus unavailable versus broken — works identically either way.

Selection is by credential, not by a mode flag: whichever key is present wins, ANTHROPIC_API_KEY wins if both are set, and SCOPE_PROVIDER pins one explicitly when you want to override that order.

This is a cost control and not merely an architecture preference. It means that if spend becomes a problem, the response is an environment variable change with no redeploy, rather than a refactor negotiated under pressure at the exact moment you are least able to think clearly. Having the cheap path built and tested before you need it is the difference between a switch and a project.

It does not abstract perfectly, and pretending otherwise would be the dishonest version of this post. The schema had to be translated: Gemini rejects additionalProperties outright, wants type names as an uppercase enum, and generates fields in the order it is given them, so the adapter carries a translation function and a propertyOrdering array. One definition of the scope shape, two dialects of it. I wrote about how that abstraction is put together separately.

The decision that matters more than all five

Every control above is a way of failing. Rate limited. Quota exhausted. Provider down. Response refused. No key configured. Five controls means five new ways for a visitor to hit a broken feature — and the more aggressively you tune them, the more often that happens.

So the actual invariant is this: every failure path resolves to a real result rather than an error.

Behind the live model sit thirteen hand-written scopes, composed by me, one per bottleneck the studio actually sells against. The client-side transport is written so that no path throws:

| What went wrong | Signal | What the visitor sees | | --- | --- | --- | | Over the hourly limit | 429 | A hand-written scope | | No key, or no limiter in prod | 503 | A hand-written scope | | Model declined | refused event | A hand-written scope | | Provider down, timeout, dropped stream | failed event | A hand-written scope | | Nothing | result event | A generated scope |

Note what is not in the right-hand column: a spinner that never resolves, a stack trace, a toast saying "something went wrong", an empty panel. There is no on-screen tell at all, which is a deliberate and slightly uncomfortable choice — I will come back to it.

This inverts the usual relationship between cost control and product quality. Normally the two fight: every cap you add is a worse experience at the boundary, so caps get set generously, so they stop being caps. With a credible fallback the fight ends, because tripping a control is no longer a visible failure. Five an hour is a limit I can set precisely because the sixth request still returns something real.

It also means the honest description of the feature is "an interactive scoping tool that usually runs live", not "an AI-powered tool". I am fine with that. The widget is the site's proof of work; it is never allowed to look broken, and a static-but-correct answer serves that better than a live one that sometimes is not there.

The short version

If you are putting an LLM behind a public endpoint:

  • Work out the worst-case token cost of one request, multiply by your rate limit, and look at the number. If you cannot produce that number, you do not have a cost control, you have a hope.
  • Rate limit per IP on a sliding window, and make the limiter mandatory rather than best-effort. Fail closed.
  • Hash the IP before it reaches storage. You need to correlate, not to identify.
  • Cap output tokens explicitly. Unbounded generation is an unbounded bill, and the cap doubles as a design check on what you are asking for.
  • Constrain the response with a real schema so malformed output does not become paid retries — and keep the schema tied to your render shape so dead fields fail the build instead of billing quietly.
  • Match reasoning effort to the actual task instead of defaulting to maximum.
  • Build the cheap provider path before you need it, behind the same interface.
  • Then make every one of those controls invisible by giving the feature something real to fall back to.

The last one is what lets the first seven be strict. Build it first, not last — in this system the fallback library predates the API call by four phases, and that ordering is the only reason the controls above could be set where they are.

02 / Contact

You already know what to build.

Send the two-paragraph version. You'll get a real technical response, not a calendar link and a discovery questionnaire.