Traditional denial of service tries to take your app down. On modern pay-per-use infrastructure there is a nastier variant: the app stays up and the bill goes vertical. It is called Denial of Wallet, and for a solo founder it can be worse than downtime.
This post covers three documented bills that each passed $80,000, the code patterns that produce them, and what the major platforms actually do when you set a "budget". The short version: most budgets are emails, and the caps that do exist are usually opt-in.
What is a denial of wallet attack?
A denial of wallet attack is any flood of requests, hostile or accidental, that targets your bill instead of your uptime. On infrastructure billed per invocation, per gigabyte or per token, the platform scales to absorb the load exactly as designed, nothing crashes, nobody is paged, and you pay for all of it.
When you pay per invocation, per request, or per token, cost becomes an attack surface. A function that calls itself, a retry with no backoff, or an unbounded fan-out does not error out. It scales, exactly as designed, and each scale-up is a line on your invoice.
OWASP uses the term too. Its Top 10 for LLM Applications lists Denial of Wallet (DoW) under LLM10:2025, Unbounded Consumption: an attacker generates a high volume of operations to exploit pay-per-use pricing until the cost is unsustainable.
An attacker does not need to breach you to hurt you. They just need to make you pay for traffic.
Can one bug really run up a $10,000 bill?
Yes, and the documented cases go well past it. A static site on Netlify drew a bill of about $104,000 in February 2024. The artist platform Cara was billed $96,280 by Vercel for a single week in June 2024. A stolen Gemini API key cost a three-person startup $82,314.44 in 48 hours in February 2026.
Bandwidth: the $104k static site. On 27 February 2024 a Hacker News thread titled "Netlify just sent me a $104k bill for a simple static site" described a free-tier site that had served 190TB of bandwidth in four days. Support first offered to reduce the bill to 20 percent, then to 5 percent, which is still more than $5,000. Once the thread spread, Netlify's CEO replied in it that the user would not be charged, and said the company's policy was to forgive bills from legitimate mistakes after the fact.
Compute: the $96,280 week. TechCrunch reported on 6 June 2024 that Cara, a portfolio and social app for artists founded by Jingna Zhang, grew from 40,000 to 650,000 users in a week. Zhang then posted that her Vercel bill for that week would come to $96,280. There was no attacker here. It was usage-based hosting doing what it was configured to do, for a crowd nobody had modelled.
LLM tokens: $82,314.44 in 48 hours. The Register reported on 3 March 2026 that a three-developer startup in Mexico, which normally spent about $180 a month, had a Google Cloud API key compromised on 11 and 12 February. Whoever held the key spent $82,314.44, mostly on Gemini 3 Pro image and text models. Google's representative pointed to the shared responsibility model, and the bill was still unresolved when the article ran. This is why a leaked API key is now a billing problem as much as a data problem.
The Netlify bill was waived, but only after it went public. Counting on that is not a plan.
The code patterns that cost money
Most runaway bills need no outside traffic. These are the shapes to look for in your own code:
- Recursive functions or event loops with no stop condition.
- Retries without exponential backoff or a cap, turning one failure into a storm.
- Unbounded fan-out, where one request spawns thousands.
- Expensive AI or third-party calls with no rate limit or budget guard.
- Missing idempotency, so a client retry doubles the work and the cost.
Here is the version that shows up most often in AI-assisted code, because it makes the error disappear during testing:
async function summarize(text) {
while (true) {
try {
const res = await client.messages.create({
model: MODEL,
max_tokens: 64000,
messages: [{ role: "user", content: text }],
});
return JSON.parse(res.content[0].text);
} catch (err) {
// bad JSON or a timeout: try again
}
}
}
If the model answers in prose instead of JSON, the call has already been billed, the parse throws, and the loop buys another one. There is no attempt limit, no delay, no input limit, and a ceiling of 64,000 output tokens per attempt. One awkward request is enough. The bounded version is not much longer:
const MAX_ATTEMPTS = 3;
const MAX_INPUT_CHARS = 20000;
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
async function summarize(text) {
if (text.length > MAX_INPUT_CHARS) throw new Error("Input too long");
for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
try {
const res = await client.messages.create({
model: MODEL,
max_tokens: 1024,
messages: [{ role: "user", content: text }],
});
return JSON.parse(res.content[0].text);
} catch (err) {
if (attempt === MAX_ATTEMPTS) throw err;
await sleep(2 ** attempt * 500 + Math.random() * 250);
}
}
}
Four changes: a hard attempt limit, exponential backoff with jitter, an input cap, and an output cap. On the Anthropic Messages API max_tokens is a required parameter, so the question is never whether you set it but how high. Size it to the feature, not to the model maximum.
AWS has built one guardrail for the first pattern on the list. According to the Lambda documentation, recursive loop detection is on by default and stops a function after roughly 16 invocations in the same chain of requests. It only covers loops through SQS, SNS, S3, EventBridge custom event buses and Lambda itself. A loop through DynamoDB is not detected.
Do AWS budgets stop you from being charged?
No. AWS Budgets tracks your spend and notifies you. It does not cap your account. The AWS documentation says budget information is updated up to three times a day, typically 8 to 12 hours apart, and warns that you can run past a threshold before the notification arrives. Budget actions exist, but they are narrow.
Per the Budgets documentation, an action can apply an IAM policy, apply a service control policy, or target specific EC2 or RDS instances. Those are useful for stopping new provisioning. They are not a pause button for a function that is already being invoked a thousand times a second.
The reporting gap matters. The Gemini incident above burned roughly $1,700 an hour, and at that rate a ten-hour gap between updates is about $17,000 before the first email. Set a budget anyway, and pair it with the tools the Lambda docs recommend: CloudWatch alarms on invocations and concurrency, a billing alarm, and AWS Cost Anomaly Detection.
Which cloud platforms have a real spending cap?
Few do by default. Going by each vendor's documentation in October 2026, AWS, Google Cloud and Cloudflare budgets are alerts only. Vercel, Firebase, OpenAI and Anthropic offer limits that actually stop usage, but most are opt-in, and the vendors warn that enforcement lags by minutes, so spend can overshoot the number you set.
- AWS. Budgets are alerts plus the optional actions described above.
- Google Cloud. The Cloud Billing docs caution that an alerts-only budget does not automatically cap usage or spending, and that cost data arrives with a delay.
- Firebase. Budget alerts do not pause services. Spend cap budgets, currently in Preview on the Blaze plan, pause a service for the rest of the month at 100 percent of its budget, but only for Firebase AI Logic, App Hosting, Cloud Functions for Firebase and Extensions. Firebase states that these are not hard caps, because enforcement can trail usage by several minutes.
- Vercel. Spend Management is available on Pro and on Enterprise with Flexible Commitment, and a changelog entry dated 9 September 2025 made it on by default for new Pro teams. It notifies at 50, 75 and 100 percent, but per the docs a spend amount does not stop usage on its own: you have to turn on Pause Production Deployments, and the check runs every few minutes.
- Cloudflare. Budget alerts arrived in April 2026 for pay-as-you-go accounts. The docs describe them as informational only: "They do not pause or cap usage."
- OpenAI. Spend limits can be a soft alert or, if you switch on enforcement, a hard limit that makes API requests fail with a 429. The guide notes that enforcement is not instantaneous.
- Anthropic. Each usage tier carries a monthly spend cap ($500 on Start, $1,000 on Build), and you can set your own lower limit for the organization or for a workspace in the Console. Requests are rejected once the limit is reached.
Wherever a real cap exists, switch it on and set it below the most you could stand to lose. Wherever it does not, the cap has to live in your own configuration.
Capping the blast radius in your own config
Provider budgets tell you what already happened. These settings limit what can happen.
Concurrency and timeouts
The default Lambda quota is 1,000 concurrent executions per Region (brand-new accounts start lower), and without a per-function limit all of that is available to your most expensive function. Reserved concurrency sets a per-function ceiling, and the Lambda docs confirm it costs nothing extra:
aws lambda put-function-concurrency \
--function-name generate-report \
--reserved-concurrent-executions 10
Setting the value to 0 throttles the function completely, which makes it your kill switch. Pair the ceiling with a short timeout, and declare both next to the function so the limit ships with the code:
resource "aws_lambda_function" "report" {
# ...
timeout = 10
reserved_concurrent_executions = 10
}
If you are editing Terraform anyway, the neighbouring defaults deserve a look too. They are covered in the Docker and Terraform traps.
Rate limits and quotas
Rate-limit anything that triggers a paid downstream call, per user and per IP. For LLM features add a per-user daily quota stored server-side, because a rate limit alone still lets one patient account spend all day. Require authentication on every endpoint that costs money to serve.
Retries and queues
Cap every retry and add backoff. Put a dead-letter queue on the source queue with a low maxReceiveCount, so a poison message is parked instead of being reprocessed indefinitely. Make handlers idempotent, so a client retry does not double the work and the cost.
The checklist to run before launch
- Turn on every real spend limit your providers offer: Vercel's Pause Production Deployments, Firebase spend caps, an OpenAI hard limit, an Anthropic spend limit.
- Create a budget and a billing alarm everywhere else, with alerts going to a channel you actually read.
- Set reserved concurrency and a timeout on every function that calls something paid.
- Rate-limit and authenticate every endpoint that triggers a paid call.
- Cap
max_tokensand input size on every LLM request, and add a per-user quota. - Search your code for
while (true), recursive invokes and retry blocks. Each needs a counter and a backoff. - Check that no function writes to the same queue, bucket or table that triggers it.
- Confirm that no provider key is in your client bundle or your git history.
For the rest of the launch surface, work through the checklist to run before you make your backend public. If an assistant wrote most of your backend, the security risks of vibe coding explains why the unbounded retry shows up so often. The full set of app security guides is collected in one place.
Scanning for this before you ship
Review your code for the loop, retry and fan-out patterns before you deploy, because the first time you find them should not be on your credit card statement. If you would rather have something read your source first, the free IOnclad browser scanner runs a 167-rule subset entirely in your browser, with no signup. Your code stays in the tab.
The IOnclad desktop app runs 24 scanners and 500+ checks across secrets and git history, dependencies, the OWASP web categories, session and auth, an API surface map, and a denial-of-wallet check. It returns one Ship-It verdict with the file, line and a fix for each finding. It does not promise your app is secure, and no scanner honestly can. The spending cap itself is still a setting in your provider's dashboard, and only you can switch it on.