Skip to content

Rate Limiting an MCP Server: Stopping Agent Loops from Scraping You

Published July 13, 2026

An agent will call a tool in a loop if the description is vague or the result is short. A nutrition server that answers search_foods is also a way to walk a catalog. The limit has to sit on the account, in the same bucket as any other call that account makes, or the client can alternate transports and reset the counter.

This post lists the caps that are configured in the product. It does not report a staging incident. We have not run an unbounded agent against this server and counted the calls, so there is no "it scraped N foods in M minutes" figure here. Publishing one would be fiction.

One bucket per user, per minute

The Redis key is rate_limit:user:{user id}:{minute}, built by one function that both the REST middleware and the MCP guard call. The minute is unix time // 60. The first increment sets a 60-second expiry.

A second API key on the same account does not get a second bucket. An MCP call after a REST call spends the same counter. That is the property that matters. If MCP used rate_limit:mcp:{user id} and REST used another key, a client could split traffic and double the allowance. We do not do that.

The paid MCP plan is configured at 20 calls per minute and 10,000 calls per month, with 25 foods per search. Those numbers are the plan, also written on MCP pricing. They are not a measurement of how fast a model can loop.

Trial is a status, not a second plan

A 7-day trial stays on the same plan row. The guard checks subscriptions.status = 'trialing' and then applies a tighter set:

  • 5 calls per minute
  • 200 calls total, not per month
  • 10 results per search
  • 0.25% distinct foods, computed in the guard because the feature column is an integer and cannot store 0.25
  • 10 photo analyses total

When the subscription becomes active, the next check uses the paid caps. We do not swap the user onto a different plan row to represent the trial. Swapping rows is how a webhook race leaves someone on the wrong limits.

If the trial-status lookup fails, the guard treats the caller as not trialing and uses the paid caps. If the trial call-count query fails, the used count is treated as 0 for that request. Both are fail-open. They keep a database blip from locking a paying path, and they are also a hole if the database is the thing you needed for the cap. That is the behavior in the code. It is not a claim that the hole has been measured under attack.

When Redis is down

If the Redis client is missing, or INCR throws, the rate check allows the call and trips a breaker. The product stays up when the counter is down. The monthly quota and the distinct-food check still read the database on the paths that use them. A Redis outage is not a free catalog export, and it is also not a hard stop. Say that plainly if you copy the pattern. A fail-closed rate limit is safer against a scrape and worse for a Tuesday outage. We chose the outage.

Photo calls have two extra caps on top of the per-minute bucket: 150 per month and 20 per day on the paid plan, 10 total during the trial. Those counts are rows in usage, not a second Redis key.

Concurrency is not the same as rate

Two semaphores sit in front of the work: 16 calls in flight for the process (MCP_MAX_CONCURRENCY) and 4 for one user (MCP_PER_USER_CONCURRENCY). Those are defaults in config, not a load-test result. They stop one account from occupying every worker while the rate limit is still allowing the minute. They do not replace the minute cap.

What is not a tool call

/mcp is on the skip list for the REST rate-limit middleware. Handshake traffic should not debit the food quota. /api/v1/mcp is on the billable prefix list, and each successful tool records its own endpoint, /api/v1/mcp/search_foods and so on. The usage chart labels those as MCP tool names.

An MCP-plan key that calls REST search, foods, calc, or vision gets 403. The response says the plan is MCP only and points at the docs and the REST pricing page. Account, billing, and auth routes still work, so the person can manage the subscription. The commercial header is rejected on the MCP guard before a rate token is taken. Personal use is the product. The short version is on the MCP page.

What the counter actually does

Take a paid limit of 20 and one user id. The key for the current minute is rate_limit:user:{id}:{minute}. INCR returns 1, then 2, then 3. The call is allowed while the returned count is less than or equal to 20. The 21st increment in that same minute is refused, and the error tells the caller to wait until the next minute boundary. A REST call and an MCP call both increment that key, because both call shared_user_rate_key. A unit test locks the format: user 7 and minute 100 produce rate_limit:user:7:100, and the REST middleware and the MCP guard both reference that helper. That test does not open 21 sockets against a running server. It checks that the two paths cannot drift into two keys. Do not read it as a load test.

The distinct-food cap is a different counter. It counts food ids recorded for the account this month and compares them with a percentage of the catalog. Paid is 2%, stored as the integer feature limit 2. Trial is 0.25%, which cannot be stored in that integer, so the guard applies 0.25 only while the subscription is trialing. Search results are also clamped: 25 on the paid plan, 10 on the trial, even if the model asks for more. limit on search_foods is documented as 1-25. Asking for 100 does not bypass the plan.

None of those sentences is a report of an agent we let run until it stopped. They are the branches in the guard. If you need a published scrape count, run the agent and write down the number. This article will not invent one for you.

Order of the checks

The guard does not start with the Redis increment. A commercial usage header is rejected first, so a rejected call does not consume a personal-use token and then fail. An account on a security hold is rejected next. Only then does the code decide trial versus paid and take the minute token.

Vision sits in the same minute bucket and then has its own counts. Paid is 150 photos in the month and 20 in the day. Trial is 10 photos total. Those counts are rows on /api/v1/mcp/analyze_food_photo, not a second Redis key. The size check on a photo looks at the base64 length before decode, and it raises a limit error. It should not be confused with the monthly photo cap, which counts calls that got that far.

The trial's 200 calls count usage rows whose endpoint matches the MCP tool prefix and whose status is below 400. A refused call is not a successful row. The per-minute Redis key still moves on attempts that reach the increment. They are two counters. A dashboard can show quota remaining while the current minute is already closed.

If you copy this, keep one function for the Redis key. The test that matters fails when someone inlines a different string. Ours checks that user 7 and minute 100 produce rate_limit:user:7:100, and that the REST middleware and the MCP guard both call that helper. It does not open 21 connections. Add that live test on a staging stack. Do not invent the result in an article.

What to tell the model

Rate-limit errors name the cap and the wait: Rate limit reached (20 calls/minute). Wait N seconds. The model can back off. A generic 429 with no wait sends it into an immediate retry, which is the loop you were trying to stop. Tool descriptions should also say to search once and reuse food_id. The schema for that is in the MCP docs.

A small client-side pause

This is not a server limit. It is the polite half of the same idea. If you are writing a script against the HTTP API, do not tight-loop.

import time

def after_limit(response):
    if response.status_code == 429:
        time.sleep(2)

The MCP server's own error already includes a wait in seconds. Prefer that number over a fixed sleep when you have it. The snippet above is only a fallback for a client that sees HTTP 429 from REST.

Frequently Asked Questions

Does each API key get its own rate limit?

No. The bucket is the user id and the current minute. A second key on the same account spends the same counter, and REST and MCP share it.

What are the paid MCP caps?

20 calls per minute, 10,000 calls per month, 25 foods per search, 150 photos per month, and 20 photos per day. The trial is 5 per minute, 200 calls total, and 10 photos total.

What happens if Redis is down?

The per-minute check allows the call. Monthly quota and distinct-food checks still use the database on the paths that read them. This is fail-open, chosen so an outage does not take the API down.

Where is the unbounded-agent case study?

It has not been run for publication. This article lists configured caps only. It does not invent a call count from a scrape.

← Back to all articles

Start building with the Calorie API

Get a free API key and access 4M+ foods with search, barcode lookup, and full macro data.