Skip to content

Giving an AI Agent a 4M-Food Nutrition Database over MCP

Published August 17, 2026

Handing an agent four million rows is not a feature. It is a context-window problem waiting to happen.

Every tool reply an agent reads is tokens in a context that also has to hold your conversation, your profile, your log, and whatever else the client injected. A nutrition catalog can return a great deal of data per food: dozens of nutrients, provenance, serving descriptions. Return all of it, on twenty-five results, and you have spent a meaningful part of the window on one search that the agent needed four numbers from.

So the interesting question about a food database over MCP is not how big it is. It is what a single response costs.

What the catalog is

Roughly four million foods, made of a few distinct populations, and the difference between them matters more than the total:

  • Whole and generic foods. Chicken breast, brown rice, olive oil. Well covered, stable, and this is what most home cooking resolves to.
  • Branded packaged products, addressable by barcode. Large but never complete: new products appear weekly and regional own-brand stock is where coverage thins.
  • Restaurant items for chains that publish nutrition data.
  • Rows with provenance tracking, which is what the verified_only flag on search_foods filters to.

A number like four million invites the assumption that everything is in there. It is not, and the honest framing is coverage by category rather than a headline count. What matters is whether your shopping basket resolves, and that varies by country more than by anything else. The data-quality machinery behind the verified flag is described in the verified foods guide.

Result caps are a feature, not a restriction

search_foods takes a limit between 1 and 25. The plan caps it: 25 foods per search on the paid plan, 10 on the trial.

Read that as a design decision rather than a paywall. An agent asked for chicken breast does not benefit from a hundred rows. It benefits from a handful it can choose between, because:

  • A hundred near-identical rows is a harder disambiguation problem than five, not an easier one.
  • Every extra row is tokens that crowd out the conversation.
  • Long result lists make a model more likely to pick a row whose name merely looks closest, rather than asking you which one you meant.

The useful behaviour is a small, well-ordered result set plus a follow-up question when the top hits are genuinely ambiguous. If an assistant silently picks row one every time, ask it to show you the candidates.

Compact responses, and why the shape is specific

search_foods returns, per food, an identifier and macros per 100 g: calories, protein, fat, carbohydrate. Not the full nutrient panel.

That is the deliberate split. Four numbers are what a food log needs, and they are what ninety-odd percent of questions need. The full panel (sodium, fibre, sugars, vitamins, minerals) is a separate call, get_food_nutrition, on a food_id you already have.

The alternative design returns everything on every search, which is more convenient for exactly one request and worse for every subsequent one. Twenty-five foods with a full nutrient panel each is a large response; twenty-five foods with four macros each is a small one. Compacting tool output is the highest-leverage thing you can do to an MCP server's usability, because it buys back context for the work.

Micronutrient coverage generally, and how sparse it gets on branded rows, is covered in the vitamin and mineral guide.

The food_id is the joint in the whole design

search_foods is the only tool that produces a food_id. get_food_nutrition, calculate_portion and calculate_recipe all consume one.

This is the constraint that makes the numbers trustworthy, and it is also the one that produces the most confusing failure. A model that has seen a few ids will sometimes produce something id-shaped without calling search. The next call errors. That is correct, because the alternative is scaling a fabricated row, and silently wrong beats loudly wrong only if you enjoy bad data.

Mitigations, in order of how well they work: the tool descriptions say ids come from search; the error message says the id was not found; and a line in your profile file telling the assistant to search first is the one that actually fixes it in practice. The general problem of writing schemas a model uses correctly is the subject of designing MCP tools an LLM actually calls correctly.

How to search a catalog this size

Four million rows means almost any query returns something. That is the risk: you very rarely get nothing back, so a bad query produces a confident wrong row rather than an obvious failure.

Three things that change the result quality more than anything else.

Qualify the state of the food. "Chicken" is ambiguous in every direction that matters: breast or thigh, raw or cooked, skin on or off, and each of those is a different row with substantially different macros. "Cooked chicken thigh, skinless" is a query with one sensible answer. The same applies to rice, oats, pasta and lentils, where dry and cooked differ enormously per gram.

Use verified_only when precision matters more than coverage. The flag restricts results to rows with provenance behind them. For a staple you will eat all year, that is the right trade. For an obscure own-brand product, restricting it will just return nothing useful, and a barcode lookup is the better route.

Use suggest_foods when you are exploring. It returns short autocomplete-style suggestions for a partial name. It is cheaper in tokens than a full search and it is the right tool for "what is in the catalog under this name", as distinct from "give me the row for this food".

When the top hits are genuinely ambiguous, the correct behaviour is a question, not a pick:

You: Show me the top five with their per-100 g calories before you choose one.

Seeing five rows side by side makes the wrong choice obvious in a way that a single confident answer never does. If two rows differ by 80 kcal per 100 g, you want to know that before you scale one of them to 300 g.

What a verified row means, and does not

The verified_only flag is about provenance: where the row came from and whether it has been reviewed, not whether it matches the specific item in your hand.

So a verified row for "cooked chicken thigh" is a well-sourced figure for cooked chicken thigh. It is not a claim about your chicken thigh, which came from a different bird, was cooked differently, and lost a different amount of water. The catalog's job is to give you a good population figure. The variance between individual foods is real and no database removes it.

This is worth internalising because it bounds how much accuracy the database can buy you. Moving from a model's recollection to a verified catalog row is a genuine improvement and it is the reason this server exists. Moving from a verified row to reality still has your portion estimate and biological variation in the way. The full treatment of where calorie figures come from and how they are derived is in food API calorie data accuracy.

What an agent session actually spends

A worked example, in calls rather than tokens, since token counts depend on the client.

Logging a four-item meal: four searches, four portion scales. Eight calls. Each search returns up to 25 compact rows; each portion call returns one small object.

Totalling a recipe: one search per ingredient, then a single calculate_recipe with up to 40 ingredient pairs and a serving count. For a ten-ingredient recipe that is eleven calls, and the recipe call replaces ten portion calls plus the summation.

Against 20 calls a minute and 10,000 successful calls a month, neither is a problem. What is a problem is an agent that re-searches the same food every time it is mentioned. The fix is not a bigger plan; it is caching resolved rows in a file, as in meal planning with Claude. A food's per-100 g values do not change between Tuesday and Thursday.

One more spending pattern worth naming: the agent that searches, gets an ambiguous result, searches again with a different phrasing, and repeats. That loop is usually a symptom of an underspecified request rather than a bad catalog. Name the state of the food in the first place, cooked or skinless or dry, and the second search does not happen.

This is not the REST API, and that is the point

The MCP plan calls MCP tools. REST search, foods, calc and vision return 403 on this plan; account and billing routes still work.

People assume that is upsell mechanics. It is mostly about what each interface is for. A REST endpoint is designed for code you wrote, which knows exactly what it wants and handles pagination, retries and caching itself. An MCP tool is designed for a model choosing between options at runtime, which needs short descriptions, small responses and forgiving arguments. The same data, genuinely different surfaces, and optimising one for the other makes both worse. That argument is worked through in MCP vs REST.

If you are writing software, you want a REST plan. If you are talking to an assistant, you want this one. If you are building a product that resells the data, you want a REST plan and a commercial licence, because the MCP plan is personal use and a commercial usage header on an MCP key is rejected.

Connect it

Claude Code:

claude mcp add --transport http calorie-api \
  https://calorieapiadmin.com/mcp \
  --header "X-API-Key: YOUR_KEY"

Cursor:

{
  "mcpServers": {
    "calorie-api": {
      "url": "https://calorieapiadmin.com/mcp",
      "headers": { "X-API-Key": "YOUR_KEY" }
    }
  }
}

Browser sign-in for Claude apps is off on this server until OAuth is enabled, so those are the two clients. The MCP page has the current position, the docs have the reference, and MCP pricing has the plan.

What we have not measured

We have not published a token-cost measurement for tool responses before and after compaction, which is the number that would justify the design choices above quantitatively. We have not measured catalog hit rate against a fixed basket, by country or otherwise. We have not measured how often an agent picks the right row from a 25-result set.

Those three measurements are the ones that matter for this post and none of them exists yet. What is stated above is the response shape, the caps, and the call arithmetic, all of which you can verify yourself against the server in a few minutes.

Frequently Asked Questions

Why does search return only macros instead of every nutrient?

Because every tool reply costs context an agent also needs for your conversation and your log. Four macros cover the large majority of questions; the full nutrient panel is a separate get_food_nutrition call on a food id you already have.

Why is search capped at 25 results?

A hundred near-identical rows is a harder disambiguation problem than five, and every extra row crowds the context window. The paid cap is 25 foods per search and the trial cap is 10. A small, well-ordered set plus a clarifying question is the better behaviour.

Where does a food_id come from?

From search_foods and nowhere else. Portion, recipe and nutrient lookups all consume one. A model that invents an id gets an error, which is correct: the alternative would be scaling a fabricated row. Telling the assistant to search first is what fixes it in practice.

Does four million foods mean everything is in there?

No. Coverage is strong for whole and generic foods, large but never complete for branded products, and thinnest on regional own-brand stock. Whether your own shopping basket resolves varies by country more than by anything else.

Can I call the REST endpoints on this plan?

No. REST search, foods, calc and vision return 403 on the MCP plan, while account and billing routes work normally. The two surfaces are optimised for different callers: code that knows what it wants, versus a model choosing at runtime.

← Back to all articles

Start building with the Calorie API

Get a free API key and access 4M+ foods with search, barcode lookup, and full macro data.