Skip to content

Designing MCP Tools an LLM Actually Calls Correctly

Published July 16, 2026

A tool the model does not call is a REST endpoint with extra steps. The model sees a name, a description, and an argument list. It does not see your route table. If the name sounds like prose, or the argument is barcode in the docs and upc in the schema, the call is wrong in a way your server then has to forgive. Forgiving it hides the bug. Rejecting it, with a description that says the right name, fixes the next call.

This is the tool surface on the Calorie API MCP server, and the checks we can run without a model. A scored accuracy run (right tool, right arguments, fewest calls) has not been published. The harness exists. The percentage does not. Do not treat anything below as a measured accuracy number.

One action, one name

Eight tools, each a verb the model already uses:

ToolRequired ideaArgument names
search_foodsFind a foodquery, optional limit, verified_only
get_food_nutritionOne food's nutrientsfood_id
suggest_foodsPrefix completionsquery
lookup_barcodePackaged itemupc
calculate_portionScale to gramsfood_id, grams
calculate_recipeSum a recipeingredients, servings
calculate_macro_targetsA day's targetsage, gender, weight_kg, height_cm, activity, goal
analyze_food_photoA photoimage_base64, content_type

lookup_barcode does not take barcode. The description says UPC or EAN digits. A client, or a model, that sends barcode fails the schema. That is better than a silent alias, because the next prompt still says barcode and you never notice. The connect docs and the MCP reference use upc.

analyze_food_photo does not take a URL. The bytes are image_base64, without a data: prefix. JPEG, PNG, or WebP. The decoded cap is 2 MB, and the length is checked before decode so a huge string is rejected without allocating the image. There is no URL fetch. The description tells the model that clients do not attach images on their own. If you leave that sentence out, the model says "I cannot see the photo" or it invents a URL argument you do not accept.

Descriptions are the prompt

Write the description as the condition for calling the tool, not as marketing. search_foods returns per-100 g values and a food_id. calculate_portion scales that id. A model that skips the search and invents an id will get a not-found from the portion tool. The description of the portion tool should say the id comes from search. The description of search should say the numbers are per 100 g, not per portion, so the model does not treat them as the finished log entry.

suggest_foods is autocomplete, not a meal plan. If the description is "find foods," the model uses it instead of search and then has no id to scale. Narrow descriptions reduce the wrong call. They do not remove it. You still want a harness.

A harness that does not call a model

app/mcp/accuracy.py holds natural-language prompts paired with a tool name and the argument names that must be present. score_call returns true only when the tool matches and those arguments are non-empty. missing_tools returns any known tool that no case covers. The case list includes portion questions, barcodes, search, nutrition detail, suggestions, recipes, macro targets, and photos. Barcode cases expect upc. Macro cases expect the body measurements, not a single calorie number. Photo cases expect image_base64.

That function can grade a log of calls you already have. It cannot tell you what Claude will do on a Tuesday. Running it against a live model, then editing names until the score moves, is the measurement this article is supposed to contain. It has not been run for publication. Until it has, the design claims above are reviews of the schema, not a leaderboard.

If you build the same check, keep the expected arguments as names, not as example food ids. An example id in a test becomes a number the model memorizes, and then a blog post repeats it. We do not put catalog ids in the marketing transcript for that reason.

Errors the model can use

Authentication, plan, limit, and unavailable are different. A limit error includes the cap. A plan error says this subscription cannot do the thing, and where to change that. An unavailable error says to retry and does not include a stack trace. A photo that is too big says 2 MB and raw base64, not "bad request."

OAuth callers without nutrition:vision get a plan error that names the scope, so a reconnect can add it. API-key callers skip that check and use the plan's photo entitlement instead. Both paths are described in Authenticating a Remote MCP Server.

Do not hide the rate limit inside a friendly empty list

An empty search is a real empty search. A rate-limit failure is not an empty search. If you return {"results": []} when the account is over its minute, the model searches again with a shorter string and you have taught it the loop that rate limiting exists to stop. Return the limit error.

The descriptions the server actually sends

These are the strings on the tools, not a rewrite for this article.

search_foods tells the model to use it first whenever the user names a food, that the numbers are per 100 g, and that each result has a food_id which calculate_portion and get_food_nutrition accept. verified_only defaults to true. The field description says to keep it true unless the user asks for broader coverage. limit says 1-25. The plan may clamp it lower. The model asking for 25 on a trial account does not get 25.

calculate_portion says to use it after search, when the user gives an amount. food_id has to be a positive integer. grams has to be greater than 0 and at most 10000. A zero or a negative is a limit error, not a silent zero-calorie row.

lookup_barcode says to use it when the user provides a barcode number. The argument is upc. Spaces are stripped. The value has to be 8 to 14 digits or the tool returns upc must be 8 to 14 digits. An argument named barcode does not match.

calculate_macro_targets says it estimates a day from age, sex, weight, height, activity, and goal, and that it does not look up foods. Activity is one of sedentary, lightly_active, moderately_active, very_active, extra_active. Goal is lose, maintain, or gain. A prompt like "build me a 2200 calorie day" still needs those inputs. The tool will not invent a body.

analyze_food_photo says the photo is passed as image_base64 and that clients do not attach images automatically. The content type is image/jpeg, image/png, or image/webp. A data: prefix is rejected. The size check happens on the base64 length before decode.

What the harness grades

Three of the cases, copied from the harness, show the shape and nothing else:

  • "how much protein in 200 g of chicken thigh" expects calculate_portion with food_id and grams. It does not expect the model to know a food id. The grade fails if those names are missing. It does not check that the id is a real chicken.
  • "what's in this barcode 012345678905" expects lookup_barcode with upc.
  • "build me a 2200 calorie day" expects calculate_macro_targets with age, gender, weight, height, activity, and goal. A single calorie number is not enough to pass.

score_call is true only when the tool name matches and every expected argument is present and non-empty. missing_tools lists any known tool that no case mentions. That is the whole metric. There is no accuracy percentage beside it, because the cases have not been run through a model for this article.

get_food_nutrition says it returns the nutrient panel for one food, per 100 g, and that the id comes from search or from a barcode lookup. A missing id is No food found for that food_id. That is a limit-style error the model can recover from by searching again. It is not an authentication error.

suggest_foods says it is typeahead for a partial name and that it returns food_id and name only. It is not a search result with macros. If the description said "find foods," the model would stop there and never scale a portion. calculate_recipe says each ingredient needs a food_id and a gram weight. Servings have to be from 1 to 100. Ingredients have to be from 1 to 40. An empty list is an error, not a zero-calorie recipe.

analyze_food_photo says to use it when the user supplies a photo rather than naming a food, and it repeats that the bytes have to be passed as image_base64 because clients do not attach images automatically. Repeating that in the description is the whole fix. A blog post the model does not see will not change the call.

Frequently Asked Questions

Why is the barcode argument named upc?

That is the schema. UPC and EAN digits both go in upc. An argument named barcode does not match the tool.

Can the model pass an image URL?

No. The tool takes image_base64, raw base64, up to 2 MB decoded. URL fetch is implemented and left disabled.

What accuracy did the harness measure?

None that we are publishing. The harness checks tool name and required argument names. It has not been scored against a live model.

Should a rate limit look like an empty search?

No. Return the limit error, including the cap and the wait, so the model backs off instead of searching again.

← Back to all articles

Start building with the Calorie API

Get a free API key and access 4M+ foods with search, barcode lookup, and full macro data.