Skip to content

Food Photo to Calories Inside Claude

Published August 13, 2026

Photo calorie estimation is the most demonstrated feature in consumer nutrition software and the least well explained. Every app shows the same thirty-second video: plate, shutter, numbers. Nobody shows you what the numbers are made of.

Over MCP the mechanism is visible, which is useful, and it is also more awkward than an app, which is honest. This post covers both.

The image has to be inside the tool call

Start here, because it is the single thing people get wrong.

Dragging a photo into Claude Code or Cursor does not send it to this server. The client shows the image to the model. MCP clients do not forward an attachment to a tool, so the server never sees a file at all. The assistant will look at your picture and may well describe it, but that description came from the model, not from analyze_food_photo, and no catalog was involved.

What the tool takes is an argument:

{
  "image_base64": "<raw base64, no data: prefix>",
  "content_type": "image/jpeg"
}

Raw base64. No data:image/jpeg;base64, prefix. image/jpeg, image/png or image/webp. Decoded size up to 2 MB. A URL is not accepted, so you cannot hand it a link to an image.

So somebody has to read the file and encode it. In Claude Code that somebody is the assistant, because it has file access:

You: Encode ./dinner.jpg and run it through the photo tool.

That works, and it is the realistic flow: photo on the phone, into a synced folder, point the assistant at the path. Two steps more than an app. If a camera-to-log flow in one tap is what you want, an app is the right tool and there is no shame in using one.

The 2 MB limit is on the decoded bytes, and base64 inflates a file by about a third on the wire. A modern phone photo is often 3 to 5 MB, so plenty of pictures need resizing first. Long side down to about 1,200 pixels at JPEG quality 80 is plenty, because the model is identifying foods and estimating volumes rather than reading fine print. If you exceed the limit the error says the image is too large. It does not return a stack trace and it does not silently truncate.

What happens after the bytes arrive

The tool detects the foods in the image and returns an estimate. Two distinct things are happening, and they fail differently.

Identification. What is on the plate. This is the part that works well. Rice, chicken, broccoli, a fried egg: these are visually distinctive and a modern vision model handles them. Where it degrades is predictable: mixed dishes where components are not visible, anything under a sauce, foods that look alike at different fat contents, and cuisines that are underrepresented in training data. A curry is one brown mass; whether it has 30 g or 80 g of fat in it is not visible.

Portion estimation. How much. This is the hard part and it is the part that dominates the error. A plate photographed from above has no depth. The same visible surface of rice can be a shallow spread or a deep mound, and the difference is most of the calories. There is no reference object unless you put one in the frame.

Which means: the identification being good is not the same as the estimate being good. When two calorie apps disagree by a third on the same photo, it is almost never because one of them failed to see the chicken.

Getting a better estimate

The interesting thing about doing this in a chat rather than an app is that you can argue with it. Four things that help:

Shoot at an angle, not from directly above. Thirty to forty-five degrees gives the model depth information. A top-down photo throws away the dimension you most need.

Put something known in the frame. A fork, a standard plate, your hand. Scale is the missing variable and any familiar object supplies some of it.

Tell it what it cannot see. This is the big one and it is free:

You: That is cooked in about 20 g of butter and the sauce is cream-based. The rice is a full cup cooked.

Hidden fat is the most underestimated thing in any photo estimate, because it is invisible by design. Restaurant food in particular carries fat you would not add at home. Saying so shifts the estimate toward the higher-fat reading, which is usually the correct one.

Then correct it against the catalog. The photo gives you a starting list of foods and rough weights. Convert that into real entries:

You: Take those four foods, search each in the catalog, and scale them to the weights you estimated. Show me the portion calls.

Now the food identification came from the image and the macros came from a lookup, which is the division of labour that makes sense. The photo is a data-entry shortcut, not an oracle. That two-step pattern is the same one used throughout calorie tracking in Claude.

Reading the reply properly

The output is a list of detected foods with estimated quantities and a calorie and macro estimate. There are three things to look at before you accept any of it.

Did it get the food list right? Scan the names. A missing component is the most consequential error, and the most common missing component is a liquid or a fat: the dressing, the oil, the sauce. If a food is absent from the list, nothing downstream includes it.

What weights did it assume? This is the number to interrogate. If it says 150 g of rice and you know that bowl holds closer to 300 g, the estimate is out by roughly a hundred and fifty calories on that line alone. Ask:

You: What gram weight did you assume for each item?

If the assumptions are not stated, ask for them. An estimate whose assumptions are hidden cannot be corrected, which makes it useless for anything except a vague sense of scale.

Is the total plausible? The 4/4/9 cross-check works here as well as on a label: protein and carbohydrate at roughly 4 kcal per gram, fat at roughly 9. If the macros do not multiply out to something near the stated calories, one of the numbers came from somewhere strange.

Where the image goes

Worth knowing, because it is a photo of your food and sometimes your kitchen.

The base64 goes to the server in the tool call, and the server passes the bytes to a vision model to get the estimate back. So the image leaves your machine. It is not stored as part of a food log, because there is no food log and the server keeps no history for you, but it is processed off-device, and that is a different privacy posture from a server that ships a local dataset and never makes an outbound call.

If that matters for a particular photo, the answer is not to send it. Describe the plate in words instead and reconstruct it from catalog lookups, which for a meal you can name is usually the better method anyway.

Where a photo is the wrong tool

Three cases where something else is strictly better, and reaching for the camera is a mistake:

Packaged food. There is a barcode. Use it. A photo of a packet is a worse input than the twelve digits printed on it, and barcode lookup gives you the manufacturer's own numbers.

Food you cooked. You know the ingredients and you could have weighed them. calculate_recipe on the actual ingredients beats guessing at the finished dish from a picture, every time, by a wide margin.

Anything you eat weekly. Resolve it once, save the row in your profile file, log it from then on with no call at all.

That leaves the case a photo is genuinely for: someone else's cooking, eaten out, that you did not weigh and cannot label. Which is a real and common case. It is just narrower than the marketing suggests.

The limits

Paid accounts get 150 photo estimates a month and 20 a day. The trial includes 10 in total, alongside its 200-call overall budget.

The daily cap is the one worth noting. Twenty a day is generous for logging your own meals and deliberately too low to process a folder of images in a loop, which is what the limits exist to prevent. The reasoning is in rate limiting an MCP server, and the numbers are on MCP pricing and in the nutrition://limits resource on the server.

Connect it

Claude Code:

claude mcp add --transport http calorie-api \
  https://calorieapiadmin.com/mcp \
  --header "X-API-Key: YOUR_KEY"

Cursor:

{
  "mcpServers": {
    "calorie-api": {
      "url": "https://calorieapiadmin.com/mcp",
      "headers": { "X-API-Key": "YOUR_KEY" }
    }
  }
}

Those two clients, with the header. Browser sign-in for Claude.ai, Claude Desktop and Claude mobile is off on this server until OAuth is enabled. Between the two clients, Claude Code has the advantage here specifically because the photo flow needs file access. The full setup is in the connect guide, the argument reference is in the docs, and the plan itself is on the MCP page.

What we have not measured, and why there is no number here

You will have noticed this post contains no accuracy figure. That is deliberate and it is worth explaining, because every competing page leads with one.

We have not run a photo benchmark. No set of meals photographed and then weighed, no error distribution, no per-cuisine breakdown. Until we do, we have nothing to report, and we are not going to borrow a percentage from another company's blog and present it as though it applied to this tool.

There is published work on photo calorie estimation, and the figures in it vary enormously depending on method, food type and what is being measured. Identification accuracy and portion accuracy are wildly different numbers, and pages that quote one headline percentage usually do not say which. If you see a single confident number for "AI calorie accuracy", that is a signal to ask what was measured, on what foods, against what ground truth.

What we can tell you from the mechanism, and stand behind:

  • Identification is the easier half; portion estimation is where the error lives.
  • A top-down photo of a mixed dish is the worst case, and hidden fat is the most commonly missed thing.
  • A barcode or a weighed recipe beats a photo whenever either is available.
  • Correcting a photo estimate against catalog rows is better than accepting the estimate whole.

When we run a benchmark it will come with the method and the raw data, so you can disagree with it.

Frequently Asked Questions

Why does dropping a photo into the chat not work?

MCP clients do not forward an attachment to a server. The client shows the image to the model, the server never receives a file, and any description you get came from the model rather than from the photo tool. The image has to be passed as an image_base64 argument.

What format and size does the photo tool accept?

Raw base64 with no data: prefix, as image/jpeg, image/png or image/webp, up to 2 MB decoded. A URL is not accepted. Resizing the long side to about 1200 pixels at quality 80 is enough, since the model is estimating volumes rather than reading small print.

What is the accuracy of the photo estimate?

We have not measured it and we do not publish a figure. Identification of visible foods is the easier half; portion estimation dominates the error because a photo has no depth. Published third-party numbers vary enormously by method and food type, so we do not quote them as ours.

How do I get a better estimate from a photo?

Shoot at an angle rather than top-down, include a familiar object for scale, and tell the assistant what it cannot see, such as the cooking fat or a cream-based sauce. Then have the identified foods searched in the catalog and scaled, so the macros come from a lookup.

When should I not use a photo?

When there is a barcode, when you cooked the food and know the ingredients, or when it is something you eat every week and could save once. A photo is for someone else's cooking that you did not weigh and cannot label.

← Back to all articles

Start building with the Calorie API

Get a free API key and access 4M+ foods with search, barcode lookup, and full macro data.