Skip to content

Can ChatGPT Count Calories? How to Check Its Numbers

Published September 29, 2026

ChatGPT can estimate calories and it cannot look them up. It produces a food number the same way it produces the rest of the sentence, from patterns in training data rather than from a food database. The estimates are often close and they are never checkable. Here is where they drift, how to test yours in five minutes, and what actually fixes it.

The short answer

Ask ChatGPT how many calories are in 180 g of cooked chicken breast and you will get a confident figure. It will usually be in the right region, because food composition tables are well represented in the text it learned from, and chicken breast is one of the most written-about foods on the internet.

What did not happen is a lookup. No database was queried. Nothing in the reply traces back to a row you can inspect. Ask the same question in a new chat and you can get a different number, which is the clearest evidence of what is going on underneath.

For one question that is fine. Over a logged week it compounds, and the errors do not cancel out, because they are not random. They lean in specific directions covered below.

Why the numbers usually look right

This is worth understanding, because it is exactly what makes the failure mode hard to spot.

Generic whole foods are the best case. Chicken breast, white rice, an egg, olive oil, a banana. These have stable, widely published values, so a recalled figure lands close. If your diet is mostly plain staples weighed on a scale, ChatGPT will do a passable job and you may never notice a problem.

The reply is also fluent and specific. It gives you 297 kcal, not "roughly 300". Precision in the wording reads as precision in the method, and the two are unrelated. A number with a decimal place is not a number that was measured.

Four places the estimate drifts

Portions. This is the largest error and it is not really ChatGPT's fault. Tell it "a bowl of pasta" and something has to convert that into grams. Whatever it picks is a population average, and your bowl is not the average. The single biggest improvement available to anyone counting calories is a kitchen scale, before any software decision.

Branded and own-brand products. A supermarket own-brand granola is not in any training set in a reliable way. You will still get a number. It will be a generic granola figure wearing the brand name you typed, and the difference between generic granola and the one in your cupboard can be substantial, because the fat and sugar vary a lot between products that look alike.

Cooked versus raw. Dry rice roughly triples in weight when cooked. Dry pasta roughly doubles. Oats absorb several times their own weight. If you say "100 g of rice" without saying which, you are asking an ambiguous question, and you will get an answer to the version the model picked, silently. This error is a factor of three, not a rounding difference.

Compounding across a day. Each meal entry carries its own error, and a day is fifteen to twenty entries. They do not cancel, because under-reporting is directional: the things people forget to mention are cooking oil, dressings, milk in coffee, and what they ate standing at the counter. A day that is quietly 300 kcal light looks like a perfectly normal day in the chat.

Test yours in five minutes

Do not take anyone's word for how accurate it is, including ours. Five checks, and you can run all of them now.

1. Ask the same question twice, in separate chats. Same food, same weight. A lookup returns the same row. A recollection wanders. If the two answers differ by more than rounding, you have your answer about the mechanism.

2. Ask where the number came from. Which database did that come from, and what is the per 100 g row? An honest reply is that it came from general knowledge. If you get a confident citation to a specific database, treat that with suspicion rather than relief, since the lookup did not happen.

3. Run the 4/4/9 check. Protein and carbohydrate supply roughly 4 kcal per gram, fat roughly 9. Multiply out the macros it gave you and compare to the calorie figure it gave you:

(protein_g x 4) + (carbs_g x 4) + (fat_g x 9)  vs  kcal

Within about 10% is normal, since fibre, sugar alcohols and label rounding all account for small gaps. A larger gap means the macro figures and the calorie figure were produced independently rather than derived from one another.

4. Check it against something in your hand. Take a packet with a nutrition label, do not show it the label, and ask for the per 100 g values. This is the most informative test available to you, because you hold the ground truth. Branded products are where the gap is widest.

5. Ask for the gram weight it assumed. For anything you described in words rather than weights: what gram weight did you use for that? If the assumption is not stated, the estimate cannot be corrected, which makes it useless for anything beyond a rough sense of scale.

Why there is no accuracy percentage in this post

You will find plenty of pages offering one. We are not going to add to them.

We have not run a controlled test of ChatGPT's calorie estimates against weighed and labelled food. Until we do, we have no figure to report, and we are not going to borrow one from another company's blog and present it as though it applied here.

The published figures that circulate also disagree with each other, for a reason worth knowing: they measure different things. Identifying the foods on a plate and estimating how much of each is there are separate problems with very different error rates, and a single headline percentage usually does not say which one it measured, on what foods, against what ground truth. When you see one confident number for "AI calorie accuracy", that is the question to ask.

What we can tell you is mechanical and checkable in five minutes, which is what the section above is for.

Using it to plan a calorie deficit

A related use, and the failure mode is different.

ChatGPT is genuinely decent at the arithmetic of a deficit. Give it your height, weight, age, activity and goal, and it will apply a standard resting-expenditure formula with an activity multiplier and subtract a sensible amount. That is a formula, it is well represented in training data, and the output is reasonable.

Two cautions. First, the formula is an estimate of a population, not a measurement of you. Two people with identical inputs genuinely differ by several hundred calories a day. Hold whatever number it gives you for two weeks, track a weekly average weight rather than daily numbers, and adjust from what the scale actually did.

Second, and more important: a deficit is only as good as the intake tracking underneath it. A target of 2,150 kcal is meaningless if the logged days are quietly running 300 kcal light. The plan will look correct and the scale will not move, and the usual conclusion people reach is that their metabolism is broken. It is much more often the logging. There is a diagnostic for exactly this, weighing everything strictly for four days, described in the fat-loss tracking post.

Macros specifically

The same split applies. Ask ChatGPT to calculate macro targets from your bodyweight and goal and it does formula work, which is fine. Ask it for the macros in a specific food and it is recalling, which is the part to verify.

One tell worth watching for: ask for a day's total macros and then add the individual entries yourself. Totals produced in prose drift from their own line items more often than you would expect, because the total is generated as text rather than computed. The 4/4/9 check catches the worst of it.

What actually fixes it

Two honest options, and which one applies depends on what you are doing.

Give an assistant a real database. This is the structural fix. Instead of asking a model to recall a food composition table, you attach a tool that queries one, so the assistant searches the catalog, gets a row with an identifier and macros per 100 g, and scales that row to your gram weight. The numbers in the reply then came from a database and you can ask which row.

That is what the Model Context Protocol is for, and it is what this nutrition MCP server does. An important limitation, stated plainly: this server authenticates with an API key header, and we have verified it with Claude Code and Cursor. We have not verified it inside ChatGPT, and browser sign-in, which is what most chat apps expect, is not enabled here. So today this is a Claude Code or Cursor setup, not a ChatGPT one. The tool list and arguments are in the docs, the plan and its personal-use limits are on MCP pricing, and a full day logged that way is in calorie tracking in Claude.

Build against an API. If you are writing software rather than chatting, a chat assistant is the wrong component entirely. You want HTTP endpoints with predictable responses, which is a REST plan.

If you are staying with ChatGPT anyway

Entirely reasonable. Most people are not going to move their food log into a terminal, and the honest position is that a rough log you actually keep beats a precise one you abandon. Four things that reduce the damage:

  • Weigh things. Grams in, not "a bowl". This fixes the largest error without changing any software.
  • Use barcodes and labels for packaged food. Read the per 100 g values off the packet and ask it to scale them. You have supplied the truth; it is only doing arithmetic.
  • Say raw or cooked, every time.
  • Ask it to state assumptions. Any weight it guessed should be visible in the reply.

Do those four and ChatGPT becomes a reasonable calculator operating on numbers you provided, which is a genuinely different and much safer thing than asking it to remember what food contains.

What we have not measured

No controlled comparison of ChatGPT's estimates against weighed and labelled food. No error distribution by food type. No catalog hit-rate figure for our own data. No measurement of how often the 4/4/9 check catches a bad reply.

Everything above is either mechanism, which you can verify yourself in five minutes with the tests in this post, or arithmetic. When we run a benchmark it will be published with its method and its raw data so you can disagree with it.

Frequently Asked Questions

Can ChatGPT count calories accurately?

It estimates rather than looks up. Figures for generic whole foods land close because those values are well represented in its training data. Branded products, portion sizes described in words, and cooked versus raw weights are where it drifts, and nothing in the reply traces to a row you can inspect.

How do I tell whether ChatGPT looked a number up?

Ask the same food and weight in two separate chats. A lookup returns the same row every time; a recollection wanders. Then ask which database it used and what the per 100 g row was. An honest answer is that it came from general knowledge.

What is the 4/4/9 check?

Protein and carbohydrate supply roughly 4 kcal per gram and fat roughly 9. Multiply the macros it gave you and compare with the calorie figure it gave you. Within about 10% is normal. A larger gap means the macros and the calories were produced independently rather than derived from each other.

How accurate is it, as a percentage?

We have not run a controlled test against weighed and labelled food, so we do not publish a figure. Published numbers elsewhere disagree because they measure different things: identifying foods and estimating portion sizes are separate problems with very different error rates.

Can I connect a nutrition database to ChatGPT?

Not to this server today. It authenticates with an API key header and we have verified it with Claude Code and Cursor. Browser sign-in, which most chat apps expect, is not enabled here. If you are building software rather than chatting, a REST plan is the right route.

← Back to all articles

Start building with the Calorie API

Get a free API key and access 4M+ foods with search, barcode lookup, and full macro data.