Food data

Where our food data comes from

Yumizo’s food search and barcode lookup use our own copy of the OpenNutrition database. For barcodes it doesn’t have, we ask Open Food Facts. Both are shared under the Open Database License.

The two databases

How we clean OpenNutrition

Before the data reaches the app, we make these changes to each food, in this order:

  1. Names and brands. Trim spaces from both ends of every name. On grocery and restaurant foods, the text after the last “ by ” in the name (the word “by” in small letters, with a space on each side) becomes the brand, and the text before it stays the name, both trimmed; an empty brand counts as none. A grocery or restaurant food named only “by” and a brand is named after the brand. A food left with no name is dropped.
  2. Rounding. Take each food’s calories, protein, carbs, fat, fibre, sugar, saturated fat and sodium per 100 g, rounded to two decimal places (whole numbers stay whole). A figure the dataset doesn’t give stays empty.
  3. Hand corrections. Drop or correct the foods in our list of hand corrections (overrides.csv, below), matched by their OpenNutrition ID. They are foods whose figures agree with their own protein, carbs and fat but not with the food, so the steps after this one can’t catch them. A correction replaces only the figures it gives, rounded the same way, and the steps after this one still apply to it.
  4. Impossible values. Drop foods with no calorie figure, with more than 902 calories per 100 g (more than pure fat), or with more than 105 g of protein, carbs and fat together per 100 g, a missing figure counting as 0.
  5. Calories that don’t add up. Work out the calories a food’s nutrients explain: 4 for each gram of protein, 4 for each gram of carbs, 9 for each gram of fat and 7 for each gram of alcohol (the dataset’s alcohol figure), a missing figure counting as 0. When that comes to more than 0, and the published calories differ from it by more than half of it and by at least 20 calories per 100 g, the published figure is replaced by the worked-out one, rounded to one decimal place, and the food is marked as recalculated. Two kinds of food keep their published calories instead, because those nutrients can’t explain their energy. Sweeteners are checked first: a food with a sugar-alcohol figure above 0, or with a sweetener word in its name. Then alcoholic drinks with no alcohol figure (or a figure of 0): a food carrying the dataset’s “alcohol” label, or with an alcohol word in its name. The name here is the one left after the first step, and a word matches only whole, in any mix of capital and small letters.

    Sweetener words: stevia, sucralose, monk fruit (or monkfruit), erythritol, xylitol, allulose, aspartame, saccharin, acesulfame, sweetener, sugar substitute, sugar-free syrup (or sugar free syrup), splenda, truvia.

    Alcohol words: beer, wine, ale, lager, stout, porter, pilsner, ipa, cider, mead, sake, soju, vodka, whisky (or whiskey), bourbon, scotch, rum, gin, tequila, mezcal, brandy, cognac, liqueur, schnapps, vermouth, sherry, port, champagne, prosecco, cava, sangria, mimosa, margarita, mojito, martini, daiquiri, cocktail, spritz, hard seltzer, extract, bitters, shochu, baijiu, makgeolli, chardonnay, merlot, cabernet, pinot, sauvignon, riesling, zinfandel, malbec, rosé, rose wine, moscato, shiraz, syrah, grenache, tempranillo, chianti, rioja.

  6. Sugar and saturated fat. Where both figures are given, lower sugar to the carbs figure when it is higher, and saturated fat to the fat figure when it is higher.
  7. Sodium. Remove a sodium figure above 40,000 mg per 100 g, more than table salt, as a unit error.
  8. Dry grains. Add “, dry” to a food’s name when all of these hold, so a cup of cooked rice isn’t counted as dry rice: the food isn’t a restaurant dish; it has more than 300 calories and under 10 g of fat per 100 g after the steps above (a missing fat figure counting as 0); its name ends in a grain or pulse word; and its name has no word that already says how the food is eaten or that it isn’t the plain grain. A word matches only whole, in any mix of capital and small letters.

    Grain and pulse words: rice, quinoa, oats, lentil (or lentils), dal, daal, dahl, pasta, spaghetti, macaroni, penne, fusilli, couscous, bulgur, barley, millet, farro, buckwheat.

    Words that leave a name alone: cooked, steamed, boiled, dry, dried, uncooked, raw, puffed, popped, crisp, crispy, crisps, roasted, toasted, fried, baked, flour, cake, cakes, cracker, crackers, chips, snack, salad, soup, pudding, canned, pouch, ready.

  9. Servings. Keep a food’s default serving amount when it is a number more than 0 and up to 5,000 in a unit we can read, ignoring capitals and spaces around it: “g”, “grm” (a misspelling of grams) and “mg” (a unit error in the data: “28 mg” of chips is 28 g) are read as grams, and “ml” as millilitres, which marks the serving as a liquid. Any other amount is dropped, and the app falls back to 100 g. The serving’s label is its household measure, quantity then unit (“1 cup”), with the measure’s own name in front when it has one (“Small (12 fl oz)”), or that name alone; with neither, it is the amount kept (“28 g”). A label is kept even when the amount isn’t, is cut to 60 characters, and writes quantities to at most two decimal places without trailing zeros. A package size is kept when the serving amount was and the package is more than 0 in the same kind of unit, read the same way: “ml” for a liquid, otherwise “g” or “grm”. Amounts are rounded to two decimal places.
  10. Other names and barcodes. Keep a food’s other names that are text, for search. Keep its barcode when it is made only of digits, stored as a number, so leading zeros are not kept; any other barcode is dropped.
  11. What we keep. Each food’s OpenNutrition ID, name, brand, kind of food (everyday, prepared, restaurant or grocery), the eight figures above per 100 g, its default serving (amount, label and whether it is a liquid), its package size, its other names and barcode, and whether its calories were recalculated. Nothing else from the dataset is kept, and nothing else is changed. Alongside the foods we keep a search index of each food’s name, other names and brand, which ignores capitals and accents and weighs a match in the name most, then the other names, then the brand; and a record of the dataset’s name, version, web address and licence, the units, the number that stands for each kind of food, when and by which script our copy was built, a fingerprint of the dataset file and of the hand corrections, and how many foods each step read, kept, dropped or changed.

That is the whole of what we change. The list of hand corrections is a file you can download:

  • overrides.csv, the foods we drop or correct by hand, by their OpenNutrition ID

Our version of the database is shared under the same licence. Nothing on this site limits your rights under the ODbL in the data itself.

Numbers are estimates

Food data is provided as it is. A label is a manufacturer’s figure, a product’s recipe may have changed since it was recorded, and a database can be wrong. You can change any food or number in your diary. See section 7 of our Terms of Service.