All datasets
246 cleaned, analysis-ready datasets and reference tables - every one QC-verified and AI-training-safe.
Showing 13 of 246 datasets
European Health System
The health-system profile of every European country in one clean panel: life expectancy at birth, hospital beds and practising physicians per 100 000 inhabitants, and total health expenditure as a share of GDP. Cleaned from Eurostat health statistics into one tidy annual per-country table - the capacity-and-outcomes view health-policy, insurance and life-science teams benchmark against.
US Drug Adverse
How abnormal each drug's adverse-event reporting is now: the z-score of monthly FAERS counts against the drug's own trailing 24-month baseline. |z| ≥ 3 = a 3σ safety-signal surge. Pre-computed from the FAERS monthly-counts product.
US Drug Adverse
Monthly counts of FDA adverse-event reports (FAERS) for ~40 major drugs by generic name - statins, GLP-1s, anticoagulants, biologics, oncology and more. Built from openFDA's COUNT endpoint, so it carries zero individual reports and zero patient data: a clean safety-signal time series per drug.
UK Food Hygiene Ratings (FSA FHRS)
The official hygiene rating of every food business in Great Britain & Northern Ireland - the regulator's version of the 'restaurant' data the review-site scrapers sell, but authoritative and open. The UK Food Standards Agency publishes ~520k establishments across ~360 local authorities as one XML file each; we assemble them into one clean table with the rating, inspection date, the three component scores (hygiene, structural, management) and a geocode. Covers both schemes - FHRS 0-5 (England/Wales/NI) and FHIS Pass/Improvement (Scotland). Open Government Licence v3.0 (fully commercial).
US County Health Prevalence (CDC PLACES)
Model-based health estimates for EVERY U.S. county in one tidy table - obesity, diabetes, smoking, high blood pressure, mental & physical health, cancer screening, insurance access and ~30 more. CDC's PLACES program publishes these as a long-format release; we pull the whole county file and ship one fat panel keyed by (county FIPS × measure × value-type), with readable measure names (not cryptic codes), both crude and age-adjusted prevalence, 95% confidence limits, the county population and a geocode. Drops straight next to any county-keyed dataset (proptech, insurance, policy) on the 5-digit FIPS. Public domain.
US Drug Directory (FDA NDC)
Every drug product marketed in the US in one clean reference table - the FDA National Drug Code directory, flattened. From openFDA's bulk NDC download (~137k products): one fat row per NDC with brand and generic names, active ingredients & strengths, dosage form, route, manufacturer/labeler, pharmacologic class, product type (prescription / OTC) and marketing category. The drug reference table health-tech, pharmacies, formularies and claims pipelines join on. Public domain / CC0.
US Food Bioactive Content
The polyphenol content of foods - the antioxidant and phytoestrogen compounds that appear in NO standard nutrition table - assembled from USDA's three special-interest bioactive databases into one clean long table. Flavonoids (Release 3.3: flavonols, flavones, flavanones, flavan-3-ols and anthocyanidins), Proanthocyanidins (Release 2.1: monomers through polymers) and Isoflavones (Release 2.1: daidzein, genistein, glycitein - the soy phytoestrogens), each mapped to its food, food group and compound class with the amount per 100 g. USDA ships these ONLY as legacy MS Access files on an ARS file server - which is exactly why no clean version exists; we read the Access databases directly, resolve the drifting column spellings and de-duplicate to one tidy (database × food × compound) table. Keyed by the legacy NDB number so it joins straight to the nutrition-composition dataset. Public domain (U.S. Government work).
US Food Glucosinolate Content
The glucosinolate content of cruciferous vegetables - the sulphur compounds in broccoli, cabbage, mustard and radish that chemoprevention and nutrition research studies, and that appear in no standard nutrition table. USDA and NIH's Office of Dietary Supplements published the first substantial public dataset as one Excel workbook of very wide tables (a Mean/SD/Min/Max block per compound, ~25 compounds, names on a merged header row); we parse every sheet into one tidy long table of (food × glucosinolate → mg per 100 g fresh weight), with the plant's scientific name. Public domain (U.S. Government work).
US Food Nutrition Composition
Every food in USDA's authoritative composition databases with the amount of every nutrient it contains, in one clean long table - the join nobody wants to do by hand. FoodData Central ships this as five separate CSVs (food, food_nutrient, nutrient, food_category and a per-type NDB crosswalk); we assemble and de-duplicate them into one tidy (food × nutrient) Parquet. Covers the three pure-USDA bundles - Foundation Foods (analytically measured reference foods), SR Legacy (the classic ~7.8k-food Standard Reference that underpins US nutrition analysis) and FNDDS (the 'as consumed' survey foods people actually report eating) - with energy, macronutrients, vitamins, minerals, amino acids and fatty acids, each with its unit and amount per 100 g. A data_type column preserves the measured-vs-survey distinction so you can filter either way. Keyed by FDC id with the legacy NDB number carried through, so it joins to recipe, dietary-intake and agricultural data. Public domain (U.S. Government work).
US Food Portion & Serving Weights
The gram weight of every household measure of a food - '1 cup', '1 tablespoon', '1 slice', '1 medium' → how many grams - the conversion every recipe, nutrition-tracking and dietary app needs and nobody ships cleanly. FoodData Central hides it in food_portion.csv behind a measure-unit lookup with the phrasing split across four columns; we join it (across Foundation, SR Legacy and FNDDS) into one tidy table of (food × household measure → gram weight). Keyed by FDC id so it drops straight next to the nutrition-composition dataset - turn any per-100g nutrient into per-serving. Public domain (U.S. Government work).
US Food Purine Content
How much purine is in each food and drink - the number people managing gout or high uric acid (and the diet apps serving them) actually need, and which exists nowhere as a clean table. USDA and NIH's Office of Dietary Supplements published the analytically-measured Purine Database (Release 2.0, 2025) as a multi-sheet Excel workbook with merged two-row headers and food-group headers interleaved with the data; we parse it into one tidy long table of (food × purine → mg per 100 g). Covers North-American and internationally-sourced foods plus alcoholic beverages, for the four measured purine bases - adenine, guanine, hypoxanthine, xanthine - and their total, each with mean, SEM, min and max. Public domain (U.S. Government work).
US Hospital Prices
What US hospitals actually charge - gross, discounted-cash AND payer-negotiated rates - for each billing code, in one tidy cross-hospital table. Since CMS's July-2024 mandate every hospital must publish a standardized machine-readable file (MRF) of its standard charges, but the files are a nightmare to use: 68 MB to 480 MB+ each, three CMS template variants (CSV 'tall', CSV 'wide', JSON), per-hospital column drift, prices as strings, and negotiated rates buried in nested payer arrays. We stream-parse a curated set of large hospitals that publish the CMS-conformant CSV-tall or JSON template and normalise everything to one schema: hospital, state, payer, plan, code type & code, item description, care setting, and the gross, discounted-cash, negotiated-dollar, min and max charges. Cleaning IS the product. Factual price data published under a federal mandate - freely reusable; shipped as a cleaned/derived aggregate, not the official files.
US Dietary Supplement Label Ingredients
Every US dietary-supplement product and what's actually in it, in one clean long table - the 200k+ label catalog nobody sells cleanly. The NIH Office of Dietary Supplements' Label Database has no bulk download and a search API that caps every result window at ~10k rows, so assembling the whole thing is genuinely tedious; we enumerate it exhaustively (slicing by entry-year × product-type) and explode each label into one row per (product × ingredient) with brand, product name, product type, market status and entry year. Answer 'which supplements contain ashwagandha / melatonin / creatine?' or 'what does brand X put in its multivitamins?' with a single query. Keyed by DSLD id. Public domain (U.S. Government work, NIH ODS).