For AI & ML teams

Data you can actually train on.

Every ReadySet dataset is built from a named, redistributable official source under a documented open licence - not scraped web data. No copyright grey zone, no personal data, provenance recorded per file. The rare training data that's legally safe, ready for your pipeline, and priced per file.

246
datasets, training-ready
336M+
rows, clean & typed
100%
licence-clean, safe to train on
0
scraped web data · 0 PII
Why it matters

Safe to train on - the part scraped data can't offer

Training on data of unknown origin is now a real legal exposure: copyright, terms-of-service and privacy claims all follow the data into your model. Scraped web datasets - the kind most "AI data" vendors sell - carry exactly that risk. ReadySet is the opposite by construction:

  • Official source, named per dataset - government agencies, statistical offices, central banks, multilateral bodies (Eurostat, OECD, World Bank, the Fed, and more). Never a scrape of unknown provenance.
  • Documented open licence - public domain / CC0 / CC-BY / open-government. Redistribution and AI training are permitted, and we ship the exact licence + attribution with every file.
  • Aggregate only - no personal data. Our pipeline rejects any source carrying individual-level personal fields. Nothing to leak into your weights.
  • Provenance you can hand to compliance - a machine-readable training-licence manifest per dataset (source, licence, permits-training, no-PII), supporting AI-Act training-data documentation.
Built for

Tabular & time-series AI - not another text scrape

Clean, gap-aware, one-schema numeric data across decades and hundreds of countries. The shape that quantitative and structured-data models actually want.

Forecasting & tabular models

Long, aligned time series with honest gaps and a documented grain - ready for forecasting, tabular foundation models and feature stores.

Quant & agentic AI

Macro, energy, finance and ESG panels an agent can query and reason over - one join key, one licence, no parsing.

RAG over structured data

Documented schemas + semantic descriptions per column, so a retrieval layer knows exactly what each field means.

Machine-ready out of the box

Load it in one line; document it for free

Every dataset ships as columnar Apache Parquet with a data dictionary and a load snippet, plus two machine-readable metadata files served with each product:

  • Croissant metadata (ML Commons) - the standard Hugging Face, Kaggle and Google read, so the dataset loads with your existing tooling and is auto-discoverable as an ML dataset.
  • AI training-licence manifest - a compact JSON stating the source, SPDX licence + URL, that training and redistribution are permitted, and that no personal data is present. Drop it straight into your data-governance record.
  • Machine-readable data dictionary (Markdown) - every column with its type and a plain-language meaning. Feed it straight into an agent or RAG system prompt so the model knows exactly what each field is, without parsing the Parquet.

Don't take our word for it - inspect the metadata before you buy (all three are public, no purchase needed). For example on european-day-ahead-prices: Croissant JSON · training-licence manifest · data dictionary. Verify it loads in Hugging Face / Kaggle before spending a cent.

Start here

Comprehensive corpora, one file each

Our flagship panels join dozens of official indicators across every country and decade into a single, coherent, licence-clean training set.

Economy

Global Country Profile

Every country on Earth, every year, ~75 indicators, one file. The worldwide flagship: World Bank development indicators (GDP, growth, trade, debt, health, education, energy, environment, demographics) and climate indicators, joined with OECD harmonised inflation, house prices and unemployment - aligned on ISO-3 country × year since 2000, ~200 countries. One coherent, licence-clean panel that replaces a dozen separate downloads - the single cross-country sheet a global-macro analyst or a forecasting model actually wants. 100% redistributable official data (World Bank CC-BY + OECD CC-BY), safe to redistribute and to train on.

524K rows· 217 zones· 1990-2026 (annual)
29 View
Economy

European Country Profile

Everything we know about every European country, in one file. This flagship panel joins ~45 annual variables per country - inflation and the full HICP breakdown, GDP, house prices, unemployment and vacancies, government revenue/deficit/debt, births/deaths and vital rates, household & industrial energy prices, energy import-dependency, construction output, tourism nights, air passengers and greenhouse-gas emissions - aligned on country × year since 2000. One coherent, licence-clean panel worth far more than the sum of its parts: the single sheet a cross-country analyst (or a forecasting model) actually wants. Built entirely from redistributable official sources (Eurostat), safe to redistribute and to train on.

121K rows· 42 zones· 1990-2026 (annual)
29 View
Finance

Global Financial Markets (OECD)

Rates, equities and currencies for the whole OECD in one clean Parquet - the most-requested cross-country macro-finance panel. Harmonised monthly short-term (3-month) and long-term (~10-year government bond) interest rates, the national share-price index and the real effective exchange rate - the US, Japan, the euro area, the UK and every member, one row per country × series × month since 2000. The comparable rates/FX/equity feed carry, curve and cross-asset models need; licence-clean and AI-training-safe.

OECD· 55K rows· 47 zones· 2000-2026 (monthly)
19 View
Energy

European Energy

The decarbonisation dashboard for every European country in one file: the renewable share of electricity generation, grid carbon intensity, retail household electricity and gas prices, energy import-dependency and total greenhouse-gas emissions - aligned on country × year. Joins our ENTSO-E power feed (rolled up from bidding zones to country) with Eurostat energy prices, import-dependency and emissions into one energy-transition sheet worth more than the parts. Licence-clean official data (ENTSO-E + Eurostat), safe to redistribute and train on.

4K rows· 43 zones· 1990-2026 (annual)
19 View

Train on data you can stand behind.

Pay per file from €9, or go All-Access for unlimited pulls across your training runs.