Data you can actually train on.
Every ReadySet dataset is built from a named, redistributable official source under a documented open licence - not scraped web data. No copyright grey zone, no personal data, provenance recorded per file. The rare training data that's legally safe, ready for your pipeline, and priced per file.
Safe to train on - the part scraped data can't offer
Training on data of unknown origin is now a real legal exposure: copyright, terms-of-service and privacy claims all follow the data into your model. Scraped web datasets - the kind most "AI data" vendors sell - carry exactly that risk. ReadySet is the opposite by construction:
- Official source, named per dataset - government agencies, statistical offices, central banks, multilateral bodies (Eurostat, OECD, World Bank, the Fed, and more). Never a scrape of unknown provenance.
- Documented open licence - public domain / CC0 / CC-BY / open-government. Redistribution and AI training are permitted, and we ship the exact licence + attribution with every file.
- Aggregate only - no personal data. Our pipeline rejects any source carrying individual-level personal fields. Nothing to leak into your weights.
- Provenance you can hand to compliance - a machine-readable training-licence manifest per dataset (source, licence, permits-training, no-PII), supporting AI-Act training-data documentation.
Tabular & time-series AI - not another text scrape
Clean, gap-aware, one-schema numeric data across decades and hundreds of countries. The shape that quantitative and structured-data models actually want.
Forecasting & tabular models
Long, aligned time series with honest gaps and a documented grain - ready for forecasting, tabular foundation models and feature stores.
Quant & agentic AI
Macro, energy, finance and ESG panels an agent can query and reason over - one join key, one licence, no parsing.
RAG over structured data
Documented schemas + semantic descriptions per column, so a retrieval layer knows exactly what each field means.
Load it in one line; document it for free
Every dataset ships as columnar Apache Parquet with a data dictionary and a load snippet, plus two machine-readable metadata files served with each product:
- Croissant metadata (ML Commons) - the standard Hugging Face, Kaggle and Google read, so the dataset loads with your existing tooling and is auto-discoverable as an ML dataset.
- AI training-licence manifest - a compact JSON stating the source, SPDX licence + URL, that training and redistribution are permitted, and that no personal data is present. Drop it straight into your data-governance record.
- Machine-readable data dictionary (Markdown) - every column with its type and a plain-language meaning. Feed it straight into an agent or RAG system prompt so the model knows exactly what each field is, without parsing the Parquet.
Don't take our word for it - inspect the metadata before you buy (all three are public, no purchase needed). For example on
european-day-ahead-prices:
Croissant JSON ·
training-licence manifest ·
data dictionary.
Verify it loads in Hugging Face / Kaggle before spending a cent.
Comprehensive corpora, one file each
Our flagship panels join dozens of official indicators across every country and decade into a single, coherent, licence-clean training set.
Global Country Profile
Every country on Earth, every year, ~75 indicators, one file. The worldwide flagship: World Bank development indicators (GDP, growth, trade, debt, health, education, energy, environment, demographics) and climate indicators, joined with OECD harmonised inflation, house prices and unemployment - aligned on ISO-3 country × year since 2000, ~200 countries. One coherent, licence-clean panel that replaces a dozen separate downloads - the single cross-country sheet a global-macro analyst or a forecasting model actually wants. 100% redistributable official data (World Bank CC-BY + OECD CC-BY), safe to redistribute and to train on.
European Country Profile
Everything we know about every European country, in one file. This flagship panel joins ~45 annual variables per country - inflation and the full HICP breakdown, GDP, house prices, unemployment and vacancies, government revenue/deficit/debt, births/deaths and vital rates, household & industrial energy prices, energy import-dependency, construction output, tourism nights, air passengers and greenhouse-gas emissions - aligned on country × year since 2000. One coherent, licence-clean panel worth far more than the sum of its parts: the single sheet a cross-country analyst (or a forecasting model) actually wants. Built entirely from redistributable official sources (Eurostat), safe to redistribute and to train on.
Global Financial Markets (OECD)
Rates, equities and currencies for the whole OECD in one clean Parquet - the most-requested cross-country macro-finance panel. Harmonised monthly short-term (3-month) and long-term (~10-year government bond) interest rates, the national share-price index and the real effective exchange rate - the US, Japan, the euro area, the UK and every member, one row per country × series × month since 2000. The comparable rates/FX/equity feed carry, curve and cross-asset models need; licence-clean and AI-training-safe.
European Energy
The decarbonisation dashboard for every European country in one file: the renewable share of electricity generation, grid carbon intensity, retail household electricity and gas prices, energy import-dependency and total greenhouse-gas emissions - aligned on country × year. Joins our ENTSO-E power feed (rolled up from bidding zones to country) with Eurostat energy prices, import-dependency and emissions into one energy-transition sheet worth more than the parts. Licence-clean official data (ENTSO-E + Eurostat), safe to redistribute and train on.
Train on data you can stand behind.
Pay per file from €9, or go All-Access for unlimited pulls across your training runs.