Trust & data quality

Every number is sourced, cleaned and checked.

We sell data other people rely on, so provenance and quality aren't a footnote - they're the product. Here is exactly where each dataset comes from, how it's cleaned, how it's quality-checked, and how we handle licensing and privacy.

246
datasets, live
336M+
rows cleaned & typed
78
official sources
100%
pass QC clean
1 · Sourcing

Primary sources only - and only ones we may resell

Every dataset is built from a named upstream authority, not a scrape of unknown origin. The catalog draws from institutions like ENTSO-E, ECB, U.S. Treasury, FRED, U.S. EIA, SEC EDGAR, Eurostat, OECD and more - and each product page states its source explicitly.

The single most important decision we make is which sources we're allowed to build on at all. We only use data with a redistributable licence - CC-BY, public domain, open government and similar. Exchange feeds, non-commercial-only and unlicensed sources never enter the catalog, regardless of how good the numbers are. Where upstream requires attribution, it passes through to you, documented on the product page and in the download. See the Licence for the full terms.

2 · Cleaning & quality control

What "analysis-ready" actually means here

The same pipeline runs on every product, so quality is a property of the factory, not of how careful we felt that day.

One documented grain

Every dataset has a single, stated grain (e.g. one row per zone per hour), UTC timestamps, consistent units, and a column dictionary describing each field and its type.

Deduplicated & gap-aware

Duplicate rows are removed deterministically; real gaps are represented as gaps, not silently forward-filled. Revisions collapse to the finest available resolution.

A QC report on every file

Each product ships a machine-generated QC report - row counts, null rates, range checks, duplicate and gap proofs. A dataset only ships when it passes with zero hard failures.

Swept as a whole catalog

A single automated audit re-checks every product in the catalog on each build, so a regression in one dataset can't hide. Today: 100% pass fully clean.

See it before you buy

Every dataset has a free sample (first rows + full schema) plus its QC summary and coverage, downloadable without an account. Confirm fit, then buy.

Fast, portable formats

Delivered as Apache Parquet with a data dictionary and loader snippet - readable instantly in pandas, DuckDB or Polars. CSV samples for Excel.

3 · Freshness

Updated on a stated cadence - and dated

Datasets are refreshed on a regular schedule (most monthly; some track their upstream's own publication cadence). Every product page shows both its update cadence and the date the data currently runs through, so you're never guessing how current a file is. Buy once and re-download the latest build anytime from your account.

4 · Privacy & security

Aggregate data. No personal data. No tracking theatre.

Our products are aggregate statistics - prices, volumes, generation, positioning, reference tables. We do not sell personal data: the pipeline actively rejects any relational source carrying individual-level personal fields, so nothing with names, emails or officer/director records reaches the catalog.

For your account we store only what's needed to sell you a file and let you re-download it - your email, your purchases, and download logs used solely to keep the service fair and available. No third-party ad trackers. Payments and EU VAT are handled by our merchant of record (Polar); we never see your card. Full detail in the Privacy policy.

Questions about a source or licence?

Ask before you buy - we'll answer plainly.

Want the exact provenance of a dataset, a licence clarification, or a source you don't see yet? Get in touch.