purposed2.

Intelligence · Desk Research

Read a market on evidence, not on the first chart you find.

Free public data can tell you a great deal about a market — and it can also mislead you in specific, well-documented ways. This page is honest about both: what each of seven verified sources actually covers, the sampling and survivorship traps that catch desk researchers, and why triangulating two or three sources beats trusting one.

7Free data sources, verified
0Invented statistics used
1943Origin year of the bias case study
100%Free & publicly accessible

Every ledger row below was checked against the source's own official page — see “About this page” under Sources.

01 — The Honest Starting Point

What public data can — and cannot — tell you.

Free public data is genuinely good at three things: showing you direction (is a market growing or shrinking), relative size (is this country's trade volume larger or smaller than that one's), and structure (which sectors, partners or demographics dominate). Every source in the ledger below can do this reliably, because it's what national statistical offices, central banks and multilateral bodies are built to publish.

It is much weaker at three other things it's frequently asked to do anyway. First, precise point-in-time figures for a narrow slice (a specific city, a specific product code, a specific month) — the finer you slice, the more a single source's sampling and reporting lag start to matter. Second, the informal economy — surveys and administrative filings systematically undercount activity that never registers with a tax authority or customs post, which matters enormously in markets where informal trade is large. Third, causal explanation — a data source can show you that two things moved together; it cannot, by itself, tell you that one caused the other.

The discipline this page argues for is simple to state and easy to skip under deadline pressure: know which of the three things you actually need before you go looking, and treat any single number as a claim to be checked, not a fact to be repeated.

💬 Chat with our AI about this →

02 — The Source Ledger

Seven verified free public data sources.

Sortable (click a column) and filterable (type below). Every row was checked against the source's own official page on 27 Jul 2026 — nothing here is a second-hand summary of a summary.

7 of 7 sources
Official link
World Bank Open Data ~16,000 development indicators across 45+ databases — economy, education, health, environment, gender, poverty, trade, technology — for essentially every country the Bank tracks. Mostly country-year; some regional/income-group aggregates. Varies by indicator; API supports quarterly/monthly/yearly queries — no single global refresh date. Coverage gaps for smaller economies and recent years (1–3-year lags common); definitions can shift between revisions. data.worldbank.org
RBI Database on Indian Economy India's real sector, corporate sector, financial sector, financial markets, external sector, public finance and socio-economic indicators. Mostly national and sector-level; some series (banking outlets) sub-national. Varies by series — many banking/monetary series weekly or monthly; national-accounts series follow India's official release calendar. India-specific only; portal migrated from dbie.rbi.org.in to data.rbi.org.in in Jun 2024, so old links may break; some series revised under the 2023 CIMS overhaul. data.rbi.org.in
MoSPI / National Statistical Office India's headline statistics — GDP/National Accounts, CPI, Index of Industrial Production, Periodic Labour Force Survey, Household Consumption, Annual Survey of Industries, Economic Census. National and state-level for most headline series; PLFS and consumption surveys split rural/urban. Set by an official release calendar — e.g. CPI monthly, GDP quarterly, Economic Census only every several years. Survey-based figures carry sampling error; India's large informal sector is harder to capture in establishment-based surveys; base-year revisions break long time series. mospi.gov.in
UN Comtrade Official merchandise trade statistics from ~200 countries/areas — by commodity, partner and trade flow, annual data back to 1962. Country-pair, commodity-code level (down to 6-digit HS in the standard dataset). Continuous, as national authorities submit — no fixed global publication date, so reporters lag by months. Mirror-data mismatches between reported exports and partner-reported imports (valuation/timing/re-export differences); services trade far less complete than goods. comtradeplus.un.org
Our World in Data Cross-topic aggregation — health, energy, climate, poverty, demographics, education — harmonising WHO, World Bank, UN agencies, IEA and peer-reviewed sources. Mostly country-year; granularity depends entirely on the underlying source being re-published. Varies by topic — each dataset re-syncs on its own schedule, documented per chart. A secondary aggregator, not a primary collector — only as reliable as its underlying source; some series (e.g. certain health/poverty figures) are modelled estimates, not direct measurements. ourworldindata.org
OpenAlex Open scholarly metadata — hundreds of millions of works, plus authors, institutions, journals and topics — aggregated from Crossref, PubMed, ORCID and ROR. Per-work record (title, authors, venue, citations, topics, open-access status). Documented as updated daily from its upstream sources. Metadata only, no full text; coverage skews toward what Crossref/PubMed index; since Feb 2025 a free API key is needed for high-volume use (100 credits/day without one). openalex.org
Google Dataset Search A search index over dataset metadata published anywhere on the web using schema.org/Dataset or DCAT markup — government portals, university repositories, data-journalism projects. Per-dataset listing (title, provider, description, format, licence) — a pointer to the data, not the data itself. Continuous web-crawl based, same as general Google Search — no fixed refresh schedule. Only surfaces datasets whose publishers added correct structured markup — absence from results is not evidence a dataset doesn't exist. datasetsearch.research.google.com

💬 Ask the AI which sources fit your question →

03 — Sampling & Survivorship Bias

The data you don't see is often the point.

Sampling bias is the simpler trap: a survey or dataset systematically over- or under-represents part of the population it claims to describe. India's household consumption and labour-force surveys, for instance, sample households — which is the right method, but means the informal, unregistered edges of the economy are structurally harder to capture than formally employed, urban respondents.

Survivorship bias is subtler and easy to miss because the data that would correct it has, by definition, already vanished. The canonical illustration is Abraham Wald's 1943 work for the U.S. Army's Statistical Research Group: asked to recommend where to add armor to bombers based on the damage pattern of planes returning from missions, Wald argued for reinforcing the spots with the fewest hits — because planes hit in the engines or other critical areas were the ones that never made it back to be counted. The damage visible in the surviving sample was, precisely because of survivorship, the least informative part of the picture.

Worth being honest about the story itself: Wald's own 1943 memorandum doesn't use the phrase "survivorship bias" or tell the vivid "bullet holes" anecdote in the form it's usually retold today — that framing is a later, simplified reconstruction. The rigorous technical treatment of what Wald actually did is Mangel & Samaniego's 1984 paper in the Journal of the American Statistical Association (source 9 below). The lesson still holds even with the popular telling softened: when you only observe the survivors of a market, a channel, or a customer segment, ask what left the sample before you could measure it.

💬 Chat with our AI about this →

04 — Triangulating Sources

One source is a claim. Two that agree are evidence.

The UN's own Fundamental Principles of Official Statistics — endorsed by the General Assembly in January 2014 — are explicit that trust in official data rests on transparency about method, not on authority alone. Practically, that means: before you build a decision on one number, check whether an independent source corroborates the shape of it.

  • Sizing a market in India? Cross-check MoSPI's national-accounts figure against a relevant RBI sectoral series and, where the market has an international angle, a World Bank indicator — three independent statistical authorities telling a consistent story is worth far more than any one of them alone.
  • Sizing a bilateral trade flow? Pull both sides of a UN Comtrade mirror pair (what the exporter reports vs. what the importer reports) — a large, unexplained gap is itself a signal worth investigating, not an error to average away.
  • Grounding a claim in research? Use OpenAlex to check whether a "well-known" statistic actually traces to a peer-reviewed study, and Google Dataset Search to see whether the underlying microdata is published anywhere you could re-check it yourself.
  • Working across many topics at once? Our World in Data is a fast way to see a harmonised chart — but its own FAQ is clear that it re-publishes others' data, so treat it as a fast first look, then verify anything decision-critical against the primary source it cites.

💬 Chat with our AI about this →

05 — FAQ

Common questions.

Is free public data good enough to make a real business decision?

Sometimes, for the shape of a market — growth direction, relative size, broad trends. Rarely for a single precise number you would stake money on without checking whether at least one other source agrees with it.

What's the single biggest mistake people make reading market data?

Treating one number, from one source, as settled fact — without checking what it actually measures, how it was collected, or whether an independent source corroborates it.

What is survivorship bias, in one sentence?

Drawing conclusions only from the cases that "survived" to become visible to you, while the cases that didn't survive — and might tell a very different story — stay invisible. Abraham Wald's WWII aircraft-armor analysis is the canonical illustration, though the popular retelling simplifies what his own 1943 memo actually argued.

Why do UN Comtrade's import and export figures for the same trade flow sometimes disagree?

Because two different national statistical authorities report each side of the same flow — the exporter and the importer — using different valuation conventions, timing and treatment of re-exports. This is a documented, known feature of mirror trade statistics, not a data error.

Do any of the seven sources in the ledger cost money?

No — all seven are free. OpenAlex now asks for a free API key for higher-volume use (100 requests/day without one, 100,000/day with a free key), but the underlying data carries no charge anywhere in the ledger.

06 — Sources

Where the ledger and the arguments come from

  1. World Bank, “About the Indicators API Documentation” — indicator count, database count and query-frequency support.datahelpdesk.worldbank.org · accessed 27 Jul 2026
  2. Reserve Bank of India, Database on Indian Economy (DBIE) — subject-area coverage and the June 2024 URL migration to data.rbi.org.in.data.rbi.org.in · accessed 27 Jul 2026
  3. Ministry of Statistics and Programme Implementation (MoSPI) & National Statistical Office, Government of India — headline indicator list and survey methodology.mospi.gov.in · accessed 27 Jul 2026
  4. United Nations Statistics Division, UN Comtrade — reporter coverage (∼200 countries/areas), history back to 1962, and mirror-statistics documentation.comtradeplus.un.org · unstats.un.org/unsd/trade · accessed 27 Jul 2026
  5. Our World in Data, “FAQs and User Guidelines” and “About” — sourcing and harmonisation methodology.ourworldindata.org/faqs · accessed 27 Jul 2026
  6. OpenAlex technical documentation — coverage, update frequency and the February 2025 API-key change.docs.openalex.org · help.openalex.org · accessed 27 Jul 2026
  7. Google, Dataset Search Help & Google Search Central — schema.org/Dataset and DCAT indexing method.datasetsearch.research.google.com/help · accessed 27 Jul 2026
  8. Wald, A. (1943). “A Method of Estimating Plane Vulnerability Based on Damage of Survivors.” Statistical Research Group, Columbia University (Center for Naval Analyses reprint, 1980) — the original aircraft-survivability memorandum.
  9. Mangel, M. & Samaniego, F.J. (1984). “Abraham Wald's Work on Aircraft Survivability.” Journal of the American Statistical Association, 79(386), 259–267 — the rigorous academic reappraisal cited above.
  10. United Nations Statistics Division, Fundamental Principles of Official Statistics (UN General Assembly resolution A/RES/68/261, 29 Jan 2014).unstats.un.org/fpos · accessed 27 Jul 2026

About this page: the ledger above lists only sources this page's author checked directly against their own official page on the date shown. No coverage figure, update cadence or limitation was invented or estimated — where a source's own documentation was ambiguous about a detail (e.g. exact update timing), the ledger says so rather than asserting a specific number.

The Concierge

Talk through your market question.

The purposed2 concierge can help you decide which sources to check first for your specific question, or explain any part of the reasoning above in plain words — and will say when a working session with Amit is the better next step.

Ask Amit directly →