PAKDATAHUB
Docs · Introduction

For AI agents: this page is available as raw Markdown at /docs/sources-provenance.md · every page in one file at /llms-full.txt · index at /docs/llms.txt

Sources & provenance

Every number in PakDataHub is an official figure published by a Pakistani public institution. PakDataHub doesn't estimate, model or re-survey anything. It collects the official releases, normalizes them into one consistent format, validates them and serves them through one API. For research and financial applications, the source of truth is the publisher. PakDataHub gives you that publisher's number, labelled with where it came from.

PakDataHub is an independent service and isn't affiliated with the institutions below.

The publishers

Publisher What PakDataHub takes from it How
State Bank of Pakistan (SBP) Exchange rates, KIBOR, policy rate, reserves, balance of payments, remittances, money and banking, external debt, public finance, GDP/real-sector compilations, social-sector data, payment systems, SME finance SBP's open-data portal EasyData (every dataset, via its API), SBP's website (KIBOR, policy rate, auctions) and SBP publications (Payment Systems Review, SME Finance Review)
Pakistan Bureau of Statistics (PBS) CPI and CPI by group, WPI, weekly SPI prices in 17 cities, Large-Scale Manufacturing, external trade by commodity Official PBS releases (PDF and Excel)
Pakistan Telecommunication Authority (PTA) Subscribers by operator, teledensity, cell sites, ARPU, broadband, sector revenue and investment PTA telecom indicators
Mutual Funds Association of Pakistan (MUFAP) Fund NAVs, returns, payouts, expense ratios, allocations; PKRV/PKISRV yield curves; debt-security prices and trades MUFAP industry statistics and valuation files

The SBP data compiles figures from other official bodies, including PAMA (autos), NEPRA (power), OCAC (petroleum), NFDC (fertilizer), APCMA (cement), the Ministry of Finance, FBR and NIPS. Each series names its publisher.

Tracing any value

Every catalog record has two provenance fields:

  • source: the publisher (SBP, PBS, PTA or MUFAP).
  • source_url: the exact official dataset or page the series is ingested from.
curl "https://api.pakdatahub.com/v1/catalog/fx.rate.avg.usd"
# ... "source": "SBP",
#     "source_url": "https://easydata.sbp.org.pk/api/v1/series/TS_GP_ER_FAERPKR_M.E00220/data"

For SBP EasyData series, the source_url contains SBP's own dataset and series code (here TS_GP_ER_FAERPKR_M.E00220), so any value can be checked against the SBP original.

How data flows

  1. Fetch: scheduled jobs download each official release as it's published: daily for KIBOR, curves and fund NAVs; weekly for SPI; monthly and quarterly for the rest.
  2. Archive: the raw file is stored before parsing, so every value can be re-derived.
  3. Parse & normalize: into (series_id, date, value, dims) with consistent ids, dates (month-end or quarter-end) and units. SBP publishes many values without stating the scale, so each dataset's scale (thousand, million, billion) was checked against known published totals and labelled in unit.
  4. Validate: each value is checked against per-series bounds and step limits. Values that fail aren't stored and are raised for review, and a single bad file never aborts a run. For example, a misread cell in a PBS price PDF is dropped rather than published.
  5. Cross-validate: derived figures are checked against the publisher's own derived figures. For example, CPI year-on-year computed from the index matches PBS's published YoY to within 0.05 percentage points.
  6. Version: when a publisher revises a value, the previous value is kept. See revision vintages.
  7. Data-quality sweep (twice a day): some glitches only show once a point has neighbours. Nothing is silently discarded; every change is kept in an audit table and the revision history.
    • Quarantined: isolated fund-NAV spikes (e.g. a NAV of 1.00 between two days of ~1,180), non-positive NAVs, impossible debt-security prices (e.g. a spreadsheet date serial where a price should be), and impossible auction yields.
    • Corrected: unambiguous unit slips in SBP level series, where a single point is published at 1,000x or 1,000,000x its usual scale and rescaling it lands exactly between its neighbours. Percent and index series are never altered.
  8. Monitor: a public status page shows freshness per module against each source's real release calendar.

PakDataHub vs going to each source directly

The official portals are authoritative, and PakDataHub is built on them. What PakDataHub adds:

Official portals (SBP EasyData, PBS, PTA, MUFAP) PakDataHub
Coverage One institution each SBP + PBS + PTA + MUFAP in one catalog
Access Separate sites and formats: API (EasyData, sign-in required), PDFs, Excel, web tables One REST API, one key, JSON or CSV
Identifiers Source codes (e.g. TS_GP_ER_FAERPKR_M.E00220) Readable, stable ids (fx.rate.avg.usd), with the source code kept for traceability
Units Scale often unstated Scale labelled (usd_mn, pkr_bn, …)
Price and fund data PDFs (SPI) and web tables (MUFAP) Clean series: 30 items × 17 cities, 567 funds with NAV history
Extras - Server-side transforms, revision vintages, webhooks, fund analytics

If you need a single series once, the official portal is fine. If you're building a product, model or dashboard that spans institutions, PakDataHub saves you the scraping, cleaning and upkeep. The numbers are the same.

What PakDataHub doesn't do

  • It doesn't produce its own estimates or forecasts. Every value is published.
  • The one exception is the clearly labelled derived layer (source: PakDataHub, with a derived object giving the formula and inputs): real interest rates, spreads, import cover, essentials inflation by city, fund-category yields, the weekly cost-of-living index (cost_of_living.*) and SPI national averages. Fund analytics and server-side transforms are computed the same way, on request.
  • It doesn't cover PSX equities. For those, see pypsx.com.
View as Markdown (.md)