Public datasets arrive with different schemas, identifiers, and refresh cycles. CLOUDSUFI built repeatable workflows to ingest, validate, and connect them at scale, with human review for matching exceptions.
Google Data Commons is an open data platform that aggregates and publishes datasets from governments, research institutions, and organizations around the world — census data, climate records, economic indicators, health statistics, and more. The mission is straightforward: make the world’s public data accessible and useful.
The execution is anything but simple. Data Commons imports and refreshes datasets from dozens of sources continuously. Each source has its own schema, its own update cadence, its own identifiers. Keeping thousands of datasets in sync — accurately, at scale, without manual intervention — was the core engineering challenge.
Cross-referencing entities across sources was slow. The same place, variable, or statistical concept could appear under different names in different datasets. Manual reconciliation ate engineering time. And as the number of datasets grew, the infrastructure couldn’t keep pace — more data coming in, no way to process it faster without throwing more people at the problem.

The answer was a knowledge graph built on BigQuery — not a static database, but a living structure where every entity, relationship, and attribute connects to everything else. Google Data Commons brought in CloudSufi to design and build the AI layer that made it work at scale.
Three innovations did the heavy lifting. First, AI schema mapping that learns new dataset structures on its own. When a new source lands, the system reads its shape, figures out how it maps to the existing graph, and plugs it in — no manual integration code required. Second, an entity resolution engine built on Vertex AI that matches statistical variables and place identifiers across sources with accuracy that manual methods couldn’t touch. It catches mismatches, merges fragments, and builds a single clean identity — even when names differ, formats don’t match, and IDs disagree. Third, an event-driven architecture on Google Cloud that updates the graph in real time. No nightly batches. No stale snapshots. The graph stays current as datasets flow in from governments and research institutions worldwide.
The result was a shift from a fragmented import pipeline to a compounding intelligence asset. Each new dataset doesn’t just add rows to BigQuery — it enriches every entity already in the graph. Connections that were invisible before surface on their own. The more data goes in, the more useful the whole platform becomes.
Data throughput grew 5× — and the Data Commons team didn’t hire a single new engineer to handle it. The knowledge graph and AI automation absorbed the volume. Dataset imports that used to require manual schema work now happen on their own.
Manual processing dropped by 40%. Schema mapping, data integration, and record matching — tasks that used to eat weeks of engineering time — are now handled by AI. Engineers moved from plumbing to product work.
Entity resolution accuracy jumped from 60–70% to 95%+. Millions of identity records resolved correctly. Ghost records stopped polluting reports. Downstream analytics finally had a foundation they could trust.
But the biggest change is structural. The graph compounds. Every new dataset imported into BigQuery doesn’t just sit alongside the old ones — it makes the old ones more valuable. Connections between statistical variables surface automatically. What started as a data engineering problem became the backbone of one of the world’s most comprehensive open data platforms.
The knowledge graph is now the foundation, not the finish line. The Data Commons team is expanding into deeper cross-domain analysis — using the graph’s connected structure to surface correlations between climate, economic, and health datasets that no single source could show alone.
New government and institutional data sources are continuously being onboarded into BigQuery. Natural language querying via Vertex AI is expanding the platform’s reach — so researchers and policymakers can ask questions in plain English instead of writing SQL against a schema they need to memorize.
The stagnant repository is gone. In its place is an engine that gets better every time new data arrives.
Let's build what's next.
Talk to us about your data and AI challenges — and how CLOUDSUFI can help solve them.
Talk to us →