// Data
Data governance: data with an owner, a definition and a change history
Nobody can tell you within an hour which reports will break if one table changes. We sort out ownership, definitions and lineage so that changes, audits and AI stop running on guesswork.
// In short
What is data governance
Data governance is the set of rules and tools that gives every important dataset an owner, a documented definition and a known origin. We set it up for companies whose reports are built outside their core systems, whose auditors ask for an asset inventory and whose teams want to feed data to AI. We start with one data flow, not a year-long program.
Without it, the data team is afraid to change anything, and the business builds its own reports in spreadsheets and private databases.
// What we do
From a catalog to data that is ready for AI
We pick the tool last. First we agree who is responsible for the data and what it means.
01Catalog
A data catalog with owners
Every important dataset has a business owner and a technical steward.
- Inventory of datasets and source systems
- An owner and a steward for each domain
- Responsibilities recorded in the catalog
02Definitions and lineage
Documented definitions and lineage
You know where a number in a report comes from and which transformations it went through.
- Business glossary
- Lineage from source to report
- One agreed definition per KPI
03Change
Impact analysis in change management
Before a table changes, you can see which reports, models and external recipients will feel it.
- Impact analysis built into the change process
- Register of data flows to third parties
- Input for the asset inventory and the DORA register of information
04AI
Access and classification ready for AI
The model and RAG search only see what each user is allowed to see.
- Data classified by sensitivity
- Permissions documented in the catalog
- Datasets approved for AI use
First step: one data flow end to end in 5 weeks
Flow map with owners
We take one flow: from the source, through the pipeline, to the report and external recipients. Every stage gets an owner.
Definitions and lineage
We document the definitions and lineage of that flow in a catalog, using a tool you already have or one we choose together.
Gaps and a 90-day roadmap
A prioritized list of control gaps and a plan for the next domains.
The result: a working data catalog for one domain
Fixed scope and date, agreed in writing before we start. Request the first step
// Scope
What we do, and what we don't
We work with the tool that fits your platform: Unity Catalog, Microsoft Purview, DataHub, OpenMetadata or dbt docs.
We do
- A data catalog with owners and documented definitions
- Lineage and impact analysis for the flows that matter
- Data classification and an access model for AI and RAG
- Handover of the process and documentation to your data stewards
We don't
- A year-long governance program on slides
- Buying a tool before anyone knows who owns the data
- Documenting every table at once
- Reselling licenses
If all you need is a document for the inspection file, we'll be upfront about that on the first call.
// Read more
Related topics and services
We write about data and AI on our blog:
// FAQ
Questions about data governance
Do we need a new tool?
Usually not to start with. A catalog can begin in Unity Catalog, Microsoft Purview, DataHub, OpenMetadata or dbt docs if you already use one of them. We choose a tool once it is clear who owns the data and what needs documenting. We don't resell licenses.
Who owns the data after the project?
You do. Domain owners are people in your company, listed in the catalog by name and role. We set up the process and the tool, show your data stewards how to work with the catalog and hand over the documentation.
How does this help with NIS2 and DORA?
If you operate in Poland, NIS2 in Poland (the KSC Act) requires asset management among other measures, and these Chapter 3 obligations apply from 3 April 2027 (as of 24 September 2026, gov.pl). A catalog with owners is a ready-made inventory of information assets. Under DORA, the same data feeds the register of information on ICT third-party arrangements, which financial entities in Poland submit to KNF.
How does this prepare data for AI?
A model cannot guess which of three tables with the same name is current, who owns the margin definition, or which data must never leave the company. The catalog records it, while classification and permissions decide what RAG search must not show a given user. Without that, AI runs on guesswork.
How long does it take?
The first step takes 5 weeks and covers one data flow. We plan the next domains in a 90-day roadmap, in the order set by your reports, audits and AI projects.
Let's start with one data flow
Describe the flow that causes the most trouble today: a board report, data for a customer or input for an AI model. We reply within one business day, with questions about scope.
Request the first stepPrefer to talk first?
30 minutes about your data, with no sales pitch.
Book a 30-minute call (opens in a new tab)