Skip to content
IT & Collaboration · 11 vendors ranked

Best Data Catalog Software in 2026

A data catalog inventories what an organization's data actually is: tables, dashboards, files, and increasingly AI models, and records what each one means, who owns it, and where it came from. Data governance software layers policy on top of that inventory: access rules, stewardship workflows, and an audit trail of who touched what, and increasingly, the governed context an AI agent reads before it answers a business question. This ranking is aimed at data teams evaluating a catalog for the first time or replacing a spreadsheet glossary, not at enterprises already standardized company-wide on an incumbent. We judged how much of the first ninety days goes to connecting real sources versus a professional-services engagement, what the bill becomes once metadata volume and named seats are added, whether the glossary and lineage you build come back out cleanly if you leave, and which company or jurisdiction ultimately holds the index of your data.

What it is: Data catalog software inventories an organization's data assets, tables, dashboards, files and models, and records what each one means, who owns it and where it came from, usually by scanning source systems automatically and letting stewards add business context on top. Data governance software adds policy to that inventory: access rules, approval workflows, data-quality checks and an audit trail, and increasingly the governed context an AI agent reads before answering a question about the business.

Visibility in this ranking can be paid for. Payment moves a vendor's position within the shortlist; it never adds a vendor, and it never changes a word of the review. The largest vendors in data catalog software cannot hold places 1 to 3. How it works: placement disclosure · editorial process.

The top three

  1. #1

    Atlan

    Data teams building AI-agent context

    Reverse-engineers a live, queryable metadata graph from a company's warehouses and pipelines, aimed as much at AI agents as at analysts.

  2. #2

    Alation

    Teams wanting the most established catalog brand

    The original data-catalog company, now repositioned around consistent, governed answers for both human analysts and AI agents asking the same question.

  3. #3

    DataGalaxy

    Business teams that need a usable glossary interface

    A Lyon-built catalog designed so business analysts and stewards can use it directly, with a separate product line tracking data and AI initiatives.

How we ranked these

Setup: how much of the metadata arrives on its own

Setup here splits on how much of the metadata arrives automatically. Every tool on this list ships a connector-based scanner that crawls a warehouse, a BI tool or a file store and proposes an initial inventory within hours. What differs sharply is what happens next: turning that scanned inventory into a glossary anyone actually trusts. Atlan, Alation and DataGalaxy are built to let a business steward add definitions directly in a familiar interface; Collibra and the hyperscaler-native catalogs, Microsoft Purview and Google Cloud Knowledge Catalog, assume a dedicated governance function will own that work, often with a systems-integrator alongside it. OpenMetadata is the outlier: self-hosting it is a real infrastructure project before a single table gets tagged.

The real price: quote-only independents against metered hyperscalers

None of the independent vendors here publish a price; a buyer has to talk to sales before seeing a number for Atlan, Alation, DataGalaxy, eccenca, erwin Data Intelligence or Collibra. The hyperscaler-native catalogs break that pattern in the opposite direction: Microsoft Purview and Google Cloud Knowledge Catalog meter by usage, per governed asset or per API call, on a rate card anyone can read without a sales call, though that transparency only helps a buyer already committed to Azure or Google Cloud. AWS Glue Data Catalog goes furthest, with a genuinely published per-object and per-request rate. OpenMetadata is free to self-host at any size, and its managed tier from Collate keeps a real free plan for small teams before quote-only pricing starts.

Getting your metadata out: the glossary is the lock-in, not the software

A catalog's real lock-in is not the software, it is the glossary definitions, lineage graph and stewardship history a team spends months building inside it. Most vendors here offer an API-based export of that content, but almost none guarantee the export preserves definitions, approval history and lineage relationships exactly as built; DataGalaxy and eccenca, as smaller independent vendors, are the most direct about what does and does not travel. OpenMetadata is the one product where the data was never locked in to begin with, since the whole metadata store is queryable and exportable by design, open source software running on infrastructure a customer controls. Getting metadata out of a hyperscaler-native catalog means leaving the cloud platform it is tied to, not just the catalog.

Independence from the vendor: who can reach your metadata index

Eight of the eleven vendors here are US-incorporated, which puts customer metadata within reach of US legal process regardless of where the underlying data itself sits, a distinction from where servers physically are. Collibra complicates that reading further: founded in Brussels in 2008 and still holding a Belgian entity for EU privacy purposes, but its own site now describes the company as headquartered in both New York and Brussels, with its website terms governed by New York law. DataGalaxy and eccenca are the two genuinely EU-only entries, incorporated and operating entirely within France and Germany respectively. OpenMetadata is the only open-source option, letting a buyer read what the software does with its metadata before trusting it, and self-host it entirely outside any vendor's cloud.

Who it's for: three shapes of buyer, not one

Three shapes of buyer show up in this category. A small or mid-sized data team evaluating its first catalog is usually best served by an independent, business-user-friendly product, Atlan, Alation or DataGalaxy, that a steward can use without a governance office already in place. A large regulated enterprise with a dedicated data-governance function is the audience Collibra and erwin Data Intelligence are actually built for, with policy engines and approval workflows a smaller team will never fully use. A company already fully committed to one cloud platform is often better served by that platform's native catalog, Microsoft Purview, Google Cloud Knowledge Catalog or AWS Glue Data Catalog, than by a third-party tool duplicating metadata the platform already tracks.

Compared at a glance

#ToolBest forPricing modelFree optionHeadquarters
1Atlan Data teams building AI-agent contextQuote onlyNoneSingapore
2Alation Teams wanting the most established catalog brandQuote onlyNoneUnited States
3DataGalaxy Business teams that need a usable glossary interfaceQuote onlyNoneFrance
4eccenca Corporate Memory Teams wanting explainable, low-code governanceQuote onlyNoneGermany
5OpenMetadata Teams that want to read the source before trusting itOpen source + paid tiersFree planUnited States
6erwin Data Intelligence Regulated teams needing unstructured-data governanceQuote onlyNoneUnited States
7Collibra Large regulated enterprises with a governance office already in placeQuote onlyNoneBelgium / United States
8data.world Organizations already committed to ServiceNowQuote onlyNoneUnited States
9Microsoft Purview Unified Catalog Organizations already committed to Azure and Microsoft 365Usage-basedNoneUnited States
10AWS Glue Data Catalog Teams standardized on the AWS analytics stackUsage-basedFree planUnited States
11Google Cloud Knowledge Catalog Organizations standardized on Google Cloud and GeminiUsage-basedNoneUnited States

The 11 tools, reviewed

#1 Atlan

Active metadata built for AI agents · Singapore · atlan.com

Active metadata AI-agent context layer GIC-backed

Atlan reverse-engineers a live metadata graph out of a company's warehouses, BI tools and pipelines, then exposes it through open APIs so both people and AI agents query the same governed context instead of guessing at a table name. It raised a $105M Series C in May 2024 at a $750M valuation led by Singapore's GIC, so it is well-capitalised and not an acquisition target in obvious distress the way several competitors here now are.

Where it falls short

There is no published price and no self-serve signup, so evaluating it always means a sales cycle before a single real screenshot. The product's roadmap is visibly pointed at AI-agent context rather than classic catalog browsing, which is a mismatch for a buyer who wants a traditional glossary-first tool.

Wrong for

A small team that wants to see a price before talking to anyone, or a buyer evaluating this purely as a passive metadata browser rather than infrastructure for AI agents querying the business.

Pricing: Quote-only; no tiers or figures published on the pricing page (none)

Visit Atlan →

#2 Alation

The catalog category's original pioneer · United States · alation.com

Catalog pioneer AI governance layer Self-improving feedback loops

Alation built the modern data-catalog category and has spent the last two years reselling the same governed metadata as AI governance: the pitch is that a chatbot and a human analyst get the same answer to the same business question because both read from the same governed glossary. That consistency argument is real, and the company's decade of catalog-specific product depth is hard for newer entrants to match quickly.

Where it falls short

What is missing from the public site is a number: there is no published price, no free tier and no self-serve signup, so evaluating it means a sales cycle before a single screenshot. As one of the older products in this category, parts of the interface carry more legacy complexity than newer, narrower competitors.

Wrong for

A buyer who wants to self-serve a trial without a sales call, or a small team whose glossary needs are simple enough that a decade of enterprise feature depth is mostly unused weight.

Pricing: Quote-only; no tiers or figures published (none)

Visit Alation →

#3 DataGalaxy

Built for business users, not just engineers · France · datagalaxy.com

EU-incorporated, Lyon Business-user-first interface Data and AI initiative tracking

DataGalaxy was built in Lyon in 2015 around a specific bet: that a data catalog fails when only engineers can use it, so the interface, glossary and workflows are designed for business analysts and stewards first. It has since added a separate Portfolio product for tracking data and AI initiatives against governance requirements, and it is one of only two vendors in this ranking incorporated entirely within the EU.

Where it falls short

Pricing is entirely quote-only across both product lines, with no published starting figure anywhere on the site. As a mid-sized independent European vendor, it does not have the AI-agent integration depth or the funding scale of the largest US and Singapore-based competitors on this list.

Wrong for

A large enterprise wanting the deepest possible AI-agent integration roadmap, or a buyer specifically wanting a published starting price before any conversation with sales.

Pricing: Quote-only across Starter, Professional and Enterprise catalog tiers, and a separate Portfolio product line (none)

Visit DataGalaxy →

#4 eccenca Corporate Memory

Governance modelled as a knowledge graph · Germany · eccenca.com

EU-incorporated, Leipzig Knowledge-graph architecture Low-code for subject-matter experts

eccenca's Corporate Memory models an organisation's data and business rules as a semantic knowledge graph rather than a flat catalog table, with a stated goal of making the result explainable enough for AI systems built on top of it to be trusted. Low-code tooling is aimed at subject-matter experts rather than only data engineers, a genuine differentiator from the more developer-first tools on this list, and it is incorporated entirely within Germany.

Where it falls short

It is a smaller company with no published pricing anywhere and a considerably smaller public profile than the venture-funded US and Singapore competitors here, which makes it harder to evaluate against peers without a direct sales conversation. Knowledge-graph modelling also carries a steeper initial learning curve than a simpler flat catalog.

Wrong for

A buyer wanting the largest possible existing user community and public case-study library, or a team unwilling to invest time in modelling data as a semantic graph rather than a flat list.

Pricing: Quote-only, demo-gated; no published tiers (none)

Visit eccenca Corporate Memory →

#5 OpenMetadata

Open source catalog you can self-host · United States · open-metadata.org

Open source Self-hostable 13,500+ community members

OpenMetadata is a genuinely open-source project covering cataloguing, lineage, data quality, observability and governance in one unified metadata graph, which a team can self-host and read the source of before trusting it with anything. Collate, Inc. is the commercial company behind it, selling a managed version with a real free tier for small teams, up to 5 users and 500 assets, before quote-only pricing starts above that.

Where it falls short

Self-hosting is genuine ongoing operational work: patching, scaling and backups become the team's own responsibility rather than a vendor's. The managed product's higher tiers are no more transparently priced than the closed-source competitors on this list, so the pricing advantage disappears once an organization outgrows the free tier.

Wrong for

A team with no engineering capacity to run and maintain its own infrastructure, or an enterprise buyer specifically wanting a single accountable vendor to escalate every issue to.

Pricing: Free self-hosted open source; Collate's managed tier has a free plan for 5 users and 500 assets, Premium and Enterprise are quote-only (free plan)

Visit OpenMetadata →

#6 erwin Data Intelligence

Governs unstructured data and AI models too · United States · quest.com

Structured + unstructured data AI-model certification Owned by Quest Software

erwin's data-modelling roots go back to 1984, and Quest Software bought the company in October 2021 and still sells it as Quest erwin. The genuine differentiator against the pure-play catalogs on this list: it explicitly governs unstructured data and can certify an AI model's lineage and explainability, not just tag warehouse tables, which matters to a regulated buyer worried about more than structured databases.

Where it falls short

It is a smaller, less AI-agent-focused brand than Atlan or Alation now, with a public profile shaped more by its legacy data-modelling business than by the catalog product itself. Being one product inside a much larger Quest portfolio means the catalog does not get the company's full attention the way it does for an independent vendor.

Wrong for

A buyer wanting the newest, most AI-native product on the market, or a team specifically wary of buying a product whose roadmap sits inside a much larger, unrelated software portfolio.

Pricing: Quote-only; no published tiers (none)

Visit erwin Data Intelligence →

#7 Collibra

The category's largest established incumbent · Belgium / United States · collibra.com

Belgian-founded, 2008 Dual New York/Brussels HQ AI Command Center / Context Engine

Collibra was founded in Brussels in 2008 and is widely credited as Belgium's first software unicorn; its own about page now describes the company as headquartered in both New York and Brussels, and its website terms are governed by New York law. The product has grown from a governance-and-catalog tool into what Collibra calls an Enterprise AI Control Plane: an ontology-based Context Engine meant to feed governed metadata to AI agents, backed by the deepest enterprise feature set in this category.

Where it falls short

It is one of the largest, most established names in this category, which is also the trade-off: an incumbent's roadmap and pricing move on incumbent timelines, nothing about pricing is public, and its own dual New York-Brussels structure means it does not cleanly satisfy a strict EU-only procurement requirement.

Wrong for

A small team without a dedicated governance function, who will pay for enterprise policy-engine depth it will not use, and any buyer requiring a vendor genuinely and solely headquartered within the EU.

Pricing: Quote-only; no published tiers (none)

Visit Collibra →

#8 data.world

Knowledge-graph catalog, now part of ServiceNow · United States · data.world

Knowledge-graph architecture Now part of ServiceNow Governance + catalog

data.world built its catalog and governance product on a knowledge-graph architecture rather than a flat metadata table, aimed at connecting business meaning to technical assets in a way a simple table-and-tag model does not capture as well. Its own site now states plainly that data.world is now part of ServiceNow, which puts a large enterprise-platform company's resources behind the product.

Where it falls short

As an acquired company, it is no longer evaluated as an independent vendor; expect its roadmap, pricing and packaging to fold into ServiceNow's much larger enterprise-platform sales motion rather than continue as a standalone product with its own pricing page. No current pricing information is published on the site.

Wrong for

A buyer wanting to evaluate a standalone, independently-run catalog company, or anyone not already inside or planning to move onto the ServiceNow platform.

Pricing: Not published on the current site (none)

Visit data.world →

#9 Microsoft Purview Unified Catalog

Microsoft's metered, Azure-native catalog · United States · learn.microsoft.com

Consumption-based, published rates Governs data, AI models and M365 content together Formerly Azure Data Catalog / Azure Purview

Unified Catalog is the layer of Microsoft Purview where a data product, a business glossary term and the technical tables behind it get connected, sitting on top of Purview's Data Map scanning engine. Pricing is genuinely metered and published rather than quote-only, which is unusual in this category and lets a technical buyer estimate cost without a sales call.

Where it falls short

It has been renamed twice, from Azure Data Catalog (2016) to Azure Purview (2021) to Microsoft Purview (2022), which makes it easy to find outdated documentation and confuses procurement conversations. It only makes sense as a purchase for an organisation already committed to Azure and Microsoft 365; it is not a portable, engine-agnostic catalog for a multi-cloud or on-premises estate.

Wrong for

A company running its data estate outside Azure, or a buyer wanting one catalog that works identically across multiple cloud providers rather than one tied to Microsoft's own stack.

Pricing: Consumption-based: billed per governed data asset plus metadata-operation Capacity Units, published rate card via the Azure pricing calculator (none)

Visit Microsoft Purview Unified Catalog →

#10 AWS Glue Data Catalog

AWS's shared technical metastore · United States · aws.amazon.com

Real published usage pricing Hive-metastore compatible Shared across the AWS analytics stack

The Glue Data Catalog is deliberately narrower than the rest of this list: a technical metadata store, not a business glossary or a stewardship workflow, that Athena, Redshift Spectrum, EMR and Glue's own ETL jobs all read from so a table's schema is defined once. That narrowness is also the honest selling point, alongside genuinely published per-object and per-request pricing rather than a quote.

Where it falls short

It has no business glossary, no stewardship workflow and no approval chain of its own, so it is the wrong tool for a buyer looking for a governance product rather than a shared technical schema store. It is also the most tightly coupled to a single cloud provider's analytics stack of anything reviewed here.

Wrong for

A business team looking for a glossary and stewardship interface, or any organization not already running its analytics workloads on AWS.

Pricing: Published, usage-based: first 1M objects stored free, then $1.00 per 100,000 objects/month; first 1M requests free, then metered; crawlers billed at $0.44 per DPU-hour (free plan)

Visit AWS Glue Data Catalog →

#11 Google Cloud Knowledge Catalog

Google's Gemini-grounded semantic catalog · United States · docs.cloud.google.com

Renamed from Dataplex, April 2026 Gemini-powered semantic context Structured + unstructured coverage

Google's catalog product has changed its name three times in six years: Data Catalog (2020), merged into Dataplex as Dataplex Universal Catalog (2022), then renamed again to Knowledge Catalog in April 2026, with the API and CLI names left unchanged underneath. The current pitch is an active context graph that grounds Gemini and other AI agents in a company's real metadata, across both structured tables and unstructured files.

Where it falls short

The naming churn is a genuine risk: documentation, tutorials and job postings referencing Dataplex are already out of step with the product's current name, which makes evaluating it and finding reliable outside guidance harder than for a stably-named competitor. Pricing is consumption-based rather than a number a buyer can budget from a single page.

Wrong for

A buyer put off by repeated rebranding and naming instability, or any organization not already standardized on Google Cloud for its underlying data estate.

Pricing: Consumption-based via standard Google Cloud pricing, no flat published figure (none)

Visit Google Cloud Knowledge Catalog →

What the data says about this market

Acquisition has reshaped this category faster than most software markets in the past eighteen months. ServiceNow bought data.world, Atlassian announced its purchase of the catalog vendor Secoda on 4 December 2025, and Quest Software has owned erwin since October 2021; none of the three continues to set its own roadmap independently. Two vendors researched for this ranking, Zeenea and Castor, were left out entirely rather than listed, because both were acquired, by HCLSoftware/Actian and by Coalesce Automation respectively, and no longer operate as standalone products under their original names. A buyer shortlisting this category in 2026 is choosing a metadata layer as much as a company, and that company's independence is worth weighing alongside its feature list.

Eight of the eleven vendors ranked here are US-incorporated; Collibra's own site describes a dual Brussels-New York headquarters rather than a single EU base, and only DataGalaxy and eccenca are incorporated entirely within the EU. ICT services have grown faster as a share of global trade than most other service categories, from 9.1% of world service exports in 2013 to 14.48% in 2023, with the European Union's share rising even further, from 10.58% to 17.55% over the same period (World Bank). Metadata management sits underneath a meaningful slice of that trade, since almost no cross-border data service gets built without some inventory of what data it touches, and the concentration of catalog vendors in the US is a narrower geographic spread than that overall growth would suggest.

The largest cloud platforms have all renamed their catalog products within the last five years, a sign of how quickly the underlying pitch has shifted from passive inventory to AI-agent context. Google's offering alone has been called Data Catalog, then Dataplex Universal Catalog, then Knowledge Catalog as of April 2026; Microsoft folded Azure Data Catalog into Azure Purview in 2021 and renamed the whole line Microsoft Purview in 2022. Independent vendors have followed the same shift in emphasis rather than name: Atlan and Alation both now describe their product primarily as infrastructure for AI agents, not analysts, which is a materially different pitch from the one the category launched with a decade ago.

The 11 ranked vendors, counted

  • Headquarters by region: North America 7, Europe 2, Asia-Pacific 1, Other 1
  • By country: United States 7, Belgium / United States 1, France 1, Germany 1, Singapore 1
  • Pricing model: Quote only 7, Usage-based 3, Open source + paid tiers 1
  • Free option: None 9, Free plan 2

Counted from the 11 vendors on this page. More in our market data.

For the wider market behind data catalog software, read our report The Global Shift to ICT Services, or browse all industry reports.

How to choose

  1. Separate the connector question from the governance-maturity question

    Decide these separately, because the answers rarely point the same direction. Every catalog on this list can scan a warehouse and propose an inventory within a day; that part is close to a commodity. What is not a commodity is whether your organization has assigned anyone to own stewardship, definitions, approvals and data-quality rules, and buying enterprise governance tooling before that role exists produces an expensive, unused catalog. Price out who will actually write the first hundred glossary definitions before signing anything, not the per-seat licence fee. A small team without a dedicated governance function gets more value from a lighter, business-user-friendly catalog than from the policy engine a regulated enterprise needs.

  2. Ask for the export before you ask for the demo

    Request an export of glossary terms, lineage relationships and stewardship history in a format you can actually open, and ask whether that export is included in the contract or billed as a separate project. A vendor that hesitates on this question is telling you what a renewal negotiation will feel like in three years. OpenMetadata is the only entry here where this is a non-issue by design, since the whole metadata store is open source and queryable without asking permission; every other vendor's answer to this question is worth getting in writing before a purchase order is signed.

  3. Match the vendor's independence to your own compliance requirement

    If your organization has no data-residency requirement, the cheapest or most usable product on this list is the right starting point regardless of where the vendor is incorporated. If a compliance or public-sector procurement rule requires an EU-only metadata layer, treat Collibra's dual New York-Brussels structure as US exposure rather than a genuine EU option, and shortlist DataGalaxy or eccenca instead, both incorporated entirely within the EU. A company already standardized on one cloud platform should also weigh that platform's native catalog, Microsoft Purview, Google Cloud Knowledge Catalog or AWS Glue Data Catalog, against a third-party tool that will duplicate metadata the platform already tracks.

Questions and answers

What is the best data catalog software in 2026?

Atlan ranks #1 in this list because it pairs a genuinely active metadata graph, built by reverse-engineering warehouses and pipelines rather than waiting for manual tagging, with real venture backing, a $105M Series C in 2024 at a $750M valuation, that suggests it will still be an independent company next year. Alation is the better pick for a team that wants the original, most established catalog brand, and DataGalaxy is the better pick for a business team that wants a glossary interface non-engineers can use unaided.

Which vendors in this category are US companies, and does that matter?

Alation, OpenMetadata's commercial operator Collate, erwin Data Intelligence's owner Quest Software, data.world, Microsoft and the company behind AWS Glue Data Catalog are all US-incorporated, which puts customer metadata within reach of US legal process regardless of where the underlying data itself is hosted. Collibra's own site describes a dual New York-Brussels headquarters rather than a single EU base. DataGalaxy and eccenca are the two vendors here incorporated entirely within the EU, in France and Germany respectively. Whether that distinction matters depends entirely on a buyer's own compliance requirements.

Which of these tools publish real pricing?

Three do. Microsoft Purview Unified Catalog and Google Cloud Knowledge Catalog both meter by usage on a published, if complex, rate card; AWS Glue Data Catalog goes furthest with a genuinely simple published rate: the first million objects stored free, then $1.00 per 100,000 objects a month, plus metered requests and $0.44 per DPU-hour for crawlers. OpenMetadata is free to self-host as open source, with a real free managed tier from Collate for small teams. The other seven, Atlan, Alation, DataGalaxy, eccenca, erwin Data Intelligence, Collibra and data.world, are quote-only.

Is there a free or open-source data catalog?

OpenMetadata is the one genuinely open-source option on this list: free to self-host at any size, with the source code available to read before trusting it with your metadata. Its commercial operator, Collate, also offers a managed cloud tier with a real free plan for up to 5 users and 500 assets before quote-only pricing starts above that. None of the other ten vendors here publish source code or offer a permanent free tier of their own hosted product.

What happened to Zeenea and Castor, two well-known European data catalog vendors?

Both were acquired and no longer operate as independent standalone products. Zeenea, originally headquartered in Paris, was bought by HCLSoftware in September 2024 and folded into Actian's Data Intelligence Platform; its old site now redirects entirely and its legal-notice page returns a Gone response. Castor, also originally Paris-based and known as CastorDoc, was acquired by Coalesce Automation, a Delaware company, in March 2025 and rebranded as Coalesce Catalog. Neither is listed in this ranking, since reviewing a product under a name and ownership structure it no longer has would mislead readers rather than help them.

Which of these tools are open to self-hosting versus hosted-only?

OpenMetadata is the only genuinely self-hostable option, and it is the only one built as open source software from the start. Every other vendor on this list, from Atlan and Alation down to the hyperscaler-native catalogs, is sold as a hosted service where the vendor decides where the metadata index physically sits, not the customer. That includes Collibra, whose product runs entirely on Collibra's own infrastructure regardless of where a customer's underlying data lives.

How is a data catalog different from data governance software?

In practice the two have merged into one product category, and every vendor on this list sells both under one roof, but the distinction is still useful. A catalog answers what data do we have and what does it mean: an inventory, a glossary, a lineage graph. Governance answers who is allowed to see it and who approved that: access policy, stewardship workflows, approval chains and an audit trail. A team evaluating this category should ask which half of that job it actually needs solved first, since the lighter catalog-first tools here are not worse products, they are aimed at a narrower job.

Why do Microsoft, Google and Amazon all rank below the independent vendors?

Because their catalogs are the most locked to a single cloud platform of anything on this list: Microsoft Purview only makes sense inside Azure and Microsoft 365, Google Cloud Knowledge Catalog inside Google Cloud, and AWS Glue Data Catalog inside the AWS analytics stack. A buyer who is not already committed to that specific platform gains little from the published, metered pricing that is otherwise a genuine advantage over the quote-only independents. Their position here reflects portability, not capability; a company fully standardized on one of these clouds may reasonably prefer its native catalog to any third-party tool on this list.

Is Collibra a European company?

Partly, and its own materials are explicit about the ambiguity. Collibra was founded in Brussels in 2008 and is widely credited as Belgium's first software unicorn, and it maintains a real Belgian entity, Collibra Belgium BV, for EU privacy purposes. But Collibra's own about page now describes the company as headquartered in both New York and Brussels, and its website terms of use are governed by New York law with exclusive jurisdiction in New York courts. For a buyer with a strict EU-only procurement requirement, that dual structure is closer to US exposure than to a genuine EU vendor.

What should a buyer with no dedicated governance team pick?

Atlan, Alation or DataGalaxy, in that order for this ranking, since all three are built to let a business analyst or steward add glossary definitions directly rather than assuming a governance office will run the rollout. Collibra, erwin Data Intelligence and the hyperscaler-native catalogs assume more organizational infrastructure already exists around the product, and a smaller team will pay for policy-engine depth it never fully uses.

How often is this data catalog software ranking updated?

Whenever a fact underneath it changes: an acquisition like ServiceNow's purchase of data.world, a pricing change, or a rename like Google's shift from Dataplex Universal Catalog to Knowledge Catalog in April 2026. The published and last-reviewed dates at the top of this guide are real, and a review means someone checked the vendor's current documentation, not that a date was moved forward without a re-check.