Baike.dev
All toolsTrendingOpen sourceNewsSubmit
Log in
< 返回工具列表
O

OpenMetadata

> 编程语言
开源

The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants,

14.6K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants,

OpenMetadata

The Open Context Layer for AI

The largest and fastest-growing open-source project for AI context, data cataloging, and metadata management.

OpenMetadata is the open platform for trusted data context, organizational memory, and business semantics for every data user, AI assistant, and agent.

OpenMetadata connects technical metadata, data quality signals, lineage, column-level lineage, ownership, usage, policies, conversations, memories, glossaries, classifications, metrics, domains, data contracts, and data products into a unified metadata knowledge graph. With 130+ connectors, open metadata standards, semantic search, APIs, SDKs, and an MCP server, OpenMetadata gives every user and AI system the governed context it needs to discover, understand, trust, remember, and use data.

AI does not need another raw database connector. AI needs context + memory.

OpenMetadata provides the context AI needs to know:

  • what data exists
  • what it means
  • who owns it
  • how it is used
  • where it came from
  • where it flows
  • whether it is fresh, tested, certified, and trusted
  • which business concepts, classifications, glossary terms, policies, contracts, and data products apply
  • what downstream dashboards, pipelines, metrics, ML models, and applications depend on it
  • what conversations, decisions, assumptions, and memory nuggets have already been captured about it

Why OpenMetadata for AI?

AI systems need more than data access. They need governed context, business meaning, trust signals, lineage, usage, ownership, standards, and organizational memory.

A direct connection to a warehouse, lake, dashboard, or pipeline exposes raw structures. It does not tell an AI assistant what the data means, whether it is certified, who owns it, which policies apply, what contract governs it, what breaks if it changes, or what the organization has already learned about it.

OpenMetadata is the open context layer that gives every data user and AI agent the full picture of enterprise data.

OpenMetadata brings together five capabilities:

  1. Context — technical, operational, trust, usage, and lineage metadata from across the data ecosystem.
  2. Semantics — business meaning through glossaries, metrics, classifications, domains, policies, ontologies, and data products.
  3. Knowledge Graph — relationships connecting assets, columns, people, teams, quality, lineage, policies, memories, contracts, and business concepts.
  4. Memory — conversations, AI threads, decisions, assumptions, runbooks, remediation notes, and reusable memory nuggets that preserve tribal knowledge.
  5. Activation — MCP, Semantic Search, APIs, SDKs, events, and workflows that make context usable by AI assistants, agents, applications, and humans.

With OpenMetadata, users and AI agents can answer:

  • What does this metric mean and how is it calculated?
  • Which datasets power this dashboard?
  • Who owns this data product?
  • Which data contract applies?
  • Is this dataset fresh, tested, certified, and trusted?
  • Which downstream dashboards, pipelines, or ML models are affected by this column change?
  • Which columns contain sensitive customer information?
  • Which glossary terms, policies, standards, and business concepts apply?
  • What decisions, assumptions, incidents, or conversations have already been captured about this asset?

The Context OpenMetadata Connects

OpenMetadata collects and connects the context AI needs to reason safely over enterprise data.

Context type What OpenMetadata captures Why it matters for AI Technical metadata Databases, schemas, tables, columns, topics, dashboards, charts, pipelines, APIs, search indexes, ML models, storage assets, data types, constraints, descriptions, joins, sample queries, service metadata, owners, teams, usage, domains, and data products Helps AI discover what exists and understand how assets are structured Quality and trust Test cases, test suites, freshness checks, volume checks, null, uniqueness, distribution, custom tests, profiling results, observability signals, incidents, alerts, and quality history Helps AI avoid treating every dataset as equally trustworthy Lineage and impact Upstream and downstream lineage, table lineage, column-level lineage, dashboard lineage, pipeline lineage, metric lineage, ML model lineage, API and topic dependencies, and OpenLineage events Helps AI explain where data came from, where it flows, and what changes may break Semantics Glossaries, business terms, synonyms, related terms, metrics, KPIs, classifications, tags, domains, data products, policies, personas, lifecycle states, and ontologies Helps AI map technical names to business meaning Governance Owners, stewards, teams, policies, roles, classifications, access context, certification, review workflows, lifecycle states, and data contracts Helps AI act with policy-aware context Memory and tribal knowledge Conversations, AI threads, decisions, assumptions, runbooks, remediation notes, incident learnings, and reusable memory nuggets attached to assets, users, teams, data products, and agent workflows Helps humans and agents inherit what the organization already learned instead of rediscovering it in every conversation Standards and interoperability DCAT, DPROD, PROV-O, OpenLineage, ODCS, RDF/OWL, JSON-LD, SHACL, JSON Schema, APIs, events, and metadata schemas Helps context move across tools, agents, catalogs, contracts, and knowledge graphs

Architecture: Context + Memory Graph

OpenMetadata is built around an open, schema-first metadata graph.

  1. Collect metadata from warehouses, lakes, BI tools, pipelines, ML platforms, messaging systems, storage systems, APIs, search systems, SaaS applications, metadata systems, documents, conversations, and agent workflows through 130+ connectors, ingestion APIs, events, and SDKs.
  2. Normalize metadata with open schemas and standards so every asset, relationship, policy, contract, lineage event, and memory can be represented consistently.
  3. Connect technical metadata, quality signals, lineage, ownership, usage, policies, conversations, memories, semantics, domains, contracts, and data products into one graph.
  4. Preserve Memory by turning conversations, AI threads, decisions, assumptions, runbooks, and remediation notes into reusable governed memory nuggets tied to data assets and business context.
  5. Govern context with open standards, classifications, policies, roles, data quality, review workflows, data contracts, and stewardship.
  6. Activate that context through Semantic Search, MCP, APIs, SDKs, events, webhooks, metadata applications, and AI workflows.

Memory is part of the architecture, not a side channel. It lets engineers use APIs, SDKs, MCP, or AI workflows to preserve conversational context and convert tribal knowledge into reusable organizational knowledge.


Context Graph, Semantics, and Memory

The OpenMetadata graph does not only store data assets. It stores the relationships between assets, columns, owners, teams, policies, quality tests, lineage, classifications, glossary terms, metrics, domains, data contracts, data products, conversations, and memory nuggets.

Example relationships:

…

This graph gives AI systems the relationships, meaning, memory, and governance they need to reason across the data estate.


Memories: Organizational Context for Humans and Agents

Memories preserve the important context that usually disappears inside chats, tickets, meetings, notebooks, and AI agent threads.

A memory is an open, governed OpenMetadata entity that can be tied to data assets, users, teams, threads, domains, data products, metrics, policies, incidents, and workflows. Engineers can capture and retrieve memories through APIs, SDKs, MCP, chat, or AI applications.

Use memories to preserve:

  • why a metric changed
  • why a column was renamed
  • what assumption was used in an analysis
  • which remediation fixed a data quality issue
  • which dashboard or data product a decision applies to
  • what an AI agent learned while investigating an incident
  • what a domain expert explained in a conversation

Memories unlock tribal knowledge by making it reusable, governed, searchable, and available to every human, assistant, and agent that touches your data.


MCP, Semantic Search, APIs, AI SDK, and Memory

OpenMetadata makes context actionable through AI- and developer-friendly interfaces.

MCP Server

OpenMetadata includes an MCP server that lets MCP-compatible assistants and agents interact with the metadata graph through natural language.

AI assistants can use OpenMetadata MCP to:

  • search metadata
  • run semantic search
  • retrieve entity details
  • inspect lineage
  • understand data contracts and policy context
  • retrieve or preserve memory nuggets
  • update descriptions, tags, owners, and other metadata
  • create glossary terms and lineage
  • list and create data quality tests
  • analyze root causes of data quality failures

Get started: OpenMetadata MCP Server Documentation

Semantic Search

Semantic Search lets users and AI assistants find data assets by meaning, not only exact keywords.

Find trusted customer purchase datasets with known data quality issues and recent remediation notes.

OpenMetadata can surface conceptually related assets, metrics, glossary terms, data products, memory nuggets, and governance context even when names differ across domains, tools, and teams.

APIs, SDKs, Events, and Webhooks

OpenMetadata exposes APIs, SDKs, events, and webhooks so teams can ingest, update, search, subscribe to, and automate metadata across their ecosystem.

Developers can use the AI SDK to build custom AI applications that use OpenMetadata context and memory programmatically.


Use It From Code

Two packages, depending on what you're building.

Goal Package Install Read/write metadata, lineage, glossary, quality openmetadata-ingestion pip install "openmetadata-ingestion" Give an LLM or agent governed access (MCP, LangChain) data-ai-sdk pip install data-ai-sdk

Also available: @openmetadata/ai-sdk (TypeScript), org.open-metadata:ai-sdk (Java).

Python SDK — connect and read metadata

Match the SDK version to your server version.

…

Entities are hierarchical — a Table belongs to a Schema, which belongs to a Database, which belongs to a DatabaseService. Every entity references its parent by fullyQualifiedName.

AI SDK — give an agent governed context via MCP

OpenMetadata exposes an MCP server at /mcp. Unlike generic connectors that only read raw database schemas, it exposes semantic search, lineage traversal, glossary/classification, and metadata mutations as tools any LLM can call.

from ai_sdk import AISdk, AISdkConfig

client = AISdk.from_config(AISdkConfig.from_env())

# Convert MCP tools to LangChain format — one line
tools = client.mcp.as_langchain_tools()

# Or call a tool directly
result = client.mcp.call_tool("search_metadata", {"query": "customers"})

Works with LangChain and OpenAI function calling out of the box.

Docs & examples

  • Python SDK reference: https://docs.open-metadata.org/latest/sdk/python
  • MCP server guide: https://docs.open-metadata.org/latest/how-to-guides/mcp
  • AI SDK (MCP tools + agents): https://github.com/open-metadata/ai-sdk
  • REST API: https://docs.open-metadata.org/latest/main-conce

核心特点

  • •what data exists
  • •what it means
  • •who owns it
  • •how it is used
  • •where it came from
  • •where it flows
  • •whether it is fresh, tested, certified, and trusted
  • •which business concepts, classifications, glossary terms, policies, contracts, and data products apply
  • •what downstream dashboards, pipelines, metrics, ML models, and applications depend on it
  • •what conversations, decisions, assumptions, and memory nuggets have already been captured about it

> 标签

TypeScriptcontextcontext-layerdata-catalogdata-collaboration

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月9日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言