Data teams increasingly expect search that understands intent, not just keywords. Embedding‑powered data catalogs combine traditional metadata with vector search so users can find tables, columns, ETL jobs, dashboards or SQL snippets by meaning — for example, “monthly churn by cohort” or “customers who upgraded after promo”. This guide walks data and analytics engineers through a practical, production-ready build: architecture, model selection, ingestion patterns, indexing, hybrid search, evaluation, operations and governance (June 2026 context).

Why embeddings for a data catalog?

Traditional catalogs rely on exact or fuzzy keyword search across table/column names, descriptions and tags. That's fine for well-documented assets, but real-world metadata is noisy: terse column names, missing descriptions, or domain words that differ from user queries. Embeddings map text (names, descriptions, samples, SQL) into vectors so semantically similar items appear near each other. Use cases that benefit immediately:

  • Natural language discovery: business users ask “how do I find active subscriptions?” and get relevant tables and dashboards.
  • Cross-asset linking: surface dashboards and ETL jobs semantically related to a table even without explicit links.
  • SQL snippet search: find SQL queries and joins by intent, not exact column names.
  • Assist analysts by suggesting joins, related dimensions, or common aggregations.

High-level architecture

At a high level the system has these parts:

  1. Metadata source(s): catalog metadata (OpenMetadata, Hive Metastore, Glue), BI catalog, Git repos with SQL, notebooks, job schedulers, and lineage (OpenLineage).
  2. Preprocessor/featurizer: assembles documents per asset (table, column, query) and applies field weighting, chunking, and normalization.
  3. Embedding model: converts text into numeric vectors (on‑prem or hosted).
  4. Vector store: Weaviate, Milvus, Qdrant, or managed services (Pinecone, Zilliz Cloud) that support ANN queries and metadata filters.
  5. Search layer and UI: combines keyword search (BM25) + vector re-ranking, applies filters/ACLs, and serves results to users.
  6. Orchestration & monitoring: incremental ingestion, freshness metrics, and observability.

Step 1 — Define scope and assets to embed

Decide what you will embed first. Start small and add asset types. Typical starting set:

  • Table-level documents: schema name, table name, description, owner, tags, lineage parents, most common query examples.
  • Column-level descriptors: column name, datatype, sample values, compute expressions.
  • SQL/Notebook snippets: full query text plus a short extracted description.
  • Dashboard panels and metrics: metric name, description, underlying table references and last modified.

For a first rollout, aim for 5k–50k documents — enough for meaningful results without high operational complexity.

Step 2 — Choose an embedding model (2026 guidance)

Choices in 2026 fall into two camps: managed high-performance embeddings (OpenAI, Cohere, Anthropic embeddings) and open-source models you can host (Meta/Llama-based embedding models, Mistral, or smaller specialized encoders like SentenceTransformers variants). Consider:

  • Privacy/regulatory: if metadata contains PII or proprietary SQL, prefer on‑prem/self‑hosted models.
  • Latency and cost: hosted services reduce ops but incur per‑request cost; open models require GPUs for low latency.
  • Dimensionality & memory: common dims today are 1,024–3,072. Larger dims improve nuance but increase storage.

Practical recommendation: start with a 1,536‑dim model (e.g., OpenAI text-embedding-3-small/large equivalents or a 1,536-dim SentenceTransformers local model), evaluate quality, then scale to higher dims if needed.

Step 3 — Prepare documents and embedding input

How you construct the text passed to the embedder matters. Create a compact, structured document for each asset; include fields and weights:

  • Title (weight high): "table_schema.table_name" or metric display name.
  • Short description (weight high): owner-written description if present.
  • Representative SQL or sample rows (weight medium): include sanitized sample values — avoid raw PII.
  • Tags and lineage snippet (weight low): parent datasets, source systems.

Example concatenated input (pseudo):

<TITLE> orders.orders_fct
<DESC> Fact table of customer orders aggregated nightly — includes status, order_total
<SQL> SELECT order_id, customer_id, SUM(price) AS order_total FROM raw.orders WHERE status > 'cancelled' GROUP BY order_id
<SAMPLES> [1001, 2002, 3003] ...

Chunk long notebooks/queries: split by function or SQL blocks and index each chunk with a link back to the parent asset. Keep chunks ≤ 1,000 tokens for reliable embeddings.

Step 4 — Select and configure a vector store

Vector stores vary primarily by features (metadata filters, hybrid search), scaling model (disk-backed vs RAM), and ops model. Common picks in 2026:

  • Weaviate — built-in schema, hybrid search, OpenSearch/Elasticsearch connectors.
  • Milvus — high performance, multiple index types, good for very large collections.
  • Qdrant — lightweight, Rust-based, easy to operate with payload filters.
  • Pinecone/Zilliz Cloud — managed services with easy scaling and SLAs.

ANN tuning (HNSW example): M (connectivity) = 16–64, efConstruction = 100–200, efSearch = 50–200. Higher efSearch increases recall at query time but costs latency. For catalog search prioritize recall > 95% for top‑k results.

Estimate storage: a 1,536-dim embedding uses 1,536 * 4 bytes ≈ 6 KB per vector. 100k vectors ≈ 600 MB + index metadata and payload. Plan for overhead (30–60%).

Step 5 — Hybrid search: combine BM25 and embeddings

Pure semantic search can surface assets that are semantically close but irrelevant by context. A pragmatic approach is hybrid search:

  1. Run a fast lexical search (Elasticsearch/Opensearch) to get a candidate set (top 100).
  2. Compute embedding for the query and ANN-search within the candidate set, or re-rank the lexical results by cosine similarity.
  3. Apply metadata filters and ACLs, and present fused results.

This keeps high precision for well-documented assets while improving recall for poorly-documented ones. Many vector stores support hybrid out-of-the-box; if not, implement re-ranking in the application layer.

Step 6 — Incremental updates and versioning

Recompute embeddings only for changed assets. Use a change feed from your metadata system (OpenMetadata, Glue/Atlas hooks, or periodic diffs) to detect updates. Recommended pattern:

  • Store a hash/ETag of the source fields used to build the embedding.
  • On change detection, run preprocessing and upsert the vector into the store.
  • For large-scale schema changes, schedule a controlled re-embed with resource throttling.

Keep a mapping table (asset_id → vector_id, embedding_version, model_id) so you can trace which model/version produced a vector and roll back if necessary.

Step 7 — Privacy, governance and access control

Metadata often references internal identifiers and sampled values. Treat it as sensitive:

  • PII: redact or hash raw sample values before embedding. Do not embed raw email addresses or SSNs.
  • Encryption: enable at-rest encryption for vector store files and enforce TLS for network traffic.
  • RBAC: the search layer must enforce catalog ACLs so vector results are filtered by user permissions. Many vector stores support payload filters; implement row-level filtering where needed.
  • Audit logs: keep search logs (queries, top results) for compliance and UX tuning.

Step 8 — Evaluation and tuning

Define success metrics and run A/B tests:

  • Quantitative: MAP@k, NDCG, CTR on clicked assets, time-to-first-value for analysts.
  • Qualitative: analyst feedback sessions, relevance judgments for a labeled query set.

Iterate on document construction (field weights), embedding model, and ANN parameters. For most catalogs the biggest wins come from improving the preprocessing and including representative SQL samples.

Step 9 — Observability and operational runbook

Track these signals:

  • Search latency (end-to-end), embedding request latency, vector store query latency.
  • Freshness: age of the vector relative to the latest metadata update.
  • Hit rate: fraction of queries returning useful results (via user feedback or clickthrough).
  • Error rates and failed upserts during reindexing.

Operational runbook should cover: model rollout (canary embedding version), reindexing procedure (batch windows, throttling), and incident playbook for a vector store outage (fallback to lexical search).

Sample rollout plan (8–12 weeks)

  1. Week 1–2: Define scope and collect labeled queries from users. Choose model and vector store PoC.
  2. Week 3–4: Implement preprocessing and a PoC pipeline for 5k assets. Build simple UI and hybrid search flow.
  3. Week 5–6: Evaluate quality, tune HNSW/BM25, implement incremental updater from metadata source.
  4. Week 7–8: Harden security, ACL enforcement, observability. Run internal beta with power users.
  5. Week 9–12: Roll out wider, monitor metrics and iterate on UI, suggestions and reranking.

Pitfalls to avoid

  • Embedding raw PII or confidential snippets — always sanitize.
  • Indexing everything as one giant document — chunk and maintain parent links.
  • Assuming higher embedding dimensions always win — tradeoffs on cost and latency.
  • Not enforcing ACLs at the vector layer — metadata leaks can occur through search results.

Conclusion and next steps

An embedding-powered catalog dramatically improves discoverability when carefully designed: thoughtful document construction, hybrid search, privacy-safe preprocessing and incremental updates are critical. Start small, evaluate with real queries and users, and use metrics to guide expansion into more asset types (dashboards, metrics, data products). Future extensions include using embeddings for automated lineage suggestions, similarity-based data quality checks, and LLM-assisted answers that use retrieved catalog contexts for safe, grounded explanations.

For practical follow-ups, test two embedding models (one hosted, one self-hosted), pilot with 5–10 power users for two weeks, and iterate on the preprocessor and hybrid ranking — that cycle produces the biggest improvement in perceived relevance.