
How to modernize Apache Hive using Google Cloud’s Lakehouse runtime catalog
Google Cloud provides a serverless Lakehouse runtime catalog to modernize Apache Hive Metastores. It enables users to transition production Hive tables using a zero-data-copy migration solution.
Why it matters
Data platform teams can reduce operational costs and eliminate performance bottlenecks associated with legacy metadata registries. This allows multiple query engines to access the same data without needing to duplicate petabytes of storage.
The details
- Supports both Iceberg REST Catalog and Hive Catalog open APIs.
- Integrates with Cloud IAM for consistent table-level security across compute engines.
- Utilizes Google's Spanner infrastructure to ensure metadata scales with data growth.
Show entities and relationshipsHide entities and relationships
In this article
Key connections
Google owns Lakehouse runtime catalog
Google Cloud introduced the serverless Lakehouse runtime catalog for modern data architectures.
Cloud IAM is Google Cloud's identity and access management service.
Lakehouse runtime catalog is built with Apache Iceberg
The Lakehouse runtime catalog is built on the open Apache Iceberg REST catalog specification.
Lakehouse runtime catalog uses Spanner
The Lakehouse runtime catalog is backed by Google Spanner infrastructure to scale metadata with data.
Lakehouse runtime catalog uses Google Cloud Storage
Lakehouse runtime catalog table definitions point directly to existing data stored in Google Cloud Storage.
Lakehouse runtime catalog is related to Knowledge Catalog
The Lakehouse runtime catalog integrates directly with Knowledge Catalog for metadata governance and trusted agent context.
Show 11 more connectionsShow fewer connections
Lakehouse runtime catalog is related to Cloud IAM
The Lakehouse runtime catalog integrates directly with Cloud IAM to enforce table-level security across compute engines.
Lakehouse runtime catalog competes with Apache Hive Metastore
The serverless Lakehouse runtime catalog modernizes and replaces legacy standalone Apache Hive Metastores.
Managed Service for Apache Spark uses Lakehouse runtime catalog
Google Cloud Managed Service for Apache Spark queries tables registered in the Lakehouse runtime catalog.
BigQuery uses Lakehouse runtime catalog
BigQuery queries tables registered in the Lakehouse runtime catalog via open standard REST interfaces.
Gemini uses Lakehouse runtime catalog
Conversational Analytics agents with Gemini tap into data registered in the Lakehouse runtime catalog.
Trino uses Lakehouse runtime catalog
Trino can query unified datasets registered in the Lakehouse runtime catalog.
Presto uses Apache Hive Metastore
Presto uses Apache Hive Metastore as a central schema registry to query raw data files.
Apache Hive Metastore uses MySQL
Standalone Apache Hive Metastore deployments rely on relational database backends like MySQL to track table schemas and partitions.
Apache Hive Metastore uses PostgreSQL
Standalone Apache Hive Metastore deployments rely on relational database backends like PostgreSQL to track table schemas and partitions.
Apache Hive Metastore uses Google Compute Engine
Self-managed Apache Hive Metastores are deployed on Google Compute Engine virtual machine instances.
Apache Spark uses Lakehouse runtime catalog
Apache Spark compute jobs query tables registered in the Lakehouse runtime catalog.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.