google/google-cloud-solution-agentic-analytics-spark-knowledge-catalog
>- Discovers requirements and generates guidance to design and deploy a governed, secure agentic-analytics solution for data that's distributed across Google Cloud, other cloud providers, or on-premises. Data that's outside Google Cloud (such as data from Databricks, Snowflake, Salesforce, SAP, or Oracle systems) is accessed through federation mechanisms such as Apache Iceberg, other "zero-copy ETL" methods, or remote query push-down. Use this skill when designing an architecture for efficient analytics across large volumes of structured and unstructured data that's located in multiple systems and environments, including other cloud providers and on-premises.
npx skills add https://github.com/google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog
This skill provides a workflow to design and implement a governed, secure
pipeline for agentic analytics solution across structured and unstructured data
that's distributed across Google Cloud, on-premises systems, and other cloud
providers.
The workflow consists of the following phases:
the cloud workload or use case that the user needs assistance for.
in Phase 1 to generate a detailed solution architecture for the cloud
workload or use case.
solution, generate validation instructions and scripts, and run the
validation.
content and present the solution.
Important notes about the workflow:
you ask the user clarifying questions, DON'T recommend, propose, or outline
any architectural designs, technical decompositions, cloud services, or
component mappings.
specific phase or task in this workflow is already completed or approved
(e.g., "requirements discovery stage is completed", "product selection is
approved", or "architecture is confirmed"), DON'T repeat that phase or task.
Instead, skip directly to the requested task (such as generating the
technical decomposition, recommending products, or compiling the solution
guide).
processes, activities, and use cases) of their workload. Ask the user the
following questions, one question at a time:
(e.g., PDF flavor recipes, invoices) or structured (e.g., historical
sales in Iceberg)?
Blob, Google Cloud Storage, or databases like AlloyDB?
sources within Google Cloud and in external locations (such as other
cloud providers)?
and run forecast models over large-scale distributed data?
operational agents expect to execute in their agentic IDE (VS Code or
Antigravity IDE)?
workload.
The following are examples of questions you can ask to gather non-functional
requirements:
regulatory compliance (e.g., GDPR, HIPAA), or data governance
requirements must the system adhere to?
fault-tolerance, and disaster recovery objectives (RTO/RPO)?
your workload require?
scientists and engineers need?
data egress/transfer cost requirements?
on-premises.
architecture of the current deployment.
products, or tools. The following are examples of questions that you can
ask to get information about the dependencies:
(e.g., identity providers, data curation platforms, CI/CD pipelines, or
active data catalogs)?
delivery lifecycle (e.g., version control, testing, data quality
assurance)? Provide the path to a directory or examples of these
artifacts.
are any ambiguities or contradictions.
If you identify any ambiguities or contradictions in the requirements that
the user has provided (e.g., zero-copy vs copying data to a repository), then
do the following for each ambiguity or contradiction that you identify:
contradicts the zero-copy requirement and also incurs data-transfer
costs).
"do what you think is best" or "you decide"), then provide a clear
suggestion to resolve the ambiguity or contradiction (e.g., suggest
prioritizing zero-copy remote queries), explain your reasoning
(e.g., to eliminate multi-cloud fees and data duplication), and ask
the user to approve your suggestion.
Critical: Until all the ambiguities and contradictions that you identify
are resolved according to the preceding guidance, you must NOT recommend or
generate any architecture design, technical decomposition, or Google Cloud
product recommendations.
or ambiguities from Step 5.
Generate a technical decomposition of the components of the workload.
components.
within the relevant layers.
which represent a standard architectural pattern for agentic analytics
solutions, flowing from user interaction through data context and
governance to core data processing:
environment.
and data warehouse in the cloud.
data processing, and external data stores.
decomposition.
decomposition.
10. After the user approves the technical decomposition, proceed to Phase 2.
Important: Don't proceed to the next phase until the user approves the
generated technical decomposition of the workload.
For each task in this phase, to ensure that the generated content aligns with
the latest and official Google Cloud guidance, ground the generated content by
using the following resources:
https://developers.google.com/knowledge/mcp.md.txt
developerknowledge:search_documentsdeveloperknowledge:get_documentsdeveloperknowledge:answer_queryacross multi-cloud data lakes, structured data warehouses, and
unstructured data stores:
https://docs.cloud.google.com/architecture/agentic-ai-cross-cloud-analytics.md.txt
the workload:
https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/decision-making-guides.md
the workload:
https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/best-practices-guides.md
appropriate Google Cloud products and features, based on the guidance in the
following resources and adjusted suitably based on the approved technical
decomposition:
references/product-selection-guidance.mdhttps://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/decision-making-guides.mdthe recommendations.
https://github.com/mermaid-js/mermaid.
architecture diagram.
relationships between the components, and the task flow or data flow.
to approve the description.
each component in the architecture based on the workload's requirements.
Important:
references/design-recommendations.md.
resources that are listed in
references/knowledge-catalog-documentation.md
following skills:
google-cloud-waf-securitygoogle-cloud-waf-reliabilitygoogle-cloud-waf-cost-optimizationgoogle-cloud-waf-operational-excellencegoogle-cloud-waf-performance-optimizationgoogle-cloud-waf-sustainabilityneeds any changes.
recommendations meet their requirements.
user to deploy the solution.
Important:
setting up the Google Cloud project, enabling billing, enabling the
required APIs, and setting up the required roles and permissions.
deployment guidance that you generate:
A plugin that provides a specialized suite of skills and MCP tools
to let you use your preferred coding agent to architect complex data
pipelines, transform data with dbt, write Spark and BigQuery SQL
notebooks, create and troubleshoot Dataflow pipelines, and
orchestrate end-to-end workflows across the Google Cloud data
ecosystem.
A codelab that provides instructions to use the Data Agent Kit
extension to efficiently analyze a cross-cloud data topology from
within your preferred agentic development environment.
codelab that provides instructions to build a data foundation in
BigQuery, apply rigid metadata tags (Knowledge Catalog Aspects) to
differentiate valid data from noise, and use the Gemini CLI to
locally test if the LLM strictly follows your governance rules.
A tutorial that shows how to establish data context in Knowledge
Catalog.
A guide that explains how to bring information about your unique,
custom data sources into Knowledge Catalog.
user needs any changes.
guidance meets their requirements.
steps to verify that the generated solution meets the workload's
requirements.
curl or gcloud to performthe steps in the approved validation plan.
deployment issues.
into a single Markdown file named solution-architecture-guide.md, based on
the template in assets/output-template.md.
workspace.
workspace.
Guide to how the Data Agent Kit extension lets you use notebooks for data
transformation and analysis.
Knowledge Catalog.
Guide to accelerating Apache Spark workloads by using Lightning Engine.
Guide to use Knowledge Catalog as a governance and agentic layer for
BigQuery.
Take google/google-cloud-solution-agentic-analytics-spark-knowledge-catalog from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.