>- Discovers requirements and generates guidance to design and deploy a governed, secure agentic-analytics solution for data that's distributed across Google Cloud, other cloud providers, or on-premises. Data that's outside Google Cloud (such as data from Databricks, Snowflake, Salesforce, SAP, or Oracle systems) is accessed through federation mechanisms such as Apache Iceberg, other "zero-copy ETL" methods, or remote query push-down. Use this skill when designing an architecture for efficient analytics across large volumes of structured and unstructured data that's located in multiple systems and environments, including other cloud providers and on-premises.
npx skills add https://github.com/google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog
This skill provides a workflow to design and implement a governed, secure
pipeline for agentic analytics solution across structured and unstructured data
that's distributed across Google Cloud, on-premises systems, and other cloud
providers.
The workflow consists of the following phases:
the cloud workload or use case that the user needs assistance for.
in Phase 1 to generate a detailed solution architecture for the cloud
workload or use case.
solution, generate validation instructions and scripts, and run the
validation.
content and present the solution.
Important notes about the workflow:
you ask the user clarifying questions, DON'T recommend, propose, or outline
any architectural designs, technical decompositions, cloud services, or
component mappings.
specific phase or task in this workflow is already completed or approved
(e.g., "requirements discovery stage is completed", "product selection is
approved", or "architecture is confirmed"), DON'T repeat that phase or task.
Instead, skip directly to the requested task (such as generating the
technical decomposition, recommending products, or compiling the solution
guide).
processes, activities, and use cases) of their workload. Ask the user the
following questions, one question at a time:
(e.g., PDF flavor recipes, invoices) or structured (e.g., historical
sales in Iceberg)?
Blob, Google Cloud Storage, or databases like AlloyDB?
sources within Google Cloud and in external locations (such as other
cloud providers)?
and run forecast models over large-scale distributed data?
operational agents expect to execute in their agentic IDE (VS Code or
Antigravity IDE)?
workload.
The following are examples of questions you can ask to gather non-functional
requirements:
regulatory compliance (e.g., GDPR, HIPAA), or data governance
requirements must the system adhere to?
fault-tolerance, and disaster recovery objectives (RTO/RPO)?
your workload require?
scientists and engineers need?
data egress/transfer cost requirements?
on-premises.
architecture of the current deployment.
products, or tools. The following are examples of questions that you can
ask to get information about the dependencies:
(e.g., identity providers, data curation platforms, CI/CD pipelines, or
active data catalogs)?
delivery lifecycle (e.g., version control, testing, data quality
assurance)? Provide the path to a directory or examples of these
artifacts.
are any ambiguities or contradictions.
If you identify any ambiguities or contradictions in the requirements that
the user has provided (e.g., zero-copy vs copying data to a repository), then
do the following for each ambiguity or contradiction that you identify:
contradicts the zero-copy requirement and also incurs data-transfer
costs).
"do what you think is best" or "you decide"), then provide a clear
suggestion to resolve the ambiguity or contradiction (e.g., suggest
prioritizing zero-copy remote queries), explain your reasoning
(e.g., to eliminate multi-cloud fees and data duplication), and ask
the user to approve your suggestion.
Critical: Until all the ambiguities and contradictions that you identify
are resolved according to the preceding guidance, you must NOT recommend or
generate any architecture design, technical decomposition, or Google Cloud
product recommendations.
or ambiguities from Step 5.
Generate a technical decomposition of the components of the workload.
components.
within the relevant layers.
which represent a standard architectural pattern for agentic analytics
solutions, flowing from user interaction through data context and
governance to core data processing:
environment.
and data warehouse in the cloud.
data processing, and external data stores.
decomposition.
decomposition.
10. After the user approves the technical decomposition, proceed to Phase 2.
Important: Don't proceed to the next phase until the user approves the
generated technical decomposition of the workload.
For each task in this phase, to ensure that the generated content aligns with
the latest and official Google Cloud guidance, ground the generated content by
using the following resources:
https://developers.google.com/knowledge/mcp.md.txt
developerknowledge:search_documentsdeveloperknowledge:get_documentsdeveloperknowledge:answer_queryacross multi-cloud data lakes, structured data warehouses, and
unstructured data stores:
https://docs.cloud.google.com/architecture/agentic-ai-cross-cloud-analytics.md.txt
the workload:
https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/decision-making-guides.md
the workload:
https://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/best-practices-guides.md
appropriate Google Cloud products and features, based on the guidance in the
following resources and adjusted suitably based on the approved technical
decomposition:
references/product-selection-guidance.mdhttps://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/decision-making-guides.mdthe recommendations.
https://github.com/mermaid-js/mermaid.
architecture diagram.
relationships between the components, and the task flow or data flow.
to approve the description.
each component in the architecture based on the workload's requirements.
Important:
references/design-recommendations.md.
resources that are listed in
references/knowledge-catalog-documentation.md
following skills:
google-cloud-waf-securitygoogle-cloud-waf-reliabilitygoogle-cloud-waf-cost-optimizationgoogle-cloud-waf-operational-excellencegoogle-cloud-waf-performance-optimizationgoogle-cloud-waf-sustainabilityneeds any changes.
recommendations meet their requirements.
user to deploy the solution.
Important:
setting up the Google Cloud project, enabling billing, enabling the
required APIs, and setting up the required roles and permissions.
deployment guidance that you generate:
A plugin that provides a specialized suite of skills and MCP tools
to let you use your preferred coding agent to architect complex data
pipelines, transform data with dbt, write Spark and BigQuery SQL
notebooks, create and troubleshoot Dataflow pipelines, and
orchestrate end-to-end workflows across the Google Cloud data
ecosystem.
A codelab that provides instructions to use the Data Agent Kit
extension to efficiently analyze a cross-cloud data topology from
within your preferred agentic development environment.
codelab that provides instructions to build a data foundation in
BigQuery, apply rigid metadata tags (Knowledge Catalog Aspects) to
differentiate valid data from noise, and use the Gemini CLI to
locally test if the LLM strictly follows your governance rules.
A tutorial that shows how to establish data context in Knowledge
Catalog.
A guide that explains how to bring information about your unique,
custom data sources into Knowledge Catalog.
user needs any changes.
guidance meets their requirements.
steps to verify that the generated solution meets the workload's
requirements.
curl or gcloud to performthe steps in the approved validation plan.
deployment issues.
into a single Markdown file named solution-architecture-guide.md, based on
the template in assets/output-template.md.
workspace.
workspace.
Guide to how the Data Agent Kit extension lets you use notebooks for data
transformation and analysis.
Knowledge Catalog.
Guide to accelerating Apache Spark workloads by using Lightning Engine.
Guide to use Knowledge Catalog as a governance and agentic layer for
BigQuery.
Comprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas
Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the xlsx in my downloads\") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.
Picks random winners from lists, spreadsheets, or Google Sheets for giveaways, raffles, and contests. Ensures fair, unbiased selection with transparency.
Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research.
MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing. Use when writing MATLAB/Octave scripts for linear algebra, signal processing, image processing, differential equations, optimization, statistics, or creating scientific visualizations. Also use when the user needs help with MATLAB syntax, functions, or wants to convert between MATLAB and Python code. Scripts can be executed with MATLAB or the open-source GNU Octave interpreter.
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
Creating interactive data visualisations using d3.js. This skill should be used when creating custom charts, graphs, network diagrams, geographic visualisations, or any complex SVG-based data visualisation that requires fine-grained control over visual elements, transitions, or interactions. Use this for bespoke visualisations beyond standard charting libraries, whether in React, Vue, Svelte, vanilla JavaScript, or any other environment.
Access AlphaFold 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.
Take google/google-cloud-solution-agentic-analytics-spark-knowledge-catalog from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.