google/agent-platform-alert-configuration
>- Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. and work across runtimes (e.g., Cloud Run, Vertex AI). Quality alerts rely on Vertex AI Online Monitors and are strictly bound to Vertex AI deployments.
npx skills add https://github.com/google/skills --skill agent-platform-alert-configuration
Before executing any commands or writing configurations on behalf of the user,
you MUST adhere to the following safety tiers based on the action requested:
check_telemetry.py / gather_agent_info.py)immediately to inspect telemetry status or gather agent configuration
details.
create_online_monitor.py /provisioning)**
additional billing charges and create cloud resources. The agent MUST
ALWAYS warn the user explicitly about the potential extra billing costs
of BOTH the Online Monitor (specifically mentioning LLM evaluations)
and Telemetry (specifically mentioning Cloud Trace/Logging export).
You MUST STOP and ask for explicit approval before proceeding with
provisioning or providing setup commands.
function, the underlying agent MUST be instrumented to emit OpenTelemetry
(OTel) metrics. If the agent does not emit these metrics, the alerting
policies will have no data stream to evaluate.
Before executing any python script in this skill you MUST install the required
dependencies in your environment. Run this command first:
pip install -r scripts/requirements.txt
telemetry, or interact with the Google Cloud Project(s) explicitly provided
by the user in the prompt. Do NOT assume or use other projects from your
environment or history unless the user explicitly directs you to do so.
file and then modify it, you MUST perform these actions sequentially (copy
first, then modify) rather than writing the final content directly.
generating or writing ANY configuration, you MUST execute these steps in
order:
gather_agent_info.py to automatically identify agent runtime, check
telemetry, metric scopes, linked datasets, and more. This script covers
most of the manual checks listed in subsequent steps.
{project_id} --agent-name {agent_name}`
doesn't produce everything you need, you MUST satisfy requirements
by running the manual fallback steps listed in Step 2 and then
perform Step 3 below. If Step 1 succeeds and provides all info,
SKIP to Step 3 (Pre-existing Policies Check).
failed to determine the metric scope.
projects/{project_id}`. If a scoping project is returned, you MUST
deploy policies there.
google_monitoring_monitored_project resources to extract the
scoping project.
a multi-project Cloud Monitoring Metric Scope? If so, what is the
scoping project ID?"
already exist targeting the same metrics (grouped by
reasoning_engine_id or gen_ai_agent_name). Use
scan_duplicates.py to verify.
references/ with names ending in _alert_policies.md to learn how to
configure alert policies based on type. By default you should configure all
of the following alert types UNLESS the user requests to generate explicit
alert policies and/or types. Follow their tables of content to help you find
the reference sections you need to read:
Alert Type | Reference File
:-------------- | :-------------
Reliability | reliability_alert_policies.md
Quality | quality_alert_policies.md
Cost | cost_alert_policies.md
Safety | safety_alert_policies.md
Security | security_alert_policies.md
policies:
policies (Requires Vertex AI Online Monitors):
policy:
alerting policy:
Analytics Alerting)
alerting policy:
Alerting)
Terraform (.tf) files (e.g., alerts.tf, variables.tf).
alerts AND there is no valid Terraform install. SQL-based alerting using
condition_sql requires the provider version >= 6.0.0 (or late 5.x
versions supporting the feature).
terraform.
NOT hardcode specific agent IDs or resource name filters (e.g.,
{gen_ai_agent_name="{agent_name}"} or
metric.labels.agent_resource_name="{agent_name}") in alerting conditions
unless explicitly requested (e.g., "ONLY for this agent"). Merely mentioning
a specific agent name or ID in the request does NOT constitute an explicit
request to pin/filter; you MUST still default to dynamic grouping to cover
all agents. To cover all active agents in the project dynamically:
aggregations. Group by gen_ai_agent_name (e.g., `by
(gen_ai_agent_name)`). Avoid filtering to a single ID/Name unless
requested.
agent_resource_name filter entirely. Configure the condition filter to
only target the monitored resource type
(aiplatform.googleapis.com/OnlineEvaluator) and metric type
(aiplatform.googleapis.com/online_evaluator/scores) globally for the
project.
ENDS_WITH filtertargeting a specific agent name. Instead, extract the agent identifier
(e.g., JSON_VALUE(resource.attributes, '$."cloud.resource_id"')) and
add it to the GROUP BY clause alongside the model or tool name.
any). Otherwise, deploy configuration files to target Terraform or SRE
folders (e.g. monitoring/, ops/, sre/). Use tools to locate where
alert policies or state pointers exist in the project, rather than blindly
writing to the root.
channels without user input. If the user explicitly provides a notification
channel in their prompt, configure the alerts to use it. If no notification
channel is provided, you MUST explicitly ask the user in your final response
if they would like to configure notification channels. **This is a mandatory
question and you MUST NOT omit it from your response. IMPORTANT** Do NOT
make assumptions about notification channels. If you search the codebase for
a notification channel you must ALWAYS confirm with the user before using
it.
what the alerts do in your response. This must explain in plain English what
the alert measures, how the algorithm works, and what a trigger indicates.
tasks that you spawn. Before completing your execution and returning your
final response, you MUST terminate or kill any active or hanging background
tasks (using the manage_task tool with action kill).
the output files are written with the correct grammar and structure. See
details about the tool in the Tooling Scripts section below.
Use the following scripts to discover agents, gather configuration details,
resolve duplicates, and validate configs:
(Metric Scopes, BQ Datasets, Notification Channels), table derivations (Log
& Trace), and Online Evaluator checks.
--agent-name {agent_name}`
folder to ensure changes are merged in-place rather than appended:
--engine-var '${var.gen_ai_agent_name}'`
HCL structure:
python3 scripts/lint_syntax.py {path_to_tf_file}errors), you MUST read the command output, locate the line/file
containing the lint error, analyze the PromQL syntax or Terraform HCL
issue, apply adjustments in-place, and re-run the lint_syntax.py
validation. Repeat this loop until the validation script passes
successfully.
request count boundaries do not scale under changing traffic throughput.
Recommend ratio-based error rate alerts instead.
metric threshold policy end-to-end, do NOT attempt to force real platform
errors. Instead, deploy the alert policy with standard safe bounds (Z-score
multiplier > 15), then temporarily update standard deviation Z-score limits
to a negative value (e.g. > -3) to trigger/verify the "Firing" state before
reverting. Always get confirmation before taking this action proactively.
scan_duplicates.py exiting with code 1: Parse the JSONoutput for duplicate resource targets. Perform in-place upgrade edits,
then re-check until it passes with 0.
utility scripts (such as gather_agent_info.py, check_telemetry.py,
create_online_monitor.py, analyze_traffic.py,
list_log_scope_table_names.py, or list_trace_scope_table_names.py)
fails unexpectedly, you MUST read and inspect the stdout/stderr logs or
error output. Analyze the error message and attempt to dynamically
correct parameters and retry execution before escalating or
falling back to manual plans. Consult the relevant domain-specific
reference file for detailed troubleshooting steps for specific scripts.
ALIGN_MEAN cannot beapplied to DELTA distribution metrics like online_evaluator/scores. You
MUST use percentile-based aligners (like ALIGN_PERCENTILE_50) to reduce
the score distribution into a comparable numeric stream.
PromQL or SQL queries (which are defined as strings), you MUST use the
${var.variable_name} syntax. Bare references like var.variable_name will
fail at deployment time.
or search commands (such as ls -R, find ., or raw recursive grep) from
the repository root if it contains a very large number of files, as this
will freeze your session. Always target specific subdirectories.
Take google/agent-platform-alert-configuration from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.