Service health monitoring, endpoint validation, and CVE source auditing.
npx skills add https://github.com/notque/vexjoy-agent --skill service-health-check
This skill provides deterministic service health monitoring using the Discover-Check-Report pattern. It finds services, gathers health signals from multiple sources (process table, health files, port binding), and produces actionable reports identifying degraded or failed services.
Core principle: Health assessment is evidence-based. Never report a service healthy without verifying process status independently of health file content. Never assume a running process is functional — always cross-check against health files and port binding.
| Signal | Load These Files | Why |
|---|---|---|
| Endpoint validation request | references/endpoint-validator.md | Full endpoint validation methodology |
| Security header WARNs, HSTS/CSP/X-Frame issues | references/security-headers.md | Deep security header reference |
| Config errors, hardcoded IPs, timeout problems | references/endpoint-config-preferred-patterns.md | Endpoint config patterns |
| 401/403 failures, Bearer/API-key/cookie auth | references/auth-endpoint-patterns.md | Auth endpoint patterns |
| CVE source audit request | references/cve-source-check.md | Full CVE source check methodology |
| CVE registry schema questions | references/registry-schema.md | Registry shape and entry format |
| CVE source URL verification | references/source-verification.md | HEAD-check semantics |
| CVE report format questions | references/output-formats.md | JSON schema and Markdown sections |
Goal: Identify all services to check before running any health probes.
Step 1: Locate service definitions
Search for service configuration in this order:
services.json in project rootStep 2: Build service manifest
For each service, establish:
## Service Manifest
| Service | Process Pattern | Health File | Port | Stale Threshold |
|---------|----------------|-------------|------|-----------------|
| api-server | gunicorn.*app:app | /tmp/api_health.json | 8000 | 300s |
| worker | celery.*worker | /tmp/worker_health.json | - | 300s |
| cache | redis-server | - | 6379 | - |
Validation constraints:
Step 3: Validate manifest
Confirm each entry passes the constraints above. If a pattern is too broad, use ps aux | grep to identify distinguishing arguments, then update the pattern.
Gate: Service manifest complete with at least one service. Proceed only when gate passes.
Goal: Gather health signals for every service in the manifest. Always check process status independently of health file content—a running process and a healthy health file are separate signals.
Step 1: Check process status
For each service, run process check:
pgrep -f "<process_pattern>"
Record: running (true/false), PIDs, process count.
Rationale: Process existence is the primary signal. A missing process always means the service is DOWN. A running process alone is insufficient—the service may have crashed or failed to bind to its port.
Step 2: Parse health files (if configured)
Read and parse JSON health files. Evaluate:
Critical constraint: Never trust health file content alone. The file could be stale from before a process crash. Always verify:
Step 3: Probe ports (if configured)
Check if expected ports are listening:
ss -tlnp "sport = :<port>"
Rationale: Verify ports are actually bound. A process can start but fail to bind to its configured port—that is effectively a DOWN state, not HEALTHY.
Step 4: Evaluate health per service
Apply this decision tree (constraints embedded in logic):
Gate: All services evaluated with evidence-based status. No status is determined without concrete signal (process check, health file, or port probe). Proceed only when gate passes.
Goal: Produce structured, actionable health report with specific remediation commands.
Step 1: Generate summary
SERVICE HEALTH REPORT
=====================
Checked: N services
Healthy: X/N
RESULTS:
service-name [OK ] HEALTHY PID 12345, uptime 2d 4h
background-worker [WARN] WARNING Health file stale (15 min)
cache-service [DOWN] DOWN Process not found
RECOMMENDATIONS:
background-worker: Restart recommended - health file not updated in 900s
cache-service: Start service - process not running
SUGGESTED ACTIONS:
systemctl restart background-worker
systemctl start cache-service
Step 2: Set exit status
Step 3: Present to user
Gate: Report delivered with actionable recommendations for all non-healthy services.
User says: "Are all services up?"
Actions:
Result: Clean report, no action needed
User says: "The background worker seems stuck"
Actions:
Result: Specific diagnosis with actionable command
Cause: No services.json, docker-compose, or systemd units discovered
Solution:
Cause: Pattern too broad (e.g., "python" matches all Python processes)
Solution:
ps aux | grep to identify distinguishing argumentsCause: Malformed JSON, permissions issue, or file being written during read
Solution:
ls -laServices should write health files as:
{
"timestamp": "ISO8601, updated every 30-60s",
"status": "healthy|degraded|error",
"connection": "connected|disconnected|reconnecting",
"last_activity": "ISO8601 of last meaningful action",
"running": true,
"uptime_seconds": 12345,
"metrics": {}
}
| Constraint | Rationale | Application |
|-----------|-----------|-------------|
| Process status verified independently of health file | Running process ≠ functional service | Always check process before trusting health file |
| Health file staleness detected by timestamp freshness | File could be stale from before crash | Check timestamp against 300s (configurable) threshold |
| Port binding verified when configured | Process running doesn't mean port is bound | Always verify expected port listening when port specified |
| No auto-restart without explicit flag | Restart masks root cause | Report findings first; only execute restart if user flags it |
| Narrow process patterns required | "python" matches all processes, giving false matches | Use full paths or specific args; validate with ps aux \| grep |
| Evidence-based status only | Status must have supporting signal | No status without concrete evidence (process, health file, or port) |
This skill should be used when the user asks to "perform cloud penetration testing", "assess Azure or AWS or GCP security", "enumerate cloud resources", "exploit cloud misconfigurations", "test O365 security", "extract secrets from cloud environments", or "audit cloud infrastructure". It provides comprehensive techniques for security assessment across major cloud platforms.
Comprehensive Flow Nexus platform management - authentication, sandboxes, app deployment, payments, and challenges
Coordinate multi-layer security scanning and hardening across application, infrastructure, and compliance controls.
Implement Kubernetes security policies including NetworkPolicy, PodSecurityPolicy, and RBAC for production-grade security. Use when securing Kubernetes clusters, implementing network isolation, or enforcing pod security standards.
Implement Kubernetes security policies including NetworkPolicy, PodSecurityPolicy, and RBAC for production-grade security. Use when securing Kubernetes clusters, implementing network isolation, or enforcing pod security standards.
Run Azure compliance and security audits with azqr plus Key Vault expiration checks. Covers best-practice assessment, resource review, policy/compliance validation, and security posture checks. WHEN: compliance scan, security audit, BEFORE running azqr (compliance cli tool), Azure best practices, Key Vault expiration check, expired certificates, expiring secrets, orphaned resources, compliance assessment.
Run Azure compliance and security audits with azqr plus Key Vault expiration checks. Covers best-practice assessment, resource review, policy/compliance validation, and security posture checks. WHEN: compliance scan, security audit, BEFORE running azqr (compliance cli tool), Azure best practices, Key Vault expiration check, expired certificates, expiring secrets, orphaned resources, compliance assessment.
Comprehensive AWS security posture assessment using AWS CLI and security best practices
Take notque/service-health-check from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.