Profiles Meta Quest and Horizon OS application CPU performance using simpleperf — workload classification, CPU hotspot recording, kernel overhead measurement. Use when diagnosing whether an app is CPU-bound, memory-bound, or I/O-bound on Quest devices.
npx skills add https://github.com/meta-quest/agentic-tools --skill hz-simpleperf-debug
Use this skill when you need hardware-level CPU performance insights on Meta Quest devices:
This skill complements hz-perfetto-debug. Perfetto shows *what* your app is doing over time. Simpleperf shows *where* the CPU is spending hardware cycles — cache misses, branch mispredictions, and instruction throughput that Perfetto can't see.
Quest devices run on mobile ARM SoCs with strict thermal and power budgets. CPU-bound apps hit frame drops when:
| Refresh Rate | CPU Frame Budget | Notes |
|-------------|-----------------|-------|
| 120 Hz | 8.3 ms | Tight — simpleperf critical for finding hotspots |
| 90 Hz | 11.1 ms | Default target for most apps |
| 72 Hz | 13.9 ms | Fallback for heavier apps |
Simpleperf's hardware counters reveal bottlenecks invisible to software tracing.
Simpleperf profiling is powered by the metavr CLI. Invoke via npx — no install required:
npx -y metavr --version
Examples below use the bare metavr command for brevity. If metavr is not on PATH, invoke the same CLI via npx -y metavr <args> (the CLI is published under the npm package metavr). Connect your Quest via USB with developer mode enabled.
Before optimizing, determine the bottleneck type:
# Classify the foreground app's workload (10-second sample)
metavr perf simpleperf classify
# Target a specific app
metavr perf simpleperf classify --app com.example.myapp
# Custom duration
metavr perf simpleperf classify --duration 15
Returns a classification with evidence:
| Classification | Indicator | Optimization Strategy |
|---------------|-----------|----------------------|
| CPU-bound | High IPC, low stall ratio | Optimize algorithms, reduce draw calls, batch work |
| Memory-bound | High stall ratio (stalled-cycles-backend / cpu-cycles) | Reduce cache misses, improve data locality, shrink working set |
| I/O-bound | High context switches per second | Reduce blocking I/O, use async, minimize thread contention |
Capture a CPU cycle profile to find the most expensive functions:
# Record CPU hotspots for the foreground app
metavr perf simpleperf record
# Custom frequency and duration
metavr perf simpleperf record --frequency 4000 --duration 10
# Target a specific app
metavr perf simpleperf record --app com.example.myapp
The recording samples CPU cycles at the specified frequency (default 4000 Hz) and generates a profile showing which functions consume the most CPU time.
Determine how much CPU time is spent in kernel vs userspace per thread:
# Measure kernel overhead for the foreground app
metavr perf simpleperf kernel-overhead
# Custom duration
metavr perf simpleperf kernel-overhead --app com.example.myapp --duration 10
Returns per-thread breakdown of user-mode vs kernel-mode CPU cycles. High kernel overhead (>20%) in a thread suggests:
Always start with classification. This prevents wasting time optimizing the wrong thing.
metavr perf simpleperf classify --app com.example.myapp --duration 10
Decision tree based on results:
metavr perf simpleperf record --app com.example.myapp --duration 10
Review the top functions by CPU cycle consumption. Common VR hotspots:
| Function Pattern | Likely Cause | Fix |
|-----------------|-------------|-----|
| Physics.* / PhysX | Complex physics simulation | Reduce collider count, simplify meshes, increase fixed timestep |
| Render* / Draw* | Too many draw calls | Batch materials, use GPU instancing, reduce unique materials |
| GC_* / gc_alloc | Garbage collection pressure | Pool allocations, avoid per-frame allocations |
| memcpy / memmove | Large data copies | Use references, reduce buffer sizes, avoid unnecessary copies |
| LZ4_* / compress | Asset decompression | Pre-decompress, use lighter compression, cache results |
If classification shows memory-bound, the issue is likely cache misses or memory bandwidth:
Use Perfetto hz-perfetto-debug to correlate memory-bound regions with specific code paths.
metavr perf simpleperf kernel-overhead --app com.example.myapp
Interpreting results by thread:
| Thread | Expected Kernel % | High Kernel % Indicates |
|--------|------------------|------------------------|
| Main/Game thread | < 5% | Excessive file I/O, logging, or allocations |
| Render thread | 5-15% | Normal (GPU driver overhead). >20% = driver issue |
| Worker threads | < 5% | Thread synchronization overhead |
| Audio thread | < 10% | Normal for audio HAL calls |
Simpleperf tells you *where* cycles go. Perfetto tells you *when* and in what context. Use together:
metavr perf capture) shows when those functions run relative to frame boundariesmetavr perf query to correlate function timing with frame dropsmetavr device info.adb shell simpleperf fails, ensure developer mode is enabled and USB debugging is authorized.For detailed guides on specific topics, see:
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
React Native and Expo best practices for building performant mobile apps. Use when building React Native components, optimizing list performance, implementing animations, or working with native modules. Triggers on tasks involving React Native, Expo, mobile performance, or native platform APIs.
React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.
Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling
Use when starting feature work that needs isolation from current workspace or before executing implementation plans - creates isolated git worktrees with smart directory selection and safety verification
Take meta-quest/hz-simpleperf-debug from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference npx.
Without those the skill loads but fails at the first command.