microsoft/android-tester
Execute Android UI workflows on a sandbox device, review results, and produce a structured execution report.
npx skills add https://github.com/microsoft/Sico --skill android-tester
Execute an Android UI test case, review the results, and generate a structured execution report.
android-tester against the sandbox device for each test casepython >= 3.11uvFrom the skill root directory ($SKILL_ROOT), run:
uv sync
sh scripts/install-adb.sh
You MUST use --precondition whenever a test case has a precondition. Never embed precondition text in --instructions.
--precondition is repeatable, once per atomic precondition, each in label: description form (label = text before the first : ; short, lowercase, hyphenated, no colon). Decompose a test case's preconditions block into atomic items:
<output-dir>/preconditions/ as sub-dirs named after their labels. Reuse existing labels whenever possible for better script caching and faster execution.--precondition "label: description" argument when invoking android-tester.Preconditions are established in the order given. On a label's first use the tool establishes the state and records a reusable script under <output-dir>/preconditions/<label>/ (action_log.json plus a description.txt); on repeat it replays the script without LLM calls. Cases without a precondition run with just a device reset.
The device is reset automatically before each test case. When running multiple test cases, execute them sequentially without pausing for confirmation between cases. The execution is expected to be uninterrupted. If any concerns or issues arise during individual runs, include them in the final report rather than stopping to ask.
These argument values are literal CLI strings. Do not write instructions, task_name, precondition, or keep_app_state to workspace files and do not pass them as {workspace_dir}/... paths; pass their parameter values directly.
--instructions: the test steps to execute (precondition text must NOT be included here)--device-id: address of the allocated Android device (e.g. 10.0.0.5:5555)--task-name: short label for the test case (max 5–10 words, e.g. "Sign in MSA" not the full test case title).--precondition "label: description": one atomic precondition in label: description form. Repeatable — supply it once per atomic precondition (see Preconditions above).--resources-path: directory holding files the test case needs staged on the device (photos, documents, etc.). Pass it whenever a precondition requires files to be present in /sdcard folders such as Pictures, Documents, etc.android-tester \
--device-id 10.0.0.5:5555 \
--precondition "home-screen: the Android device shows the home screen" \
--precondition "signed-out: User is signed out of Copilot" \
--instructions "Launch Copilot and sign in with MSA account" \
--task-name "Sign in MSA"
android-tester enforces its own end-to-end execution timeout of 3600 seconds (1 hour) per test case. When invoking it through run_command (or any other shell-execution tool), do not pass a timeout lower than this value — use timeout=0 (no external timeout) or a value of at least 3600. When running multiple cases sequentially, use timeout=0 since the wall-clock time scales with the number of cases.
android-tester wipes the app data of open apps before and after each test case to ensure a predictable initial state. If the user explicitly requests to not wipe certain apps, pass --keep-app-state with a comma-separated list of fully qualified Android package names (e.g. com.android.chromium,com.android.settings).
Do not guess the package name. Resolve it on the actual device first, e.g., using
# List all packages whose name contains 'chrome' (case-insensitive)
adb -s <device-id> shell pm list packages | grep -i chrome
# List only user-installed (third-party) packages
adb -s <device-id> shell pm list packages -3
Pick the exact package:<name> line(s) and pass the <name> portion to --keep-app-state. If the user requests to preserve an app but you cannot find its package name, include that in the final report as a blocker and proceed without preservation.
android-tester outputs one JSON line per event to stdout. Each line contains an event field, a timestamp, and event-specific data.
All per-step events go to stdout. The output directory path is logged to stdout at startup.
0: test case executed successfully1: test case failed2: test case blockedREADME.md contains full documentation on android-tester arguments, environment variables, and output event types. Its lecture is optional.
After all test cases have been executed, review the output of each test case and produce a single Regression Testing Delivery Report named Regression Testing Report.md in your workspace directory. Base the report on the stdout JSON logs and report files from each test case run. If no report file is generated for a test case, leave the "execution log" field for that test case blank. All links must be valid http(s) links.
If multiple test cases were executed, summarize all of them in this single report. Do not output the *.html, *.jsonl and *.stderr files directly, unless explicitly requested by the user.
The report must strictly follow this template:
# Regression Testing Delivery Report
## 1. Test Details
| # | Field | Content |
|---|------------------|----------------------------------------------------|
| 1 | Report Date | [YYYY.MM.DD] |
| 2 | Test Pass Status | Completed / Blocked |
| 3 | Testing Minutes | [X] mins |
| 4 | Test Environment | - Sandbox/device Name 1<br>- Sandbox/device Name 2 |
## 2. Execution Summary
| # | Metric | Value |
|---|------------------------|-------------------------------|
| 1 | Total Cases | [Total] |
| 2 | Executed | [Executed] |
| 3 | Passed | [Passed] |
| 4 | Failed | [Failed] |
| 5 | Blocked | [Blocked] |
| 6 | Execution Rate | [XX%] |
| 7 | Pass Rate | [XX%] |
| 8 | Release Recommendation | [Go / Conditional Go / No-Go] |
## 3. Test Scope
| # | Module / Feature | Total Cases | Executed | Passed | Failed | Blocked |
|---|----------------------|-------------|--------------------|-------------------|----------|-----------|
| 1 | [Module / feature 1] | [Total] | [Executed] ([XX%]) | [Passed] ([XX%]) | [Failed] | [Blocked] |
| 2 | ... | | | | | |
## 4. Execution Analysis
### 4.1 Bug Analysis
- **Total Unique Bugs**: [X]
- **Total Failed Cases**: [X]
| # | Bug Title | Description | Impacted Cases |
|---|---------------|----------------------------------------------------------------------------|-------------------------------------------------------|
| 1 | [Short Title] | Clear, concise explanation of the bug (what is wrong vs expected behavior) | [N] cases:<br>- [Case Name 1]<br>- [Case Name 2] |
| 2 | ... | | |
### 4.2 Block Analysis
- **Total Unique Blockers**: [X]
- **Total Blocked Cases**: [X]
| # | Blocker Type | Description | Impacted Cases |
|---|---------------|---------------------------------------------------------------|-------------------------------------------------------|
| 1 | [Short Title] | Clear, concise explanation of the blocker (the environment) | [N] cases:<br>- [Case Name 1]<br>- [Case Name 2] |
| 2 | ... | | |
## 5. Release Recommendation
### Recommendation
[Go / Conditional Go / No-Go]
### Rationale
- Reason 1
- Reason 2
## 6. Detailed Case Result Reference
| # | Case ID | Case Title | Module | Case Step | Execution Time | Result | Execution Log |
|---|----------|-------------|----------|-----------|----------------|-----------------------|-----------------------------------------------------------------|
| 1 | [TC-001] | [Case Name] | [Module] | [X] | X mins | Pass / Fail / Blocked | [TC-001 Execution Log](https://example.org/link/to/report.html) |
| 2 | ... | | | | | | |
After producing the report, give the user a short summary of the testing outcome(s) and your recommendation for release. The summary should be concise and highlight the key points from the report, such as overall pass/fail rates, major blockers, and your final recommendation on whether to proceed with the release. If you include any files in the summary, make sure to provide http(s) links with readable short names, e.g., My File.
CRITICAL: All links must be valid. Only take links from information sources available to you (e.g., plan, deliverables, reports, logs, etc.). Avoid links to workspace-local files because the user cannot access them. If the user requests workspace-local files, use the report tool to upload them first and provide the corresponding http(s) link.
Take microsoft/android-tester from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.