Runs Python unit tests with coverage, analyzes coverage reports, and implements meaningful tests to increase coverage by ~0.2%. Use when you want to systematically improve Python test coverage with high-value test cases.
npx skills add https://github.com/streamlit/streamlit --skill improving-python-coverage
Increase Python unit test coverage by ~0.2% through meaningful tests that add real value.
Be fully autonomous — Do NOT stop or pause to ask for confirmation. Keep iterating (analyze → implement → verify) until the 0.2% coverage target is reached. If you encounter ambiguities about what to test, make a reasonable choice and proceed.
Step 1: Run tests with coverage
make python-tests # ~3 min, creates .coverage file
Generate JSON report for analysis:
uv run coverage json -o coverage.json
The JSON contains per-file missing_lines arrays showing uncovered line numbers.
Step 2: Analyze and prioritize
Read coverage.json to find files with:
percent_covered (high impact)lib/streamlit/elements/ or lib/streamlit/runtime/Skip: >97% coverage, proto/*, vendor/*, static/*, test files.
Step 3: Implement tests (in subagent)
Launch a subagent to implement tests for each prioritized file. Provide the subagent with:
missing_lines from coverageThe subagent should:
lib/tests/streamlit/<path>/<module>_test.pylib/tests/AGENTS.md: prefer pytest-style standalone functions over unittest.TestCase classes, use @pytest.mark.parametrize to consolidate tests that only differ in inputs/expected outputs, add numpydoc docstrings and type annotationsuv run pytest lib/tests/streamlit/path/to/module_test.py -vStep 4: Verify and iterate
uv run pytest lib/tests/streamlit/path/to/module_test.py -v # Run new tests
make python-tests # Measure progress
Repeat steps 2-4 until coverage improves by ≥0.2%, then run make check.
Step 5: Simplify, review, and address feedback
Once all tests pass and coverage target is met:
simplifying-local-changes subagent to clean up and simplify the code changes. Wait for completion.reviewing-local-changes subagent to review the changes. Wait for completion and read the review output.DO test: Conditional logic, error handling, edge cases (None, empty, zero, max), public API functions, complex branches.
DON'T test: Simple accessors, protobufs, implementation details, already well-covered code.
Coverage exclusions: Use # pragma: no cover sparingly for code that genuinely doesn't need testing. Always include a reason (e.g., # pragma: no cover - defensive):
# pragma: no cover - platform-specific)# pragma: no cover - defensive)# pragma: no cover - abstract)Integration dependencies: Packages listed under [dependency-groups] integration in pyproject.toml (e.g., pydantic, sympy, polars, sqlalchemy) are only installed for integration tests, not regular unit tests. When writing tests that use these packages:
@pytest.mark.require_integration marker to the testlib/tests/streamlit/<package>/<module>_test.py mirrors lib/streamlit/<package>/<module>.py
lib/tests/AGENTS.md/checking-changes after implementing testspytest.mark.require_integration. These integration tests are not included in the coverage numbers from make python-tests. When analyzing missing lines, check whether the uncovered code is exercised by integration tests before adding unit tests or # pragma: no cover annotations.if sys.version_info >= (3, 14)) run only on matching CI jobs@pytest.mark.require_integration) run in separate CI jobs with those packages installedBefore adding tests or # pragma: no cover for such code, verify whether it's already exercised in CI.
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Use when implementing any feature or bugfix, before writing implementation code
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes
Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always
Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
Take streamlit/improving-python-coverage from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.