mcpbeat Sign in

Bigquery Bigframes Agent Skill

>- Generates Python code using BigQuery DataFrames (BigFrames), the pandas/scikit-learn-style API over BigQuery. Use when writing BigFrames code or doing pandas-style dataframe/ML work against BigQuery (e.g. in a notebook). Don't use for SQL-first workflows or the google-cloud-bigquery client library — use bigquery-basics.

2k tokens
context cost
the whole folder, loaded on every use
3
files
instructions only
0
copies elsewhere
how many repositories repackaged it
15506
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/google/skills --skill bigquery-bigframes

What comes with it

2 468 bytes besides the instruction
references/linear_regression.md
references/logistic_regression.md

The instruction itself

5 sections, as written by the author

BigFrames (BigQuery DataFrame) basics

BigFrames is a Python library that lets you take advantage of BigQuery

data processing by using familiar Python APIs.

Dataframe API best practices

  • Stay in the Cloud: Perform data cleaning, transformation, and analysis

via BigFrames methods to leverage BigQuery's scale rather than downloading

data.

  • Prefer partial ordering mode: Enable partial ordering mode right after

importing BigFrames. This speeds up data processing significantly by relaxing

row-sequence constraints.

  import bigframes.pandas as bpd
  bpd.options.bigquery.ordering_mode = 'partial'
  • Use peek() for data preview: Use peek(n) to preview data instead of

head(n). peek(n) randomly samples n rows and is significantly faster.

head(n) returns rows in strict order and fails in partial ordering mode

unless the DataFrame has been explicitly sorted.

  • Avoid materializing data locally: Methods like to_pandas() download all

data to client memory, bypassing BigQuery’s distributed computation and

risking Out of Memory (OOM) errors. Do not materialize data locally unless:

  • The dataset is small enough to fit safely in memory.
  • An error message explicitly requires local materialization.
  • Prefer Dataframe API over SQL queries: Do not write raw SQL queries via

read_gbq() if a DataFrame/Series method achieves the same result, as it

breaks the Pandas abstraction and prevents lazy query execution.

  • Accessors over UDFs/Lambdas:
  • Use built-in accessors (e.g., df.col.str.*, df.col.dt.*) instead of

remote User Defined Functions (UDFs). UDFs require extra resources and

time to deploy.

  • Do not use lambdas with Series.map() or DataFrame.apply(). These

methods do not accept functions without udf or remote_function

decorators.

    # Avoid:
    df["upper"] = df["name"].map(lambda x: x.upper())

    # Prefer:
    df["upper"] = df["name"].str.upper()
  • Schema Verification: Do not assume the schema of intermediate outputs.

Proactively verify schemas using .dtypes and inspect sample records using

display() with .peek().

  • Visualization: Plot directly from the BigFrames DataFrame/Series when

possible. BigFrames is compatible with Matplotlib and Seaborn. If direct

plotting fails, use the .plot accessor. If the dataset is too large to plot,

aggregate or sample the data before calling

.to_pandas() to plot locally.

Machine Learning

  • Use bigframes.bigquery.ml package: Do not use Scikit-learn or other ML

libraries with BigQuery DataFrames. Standard Scikit-learn models require

bringing data into local client memory, whereas bigframes.bigquery.ml

delegates training directly to BigQuery's scalable ML engine. Import functions

from bigframes.bigquery.ml.

Reference Directory

  • Linear Regression: Train a linear

regression model to predict numerical values.

  • Logistic Regression: Train a logistic

regression model to predict boolean values.

BigFrames ML (Legacy)

The BigFrames ML package (bigframes.ml) is a legacy package that mimics the

scikit-learn API but is no longer recommended for new projects. Only use this

package if the user explicitly requests BigFrames ML.

  • Legacy Imports: When legacy BigFrames ML is requested, import tools and

classes from bigframes.ml instead of bigframes.bigquery.ml.

  • DataFrame Return on Prediction: Unlike Scikit-learn, BigFrames'

predict() method always returns a DataFrame containing both predictions

and features, rather than a single series of predictions.

  • No random_state: Do not pass a random_state argument when

instantiating BigFrames ML models, as this parameter is not supported in the

BigFrames ML package.

  • Automatic Scaling: Do not use OneHotEncoder or StandardScaler unless

explicitly requested, as scaling is handled automatically.

  • Hyperparameter Tuning: Write custom loops for hyperparameter tuning, as

BigFrames lacks GridSearchCV or RandomizedSearchCV.

  • ARIMA Plus (Forecasting):
  • Import from bigframes.ml.forecasting.
  • Sort data chronologically and split around a timepoint before training.
  • Ensure the prediction horizon is less than or equal to the training

horizon.

  • PCA: BigFrames' PCA class lacks a transform() method. Use predict()

instead.

  • Model Persistence: To persist a model, use model.to_gbq(). To load a

persisted model, use bpd.read_gbq_model().

Other skills for the same job

different authors, same section of the catalogue
D3 Viz
by chrisvoncsefalvay
×3

Creating interactive data visualisations using d3.js. This skill should be used when creating custom charts, graphs, network diagrams, geographic visualisations, or any complex SVG-based data visualisation that requires fine-grained control over visual elements, transitions, or interactions. Use this for bespoke visualisations beyond standard charting libraries, whether in React, Vue, Svelte, vanilla JavaScript, or any other environment.

20k tokens
Astropy
by christophacham
×3

Comprehensive Python library for astronomy and astrophysics. This skill should be used when working with astronomical data including celestial coordinates, physical units, FITS files, cosmological calculations, time systems, tables, world coordinate systems (WCS), and astronomical data analysis. Use when tasks involve coordinate transformations, unit conversions, FITS file manipulation, cosmological distance calculations, time scale conversions, or astronomical data processing.

16k tokens
Instrument Data To Allotrope
by anthropics
vendor ×2

Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV. Use this skill when scientists need to standardize instrument data for LIMS systems, data lakes, or downstream analysis. Supports auto-detection of instrument types. Outputs include full ASM JSON, flattened CSV for easy import, and exportable Python code for data engineers. Common triggers include converting instrument files, standardizing lab data, preparing data for upload to LIMS/ELN systems, or generating parser code for production pipelines.

33k tokens scripts
Qutip
by ComeOnOliver
×2

Quantum mechanics simulations and analysis using QuTiP (Quantum Toolbox in Python). Use when working with quantum systems including: (1) quantum states (kets, bras, density matrices), (2) quantum operators and gates, (3) time evolution and dynamics (Schrödinger, master equations, Monte Carlo), (4) open quantum systems with dissipation, (5) quantum measurements and entanglement, (6) visualization (Bloch sphere, Wigner functions), (7) steady states and correlation functions, or (8) advanced methods (Floquet theory, HEOM, stochastic solvers). Handles both closed and open quantum systems across various domains including quantum optics, quantum computing, and condensed matter physics.

27k tokens
Copilot Usage Metrics
by github
vendor ×1

Retrieve and display GitHub Copilot usage metrics for organizations and enterprises using the GitHub CLI and REST API.

1k tokens scripts
Mentoring Juniors
by github
vendor ×1

Socratic mentoring for junior developers and AI newcomers. Guides through questions, never answers. Triggers: "help me understand", "explain this code", "I''m stuck", "Im stuck", "I''m confused", "Im confused", "I don''t understand", "I dont understand", "can you teach me", "teach me", "mentor me", "guide me", "what does this error mean", "why doesn''t this work", "why does not this work", "I''m a beginner", "Im a beginner", "I''m learning", "Im learning", "I''m new to this", "Im new to this", "walk me through", "how does this work", "what''s wrong with my code", "what''s wrong", "can you break this down", "ELI5", "step by step", "where do I start", "what am I missing", "newbie here", "junior dev", "first time using", "how do I", "what is", "is this right", "not sure", "need help", "struggling", "show me", "help me debug", "best practice", "too complex", "overwhelmed", "lost", "debug this", "/socratic", "/hint", "/concept", "/pseudocode". Progressive clue systems, teaching techniques, and success metrics.

4k tokens
Astropy
by K-Dense-AI
×1

Core Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.

18k tokens
Polars
by K-Dense-AI
×1

High-performance DataFrame library for Python ETL, analytics, and pandas migration. Use for expression-based data manipulation with lazy query optimization, parallel execution, streaming out-of-core processing, Arrow interoperability, and optional GPU execution.

20k tokens

How to use it

Copy the folder

Take google/bigquery-bigframes from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.